Sune Lehmann

dblp:33/7479 · DBLP profile ↗
← Back
7ranked-venue papers
0as first author
2since 2021 · last 2024
0000-0001-6099-2345ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 4 · 1 since 2021Artificial intelligence and machine learning · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 2Applied, interdisciplinary, general and emerging computing · 2Computer networks · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Graph learning · 67% Representation and self-supervised learning · 22% Information extraction and text analysis · 11%
Computer networks
1 paper
Wireless sensing and localization · 50% Network measurement and analytics · 50%

Topics — the 8 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Graph learning
graph representation learning
0.812024
A Hierarchical Block Distance Model for Ultra Low-Dimensional Graph Representations · IEEE Trans. Knowl. Data Eng. 2024
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction
low-dimensional embedding
0.812024
A Hierarchical Block Distance Model for Ultra Low-Dimensional Graph Representations · IEEE Trans. Knowl. Data Eng. 2024
Machine learning › Graph learning › network embedding
scalable embedding
0.812024
A Hierarchical Block Distance Model for Ultra Low-Dimensional Graph Representations · IEEE Trans. Knowl. Data Eng. 2024
Machine learning › Graph learning
stochastic block model
0.812024
A Hierarchical Block Distance Model for Ultra Low-Dimensional Graph Representations · IEEE Trans. Knowl. Data Eng. 2024
Natural language and speech › Information extraction and text analysis
sentiment analysis
0.312017
Using millions of emoji occurrences to learn any-domain representations for detecting sentiment, emotion and sarcasm · EMNLP 2017
Wireless sensing and localization › indoor localization
wifi localization
0.212015
Opportunities and Challenges in Crowdsourced Wardriving · Internet Measurement Conference 2015
Natural language and speech › Information extraction and text analysis › relation extraction
distant supervision
0.112017
Using millions of emoji occurrences to learn any-domain representations for detecting sentiment, emotion and sarcasm · EMNLP 2017
Ubiquitous computing and smart environments
location-based services
0.112015
Opportunities and Challenges in Crowdsourced Wardriving · Internet Measurement Conference 2015

Methods — techniques the papers use, named apart from their topics

pre-training · 0.3emoji prediction · 0.3
YearPublicationVenuePosition
2024 Time to Cite: Modeling Citation Networks using the Dynamic Impact Single-Event Embedding Model
abstract
Understanding the structure and dynamics of scientific research, i.e., the science of science (SciSci), has become an important area of research in order to address imminent questions including how scholars interact to advance science, how disciplines are related and evolve, and how research impact can be quantified and predicted. Central to the study of SciSci has been the analysis of citation networks. Here, two prominent modeling methodologies have been employed: one is to assess the citation impact dynamics of papers using parametric distributions, and the other is to embed the citation networks in a latent space optimal for characterizing the static relations between papers in terms of their citations. Interestingly, citation networks are a prominent example of single-event dynamic networks, i.e., networks for which each dyad only has a single event (i.e., the point in time of citation). We presently propose a novel likelihood function for the characterization of such single-event networks. Using this likelihood, we propose the Dynamic Impact Single-Event Embedding model (DISEE). The DISEE model characterizes the scientific interactions in terms of a latent distance model in which random effects account for citation heterogeneity while the time-varying impact is characterized using existing parametric representations for assessment of dynamic impact. We highlight the proposed approach on several real citation networks finding that DISEE well reconciles static latent distance network embedding approaches with classical dynamic impact assessments.
Nikolaos Nakis, Abdulkadir Çelikkanat, Louis Boucherie, Sune Lehmann, Morten Mørup
AISTATS4
2024 A Hierarchical Block Distance Model for Ultra Low-Dimensional Graph Representations
abstract
Graph Representation Learning (GRL) has become central for characterizing structures of complex networks and performing tasks such as link prediction, node classification, network reconstruction, and community detection. Whereas numerous generative GRL models have been proposed, many approaches have prohibitive computational requirements hampering large-scale network analysis, fewer are able to explicitly account for structure emerging at multiple scales, and only a few explicitly respect important network properties such as homophily and transitivity. This paper proposes a novel scalable graph representation learning method named the Hierarchical Block Distance Model (HBDM). The HBDM imposes a multiscale block structure akin to stochastic block modeling (SBM) and accounts for homophily and transitivity by accurately approximating the latent distance model (LDM) throughout the inferred hierarchy. The HBDM naturally accommodates unipartite, directed, and bipartite networks whereas the hierarchy is designed to ensure linearithmic time and space complexity enabling the analysis of very large-scale networks. We evaluate the performance of the HBDM on massive networks consisting of millions of nodes. Importantly, we find that the proposed HBDM framework significantly outperforms recent scalable approaches in all considered downstream tasks. Surprisingly, we observe superior performance even imposing ultra-low two-dimensional embeddings facilitating accurate direct and hierarchical-aware network visualization and interpretation.
Nikolaos Nakis, Abdulkadir Çelikkanat, Sune Lehmann, Morten Mørup
IEEE Trans. Knowl. Data Eng.3
2017 Using millions of emoji occurrences to learn any-domain representations for detecting sentiment, emotion and sarcasm
abstract
NLP tasks are often limited by scarcity of manually annotated data. In social media sentiment analysis and related tasks, researchers have therefore used binarized emoticons and specific hashtags as forms of distant supervision. Our paper shows that by extending the distant supervision to a more diverse set of noisy labels, the models can learn richer representations. Through emoji prediction on a dataset of 1246 million tweets containing one of 64 common emojis we obtain state-of-the-art performance on 8 benchmark datasets within sentiment, emotion and sarcasm detection using a single pretrained model. Our analyses confirm that the diversity of our emotional labels yield a performance improvement over previous distant supervision approaches.
Bjarke Felbo, Alan Mislove, Anders Søgaard, Iyad Rahwan, Sune Lehmann
EMNLP5
2017 Modeling the Temporal Nature of Human Behavior for Demographics Prediction
Bjarke Felbo, Pål Roe Sundsøy, Alex Pentland, Sune Lehmann, Yves-Alexandre de Montjoye
ECML/PKDD (3)4
2015 Opportunities and Challenges in Crowdsourced Wardriving
abstract
Knowing the physical location of a mobile device is crucial for a number of context-aware applications. This information is usually obtained using the Global Positioning System (GPS), or by calculating the position based on proximity of WiFi access points with known location (where the position of the access points is stored in a database at a central server). To date, most of the research regarding the creation of such a database has investigated datasets collected both artificially and over short periods of time (e.g., during a one-day drive around a city). In contrast, most in-use databases are collected by mobile devices automatically, and are maintained by large mobile OS providers.
Piotr Sapiezynski, Radu Gatej, Alan Mislove, Sune Lehmann
Internet Measurement Conference4
2012 Tweetin' in the Rain: Exploring Societal-Scale Effects of Weather on Mood
Aniko Hannak, Eric Anderson 0001, Lisa Feldman Barrett, Sune Lehmann, Alan Mislove, Mirek Riedewald
ICWSM4
2011 Understanding the Demographics of Twitter Users
Alan Mislove, Sune Lehmann, Yong-Yeol Ahn, Jukka-Pekka Onnela, J. Niels Rosenquist
ICWSM2