Nicolas Lell

dblp:289/1307 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
8since 2021 · last 2025
0000-0002-6079-6480ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 3 first-author · 7 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 iN2V: Bringing Transductive Node Embeddings to Inductive Graphs
abstract
Shallow node embeddings like node2vec (N2V) can be used for nodes without features or to supplement existing features with structure-based information. Embedding methods like N2V are limited in their application on new nodes, which restricts them to the transductive setting where the entire graph, including the test nodes, is available during training. We propose inductive node2vec (iN2V), which combines a post-hoc procedure to compute embeddings for nodes unseen during training and modifications to the original N2V training procedure to prepare the embeddings for this post-hoc procedure. We conduct experiments on several benchmark datasets and demonstrate that iN2V is an effective approach to bringing transductive embeddings to an inductive setting. Using iN2V embeddings improves node classification by 1 point on average, with up to 6 points of improvement depending on the dataset and the number of unseen nodes. Our iN2V is a plug-in approach to create new or enrich existing embeddings. It can also be combined with other embedding methods, making it a versatile approach for inductive node representation learning. Code to reproduce the results is available at https://github.com/Foisunt/iN2V.
Nicolas Lell, Ansgar Scherp
ICML1
2024 HyperAggregation: Aggregating over Graph Edges with Hypernetworks
abstract
HyperAggregation is a hypernetwork-based aggregation function for Graph Neural Networks. It uses a hyper-network to dynamically generate weights in the size of the current neighborhood, which are then used to aggregate this neighborhood. This aggregation with the generated weights is done like an MLP-Mixer channel mixing over variable-sized vertex neighborhoods. We demonstrate HyperAggregation in two models, GraphHyperMixer is a model based on MLP-Mixer while GraphHyperConv is derived from a GCN but with a hypernetwork-based aggregation function. We perform experiments on diverse benchmark datasets for the vertex classification, graph classification, and graph regression tasks. The results show that HyperAggregation can be effectively used for homophilic and heterophilic datasets in both inductive and transductive settings. GraphHyperConv performs better than GraphHyperMixer and is especially strong in the transductive setting. On the heterophilic dataset Roman-Empire it reaches a new state of the art. On the graph-level tasks our models perform in line with similarly sized models. Ablation studies investigate the robustness against various hyperparameter choices. The implementation of HyperAggregation as well code to reproduce all experiments is available under https://github.com/Foisunt/HyperAggregation.
Nicolas Lell, Ansgar Scherp
IJCNN1
2024 Text Role Classification in Scientific Charts Using Multimodal Transformers
Hye Jin Kim, Nicolas Lell, Ansgar Scherp
NLDB (1)2
2023 Memorization of Named Entities in Fine-Tuned BERT Models
Andor Diera, Nicolas Lell, Aygul Garifullina, Ansgar Scherp
CD-MAKE2
2023 The Split Matters: Flat Minima Methods for Improving the Performance of GNNs
Nicolas Lell, Ansgar Scherp
CD-MAKE1
2023 Fine-Tuning Language Models for Scientific Writing Support
Justin Mücke, Daria Waldow, Luise Metzger, Philipp Schauz, Marcel Hoffmann 0002, Nicolas Lell, Ansgar Scherp
CD-MAKE6
2023 On the Rule-Based Extraction of Statistics Reported in Scientific Papers
Tobias Kalmbach, Marcel Hoffmann 0002, Nicolas Lell, Ansgar Scherp
NLDB3
2021 STEREO: A Pipeline for Extracting Experiment Statistics, Conditions, and Topics from Scientific Papers
abstract
A common writing style for statistical results are the recommendations of the American Psychology Association (APA). In practice, writing styles vary as reports are not 100% following APA-style or parameters are not reported despite being mandatory. In addition, the statistics are not reported in isolation but in context of experiment conditions investigated and the general experiment topic. We address these challenges by proposing a flexible pipeline STEREO based on wrapper induction and unsupervised aspect detection to extract experiment statistics, conditions, and topics. Thus, in contrast to existing rule-based tools like statcheck with a pre-defined set of rules, we learn rules via induction. It required only 0.25% of the CORD-19 corpus (about 500 documents) to learn statistics extraction rules that cover 95% of the sentences in CORD-19. The statistic extraction has 100% precision on APA-conform statistics, which is identical with statcheck. In addition, STEREO can extract non-APA writing styles with precision, which statcheck does not support. Extracting non-APA conform statistics is important as they make more than 99% of all 113k extracted statistics. We could extract in 46% the correct conditions from APA-conform reports (30% for non-APA). The best model for topic extraction achieves a precision of 75% on statistics reported in APA style (73% for non-APA conform).
Steffen Epp, Marcel Hoffmann 0002, Nicolas Lell, Michael Mohr, Ansgar Scherp
iiWAS3