EDBT 2026 Demo / reviewers in the wild / expert
Dexiong Chen
dblp:240/6347
· DBLP profile ↗
21ranked-venue papers
9as first author
16since 2021 · last 2025
0009-0009-7075-7483ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 7 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Learning Laplacian Positional Encodings for Heterophilous GraphsabstractIn this work, we theoretically demonstrate that current graph positional encodings (PEs) are not beneficial and could potentially hurt performance in tasks involving heterophilous graphs, where nodes that are close tend to have different labels. This limitation is critical as many real-world networks exhibit heterophily, and even highly homophilous graphs can contain local regions of strong heterophily. To address this limitation, we propose Learnable Laplacian Positional Encodings (LLPE), a new PE that leverages the full spectrum of the graph Laplacian, enabling them to capture graph structure on both homophilous and heterophilous graphs. Theoretically, we prove LLPE’s ability to approximate a general class of graph distances and demonstrate its generalization properties. Empirically, our evaluation on 12 benchmarks demonstrates that LLPE improves accuracy across a variety of GNNs, including graph transformers, by up to 35% and 14% on synthetic and real-world graphs, respectively. Going forward, our work represents a significant step towards developing PEs that effectively capture complex structures in heterophilous graphs. Michael Ito, Jiong Zhu, Dexiong Chen, Danai Koutra, Jenna Wiens |
AISTATS | 3 |
| 2025 | Learning Long Range Dependencies on Graphs via Random WalksabstractMessage-passing graph neural networks (GNNs) excel at capturing local relationships but struggle with long-range dependencies in graphs. In contrast, graph transformers (GTs) enable global information exchange but often oversimplify the graph structure by representing graphs as sets of fixed-length vectors. This work introduces a novel architecture that overcomes the shortcomings of both approaches by combining the long-range information of random walks with local message passing. By treating random walks as sequences, our architecture leverages recent advances in sequence models to effectively capture long-range dependencies within these walks. Based on this concept, we propose a framework that offers (1) more expressive graph representations through random walk sequences, (2) the ability to utilize any sequence model for capturing long-range dependencies, and (3) the flexibility by integrating various GNN and GT architectures. Our experimental evaluations demonstrate that our approach achieves competitive performance on 19 graph and node benchmark datasets, notably outperforming existing methods by up to 13\% on the PascalVoc-SP and COCO-SP datasets.
Code: https://github.com/BorgwardtLab/NeuralWalker Dexiong Chen, Till Hendrik Schulz, Karsten M. Borgwardt |
ICLR | 1 |
| 2025 | Flatten Graphs as Sequences: Transformers are Scalable Graph GeneratorsabstractWe introduce AutoGraph, a scalable autoregressive model for attributed graph generation using decoder-only transformers. By flattening graphs into random sequences of tokens through a reversible process, AutoGraph enables modeling graphs as sequences without relying on additional node features that are expensive to compute, in contrast to diffusion-based approaches. This results in sampling complexity and sequence lengths that scale optimally linearly with the number of edges, making it scalable and efficient for large, sparse graphs. A key success factor of AutoGraph is that its sequence prefixes represent induced subgraphs, creating a direct link to sub-sentences in language modeling. Empirically, AutoGraph achieves state-of-the-art performance on synthetic and molecular benchmarks, with up to 100x faster generation and 3x faster training than leading diffusion models. It also supports substructure-conditioned generation without fine-tuning and shows promising transferability, bridging language modeling and graph generation to lay the groundwork for graph foundation models. Our code is available at https://github.com/BorgwardtLab/AutoGraph. Dexiong Chen, Markus Krimmel, Karsten M. Borgwardt |
NeurIPS | 1 |
| 2025 | Detecting Antimicrobial Resistance Through MALDI-TOF Mass Spectrometry with Statistical Guarantees Using Conformal Prediction
Nina Corvelo Benz, Lucas Miranda 0002, Dexiong Chen, Janko Sattler, Karsten M. Borgwardt |
RECOMB | 3 |
| 2025 | Endowing protein language models with structural knowledgeabstractMOTIVATION: Protein language models (PLMs) have transformed protein research by learning rich representations from sequence data alone, yet they largely ignore the wealth of structural information now available through advances in structure prediction. Current methods that incorporate structural data often require substantial computational resources and complex architectures, limiting their practical adoption. We present a novel joint sequence and structure embedding method that achieves computational and parameter efficiency while maintaining high performance. Our approach introduces a lightweight integration framework that combines pretrained sequence transformers' self-attention with specialized structural adapters, enabling seamless incorporation of structural knowledge into existing PLMs through these enhanced self-attention mechanisms. RESULTS: The method demonstrates remarkable efficiency, requiring only modest pretraining on 542K protein structures, three orders of magnitude less than the data used to train PLMs, using standard masked language modeling objectives. Despite this lightweight approach, our joint embeddings consistently outperform sequence-only models like ESM-2 while achieving comparable results to more complex structure-based methods that use significantly more parameters and computational resources. This work establishes a new paradigm for protein representation learning that balances performance with practical constraints. By providing computationally efficient joint sequence-structure embeddings, we offer the scientific community an accessible tool that captures both sequential and structural protein information without the computational overhead typically associated with structure-aware models. AVAILABILITY AND IMPLEMENTATION: code and links to checkpoints are available at https://github.com/BorgwardtLab/PST. Philip Hartout, Dexiong Chen, Paolo Pellizzoni, Carlos G. Oliver, Karsten M. Borgwardt |
Bioinform. | 2 |
| 2025 | HTR-VT: Handwritten text recognition with vision transformer
Yuting Li 0001, Dexiong Chen, Tinglong Tang, Xi Shen 0001 |
Pattern Recognit. | 2 |
| 2025 | AdvMixUp: Adversarial MixUp Regularization for Deep LearningabstractDeep neural networks (DNNs) have shown significant progress in many application fields. However, overfitting remains a significant challenge in their development. While existing data-augmentation techniques such as MixUp have been successful in preventing overfitting, they often fail to generate hard mixed samples near the decision boundary, impeding model optimization. In this article, we present adversarial MixUp (AdvMixUp), a novel sample-dependent method for regularizing DNNs. AdvMixUp addresses this issue by incorporating adversarial training (AT) to create sample-dependent and feature-level interpolation masks, generating more challenging mixed samples. These virtual samples enable DNNs to learn more robust features, ultimately reducing overfitting. Empirical evaluations on CIFAR-10, CIFAR-100, Tiny-ImageNet, and ImageNet demonstrate that AdvMixUp outperforms existing MixUp variants. Jun Fu 0001, Xianrui Ji, Dexiong Chen, Guosheng Hu, Shuang Li 0008, Xiating Feng |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | SURE: SUrvey REcipes for Building Reliable and Robust Deep NetworksabstractIn this paper, we revisit techniques for uncertainty estimation within deep neural networks and consolidate a suite of techniques to enhance their reliability. Our investigation reveals that an integrated application of diverse techniques-spanning model regularization, classifier and optimization-substantially improves the accuracy of uncertainty predictions in image classification tasks. The synergistic effect of these techniques culminates in our novel SURE approach. We rigorously evaluate SURE against the benchmark of failure prediction, a critical testbed for uncertainty estimation efficacy. Our results showcase a consistently better performance than models that individually deploy each technique, across various datasets and model architectures. When applied to real-world challenges, such as data corruption, label noise, and long-tailed class distribution, SURE exhibits remarkable robustness, delivering results that are superior or on par with current state-of-the-art specialized methods. Particularly on Animal-10N and Food-101N for learning with noisy labels, SURE achieves state-of-the-art performance without any task-specific adjustments. This work not only sets a new benchmark for robust uncertainty estimation but also paves the way for its application in diverse, real-world scenarios where reliability is paramount. Our code is available at https://yutingli0606.github.io/SURE/. Yuting Li 0001, Yingyi Chen, Xuanlong Yu, Dexiong Chen, Xi Shen 0001 |
CVPR | 4 |
| 2024 | On the Expressivity and Sample Complexity of Node-Individualized Graph Neural NetworksabstractGraph neural networks (GNNs) employing message passing for graph classification are inherently limited by the expressive power of the Weisfeiler-Leman (WL) test for graph isomorphism. Node individualization schemes, which assign unique identifiers to nodes (e.g., by adding random noise to features), are a common approach for achieving universal expressiveness. However, the ability of GNNs endowed with individualization schemes to generalize beyond the training data is still an open question. To address this question, this paper presents a theoretical analysis of the sample complexity of such GNNs from a statistical learning perspective, employing Vapnik–Chervonenkis (VC) dimension and covering number bounds. We demonstrate that node individualization schemes that are permutation-equivariant result in lower sample complexity, and design novel individualization schemes that exploit these results. As an application of this analysis, we also develop a novel architecture that can perform substructure identification (i.e., subgraph isomorphism) while having a lower VC dimension compared to competing methods. Finally, our theoretical findings are validated experimentally on both synthetic and real-world datasets. Paolo Pellizzoni, Till Hendrik Schulz, Dexiong Chen, Karsten M. Borgwardt |
NeurIPS | 3 |
| 2024 | Biomarker identification by interpretable maximum mean discrepancyabstractMOTIVATION: In many biomedical applications, we are confronted with paired groups of samples, such as treated versus control. The aim is to detect discriminating features, i.e. biomarkers, based on high-dimensional (omics-) data. This problem can be phrased more generally as a two-sample problem requiring statistical significance testing to establish differences, and interpretations to identify distinguishing features. The multivariate maximum mean discrepancy (MMD) test quantifies group-level differences, whereas statistically significantly associated features are usually found by univariate feature selection. Currently, few general-purpose methods simultaneously perform multivariate feature selection and two-sample testing. RESULTS: We introduce a sparse, interpretable, and optimized MMD test (SpInOpt-MMD) that enables two-sample testing and feature selection in the same experiment. SpInOpt-MMD is a versatile method and we demonstrate its application to a variety of synthetic and real-world data types including images, gene expression measurements, and text data. SpInOpt-MMD is effective in identifying relevant features in small sample sizes and outperforms other feature selection methods such as SHapley Additive exPlanations and univariate association analysis in several experiments. AVAILABILITY AND IMPLEMENTATION: The code and links to our public data are available at https://github.com/BorgwardtLab/spinoptmmd. Michael F. Adamer, Sarah C. Brüningk, Dexiong Chen, Karsten M. Borgwardt |
Bioinform. | 3 |
| 2023 | Unsupervised Manifold Alignment with Joint Multidimensional Scaling
Dexiong Chen, Bowen Fan, Carlos G. Oliver, Karsten M. Borgwardt |
ICLR | 1 |
| 2023 | Fisher Information Embedding for Node and Graph LearningabstractAttention-based graph neural networks (GNNs), such as graph attention networks (GATs), have become popular neural architectures for processing graph-structured data and learning node embeddings. Despite their empirical success, these models rely on labeled data and the theoretical properties of these models have yet to be fully understood. In this work, we propose a novel attention-based node embedding framework for graphs. Our framework builds upon a hierarchical kernel for multisets of subgraphs around nodes (e.g. neighborhoods) and each kernel leverages the geometry of a smooth statistical manifold to compare pairs of multisets, by ``projecting'' the multisets onto the manifold. By explicitly computing node embeddings with a manifold of Gaussian mixtures, our method leads to a new attention mechanism for neighborhood aggregation. We provide theoretical insights into generalizability and expressivity of our embeddings, contributing to a deeper understanding of attention-based GNNs. We propose both efficient unsupervised and supervised methods for learning the embeddings. Through experiments on several node classification benchmarks, we demonstrate that our proposed method outperforms existing attention-based graph models like GATs. Our code is available at https://github.com/BorgwardtLab/fisher_information_embedding. Dexiong Chen, Paolo Pellizzoni, Karsten M. Borgwardt |
ICML | 1 |
| 2023 | ProteinShake: Building datasets and benchmarks for deep learning on protein structuresabstractWe present ProteinShake, a Python software package that simplifies datasetcreation and model evaluation for deep learning on protein structures. Users cancreate custom datasets or load an extensive set of pre-processed datasets fromthe Protein Data Bank (PDB) and AlphaFoldDB. Each dataset is associated withprediction tasks and evaluation functions covering a broad array of biologicalchallenges. A benchmark on these tasks shows that pre-training almost alwaysimproves performance, the optimal data modality (graphs, voxel grids, or pointclouds) is task-dependent, and models struggle to generalize to new structures.ProteinShake makes protein structure data easily accessible and comparisonamong models straightforward, providing challenging benchmark settings withreal-world implications.ProteinShake is available at: https://proteinshake.ai Tim Kucera, Carlos G. Oliver, Dexiong Chen, Karsten M. Borgwardt |
NeurIPS | 3 |
| 2022 | Structure-Aware Transformer for Graph Representation LearningabstractThe Transformer architecture has gained growing attention in graph representation learning recently, as it naturally overcomes several limitations of graph neural networks (GNNs) by avoiding their strict structural inductive biases and instead only encoding the graph structure via positional encoding. Here, we show that the node representations generated by the Transformer with positional encoding do not necessarily capture structural similarity between them. To address this issue, we propose the Structure-Aware Transformer, a class of simple and flexible graph Transformers built upon a new self-attention mechanism. This new self-attention incorporates structural information into the original self-attention by extracting a subgraph representation rooted at each node before computing the attention. We propose several methods for automatically generating the subgraph representation and show theoretically that the resulting representations are at least as expressive as the subgraph representations. Empirically, our method achieves state-of-the-art performance on five graph prediction benchmarks. Our structure-aware framework can leverage any existing GNN to extract the subgraph representation, and we show that it systematically improves performance relative to the base GNN model, successfully combining the advantages of GNNs and Transformers. Our code is available at https://github.com/BorgwardtLab/SAT. Dexiong Chen, Leslie O'Bray, Karsten M. Borgwardt |
ICML | 1 |
| 2022 | MetaMixUp: Learning Adaptive Interpolation Policy of MixUp With MetalearningabstractMixUp is an effective data augmentation method to regularize deep neural networks via random linear interpolations between pairs of samples and their labels. It plays an important role in model regularization, semisupervised learning (SSL), and domain adaption. However, despite its empirical success, its deficiency of randomly mixing samples has poorly been studied. Since deep networks are capable of memorizing the entire data set, the corrupted samples generated by vanilla MixUp with a badly chosen interpolation policy will degrade the performance of networks. To overcome overfitting to corrupted samples, inspired by metalearning (learning to learn), we propose a novel technique of learning to a mixup in this work, namely, MetaMixUp. Unlike the vanilla MixUp that samples interpolation policy from a predefined distribution, this article introduces a metalearning-based online optimization approach to dynamically learn the interpolation policy in a data-adaptive way (learning to learn better). The validation set performance via metalearning captures the noisy degree, which provides optimal directions for interpolation policy learning. Furthermore, we adapt our method for pseudolabel-based SSL along with a refined pseudolabeling strategy. In our experiments, our method achieves better performance than vanilla MixUp and its variants under SL configuration. In particular, extensive experiments show that our MetaMixUp adapted SSL greatly outperforms MixUp and many state-of-the-art methods on CIFAR-10 and SVHN benchmarks under the SSL configuration. Zhijun Mai, Guosheng Hu, Dexiong Chen, Fumin Shen, Heng Tao Shen |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2021 | A Trainable Optimal Transport Embedding for Feature Aggregation and its Relationship to Attention
Grégoire Mialon, Dexiong Chen, Alexandre d'Aspremont, Julien Mairal |
ICLR | 2 |
| 2020 | Convolutional Kernel Networks for Graph-Structured DataabstractWe introduce a family of multilayer graph kernels and establish new links between graph convolutional neural networks and kernel methods. Our approach generalizes convolutional kernel networks to graph-structured data, by representing graphs as a sequence of kernel feature maps, where each node carries information about local graph substructures. On the one hand, the kernel point of view offers an unsupervised, expressive, and easy-to-regularize data representation, which is useful when limited samples are available. On the other hand, our model can also be trained end-to-end on large-scale data, leading to new types of graph convolutional neural networks. We show that our method achieves competitive performance on several graph classification benchmarks, while offering simple model interpretation. Our code is freely available at https://github.com/claying/GCKN. Dexiong Chen, Laurent Jacob, Julien Mairal |
ICML | 1 |
| 2019 | A Kernel Perspective for Regularizing Deep Neural NetworksabstractWe propose a new point of view for regularizing deep neural networks by using the norm of a reproducing kernel Hilbert space (RKHS). Even though this norm cannot be computed, it admits upper and lower approximations leading to various practical strategies. Specifically, this perspective (i) provides a common umbrella for many existing regularization principles, including spectral norm and gradient penalties, or adversarial training, (ii) leads to new effective regularization penalties, and (iii) suggests hybrid strategies combining lower and upper bounds to get better approximations of the RKHS norm. We experimentally show this approach to be effective when learning on small datasets, or to obtain adversarially robust models. Alberto Bietti, Grégoire Mialon, Dexiong Chen, Julien Mairal |
ICML | 3 |
| 2019 | Recurrent Kernel NetworksabstractSubstring kernels are classical tools for representing biological sequences or text. However, when large amounts of annotated data is available, models that allow end-to-end training such as neural networks are often prefered. Links between recurrent neural networks (RNNs) and substring kernels have recently been drawn, by formally showing that RNNs with specific activation functions were points in a reproducing kernel Hilbert space (RKHS). In this paper, we revisit this link by generalizing convolutional kernel networks---originally related to a relaxation of the mismatch kernel---to model gaps in sequences. It results in a new type of recurrent neural network which can be trained end-to-end with backpropagation, or without supervision by using kernel approximation techniques. We experimentally show that our approach is well suited to biological sequences, where it outperforms existing methods for protein classification tasks. Dexiong Chen, Laurent Jacob, Julien Mairal |
NeurIPS | 1 |
| 2019 | Biological Sequence Modeling with Convolutional Kernel Networks
Dexiong Chen, Laurent Jacob, Julien Mairal |
RECOMB | 1 |
| 2019 | Biological sequence modeling with convolutional kernel networksabstractMOTIVATION: The growing number of annotated biological sequences available makes it possible to learn genotype-phenotype relationships from data with increasingly high accuracy. When large quantities of labeled samples are available for training a model, convolutional neural networks can be used to predict the phenotype of unannotated sequences with good accuracy. Unfortunately, their performance with medium- or small-scale datasets is mitigated, which requires inventing new data-efficient approaches. RESULTS: We introduce a hybrid approach between convolutional neural networks and kernel methods to model biological sequences. Our method enjoys the ability of convolutional neural networks to learn data representations that are adapted to a specific task, while the kernel point of view yields algorithms that perform significantly better when the amount of training data is small. We illustrate these advantages for transcription factor binding prediction and protein homology detection, and we demonstrate that our model is also simple to interpret, which is crucial for discovering predictive motifs in sequences. AVAILABILITY AND IMPLEMENTATION: Source code is freely available at https://gitlab.inria.fr/dchen/CKN-seq. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Dexiong Chen, Laurent Jacob, Julien Mairal |
Bioinform. | 1 |