EDBT 2026 Demo / reviewers in the wild / expert
Samet Demir
dblp:254/1845
· DBLP profile ↗
3ranked-venue papers
3as first author
3since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 3 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Learning theory · 48% Deep learning architectures and training · 16% Optimization for machine learning · 16% |
Topics — the 7 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Learning theory › statistical learning theory
asymptotic analysis |
0.9 | 1 | 2025 | How Data Mixing Shapes In-Context Learning: Asymptotic Equivalence for Transformers with MLPs · NeurIPS 2025 |
Machine learning › Learning theory › generalization
generalization theory |
0.9 | 1 | 2025 | Asymptotic Analysis of Two-Layer Neural Networks after One Gradient Step under Gaussian Mixtures Data with Structure · ICLR 2025 |
Machine learning › Optimization for machine learning › convergence guarantees
gradient descent convergence |
0.9 | 1 | 2025 | Asymptotic Analysis of Two-Layer Neural Networks after One Gradient Step under Gaussian Mixtures Data with Structure · ICLR 2025 |
Natural language and speech › Language models and text generation
in-context learning |
0.9 | 1 | 2025 | How Data Mixing Shapes In-Context Learning: Asymptotic Equivalence for Transformers with MLPs · NeurIPS 2025 |
Machine learning › Learning theory › neural network theory
neural network asymptotics |
0.9 | 1 | 2025 | Asymptotic Analysis of Two-Layer Neural Networks after One Gradient Step under Gaussian Mixtures Data with Structure · ICLR 2025 |
Machine learning › Deep learning architectures and training
training dynamics |
0.9 | 1 | 2025 | Asymptotic Analysis of Two-Layer Neural Networks after One Gradient Step under Gaussian Mixtures Data with Structure · ICLR 2025 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model › mixture model
gaussian mixture model |
0.3 | 1 | 2025 | Asymptotic Analysis of Two-Layer Neural Networks after One Gradient Step under Gaussian Mixtures Data with Structure · ICLR 2025 |
Methods — techniques the papers use, named apart from their topics
gaussian universality · 1.7polynomial approximation · 0.9orthogonal polynomials · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Asymptotic Analysis of Two-Layer Neural Networks after One Gradient Step under Gaussian Mixtures Data with StructureabstractIn this work, we study the training and generalization performance of two-layer neural networks (NNs) after one gradient descent step under structured data modeled by Gaussian mixtures. While previous research has extensively analyzed this model under isotropic data assumption, such simplifications overlook the complexities inherent in real-world datasets. Our work addresses this limitation by analyzing two-layer NNs under Gaussian mixture data assumption in the asymptotically proportional limit, where the input dimension, number of hidden neurons, and sample size grow with finite ratios. We characterize the training and generalization errors by leveraging recent advancements in Gaussian universality. Specifically, we prove that a high-order polynomial model performs equivalent to the non-linear neural networks under certain conditions. The degree of the equivalent model is intricately linked to both the "data spread" and the learning rate employed during one gradient step. Through extensive simulations, we demonstrate the equivalence between the original model and its polynomial counterpart across various regression and classification tasks. Additionally, we explore how different properties of Gaussian mixtures affect learning outcomes. Finally, we illustrate experimental results on Fashion-MNIST classification, indicating that our findings can translate to realistic data. Samet Demir, Zafer Dogan |
ICLR | 1 |
| 2025 | How Data Mixing Shapes In-Context Learning: Asymptotic Equivalence for Transformers with MLPsabstractPretrained Transformers demonstrate remarkable in-context learning (ICL) capabilities, enabling them to adapt to new tasks from demonstrations without parameter updates. However, theoretical studies often rely on simplified architectures (e.g., omitting MLPs), plain data models (e.g., linear regression with isotropic inputs), and single-source training—limiting their relevance to realistic settings. In this work, we study ICL in pretrained Transformers with nonlinear MLP heads on nonlinear tasks drawn from multiple data sources with heterogeneous input, task, and noise distributions. We analyze a model where the MLP comprises two layers, with the first layer trained via a single gradient step and the second layer fully optimized. Under high-dimensional asymptotics, we prove that such models are equivalent in ICL error to structured polynomial predictors, leveraging results from the theory of Gaussian universality and orthogonal polynomials. This equivalence reveals that nonlinear MLPs meaningfully enhance ICL performance—particularly on nonlinear tasks—compared to linear baselines. It also enables a precise analysis of data mixing effects: we identify key properties of high-quality data sources (low noise, structured covariances) and show that feature learning emerges only when the task covariance exhibits sufficient structure. These results are validated empirically across various activation functions, model sizes, and data distributions. Finally, we experiment with a real-world scenario involving multilingual sentiment analysis where each language is treated as a different source. Our experimental results for this case exemplify how our findings extend to real-world cases. Overall, our work advances the theoretical foundations of ICL in Transformers and provides actionable insight into the role of architecture and data in ICL. Samet Demir, Zafer Dogan |
NeurIPS | 1 |
| 2023 | Optimal Nonlinearities Improve Generalization Performance of Random Features
Samet Demir, Zafer Dogan |
ACML | 1 |