Oguz Kaan Yüksel

dblp:283/8205 · DBLP profile ↗
← Back
5ranked-venue papers
5as first author
5since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 5 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Transfer learning and domain adaptation · 25% Generative modeling · 19% Learning theory · 18%
Theoretical computer science
1 paper
Mathematical optimization · 100%

Topics — the 12 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Learning theory
sample complexity
0.912025
Long-Context Linear System Identification · ICLR 2025
Robotics › Motion planning and robot control
system identification
0.912025
Long-Context Linear System Identification · ICLR 2025
Machine learning › Transfer learning and domain adaptation
meta-learning
0.812024
First-order ANIL provably learns representations despite overparametrisation · ICLR 2024
Machine learning › Transfer learning and domain adaptation › meta-learning › gradient-based meta-learning
model-agnostic meta-learning
0.812024
First-order ANIL provably learns representations despite overparametrisation · ICLR 2024
Machine learning › Representation and self-supervised learning › representation learning › joint representation learning › multi-task representation learning
shared representation learning
0.812024
First-order ANIL provably learns representations despite overparametrisation · ICLR 2024
Machine learning › Deep learning architectures and training
data augmentation
0.512021
Semantic Perturbations with Normalizing Flows for Improved Generalization · ICCV 2021
Machine learning › Generative modeling
generative adversarial network
0.512021
LatentCLR: A Contrastive Learning Approach for Unsupervised Discovery of Interpretable Directions · ICCV 2021
Machine learning › Generative modeling › generative adversarial network
latent direction discovery
0.512021
LatentCLR: A Contrastive Learning Approach for Unsupervised Discovery of Interpretable Directions · ICCV 2021
Mathematical optimization › continuous optimization › matrix optimization
matrix recovery
0.312025
Long-Context Linear System Identification · ICLR 2025
Machine learning › Learning theory › generalization bounds
meta-learning for domain generalization
0.212024
First-order ANIL provably learns representations despite overparametrisation · ICLR 2024
Machine learning › Representation and self-supervised learning
contrastive learning
0.112021
LatentCLR: A Contrastive Learning Approach for Unsupervised Discovery of Interpretable Directions · ICCV 2021
Machine learning › Generative modeling
normalizing flow
0.112021
Semantic Perturbations with Normalizing Flows for Improved Generalization · ICCV 2021

Methods — techniques the papers use, named apart from their topics

sample complexity analysis · 1.7rank-regularized estimation · 1.7overparametrization analysis · 0.8gradient descent analysis · 0.8self-supervised learning · 0.5normalizing flow · 0.5contrastive learning · 0.5adversarial perturbation · 0.5
YearPublicationVenuePosition
2025 On the Sample Complexity of Next-Token Prediction
abstract
Next-token prediction with cross-entropy loss is the objective of choice in sequence and language modeling. Despite its widespread use, there is a lack of theoretical analysis regarding the generalization of models trained using this objective. In this work, we provide an analysis of empirical risk minimization for sequential inputs generated by order-$k$ Markov chains. Assuming bounded and Lipschitz logit functions, our results show that in-sample prediction error decays optimally with the number of tokens, whereas out-of-sample error incurs an additional term related to the mixing properties of the Markov chain. These rates depend on the statistical complexity of the hypothesis class and can lead to generalization errors that do not scale exponentially with the order of the Markov chain—unlike classical $k$-gram estimators. Finally, we discuss the possibility of achieving generalization rates independent of mixing.
Oguz Kaan Yüksel, Nicolas Flammarion
AISTATS1
2025 Long-Context Linear System Identification
abstract
This paper addresses the problem of long-context linear system identification, where the state $x_t$ of the system at time $t$ depends linearly on previous states $x_s$ over a fixed context window of length $p$. We establish a sample complexity bound that matches the _i.i.d._ parametric rate, up to logarithmic factors for a broad class of systems, extending previous work that considered only first-order dependencies. Our findings reveal a ``learning-without-mixing'' phenomenon, indicating that learning long-context linear autoregressive models is not hindered by slow mixing properties potentially associated with extended context windows. Additionally, we extend these results to _(i)_ shared low-rank feature representations, where rank-regularized estimators improve rates with respect to dimensionality, and _(ii)_ misspecified context lengths in strictly stable systems, where shorter contexts offer statistical advantages.
Oguz Kaan Yüksel, Mathieu Even, Nicolas Flammarion
ICLR1
2024 First-order ANIL provably learns representations despite overparametrisation
abstract
Due to its empirical success in few-shot classification and reinforcement learning, meta-learning has recently received significant interest. Meta-learning methods leverage data from previous tasks to learn a new task in a sample-efficient manner. In particular, model-agnostic methods look for initialization points from which gradient descent quickly adapts to any new task. Although it has been empirically suggested that such methods perform well by learning shared representations during pretraining, there is limited theoretical evidence of such behavior. More importantly, it has not been shown that these methods still learn a shared structure, despite architectural misspecifications. In this direction, this work shows, in the limit of an infinite number of tasks, that first-order ANIL with a linear two-layer network architecture successfully learns linear shared representations. This result even holds with _overparametrization_; having a width larger than the dimension of the shared representations results in an asymptotically low-rank solution. The learned solution then yields a good adaptation performance on any new task after a single gradient step. Overall, this illustrates how well model-agnostic methods such as first-order ANIL can learn shared representations.
Oguz Kaan Yüksel, Etienne Boursier, Nicolas Flammarion
ICLR1
2021 LatentCLR: A Contrastive Learning Approach for Unsupervised Discovery of Interpretable Directions
abstract
Recent research has shown that it is possible to find interpretable directions in the latent spaces of pre-trained Generative Adversarial Networks (GANs). These directions enable controllable image generation and support a wide range of semantic editing operations, such as zoom or rotation. The discovery of such directions is often done in a supervised or semi-supervised manner and requires manual annotations which limits their use in practice. In comparison, unsupervised discovery allows finding subtle directions that are difficult to detect a priori. In this work, we propose a contrastive learning-based approach to discover semantic directions in the latent space of pre-trained GANs in a self-supervised manner. Our approach finds semantically meaningful dimensions compatible with state-of-the-art methods.
Oguz Kaan Yüksel, Enis Simsar, Ezgi Gülperi Er, Pinar Yanardag Delul
ICCV1
2021 Semantic Perturbations with Normalizing Flows for Improved Generalization
abstract
Data augmentation is a widely adopted technique for avoiding overfitting when training deep neural networks. However, this approach requires domain-specific knowledge and is often limited to a fixed set of hard-coded transformations. Recently, several works proposed to use generative models for generating semantically meaningful perturbations to train a classifier. However, because accurate encoding and decoding are critical, these methods, which use architectures that approximate the latent-variable inference, remained limited to pilot studies on small datasets.Exploiting the exactly reversible encoder-decoder structure of normalizing flows, we perform on-manifold perturbations in the latent space to define fully unsupervised data augmentations. We demonstrate that such perturbations match the performance of advanced data augmentation techniques—reaching 96.6% test accuracy for CIFAR10 using ResNet-18 and outperform existing methods, particularly in low data regimes—yielding 10–25% relative improvement of test accuracy from classical training. We find that our latent adversarial perturbations adaptive to the classifier throughout its training are most effective, yielding the first test accuracy improvement results on real-world datasets—CIFAR-10/100—via latent-space perturbations.
Oguz Kaan Yüksel, Sebastian U. Stich, Martin Jaggi, Tatjana Chavdarova
ICCV1