VLDB 2026 Research / reviewers in the wild / expert
Jiayi Ni
dblp:357/5198
· DBLP profile ↗
4ranked-venue papers
2as first author
4since 2021 · last 2025
0000-0002-1941-8170ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Representation and self-supervised learning · 52% Transfer learning and domain adaptation · 19% Efficient and distributed learning · 14% |
Topics — the 12 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning
self-supervised representation learning |
1.6 | 2 | 2025 | Dataset Distillation via Knowledge Distillation: Towards Efficient Self-Supervised Pre-training of Deep Networks · ICLR 2025 Investigating the Benefits of Projection Head for Representation Learning · ICLR 2024 |
Machine learning › Efficient and distributed learning
dataset distillation |
0.9 | 1 | 2025 | Dataset Distillation via Knowledge Distillation: Towards Efficient Self-Supervised Pre-training of Deep Networks · ICLR 2025 |
Machine learning › Transfer learning and domain adaptation › test-time adaptation
continual test-time adaptation |
0.8 | 1 | 2024 | Distribution-Aware Continual Test-Time Adaptation for Semantic Segmentation · ICRA 2024 |
Machine learning › Representation and self-supervised learning
contrastive learning |
0.8 | 1 | 2024 | Investigating the Benefits of Projection Head for Representation Learning · ICLR 2024 |
Machine learning › Representation and self-supervised learning › contrastive learning
projection head |
0.8 | 1 | 2024 | Investigating the Benefits of Projection Head for Representation Learning · ICLR 2024 |
Machine learning › Representation and self-supervised learning › representation analysis
representation learning theory |
0.8 | 1 | 2024 | Investigating the Benefits of Projection Head for Representation Learning · ICLR 2024 |
Computer vision › Segmentation and scene understanding
semantic segmentation |
0.8 | 1 | 2024 | Distribution-Aware Continual Test-Time Adaptation for Semantic Segmentation · ICRA 2024 |
Machine learning › Transfer learning and domain adaptation
test-time adaptation |
0.8 | 1 | 2024 | Distribution-Aware Continual Test-Time Adaptation for Semantic Segmentation · ICRA 2024 |
Machine learning › Efficient and distributed learning › model compression
knowledge distillation |
0.3 | 1 | 2025 | Dataset Distillation via Knowledge Distillation: Towards Efficient Self-Supervised Pre-training of Deep Networks · ICLR 2025 |
Machine learning › Representation and self-supervised learning
representation matching |
0.3 | 1 | 2025 | Dataset Distillation via Knowledge Distillation: Towards Efficient Self-Supervised Pre-training of Deep Networks · ICLR 2025 |
Robotics › Autonomous driving
perception |
0.2 | 1 | 2024 | Distribution-Aware Continual Test-Time Adaptation for Semantic Segmentation · ICRA 2024 |
Machine learning › Trustworthy machine learning
robustness |
0.2 | 1 | 2024 | Investigating the Benefits of Projection Head for Representation Learning · ICLR 2024 |
Methods — techniques the papers use, named apart from their topics
trajectory matching · 0.9self-supervised learning · 0.9knowledge distillation · 0.9self-training · 0.8pseudo-labeling · 0.8parameter-efficient fine-tuning · 0.8layer-wise feature weighting analysis · 0.8contrastive loss · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Dataset Distillation via Knowledge Distillation: Towards Efficient Self-Supervised Pre-training of Deep NetworksabstractDataset distillation (DD) generates small synthetic datasets that can efficiently train deep networks with a limited amount of memory and compute. Despite the success of DD methods for supervised learning, DD for self-supervised pre-training of deep models has remained unaddressed. Pre-training on unlabeled data is crucial for efficiently generalizing to downstream tasks with limited labeled data. In this work, we propose the first effective DD method for SSL pre-training. First, we show, theoretically and empirically, that naiive application of supervised DD methods to SSL fails, due to the high variance of the SSL gradient. Then, we address this issue by relying on insights from knowledge distillation (KD) literature. Specifically, we train a small student model to match the representations of a larger teacher model trained with SSL. Then, we generate a small synthetic dataset by matching the training trajectories of the student models. As the KD objective has considerably lower variance than SSL, our approach can generate synthetic datasets that can successfully pre-train high-quality encoders. Through extensive experiments, we show that our distilled sets lead to up to 13% higher accuracy than prior work, on a variety of downstream tasks, in the presence of limited labeled data. Code at https://github.com/BigML-CS-UCLA/MKDT. Siddharth Joshi 0004, Jiayi Ni, Baharan Mirzasoleiman |
ICLR | 2 |
| 2024 | Investigating the Benefits of Projection Head for Representation LearningabstractAn effective technique for obtaining high-quality representations is adding a projection head on top of the encoder during training, then discarding it and using the pre-projection representations. Despite its proven practical effectiveness, the reason behind the success of this technique is poorly understood. The pre-projection representations are not directly optimized by the loss function, raising the question: what makes them better? In this work, we provide a rigorous theoretical answer to this question. We start by examining linear models trained with self-supervised contrastive loss. We reveal that the implicit bias of training algorithms leads to layer-wise progressive feature weighting, where features become increasingly unequal as we go deeper into the layers. Consequently, lower layers tend to have more normalized and less specialized representations. We theoretically characterize scenarios where such representations are more beneficial, highlighting the intricate interplay between data augmentation and input features. Additionally, we demonstrate that introducing non-linearity into the network allows lower layers to learn features that are completely absent in higher layers. Finally, we show how this mechanism improves the robustness in supervised contrastive learning and supervised learning. We empirically validate our results through various experiments on CIFAR-10/100, UrbanCars and shifted versions of ImageNet. We also introduce a potential alternative to projection head, which offers a more interpretable and controllable design. Yihao Xue, Eric Gan, Jiayi Ni, Siddharth Joshi 0004, Baharan Mirzasoleiman |
ICLR | 3 |
| 2024 | Distribution-Aware Continual Test-Time Adaptation for Semantic SegmentationabstractSince autonomous driving systems usually face dynamic and ever-changing environments, continual test-time adaptation (CTTA) has been proposed as a strategy for transferring deployed models to continually changing target domains. However, the pursuit of long-term adaptation often introduces catastrophic forgetting and error accumulation problems, which impede the practical implementation of CTTA in the real world. Recently, existing CTTA methods mainly focus on utilizing a majority of parameters to fit target domain knowledge through self-training. Unfortunately, these approaches often amplify the challenge of error accumulation due to noisy pseudo-labels, and pose practical limitations stemming from the heavy computational costs associated with entire model updates. In this paper, we propose a distribution-aware tuning (DAT) method to make the semantic segmentation CTTA efficient and practical in real-world applications. DAT adaptively selects and updates two small groups of trainable parameters based on data distribution during the continual adaptation process, including domain-specific parameters (DSP) and task-relevant parameters (TRP). Specifically, DSP exhibits sensitivity to outputs with substantial distribution shifts, effectively mitigating the problem of error accumulation. In contrast, TRP are allocated to positions that are responsive to outputs with minor distribution shifts, which are fine-tuned to avoid the catastrophic forgetting problem. In addition, since CTTA is a temporal task, we introduce the Parameter Accumulation Update (PAU) strategy to collect the updated DSP and TRP in target domain sequences. We conducted extensive experiments on two widely-used semantic segmentation CTTA benchmarks, achieving competitive performance and efficiency compared to previous state-of-the-art methods. Jiayi Ni, Senqiao Yang, Ran Xu 0013, Jiaming Liu 0003, Xiaoqi Li 0020, Wenyu Jiao, Shanghang Zhang |
ICRA | 1 |
| 2023 | High-speed anomaly traffic detection based on staged frequency domain features
Jiayi Ni, Wei Chen 0006, Jiacheng Tong, Haiyong Wang, Lifa Wu |
J. Inf. Secur. Appl. | 1 |