Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Jiayi Ni

dblp:357/5198 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
4since 2021 · last 2025
0000-0002-1941-8170ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Representation and self-supervised learning · 52% Transfer learning and domain adaptation · 19% Efficient and distributed learning · 14%

Topics — the 12 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning
self-supervised representation learning
1.622025
Dataset Distillation via Knowledge Distillation: Towards Efficient Self-Supervised Pre-training of Deep Networks · ICLR 2025
Investigating the Benefits of Projection Head for Representation Learning · ICLR 2024
Machine learning › Efficient and distributed learning
dataset distillation
0.912025
Dataset Distillation via Knowledge Distillation: Towards Efficient Self-Supervised Pre-training of Deep Networks · ICLR 2025
Machine learning › Transfer learning and domain adaptation › test-time adaptation
continual test-time adaptation
0.812024
Distribution-Aware Continual Test-Time Adaptation for Semantic Segmentation · ICRA 2024
Machine learning › Representation and self-supervised learning
contrastive learning
0.812024
Investigating the Benefits of Projection Head for Representation Learning · ICLR 2024
Machine learning › Representation and self-supervised learning › contrastive learning
projection head
0.812024
Investigating the Benefits of Projection Head for Representation Learning · ICLR 2024
Machine learning › Representation and self-supervised learning › representation analysis
representation learning theory
0.812024
Investigating the Benefits of Projection Head for Representation Learning · ICLR 2024
Computer vision › Segmentation and scene understanding
semantic segmentation
0.812024
Distribution-Aware Continual Test-Time Adaptation for Semantic Segmentation · ICRA 2024
Machine learning › Transfer learning and domain adaptation
test-time adaptation
0.812024
Distribution-Aware Continual Test-Time Adaptation for Semantic Segmentation · ICRA 2024
Machine learning › Efficient and distributed learning › model compression
knowledge distillation
0.312025
Dataset Distillation via Knowledge Distillation: Towards Efficient Self-Supervised Pre-training of Deep Networks · ICLR 2025
Machine learning › Representation and self-supervised learning
representation matching
0.312025
Dataset Distillation via Knowledge Distillation: Towards Efficient Self-Supervised Pre-training of Deep Networks · ICLR 2025
Robotics › Autonomous driving
perception
0.212024
Distribution-Aware Continual Test-Time Adaptation for Semantic Segmentation · ICRA 2024
Machine learning › Trustworthy machine learning
robustness
0.212024
Investigating the Benefits of Projection Head for Representation Learning · ICLR 2024

Methods — techniques the papers use, named apart from their topics

trajectory matching · 0.9self-supervised learning · 0.9knowledge distillation · 0.9self-training · 0.8pseudo-labeling · 0.8parameter-efficient fine-tuning · 0.8layer-wise feature weighting analysis · 0.8contrastive loss · 0.8
YearPublicationVenuePosition
2025 Dataset Distillation via Knowledge Distillation: Towards Efficient Self-Supervised Pre-training of Deep Networks
abstract
Dataset distillation (DD) generates small synthetic datasets that can efficiently train deep networks with a limited amount of memory and compute. Despite the success of DD methods for supervised learning, DD for self-supervised pre-training of deep models has remained unaddressed. Pre-training on unlabeled data is crucial for efficiently generalizing to downstream tasks with limited labeled data. In this work, we propose the first effective DD method for SSL pre-training. First, we show, theoretically and empirically, that naiive application of supervised DD methods to SSL fails, due to the high variance of the SSL gradient. Then, we address this issue by relying on insights from knowledge distillation (KD) literature. Specifically, we train a small student model to match the representations of a larger teacher model trained with SSL. Then, we generate a small synthetic dataset by matching the training trajectories of the student models. As the KD objective has considerably lower variance than SSL, our approach can generate synthetic datasets that can successfully pre-train high-quality encoders. Through extensive experiments, we show that our distilled sets lead to up to 13% higher accuracy than prior work, on a variety of downstream tasks, in the presence of limited labeled data. Code at https://github.com/BigML-CS-UCLA/MKDT.
Siddharth Joshi 0004, Jiayi Ni, Baharan Mirzasoleiman
ICLR2
2024 Investigating the Benefits of Projection Head for Representation Learning
abstract
An effective technique for obtaining high-quality representations is adding a projection head on top of the encoder during training, then discarding it and using the pre-projection representations. Despite its proven practical effectiveness, the reason behind the success of this technique is poorly understood. The pre-projection representations are not directly optimized by the loss function, raising the question: what makes them better? In this work, we provide a rigorous theoretical answer to this question. We start by examining linear models trained with self-supervised contrastive loss. We reveal that the implicit bias of training algorithms leads to layer-wise progressive feature weighting, where features become increasingly unequal as we go deeper into the layers. Consequently, lower layers tend to have more normalized and less specialized representations. We theoretically characterize scenarios where such representations are more beneficial, highlighting the intricate interplay between data augmentation and input features. Additionally, we demonstrate that introducing non-linearity into the network allows lower layers to learn features that are completely absent in higher layers. Finally, we show how this mechanism improves the robustness in supervised contrastive learning and supervised learning. We empirically validate our results through various experiments on CIFAR-10/100, UrbanCars and shifted versions of ImageNet. We also introduce a potential alternative to projection head, which offers a more interpretable and controllable design.
Yihao Xue, Eric Gan, Jiayi Ni, Siddharth Joshi 0004, Baharan Mirzasoleiman
ICLR3
2024 Distribution-Aware Continual Test-Time Adaptation for Semantic Segmentation
abstract
Since autonomous driving systems usually face dynamic and ever-changing environments, continual test-time adaptation (CTTA) has been proposed as a strategy for transferring deployed models to continually changing target domains. However, the pursuit of long-term adaptation often introduces catastrophic forgetting and error accumulation problems, which impede the practical implementation of CTTA in the real world. Recently, existing CTTA methods mainly focus on utilizing a majority of parameters to fit target domain knowledge through self-training. Unfortunately, these approaches often amplify the challenge of error accumulation due to noisy pseudo-labels, and pose practical limitations stemming from the heavy computational costs associated with entire model updates. In this paper, we propose a distribution-aware tuning (DAT) method to make the semantic segmentation CTTA efficient and practical in real-world applications. DAT adaptively selects and updates two small groups of trainable parameters based on data distribution during the continual adaptation process, including domain-specific parameters (DSP) and task-relevant parameters (TRP). Specifically, DSP exhibits sensitivity to outputs with substantial distribution shifts, effectively mitigating the problem of error accumulation. In contrast, TRP are allocated to positions that are responsive to outputs with minor distribution shifts, which are fine-tuned to avoid the catastrophic forgetting problem. In addition, since CTTA is a temporal task, we introduce the Parameter Accumulation Update (PAU) strategy to collect the updated DSP and TRP in target domain sequences. We conducted extensive experiments on two widely-used semantic segmentation CTTA benchmarks, achieving competitive performance and efficiency compared to previous state-of-the-art methods.
Jiayi Ni, Senqiao Yang, Ran Xu 0013, Jiaming Liu 0003, Xiaoqi Li 0020, Wenyu Jiao, Shanghang Zhang
ICRA1
2023 High-speed anomaly traffic detection based on staged frequency domain features
Jiayi Ni, Wei Chen 0006, Jiacheng Tong, Haiyong Wang, Lifa Wu
J. Inf. Secur. Appl.1