Kento Nishi

dblp:286/8233 · DBLP profile ↗
← Back
6ranked-venue papers
4as first author
6since 2021 · last 2025
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 4 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Language models and text generation · 50% Learning paradigms · 17% Trustworthy machine learning · 17%

Topics — the 16 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
in-context learning
1.622025
ICLR: In-Context Learning of Representations · ICLR 2025
Towards an Understanding of Stepwise Inference in Transformers: A Synthetic Graph Navigation Model · ICML 2024
Machine learning › Trustworthy machine learning › robustness › learning with noisy labels
label noise robustness
1.022021
Augmentation Strategies for Learning With Noisy Labels · CVPR 2021
Improving Label Noise Robustness with Data Augmentation and Semi-Supervised Learning (Student Abstract) · AAAI 2021
Natural language and speech › Language models and text generation
knowledge editing
0.912025
Representation Shattering in Transformers: A Synthetic Study with Knowledge Editing · ICML 2025
Natural language and speech › Language models and text generation
large language model
0.912025
ICLR: In-Context Learning of Representations · ICLR 2025
Natural language and speech › Language models and text generation
chain-of-thought reasoning
0.812024
Towards an Understanding of Stepwise Inference in Transformers: A Synthetic Graph Navigation Model · ICML 2024
Natural language and speech › Language models and text generation
compositional generalization
0.812024
Towards an Understanding of Stepwise Inference in Transformers: A Synthetic Graph Navigation Model · ICML 2024
Machine learning › Learning paradigms
multi-task learning
0.812024
Joint-Task Regularization for Partially Labeled Multi-Task Learning · CVPR 2024
Machine learning › Learning paradigms › multi-task learning
partially annotated multi-task learning
0.812024
Joint-Task Regularization for Partially Labeled Multi-Task Learning · CVPR 2024
Machine learning › Deep learning architectures and training
data augmentation
0.512021
Augmentation Strategies for Learning With Noisy Labels · CVPR 2021
Machine learning › Trustworthy machine learning › robustness
learning with noisy labels
0.512021
Augmentation Strategies for Learning With Noisy Labels · CVPR 2021
Machine learning › Trustworthy machine learning
robustness
0.512021
Improving Label Noise Robustness with Data Augmentation and Semi-Supervised Learning (Student Abstract) · AAAI 2021
Machine learning › Learning paradigms
semi-supervised learning
0.512021
Improving Label Noise Robustness with Data Augmentation and Semi-Supervised Learning (Student Abstract) · AAAI 2021
Knowledge, reasoning and agents › Knowledge representation and reasoning › semantic representation
concept representation
0.312025
ICLR: In-Context Learning of Representations · ICLR 2025
Natural language and speech › Language models and text generation › neural language model
autoregressive transformer
0.212024
Towards an Understanding of Stepwise Inference in Transformers: A Synthetic Graph Navigation Model · ICML 2024
Machine learning › Deep learning architectures and training
transformer
0.212024
Towards an Understanding of Stepwise Inference in Transformers: A Synthetic Graph Navigation Model · ICML 2024
Computer vision › Image recognition and object detection
image classification
0.112021
Improving Label Noise Robustness with Data Augmentation and Semi-Supervised Learning (Student Abstract) · AAAI 2021

Methods — techniques the papers use, named apart from their topics

mechanistic analysis · 1.6synthetic knowledge graph task · 0.9graph tracing task · 0.9energy minimization · 0.9synthetic graph navigation task · 0.8joint-task regularization · 0.8semi-supervised learning · 0.5pseudo-labeling · 0.5loss filtering · 0.5data augmentation · 0.5
YearPublicationVenuePosition
2025 ICLR: In-Context Learning of Representations
abstract
Recent work demonstrates that structured patterns in pretraining data influence how representations of different concepts are organized in a large language model’s (LLM) internals, with such representations then driving downstream abilities. Given the open-ended nature of LLMs, e.g., their ability to in-context learn novel tasks, we ask whether models can flexibly alter their semantically grounded organization of concepts. Specifically, if we provide in-context exemplars wherein a concept plays a different role than what the pretraining data suggests, can models infer these novel semantics and reorganize representations in accordance with them? To answer this question, we define a toy “graph tracing” task wherein the nodes of the graph are referenced via concepts seen during training (e.g., apple, bird, etc.), and the connectivity of the graph is defined via some predefined structure (e.g., a square grid). Given exemplars that indicate traces of random walks on the graph, we analyze intermediate representations of the model and find that as the amount of context is scaled, there is a sudden re-organization of representations according to the graph’s structure. Further, we find that when reference concepts have correlations in their semantics (e.g., Monday, Tuesday, etc.), the context-specified graph structure is still present in the representations, but is unable to dominate the pretrained structure. To explain these results, we analogize our task to energy minimization for a predefined graph topology, which shows getting non-trivial performance on the task requires for the model to infer a connected component. Overall, our findings indicate context-size may be an underappreciated scaling axis that can flexibly re-organize model representations, unlocking novel capabilities.
Core Francisco Park, Andrew Lee 0001, Ekdeep Singh Lubana, Yongyi Yang, Maya Okawa, Kento Nishi, Martin Wattenberg, Hidenori Tanaka
ICLR6
2025 Representation Shattering in Transformers: A Synthetic Study with Knowledge Editing
abstract
Knowledge Editing (KE) algorithms alter models’ weights to perform targeted updates to incorrect, outdated, or otherwise unwanted factual associations. However, recent work has shown that applying KE can adversely affect models’ broader factual recall accuracy and diminish their reasoning abilities. Although these studies give insights into the potential harms of KE algorithms, e.g., performance evaluations on benchmarks, little is understood about why such destructive failures occur. Motivated by this, we define a novel synthetic task in which a Transformer is trained from scratch to internalize a "structured" knowledge graph. The structure enforces relationships between entities of the graph, such that editing a factual association has "trickling effects" on other entities (e.g., altering X’s parent is Y to Z affects who X’s siblings’ parent is). Through evaluations of edited models on this task, we show that KE inadvertently affects representations of entities beyond the targeted one, distorting relevant structures that allow a model to infer unseen knowledge about an entity. We call this phenomenon representation shattering and demonstrate that it degrades models’ factual recall and reasoning performance. We further corroborate our findings in naturalistic settings with pre-trained Llama and Mamba models as well. Overall, our work yields a precise mechanistic hypothesis to explain why KE has adverse effects on model abilities.
Kento Nishi, Rahul Ramesh, Maya Okawa, Mikail Khona, Hidenori Tanaka, Ekdeep Singh Lubana
ICML1
2024 Joint-Task Regularization for Partially Labeled Multi-Task Learning
abstract
Multi-task learning has become increasingly popular in the machine learning field, but its practicality is hindered by the need for large, labeled datasets. Most multi-task learning methods depend on fully labeled datasets wherein each input example is accompanied by ground-truth labels for all target tasks. Unfortunately, curating such datasets can be prohibitively expensive and impractical, especially for dense prediction tasks which require per-pixel labels for each image. With this in mind, we propose Joint-Task Regularization (JTR), an intuitive technique which leverages cross-task relations to simultaneously regularize all tasks in a single joint-task latent space to improve learning when data is not fully labeled for all tasks. JTR stands out from existing approaches in that it regularizes all tasks jointly rather than separately in pairs-therefore, it achieves linear complexity relative to the number of tasks while previous methods scale quadratically. To demonstrate the validity of our approach, we extensively benchmark our method across a wide variety of partially labeled scenarios based on NYU-v2, Cityscapes, and Taskonomy.
Kento Nishi, Junsik Kim 0001, Wanhua Li 0001, Hanspeter Pfister
CVPR1
2024 Towards an Understanding of Stepwise Inference in Transformers: A Synthetic Graph Navigation Model
abstract
Stepwise inference protocols, such as scratchpads and chain-of-thought, help language models solve complex problems by decomposing them into a sequence of simpler subproblems. To unravel the underlying mechanisms of stepwise inference we propose to study autoregressive Transformer models on a synthetic task that embodies the multi-step nature of problems where stepwise inference is generally most useful. Specifically, we define a graph navigation problem wherein a model is tasked with traversing a path from a start to a goal node on the graph. We find we can empirically reproduce and analyze several phenomena observed at scale: (i) the stepwise inference reasoning gap, the cause of which we find in the structure of the training data; (ii) a diversity-accuracy trade-off in model generations as sampling temperature varies; (iii) a simplicity bias in the model’s output; and (iv) compositional generalization and a primacy bias with in-context exemplars. Overall, our work introduces a grounded, synthetic framework for studying stepwise inference and offers mechanistic hypotheses that can lay the foundation for a deeper understanding of this phenomenon.
Mikail Khona, Maya Okawa, Jan Hula, Rahul Ramesh, Kento Nishi, Robert P. Dick, Ekdeep Singh Lubana, Hidenori Tanaka
ICML5
2021 Improving Label Noise Robustness with Data Augmentation and Semi-Supervised Learning (Student Abstract)
abstract
Modern machine learning algorithms typically require large amounts of labeled training data to fit a reliable model. To minimize the cost of data collection, researchers often employ techniques such as crowdsourcing and web scraping. However, web data and human annotations are known to exhibit high margins of error, resulting in sizable amounts of incorrect labels. Poorly labeled training data can cause models to overfit to the noise distribution, crippling performance in real-world applications. In this work, we investigate the viability of using data augmentation in conjunction with semi-supervised learning to improve the label noise robustness of image classification models. We conduct several experiments using noisy variants of the CIFAR-10 image classification dataset to benchmark our method against existing algorithms. Experimental results show that our augmentative SSL approach improves upon the state-of-the-art.
Kento Nishi, Yi Ding 0010, Alexander Rich 0001, Tobias Höllerer
AAAI1
2021 Augmentation Strategies for Learning With Noisy Labels
abstract
Imperfect labels are ubiquitous in real-world datasets. Several recent successful methods for training deep neural networks (DNNs) robust to label noise have used two primary techniques: filtering samples based on loss during a warm-up phase to curate an initial set of cleanly labeled samples, and using the output of a network as a pseudo-label for subsequent loss calculations. In this paper, we evaluate different augmentation strategies for algorithms tackling the "learning with noisy labels" problem. We propose and examine multiple augmentation strategies and evaluate them using synthetic datasets based on CIFAR-10 and CIFAR-100, as well as on the real-world dataset Clothing1M. Due to several commonalities in these algorithms, we find that using one set of augmentations for loss modeling tasks and another set for learning is the most effective, improving results on the state-of-the-art and other previous methods. Furthermore, we find that applying augmentation during the warm-up period can negatively impact the loss convergence behavior of correctly versus incorrectly labeled samples. We introduce this augmentation strategy to the state-of-the-art technique and demonstrate that we can improve performance across all evaluated noise levels. In particular, we improve accuracy on the CIFAR-10 benchmark at 90% symmetric noise by more than 15% in absolute accuracy, and we also improve performance on the Clothing1M dataset.
Kento Nishi, Yi Ding 0010, Alexander Rich 0001, Tobias Höllerer
CVPR1