Zhixuan Lin

dblp:254/1249 · DBLP profile ↗
← Back
5ranked-venue papers
4as first author
2since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 4 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Reinforcement learning · 34% Deep learning architectures and training · 32% 3D vision · 10%

Topics — the 17 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Deep learning architectures and training
attention mechanism
0.912025
Forgetting Transformer: Softmax Attention with a Forget Gate · ICLR 2025
Natural language and speech › Language models and text generation › language modeling
long-context language modeling
0.912025
Forgetting Transformer: Softmax Attention with a Forget Gate · ICLR 2025
Machine learning › Deep learning architectures and training
transformer
0.912025
Forgetting Transformer: Softmax Attention with a Forget Gate · ICLR 2025
Machine learning › Reinforcement learning
exploration
0.812024
The Curse of Diversity in Ensemble-Based Exploration · ICLR 2024
Machine learning › Reinforcement learning
off-policy reinforcement learning
0.812024
The Curse of Diversity in Ensemble-Based Exploration · ICLR 2024
Machine learning › Reinforcement learning › function approximation
representation learning for reinforcement learning
0.812024
The Curse of Diversity in Ensemble-Based Exploration · ICLR 2024
Machine learning › Reinforcement learning › model-based reinforcement learning › world model
object-centric world model
0.412020
Improving Generative Imagination in Object-Centric World Models · ICML 2020
Computer vision › 3D vision › 3d scene understanding
scene decomposition
0.412020
SPACE: Unsupervised Object-Oriented Scene Representation via Spatial Attention and Decomposition · ICLR 2020
Machine learning › Deep learning architectures and training › attention mechanism › visual attention
spatial attention
0.412020
SPACE: Unsupervised Object-Oriented Scene Representation via Spatial Attention and Decomposition · ICLR 2020
Machine learning › Learning paradigms
unsupervised learning
0.412020
SPACE: Unsupervised Object-Oriented Scene Representation via Spatial Attention and Decomposition · ICLR 2020
Machine learning › Reinforcement learning › model-based reinforcement learning
world model
0.412020
Improving Generative Imagination in Object-Centric World Models · ICML 2020
Machine learning › Deep learning architectures and training › convolutional neural network › convolution design
group convolution
0.412019
GIFT: Learning Transformation-Invariant Dense Visual Descriptors via Group CNNs · NeurIPS 2019
Machine learning › Deep learning architectures and training › equivariant neural network
group equivariant neural network
0.412019
GIFT: Learning Transformation-Invariant Dense Visual Descriptors via Group CNNs · NeurIPS 2019
Computer vision › 3D vision
local feature descriptor
0.412019
GIFT: Learning Transformation-Invariant Dense Visual Descriptors via Group CNNs · NeurIPS 2019
Machine learning › Representation and self-supervised learning › representation learning › invariant representation learning
transformation-invariant representation
0.412019
GIFT: Learning Transformation-Invariant Dense Visual Descriptors via Group CNNs · NeurIPS 2019
Machine learning › Representation and self-supervised learning › representation learning
object-centric representation learning
0.112020
Improving Generative Imagination in Object-Centric World Models · ICML 2020
Computer vision › 3D vision › camera pose estimation
relative pose estimation
0.112019
GIFT: Learning Transformation-Invariant Dense Visual Descriptors via Group CNNs · NeurIPS 2019

Methods — techniques the papers use, named apart from their topics

forget gate · 0.9flashattention · 0.9replay buffer · 0.8ensemble · 0.8cross-ensemble representation learning · 0.8variational inference · 0.4spatial attention · 0.4scene decomposition · 0.4generative model · 0.4feature pooling · 0.4
YearPublicationVenuePosition
2025 Forgetting Transformer: Softmax Attention with a Forget Gate
abstract
An essential component of modern recurrent sequence models is the forget gate. While Transformers do not have an explicit recurrent form, we show that a forget gate can be naturally incorporated into Transformers by down-weighting the unnormalized attention scores in a data-dependent way. We name this attention mechanism Forgetting Attention and the resulting model the Forgetting Transformer (FoX). We show that FoX outperforms the Transformer on long-context language modeling, length extrapolation, and short-context downstream tasks, while performing on par with the Transformer on long-context downstream tasks. Moreover, it is compatible with the FlashAttention algorithm and does not require any positional embeddings. Several analyses, including the needle-in-the-haystack test, show that FoX also retains the Transformer's superior long-context capabilities over recurrent sequence models such as Mamba-2, HGRN2, and DeltaNet. We also introduce a "Pro" block design that incorporates some common architectural components in recurrent sequence models and find it significantly improves the performance of both FoX and the Transformer. Our code is available at [`https://github.com/zhixuan-lin/forgetting-transformer`](https://github.com/zhixuan-lin/forgetting-transformer).
Zhixuan Lin, Evgenii Nikishin, Xu Owen He, Aaron C. Courville
ICLR1
2024 The Curse of Diversity in Ensemble-Based Exploration
abstract
We uncover a surprising phenomenon in deep reinforcement learning: training a diverse ensemble of data-sharing agents -- a well-established exploration strategy -- can significantly impair the performance of the individual ensemble members when compared to standard single-agent training. Through careful analysis, we attribute the degradation in performance to the low proportion of self-generated data in the shared training data for each ensemble member, as well as the inefficiency of the individual ensemble members to learn from such highly off-policy data. We thus name this phenomenon *the curse of diversity*. We find that several intuitive solutions -- such as a larger replay buffer or a smaller ensemble size -- either fail to consistently mitigate the performance loss or undermine the advantages of ensembling. Finally, we demonstrate the potential of representation learning to counteract the curse of diversity with a novel method named Cross-Ensemble Representation Learning (CERL) in both discrete and continuous control domains. Our work offers valuable insights into an unexpected pitfall in ensemble-based exploration and raises important caveats for future applications of similar approaches.
Zhixuan Lin, Pierluca D'Oro, Evgenii Nikishin, Aaron C. Courville
ICLR1
2020 SPACE: Unsupervised Object-Oriented Scene Representation via Spatial Attention and Decomposition
Zhixuan Lin, Yi-Fu Wu, Skand Vishwanath Peri, Weihao Sun, Gautam Singh, Fei Deng 0001, Jindong Jiang, Sungjin Ahn
ICLR1
2020 Improving Generative Imagination in Object-Centric World Models
abstract
The remarkable recent advances in object-centric generative world models raise a few questions. First, while many of the recent achievements are indispensable for making a general and versatile world model, it is quite unclear how these ingredients can be integrated into a unified framework. Second, despite using generative objectives, abilities for object detection and tracking are mainly investigated, leaving the crucial ability of temporal imagination largely under question. Third, a few key abilities for more faithful temporal imagination such as multimodal uncertainty and situation-awareness are missing. In this paper, we introduce Generative Structured World Models (G-SWM). The G-SWM achieves the versatile world modeling not only by unifying the key properties of previous models in a principled framework but also by achieving two crucial new abilities, multimodal uncertainty and situation-awareness. Our thorough investigation on the temporal generation ability in comparison to the previous models demonstrates that G-SWM achieves the versatility with the best or comparable performance for all experiment settings including a few complex settings that have not been tested before. https://sites.google.com/view/gswm
Zhixuan Lin, Yi-Fu Wu, Skand Vishwanath Peri, Bofeng Fu, Jindong Jiang, Sungjin Ahn
ICML1
2019 GIFT: Learning Transformation-Invariant Dense Visual Descriptors via Group CNNs
abstract
Finding local correspondences between images with different viewpoints requires local descriptors that are robust against geometric transformations. An approach for transformation invariance is to integrate out the transformations by pooling the features extracted from transformed versions of an image. However, the feature pooling may sacrifice the distinctiveness of the resulting descriptors. In this paper, we introduce a novel visual descriptor named Group Invariant Feature Transform (GIFT), which is both discriminative and robust to geometric transformations. The key idea is that the features extracted from the transformed versions of an image can be viewed as a function defined on the group of the transformations. Instead of feature pooling, we use group convolutions to exploit underlying structures of the extracted features on the group, resulting in descriptors that are both discriminative and provably invariant to the group of transformations. Extensive experiments show that GIFT outperforms state-of-the-art methods on several benchmark datasets and practically improves the performance of relative pose estimation.
Yuan Liu 0025, Zehong Shen, Zhixuan Lin, Sida Peng, Hujun Bao, Xiaowei Zhou 0001
NeurIPS3