Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Tianzhe Chu

dblp:348/8957 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
5since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
3D vision · 35% Representation and self-supervised learning · 19% Deep learning architectures and training · 12%
Databases, data mining, and information retrieval
1 paper
Data mining · 100%

Topics — the 13 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › sparse coding › sparse feature learning
sparse rate reduction
1.422024
White-Box Transformers via Sparse Rate Reduction: Compression Is All There Is? · J. Mach. Learn. Res. 2024
White-Box Transformers via Sparse Rate Reduction · NeurIPS 2023
Machine learning › Deep learning architectures and training
transformer
1.422024
White-Box Transformers via Sparse Rate Reduction: Compression Is All There Is? · J. Mach. Learn. Res. 2024
White-Box Transformers via Sparse Rate Reduction · NeurIPS 2023
Computer vision › 3D vision
camera pose estimation
1.012026
Seeing from Another Perspective: Evaluating Multi-View Understanding in MLLMs · AAAI 2026
Computer vision › 3D vision › visual localization
cross-view matching
1.012026
Seeing from Another Perspective: Evaluating Multi-View Understanding in MLLMs · AAAI 2026
Computer vision › Vision and language › vision-language model
multimodal large language model
1.012026
Seeing from Another Perspective: Evaluating Multi-View Understanding in MLLMs · AAAI 2026
Computer vision › 3D vision
multi-view geometry
1.012026
Seeing from Another Perspective: Evaluating Multi-View Understanding in MLLMs · AAAI 2026
Computer vision › 3D vision › 3d scene understanding
multi-view understanding
1.012026
Seeing from Another Perspective: Evaluating Multi-View Understanding in MLLMs · AAAI 2026
Machine learning › Learning theory
generalization
0.912025
SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training · ICML 2025
Machine learning › Reinforcement learning
reinforcement learning from human feedback
0.912025
SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training · ICML 2025
Machine learning › Representation and self-supervised learning › pre-training
pre-trained representations
0.812024
Image Clustering via the Principle of Rate Reduction in the Age of Pretrained Models · ICLR 2024
Data mining
clustering
0.812024
Image Clustering via the Principle of Rate Reduction in the Age of Pretrained Models · ICLR 2024
Data mining › clustering
image clustering
0.812024
Image Clustering via the Principle of Rate Reduction in the Age of Pretrained Models · ICLR 2024
Machine learning › Trustworthy machine learning
interpretability
0.212023
White-Box Transformers via Sparse Rate Reduction · NeurIPS 2023

Methods — techniques the papers use, named apart from their topics

rate reduction · 2.3self-labeling · 1.5CLIP · 1.5sparse coding · 1.4alternating optimization · 1.4question answering · 1.0benchmark evaluation · 1.0supervised fine-tuning · 0.9reinforcement learning · 0.9
YearPublicationVenuePosition
2026 Seeing from Another Perspective: Evaluating Multi-View Understanding in MLLMs
abstract
Multi-view understanding, the ability to reconcile visual information across diverse viewpoints for effective navigation, manipulation, and 3D scene comprehension, is a fundamental challenge in Multi-Modal Large Language Models (MLLMs) to be used as embodied agents. While recent MLLMs have shown impressive advances in high-level reasoning and planning, they frequently fall short when confronted with multi-view geometric consistency and cross-view correspondence. To comprehensively evaluate the challenges of MLLMs in multi-view scene reasoning, we introduce All-Angles Bench, a human carefully benchmark with over 2,100 question-answer pairs from 90 diverse, real-world scenes. Our broad evaluation across 38 general-purpose and 3D spatial reasoning MLLMs reveals a substantial performance gap compared to humans. More critically, our analysis identifies two root failure modes: (1) cross-view object mismatch—the inability to establish consistent object correspondence across views; and (2) cross-view spatial misalignment—the failure to infer accurate camera poses and spatial layouts. These findings underscore a lack of multi-view awareness in current MLLMs, calling for architectural innovations beyond prompt tuning alone. We believe that our benchmark offers valuable insights toward building spatially-intelligent MLLMs.
Chun-Hsiao Yeh, Shengbang Tong, Ta Ying Cheng, Ruoyu Wang 0014, Tianzhe Chu, Yuexiang Zhai, Yubei Chen, Shenghua Gao, Yi Ma 0001
AAAI6
2025 SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training
abstract
Supervised fine-tuning (SFT) and reinforcement learning (RL) are widely used post-training techniques for foundation models. However, their roles in enhancing model generalization capabilities remain unclear. This paper studies the difference between SFT and RL on generalization and memorization, focusing on text-based rule variants and visual variants. We introduce GeneralPoints, an arithmetic reasoning card game, and adopt V-IRL, a real-world navigation environment, to assess how models trained with SFT and RL generalize to unseen variants in both textual and visual domains. We show that RL, especially when trained with an outcome-based reward, generalizes across both rule-based textual and visual variants. SFT, in contrast, tends to memorize training data and struggles to generalize out-of-distribution scenarios. Further analysis reveals that RL improves the model's underlying visual recognition capabilities, contributing to its enhanced generalization in the visual domain. Despite RL's superior generalization, we show that SFT remains essential for effective RL training; SFT stabilizes the model's output format, enabling subsequent RL to achieve its performance gains. These findings demonstrates the capability of RL for acquiring generalizable knowledge in complex, multi-modal tasks.
Tianzhe Chu, Yuexiang Zhai, Jihan Yang, Shengbang Tong, Saining Xie, Dale Schuurmans, Quoc V. Le, Sergey Levine, Yi Ma 0001
ICML1
2024 Image Clustering via the Principle of Rate Reduction in the Age of Pretrained Models
abstract
The advent of large pre-trained models has brought about a paradigm shift in both visual representation learning and natural language processing. However, clustering unlabeled images, as a fundamental and classic machine learning problem, still lacks an effective solution, particularly for large-scale datasets. In this paper, we propose a novel image clustering pipeline that leverages the powerful feature representation of large pre-trained models such as CLIP and cluster images effectively and efficiently at scale. We first developed a novel algorithm to estimate the number of clusters in a given dataset. We then show that the pre-trained features are significantly more structured by further optimizing the rate reduction objective. The resulting features may significantly improve the clustering accuracy, e.g., from 57\% to 66\% on ImageNet-1k. Furthermore, by leveraging CLIP's multimodality bridge between image and text, we develop a simple yet effective self-labeling algorithm that produces meaningful text labels for the clusters. Through extensive experiments, we show that our pipeline works well on standard datasets such as CIFAR-10, CIFAR-100, and ImageNet-1k. It also extends to datasets without predefined labels, such as LAION-Aesthetics and WikiArts.
Tianzhe Chu, Shengbang Tong, Tianjiao Ding, Xili Dai, Benjamin D. Haeffele, René Vidal, Yi Ma 0001
ICLR1
2024 White-Box Transformers via Sparse Rate Reduction: Compression Is All There Is?
abstract
In this paper, we contend that a natural objective of representation learning is to compress and transform the distribution of the data, say sets of tokens, towards a low-dimensional Gaussian mixture supported on incoherent subspaces. The goodness of such a representation can be evaluated by a principled measure, called sparse rate reduction, that simultaneously maximizes the intrinsic information gain and extrinsic sparsity of the learned representation. From this perspective, popular deep network architectures, including transformers, can be viewed as realizing iterative schemes to optimize this measure. Particularly, we derive a transformer block from alternating optimization on parts of this objective: the multi-head self-attention operator compresses the representation by implementing an approximate gradient descent step on the coding rate of the features, and the subsequent multi-layer perceptron sparsifies the features. This leads to a family of white-box transformer-like deep network architectures, named CRATE, which are mathematically fully interpretable. We show, by way of a novel connection between denoising and compression, that the inverse to the aforementioned compressive encoding can be realized by the same class of CRATE architectures. Thus, the so-derived white-box architectures are universal to both encoders and decoders. Experiments show that these networks, despite their simplicity, indeed learn to compress and sparsify representations of large-scale real-world image and text datasets, and achieve strong performance across different settings: ViT, MAE, DINO, BERT, and GPT2. We believe the proposed computational framework demonstrates great potential in bridging the gap between theory and practice of deep learning, from a unified perspective of data compression.
Yaodong Yu, Sam Buchanan, Druv Pai, Tianzhe Chu, Ziyang Wu, Shengbang Tong, Yuexiang Zhai, Benjamin D. Haeffele, Yi Ma 0001
J. Mach. Learn. Res.4
2023 White-Box Transformers via Sparse Rate Reduction
abstract
In this paper, we contend that the objective of representation learning is to compress and transform the distribution of the data, say sets of tokens, towards a mixture of low-dimensional Gaussian distributions supported on incoherent subspaces. The quality of the final representation can be measured by a unified objective function called sparse rate reduction. From this perspective, popular deep networks such as transformers can be naturally viewed as realizing iterative schemes to optimize this objective incrementally. Particularly, we show that the standard transformer block can be derived from alternating optimization on complementary parts of this objective: the multi-head self-attention operator can be viewed as a gradient descent step to compress the token sets by minimizing their lossy coding rate, and the subsequent multi-layer perceptron can be viewed as attempting to sparsify the representation of the tokens. This leads to a family of white-box transformer-like deep network architectures which are mathematically fully interpretable. Despite their simplicity, experiments show that these networks indeed learn to optimize the designed objective: they compress and sparsify representations of large-scale real-world vision datasets such as ImageNet, and achieve performance very close to thoroughly engineered transformers such as ViT. Code is at https://github.com/Ma-Lab-Berkeley/CRATE.
Yaodong Yu, Sam Buchanan, Druv Pai, Tianzhe Chu, Ziyang Wu, Shengbang Tong, Benjamin D. Haeffele, Yi Ma 0001
NeurIPS4