VLDB 2026 Research / reviewers in the wild / expert
Tianzhe Chu
dblp:348/8957
· DBLP profile ↗
5ranked-venue papers
2as first author
5since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
3D vision · 35% Representation and self-supervised learning · 19% Deep learning architectures and training · 12% | |
| Databases, data mining, and information retrieval
1 paper |
Data mining · 100% |
Topics — the 13 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › sparse coding › sparse feature learning
sparse rate reduction |
1.4 | 2 | 2024 | White-Box Transformers via Sparse Rate Reduction: Compression Is All There Is? · J. Mach. Learn. Res. 2024 White-Box Transformers via Sparse Rate Reduction · NeurIPS 2023 |
Machine learning › Deep learning architectures and training
transformer |
1.4 | 2 | 2024 | White-Box Transformers via Sparse Rate Reduction: Compression Is All There Is? · J. Mach. Learn. Res. 2024 White-Box Transformers via Sparse Rate Reduction · NeurIPS 2023 |
Computer vision › 3D vision
camera pose estimation |
1.0 | 1 | 2026 | Seeing from Another Perspective: Evaluating Multi-View Understanding in MLLMs · AAAI 2026 |
Computer vision › 3D vision › visual localization
cross-view matching |
1.0 | 1 | 2026 | Seeing from Another Perspective: Evaluating Multi-View Understanding in MLLMs · AAAI 2026 |
Computer vision › Vision and language › vision-language model
multimodal large language model |
1.0 | 1 | 2026 | Seeing from Another Perspective: Evaluating Multi-View Understanding in MLLMs · AAAI 2026 |
Computer vision › 3D vision
multi-view geometry |
1.0 | 1 | 2026 | Seeing from Another Perspective: Evaluating Multi-View Understanding in MLLMs · AAAI 2026 |
Computer vision › 3D vision › 3d scene understanding
multi-view understanding |
1.0 | 1 | 2026 | Seeing from Another Perspective: Evaluating Multi-View Understanding in MLLMs · AAAI 2026 |
Machine learning › Learning theory
generalization |
0.9 | 1 | 2025 | SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training · ICML 2025 |
Machine learning › Reinforcement learning
reinforcement learning from human feedback |
0.9 | 1 | 2025 | SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training · ICML 2025 |
Machine learning › Representation and self-supervised learning › pre-training
pre-trained representations |
0.8 | 1 | 2024 | Image Clustering via the Principle of Rate Reduction in the Age of Pretrained Models · ICLR 2024 |
Data mining
clustering |
0.8 | 1 | 2024 | Image Clustering via the Principle of Rate Reduction in the Age of Pretrained Models · ICLR 2024 |
Data mining › clustering
image clustering |
0.8 | 1 | 2024 | Image Clustering via the Principle of Rate Reduction in the Age of Pretrained Models · ICLR 2024 |
Machine learning › Trustworthy machine learning
interpretability |
0.2 | 1 | 2023 | White-Box Transformers via Sparse Rate Reduction · NeurIPS 2023 |
Methods — techniques the papers use, named apart from their topics
rate reduction · 2.3self-labeling · 1.5CLIP · 1.5sparse coding · 1.4alternating optimization · 1.4question answering · 1.0benchmark evaluation · 1.0supervised fine-tuning · 0.9reinforcement learning · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Seeing from Another Perspective: Evaluating Multi-View Understanding in MLLMsabstractMulti-view understanding, the ability to reconcile visual information across diverse viewpoints for effective navigation, manipulation, and 3D scene comprehension, is a fundamental challenge in Multi-Modal Large Language Models (MLLMs) to be used as embodied agents. While recent MLLMs have shown impressive advances in high-level reasoning and planning, they frequently fall short when confronted with multi-view geometric consistency and cross-view correspondence. To comprehensively evaluate the challenges of MLLMs in multi-view scene reasoning, we introduce All-Angles Bench, a human carefully benchmark with over 2,100 question-answer pairs from 90 diverse, real-world scenes. Our broad evaluation across 38 general-purpose and 3D spatial reasoning MLLMs reveals a substantial performance gap compared to humans. More critically, our analysis identifies two root failure modes: (1) cross-view object mismatch—the inability to establish consistent object correspondence across views; and (2) cross-view spatial misalignment—the failure to infer accurate camera poses and spatial layouts. These findings underscore a lack of multi-view awareness in current MLLMs, calling for architectural innovations beyond prompt tuning alone. We believe that our benchmark offers valuable insights toward building spatially-intelligent MLLMs. Chun-Hsiao Yeh, Shengbang Tong, Ta Ying Cheng, Ruoyu Wang 0014, Tianzhe Chu, Yuexiang Zhai, Yubei Chen, Shenghua Gao, Yi Ma 0001 |
AAAI | 6 |
| 2025 | SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-trainingabstractSupervised fine-tuning (SFT) and reinforcement learning (RL) are widely used post-training techniques for foundation models. However, their roles in enhancing model generalization capabilities remain unclear. This paper studies the difference between SFT and RL on generalization and memorization, focusing on text-based rule variants and visual variants. We introduce GeneralPoints, an arithmetic reasoning card game, and adopt V-IRL, a real-world navigation environment, to assess how models trained with SFT and RL generalize to unseen variants in both textual and visual domains. We show that RL, especially when trained with an outcome-based reward, generalizes across both rule-based textual and visual variants. SFT, in contrast, tends to memorize training data and struggles to generalize out-of-distribution scenarios. Further analysis reveals that RL improves the model's underlying visual recognition capabilities, contributing to its enhanced generalization in the visual domain. Despite RL's superior generalization, we show that SFT remains essential for effective RL training; SFT stabilizes the model's output format, enabling subsequent RL to achieve its performance gains. These findings demonstrates the capability of RL for acquiring generalizable knowledge in complex, multi-modal tasks. Tianzhe Chu, Yuexiang Zhai, Jihan Yang, Shengbang Tong, Saining Xie, Dale Schuurmans, Quoc V. Le, Sergey Levine, Yi Ma 0001 |
ICML | 1 |
| 2024 | Image Clustering via the Principle of Rate Reduction in the Age of Pretrained ModelsabstractThe advent of large pre-trained models has brought about a paradigm shift in both visual representation learning and natural language processing. However, clustering unlabeled images, as a fundamental and classic machine learning problem, still lacks an effective solution, particularly for large-scale datasets. In this paper, we propose a novel image clustering pipeline that leverages the powerful feature representation of large pre-trained models such as CLIP and cluster images effectively and efficiently at scale. We first developed a novel algorithm to estimate the number of clusters in a given dataset. We then show that the pre-trained features are significantly more structured by further optimizing the rate reduction objective. The resulting features may significantly improve the clustering accuracy, e.g., from 57\% to 66\% on ImageNet-1k. Furthermore, by leveraging CLIP's multimodality bridge between image and text, we develop a simple yet effective self-labeling algorithm that produces meaningful text labels for the clusters. Through extensive experiments, we show that our pipeline works well on standard datasets such as CIFAR-10, CIFAR-100, and ImageNet-1k. It also extends to datasets without predefined labels, such as LAION-Aesthetics and WikiArts. Tianzhe Chu, Shengbang Tong, Tianjiao Ding, Xili Dai, Benjamin D. Haeffele, René Vidal, Yi Ma 0001 |
ICLR | 1 |
| 2024 | White-Box Transformers via Sparse Rate Reduction: Compression Is All There Is?abstractIn this paper, we contend that a natural objective of representation learning is to compress and transform the distribution of the data, say sets of tokens, towards a low-dimensional Gaussian mixture supported on incoherent subspaces. The goodness of such a representation can be evaluated by a principled measure, called sparse rate reduction, that simultaneously maximizes the intrinsic information gain and extrinsic sparsity of the learned representation. From this perspective, popular deep network architectures, including transformers, can be viewed as realizing iterative schemes to optimize this measure. Particularly, we derive a transformer block from alternating optimization on parts of this objective: the multi-head self-attention operator compresses the representation by implementing an approximate gradient descent step on the coding rate of the features, and the subsequent multi-layer perceptron sparsifies the features. This leads to a family of white-box transformer-like deep network architectures, named CRATE, which are mathematically fully interpretable. We show, by way of a novel connection between denoising and compression, that the inverse to the aforementioned compressive encoding can be realized by the same class of CRATE architectures. Thus, the so-derived white-box architectures are universal to both encoders and decoders. Experiments show that these networks, despite their simplicity, indeed learn to compress and sparsify representations of large-scale real-world image and text datasets, and achieve strong performance across different settings: ViT, MAE, DINO, BERT, and GPT2. We believe the proposed computational framework demonstrates great potential in bridging the gap between theory and practice of deep learning, from a unified perspective of data compression. Yaodong Yu, Sam Buchanan, Druv Pai, Tianzhe Chu, Ziyang Wu, Shengbang Tong, Yuexiang Zhai, Benjamin D. Haeffele, Yi Ma 0001 |
J. Mach. Learn. Res. | 4 |
| 2023 | White-Box Transformers via Sparse Rate ReductionabstractIn this paper, we contend that the objective of representation learning is to compress and transform the distribution of the data, say sets of tokens, towards a mixture of low-dimensional Gaussian distributions supported on incoherent subspaces. The quality of the final representation can be measured by a unified objective function called sparse rate reduction. From this perspective, popular deep networks such as transformers can be naturally viewed as realizing iterative schemes to optimize this objective incrementally. Particularly, we show that the standard transformer block can be derived from alternating optimization on complementary parts of this objective: the multi-head self-attention operator can be viewed as a gradient descent step to compress the token sets by minimizing their lossy coding rate, and the subsequent multi-layer perceptron can be viewed as attempting to sparsify the representation of the tokens. This leads to a family of white-box transformer-like deep network architectures which are mathematically fully interpretable. Despite their simplicity, experiments show that these networks indeed learn to optimize the designed objective: they compress and sparsify representations of large-scale real-world vision datasets such as ImageNet, and achieve performance very close to thoroughly engineered transformers such as ViT.
Code is at https://github.com/Ma-Lab-Berkeley/CRATE. Yaodong Yu, Sam Buchanan, Druv Pai, Tianzhe Chu, Ziyang Wu, Shengbang Tong, Benjamin D. Haeffele, Yi Ma 0001 |
NeurIPS | 4 |