Kaiwen Zha

dblp:213/6159 · DBLP profile ↗
← Back
9ranked-venue papers
3as first author
7since 2021 · last 2025
0000-0002-1552-5729ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 3 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
9 papers
Generative modeling · 23% Representation and self-supervised learning · 19% Language models and text generation · 12%
Network and information security
1 paper
Security and privacy of machine learning · 100%

Topics — the 23 heaviest of 27, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Representation and self-supervised learning
contrastive learning
1.932023
Rank-N-Contrast: Learning Continuous Representations for Regression · NeurIPS 2023
Indiscriminate Poisoning Attacks on Unsupervised Contrastive Learning · ICLR 2023
Unsupervised Representation for Semantic Segmentation by Implicit Cycle-Attention Contrastive Learning · AAAI 2022
Machine learning › Representation and self-supervised learning › contrastive learning
self-supervised contrastive learning
1.222023
Indiscriminate Poisoning Attacks on Unsupervised Contrastive Learning · ICLR 2023
Unsupervised Representation for Semantic Segmentation by Implicit Cycle-Attention Contrastive Learning · AAAI 2022
Machine learning › Generative modeling › diffusion model
conditional generation
0.912025
REG: Rectified Gradient Guidance for Conditional Diffusion Models · ICML 2025
Machine learning › Generative modeling
diffusion model
0.912025
REG: Rectified Gradient Guidance for Conditional Diffusion Models · ICML 2025
Machine learning › Generative modeling
image generation
0.912025
Language-Guided Image Tokenization for Generation · CVPR 2025
Machine learning › Generative modeling
image tokenization
0.912025
Language-Guided Image Tokenization for Generation · CVPR 2025
Natural language and speech › Language models and text generation
large language model reasoning
0.912025
RL Tango: Reinforcing Generator and Verifier Together for Language Reasoning · NeurIPS 2025
Natural language and speech › Language models and text generation › large language model training › post-training
reinforcement learning post-training
0.912025
RL Tango: Reinforcing Generator and Verifier Together for Language Reasoning · NeurIPS 2025
Machine learning › Reinforcement learning › reward learning
reward modeling
0.912025
RL Tango: Reinforcing Generator and Verifier Together for Language Reasoning · NeurIPS 2025
Machine learning › Learning paradigms › supervised learning › neural network regression
deep regression
0.712023
Rank-N-Contrast: Learning Continuous Representations for Regression · NeurIPS 2023
Security and privacy of machine learning
poisoning attack
0.712023
Indiscriminate Poisoning Attacks on Unsupervised Contrastive Learning · ICLR 2023
Computer vision › Segmentation and scene understanding
semantic segmentation
0.612022
Unsupervised Representation for Semantic Segmentation by Implicit Cycle-Attention Contrastive Learning · AAAI 2022
Computer vision › Segmentation and scene understanding › image segmentation
unsupervised segmentation
0.612022
Unsupervised Representation for Semantic Segmentation by Implicit Cycle-Attention Contrastive Learning · AAAI 2022
Machine learning › Learning paradigms › imbalanced learning
deep imbalanced regression
0.512021
Delving into Deep Imbalanced Regression · ICML 2021
Machine learning › Trustworthy machine learning
fairness
0.512021
Delving into Deep Imbalanced Regression · ICML 2021
Machine learning › Learning paradigms
imbalanced learning
0.512021
Delving into Deep Imbalanced Regression · ICML 2021
Machine learning › Trustworthy machine learning › robustness › distributional robustness
subgroup robustness
0.512021
Delving into Deep Imbalanced Regression · ICML 2021
Computer vision › Video understanding and tracking
action recognition
0.412020
Further Understanding Videos through Adverbs: A New Video Task · AAAI 2020
Machine learning › Deep learning architectures and training
recurrent neural network
0.412019
Deep RNN Framework for Visual Sequential Applications · CVPR 2019
Computer vision › Video understanding and tracking
video classification
0.412019
Deep RNN Framework for Visual Sequential Applications · CVPR 2019
Natural language and speech › Language models and text generation
mathematical reasoning
0.312025
RL Tango: Reinforcing Generator and Verifier Together for Language Reasoning · NeurIPS 2025
Machine learning › Generative modeling › diffusion model
text-to-image generation
0.312025
REG: Rectified Gradient Guidance for Conditional Diffusion Models · ICML 2025
Machine learning › Trustworthy machine learning › robustness › distribution shift
robustness to distribution shift
0.212023
Rank-N-Contrast: Learning Continuous Representations for Regression · NeurIPS 2023

Methods — techniques the papers use, named apart from their topics

contrastive learning · 2.0poisoning attack · 1.3vector quantization · 0.9supervised fine-tuning · 0.9scaled joint distribution objective · 0.9reward hacking mitigation · 0.9reinforcement learning · 0.9gradient guidance · 0.9diffusion transformer · 0.9ranking-based contrast · 0.7
YearPublicationVenuePosition
2025 Language-Guided Image Tokenization for Generation
abstract
Image tokenization, the process of transforming raw image pixels into a compact low-dimensional latent representation, has proven crucial for scalable and efficient image generation. However, mainstream image tokenization methods generally have limited compression rates, making high-resolution image generation computationally expensive. To address this challenge, we propose to leverage language for efficient image tokenization, and we call our method Text-Conditioned Image Tokenization (TexTok). TexTok is a simple yet effective tokenization framework that leverages language to provide a compact, high-level semantic representation. By conditioning the tokenization process on descriptive text captions, TexTok simplifies semantic learning, allowing more learning capacity and token space to be allocated to capture fine-grained visual details, leading to enhanced reconstruction quality and higher compression rates. Compared to the conventional tokenizer without text conditioning, TexTok achieves average reconstruction FID improvements of 29.2% and 48.1% on ImageNet-256 and -512 benchmarks respectively, across varying numbers of tokens. These tokenization improvements consistently translate to 16.3% and 34.3% average improvements in generation FID. By simply replacing the tokenizer in Diffusion Transformer (DiT) with TexTok, our system can achieve a 93.5× inference speedup while still outperforming the original DiT using only 32 tokens on ImageNet-512. TexTok with a vanilla DiT generator achieves state-of-the-art FID scores of 1.46 and 1.62 on ImageNet-256 and -512 respectively. Furthermore, we demonstrate TexTok’s superiority on the text-to-image generation task, effectively utilizing the off-the-shelf text captions in tokenization.
Kaiwen Zha, Lijun Yu, Alireza Fathi, David A. Ross, Cordelia Schmid, Dina Katabi, Xiuye Gu
CVPR1
2025 REG: Rectified Gradient Guidance for Conditional Diffusion Models
abstract
Guidance techniques are simple yet effective for improving conditional generation in diffusion models. Albeit their empirical success, the practical implementation of guidance diverges significantly from its theoretical motivation. In this paper, we reconcile this discrepancy by replacing the scaled marginal distribution target, which we prove theoretically invalid, with a valid scaled joint distribution objective. Additionally, we show that the established guidance implementations are approximations to the intractable optimal solution under no future foresight constraint. Building on these theoretical insights, we propose rectified gradient guidance (REG), a versatile enhancement designed to boost the performance of existing guidance methods. Experiments on 1D and 2D demonstrate that REG provides a better approximation to the optimal solution than prior guidance techniques, validating the proposed theoretical framework. Extensive experiments on class-conditional ImageNet and text-to-image generation tasks show that incorporating REG consistently improves FID and Inception/CLIP scores across various settings compared to its absence.
Zhengqi Gao, Kaiwen Zha, Zihui Xue, Duane S. Boning
ICML2
2025 RL Tango: Reinforcing Generator and Verifier Together for Language Reasoning
abstract
Reinforcement learning (RL) has recently emerged as a compelling approach for enhancing the reasoning capabilities of large language models (LLMs), where an LLM generator serves as a policy guided by a verifier (reward model). However, current RL post-training methods for LLMs typically use verifiers that are fixed (rule-based or frozen pretrained) or trained discriminatively via supervised fine-tuning (SFT). Such designs are susceptible to reward hacking and generalize poorly beyond their training distributions. To overcome these limitations, we propose Tango, a novel framework that uses RL to concurrently train both an LLM generator and a verifier in an interleaved manner. A central innovation of Tango is its generative, process-level LLM verifier, which is trained via RL and co-evolves with the generator. Importantly, the verifier is trained solely based on outcome-level verification correctness rewards without requiring explicit process-level annotations. This generative RL-trained verifier exhibits improved robustness and superior generalization compared to deterministic or SFT-trained verifiers, fostering effective mutual reinforcement with the generator. Extensive experiments demonstrate that both components of Tango achieve state-of-the-art results among 7B/8B-scale models: the generator attains best-in-class performance across five competition-level math benchmarks and four challenging out-of-domain reasoning tasks, while the verifier leads on the ProcessBench dataset. Remarkably, both components exhibit particularly substantial improvements on the most difficult mathematical reasoning problems.
Kaiwen Zha, Zhengqi Gao, Maohao Shen, Zhang-Wei Hong, Duane S. Boning, Dina Katabi
NeurIPS1
2023 Indiscriminate Poisoning Attacks on Unsupervised Contrastive Learning
Hao He 0011, Kaiwen Zha, Dina Katabi
ICLR2
2023 Rank-N-Contrast: Learning Continuous Representations for Regression
abstract
Deep regression models typically learn in an end-to-end fashion without explicitly emphasizing a regression-aware representation. Consequently, the learned representations exhibit fragmentation and fail to capture the continuous nature of sample orders, inducing suboptimal results across a wide range of regression tasks. To fill the gap, we propose Rank-N-Contrast (RNC), a framework that learns continuous representations for regression by contrasting samples against each other based on their rankings in the target space. We demonstrate, theoretically and empirically, that RNC guarantees the desired order of learned representations in accordance with the target orders, enjoying not only better performance but also significantly improved robustness, efficiency, and generalization. Extensive experiments using five real-world regression datasets that span computer vision, human-computer interaction, and healthcare verify that RNC achieves state-of-the-art performance, highlighting its intriguing properties including better data efficiency, robustness to spurious targets and data corruptions, and generalization to distribution shifts.
Kaiwen Zha, Jeany Son, Yuzhe Yang 0003, Dina Katabi
NeurIPS1
2022 Unsupervised Representation for Semantic Segmentation by Implicit Cycle-Attention Contrastive Learning
abstract
We study the unsupervised representation learning for the semantic segmentation task. Different from previous works that aim at providing unsupervised pre-trained backbones for segmentation models which need further supervised fine-tune, here, we focus on providing representation that is only trained by unsupervised methods. This means models need to directly generate pixel-level, linearly separable semantic results. We first explore and present two factors that have significant effects on segmentation under the contrastive learning framework: 1) the difficulty and diversity of the positive contrastive pairs, 2) the balance of global and local features. With the intention of optimizing these factors, we propose the cycle-attention contrastive learning (CACL). CACL makes use of semantic continuity of video frames, adopting unsupervised cycle-consistent attention mechanism to implicitly conduct contrastive learning with difficult, global-local-balanced positive pixel pairs. Compared with baseline model MoCo-v2 and other unsupervised methods, CACL demonstrates consistently superior performance on PASCAL VOC (+4.5 mIoU) and Cityscapes (+4.5 mIoU) datasets.
Bo Pang 0003, Yizhuo Li 0001, Gao Peng, Kaiwen Zha, Cewu Lu
AAAI6
2021 Delving into Deep Imbalanced Regression
abstract
Real-world data often exhibit imbalanced distributions, where certain target values have significantly fewer observations. Existing techniques for dealing with imbalanced data focus on targets with categorical indices, i.e., different classes. However, many tasks involve continuous targets, where hard boundaries between classes do not exist. We define Deep Imbalanced Regression (DIR) as learning from such imbalanced data with continuous targets, dealing with potential missing data for certain target values, and generalizing to the entire target range. Motivated by the intrinsic difference between categorical and continuous label space, we propose distribution smoothing for both labels and features, which explicitly acknowledges the effects of nearby targets, and calibrates both label and learned feature distributions. We curate and benchmark large-scale DIR datasets from common real-world tasks in computer vision, natural language processing, and healthcare domains. Extensive experiments verify the superior performance of our strategies. Our work fills the gap in benchmarks and techniques for practical imbalanced regression problems. Code and data are available at: https://github.com/YyzHarry/imbalanced-regression.
Yuzhe Yang 0003, Kaiwen Zha, Ying-Cong Chen, Hao Wang 0014, Dina Katabi
ICML2
2020 Further Understanding Videos through Adverbs: A New Video Task
abstract
Video understanding is a research hotspot of computer vision and significant progress has been made on video action recognition recently. However, the semantics information contained in actions is not rich enough to build powerful video understanding models. This paper first introduces a new video semantics: the Behavior Adverb (BA), which is a more expressive and difficult one covering subtle and inherent characteristics of human action behavior. To exhaustively decode this semantics, we construct the Videos with Action and Adverb Dataset (VAAD), which is a large-scale dataset with a semantically complete set of BAs. The dataset will be released to the public with this paper. We benchmark several representative video understanding methods (originally for action recognition) on BA and action recognition. The results show that BA recognition task is more challenging than conventional action recognition. Accordingly, we propose the BA Understanding Network (BAUN) to solve this problem and the experiments reveal that our BAUN is more suitable for BA recognition (11% better than I3D). Furthermore, we find these two semantics (action and BA) can propel each other forward to better performance: promoting action recognition results by 3.4% averagely on three standard action recognition datasets (UCF-101, HMDB-51, Kinetics).
Bo Pang 0003, Kaiwen Zha, Cewu Lu
AAAI2
2019 Deep RNN Framework for Visual Sequential Applications
abstract
Extracting temporal and representation features efficiently plays a pivotal role in understanding visual sequence information. To deal with this, we propose a new recurrent neural framework that can be stacked deep effectively. There are mainly two novel designs in our deep RNN framework: one is a new RNN module called Context Bridge Module (CBM) which splits the information flowing along the sequence (temporal direction) and along depth (spatial representation direction), making it easier to train when building deep by balancing these two directions; the other is the Overlap Coherence Training Scheme that reduces the training complexity for long visual sequential tasks on account of the limitation of computing resources. We provide empirical evidence to show that our deep RNN framework is easy to optimize and can gain accuracy from the increased depth on several visual sequence problems. On these tasks, we evaluate our deep RNN framework with 15 layers, 7× than conventional RNN networks, but it is still easy to train. Our deep framework achieves more than 11% relative improvements over shallow RNN models on Kinetics, UCF-101, and HMDB-51 for video classification. For auxiliary annotation, after replacing the shallow RNN part of Polygon-RNN with our 15-layer deep CBM, the performance improves by 14.7%. For video future prediction, our deep RNN improves the state-of-the-art shallow model's performance by 2.4% on PSNR and SSIM.
Bo Pang 0003, Kaiwen Zha, Cewu Lu
CVPR2