Yujin Han

dblp:317/6852 · DBLP profile ↗
← Back
7ranked-venue papers
4as first author
7since 2021 · last 2025
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 4 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Generative modeling · 34% Trustworthy machine learning · 28% Language models and text generation · 11%
Computer graphics and multimedia
1 paper
Image and video coding · 100%

Topics — the 21 heaviest of 23, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
diffusion model
2.532025
Can Diffusion Models Learn Hidden Inter-Feature Rules Behind Images? · ICML 2025
Masked Autoencoders Are Effective Tokenizers for Diffusion Models · ICML 2025
Slight Corruption in Pre-training Data Makes Better Diffusion Models · NeurIPS 2024
Machine learning › Trustworthy machine learning
interpretability
1.122025
Beyond Surface Structure: A Causal Assessment of LLMs' Comprehension ability · ICLR 2025
Can Diffusion Models Learn Hidden Inter-Feature Rules Behind Images? · ICML 2025
Machine learning › Trustworthy machine learning
robustness
1.022024
Improving Group Robustness on Spurious Correlation Requires Preciser Group Inference · ICML 2024
Slight Corruption in Pre-training Data Makes Better Diffusion Models · NeurIPS 2024
Machine learning › Deep learning architectures and training
autoencoder
0.912025
Masked Autoencoders Are Effective Tokenizers for Diffusion Models · ICML 2025
Machine learning › Generative modeling › autoregressive model
autoregressive image generation
0.912025
Parallelized Autoregressive Visual Generation · CVPR 2025
Machine learning › Generative modeling › score matching
denoising score matching
0.912025
Can Diffusion Models Learn Hidden Inter-Feature Rules Behind Images? · ICML 2025
Machine learning › Efficient and distributed learning
inference acceleration
0.912025
Parallelized Autoregressive Visual Generation · CVPR 2025
Natural language and speech › Language models and text generation
large language model evaluation
0.912025
Beyond Surface Structure: A Causal Assessment of LLMs' Comprehension ability · ICLR 2025
Machine learning › Generative modeling › diffusion model
latent diffusion model
0.912025
Masked Autoencoders Are Effective Tokenizers for Diffusion Models · ICML 2025
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning
masked autoencoder
0.912025
Masked Autoencoders Are Effective Tokenizers for Diffusion Models · ICML 2025
Machine learning › Probabilistic and Bayesian machine learning › causal inference › causal effect estimation
mediation analysis
0.912025
Beyond Surface Structure: A Causal Assessment of LLMs' Comprehension ability · ICLR 2025
Natural language and speech › Language models and text generation › decoding › decoding strategy
parallel decoding
0.912025
Parallelized Autoregressive Visual Generation · CVPR 2025
Machine learning › Trustworthy machine learning
fairness
0.812024
Improving Group Robustness on Spurious Correlation Requires Preciser Group Inference · ICML 2024
Machine learning › Trustworthy machine learning › fairness
group robustness
0.812024
Improving Group Robustness on Spurious Correlation Requires Preciser Group Inference · ICML 2024
Machine learning › Trustworthy machine learning › robustness
spurious correlation
0.812024
Improving Group Robustness on Spurious Correlation Requires Preciser Group Inference · ICML 2024
Machine learning › Deep learning architectures and training › data-centric deep learning
training data quality
0.812024
Slight Corruption in Pre-training Data Makes Better Diffusion Models · NeurIPS 2024
Knowledge, reasoning and agents › Knowledge representation and reasoning
commonsense reasoning
0.312025
Beyond Surface Structure: A Causal Assessment of LLMs' Comprehension ability · ICLR 2025
Machine learning › Generative modeling
image generation
0.312025
Parallelized Autoregressive Visual Generation · CVPR 2025
Machine learning › Generative modeling
video generation
0.312025
Parallelized Autoregressive Visual Generation · CVPR 2025
Image and video coding
image tokenization
0.312025
Masked Autoencoders Are Effective Tokenizers for Diffusion Models · ICML 2025
Machine learning › Trustworthy machine learning › out-of-distribution generalization
invariant learning
0.212024
Improving Group Robustness on Spurious Correlation Requires Preciser Group Inference · ICML 2024

Methods — techniques the papers use, named apart from their topics

variational autoencoder · 1.7masked modeling · 1.7token dependency analysis · 0.9synthetic tasks · 0.9classifier guidance · 0.9causal mediation analysis · 0.9causal effect estimation · 0.9autoregressive modeling · 0.9spurious attribute classifier · 0.8group inference · 0.8
YearPublicationVenuePosition
2025 Parallelized Autoregressive Visual Generation
abstract
Autoregressive models have emerged as a powerful approach for visual generation but suffer from slow inference speed due to their sequential token-by-token prediction process. In this paper, we propose a simple yet effective approach for parallelized autoregressive visual generation that improves generation efficiency while preserving the advantages of autoregressive modeling. Our key insight is that parallel generation depends on visual token dependencies—tokens with weak dependencies can be generated in parallel, while strongly dependent adjacent tokens are difficult to generate together, as their independent sampling may lead to inconsistencies. Based on this observation, we develop a parallel generation strategy that generates distant tokens with weak dependencies in parallel while maintaining sequential generation for strongly dependent local tokens. Our approach can be seamlessly integrated into standard autoregressive models without modifying the architecture or tokenizer. Experiments on ImageNet and UCF-101 demonstrate that our method achieves a 3.6× speedup with comparable quality and up to 9.5× speedup with minimal quality degradation across both image and video generation tasks. We hope this work will inspire future research in efficient visual generation and unified autoregressive modeling. Project page: https://yuqingwang1029.github.io/PAR-project.
Shuhuai Ren, Zhijie Lin 0001, Yujin Han, Haoyuan Guo, Zhenheng Yang, Difan Zou, Jiashi Feng, Xihui Liu
CVPR4
2025 Beyond Surface Structure: A Causal Assessment of LLMs' Comprehension ability
abstract
Large language models (LLMs) have shown remarkable capability in natural language tasks, yet debate persists on whether they truly comprehend deep structure (i.e., core semantics) or merely rely on surface structure (e.g., presentation format). Prior studies observe that LLMs' performance declines when intervening on surface structure, arguing their success relies on surface structure recognition. However, surface structure sensitivity does not prevent deep structure comprehension. Rigorously evaluating LLMs' capability requires analyzing both, yet deep structure is often overlooked. To this end, we assess LLMs' comprehension ability using causal mediation analysis, aiming to fully discover the capability of using both deep and surface structures. Specifically, we formulate the comprehension of deep structure as direct causal effect (DCE) and that of surface structure as indirect causal effect (ICE), respectively. To address the non-estimability of original DCE and ICE --- stemming from the infeasibility of isolating mutual influences of deep and surface structures, we develop the corresponding quantifiable surrogates, including approximated DCE (ADCE) and approximated ICE (AICE). We further apply the ADCE to evaluate a series of mainstream LLMs (and the one with random weights), showing that most of them exhibit deep structure comprehension ability, which grows along with the prediction accuracy. Comparing ADCE and AICE demonstrates closed-source LLMs (e.g., GPT) rely more on deep structure, while open-source LLMs (e.g., Llama) are more surface-sensitive, which decreases with model scale. Theoretically, ADCE is a bidirectional evaluation, which measures both the sufficiency and necessity of deep structure changes in causing output variations, thus offering a more comprehensive assessment than accuracy, a common evaluation in LLMs. Our work provides new insights into LLMs' deep structure comprehension and offers novel methods for LLMs evaluation. The code for our project is available at [ADCE Project](https://github.com/OpenCausaLab/ADCE).
Yujin Han, Difan Zou, Chaochao Lu
ICLR1
2025 Masked Autoencoders Are Effective Tokenizers for Diffusion Models
abstract
Recent advances in latent diffusion models have demonstrated their effectiveness for high-resolution image synthesis. However, the properties of the latent space from tokenizer for better learning and generation of diffusion models remain under-explored. Theoretically and empirically, we find that improved generation quality is closely tied to the latent distributions with better structure, such as the ones with fewer Gaussian Mixture modes and more discriminative features. Motivated by these insights, we propose MAETok, an autoencoder (AE) leveraging mask modeling to learn semantically rich latent space while maintaining reconstruction fidelity. Extensive experiments validate our analysis, demonstrating that the variational form of autoencoders is not necessary, and a discriminative latent space from AE alone enables state-of-the-art performance on ImageNet generation using only 128 tokens. MAETok achieves significant practical improvements, enabling a gFID of 1.69 with 76× faster training and 31× higher inference throughput for 512×512 generation. Our findings show that the structure of the latent space, rather than variational constraints, is crucial for effective diffusion models. Code and trained models will be released.
Hao Chen 0102, Yujin Han, Fangyi Chen, Xiang Li 0106, Yidong Wang 0003, Jindong Wang 0001, Ze Wang 0008, Zicheng Liu 0001, Difan Zou, Bhiksha Raj
ICML2
2025 Can Diffusion Models Learn Hidden Inter-Feature Rules Behind Images?
abstract
Despite the remarkable success of diffusion models (DMs) in data generation, they exhibit specific failure cases with unsatisfactory outputs. We focus on one such limitation: the ability of DMs to learn hidden rules between image features. Specifically, for image data with dependent features ($\mathbf{x}$) and ($\mathbf{y}$) (e.g., the height of the sun ($\mathbf{x}$) and the length of the shadow ($\mathbf{y}$)), we investigate whether DMs can accurately capture the inter-feature rule ($p(\mathbf{y}|\mathbf{x})$). Empirical evaluations on mainstream DMs (e.g., Stable Diffusion 3.5) reveal consistent failures, such as inconsistent lighting-shadow relationships and mismatched object-mirror reflections. Inspired by these findings, we design four synthetic tasks with strongly correlated features to assess DMs’ rule-learning abilities. Extensive experiments show that while DMs can identify coarse-grained rules, they struggle with fine-grained ones. Our theoretical analysis demonstrates that DMs trained via denoising score matching (DSM) exhibit constant errors in learning hidden rules, as the DSM objective is not compatible with rule conformity. To mitigate this, we introduce a common technique - incorporating additional classifier guidance during sampling, which achieves (limited) improvements. Our analysis reveals that the subtle signals of fine-grained rules are challenging for the classifier to capture, providing insights for future exploration.
Yujin Han, Andi Han, Chaochao Lu, Difan Zou
ICML1
2024 Conformalized Semi-supervised Random Forest for Classification and Abnormality Detection
abstract
The Random Forests classifier, a widely utilized off-the-shelf classification tool, assumes training and test samples come from the same distribution as other standard classifiers. However, in safety-critical scenarios like medical diagnosis and network attack detection, discrepancies between the training and test sets, including the potential presence of novel outlier samples not appearing during training, can pose significant challenges. To address this problem, we introduce the Conformalized Semi-Supervised Random Forest (CSForest), which couples the conformalization technique Jackknife+aB with semi-supervised tree ensembles to construct a set-valued prediction $C(x)$. Instead of optimizing over the training distribution, CSForest employs unlabeled test samples to enhance accuracy and flag unseen outliers by generating an empty set. Theoretically, we establish CSForest to cover true labels for previously observed inlier classes under arbitrarily label-shift in the test data. We compare CSForest with state-of-the-art methods using synthetic examples and various real-world datasets, under different types of distribution changes in the test domain. Our results highlight CSForest’s effective prediction of inliers and its ability to detect outlier samples unique to the test data. In addition, CSForest shows persistently good performance as the sizes of the training and test sets vary. Codes of CSForest are available at https://github.com/yujinhan98/CSForest.
Yujin Han, Mingwenchan Xu, Leying Guan
AISTATS1
2024 Improving Group Robustness on Spurious Correlation Requires Preciser Group Inference
abstract
Standard empirical risk minimization (ERM) models may prioritize learning spurious correlations between spurious features and true labels, leading to poor accuracy on groups where these correlations do not hold. Mitigating this issue often requires expensive spurious attribute (group) labels or relies on trained ERM models to infer group labels when group information is unavailable. However, the significant performance gap in worst-group accuracy between using pseudo group labels and using oracle group labels inspires us to consider further improving group robustness through preciser group inference. Therefore, we propose GIC, a novel method that accurately infers group labels, resulting in improved worst-group performance. GIC trains a spurious attribute classifier based on two key properties of spurious correlations: (1) high correlation between spurious attributes and true labels, and (2) variability in this correlation between datasets with different group distributions. Empirical studies on multiple datasets demonstrate the effectiveness of GIC in inferring group labels, and combining GIC with various downstream invariant learning methods improves worst-group accuracy, showcasing its powerful flexibility. Additionally, through analyzing the misclassifications in GIC, we identify an interesting phenomenon called semantic consistency, which may contribute to better decoupling the association between spurious attributes and labels, thereby mitigating spurious correlation. The code for GIC is available at https://github.com/yujinhanml/GIC9.
Yujin Han, Difan Zou
ICML1
2024 Slight Corruption in Pre-training Data Makes Better Diffusion Models
abstract
Diffusion models (DMs) have shown remarkable capabilities in generating realistic high-quality images, audios, and videos. They benefit significantly from extensive pre-training on large-scale datasets, including web-crawled data with paired data and conditions, such as image-text and image-class pairs. Despite rigorous filtering, these pre-training datasets often inevitably contain corrupted pairs where conditions do not accurately describe the data. This paper presents the first comprehensive study on the impact of such corruption in pre-training data of DMs. We synthetically corrupt ImageNet-1K and CC3M to pre-train and evaluate over $50$ conditional DMs. Our empirical findings reveal that various types of slight corruption in pre-training can significantly enhance the quality, diversity, and fidelity of the generated images across different DMs, both during pre-training and downstream adaptation stages. Theoretically, we consider a Gaussian mixture model and prove that slight corruption in the condition leads to higher entropy and a reduced 2-Wasserstein distance to the ground truth of the data distribution generated by the corruptly trained DMs. Inspired by our analysis, we propose a simple method to improve the training of DMs on practical datasets by adding condition embedding perturbations (CEP). CEP significantly improves the performance of various DMs in both pre-training and downstream tasks. We hope that our study provides new insights into understanding the data and pre-training processes of DMs.
Hao Chen 0102, Yujin Han, Diganta Misra, Xiang Li 0106, Kai Hu 0010, Difan Zou, Masashi Sugiyama, Jindong Wang 0001, Bhiksha Raj
NeurIPS2