EDBT 2026 Demo / reviewers in the wild / expert
Huayu Chen
dblp:259/3113
· DBLP profile ↗
19ranked-venue papers
9as first author
18since 2021 · last 2026
0000-0003-0563-4064ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 9 first-author · 16 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SS-NeRF: Shine-sphere rendering for neural radiance fields
Tongyuan Huang, Chunyuan Liu, Huayu Chen, Zipeng Wu |
Pattern Recognit. | 5 |
| 2026 | Quantifying Emotional Patterns for EEG-Based Emotion Recognition: An Interpretable Study on EEG Individual DifferencesabstractElectroencephalogram (EEG) individual differences are a critical factor influencing EEG-based emotion recognition, yet they have not been thoroughly investigated, hindering the development of affective Brain-Computer Interfaces (aBCI). Facing the lack of EEG information decoding research, we conducted an interpretable study on EEG individual differences using five datasets (SEED, SEED-IV, SEED-V, RCLS, and MPED). We analyzed the impact of different EEG information (individual, session, emotion, and trial) through sample space visualization, aggregation phenomena quantification, and energy pattern analysis. By examining emotional difference feature distribution patterns, we identified the Cross-Session Consistency of Individual Emotional Patterns (CCIEP) and the Individual Emotional Pattern Difference (IEPD). These characteristics are the main factors impacting emotion recognition stability. To quantify emotional patterns, we proposed the Correction T-test (CT) weight extraction method. Leveraging individual emotional pattern and trial information, we developed the Weight-based Channel-model Matrix Framework (WCMF) to address limitations of traditional modeling approaches caused by IEPD. Finally, WCMF was validated on cross-dataset tasks through two practical scenario experiments. The results demonstrated that WCMF achieves more stable and superior performance compared to traditional methods. This study provides a deeper understanding of EEG individual differences and offers a robust framework to advance aBCI systems. Huayu Chen, Xiaowei Li 0005, Xuexiao Shao, Huanhuan He, Jing Zhu 0003, Bin Hu 0001 |
IEEE Trans. Affect. Comput. | 1 |
| 2025 | Assessing Stability and Consistency in EEG-Based Functional Brain Network Construction MethodsabstractFunctional brain network construction methods are crucial for understanding neurological and psychiatric disorders, yet their stability and consistency across different parameters remain poorly understood. This study systematically evaluates the robustness of two widely used EEG-based functional connectivity methods: imaginary part of coherency combined with cluster span threshold (ICoh+CST) and phase lag index combined with cluster span threshold (PLI+CST). We analyzed EEG data from two independent datasets across multiple epoch lengths (120 seconds) at both scalp and source levels. Network metrics including global efficiency and local efficiency were computed across four frequency bands (delta, theta, alpha, beta). Results demonstrate that the ICoh+CST method achieves stability at significantly shorter epoch lengths compared to PLI+CST (10s vs. 14s for scalp-level analysis, 6s vs. 12s for source-level analysis). Linear regression analysis revealed that ICoh+CST exhibits lower correlation between network metric variability and epoch length changes, indicating superior robustness. Importantly, the ICoh+CST method yielded more consistent findings across datasets, with major depressive disorder patients showing significantly increased global efficiency in the theta band at the source level in both datasets. Our results have important implications for standardizing neuroimaging protocols and improving the reproducibility of functional connectivity studies in clinical neuroscience research. Fudi Qin, Huayu Chen, Huiwen Guo |
BIBM | 2 |
| 2025 | Toward Guidance-Free AR Visual Generation via Condition Contrastive AlignmentabstractClassifier-Free Guidance (CFG) is a critical technique for enhancing the sample quality of visual generative models. However, in autoregressive (AR) multi-modal generation, CFG introduces design inconsistencies between language and visual content, contradicting the design philosophy of unifying different modalities for visual AR. Motivated by language model alignment methods, we propose Condition Contrastive Alignment (CCA) to facilitate guidance-free AR visual generation. Unlike guidance methods that alter the sampling process to achieve the ideal sampling distribution, CCA directly fine-tunes pretrained models to fit the same distribution target. Experimental results show that CCA can significantly enhance the guidance-free performance of all tested models with just one epoch of fine-tuning (1% of pretraining epochs) on the pretraining dataset. This largely removes the need for guided sampling in AR visual generation and cuts the sampling cost by half. Moreover, by adjusting training parameters, CCA can achieve trade-offs between sample diversity and fidelity similar to CFG. This experimentally confirms the strong theoretical connection between language-targeted alignment and visual-targeted guidance methods, unifying two previously independent research fields. Huayu Chen, Hang Su 0006, Peize Sun, Jun Zhu 0001 |
ICLR | 1 |
| 2025 | RDT-1B: a Diffusion Foundation Model for Bimanual ManipulationabstractBimanual manipulation is essential in robotics, yet developing foundation models is extremely challenging due to the inherent complexity of coordinating two robot arms (leading to multi-modal action distributions) and the scarcity of training data. In this paper, we present the Robotics Diffusion Transformer (RDT), a pioneering diffusion foundation model for bimanual manipulation. RDT builds on diffusion models to effectively represent multi-modality, with innovative designs of a scalable Transformer to deal with the heterogeneity of multi-modal inputs and to capture the nonlinearity and high frequency of robotic data. To address data scarcity, we further introduce a Physically Interpretable Unified Action Space, which can unify the action representations of various robots while preserving the physical meanings of original actions, facilitating learning transferrable physical knowledge. With these designs, we managed to pre-train RDT on the largest collection of multi-robot datasets to date and scaled it up to $1.2$B parameters, which is the largest diffusion-based foundation model for robotic manipulation. We finally fine-tuned RDT on a self-created multi-task bimanual dataset with over $6$K+ episodes to refine its manipulation capabilities. Experiments on real robots demonstrate that RDT significantly outperforms existing methods. It exhibits zero-shot generalization to unseen objects and scenes, understands and follows language instructions, learns new skills with just 1$\sim$5 demonstrations, and effectively handles complex, dexterous tasks. We refer to https://rdt-robotics.github.io/rdt-robotics/ for the code and videos. Songming Liu, Lingxuan Wu, Bangguo Li, Hengkai Tan, Huayu Chen, Hang Su 0006, Jun Zhu 0001 |
ICLR | 5 |
| 2025 | Visual Generation Without GuidanceabstractClassifier-Free Guidance (CFG) has been a default technique in various visual generative models, yet it requires inference from both conditional and unconditional models during sampling. We propose to build visual models that are free from guided sampling. The resulting algorithm, Guidance-Free Training (GFT), matches the performance of CFG while reducing sampling to a single model, halving the computational cost. Unlike previous distillation-based approaches that rely on pretrained CFG networks, GFT enables training directly from scratch. GFT is simple to implement. It retains the same maximum likelihood objective as CFG and differs mainly in the parameterization of conditional models. Implementing GFT requires only minimal modifications to existing codebases, as most design choices and hyperparameters are directly inherited from CFG. Our extensive experiments across five distinct visual models demonstrate the effectiveness and versatility of GFT. Across domains of diffusion, autoregressive, and masked-prediction modeling, GFT consistently achieves comparable or even lower FID scores, with similar diversity-fidelity trade-offs compared with CFG baselines, all while being guidance-free. Huayu Chen, Kaiwen Zheng 0003, Jianfei Chen 0001, Hang Su 0006, Jun Zhu 0001 |
ICML | 1 |
| 2025 | Free Process Rewards without Process LabelsabstractDifferent from its counterpart outcome reward models (ORMs), which evaluate the entire responses, a process reward model (PRM) scores a reasoning trajectory step by step, providing denser and more fine-grained rewards. However, training a PRM requires labels annotated at every intermediate step, presenting significant challenges for both manual and automatic data collection. This paper aims to address this challenge. Both theoretically and empirically, we show that an implicit PRM can be obtained at no additional cost, by simply training an ORM on the cheaper response-level labels. The only assumption is to parameterize the outcome reward as the log-likelihood ratios of the policy and reference models r$\phi$(y) = $\beta$ log $\pi$$\phi$(y) $\pi$ref(y) , which can be optimized regardless of the specific choice of loss objectives. In experiments, we instantiate our implicit PRMs with various objectives and evaluate their performance on MATH. We show that our implicit PRM outperforms a strong MCTS-based baseline á la Math-Shepherd (Wang et al., 2023) using less than 1/38 of the training data. Its performance can be further improved with majority voting. We further find that scaling up instructions and responses benefits our implicit PRM, and the latter brings a larger gain. Particularly, we find that our implicit PRM, when instantiated with the cross-entropy (CE) loss, is more data-efficient and can keep improving generation models even when trained with only one response per instruction, the setup that suffers from extreme data scarcity and imbalance. Further, instructions should be relevant to downstream tasks while the diversity of responses does not bring gains. Surprisingly, training on extra Math-Shepherd step labels brings no further improvements to our implicit PRM trained on only outcome data. We hope that our work will encourage a rethinking of PRM training approaches and contribute to making training PRMs more accessible. Lifan Yuan, Wendi Li, Huayu Chen, Ganqu Cui, Ning Ding 0002, Bowen Zhou 0002, Zhiyuan Liu 0001, Hao Peng 0001 |
ICML | 3 |
| 2025 | Direct Discriminative Optimization: Your Likelihood-Based Visual Generative Model is Secretly a GAN DiscriminatorabstractWhile likelihood-based generative models, particularly diffusion and autoregressive models, have achieved remarkable fidelity in visual generation, the maximum likelihood estimation (MLE) objective, which minimizes the forward KL divergence, inherently suffers from a mode-covering tendency that limits the generation quality under limited model capacity. In this work, we propose Direct Discriminative Optimization (DDO) as a unified framework that integrates likelihood-based generative training and GAN-type discrimination to bypass this fundamental constraint by exploiting reverse KL and self-generated negative signals. Our key insight is to parameterize a discriminator implicitly using the likelihood ratio between a learnable target model and a fixed reference model, drawing parallels with the philosophy of Direct Preference Optimization (DPO). Unlike GANs, this parameterization eliminates the need for joint training of generator and discriminator networks, allowing for direct, efficient, and effective finetuning of a well-trained model to its full potential beyond the limits of MLE. DDO can be performed iteratively in a self-play manner for progressive model refinement, with each round requiring less than 1\% of pretraining epochs. Our experiments demonstrate the effectiveness of DDO by significantly advancing the previous SOTA diffusion model EDM, reducing FID scores from 1.79/1.58/1.96 to new records of 1.30/0.97/1.26 on CIFAR-10/ImageNet-64/ImageNet 512$\times$512 datasets without any guidance mechanisms, and by consistently improving both guidance-free and CFG-enhanced FIDs of visual autoregressive models on ImageNet 256$\times$256. Kaiwen Zheng 0003, Yongxin Chen 0002, Huayu Chen, Guande He, Ming-Yu Liu 0001, Jun Zhu 0001, Qinsheng Zhang |
ICML | 3 |
| 2025 | VAE-CapsNet: A common emotion information extractor for cross-subject emotion recognition
Huayu Chen, Huanhuan He, Jing Zhu 0003, Xiaowei Li 0005, Bin Hu 0001 |
Knowl. Based Syst. | 1 |
| 2024 | Score Regularized Policy Optimization through Diffusion BehaviorabstractRecent developments in offline reinforcement learning have uncovered the immense potential of diffusion modeling, which excels at representing heterogeneous behavior policies. However, sampling from diffusion policies is considerably slow because it necessitates tens to hundreds of iterative inference steps for one action. To address this issue, we propose to extract an efficient deterministic inference policy from critic models and pretrained diffusion behavior models, leveraging the latter to directly regularize the policy gradient with the behavior
distribution’s score function during optimization. Our method enjoys powerful generative capabilities of diffusion modeling while completely circumventing the computationally intensive and time-consuming diffusion sampling scheme, both during training and evaluation. Extensive results on D4RL tasks show that our method boosts action sampling speed by more than 25 times compared with various leading diffusion-based methods in locomotion tasks, while still maintaining state-of-the-art performance. Huayu Chen, Cheng Lu 0011, Hang Su 0006, Jun Zhu 0001 |
ICLR | 1 |
| 2024 | Noise Contrastive Alignment of Language Models with Explicit RewardsabstractUser intentions are typically formalized as evaluation rewards to be maximized when fine-tuning language models (LMs). Existing alignment methods, such as Direct Preference Optimization (DPO), are mainly tailored for pairwise preference data where rewards are implicitly defined rather than explicitly given. In this paper, we introduce a general framework for LM alignment, leveraging Noise Contrastive Estimation (NCE) to bridge the gap in handling reward datasets explicitly annotated with scalar evaluations. Our framework comprises two parallel algorithms, NCA and InfoNCA, both enabling the direct extraction of an LM policy from reward data as well as preference data. Notably, we show that the DPO loss is a special case of our proposed InfoNCA objective under pairwise preference settings, thereby integrating and extending current alignment theories. By comparing NCA and InfoNCA, we demonstrate that the well-observed decreasing-likelihood trend of DPO/InfoNCA is caused by their focus on adjusting relative likelihood across different responses.
In contrast, NCA optimizes the absolute likelihood for each response, thereby effectively preventing the chosen likelihood from decreasing. We evaluate our methods in both reward and preference settings with Mistral-8$\times$7B and 7B models. Experiments suggest that InfoNCA/NCA surpasses various preference baselines when reward datasets are available. We also find NCA significantly outperforms DPO in complex reasoning tasks like math and coding. Huayu Chen, Guande He, Lifan Yuan, Ganqu Cui, Hang Su 0006, Jun Zhu 0001 |
NeurIPS | 1 |
| 2024 | Aligning Diffusion Behaviors with Q-functions for Efficient Continuous ControlabstractDrawing upon recent advances in language model alignment, we formulate offline Reinforcement Learning as a two-stage optimization problem: First pretraining expressive generative policies on reward-free behavior datasets, then finetuning these policies to align with task-specific annotations like Q-values. This strategy allows us to leverage abundant and diverse behavior data to enhance generalization and enable rapid adaptation to downstream tasks using minimal annotations. In particular, we introduce Efficient Diffusion Alignment (EDA) for solving continuous control problems. EDA utilizes diffusion models for behavior modeling. However, unlike previous approaches, we represent diffusion policies as the derivative of a scalar neural network with respect to action inputs. This representation is critical because it enables direct density calculation for diffusion models, making them compatible with existing LLM alignment theories. During policy fine-tuning, we extend preference-based alignment methods like Direct Preference Optimization (DPO) to align diffusion behaviors with continuous Q-functions. Our evaluation on the D4RL benchmark shows that EDA exceeds all baseline methods in overall performance. Notably, EDA maintains about 95\% of performance and still outperforms several baselines given only 1\% of Q-labelled data during fine-tuning. Huayu Chen, Kaiwen Zheng 0003, Hang Su 0006, Jun Zhu 0001 |
NeurIPS | 1 |
| 2024 | C-GAIL: Stabilizing Generative Adversarial Imitation Learning with Control TheoryabstractGenerative Adversarial Imitation Learning (GAIL) provides a promising approach to training a generative policy to imitate a demonstrator. It uses on-policy Reinforcement Learning (RL) to optimize a reward signal derived from an adversarial discriminator. However, optimizing GAIL is difficult in practise, with the training loss oscillating during training, slowing convergence. This optimization instability can prevent GAIL from finding a good policy, harming its final performance. In this paper, we study GAIL’s optimization from a control-theoretic perspective. We show that GAIL cannot converge to the desired equilibrium. In response, we analyze the training dynamics of GAIL in function space and design a novel controller that not only pushes GAIL to the desired equilibrium but also achieves asymptotic stability in a simplified “one-step” setting. Going from theory to practice, we propose Controlled-GAIL (C-GAIL), which adds a differentiable regularization term on the GAIL objective to stabilize training. Empirically, the C-GAIL regularizer improves the training of various existing GAIL methods, including the popular GAIL-DAC, by speeding up the convergence, reducing the range of oscillation, and matching the expert distribution more closely. Tianjiao Luo, Tim Pearce, Huayu Chen, Jianfei Chen 0001, Jun Zhu 0001 |
NeurIPS | 3 |
| 2023 | Offline Reinforcement Learning via High-Fidelity Generative Behavior Modeling
Huayu Chen, Cheng Lu 0011, Chengyang Ying, Hang Su 0006, Jun Zhu 0001 |
ICLR | 1 |
| 2023 | Contrastive Energy Prediction for Exact Energy-Guided Diffusion Sampling in Offline Reinforcement LearningabstractGuided sampling is a vital approach for applying diffusion models in real-world tasks that embeds human-defined guidance during the sampling procedure. This paper considers a general setting where the guidance is defined by an (unnormalized) energy function. The main challenge for this setting is that the intermediate guidance during the diffusion sampling procedure, which is jointly defined by the sampling distribution and the energy function, is unknown and is hard to estimate. To address this challenge, we propose an exact formulation of the intermediate guidance as well as a novel training objective named contrastive energy prediction (CEP) to learn the exact guidance. Our method is guaranteed to converge to the exact guidance under unlimited model capacity and data samples, while previous methods can not. We demonstrate the effectiveness of our method by applying it to offline reinforcement learning (RL). Extensive experiments on D4RL benchmarks demonstrate that our method outperforms existing state-of-the-art algorithms. We also provide some examples of applying CEP for image synthesis to demonstrate the scalability of CEP on high-dimensional data. Cheng Lu 0011, Huayu Chen, Jianfei Chen 0001, Hang Su 0006, Chongxuan Li, Jun Zhu 0001 |
ICML | 2 |
| 2023 | Personal-Zscore: Eliminating Individual Difference for EEG-Based Cross-Subject Emotion RecognitionabstractIt was observed that accuracy of the Subject-Dependent emotion recognition model was much higher than that of the Subject-Independent model in the field of electroencephalogram (EEG) based affective computing. This phenomenon is mainly caused by the individual difference of EEG, which is the key issue to be solved for the application of emotion recognition. In this work, 14 subjects from the SEED were selected for individual difference analysis. Through individual aggregation features evaluation, sample space visualization, and correlation analysis, we proposed four quantification indicators to analyze individual difference phenomenon. Finally, we presented the Personal-Zscore (PZ) feature processing method, and it was found that the data set processed with PZ method could represent emotion better than the original data set, and the conventional model with the PZ method was more robust. The accuracies of emotion recognition models trained with PZ processing have been improved to some extent, which showed that the PZ method could effectively eliminate the individual aggregation of feature space and improve the emotional representation ability of data sets. Hence, our findings may provide a new insight into the foundation for universal implementation of EEG-based application, and the Personal-Zscore feature processing method is of great significance for the development of effective emotion recognition system. Huayu Chen, Jianxiu Li, Ruilan Yu, Xiaowei Li 0005, Bin Hu 0001 |
IEEE Trans. Affect. Comput. | 1 |
| 2023 | Clustering-Fusion Feature Selection Method in Identifying Major Depressive Disorder Based on Resting State EEG SignalsabstractDepression is a heterogeneous syndrome with certain individual differences among subjects. Exploring a feature selection method that can effectively mine the commonness intra-groups and the differences inter-groups in depression recognition is therefore of great significance. This study proposed a new clustering-fusion feature selection method. Hierarchical clustering (HC) algorithm was used to capture the heterogeneity distribution of subjects. Average and similarity network fusion (SNF) algorithms were adopted to characterize the brain network atlas of different populations. Differences analysis was also utilized to obtain the features with discriminant performance. Experiments showed that compared with traditional feature selection methods, HCSNF method yielded the optimal classification results of depression recognition in both sensor and source layers of electroencephalography (EEG) data. Especially in the beta band of EEG data at sensor layer, the classification performance was improved by more than 6%. Moreover, the long-distance connections between parietal-occipital lobe and other brain regions not only have high discriminative power, but also significantly correlate with depressive symptoms, indicating the important role of these features in depression recognition. Therefore, this study may provide methodological guidance for the discovery of reproducible electrophysiological biomarkers and new insights into common neuropathological mechanisms of heterogeneous depression diseases. Huayu Chen, Chang Yan, Qunxi Dong, Xuexiao Shao, Xiaowei Li 0005, Bin Hu 0001 |
IEEE J. Biomed. Health Informatics | 2 |
| 2022 | Tianshou: A Highly Modularized Deep Reinforcement Learning LibraryabstractIn this paper, we present Tianshou, a highly modularized Python library for deep reinforcement learning (DRL) that uses PyTorch as its backend. Tianshou intends to be research-friendly by providing a flexible and reliable infrastructure of DRL algorithms. It supports online and offline training with more than 20 classic algorithms through a unified interface. To facilitate related research and prove Tianshou's reliability, we have released Tianshou's benchmark of MuJoCo environments, covering eight classic algorithms with state-of-the-art performance. We open-sourced Tianshou at https://github.com/thu-ml/tianshou/. Jiayi Weng, Huayu Chen, Kaichao You, Alexis Duburcq, Hang Su 0006, Jun Zhu 0001 |
J. Mach. Learn. Res. | 2 |
| 2020 | EEG Based Depression Recognition by Combining Functional Brain Network and Traditional BiomarkersabstractThis Electroencephalography (EEG)-based research is to explore the effective biomarkers for depression recognition. Resting-state EEG data were collected from 24 major depressive patients (MDD) and 29 normal controls using 128-electrode geodesic sensor net. To better identify depression, we extracted multi-type of EEG features including linear features (L), nonlinear features (NL), functional connectivity features phase lagging index (PLI) and network measures (NM) to comprehensively characterize the EEG signals in patients with MDD. And machine learning algorithms and statistical analysis were used to evaluate the EEG features. Combined multi-types features (All: L+ NL + PLI + NM) outperformed single-type features for classifying depression. Analyzing the optimal features set we found that compared to other type features, PLI occupied the largest proportion of which functional connections in intra-hemisphere were much more than that of in inter-hemisphere. In addition, when using PLI features and All features, high frequency bands (alpha, beta) could achieve obviously higher classification accuracy than low frequency bands (delta, theta). Parietal-occipital lobe in the high frequency bands had great effect in depression identification. In conclusion, combined multi-types EEG features along with a robust classifier can better distinguish depressive patients from normal controls. And intra-hemispheric functional connections might be an effective biomarker to detect depression. Hence, this paper may provide objective and potential electrophysiological characteristics in depression recognition. Huayu Chen, Xuexiao Shao, Liangliang Liu 0002, Xiaowei Li 0005, Bin Hu 0001 |
BIBM | 2 |