EDBT 2026 Demo / reviewers in the wild / expert
Chuan Wen
dblp:239/8286
· DBLP profile ↗
16ranked-venue papers
7as first author
14since 2021 · last 2025
0009-0007-8594-7314ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 6 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Data Scaling Laws in Imitation Learning for Robotic ManipulationabstractData scaling has revolutionized fields like natural language processing and computer vision, providing models with remarkable generalization capabilities. In this paper, we investigate whether similar data scaling laws exist in robotics, particularly in robotic manipulation, and whether appropriate data scaling can yield single-task robot policies that can be deployed zero-shot for any object within the same category in any environment. To this end, we conduct a comprehensive empirical study on data scaling in imitation learning. By collecting data across numerous environments and objects, we study how a policy’s generalization performance changes with the number of training environments, objects, and demonstrations. Throughout our research, we collect over 40,000 demonstrations and execute more than 15,000 real-world robot rollouts under a rigorous evaluation protocol. Our findings reveal several intriguing results: the generalization performance of the policy follows a roughly power-law relationship with the number of environments and objects. The diversity of environments and objects is far more important than the absolute number of demonstrations; once the number of demonstrations per environment or object reaches a certain threshold, additional demonstrations have minimal effect. Based on these insights, we propose an efficient data collection strategy. With four data collectors working for one afternoon, we collect sufficient data to enable the policies for two tasks to achieve approximately 90\% success rates in novel environments with unseen objects. Fanqi Lin, Yingdong Hu, Pingyue Sheng, Chuan Wen, Jiacheng You, Yang Gao 0029 |
ICLR | 4 |
| 2025 | Individualized speech enhancement for hearing-impaired listenersabstractDespite significant progress in speech enhancement (SE), improving speech intelligibility and perceptual quality in noisy environments for hearing-impaired individuals remains challenging. This paper presents an individualized speech-enhancement (ISE) framework that integrates noise reduction (NR) and sound amplification for hearing loss compensation within a unified system. Our fully differentiable, closed-loop design incorporates biophysically realistic auditory models. The framework features two pathways: one simulating the auditory response of a normal-hearing (NH) system to denoised speech and the other modeling a hearing-impaired (HI) system's response to noisy speech. The ISE model is trained by minimizing the difference between NH and HI auditory responses. Experimental results show that the ISE model enhances speech intelligibility and perceptual quality in noisy conditions for HI listeners, offering a promising foundation for advancing personalized noise reduction strategies. Chuan Wen, Sarah Verhulst |
INTERSPEECH | 1 |
| 2025 | Accelerating Feature Conformal Prediction via Taylor ApproximationabstractConformal prediction is widely adopted in uncertainty quantification, due to its post-hoc, distribution-free, and model-agnostic properties.
In the realm of modern deep learning, researchers have proposed Feature Conformal Prediction (FCP), which deploys conformal prediction in a feature space, yielding reduced band lengths.
However, the practical utility of FCP is limited due to the time-consuming non-linear operations required to transform confidence bands from feature space to output space.
In this paper, we present Fast Feature Conformal Prediction (FFCP), a method that accelerates FCP by leveraging a first-order Taylor expansion to approximate these non-linear operations.
The proposed FFCP introduces a novel non-conformity score that is both effective and efficient for real-world applications.
Empirical validations showcase that FFCP performs comparably with FCP (both outperforming the vanilla version) while achieving a significant reduction in computational time by approximately 50x in both regression and classification tasks. Zihao Tang 0001, Chuan Wen, Jiaye Teng |
NeurIPS | 3 |
| 2025 | EfficientVLA: Training-Free Acceleration and Compression for Vision-Language-Action ModelsabstractVision-Language-Action (VLA) models, particularly diffusion-based architectures, demonstrate transformative potential for embodied intelligence but are severely hampered by high computational and memory demands stemming from extensive inherent and inference-time redundancies. While existing acceleration efforts often target isolated inefficiencies, such piecemeal solutions typically fail to holistically address the varied computational and memory bottlenecks across the entire VLA pipeline, thereby limiting practical deployability. We introduce EfficientVLA, a structured and training-free inference acceleration framework that systematically eliminates these barriers by cohesively exploiting multifaceted redundancies. EfficientVLA synergistically integrates three targeted strategies: (1) pruning of functionally inconsequential layers from the language module, guided by an analysis of inter-layer redundancies; (2) optimizing the visual processing pathway through a task-aware strategy that selects a compact, diverse set of visual tokens, balancing task-criticality with informational coverage; and (3) alleviating temporal computational redundancy within the iterative diffusion-based action head by strategically caching and reusing key intermediate features.
We apply our method to a standard VLA model CogACT, yielding a $1.93\times$ inference speedup and reduces FLOPs to 28.9%, with only a 0.6% success rate drop in the SIMPLER benchmark. Yantai Yang, Zichen Wen, Luo Zhongwei, Chang Zou, Chuan Wen, Linfeng Zhang 0001 |
NeurIPS | 7 |
| 2025 | Self-supervised vision transformers for semantic segmentation
Xianfan Gu, Yingdong Hu, Chuan Wen, Yang Gao 0029 |
Comput. Vis. Image Underst. | 3 |
| 2024 | Seer: Language Instructed Video Prediction with Latent Diffusion ModelsabstractImagining the future trajectory is the key for robots to make sound planning and successfully reach their goals. Therefore, text-conditioned video prediction (TVP) is an essential task to facilitate general robot policy learning.
To tackle this task and empower robots with the ability to foresee the future, we propose a sample and computation-efficient model, named Seer, by inflating the pretrained text-to-image (T2I) stable diffusion models along the temporal axis. We enhance the U-Net and language conditioning model by incorporating computation-efficient spatial-temporal attention. Furthermore, we introduce a novel Frame Sequential Text Decomposer module that dissects a sentence's global instruction into temporally aligned sub-instructions, ensuring precise integration into each frame of generation. Our framework allows us to effectively leverage the extensive prior knowledge embedded in pretrained T2I models across the frames.
With the adaptable-designed architecture, Seer makes it possible to generate high-fidelity, coherent, and instruction-aligned video frames by fine-tuning a few layers on a small amount of data. The experimental results on Something Something V2 (SSv2), Bridgedata and EpicKitchens-100 datasets demonstrate our superior video prediction performance with around 480-GPU hours versus CogVideo with over 12,480-GPU hours: achieving the 31\% FVD improvement compared to the current SOTA model on SSv2 and 83.7\% average preference in the human evaluation. Our project is available at https://seervideodiffusion.github.io/ Xianfan Gu, Chuan Wen, Weirui Ye, Jiaming Song, Yang Gao 0029 |
ICLR | 2 |
| 2024 | Imitation Learning from Observation with Automatic Discount SchedulingabstractHumans often acquire new skills through observation and imitation. For robotic agents, learning from the plethora of unlabeled video demonstration data available on the Internet necessitates imitating the expert without access to its action, presenting a challenge known as Imitation Learning from Observation (ILfO). A common approach to tackle ILfO problems is to convert them into inverse reinforcement learning problems, utilizing a proxy reward computed from the agent's and the expert's observations. Nonetheless, we identify that tasks characterized by a progress dependency property pose significant challenges for such approaches; in these tasks, the agent needs to initially learn the expert's preceding behaviors before mastering the subsequent ones. Our investigation reveals that the main cause is that the reward signals assigned to later steps hinder the learning of initial behaviors. To address this challenge, we present a novel ILfO framework that enables the agent to master earlier behaviors before advancing to later ones. We introduce an Automatic Discount Scheduling (ADS) mechanism that adaptively alters the discount factor in reinforcement learning during the training phase, prioritizing earlier rewards initially and gradually engaging later rewards only when the earlier behaviors have been mastered. Our experiments, conducted on nine Meta-World tasks, demonstrate that our method significantly outperforms state-of-the-art methods across all tasks, including those that are unsolvable by them. Our code is available at https://il-ads.github.io. Weijun Dong, Yingdong Hu, Chuan Wen, Zhao-Heng Yin, Chongjie Zhang, Yang Gao 0029 |
ICLR | 4 |
| 2024 | Can Transformers Capture Spatial Relations between Objects?abstractSpatial relationships between objects represent key scene information for humans to understand and interact with the world. To study the capability of current computer vision systems to recognize physically grounded spatial relations, we start by proposing precise relation definitions that permit consistently annotating a benchmark dataset. Despite the apparent simplicity of this task relative to others in the recognition literature, we observe that existing approaches perform poorly on this benchmark. We propose new approaches exploiting the long-range attention capabilities of transformers for this task, and evaluating key design principles. We identify a simple ``RelatiViT'' architecture and demonstrate that it outperforms all current approaches. To our knowledge, this is the first method to convincingly outperform naive baselines on spatial relation prediction in in-the-wild settings. The code and datasets are available in \url{https://sites.google.com/view/spatial-relation}. Chuan Wen, Dinesh Jayaraman, Yang Gao 0029 |
ICLR | 1 |
| 2023 | Predictive Inference with Feature Conformal Prediction
Jiaye Teng, Chuan Wen, Dinghuai Zhang, Yoshua Bengio, Yang Gao 0029, Yang Yuan 0010 |
ICLR | 2 |
| 2023 | Biophysically-inspired single-channel speech enhancement in the time domainabstractDeep neural networks (DNN) based speech enhancement approaches have recently achieved great performance. There are a numerous applications that benefit from the speech enhancement model, including automatic speech recognition (ASR) and hearing aids. The majority of these previous methods were developed in the time-frequency (T-F) domain. However, the T-F domain approach has some limitations, including a high minimum delay in reconstructing the signal from T-F domain representation, poor generalizability in unseen noise, and bad performance for negative signal-to-noise ratio’s (SNRs). To address these problems, we propose a biophysically inspired end-to-end time-domain neural network, which adopts bio-inspired features from CoNNear, a neural network that accurately simulates biophysical properties of the human auditory system such as sharp and, level-dependent filter tuning. We first generated biophysical speech feature using CoNNear were subsequently fed into the U-Net-based speech enhancement module. The latter module consisted of generator network without the discriminator from SERGAN (Baby and Verhulst, 2019) and the training dataset we use was INTERSPEECH 2021 DNS Challenge dataset. An objective evaluation was performed using perceptual evaluation of speech quality (PESQ), segmental SNR (segSNR), cepstral distance (CD) and log-likelihood ratio (LLR) with unseen samples from DNS challenge, which is different from training noise scenarios. Results of objective evaluation reveal that bio-inspired features show comparable performance with T-F features at positive SNRs, with improved generalizability in negative SNR and for mismatched noises. Additionally, our time-domain CoNNear features dramatically decreased the minimum latency of the whole system towards 4ms, making it suitable for real-time applications with high constraints on signal delay. The good generalizability in adverse noise conditions and unseen noise, as well as the low latency of our DNN-based model ensure the promising applicability in hearing aids. Chuan Wen, Sarah Verhulst |
INTERSPEECH | 1 |
| 2022 | Resolving Copycat Problems in Visual Imitation Learning via Residual Action Prediction
Chia-Chi Chuang, Donglin Yang, Chuan Wen, Yang Gao 0029 |
ECCV (39) | 3 |
| 2022 | Fighting Fire with Fire: Avoiding DNN Shortcuts through PrimingabstractAcross applications spanning supervised classification and sequential control, deep learning has been reported to find “shortcut” solutions that fail catastrophically under minor changes in the data distribution. In this paper, we show empirically that DNNs can be coaxed to avoid poor shortcuts by providing an additional “priming” feature computed from key input features, usually a coarse output estimate. Priming relies on approximate domain knowledge of these task-relevant key input features, which is often easy to obtain in practical settings. For example, one might prioritize recent frames over past frames in a video input for visual imitation learning, or salient foreground over background pixels for image classification. On NICO image classification, MuJoCo continuous control, and CARLA autonomous driving, our priming strategy works significantly better than several popular state-of-the-art approaches for feature selection and data augmentation. We connect these empirical findings to recent theoretical results on DNN optimization, and argue theoretically that priming distracts the optimizer away from poor shortcuts by creating better, simpler shortcuts. Chuan Wen, Jianing Qian, Jierui Lin, Jiaye Teng, Dinesh Jayaraman, Yang Gao 0029 |
ICML | 1 |
| 2021 | Keyframe-Focused Visual Imitation LearningabstractImitation learning trains control policies by mimicking pre-recorded expert demonstrations. In partially observable settings, imitation policies must rely on observation histories, but many seemingly paradoxical results show better performance for policies that only access the most recent observation. Recent solutions ranging from causal graph learning to deep information bottlenecks have shown promising results, but failed to scale to realistic settings such as visual imitation. We propose a solution that outperforms these prior approaches by upweighting demonstration keyframes corresponding to expert action changepoints. This simple approach easily scales to complex visual imitation settings. Our experimental results demonstrate consistent performance improvements over all baselines on image-based Gym MuJoCo continuous control tasks. Finally, on the CARLA photorealistic vision-based urban driving simulator, we resolve a long-standing issue in behavioral cloning for driving by demonstrating effective imitation from observation histories. Supplementary materials and code at: \url{https://tinyurl.com/imitation-keyframes}. Chuan Wen, Jierui Lin, Jianing Qian, Yang Gao 0029, Dinesh Jayaraman |
ICML | 1 |
| 2021 | Handwritten Chinese Font Generation with Collaborative Stroke RefinementabstractAutomatic character generation is an appealing solution for typeface design, especially for Chinese fonts with over 3700 most commonly-used characters. This task is particularly challenging for handwritten characters with thin strokes which are error-prone during deformation. To handle the generation of thin strokes, we introduce an auxiliary branch for stroke refinement. The auxiliary branch is trained to generate the bold version of target characters which are then fed to the dominating branch to guide the stroke refinement. The two branches are jointly trained in a collaborative fashion. In addition, for practical use, it is desirable to train the character synthesis model with a small set of manually designed characters. Taking advantage of content-reuse phenomenon in Chinese characters, we further propose an online zoom-augmentation strategy to reduce the dependency on large size training sets. The proposed model is trained end-to-end and can be added on top of any method for font synthesis. Experimental results on handwritten font synthesis have shown that the proposed method significantly outperforms the state-of-the-art methods under practical setting, i.e. with only 750 paired training samples. Chuan Wen, Yujie Pan 0001, Ya Zhang 0002, Siheng Chen, Yanfeng Wang 0001, Qi Tian 0001 |
WACV | 1 |
| 2020 | Author Name Disambiguation on Heterogeneous Information Network with Adversarial Representation LearningabstractAuthor name ambiguity causes inadequacy and inconvenience in academic information retrieval, which raises the necessity of author name disambiguation (AND). Existing AND methods can be divided into two categories: the models focusing on content information to distinguish whether two papers are written by the same author, the models focusing on relation information to represent information as edges on the network and to quantify the similarity among papers. However, the former requires adequate labeled samples and informative negative samples, and are also ineffective in measuring the high-order connections among papers, while the latter needs complicated feature engineering or supervision to construct the network. We propose a novel generative adversarial framework to grow the two categories of models together: (i) the discriminative module distinguishes whether two papers are from the same author, and (ii) the generative module selects possibly homogeneous papers directly from the heterogeneous information network, which eliminates the complicated feature engineering. In such a way, the discriminative module guides the generative module to select homogeneous papers, and the generative module generates high-quality negative samples to train the discriminative module to make it aware of high-order connections among papers. Furthermore, a self-training strategy for the discriminative module and a random walk based generating algorithm are designed to make the training stable and efficient. Extensive experiments on two real-world AND benchmarks demonstrate that our model provides significant performance improvement over the state-of-the-art methods. Haiwen Wang, Ruijie Wang 0004, Chuan Wen, Yuting Jia, Weinan Zhang 0001, Xinbing Wang |
AAAI | 3 |
| 2020 | Fighting Copycat Agents in Behavioral Cloning from Observation HistoriesabstractImitation learning trains policies to map from input observations to the actions that an expert would choose. In this setting, distribution shift frequently exacerbates the effect of misattributing expert actions to nuisance correlates among the observed variables. We observe that a common instance of this causal confusion occurs in partially observed settings when expert actions are strongly correlated over time: the imitator learns to cheat by predicting the expert's previous action, rather than the next action. To combat this "copycat problem", we propose an adversarial approach to learn a feature representation that removes excess information about the previous expert action nuisance correlate, while retaining the information necessary to predict the next action. In our experiments, our approach improves performance significantly across a variety of partially observed imitation learning tasks. Chuan Wen, Jierui Lin, Trevor Darrell, Dinesh Jayaraman, Yang Gao 0029 |
NeurIPS | 1 |