EDBT 2026 Demo / reviewers in the wild / expert
Feifei Feng
dblp:27/4916
· DBLP profile ↗
20ranked-venue papers
1as first author
17since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 16 since 2021Systems, architecture and hardware · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | On Support Samples of Next Word PredictionabstractLanguage models excel in various tasks by making complex decisions, yet understanding the rationale behind these decisions remains a challenge.This paper investigates data-centric interpretability in language models, focusing on the next-word prediction task.Using representer theorem, we identify two types of support samples-those that either promote or deter specific predictions.Our findings reveal that being a support sample is an intrinsic property, predictable even before training begins.Additionally, while non-support samples are less influential in direct predictions, they play a critical role in preventing overfitting and shaping generalization and representation learning.Notably, the importance of non-support samples increases in deeper layers, suggesting their significant role in intermediate representation formation.These insights shed light on the interplay between data and model decisions, offering a new dimension to understanding language model behavior and interpretability. 1 * Equal contribution. 1 Our source code is publicly available at https://github.com/liyuqian44/ On-Support-Samples-of-Next-Word-Prediction. Yupei Du, Yufang Liu, Feifei Feng, Mou Xiao Feng, Yuanbin Wu |
ACL (1) | 4 |
| 2025 | ChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action ModelabstractZhongyi Zhou, Yichen Zhu, Minjie Zhu, Junjie Wen, Ning Liu, Zhiyuan Xu, Weibin Meng, Yaxin Peng, Chaomin Shen, Feifei Feng, Yi Xu. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Zhongyi Zhou, Yichen Zhu 0001, Minjie Zhu, Ning Liu 0007, Weibin Meng, Yaxin Peng, Chaomin Shen 0001, Feifei Feng |
EMNLP | 10 |
| 2025 | CoA-VLA: Improving Vision-Language-Action Models via Visual-Textual Chain-of-Affordance
Yichen Zhu 0001, Zhibin Tang, Minjie Zhu, Chengmeng Li, Yaxin Peng, Yan Peng 0001, Feifei Feng |
ICCV | 11 |
| 2025 | DiffusionVLA: Scaling Robot Foundation Models via Unified Diffusion and AutoregressionabstractIn this paper, we present DiffusionVLA, a novel framework that integrates autoregressive reasoning with diffusion policies to address the limitations of existing methods: while autoregressive Vision-Language-Action (VLA) models lack precise and robust action generation, diffusion-based policies inherently lack reasoning capabilities. Central to our approach is autoregressive reasoning — a task decomposition and explanation process enabled by a pre-trained VLM — to guide diffusion-based action policies. To tightly couple reasoning with action generation, we introduce a reasoning injection module that directly embeds self-generated reasoning phrases into the policy learning process. The framework is simple, flexible, and efficient, enabling seamless deployment across diverse robotic platforms.
We conduct extensive experiments using multiple real robots to validate the effectiveness of DiVLA. Our tests include a challenging factory sorting task, where DiVLA successfully categorizes objects, including those not seen during training. The reasoning injection module enhances interpretability, enabling explicit failure diagnosis by visualizing the model’s decision process. Additionally, we test DiVLA on a zero-shot bin-picking task, achieving \textbf{63.7\% accuracy on 102 previously unseen objects}. Our method demonstrates robustness to visual changes, such as distractors and new backgrounds, and easily adapts to new embodiments. Furthermore, DiVLA can follow novel instructions and retain conversational ability. Notably, DiVLA is data-efficient and fast at inference; our smallest DiVLA-2B runs 82Hz on a single A6000 GPU. Finally, we scale the model from 2B to 72B parameters, showcasing improved generalization capabilities with increased model size. Yichen Zhu 0001, Minjie Zhu, Zhibin Tang, Zhongyi Zhou, Chaomin Shen 0001, Yaxin Peng, Feifei Feng |
ICML | 10 |
| 2025 | Scaling Diffusion Policy in Transformer to 1 Billion Parameters for Robotic ManipulationabstractDiffusion Policy is a powerful technique tool for learning end-to-end visuomotor robot control. It is expected that Diffusion Policy possesses scalability, a key attribute for deep neural networks, typically suggesting that increasing model size would lead to enhanced performance. However, our observations indicate that Diffusion Policy in transformer architecture (DP-T) struggles to scale effectively; even minor additions of layers can deteriorate training outcomes. To address this issue, we introduce Scalable Diffusion Transformer Policy for visuomotor learning. Our proposed method, namely ScaleDP, introduces two modules that improve the training dynamic of Diffusion Policy and allow the network to better handle multimodal action distribution. First, we identify that DPT suffers from large gradient issues, making the optimization of Diffusion Policy unstable. To resolve this issue, we factorize the feature embedding of observation into multiple affine layers, and integrate it into the transformer blocks. Additionally, our utilize non-causal attention which allows the policy network to “see” future actions during prediction, helping to reduce compounding errors. We demonstrate that our proposed method successfully scales the Diffusion Policy from 10 million to 1 billion parameters. This new model, named ScaleDP, can effectively scale up the model size with improved performance and generalization. We benchmark ScaleDP across 50 different tasks from MetaWorld and find that our largest ScaleDP outperforms DP-T with an average improvement of 21.6%. Across 7 real-world robot tasks, our ScaleDP demonstrates an average improvement of 36. 25% over DP-T on four single-arm tasks and 75% on three bimanual tasks. We believe our work paves the way for scaling up models for visuomotor learning. The project page is available at https://scaling-diffusion-policy.github.io/. Minjie Zhu, Yichen Zhu 0001, Ning Liu 0007, Chaomin Shen 0001, Yaxin Peng, Feifei Feng, Jian Tang 0008 |
ICRA | 10 |
| 2025 | Let Me Show You: Learning by Retrieving from Egocentric Video for Robotic ManipulationabstractRobots operating in complex and uncertain environments face considerable challenges. Advanced robotic systems often rely on extensive datasets to learn manipulation tasks. In contrast, when humans are faced with unfamiliar tasks, such as assembling a chair, a common approach is to learn by watching video demonstrations. In this paper, we propose a novel method for learning robot policies by Retrieving-from-Video (RfV), using analogies from human demonstrations to address manipulation tasks. Our system constructs a video bank comprising recordings of humans performing diverse daily tasks. To enrich the knowledge from these videos, we extract mid-level information, such as object affordance masks and hand motion trajectories, which serve as additional inputs to enhance the robot model's learning and generalization capabilities. We further feature a dual-component system: a video retriever that taps into an external video bank to fetch task-relevant video based on task specification, and a policy generator that integrates this retrieved knowledge into the learning cycle. This approach enables robots to craft adaptive responses to various scenarios and generalize to tasks beyond those in the training data. Through rigorous testing in multiple simulated and real-world settings, our system demonstrates a marked improvement in performance over conventional robotic systems, showcasing a significant breakthrough in the field of robotics. Yichen Zhu 0001, Feifei Feng |
IROS | 2 |
| 2025 | ACL-QL: Adaptive Conservative Level in Q-Learning for Offline Reinforcement LearningabstractOffline reinforcement learning (RL), which operates solely on static datasets without further interactions with the environment, provides an appealing alternative to learning a safe and promising control policy. The prevailing methods typically learn a conservative policy to mitigate the problem of Q-value overestimation, but it is prone to overdo it, leading to an overly conservative policy. Moreover, they optimize all samples equally with fixed constraints, lacking the nuanced ability to control conservative levels in a fine-grained manner. Consequently, this limitation results in a performance decline. To address the above two challenges in a united way, we propose a framework, adaptive conservative level in Q-learning (ACL-QL), which limits the Q-values in a mild range and enables adaptive control on the conservative level over each state-action pair, i.e., lifting the Q-values more for good transitions and less for bad transitions. We theoretically analyze the conditions under which the conservative level of the learned Q-function can be limited in a mild range and how to optimize each transition adaptively. Motivated by the theoretical analysis, we propose a novel algorithm, ACL-QL, which uses two learnable adaptive weight functions to control the conservative level over each transition. Subsequently, we design a monotonicity loss and surrogate losses to train the adaptive weight functions, Q-function, and policy network alternatively. We evaluate ACL-QL on the commonly used datasets for deep data-driven reinforcement learning (D4RL) benchmark and conduct extensive ablation studies to illustrate the effectiveness and state-of-the-art performance compared with existing offline DRL baselines. Kun Wu 0001, Yinuo Zhao, Zhengping Che, Chengxiang Yin 0001, Chi Harold Liu, Feifei Feng, Jian Tang 0008 |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2024 | Object-Centric Instruction Augmentation for Robotic ManipulationabstractHumans interpret scenes by recognizing both the identities and positions of objects in their observations. For a robot to perform tasks such as "pick and place", understanding both what the objects are and where they are located is crucial. While the former has been extensively discussed in the literature that uses the large language model to enrich the text descriptions, the latter remains underexplored. In this work, we introduce the Object-Centric Instruction Augmentation (OCI) framework to augment highly semantic and information-dense language instruction with position cues. We utilize a Multi-modal Large Language Model (MLLM) to weave knowledge of object locations into natural language instruction, thus aiding the policy network in mastering actions for versatile manipulation. Additionally, we present a feature reuse mechanism to integrate the vision-language features from off-the-shelf pre-trained MLLM into policy networks. Through a series of simulated and real-world robotic tasks, we demonstrate that robotic manipulator imitation policies trained with our enhanced instructions outperform those relying solely on traditional language instructions. Yichen Zhu 0001, Minjie Zhu, Zhengping Che, Chaomin Shen 0001, Yaxin Peng, Dong Liu 0058, Feifei Feng, Jian Tang 0008 |
ICRA | 10 |
| 2024 | Language-Conditioned Robotic Manipulation with Fast and Slow ThinkingabstractThe language-conditioned robotic manipulation aims to transfer natural language instructions into executable actions, from simple "pick-and-place" to tasks requiring intent recognition and visual reasoning. Inspired by the dual-process theory in cognitive science—which suggests two parallel systems of fast and slow thinking in human decision-making—we introduce Robotics with Fast and Slow Thinking (RFST), a framework that mimics human cognitive architecture to classify tasks and makes decisions on two systems based on instruction types. Our RFST consists of two key components: 1) an instruction discriminator to determine which system should be activated based on the current user’s instruction, and 2) a slow-thinking system that is comprised of a fine-tuned vision-language model aligned with the policy networks, which allow the robot to recognize user’s intention or perform reasoning tasks. To assess our methodology, we built a dataset featuring real-world trajectories, capturing actions ranging from spontaneous impulses to tasks requiring deliberate contemplation. Our results, both in simulation and real-world scenarios, confirm that our approach adeptly manages intricate tasks that demand intent recognition and reasoning. Minjie Zhu, Yichen Zhu 0001, Zhengping Che, Chaomin Shen 0001, Yaxin Peng, Dong Liu 0058, Feifei Feng, Jian Tang 0008 |
ICRA | 10 |
| 2024 | EDT: An Efficient Diffusion Transformer Framework Inspired by Human-like SketchingabstractTransformer-based Diffusion Probabilistic Models (DPMs) have shown more potential than CNN-based DPMs, yet their extensive computational requirements hinder widespread practical applications. To reduce the computation budget of transformer-based DPMs, this work proposes the Efficient Diffusion Transformer (EDT) framework. This framework includes a lightweight-design diffusion model architecture, and a training-free Attention Modulation Matrix and its alternation arrangement in EDT inspired by human-like sketching. Additionally, we propose a token relation-enhanced masking training strategy tailored explicitly for EDT to augment its token relation learning capability. Our extensive experiments demonstrate the efficacy of EDT. The EDT framework reduces training and inference costs and surpasses existing transformer-based diffusion models in image synthesis performance, thereby achieving a significant overall enhancement. With lower FID, EDT-S, EDT-B, and EDT-XL attained speed-ups of 3.93x, 2.84x, and 1.92x respectively in the training phase, and 2.29x, 2.29x, and 2.22x respectively in inference, compared to the corresponding sizes of MDTv2. Our code is available at https://github.com/xinwangChen/EDT. Xinwang Chen, Ning Liu 0007, Yichen Zhu 0001, Feifei Feng, Jian Tang 0008 |
NeurIPS | 4 |
| 2024 | Any2Policy: Learning Visuomotor Policy with Any-ModalityabstractHumans can communicate and observe media with different modalities, such as texts, sounds, and images. For robots to be more generalizable embodied agents, they should be capable of following instructions and perceiving the world with adaptation to diverse modalities. Current robotic learning methodologies often focus on single-modal task specification and observation, thereby limiting their ability to process rich multi-modal information. Addressing this limitation, we present an end-to-end general-purpose multi-modal system named Any-to-Policy Embodied Agents. This system empowers robots to handle tasks using various modalities, whether in combinations like text-image, audio-image, text-point cloud, or in isolation. Our innovative approach involves training a versatile modality network that adapts to various inputs and connects with policy networks for effective control. Because of the lack of existing multi-modal robotics datasets for evaluation, we assembled a comprehensive real-world dataset encompassing 30 robotic tasks. Each task in this dataset is richly annotated across multiple modalities, providing a robust foundation for assessment. We conducted extensive validation of our proposed unified modality embodied agent using several simulation benchmarks, including Franka Kitchen, Meta-World, and Maniskill2, as well as in our real-world settings. Our experiments showcase the promising capability of building embodied agents that can adapt to diverse multi-modal in a unified framework. Yichen Zhu 0001, Zhicai Ou, Feifei Feng, Jian Tang 0008 |
NeurIPS | 3 |
| 2024 | RDFC-GAN: RGB-Depth Fusion CycleGAN for Indoor Depth CompletionabstractRaw depth images captured in indoor scenarios frequently exhibit extensive missing values due to the inherent limitations of the sensors and environments. For example, transparent materials frequently elude detection by depth sensors; surfaces may introduce measurement inaccuracies due to their polished textures, extended distances, and oblique incidence angles from the sensor. The presence of incomplete depth maps imposes significant challenges for subsequent vision applications, prompting the development of numerous depth completion techniques to mitigate this problem. Numerous methods excel at reconstructing dense depth maps from sparse samples, but they often falter when faced with extensive contiguous regions of missing depth values, a prevalent and critical challenge in indoor environments. To overcome these challenges, we design a novel two-branch end-to-end fusion network named RDFC-GAN, which takes a pair of RGB and incomplete depth images as input to predict a dense and completed depth map. The first branch employs an encoder-decoder structure, by adhering to the Manhattan world assumption and utilizing normal maps from RGB-D information as guidance, to regress the local dense depth values from the raw depth map. The other branch applies an RGB-depth fusion CycleGAN, adept at translating RGB imagery into detailed, textured depth maps while ensuring high fidelity through cycle consistency. We fuse the two branches via adaptive fusion modules named W-AdaIN and train the model with the help of pseudo depth maps. Comprehensive evaluations on NYU-Depth V2 and SUN RGB-D datasets show that our method significantly enhances depth completion performance particularly in realistic indoor settings. Haowen Wang 0001, Zhengping Che, Mingyuan Wang 0003, Xiuquan Qiao, Mengshi Qi, Feifei Feng, Jian Tang 0008 |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2023 | CP3: Channel Pruning Plug-in for Point-Based NetworksabstractChannel pruning can effectively reduce both computational cost and memory footprint of the original network while keeping a comparable accuracy performance. Though great success has been achieved in channel pruning for 2D image-based convolutional networks (CNNs), existing works seldom extend the channel pruning methods to 3D point-based neural networks (PNNs). Directly implementing the 2D CNN channel pruning methods to PNNs undermine the performance of PNNs because of the different representations of 2D images and 3D point clouds as well as the network architecture disparity. In this paper, we proposed CP3, which is a Channel Pruning Plugin for Point-based network. CP3is elaborately designed to leverage the characteristics of point clouds and PNNs in order to enable 2D channel pruning methods for PNNs. Specifically, it presents a coordinate-enhanced channel importance metric to reflect the correlation between dimensional information and individual channel features, and it recycles the discarded points in PNN's sampling process and reconsiders their potentially-exclusive information to enhance the robustness of channel pruning. Experiments on various PNN architectures show that CP3constantly improves state-of-the-art 2D CNN pruning approaches on different point cloud tasks. For instance, our compressed PointNeXt-S on ScanObjectNN achieves an accuracy of 88.52% with a pruning rate of 57.8%, outperforming the baseline pruning methods with an accuracy gain of 1.94%. Yaomin Huang, Ning Liu 0007, Zhengping Che, Chaomin Shen 0001, Yaxin Peng, Guixu Zhang, Xinmei Liu, Feifei Feng, Jian Tang 0008 |
CVPR | 9 |
| 2023 | CMG-Net: An End-to-End Contact-based Multi-Finger Dexterous Grasping NetworkabstractIn this paper, we propose a novel representation for grasping using contacts between multi-finger robotic hands and objects to be manipulated. This representation significantly reduces the prediction dimensions and accelerates the learning process. We present an effective end-to-end network, CMG-Net, for grasping unknown objects in a cluttered environment by efficiently predicting multi-finger grasp poses and hand configurations from a single-shot point cloud. Moreover, we create a synthetic grasp dataset that consists of five thousand cluttered scenes, 80 object categories, and 20 million annotations. We perform a comprehensive empirical study and demonstrate the effectiveness of our grasping representation and CMG-Net. Our work significantly outperforms the state-of-the-art for three-finger robotic hands. We also demonstrate that the model trained using synthetic data perform very well for real robots. Mingze Wei, Yaomin Huang, Ning Liu 0007, Zhengping Che, Chaomin Shen 0001, Feifei Feng, Chun Shan, Jian Tang 0008 |
ICRA | 8 |
| 2023 | DTF-Net: Category-Level Pose Estimation and Shape Reconstruction via Deformable Template FieldabstractEstimating 6D poses and reconstructing 3D shapes of objects in open-world scenes from RGB-depth image pairs is challenging. Many existing methods rely on learning geometric features that correspond to specific templates while disregarding shape variations and pose differences among objects in the same category. As a result, these methods underperform when handling unseen object instances in complex environments. In contrast, other approaches aim to achieve category-level estimation and reconstruction by leveraging normalized geometric structure priors, but the static prior-based reconstruction struggles with substantial intra-class variations. To solve these problems, we propose the DTF-Net, a novel framework for pose estimation and shape reconstruction based on implicit neural fields of object categories. In DTF-Net, we design a deformable template field to represent the general category-wise shape latent features and intra-category geometric deformation features. The field establishes continuous shape correspondences, deforming the category template into arbitrary observed instances to accomplish shape reconstruction. We introduce a pose regression module that shares the deformation features and template codes from the fields to estimate the accurate 6D pose of each object in the scene. We integrate a multi-modal representation extraction module to extract object features and semantic masks, enabling end-to-end inference. Moreover, during training, we implement a shape-invariant training strategy and a viewpoint sampling method to further enhance the model's capability to extract object pose features. Extensive experiments on the REAL275 and CAMERA25 datasets demonstrate the superiority of DTF-Net in both synthetic and real scenes. Furthermore, we show that DTF-Net effectively supports grasping tasks with a real robot arm. Haowen Wang 0001, Zhengping Che, Dong Liu 0058, Feifei Feng, Yakun Huang, Xiuquan Qiao, Jian Tang 0008 |
ACM Multimedia | 7 |
| 2022 | RGB-Depth Fusion GAN for Indoor Depth CompletionabstractThe raw depth image captured by the indoor depth sen-sor usually has an extensive range of missing depth values due to inherent limitations such as the inability to perceive transparent objects and limited distance range. The incomplete depth map burdens many downstream vision tasks, and a rising number of depth completion methods have been proposed to alleviate this issue. While most existing meth-ods can generate accurate dense depth maps from sparse and uniformly sampled depth maps, they are not suitable for complementing the large contiguous regions of missing depth values, which is common and critical. In this paper, we design a novel two-branch end-to-end fusion network, which takes a pair of RGB and incomplete depth images as input to predict a dense and completed depth map. The first branch employs an encoder-decoder structure to regress the local dense depth values from the raw depth map, with the help of local guidance information extracted from the RGB image. In the other branch, we propose an RGB-depth fusion GAN to transfer the RGB image to the fine-grained textured depth map. We adopt adaptive fusion modules named W-AdaIN to propagate the features across the two branches, and we append a confidence fusion head to fuse the two out-puts of the branches for the final depth map. Extensive ex-periments on NYU-Depth V2 and SUN RGB-D demonstrate that our proposed method clearly improves the depth completion performance, especially in a more realistic setting of indoor environments with the help of the pseudo depth map. Haowen Wang 0001, Mingyuan Wang 0003, Zhengping Che, Xiuquan Qiao, Mengshi Qi, Feifei Feng, Jian Tang 0008 |
CVPR | 7 |
| 2022 | Label-Guided Auxiliary Training Improves 3D Object Detector
Yaomin Huang, Xinmei Liu, Yichen Zhu 0001, Chaomin Shen 0001, Zhengping Che, Guixu Zhang, Yaxin Peng, Feifei Feng, Jian Tang 0008 |
ECCV (9) | 9 |
| 2006 | Instant Service Policy and Its Application to Deficit Round Robin
Jinoo Joung, Dongha Shin, Feifei Feng, Hongkyu Jeong |
AAIM | 3 |
| 2006 | End-to-end stream establishment in consumer home networksabstractThis paper proposes a scheme for end-to-end stream establishment across a layer-2 Residential Ethernet (ResE) and the upper layer UPnP stack in consumer home networks. We first introduce a new proposal for a ResE subscription protocol. This protocol is used to set up guaranteed QoS layer-2 connections between ResE stations. We then propose an extension to the UPnP-AV architecture that enables a seamless integration of UPnP-AV applications and ResE layer 2 technologies. Our subscription protocol proposal is found to be suitable for this integration. The signaling to establish end-to-end AV streams in UPnP/ResE networks is described, and an example usage scenario is demonstrated. Feifei Feng, Hyunsurk Ryu, Kees den Hollander |
CCNC | 1 |
| 2006 | Timing and synchronization for audio/video applications in a converged residential ethernet networkabstractFuture residential networks will be converged, i.e., will carry multimedia, data, and voice applications in a single network infrastructure. These networks will have several features to ensure acceptable Quality of Service (QoS) and minimal administration by users. The features include bandwidth reservation, admission control, and network synchronization. The latter is needed mainly to ensure acceptable jitter, wander, and time synchronization performance for the multimedia applications (these are part of the QoS requirements). This paper describes the jitter, wander, and synchronization requirements for time-sensitive audio and video applications. A general scheme for providing network synchronization, which is being considered for Residential Ethernet, is then described, along with several specific approaches. Geoffrey M. Garner, Feifei Feng, Eric H. S. Ryu, Kees den Hollander |
CCNC | 2 |