EDBT 2026 Demo / reviewers in the wild / expert
Zhengping Che
dblp:160/5944
· DBLP profile ↗
40ranked-venue papers
5as first author
30since 2021 · last 2026
0000-0001-6818-1125ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 31 · 3 first-author · 25 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 10 since 2021Systems, architecture and hardware · 6 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Sparse-View 3-D Language Gaussian Splatting for Zero-Shot Robotic Graspingabstract3-D language Gaussian splatting has recently shown strong potential for open-vocabulary scene understanding and robotic manipulation. However, most existing methods require dense multiview observations to achieve accurate geometry reconstruction and reliable semantic alignment, which limits their applicability in scenarios where only sparse-view observations are available. In this work, we propose SparseGrasper, a framework for language-guided zero-shot robotic grasping under sparse-view conditions. SparseGrasper constructs a 3-D Gaussian language field from as few as three RGB images, enabling joint reasoning over geometry and semantics without the need for dense observations. To improve representation learning under sparse observations, we introduce a dual feature distillation module that fuses local object features with global contextual cues. We further design a language-guided grasp pose generation strategy that incorporates semantic grounding into grasp candidate selection, encouraging grasps that are both semantically relevant and geometrically feasible. Real-world experiments on a 7-DoF robotic manipulator validate that SparseGrasper effectively performs language-guided grasping of diverse, previously unseen objects from sparse observations. Yaonan Wang 0001, Wenrui Chen, He Xie, Zhengping Che, Pei Ren, Jian Tang 0008 |
IEEE Trans. Ind. Informatics | 6 |
| 2025 | Training-Free Generation of Temporally Consistent Rewards from VLMsabstractRecent advances in vision-language models (VLMs) have significantly improved performance in embodied tasks such as goal decomposition and visual comprehension. However, providing accurate rewards for robotic manipulation without fine-tuning VLMs remains challenging due to the absence of domain-specific robotic knowledge in pre-trained datasets and high computational costs that hinder real-time applicability. To address this, we propose $\mathrm{T}^2$-VLM, a novel training-free, temporally consistent framework that generates accurate rewards through tracking the status changes in VLM-derived subgoals. Specifically, our method first queries the VLM to establish spatially aware subgoals and an initial completion estimate before each round of interaction. We then employ a Bayesian tracking algorithm to update the goal completion status dynamically, using subgoal hidden states to generate structured rewards for reinforcement learning (RL) agents. This approach enhances long-horizon decision-making and improves failure recovery capabilities with RL. Extensive experiments indicate that $\mathrm{T}^2$-VLM achieves state-of-the-art performance in two robot manipulation benchmarks, demonstrating superior reward accuracy with reduced computation consumption. We believe our approach not only advances reward generation techniques but also contributes to the broader field of embodied AI. Project website: https://t2-vlm.github.io/. Yinuo Zhao, Jiale Yuan, Xiaoshuai Hao, Xinyi Zhang 0009, Kun Wu 0001, Zhengping Che, Chi Harold Liu, Jian Tang 0008 |
ICCV | 7 |
| 2025 | Learning From Imperfect Demonstrations With Self-Supervision for Robotic ManipulationabstractImproving data utilization, especially for imperfect data from task failures, is crucial for robotic manipulation due to the challenging, time-consuming, and expensive data collection process in the real world. Current imitation learning (IL) typically discards imperfect data, focusing solely on successful expert data. While reinforcement learning (RL) can learn from explorations and failures, the sim2real gap and its reliance on dense reward and online exploration make it difficult to apply effectively in real-world scenarios. In this work, we aim to conquer the challenge of leveraging imperfect data without the need for reward information to improve the model performance for robotic manipulation in an offline manner. Specifically, we introduce a Self-Supervised Data Filtering framework (SSDF) that combines expert and imperfect data to compute quality scores for failed trajectory segments. High-quality segments from the failed data are used to expand the training dataset. Then, the enhanced dataset can be used with any downstream policy learning method for robotic manipulation tasks. Extensive experiments on the ManiSkill2 benchmark built on the high-fidelity Sapien simulator and real-world robotic manipulation tasks using the Franka robot arm demonstrated that the SSDF can accurately expand the training dataset with high-quality imperfect data and improve the success rates for all robotic manipulation tasks. Kun Wu 0001, Ning Liu 0007, Di Qiu, Zhengping Che, Jian Tang 0008 |
ICRA | 6 |
| 2025 | Efficient Training of Generalizable Visuomotor Policies via Control-Aware Augmentation
Yinuo Zhao, Kun Wu 0001, Tianjiao Yi, Zhengping Che, Chi Harold Liu, Jian Tang 0008 |
AAMAS | 5 |
| 2025 | HACTS: a Human-As-Copilot Teleoperation System for Robot LearningabstractTeleoperation is essential for autonomous robot learning, especially in manipulation tasks that require human demonstrations or corrections. However, most existing systems only offer unilateral robot control and lack the ability to synchronize the robot’s status with the teleoperation hardware, preventing real-time, flexible intervention. In this work, we introduce HACTS (Human-As-Copilot Teleoperation System), a novel system that establishes bilateral, real-time joint synchronization between a robot arm and teleoperation hardware. This simple yet effective feedback mechanism, akin to a steering wheel in autonomous vehicles, enables the human copilot to intervene seamlessly while collecting action-correction data for future learning. Implemented using 3D-printed components and low-cost, off-the-shelf motors, HACTS is both accessible and scalable. Our experiments show that HACTS significantly enhances performance in imitation learning (IL) and reinforcement learning (RL) tasks, boosting IL recovery capabilities and data efficiency, and facilitating human-in-the-loop RL. HACTS paves the way for more effective and interactive human-robot collaboration and data-collection, advancing the capabilities of robot manipulation. Yinuo Zhao, Kun Wu 0001, Ning Liu 0007, Junjie Ji, Zhengping Che, Chi Harold Liu, Jian Tang 0008 |
IROS | 6 |
| 2025 | FreqPolicy: Efficient Flow-based Visuomotor Policy via Frequency ConsistencyabstractGenerative modeling-based visuomotor policies have been widely adopted in robotic manipulation, attributed to their ability to model multimodal action distributions. However, the high inference cost of multi-step sampling limits its applicability in real-time robotic systems.
Existing approaches accelerate sampling in generative modeling-based visuomotor policies by adapting techniques originally developed to speed up image generation. However, a major distinction exists: image generation typically produces independent samples without temporal dependencies, while robotic manipulation requires generating action trajectories with continuity and temporal coherence.
To this end, we propose FreqPolicy, a novel approach that first imposes frequency consistency constraints on flow-based visuomotor policies.
Our work enables the action model to capture temporal structure effectively while supporting efficient, high-quality one-step action generation.
Concretely, we introduce a frequency consistency constraint objective that enforces alignment of frequency-domain action features across different timesteps along the flow, thereby promoting convergence of one-step action generation toward the target distribution.
In addition, we design an adaptive consistency loss to capture structural temporal variations inherent in robotic manipulation tasks.
We assess FreqPolicy on $53$ tasks across $3$ simulation benchmarks, proving its superiority over existing one-step action generators.
We further integrate FreqPolicy into the vision-language-action (VLA) model and achieve acceleration without performance degradation on $40$ tasks of Libero.
Besides, we show efficiency and effectiveness in real-world robotic scenarios with an inference frequency of $93.5$ Hz. Yifei Su, Ning Liu 0007, Dong Chen 0044, Kun Wu 0001, Zhengping Che, Jian Tang 0008 |
NeurIPS | 8 |
| 2025 | ACL-QL: Adaptive Conservative Level in Q-Learning for Offline Reinforcement LearningabstractOffline reinforcement learning (RL), which operates solely on static datasets without further interactions with the environment, provides an appealing alternative to learning a safe and promising control policy. The prevailing methods typically learn a conservative policy to mitigate the problem of Q-value overestimation, but it is prone to overdo it, leading to an overly conservative policy. Moreover, they optimize all samples equally with fixed constraints, lacking the nuanced ability to control conservative levels in a fine-grained manner. Consequently, this limitation results in a performance decline. To address the above two challenges in a united way, we propose a framework, adaptive conservative level in Q-learning (ACL-QL), which limits the Q-values in a mild range and enables adaptive control on the conservative level over each state-action pair, i.e., lifting the Q-values more for good transitions and less for bad transitions. We theoretically analyze the conditions under which the conservative level of the learned Q-function can be limited in a mild range and how to optimize each transition adaptively. Motivated by the theoretical analysis, we propose a novel algorithm, ACL-QL, which uses two learnable adaptive weight functions to control the conservative level over each transition. Subsequently, we design a monotonicity loss and surrogate losses to train the adaptive weight functions, Q-function, and policy network alternatively. We evaluate ACL-QL on the commonly used datasets for deep data-driven reinforcement learning (D4RL) benchmark and conduct extensive ablation studies to illustrate the effectiveness and state-of-the-art performance compared with existing offline DRL baselines. Kun Wu 0001, Yinuo Zhao, Zhengping Che, Chengxiang Yin 0001, Chi Harold Liu, Feifei Feng, Jian Tang 0008 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | EPSD: Early Pruning with Self-Distillation for Efficient Model CompressionabstractNeural network compression techniques, such as knowledge distillation (KD) and network pruning, have received increasing attention. Recent work `Prune, then Distill' reveals that a pruned student-friendly teacher network can benefit the performance of KD. However, the conventional teacher-student pipeline, which entails cumbersome pre-training of the teacher and complicated compression steps, makes pruning with KD less efficient. In addition to compressing models, recent compression techniques also emphasize the aspect of efficiency. Early pruning demands significantly less computational cost in comparison to the conventional pruning methods as it does not require a large pre-trained model. Likewise, a special case of KD, known as self-distillation (SD), is more efficient since it requires no pre-training or student-teacher pair selection. This inspires us to collaborate early pruning with SD for efficient model compression. In this work, we propose the framework named Early Pruning with Self-Distillation (EPSD), which identifies and preserves distillable weights in early pruning for a given SD task. EPSD efficiently combines early pruning and self-distillation in a two-step process, maintaining the pruned network's trainability for compression. Instead of a simple combination of pruning and SD, EPSD enables the pruned network to favor SD by keeping more distillable weights before training to ensure better distillation of the pruned network. We demonstrated that EPSD improves the training of pruned networks, supported by visual and quantitative analyses. Our evaluation covered diverse benchmarks (CIFAR-10/100, Tiny-ImageNet, full ImageNet, CUB-200-2011, and Pascal VOC), with EPSD outperforming advanced pruning and SD techniques. Dong Chen 0044, Ning Liu 0007, Yichen Zhu 0001, Zhengping Che, Rui Ma 0011, Fachao Zhang, Xiaofeng Mou, Jian Tang 0008 |
AAAI | 4 |
| 2024 | SM3: Self-supervised Multi-task Modeling with Multi-view 2D Images for Articulated ObjectsabstractReconstructing real-world objects and estimating their movable joint structures are pivotal technologies within the field of robotics. Previous research has predominantly focused on supervised approaches, relying on annotated datasets to model articulated objects within limited categories. However, these approaches fall short of effectively addressing the diversity present in the real world. To tackle this issue, we propose a self-supervised interaction perception method, referred to as SM3, which leverages multi-view RGB images captured before and after interaction to model articulated objects, identify the movable parts, and infer the parameters of their rotating joints. By constructing 3D geometries and textures from the captured 2D images, SM3achieves integrated optimization of movable part and joint parameters during the reconstruction process, obviating the need for annotations. Furthermore, we introduce the MMArt dataset, an extension of PartNet-Mobility, encompassing multi-view and multi-modal data of articulated objects spanning diverse categories. Evaluations demonstrate that SM3surpasses existing benchmarks across various categories and objects, and its adaptability in real-world scenarios has been thoroughly validated. Haowen Wang 0001, Zhengping Che, Yakun Huang, Xiuquan Qiao, Jian Tang 0008 |
ICRA | 4 |
| 2024 | Object-Centric Instruction Augmentation for Robotic ManipulationabstractHumans interpret scenes by recognizing both the identities and positions of objects in their observations. For a robot to perform tasks such as "pick and place", understanding both what the objects are and where they are located is crucial. While the former has been extensively discussed in the literature that uses the large language model to enrich the text descriptions, the latter remains underexplored. In this work, we introduce the Object-Centric Instruction Augmentation (OCI) framework to augment highly semantic and information-dense language instruction with position cues. We utilize a Multi-modal Large Language Model (MLLM) to weave knowledge of object locations into natural language instruction, thus aiding the policy network in mastering actions for versatile manipulation. Additionally, we present a feature reuse mechanism to integrate the vision-language features from off-the-shelf pre-trained MLLM into policy networks. Through a series of simulated and real-world robotic tasks, we demonstrate that robotic manipulator imitation policies trained with our enhanced instructions outperform those relying solely on traditional language instructions. Yichen Zhu 0001, Minjie Zhu, Zhengping Che, Chaomin Shen 0001, Yaxin Peng, Dong Liu 0058, Feifei Feng, Jian Tang 0008 |
ICRA | 6 |
| 2024 | Language-Conditioned Robotic Manipulation with Fast and Slow ThinkingabstractThe language-conditioned robotic manipulation aims to transfer natural language instructions into executable actions, from simple "pick-and-place" to tasks requiring intent recognition and visual reasoning. Inspired by the dual-process theory in cognitive science—which suggests two parallel systems of fast and slow thinking in human decision-making—we introduce Robotics with Fast and Slow Thinking (RFST), a framework that mimics human cognitive architecture to classify tasks and makes decisions on two systems based on instruction types. Our RFST consists of two key components: 1) an instruction discriminator to determine which system should be activated based on the current user’s instruction, and 2) a slow-thinking system that is comprised of a fine-tuned vision-language model aligned with the policy networks, which allow the robot to recognize user’s intention or perform reasoning tasks. To assess our methodology, we built a dataset featuring real-world trajectories, capturing actions ranging from spontaneous impulses to tasks requiring deliberate contemplation. Our results, both in simulation and real-world scenarios, confirm that our approach adeptly manages intricate tasks that demand intent recognition and reasoning. Minjie Zhu, Yichen Zhu 0001, Zhengping Che, Chaomin Shen 0001, Yaxin Peng, Dong Liu 0058, Feifei Feng, Jian Tang 0008 |
ICRA | 6 |
| 2024 | RDFC-GAN: RGB-Depth Fusion CycleGAN for Indoor Depth CompletionabstractRaw depth images captured in indoor scenarios frequently exhibit extensive missing values due to the inherent limitations of the sensors and environments. For example, transparent materials frequently elude detection by depth sensors; surfaces may introduce measurement inaccuracies due to their polished textures, extended distances, and oblique incidence angles from the sensor. The presence of incomplete depth maps imposes significant challenges for subsequent vision applications, prompting the development of numerous depth completion techniques to mitigate this problem. Numerous methods excel at reconstructing dense depth maps from sparse samples, but they often falter when faced with extensive contiguous regions of missing depth values, a prevalent and critical challenge in indoor environments. To overcome these challenges, we design a novel two-branch end-to-end fusion network named RDFC-GAN, which takes a pair of RGB and incomplete depth images as input to predict a dense and completed depth map. The first branch employs an encoder-decoder structure, by adhering to the Manhattan world assumption and utilizing normal maps from RGB-D information as guidance, to regress the local dense depth values from the raw depth map. The other branch applies an RGB-depth fusion CycleGAN, adept at translating RGB imagery into detailed, textured depth maps while ensuring high fidelity through cycle consistency. We fuse the two branches via adaptive fusion modules named W-AdaIN and train the model with the help of pseudo depth maps. Comprehensive evaluations on NYU-Depth V2 and SUN RGB-D datasets show that our method significantly enhances depth completion performance particularly in realistic indoor settings. Haowen Wang 0001, Zhengping Che, Mingyuan Wang 0003, Xiuquan Qiao, Mengshi Qi, Feifei Feng, Jian Tang 0008 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | Human Pose Transfer with Augmented Disentangled Feature ConsistencyabstractDeep generative models have made great progress in synthesizing images with arbitrary human poses and transferring the poses of one person to others. Though many different methods have been proposed to generate images with high visual fidelity, the main challenge remains and comes from two fundamental issues: pose ambiguity and appearance inconsistency. To alleviate the current limitations and improve the quality of the synthesized images, we propose a pose transfer network with augmented D isentangled F eature C onsistency (DFC-Net) to facilitate human pose transfer. Given a pair of images containing the source and target person, DFC-Net extracts pose and static information from the source and target respectively, then synthesizes an image of the target person with the desired pose from the source. Moreover, DFC-Net leverages disentangled feature consistency losses in the adversarial training to strengthen the transfer coherence and integrates a keypoint amplifier to enhance the pose feature extraction. With the help of the disentangled feature consistency losses, we further propose a novel data augmentation scheme that introduces unpaired support data with the augmented consistency constraints to improve the generality and robustness of DFC-Net. Extensive experimental results on Mixamo-Pose and EDN-10k have demonstrated DFC-Net achieves state-of-the-art performance on pose transfer. Kun Wu 0001, Chengxiang Yin 0001, Zhengping Che, Jian Tang 0008 |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2024 | CATRO: Channel Pruning via Class-Aware Trace Ratio OptimizationabstractDeep convolutional neural networks are shown to be overkill with high parametric and computational redundancy in many application scenarios, and an increasing number of works have explored model pruning to obtain lightweight and efficient networks. However, most existing pruning approaches are driven by empirical heuristics and rarely consider the joint impact of channels, leading to unguaranteed and suboptimal performance. In this article, we propose a novel channel pruning method via c lass-aware t race r atio o ptimization (CATRO) to reduce the computational burden and accelerate the model inference. Utilizing class information from a few samples, CATRO measures the joint impact of multiple channels by feature space discriminations and consolidates the layerwise impact of preserved channels. By formulating channel pruning as a submodular set function maximization problem, CATRO solves it efficiently via a two-stage greedy iterative optimization procedure. More importantly, we present theoretical justifications on convergence of CATRO and performance of pruned networks. Experimental results demonstrate that CATRO achieves higher accuracy with similar computation cost or lower computation cost with similar accuracy than other state-of-the-art channel pruning algorithms. In addition, because of its class-aware property, CATRO is suitable to prune efficient networks adaptively for various classification subtasks, enhancing handy deployment and usage of deep networks in real-world applications. Wenzheng Hu, Zhengping Che, Ning Liu 0007, Jian Tang 0008, Changshui Zhang, Jianqiang Wang 0003 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | CP3: Channel Pruning Plug-in for Point-Based NetworksabstractChannel pruning can effectively reduce both computational cost and memory footprint of the original network while keeping a comparable accuracy performance. Though great success has been achieved in channel pruning for 2D image-based convolutional networks (CNNs), existing works seldom extend the channel pruning methods to 3D point-based neural networks (PNNs). Directly implementing the 2D CNN channel pruning methods to PNNs undermine the performance of PNNs because of the different representations of 2D images and 3D point clouds as well as the network architecture disparity. In this paper, we proposed CP3, which is a Channel Pruning Plugin for Point-based network. CP3is elaborately designed to leverage the characteristics of point clouds and PNNs in order to enable 2D channel pruning methods for PNNs. Specifically, it presents a coordinate-enhanced channel importance metric to reflect the correlation between dimensional information and individual channel features, and it recycles the discarded points in PNN's sampling process and reconsiders their potentially-exclusive information to enhance the robustness of channel pruning. Experiments on various PNN architectures show that CP3constantly improves state-of-the-art 2D CNN pruning approaches on different point cloud tasks. For instance, our compressed PointNeXt-S on ScanObjectNN achieves an accuracy of 88.52% with a pruning rate of 57.8%, outperforming the baseline pruning methods with an accuracy gain of 1.94%. Yaomin Huang, Ning Liu 0007, Zhengping Che, Chaomin Shen 0001, Yaxin Peng, Guixu Zhang, Xinmei Liu, Feifei Feng, Jian Tang 0008 |
CVPR | 3 |
| 2023 | CMG-Net: An End-to-End Contact-based Multi-Finger Dexterous Grasping NetworkabstractIn this paper, we propose a novel representation for grasping using contacts between multi-finger robotic hands and objects to be manipulated. This representation significantly reduces the prediction dimensions and accelerates the learning process. We present an effective end-to-end network, CMG-Net, for grasping unknown objects in a cluttered environment by efficiently predicting multi-finger grasp poses and hand configurations from a single-shot point cloud. Moreover, we create a synthetic grasp dataset that consists of five thousand cluttered scenes, 80 object categories, and 20 million annotations. We perform a comprehensive empirical study and demonstrate the effectiveness of our grasping representation and CMG-Net. Our work significantly outperforms the state-of-the-art for three-finger robotic hands. We also demonstrate that the model trained using synthetic data perform very well for real robots. Mingze Wei, Yaomin Huang, Ning Liu 0007, Zhengping Che, Chaomin Shen 0001, Feifei Feng, Chun Shan, Jian Tang 0008 |
ICRA | 5 |
| 2023 | DTF-Net: Category-Level Pose Estimation and Shape Reconstruction via Deformable Template FieldabstractEstimating 6D poses and reconstructing 3D shapes of objects in open-world scenes from RGB-depth image pairs is challenging. Many existing methods rely on learning geometric features that correspond to specific templates while disregarding shape variations and pose differences among objects in the same category. As a result, these methods underperform when handling unseen object instances in complex environments. In contrast, other approaches aim to achieve category-level estimation and reconstruction by leveraging normalized geometric structure priors, but the static prior-based reconstruction struggles with substantial intra-class variations. To solve these problems, we propose the DTF-Net, a novel framework for pose estimation and shape reconstruction based on implicit neural fields of object categories. In DTF-Net, we design a deformable template field to represent the general category-wise shape latent features and intra-category geometric deformation features. The field establishes continuous shape correspondences, deforming the category template into arbitrary observed instances to accomplish shape reconstruction. We introduce a pose regression module that shares the deformation features and template codes from the fields to estimate the accurate 6D pose of each object in the scene. We integrate a multi-modal representation extraction module to extract object features and semantic masks, enabling end-to-end inference. Moreover, during training, we implement a shape-invariant training strategy and a viewpoint sampling method to further enhance the model's capability to extract object pose features. Extensive experiments on the REAL275 and CAMERA25 datasets demonstrate the superiority of DTF-Net in both synthetic and real scenes. Furthermore, we show that DTF-Net effectively supports grasping tasks with a real robot arm. Haowen Wang 0001, Zhengping Che, Dong Liu 0058, Feifei Feng, Yakun Huang, Xiuquan Qiao, Jian Tang 0008 |
ACM Multimedia | 4 |
| 2023 | Distributional generative adversarial imitation learning with reproducing kernel generalization
Yirui Zhou, Mengxiao Lu, Zhengping Che, Jian Tang 0008, Yangchun Zhang, Yan Peng 0001, Yaxin Peng |
Neural Networks | 4 |
| 2023 | SRRNet: A Semantic Representation Refinement Network for Image SegmentationabstractSemantic context has raised concerns in semantic segmentation. In most cases, it is applied to guide feature learning. Instead, this paper applies it to extract the semantic representation, which records the global feature information of each category with a memory tensor. Specifically, we propose a novel semantic representation (SR) module, which consists of semantic embedding (SE) and semantic attention (SA) blocks. The SE block adaptively embeds features into the semantic representation by calculating the memory similarity, and the SA block aggregates the embedded features with semantic attention. The main advantages of the SR module lie in three aspects: i) it enhances the representation ability of semantic context by employing global (cross-image) semantic information; ii) it improves the consistency of intraclass features by aggregating global features of the same categories; and iii) it can be extended to build a semantic representation refinement network (SRRNet) by iteratively applying the SR module across multiple scales, shrinking the semantic gap and enhancing the structural reasoning of the model. Extensive experiments demonstrate that our method significantly improves the segmentation results and achieves superior performance on the PASCAL VOC 2012, Cityscapes, and PASCAL Context datasets. Xiaofeng Ding 0003, Tieyong Zeng, Jian Tang 0008, Zhengping Che, Yaxin Peng |
IEEE Trans. Multim. | 4 |
| 2023 | DAIS: Automatic Channel Pruning via Differentiable Annealing Indicator SearchabstractThe convolutional neural network (CNN) has achieved great success in fulfilling computer vision tasks despite large computation overhead against efficient deployment. Channel pruning is usually applied to reduce the model redundancy while preserving the network structure, such that the pruned network can be easily deployed in practice. However, existing channel pruning methods require hand-crafted rules, which can result in a degraded model performance with respect to the tremendous potential pruning space given large neural networks. In this article, we introduce differentiable annealing indicator search (DAIS) that leverages the strength of neural architecture search in the channel pruning and automatically searches for the effective pruned model with given constraints on computation overhead. Specifically, DAIS relaxes the binarized channel indicators to be continuous and then jointly learns both indicators and model parameters via bi-level optimization. To bridge the non-negligible discrepancy between the continuous model and the target binarized model, DAIS proposes an annealing-based procedure to steer the indicator convergence toward binarized states. Moreover, DAIS designs various regularizations based on a priori structural knowledge to control the pruning sparsity and to improve model performance. Experimental results show that DAIS outperforms state-of-the-art pruning methods on CIFAR-10, CIFAR-100, and ImageNet. Yushuo Guan, Ning Liu 0007, Zhengping Che, Kaigui Bian, Yanzhi Wang 0001, Jian Tang 0008 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2022 | I-SEA: Importance Sampling and Expected Alignment-Based Deep Distance Metric Learning for Time Series Analysis and EmbeddingabstractLearning effective embeddings for potentially irregularly sampled time-series, evolving at different time scales, is fundamental for machine learning tasks such as classification and clustering. Task-dependent embeddings rely on similarities between data samples to learn effective geometries. However, many popular time-series similarity measures are not valid distance metrics, and as a result they do not reliably capture the intricate relationships between the multi-variate time-series data samples for learning effective embeddings. One of the primary ways to formulate an accurate distance metric is by forming distance estimates via Monte-Carlo-based expectation evaluations. However, the high-dimensionality of the underlying distribution, and the inability to sample from it, pose significant challenges. To this end, we develop an Importance Sampling based distance metric -- I-SEA -- which enjoys the properties of a metric while consistently achieving superior performance for machine learning tasks such as classification and representation learning. I-SEA leverages Importance Sampling and Non-parametric Density Estimation to adaptively estimate distances, enabling implicit estimation from the underlying high-dimensional distribution, resulting in improved accuracy and reduced variance. We theoretically establish the properties of I-SEA and demonstrate its capabilities via experimental evaluations on real-world healthcare datasets. Sirisha Rambhatla, Zhengping Che, Yan Liu 0002 |
AAAI | 2 |
| 2022 | CADRE: A Cascade Deep Reinforcement Learning Framework for Vision-Based Autonomous Urban DrivingabstractVision-based autonomous urban driving in dense traffic is quite challenging due to the complicated urban environment and the dynamics of the driving behaviors. Widely-applied methods either heavily rely on hand-crafted rules or learn from limited human experience, which makes them hard to generalize to rare but critical scenarios. In this paper, we present a novel CAscade Deep REinforcement learning framework, CADRE, to achieve model-free vision-based autonomous urban driving. In CADRE, to derive representative latent features from raw observations, we first offline train a Co-attention Perception Module (CoPM) that leverages the co-attention mechanism to learn the inter-relationships between the visual and control information from a pre-collected driving dataset. Cascaded by the frozen CoPM, we then present an efficient distributed proximal policy optimization framework to online learn the driving policy under the guidance of particularly designed reward functions. We perform a comprehensive empirical study with the CARLA NoCrash benchmark as well as specific obstacle avoidance scenarios in autonomous urban driving tasks. The experimental results well justify the effectiveness of CADRE and its superiority over the state-of-the-art by a wide margin. Yinuo Zhao, Kun Wu 0001, Zhengping Che, Jian Tang 0008, Chi Harold Liu |
AAAI | 4 |
| 2022 | RGB-Depth Fusion GAN for Indoor Depth CompletionabstractThe raw depth image captured by the indoor depth sen-sor usually has an extensive range of missing depth values due to inherent limitations such as the inability to perceive transparent objects and limited distance range. The incomplete depth map burdens many downstream vision tasks, and a rising number of depth completion methods have been proposed to alleviate this issue. While most existing meth-ods can generate accurate dense depth maps from sparse and uniformly sampled depth maps, they are not suitable for complementing the large contiguous regions of missing depth values, which is common and critical. In this paper, we design a novel two-branch end-to-end fusion network, which takes a pair of RGB and incomplete depth images as input to predict a dense and completed depth map. The first branch employs an encoder-decoder structure to regress the local dense depth values from the raw depth map, with the help of local guidance information extracted from the RGB image. In the other branch, we propose an RGB-depth fusion GAN to transfer the RGB image to the fine-grained textured depth map. We adopt adaptive fusion modules named W-AdaIN to propagate the features across the two branches, and we append a confidence fusion head to fuse the two out-puts of the branches for the final depth map. Extensive ex-periments on NYU-Depth V2 and SUN RGB-D demonstrate that our proposed method clearly improves the depth completion performance, especially in a more realistic setting of indoor environments with the help of the pseudo depth map. Haowen Wang 0001, Mingyuan Wang 0003, Zhengping Che, Xiuquan Qiao, Mengshi Qi, Feifei Feng, Jian Tang 0008 |
CVPR | 3 |
| 2022 | Label-Guided Auxiliary Training Improves 3D Object Detector
Yaomin Huang, Xinmei Liu, Yichen Zhu 0001, Chaomin Shen 0001, Zhengping Che, Guixu Zhang, Yaxin Peng, Feifei Feng, Jian Tang 0008 |
ECCV (9) | 6 |
| 2022 | Generalization and Computation for Policy Classes of Generative Adversarial Imitation Learning
Yirui Zhou, Yangchun Zhang, Wanying Wang, Zhengping Che, Jian Tang 0008, Yaxin Peng |
PPSN (1) | 5 |
| 2022 | Alleviating Data Sparsity Problems in Estimated Time of Arrival via Auxiliary Metric LearningabstractWith millions of people using ride-hailing platforms for daily travel, estimated time of arrival (ETA) has become a significant problem in intelligent transportation systems and attracted considerable attention recently. Deep learning-based ETA methods have achieved promising results using massive spatial-temporal data. However, we find that the prediction accuracy is not satisfactory in practical applications due to the prevalent data sparsity problems. Instead of focusing on the average prediction performance as many other methods, this study aims to alleviate the data sparsity problems in ETA to enhance user experience. In general, the data sparsity problems arise from two aspects. The first is the road network, where many links are only traversed by few floating cars. The second aspect is drivers, where many drivers’ trajectories are too scarce (e.g., with only 3 trip records). To alleviate the sparsity in road network, we propose a Road Network Metric Learning framework for ETA (RNML-ETA), where an auxiliary metric learning task is used to improve the link-embedding, especially for links with insufficient data. A novel triangle loss is proposed to improve metric learning effectiveness for links. Experiments on massive real-world data show that RNML-ETA outperforms competing methods by promoting the cold links with limited data. Furthermore, we propose a novel unified framework to Alleviate Data Sparsity problems in ETA (ADS-ETA) by extending RNML-ETA with an additional auxiliary task for driver ID embedding. Results with extensive experiments demonstrate that ADS-ETA can effectively alleviate the data sparsity problems caused by road network and driver sparsity. Wenzheng Hu, Donghua Zhou, Baichuan Mo, Kun Fu 0002, Zhengping Che, Zheng Wang 0010, Shenhao Wang, Jinhua Zhao 0001, Jieping Ye, Jian Tang 0008, Changshui Zhang |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2022 | Robust Unsupervised Video Anomaly Detection by Multipath Frame PredictionabstractVideo anomaly detection is commonly used in many applications, such as security surveillance, and is very challenging. A majority of recent video anomaly detection approaches utilize deep reconstruction models, but their performance is often suboptimal because of insufficient reconstruction error differences between normal and abnormal video frames in practice. Meanwhile, frame prediction-based anomaly detection methods have shown promising performance. In this article, we propose a novel and robust unsupervised video anomaly detection method by frame prediction with a proper design which is more in line with the characteristics of surveillance videos. The proposed method is equipped with a multipath ConvGRU-based frame prediction network that can better handle semantically informative objects and areas of different scales and capture spatial-temporal dependencies in normal videos. A noise tolerance loss is introduced during training to mitigate the interference caused by background noise. Extensive experiments have been conducted on the CUHK Avenue, ShanghaiTech Campus, and UCSD Pedestrian datasets, and the results show that our proposed method outperforms existing state-of-the-art approaches. Remarkably, our proposed method obtains the frame-level AUROC score of 88.3% on the CUHK Avenue dataset. Xuanzhao Wang, Zhengping Che, Jian Tang 0008, Jieping Ye, Jingyu Wang 0001, Qi Qi 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2021 | Hierarchical Graph Attention Network for Few-shot Visual-Semantic LearningabstractDeep learning has made tremendous success in computer vision, natural language processing and even visual-semantic learning, which requires a huge amount of labeled training data. Nevertheless, the goal of human-level intelligence is to enable a model to quickly obtain an in-depth understanding given a small number of samples, especially with heterogeneity in the multi-modal scenarios such as visual question answering and image captioning. In this paper, we study the few-shot visual-semantic learning and present the Hierarchical Graph ATtention network (HGAT). This two-stage network models the intra- and inter-modal relationships with limited image-text samples. The main contributions of HGAT can be summarized as follows: 1) it sheds light on tackling few-shot multi-modal learning problems, which focuses primarily, but not exclusively on visual and semantic modalities, through better exploitation of the intra-relationship of each modality and an attention-based co-learning framework between modalities using a hierarchical graph-based architecture; 2) it achieves superior performance on both visual question answering and image captioning in the few-shot setting; 3) it can be easily extended to the semi-supervised setting where image-text samples are partially unlabeled. We show via extensive experiments that HGAT delivers state-of-the-art performance on three widely-used benchmarks of two visual-semantic learning tasks. Chengxiang Yin 0001, Kun Wu 0001, Zhengping Che, Jian Tang 0008 |
ICCV | 3 |
| 2021 | Lottery Ticket Preserves Weight Correlation: Is It Desirable or Not?abstractIn deep model compression, the recent finding "Lottery Ticket Hypothesis" (LTH) pointed out that there could exist a winning ticket (i.e., a properly pruned sub-network together with original weight initialization) that can achieve competitive performance than the original dense network. However, it is not easy to observe such winning property in many scenarios, where for example, a relatively large learning rate is used even if it benefits training the original dense model. In this work, we investigate the underlying condition and rationale behind the winning property, and find that the underlying reason is largely attributed to the correlation between initialized weights and final-trained weights when the learning rate is not sufficiently large. Thus, the existence of winning property is correlated with an insufficient DNN pretraining, and is unlikely to occur for a well-trained DNN. To overcome this limitation, we propose the "pruning & fine-tuning" method that consistently outperforms lottery ticket sparse training under the same pruning algorithm and the same total training epochs. Extensive experiments over multiple deep models (VGG, ResNet, MobileNet-v2) on different datasets have been conducted to justify our proposals. Ning Liu 0007, Geng Yuan, Zhengping Che, Xuan Shen, Qing Jin, Jian Ren 0005, Jian Tang 0008, Sijia Liu 0001, Yanzhi Wang 0001 |
ICML | 3 |
| 2021 | MiLeTS'21: 7th KDD Workshop on Mining and Learning from Time SeriesabstractTime series data are ubiquitous. Rapid advances in diverse sensing technologies, ranging from remote sensors to wearables and social sensing, are generating a rapid growth in the size and complexity of time series archives. This has resulted in a fundamental shift away from parsimonious, infrequent measurement to nearly continuous monitoring and recording. This demands development of new tools and solutions. The goals of this workshop are to: (1) highlight the significant challenges that underpin learning and mining from time series data (e.g. irregular sampling, spatiotemporal structure, and uncertainty quantification), (2) discuss recent algorithmic, theoretical, statistical, or systems-based developments for tackling these problems, and (3) synergize the research activities and discuss both new and open problems in time series analysis and mining. Sanjay Purushotham, Zhengping Che |
KDD | 3 |
| 2020 | Generative Attention Networks for Multi-Agent Behavioral ModelingabstractUnderstanding and modeling behavior of multi-agent systems is a central step for artificial intelligence. Here we present a deep generative model which captures behavior generating process of multi-agent systems, supports accurate predictions and inference, infers how agents interact in a complex system, as well as identifies agent groups and interaction types. Built upon advances in deep generative models and a novel attention mechanism, our model can learn interactions in highly heterogeneous systems with linear complexity in the number of agents. We apply this model to three multi-agent systems in different domains and evaluate performance on a diverse set of tasks including behavior prediction, interaction analysis and system identification. Experimental results demonstrate its ability to model multi-agent systems, yielding improved performance over competitive baselines. We also show the model can successfully identify agent groups and interaction types in these systems. Our model offers new opportunities to predict complex multi-agent behaviors and takes a step forward in understanding interactions in multi-agent systems. Max Guangyu Li, Bo Jiang 0001, Zhengping Che, Yan Liu 0002 |
AAAI | 4 |
| 2020 | Knowledge Transfer in Multi-Task Deep Reinforcement Learning for Continuous ControlabstractWhile Deep Reinforcement Learning (DRL) has emerged as a promising approach to many complex tasks, it remains challenging to train a single DRL agent that is capable of undertaking multiple different continuous control tasks. In this paper, we present a Knowledge Transfer based Multi-task Deep Reinforcement Learning framework (KTM-DRL) for continuous control, which enables a single DRL agent to achieve expert-level performance in multiple different tasks by learning from task-specific teachers. In KTM-DRL, the multi-task agent first leverages an offline knowledge transfer algorithm designed particularly for the actor-critic architecture to quickly learn a control policy from the experience of task-specific teachers, and then it employs an online learning algorithm to further improve itself by learning from new online transition samples under the guidance of those teachers. We perform a comprehensive empirical study with two commonly-used benchmarks in the MuJoCo continuous control task suite. The experimental results well justify the effectiveness of KTM-DRL and its knowledge transfer and online learning algorithms, as well as its superiority over the state-of-the-art by a large margin. Kun Wu 0001, Zhengping Che, Jian Tang 0008, Jieping Ye |
NeurIPS | 3 |
| 2018 | Hierarchical Deep Generative Models for Multi-Rate Multivariate Time SeriesabstractMulti-Rate Multivariate Time Series (MR-MTS) are the multivariate time series observations which come with various sampling rates and encode multiple temporal dependencies. State-space models such as Kalman filters and deep learning models such as deep Markov models are mainly designed for time series data with the same sampling rate and cannot capture all the dependencies present in the MR-MTS data. To address this challenge, we propose the Multi-Rate Hierarchical Deep Markov Model (MR-HDMM), a novel deep generative model which uses the latent hierarchical structure with a learnable switch mechanism to capture the temporal dependencies of MR-MTS. Experimental results on two real-world datasets demonstrate that our MR-HDMM model outperforms the existing state-of-the-art deep learning and state-space models on forecasting and interpolation tasks. In addition, the latent hierarchies in our model provide a way to show and interpret the multiple temporal dependencies. Zhengping Che, Sanjay Purushotham, Max Guangyu Li, Bo Jiang 0001, Yan Liu 0002 |
ICML | 1 |
| 2018 | Benchmarking deep learning models on large healthcare datasetsabstractDeep learning models (aka Deep Neural Networks) have revolutionized many fields including computer vision, natural language processing, speech recognition, and is being increasingly used in clinical healthcare applications. However, few works exist which have benchmarked the performance of the deep learning models with respect to the state-of-the-art machine learning models and prognostic scoring systems on publicly available healthcare datasets. In this paper, we present the benchmarking results for several clinical prediction tasks such as mortality prediction, length of stay prediction, and ICD-9 code group prediction using Deep Learning models, ensemble of machine learning models (Super Learner algorithm), SAPS II and SOFA scores. We used the Medical Information Mart for Intensive Care III (MIMIC-III) (v1.4) publicly available dataset, which includes all patients admitted to an ICU at the Beth Israel Deaconess Medical Center from 2001 to 2012, for the benchmarking tasks. Our results show that deep learning models consistently outperform all the other approaches especially when the 'raw' clinical time series data is used as input features to the models. Sanjay Purushotham, Chuizheng Meng, Zhengping Che, Yan Liu 0002 |
J. Biomed. Informatics | 3 |
| 2017 | Deep Learning Solutions for Classifying Patients on Opioid Use
Zhengping Che, Jennifer L. St. Sauver, Yan Liu 0002 |
AMIA | 1 |
| 2017 | Boosting Deep Learning Risk Prediction with Generative Adversarial Networks for Electronic Health RecordsabstractThe rapid growth of Electronic Health Records (EHRs), as well as the accompanied opportunities in Data-Driven Healthcare (DDH), has been attracting widespread interests and attentions. Recent progress in the design and applications of deep learning methods has shown promising results and is forcing massive changes in healthcare academia and industry, but most of these methods rely on massive labeled data. In this work, we propose a general deep learning framework which is able to boost risk prediction performance with limited EHR data. Our model takes a modified generative adversarial network namely ehrGAN, which can provide plausible labeled EHR data by mimicking real patient records, to augment the training dataset in a semi-supervised learning manner. We use this generative model together with a convolutional neural network (CNN) based prediction model to improve the onset prediction performance. Experiments on two real healthcare datasets demonstrate that our proposed framework produces realistic data samples and achieves significant improvements on classification tasks with the generated data over several stat-of-the-art baselines. Zhengping Che, Yu Cheng 0001, Shuangfei Zhai, Zhaonan Sun, Yan Liu 0002 |
ICDM | 1 |
| 2016 | Interpretable Deep Models for ICU Outcome Prediction
Zhengping Che, Sanjay Purushotham, Robinder G. Khemani, Yan Liu 0002 |
AMIA | 1 |
| 2015 | Causal Phenotype Discovery via Deep Networks
David C. Kale, Zhengping Che, Mohammad Taha Bahadori, Yan Liu 0002, Randall C. Wetzel |
AMIA | 2 |
| 2015 | Deep Computational PhenotypingabstractWe apply deep learning to the problem of discovery and detection of characteristic patterns of physiology in clinical time series data. We propose two novel modifications to standard neural net training that address challenges and exploit properties that are peculiar, if not exclusive, to medical data. First, we examine a general framework for using prior knowledge to regularize parameters in the topmost layers. This framework can leverage priors of any form, ranging from formal ontologies (e.g., ICD9 codes) to data-derived similarity. Second, we describe a scalable procedure for training a collection of neural networks of different sizes but with partially shared architectures. Both of these innovations are well-suited to medical applications, where available data are not yet Internet scale and have many sparse outputs (e.g., rare diagnoses) but which have exploitable structure (e.g., temporal order and relationships between labels). However, both techniques are sufficiently general to be applied to other problems and domains. We demonstrate the empirical efficacy of both techniques on two real-world hospital data sets and show that the resulting neural nets learn interpretable and clinically relevant features. Zhengping Che, David C. Kale, Mohammad Taha Bahadori, Yan Liu 0002 |
KDD | 1 |
| 2014 | An Examination of Multivariate Time Series Hashing with Applications to Health CareabstractAs large-scale multivariate time series data become increasingly common in application domains, such as health care and traffic analysis, researchers are challenged to build efficient tools to analyze it and provide useful insights. Similarity search, as a basic operator for many machine learning and data mining algorithms, has been extensively studied before, leading to several efficient solutions. However, similarity search for multivariate time series data is intrinsically challenging because (1) there is no conclusive agreement on what is a good similarity metric for multivariate time series data and (2) calculating similarity scores between two time series is often computationally expensive. In this paper, we address this problem by applying a generalized hashing framework, namely kernelized locality sensitive hashing, to accelerate time series similarity search with a series of representative similarity metrics. Experiment results on three large-scale clinical data sets demonstrate the effectiveness of the proposed approach. David C. Kale, Dian Gong, Zhengping Che, Yan Liu 0002, Gérard G. Medioni, Randall C. Wetzel, Patrick Ross |
ICDM | 3 |