VLDB 2026 Research / reviewers in the wild / expert
Huaping Liu 0001
dblp:69/1097-1
· DBLP profile ↗
213ranked-venue papers
44as first author
95since 2021 · last 2026
0000-0002-4042-6044ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 134 · 23 first-author · 57 since 2021Graphics, computer vision, multimedia, augmented reality and games · 50 · 5 first-author · 26 since 2021Applied, interdisciplinary, general and emerging computing · 39 · 12 first-author · 21 since 2021Systems, architecture and hardware · 36 · 8 first-author · 20 since 2021Databases, data management, data science and information retrieval · 8 · 3 first-authorHuman-computer interaction and ubiquitous computing · 6 · 4 first-author · 1 since 2021Computer networks · 3 · 3 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multi-robot coordination for effective anomaly handling in self-driving laboratories
Jiankun Yang, Haibo Lu, Huaping Liu 0001, Wen Gao 0001 |
Eng. Appl. Artif. Intell. | 4 |
| 2026 | Transfer learning-based sparse-reward meta-Q-learning algorithm for active SLAM
Xin Liu 0068, Shuhuan Wen, Zhengzheng Guo, Huaping Liu 0001 |
Expert Syst. Appl. | 4 |
| 2026 | Multimodal Large Language Models for Perception in Autonomous Driving: Architecture, Taxonomy, and ChallengesabstractAutonomous vehicles rely on continuous environmental perception to assess obstacle distribution to ensure safe driving. However, current perception technologies face substantial challenges, particularly under adverse weather conditions and when encountering long-tail scenarios. With the advent of transformer-based attention mechanisms, large language models (LLMs), exemplified by GPT, have exhibited emergent intelligence, offering new possibilities for achieving high-performance perception. This technological advancement has led to the development of multi-modal large language models (MLLMs), which incorporate multi-modal encoders. These models enable a single LLM to process multi-source data while performing advanced understanding and reasoning tasks, enhancing complex environmental perception capabilities. Despite significant progress in MLLMs, there remains a notable gap in systematic research on their optimal application to perception tasks. Therefore, this paper presents a comprehensive survey of recent advancements in MLLM-based perception. First, we introduce the mainstream vision–language perception tasks, widely adopted evaluation metrics, and existing language-enhanced autonomous driving datasets. Next, we outline the general architectural design principles of MLLMs. Subsequently, we provide a taxonomy and indepth analysis of current MLLMs, focusing on three dimensions: input modality, alignment technique, and scene representation, elucidating their underlying implementation paradigms. Finally, we summarize the key challenges and emerging research directions in MLLM-driven perception. This survey aims to facilitate further progress in MLLMs by synthesizing these insights. Xinyu Zhang 0001, Yuchuan Ji, Yanchao Ding, Jialun Yin, Ruizhi Jia, Yijin Xiong, Jun Li 0082, Huaping Liu 0001 |
IEEE Internet Things J. | 12 |
| 2026 | Distributed Resilient Connectivity Maintenance via Composite Control Barrier Functions
Bochen Li, Junwu Li, Lei Song 0005, Dan Huang 0002, Chenguang Yang 0001, Huaping Liu 0001 |
IEEE Internet Things J. | 8 |
| 2026 | Boosting Learning Efficiency in Few-Shot Tasks With Layer-Adaptive PID ControlabstractFew-shot learning seeks to recognize novel classes from limited examples. Model-agnostic meta-learning (MAML), known for its simplicity and flexibility, learns an effective initialization for fast adaptation in data-scarce settings. However, MAML-based methods face challenges when there is a significant distributional shift between training and testing tasks, leading to inefficient learning and poor generalization across domains. In this work, we identify the core issues: inflexible weight update rules and limited adaptive learning capabilities. Instead of focusing solely on better initialization, we aim to enhance the adaptation process. Consequently, we propose a novel Layer-Adaptive Proportional-Integral-Derivative (LA-PID) optimizer integrated into a meta-learning framework. This design incorporates classical control theory, utilizing PID control to dynamically adjust task-specific gains at each network layer. Additionally, the theoretical conditions for optimal hyperparameter initialization and global model convergence are addressed from both control and optimization perspectives. Experiments on benchmark datasets show that LA-PID achieves state-of-the-art performance in few-shot classification, cross-domain, and regression tasks, while requiring fewer training steps. Xinde Li, Zhentong Zhang, Fir Dunkin, Huaping Liu 0001, Zhijun Li 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2026 | Facial 3D Regional Structural Motion Representation Using Lightweight Point Cloud Networks for Micro-Expression RecognitionabstractHuman-computer interaction (HCI) relies on understanding and adapting to users' emotional states. Micro-expressions (MEs), a critical component of emotional perception, are characterized by their spontaneity, rapidity, subtlety, and difficulty to control. They often reveal an individual's true emotions. A comprehensive and detailed representation of motion is necessary to capture the nuances of facial dynamics effectively. Presently, motion representation methods are predominantly confined to 2D analysis within RGB images, overlooking the critical role of facial structure and its movements in conveying emotions. To overcome this limitation, we introduce an innovative facial motion representation that encompasses 3D facial structure, regionalized RGB and structural motion features. Furthermore, we segment the face into eight distinct regions, selecting only the most significant motion points to delineate the primary motion characteristics of each area. To model the interactions among crucial facial motion regions, we employ an advanced, lightweight point cloud and graph convolution network (Lite-Point-GCN). Comprehensive testing on the$\mathrm{CAS(ME)^{3}}$dataset, using leave-one-subject-out (LOSO), demonstrates that our method outperforms existing state-of-the-art methods. Jianqin Yin, Yonghao Dang, Huaping Liu 0001 |
IEEE Trans. Affect. Comput. | 7 |
| 2026 | Trust Under Attack: Environment Perturbation-Based Attacks on Trust-Aware Human-Robot CollaborationabstractIn trust-aware human-robot collaboration (HRC), incorporating human trust into robot decision-making has shown promise in improving collaboration performance. However, most existing research overlooks the potential security vulnerabilities in trust-aware human-robot collaboration methods, which may compromise both collaboration efficiency and human trust. To directly reveal these vulnerabilities and evaluate their impact, this paper proposes an environment perturbation-based trust attack method for human-robot collaboration. The proposed method models the human-robot collaboration process under black-box assumptions and generates adversarial environmental perturbations using the Fast Gradient Sign Method (FGSM). Extensive simulation and real-world experiments show that both traditional and large language model (LLM)-based trust-aware human-robot collaboration methods are vulnerable to such attacks, leading to reductions in collaboration efficiency and human trust. Furthermore, the analysis reveals characteristic phenomena in trust dynamics and system performance under adversarial conditions. These results highlight the importance of considering security vulnerabilities in trust-aware human-robot collaboration. Boyuan Du, Huaping Liu 0001, Yuanlong Yu 0001 |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2026 | BioSynGrasp: Bio-Inspired Structured Reinforcement Learning With Synergetic Exploration and Human Demonstration for Dexterous Grasping of Musculoskeletal RobotabstractMusculoskeletal robots with anthropomorphic structures show great potential for human-like dexterous manipulation. However, their high-dimensional, coupled and nonlinear dynamics pose significant challenges to control and manipulation. Despite recent advances, most existing methods still struggle to achieve flexible multi-object manipulation without long-sequence expert demonstrations or manual parameter tuning based on expert knowledge. To tackle these limitations, a bio-inspired reinforcement learning framework named BioSynGrasp is proposed. It comprises three core components: 1) A neuromuscular controller that integrates shared perception and planning modules with distinct arm and hand synergies sub-networks; 2) A proximal policy optimization-based exploration strategy driven by hand and arm muscle synergies with correlated actuator perturbations; 3) Hand preshapes measured via motion capture provide efficient initialization for functional grasping and reaching with minimal data. BioSynGrasp is validated on a simulated musculoskeletal robot with 63 actuated muscles and on a hardware articulated dexterous hand system. Results demonstrate that it successfully completes 28 multi-object grasping and reaching tasks with an average success rate of 85.9%, outperforming baseline methods in success rate, energy efficiency, position error, and redundant actuator recruitment. These results highlight that integrating bio-inspired motor control with hand preshapes enhances human-like dexterous manipulation and provides insights into human neuromechanics for more adaptable robotic systems. Shanlin Zhong, Huaping Liu 0001, Jiahao Chen 0003 |
IEEE Trans Autom. Sci. Eng. | 6 |
| 2026 | Guest Editorial: Artificial Intelligence Generated Content (AIGC) for Industrial Manufacturing
Huaping Liu 0001, Weiwei Wan, Jason Gu, Valeria Villani, Giulia Pedrielli, Yiannis Aloimonos |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2026 | Semantic Gaussians: Open-Vocabulary Scene Understanding With 3D Gaussian SplattingabstractOpen-vocabulary 3D scene understanding presents a significant challenge in computer vision, with wide-ranging applications in embodied agents and augmented reality systems. Existing methods adopt neural rendering methods as 3D representations and jointly optimize color and semantic features to achieve rendering and scene understanding simultaneously. In this paper, we introduce Semantic Gaussians, a novel open-vocabulary scene understanding approach based on 3D Gaussian Splatting. Our key idea is to distill knowledge from 2D pretrained models to 3D Gaussians. Unlike existing methods, we design a versatile projection approach that maps various 2D semantic features from pre-trained image encoders into a novel semantic component of 3D Gaussians, which is based on spatial relationships and needs no additional training. We further build a 3D semantic network that directly predicts the semantic component from raw 3D Gaussians for fast inference. The quantitative results on ScanNet segmentation and LERF object localization demonstrate the superior performance of our method. Additionally, we explore several applications of Semantic Gaussians, including object part segmentation, instance segmentation, scene editing, and spatiotemporal segmentation, with better qualitative results over 2D and 3D baselines, highlighting its versatility and effectiveness in supporting diverse downstream tasks. Jun Guo 0009, Xiaojian Ma 0001, Huaping Liu 0001, Qing Li 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | CrossRay3D: Geometry and Distribution Guidance for Efficient Multimodal 3D DetectionabstractThe sparse cross-modality detector offers more advantages than its counterpart, the Bird’s-Eye-View (BEV) detector, particularly in terms of adaptability for downstream tasks and computational cost savings. However, existing sparse detectors overlook the quality of token representation, leaving it with a sub-optimal foreground quality and limited performance. In this paper, we identify that the geometric structure preserved and the class distribution are the key to improving the performance of the sparse detector, and propose a Sparse Selector (SS). The core module of SS is Ray-Aware Supervision (RAS), which preserves rich geometric information during the training stage, and Class-Balanced Supervision, which adaptively reweights the salience of class semantics, ensuring that tokens associated with small objects are retained during token sampling. Thereby, outperforming other sparse multi-modal detectors in the representation of tokens. Additionally, we design Ray Positional Encoding (Ray PE) to address the distribution differences between the LiDAR modality and the image. Finally, we integrate the aforementioned module into an end-to-end sparse multi-modality detector, dubbed CrossRay3D. Experiments show that, on the challenging nuScenes benchmark, CrossRay3D achieves state-of-the-art performance with 72.4% mAP and 74.7% NDS, while running$1.84\times $faster than other leading methods. Moreover, CrossRay3D demonstrates strong robustness even in scenarios where LiDAR or camera data are partially or entirely missing. The code is available onhttps://github.com/xuehaipiaoxiang/CrossRay3D Huiming Yang, Wenzhuo Liu, Yicheng Qiao, Lei Yang 0060, Xianzhu Zeng, Li Wang 0092, Zhiwei Li 0011, Zijian Zeng 0001, Zhiying Jiang, Huaping Liu 0001, Kunfeng Wang |
IEEE Trans. Intell. Transp. Syst. | 10 |
| 2026 | Weighted Fusion of Classifiers With Approximate Reasoning and Reliability Evaluation for Multisource Information FusionabstractClassifiers fusion can be seen as a kind of multisource information fusion (MSIF), and classifiers fusion based on Dempster–Shafer (DS) evidence theory is an effective approach to improve the accuracy of classification tasks. However, different classifiers usually exhibit varying performances, making it challenging to achieve enhanced classification accuracy through direct fusion. Simultaneously, when the frame of discernment (FoD) of the target class expands, the number of focal elements involved in the fusion increases, resulting in a rapid growth in computational complexity. To enhance the classification performance while reducing the time cost of fusion, a novel weighted fusion of classifiers method based on approximate reasoning and reliability evaluation (WFC-AR-RE) is proposed in this article. Specifically, at first, the key focal elements are determined based on the outputs of classifiers, and an approximate basic belief assignment (BBA) is generated. Subsequently, the validation set is utilized to evaluate the performance of each classifier, thus obtaining the self-reliability of each BBA. Afterward, a novel divergence measure is introduced to quantify the discrepancy between BBAs, determining the relative reliability of each BBA. Finally, the fusion weight of each BBA is derived from its self-reliability and relative reliability, and Dempster’s rule is applied to combine the weighted BBA. The proposed WFC-AR-RE algorithm is applied to the MSIF system, and its effectiveness is demonstrated on 12 public datasets. Kezhu Zuo, Xinde Li, Huaping Liu 0001, Yilin Dong 0001, Jean Dezert, Tao Shen 0004, Shuzhi Sam Ge |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2025 | Bootstrapping Heterogeneous Graph Representation Learning via Large Language Models: A Generalized ApproachabstractGraph representation learning methods are highly effective in handling complex non-Euclidean data by capturing intricate relationships and features within graph structures. However, traditional methods face challenges when dealing with heterogeneous graphs that contain various types of nodes and edges due to the diverse sources and complex nature of the data. Existing heterogeneous graph neural networks (HGNNs) have shown promising results but require prior knowledge of node and edge types and unified node feature formats, which limits their applicability. Recent advancements in graph representation learning using large language models (LLMs) offer new solutions by integrating LLMs' data processing capabilities, enabling the alignment of various graph representations. Nevertheless, these methods often overlook heterogeneous graph data and require extensive preprocessing. To address these limitations, we propose an LLM-enhanced Heterogeneous Graph Neural Network (LHGNN). LHGNN leverages the strengths of both LLM and GNN, allowing for the processing of graph data with any format and type of nodes and edges without the need for type information or special preprocessing. LHGNN employs LLM to automatically summarize and classify different data formats and types, aligns node features, and uses a specialized GNN for targeted learning, thus obtaining effective graph representations for downstream tasks. Theoretical analysis and experimental validation have demonstrated the effectiveness of our method. Hang Gao 0004, Fengge Wu, Changwen Zheng, Junsuo Zhao, Huaping Liu 0001 |
AAAI | 6 |
| 2025 | MeshGen: Generating PBR Textured Mesh with Render-Enhanced Auto-Encoder and Generative Data AugmentationabstractIn this paper, we introduce MeshGen, an advanced image-to-3D pipeline that generates high-quality 3D meshes with detailed geometry and physically based rendering (PBR) textures. Addressing the challenges faced by existing 3D native diffusion models, such as suboptimal auto-encoder performance, limited controllability, poor generalization, and inconsistent image-based PBR texturing, Mesh-Gen employs several key innovations to overcome these limitations. We pioneer a render-enhanced point-to-shape auto-encoder that compresses meshes into a compact latent space by designing perceptual optimization with ray-based regularization. This ensures that the 3D shapes are accurately represented and reconstructed to preserve geometric details within the latent space. To address data scarcity and image-shape misalignment, we further propose geometric augmentation and generative rendering augmentation techniques, which enhance the model’s controllability and gen-eralization ability, allowing it to perform well even with limited public datasets. For the texture generation, Mesh-Gen employs a reference attention-based multi-view Con-trolNet for consistent appearance synthesis. This is further complemented by our multi-view PBR decomposer that estimates PBR components and a UV inpainter that fills invisible areas, ensuring a seamless and consistent texture across the 3D mesh. Our extensive experiments demonstrate that MeshGen largely outperforms previous methods in both shape and texture generation, setting a new standard for the quality of 3D meshes generated with PBR textures. Zilong Chen, Yikai Wang 0001, Wenqiang Sun, Feng Wang 0034, Huaping Liu 0001 |
CVPR | 6 |
| 2025 | MMTL-UniAD: A Unified Framework for Multimodal and Multi-Task Learning in Assistive Driving PerceptionabstractAdvanced driver assistance systems require a comprehensive understanding of the driver’s mental/physical state and traffic context but existing works often neglect the potential benefits of joint learning between these tasks. This paper proposes MMTL-UniAD, a unified multi-modal multitask learning framework that simultaneously recognizes driver behavior (e.g., looking around, talking), driver emotion (e.g., anxiety, happiness), vehicle behavior (e.g., parking, turning), and traffic context (e.g., traffic jam, traffic smooth). A key challenge is avoiding negative transfer between tasks, which can impair learning performance. To address this, we introduce two key components into the framework: one is the multi-axis region attention network to extract global context-sensitive features, and the other is the dual-branch multimodal embedding to learn multi-modal embeddings from both task-shared and task-specific features. The former uses a multi-attention mechanism to extract task-relevant features, mitigating negative transfer caused by task-unrelated features. The latter employs a dual-branch structure to adaptively adjust task-shared and task-specific parameters, enhancing cross-task knowledge transfer while reducing task conflicts. We assess MMTL-UniAD on the AIDE dataset, using a series of ablation studies, and show that it outperforms state-of-the-art methods across all four tasks. The code is available on https://github.com/Wenzhuo-Liu/MMTL-UniAD. Wenzhuo Liu, Wenshuo Wang 0001, Yicheng Qiao, Qiannan Guo, Jiayin Zhu, Zilong Chen, Huiming Yang, Zhiwei Li 0011, Tiao Tan, Huaping Liu 0001 |
CVPR | 12 |
| 2025 | ProcWorld: Benchmarking Large Model Planning in Reachability-Constrained EnvironmentsabstractWe introduce ProcWORLD, a large-scale benchmark for partially observable embodied spatial reasoning and long-term planning with large language models (LLM) and vision language models (VLM).ProcWORLD features a wide range of challenging embodied navigation and object manipulation tasks, covering 16 task types, 5,000 rooms, and over 10 million evaluation trajectories with diverse data distribution.ProcWORLD supports configurable observation modes, ranging from text-only descriptions to vision-only observations.It enables text-based actions to control the agent following language instructions.ProcWORLD has presented significant challenges for LLMs and VLMs: (1) active information gathering given partial observations for disambiguation; (2) simultaneous localization and decision-making by tracking the spatio-temporal state-action distribution; (3) constrained reasoning with dynamic states subject to physical reachability.Our extensive evaluation of 15 foundation models and 5 reasoning algorithms (with over 1 million rollouts) indicates larger models perform better.However, ProcWORLD remains highly challenging for existing state-of-the-art models and in-context learning methods due to constrained reachability and the need of combinatorial spatial reasoning. Xinghang Li, Zhengshen Zhang, Jirong Liu, Xiao Ma 0006, Hanbo Zhang, Tao Kong, Huaping Liu 0001 |
EMNLP | 8 |
| 2025 | LLM Enhancers for GNNs: An Analysis from the Perspective of Causal Mechanism IdentificationabstractThe use of large language models (LLMs) as feature enhancers to optimize node representations, which are then used as inputs for graph neural networks (GNNs), has shown significant potential in graph representation learning. However, the fundamental properties of this approach remain underexplored. To address this issue, we propose conducting a more in-depth analysis of this issue based on the interchange intervention method. First, we construct a synthetic graph dataset with controllable causal relationships, enabling precise manipulation of semantic relationships and causal modeling to provide data for analysis. Using this dataset, we conduct interchange interventions to examine the deeper properties of LLM enhancers and GNNs, uncovering their underlying logic and internal mechanisms. Building on the analytical results, we design a plug-and-play optimization module to improve the information transfer between LLM enhancers and GNNs. Experiments across multiple datasets and models validate the proposed module. Hang Gao 0004, Fengge Wu, Junsuo Zhao, Changwen Zheng, Huaping Liu 0001 |
ICML | 6 |
| 2025 | Sparse Spectral Training and Inference on Euclidean and Hyperbolic Neural NetworksabstractThe growing demands on GPU memory posed by the increasing number of neural network parameters call for training approaches that are more memory-efficient. Previous memory reduction training techniques, such as Low-Rank Adaptation (LoRA) and ReLoRA, face challenges, with LoRA being constrained by its low-rank structure, particularly during intensive tasks like pre-training, and ReLoRA suffering from saddle point issues. In this paper, we propose Sparse Spectral Training (SST) to optimize memory usage for pre-training. SST updates all singular values and selectively updates singular vectors through a multinomial sampling method weighted by the magnitude of the singular values. Furthermore, SST employs singular value decomposition to initialize and periodically reinitialize low-rank parameters, reducing distortion relative to full-rank training compared to other low-rank methods. Through comprehensive testing on both Euclidean and hyperbolic neural networks across various tasks, SST demonstrates its ability to outperform existing memory reduction training methods and is comparable to full-rank training in various cases. On LLaMA-1.3B, with only 18.7% of the parameters trainable compared to full-rank training (using a rank equivalent to 6% of the embedding dimension), SST reduces the perplexity gap between other low-rank methods and full-rank training by 97.4%. This result highlights SST as an effective parameter-efficient technique for model pre-training. Jialin Zhao 0004, Yingtao Zhang, Xinghang Li, Huaping Liu 0001, Carlo V. Cannistraci |
ICML | 4 |
| 2025 | Learn to Think: Bootstrapping LLM Logic Through Graph Representation LearningabstractLarge Language Models (LLMs) have achieved remarkable success across various domains. However, they still face significant challenges, including high computational costs for training and limitations in solving complex reasoning problems. Although existing methods have extended the reasoning capabilities of LLMs through structured paradigms, these approaches often rely on task-specific prompts and predefined reasoning processes, which constrain their flexibility and generalizability. To address these limitations, we propose a novel framework that leverages graph learning to enable more flexible and adaptive reasoning capabilities for LLMs. Specifically, this approach models the reasoning process of a problem as a graph and employs LLM-based graph learning to guide the adaptive generation of each reasoning step. To further enhance the adaptability of the model, we introduce a Graph Neural Network (GNN) module to perform representation learning on the generated reasoning process, enabling real-time adjustments to both the model and the prompt. Experimental results demonstrate that this method significantly improves reasoning performance across multiple tasks without requiring additional training or task-specific prompt design. Code can be found in https://github.com/zch65458525/L2T. Hang Gao 0004, Junsuo Zhao, Fengge Wu, Changwen Zheng, Huaping Liu 0001 |
IJCAI | 7 |
| 2025 | A Deep Reinforcement Learning-based Autonomous Robotic Operation Framework for Blood Gas AnalyzersabstractTo tackle the technical challenges of precisely aligning and inserting test tubes into sampling needles, this paper proposes an autonomous operation method for blood gas analyzers based on reinforcement learning. To simplify the complexity of the alignment and insertion task, it is decomposed into two independent subtasks, which are learned in a distributed manner and executed sequentially to complete the overall needle insertion process. To enable accurate perception of dynamic environmental states, a reinforcement learning model is developed that defines the state space, action space, and multi-task reward functions. Building on this framework, we enhance the Actor-Critic structure of the Proximal Policy Optimization (PPO) algorithm by introducing a dual-network architecture that integrates Long Short-Term Memory (LSTM) and Wavelet Transform Convolution (WTC) networks. This enhancement significantly improves both learning efficiency and policy stability. Extensive comparative experiments demonstrate that the proposed LSTM-WTC-PPO algorithm achieves a high success rate, stable convergence, and efficient policy optimization. Haiyang Jiang 0018, Tong Liu 0014, Huaping Liu 0001, Visakan Kadirkamanathan, Kai Wang 0003 |
INDIN | 3 |
| 2025 | TEM3-Learning: Time-Efficient Multimodal Multi-Task Learning for Advanced Assistive DrivingabstractMulti-task learning (MTL) can advance assistive driving by exploring inter-task correlations through shared representations. However, existing methods face two critical limitations: single-modality constraints limiting comprehensive scene understanding and inefficient architectures impeding real-time deployment. This paper proposes TEM3-Learning (Time-Efficient Multimodal Multi-task Learning), a novel framework that jointly optimizes driver emotion recognition, driver behavior recognition, traffic context recognition, and vehicle behavior recognition through a two-stage architecture. The first component, the mamba-based multi-view temporal-spatial feature extraction subnetwork (MTS-Mamba), introduces a forward-backward temporal scanning mechanism and global-local spatial attention to efficiently extract low-cost temporal-spatial features from multi-view sequential images. The second component, the MTL-based gated multimodal feature integrator (MGMI), employs task-specific multi-gating modules to adaptively highlight the most relevant modality features for each task, effectively alleviating the negative transfer problem in MTL. Evaluation on the AIDE dataset, our proposed model achieves state-of-the-art accuracy across all four tasks, maintaining a lightweight architecture with fewer than 6 million parameters and delivering an impressive 142.32 FPS inference speed. Rigorous ablation studies further validate the effectiveness of the proposed framework and the independent contributions of each module. The code is available on https://github.com/Wenzhuo-Liu/TEM3-Learning. Wenzhuo Liu, Yicheng Qiao, Qiannan Guo, Zilong Chen, Meihua Zhou, Zhiwei Li 0011, Huaping Liu 0001, Wenshuo Wang 0001 |
IROS | 10 |
| 2025 | AssistantX: An LLM-Powered Proactive Assistant in Collaborative Human-Populated EnvironmentsabstractCurrent service robots suffer from limited natural language communication abilities, heavy reliance on predefined commands, ongoing human intervention, and, most notably, a lack of proactive collaboration awareness in human-populated environments. This results in narrow applicability and low utility. In this paper, we introduce AssistantX, an LLM-powered proactive assistant designed for autonomous operation in real-world scenarios with high accuracy. AssistantX employs a multi-agent framework consisting of 4 specialized LLM agents, each dedicated to perception, planning, decision-making, and reflective review, facilitating advanced inference capabilities and comprehensive collaboration awareness, much like a human assistant by your side. We built a dataset of 210 real-world tasks to validate AssistantX, which includes instruction content and status information on whether relevant personnel are available. Extensive experiments were conducted in both text-based simulations and a real office environment over the course of a month and a half. Our experiments demonstrate the effectiveness of the proposed framework, showing that AssistantX can reactively respond to user instructions, actively adjust strategies to adapt to contingencies, and proactively seek assistance from humans to ensure successful task completion. More details and videos can be found at https://assistantx-agent.github.io/AssistantX/. Yongchang Li, Di Guo 0002, Huaping Liu 0001 |
IROS | 5 |
| 2025 | A Patch-Based Transformer Method for Electrical Capacitance Tomography Image ReconstructionabstractElectrical capacitance tomography (ECT) is a contactless and non-invasive imaging technique, which visualizes the internal permittivity distribution around a region utilizing boundary capacitance measurements. It has been widely used in the fields of object classification, tactile sensing and multiphase flows monitoring. However, due to the inherent nonlinearity and ill-conditioned nature of the ECT inverse problem, its practical implementation remains limited by challenges in the image reconstruction accuracy. To tackle the above problems, we propose a patch-based transformer method (PT) for an accurate reconstruction of ECT images. Specifically, the complex capacitance-to-image mapping is systematically decoupled into the capacitance-to-patch feature extraction and patch-to-image reconstruction, enabling more efficient and accurate permittivity distribution recovery through localized feature learning and global context integration. Additionally, a simulation ECT dataset for objects with varying sizes and positions is established. Duanpeng Shi, Huaping Liu 0001, Di Guo 0002 |
IROS | 3 |
| 2025 | A Novel Terrain Classification System with Planar ECT SensorabstractTerrain classification is crucial for robotic navigation especially in unknown environment. Existing terrain classification methods usually have high requirements for environment conditions and robot motions, making them challenging to apply to real-world scenarios. In this paper, we develop a novel terrain classification system with the planar electrical capacitance tomography (ECT) sensor, which provides a non-contact, real-time, and cost-effective way for terrain classification. Specifically, we design a planar ECT sensor and integrate it at the bottom of a mobile robot. The proposed system leverages the collected capacitance measurements to reflect the inherent differences in dielectric permittivity across various terrain types. And a multilayer perception networks is used to fuse the collected capacitance and IMU measurements for classification. Additionally, a large scale ECT dataset including 10 different types of terrains is collected with the proposed system. Extensive experiments are conducted demonstrating the effectiveness and robustness of the proposed system. Wenju Yang, Duanpeng Shi, Wuqiang Yang, Tengchen Sun, Huaping Liu 0001, Di Guo 0002 |
IROS | 5 |
| 2025 | SAMOccNet:Refined SAM-based surrounding semantic occupancy perception for autonomous driving
Qifan Tan, Wenzhuo Liu, Han Bi, Lei Yang 0060, Yicheng Qiao, Zhuo Zhao, Yanhuan Jiang, Qiannan Guo, Huaping Liu 0001, Zhiwei Li 0011 |
Neurocomputing | 10 |
| 2025 | CPL-SLAM: Centralized Collaborative Multirobot Visual-Inertial SLAM Using Point-and-Line FeaturesabstractTraditional visual-inertial Simultaneous Localization and Mapping (SLAM) systems predominantly rely on feature point matching from a single robot to realize the robot pose estimation and environment map construction. However, in complex scenarios, these traditional systems struggle with issues, such as tracking failures due to illumination changes, rapid movements, and low-texture environments, and they perform poorly in terms of mapping efficiency and global consistency. To address these challenges, we propose a centralized collaborative SLAM system that employs both point and line features for tracking in the robot and map fusion in the cloud. The proposed system leverages the fusion of point and line features across all instances in the process, which allows our method to achieve higher localization accuracy in structured, low-texture scenes. With the aid of classifying gravity-aligned vertical lines and spatial parallel lines, the proposed system can deliver faster and more accurate odometry in complex scenes. Furthermore, we developed intrarobot and interrobot loop closure detection methods based on point and line features, generating a globally consistent sparse point cloud and structured scene map in the cloud. Our method is able to build richer maps while improving accuracy compared to existing methods. Experimental results on public datasets and in real-world environments show that, compared to existing advanced methods, our approach demonstrates better performance. Xin Liu 0068, Shuhuan Wen, Huaping Liu 0001, F. Richard Yu |
IEEE Internet Things J. | 3 |
| 2025 | Risk-Aware Decision Making for Human-Robot Collaboration With Effective CommunicationabstractHuman-robot collaboration is crucial for integrating robots into intelligent manufacturing (IM). However, a significant challenge is rational decision-making for the human-cyber-physical system (HCPS) in IM to enhance cognitive limits of human operators and overcome the potential irrationality. Since humans dominate the collaboration in IM, it is essential to address two critical issues: determining the next task for the robot and deciding whether the human operator should be informed. We propose a risk-aware decision-making framework for task allocation and human-robot interaction (HRI) to achieve a balance between the autonomy level of the human operator and task efficiency. To quantify the efficiency risk, we utilize conditional value-at-risk (CVaR) considering the uncertainty of human operators. We then obtain the optimal task allocation and selection for the robot by minimizing the efficiency risk. We also establish a necessary collection of tasks that must be performed by the human operator. Furthermore, we develop two criteria to quantify the necessity of explicit HRI. Experiments with a real mechanical arm platform demonstrate that our methods can enhance human-robot collaboration (HRC), reduce the need for extensive communication, and grant human operators greater execution freedom.Note to Practitioners—This research is motivated by the fact that the imperfect information sharing between human operators and robots in IM. It results in the inability of human operators to accurately comprehend the robot’s capability constraints and the robot’s limitations in perceiving the current or next action of the human operator. As a result, the effectiveness of HRC then inevitably decreases. Our proposed method introduces the efficiency risk to quantify the need of robot active intervention in HCPS. This approach aids in determining when and what to communicate with the human operator, thereby empirically ensuring efficiency and providing the human operator with a greater degree of freedom. Experimental studies suggest that the efficiency risk can serve as a new metric for balancing the amount of human-robot communication and the improvement of efficiency. Our research can also find utility in human-centric scenarios where communication between the human operator and the robot is constrained or costly. Yang Li 0029, Hongbo Li 0001, Huaping Liu 0001 |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2025 | Trust-Aware Human-Robot Fusion Decision-Making for Emergency Indoor PatrollingabstractTrust plays a crucial role in decision-making during human-robot collaboration, particularly in emergency scenarios where it becomes more susceptible due to dynamic factors. Misalignment of human-robot trust significantly hampers the efficiency of the collaboration. Therefore, it is imperative to establish effective human-robot interaction and decision support mechanisms that mitigate biases in human confidence levels regarding the robot’s capabilities in dynamic environments. Additionally, online correction of human-robot trust based on human behavioral feedback is vital. This paper focuses on a specific type of emergency task scenario, specifically, indoor human-robot collaborative patrolling in the event of sudden power outages. We propose a trust model based on linear Gaussian and sparse Gaussian processes (sparse GP). We also employ Monte Carlo Tree Search (MCTS) method to determine the robot’s optimal fusion decision-making. Through VR-based human-robot collaborative experiments, we ascertain that the robot prioritizes enhancing human-robot trust in emergency scenarios to mitigate the long-term costs of human-robot collaboration.Note to Practitioners—The primary motivation of this paper lies in tackling the issue of decreased collaboration efficiency in human-robot cooperation, stemming from irrational human decision-making, a problem exacerbated in emergency scenarios where comprehending all available information proves challenging. Grounded in the realm of human-robot trust, this study evaluates the temporal and substantive aspects of human-robot interactions contingent upon the degree of trust. Moreover, it adapts the fusion decision-making process within the human-robot team in accordance with the decision-making strategies adopted by humans at varying trust levels. The optimized decisions of the human-robot ensemble are conveyed to the human participant via decision support. This approach ensures the maintenance of trust levels, facilitating the acceptance of decision support that may seem intuitively irrational but holds objective superiority. Furthermore, it mitigates redundant human-robot interactions once a sufficient level of trust is established. Yang Li 0029, Jiaxin Xu, Di Guo 0002, Huaping Liu 0001 |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2025 | PDDepth: Pose Decoupled Monocular Depth Estimation for Roadside Perception SystemabstractAccurate depth information is crucial for roadside perception in Cooperative Vehicle Infrastructure Systems. Beyond existing radar and LiDAR solutions, monocular depth estimation using surveillance cameras is emerging as a superior approach due to its cost-effectiveness and dense depth output. Unlike onboard cameras, roadside cameras are relatively fixed in position. Many existing monocular depth estimation methods, which do not independently model camera pose, tend to overfit to training data and produce suboptimal results when confronted with slight variations in camera poses, which may be caused by external forces within the same camera or across different cameras. To address this issue, a pose decoupled monocular depth estimation method specifically designed for roadside perception systems is proposed. This method separates depth estimation into two components: a pose-dependent modeling portion that recovers ground depth based on the current camera pose, and a pose-agnostic portion that estimates pixel height relative to the ground plane. Additionally, a knowledge distillation framework is introduced to improve the robustness of the proposed method against variations in roadside cameras. To validate the method, we propose the first open source dataset for roadside monocular depth estimation, DAIR-MDE, and a roadside instance segmentation dataset, DAIR-Ins, both derived from the DAIR dataset. The proposed method demonstrates significant advances over the state-of-the-art methods on DAIR-MDE. The proposed dataset and source code are publicly available athttps://github.com/441599828/PDDepth. Huanan Wang, Xinyu Zhang 0001, Zhengxian Chen, Jun Li 0082, Huaping Liu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | MIPD: A Multi-Sensory Interactive Perception Dataset for Embodied Intelligent DrivingabstractDuring the process of driving, humans usually rely on multiple senses to gather information and make decisions. Analogously, in order to achieve embodied intelligence in autonomous driving, it is essential to integrate multidimensional sensory information in order to facilitate interaction with the environment. However, the current multi-modal fusion sensing schemes often neglect these additional sensory inputs, hindering the realization of fully autonomous driving. This paper considers multi-sensory information and proposes a multi-modal interactive perception dataset named MIPD, enabling expanding the current autonomous driving algorithm framework, for supporting the research on embodied intelligent driving. In addition to the conventional camera, lidar, and 4D radar data, our dataset incorporates multiple sensor inputs including sound, light intensity, vibration intensity and vehicle speed to enrich the dataset comprehensiveness. Comprising 126 consecutive sequences, many exceeding twenty seconds, MIPD features over 8,500 meticulously synchronized and annotated frames. Moreover, it encompasses many challenging scenarios, covering various road and lighting conditions. The dataset has undergone thorough experimental validation, producing valuable insights for the exploration of next-generation autonomous driving frameworks. Data, development kit and more details will be available athttps://github.com/BUCT-IUSRC/Dataset__MIPD Zhiwei Li 0011, Tingzhen Zhang, Meihua Zhou, Dandan Tang, Wenzhuo Liu, Qiaoning Yang, Tianyu Shen, Kunfeng Wang, Huaping Liu 0001 |
IEEE Trans. Intell. Transp. Syst. | 10 |
| 2025 | UMD-Net: A Unified Multi-Task Assistive Driving Network Based on Multimodal FusionabstractIn recent years, researchers have focused on identifying tasks related to driver state, traffic environment, and others to enhance the safety of autonomous driving assistance systems. However, current research on these tasks is conducted independently, neglecting the interconnections between the driver, traffic environment, and vehicle. In this paper, we propose a Unified Multi-task Assistive Driving Network Based on Multimodal Fusion (UMD-Net), the first unified model capable of recognizing four tasks simultaneously by utilizing multimodal data: driver behavior recognition, driver emotion recognition, traffic context recognition, and vehicle behavior recognition. In order to better enhance the synergistic effects between multiple tasks, we designed the position-sensitive multi-directional attention feature extraction subnetwork and recursive dynamic feature fusion module. The former captures the key features of multi-view images by different directions of attention mechanism to improve the generalization of the model across multiple tasks. The latter dynamically adjusts the fusion weight according to the multimodal features to enhance the representation ability of important features in multi-task learning. Our model was evaluated on the public dataset AIDE, achieving the best performance across all four tasks and a high accuracy of 95.31% in the traffic context recognition task, demonstrating the superiority of our approach. The code is available on https://github.com/Wenzhuo-Liu/UMD-Net. Wenzhuo Liu, Yicheng Qiao, Zhiwei Li 0011, Wenshuo Wang 0001, Wei Zhang 0012, Jiayin Zhu, Yanhuan Jiang, Li Wang 0092, Hong Wang 0014, Huaping Liu 0001, Kunfeng Wang |
IEEE Trans. Intell. Transp. Syst. | 10 |
| 2025 | Semantics-Aware Hierarchical Decision Framework for Embodied Visual Room RearrangementabstractIn embodied visual room rearrangement, the agent needs to recover the scene state to the goal state through interacting with the environment based on the egocentric visual observations after the locations and states of some objects are changed. It has important application potential in the field of robotics. This task is challenging in visual perception, scene understanding, and action execution. Existing methods do not take full advantage of the semantic information and spatial relationship of objects in the scene perception and understanding process. To tackle the challenges and shortcomings of the current methods, we build a hierarchical decision framework based on the pretrained semantic scene representation and transformer-based scene memory to solve this task. The results in the unseen scenes demonstrate the effectiveness of the proposed model compared with other methods. Xinzhu Liu, Di Guo 0002, Huaping Liu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2025 | Goal-Conditioned Hierarchical Reinforcement Learning With High-Level Model ApproximationabstractHierarchical reinforcement learning (HRL) exhibits remarkable potential in addressing large-scale and long-horizon complex tasks. However, a fundamental challenge, which arises from the inherently entangled nature of hierarchical policies, has not been understood well, consequently compromising the training stability and exploration efficiency of HRL. In this article, we propose a novel HRL algorithm, high-level model approximation (HLMA), presenting both theoretical foundations and practical implementations. In HLMA, a Planner constructs an innovative high-level dynamic model to predict the -step transition of the Controller in a subtask. This allows for the estimation of the evolving performance of the Controller. At low level, we leverage the initial state of each subtask, transforming absolute states into relative deviations by a designed operator as Controller input. This approach facilitates the reuse of subtask domain knowledge, enhancing data efficiency. With this designed structure, we establish the local convergence of each component within HLMA and subsequently derive regret bounds to ensure global convergence. Abundant experiments conducted on complex locomotion and navigation tasks demonstrate that HLMA surpasses other state-of-the-art single-level RL and HRL algorithms in terms of sample efficiency and asymptotic performance. In addition, thorough ablation studies validate the effectiveness of each component of HLMA. Yu Luo 0021, Tianying Ji, Fuchun Sun 0001, Huaping Liu 0001, Jianwei Zhang 0001, Mingxuan Jing, Wenbing Huang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2025 | Self-Supervised 3-D Semantic Representation Learning for Vision-and-Language NavigationabstractIn vision-and-language navigation (VLN) tasks, most current methods primarily utilize RGB images, overlooking the rich 3-D semantic data inherent to environments. To rectify this, we introduce a novel VLN framework that integrates 3-D semantic information into the navigation process. Our approach features a self-supervised training scheme that incorporates voxel-level 3-D semantic reconstruction to create a detailed 3-D semantic representation. A key component of this framework is a pretext task focused on region queries, which determines the presence of objects in specific 3-D areas. Following this, we devise an long short-term memory (LSTM)-based navigation model that is trained using our 3-D semantic representations. To maximize the utility of these 3-D semantic representations, we implement a cross-modal distillation strategy. This strategy encourages the RGB model's outputs to emulate those from the 3-D semantic feature network, enabling the concurrent training of both branches to merge RGB and 3-D semantic data effectively. Comprehensive evaluations on both the R2R and R4R datasets reveal that our method significantly enhances performance in VLN tasks. Sinan Tan, Kuankuan Sima, Dunzheng Wang, Mengmeng Ge 0002, Di Guo 0002, Huaping Liu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2025 | iLoc: An Adaptive, Efficient, and Robust Visual Localization SystemabstractIn this article, we introduceiLoc, an innovative visual localization system designed to enhance the autonomy and adaptability of robotic agents in long-term and large-scale applications.iLocspecializes in: 1) extracting stable and consistent descriptors for place recognition, unaffected by changes in viewpoint and illumination; 2) performing swift and precise global relocalization to establish a robot's position within a large and complex environment; and 3) generating real-time tracking trajectories aligned with reference maps, ensuring continual orientation within known spaces. Distinctively,iLocincorporates a transformer-based learning module and an attention-enhanced recognition approach, enabling it to adapt to diverse environmental and viewpoint conditions.iLocleverages a coarse-to-fine global feature matching technique for enhanced localization and integrates robust state estimation combining visual odometry and loop closures through local refinement and pose graph optimization.iLocdemonstrates remarkable proficiency in place recognition, achieving localization over distances of up to 2 km within 0.5 s with average accuracy at 1 m. It maintains stable localization accuracy, even under variable conditions. Its versatile design allows integration across various environments, significantly broadening the scope of universal localization capabilities in robotics.iLocrepresents a substantial step forward in visual-based localization systems, delivering unparalleled speed and accuracy in place recognition. Its ability to adapt and respond to diverse environmental stimuli marks it as a crucial tool in advancing the field of robotic localization. Peng Yin 0001, Jing Wang 0193, Ruohai Ge, Jianmin Ji, Yeping Hu, Huaping Liu 0001, Jianda Han |
IEEE Trans. Robotics | 7 |
| 2024 | Rethinking Causal Relationships Learning in Graph Neural NetworksabstractGraph Neural Networks (GNNs) demonstrate their significance by effectively modeling complex interrelationships within graph-structured data. To enhance the credibility and robustness of GNNs, it becomes exceptionally crucial to bolster their ability to capture causal relationships. However, despite recent advancements that have indeed strengthened GNNs from a causal learning perspective, conducting an in-depth analysis specifically targeting the causal modeling prowess of GNNs remains an unresolved issue. In order to comprehensively analyze various GNN models from a causal learning perspective, we constructed an artificially synthesized dataset with known and controllable causal relationships between data and labels. The rationality of the generated data is further ensured through theoretical foundations. Drawing insights from analyses conducted using our dataset, we introduce a lightweight and highly adaptable GNN module designed to strengthen GNNs' causal learning capabilities across a diverse range of tasks. Through a series of experiments conducted on both synthetic datasets and other real-world datasets, we empirically validate the effectiveness of the proposed module. The codes are available at https://github.com/yaoyao-yaoyao-cell/CRCG. Hang Gao 0004, Chengyu Yao, Jiangmeng Li, Lingyu Si, Fengge Wu, Changwen Zheng, Huaping Liu 0001 |
AAAI | 8 |
| 2024 | Text-to-3D using Gaussian SplattingabstractAutomatic text-to-3D generation that combines Score Distillation Sampling (SDS) with the optimization of volume rendering has achieved remarkable progress in synthesizing realistic 3D objects. Yet most existing text-to-3D methods by SDS and volume rendering suffer from inaccurate geometry, e.g., the Janus issue, since it is hard to explicitly integrate 3D priors into implicit 3D representations. Besides, it is usually time-consuming for them to generate elaborate 3D models with rich colors. In response, this paper proposes GSGEN, a novel method that adopts Gaussian Splatting, a recent state-of-the-art representation, to text-to-3D generation. GSGEN aims at generating high-quality 3D objects and addressing existing shortcomings by exploiting the explicit nature of Gaussian Splatting that enables the incorporation of 3D prior. Specifically, our method adopts a progressive optimization strategy, which includes a geometry optimization stage and an appearance refinement stage. In geometry optimization, a coarse representation is established under 3D point cloud diffusion prior along with the ordinary 2D SDS optimization, ensuring a sensible and 3D-consistent rough shape. Subsequently, the obtained Gaussians undergo an iterative appearance refinement to enrich texture details. In this stage, we increase the number of Gaussians by compactness-based densification to enhance continuity and improve fidelity. With these designs, our approach can generate 3D assets with delicate details and accurate geometry. Extensive evaluations demonstrate the effectiveness of our method, especially for capturing high-frequency components. Our code is available at https://github.com/gsgen3d/gsgen. Zilong Chen, Feng Wang 0034, Yikai Wang 0001, Huaping Liu 0001 |
CVPR | 4 |
| 2024 | GaussianEditor: Swift and Controllable 3D Editing with Gaussian Splattingabstract3D editing plays a crucial role in many areas such as gaming and virtual reality. Traditional 3D editing methods, which rely on representations like meshes and point clouds, often fall short in realistically depicting complex scenes. On the other hand, methods based on implicit 3D representations, like Neural Radiance Field (NeRF), render complex scenes effectively but suffer from slow processing speeds and limited control over specific scene areas. In response to these challenges, our paper presents GaussianEditor, the first 3D editing algorithm based on Gaussian Splatting (GS), a novel 3D representation. GaussianEditor enhances precision and control in editing through our proposed Gaussian semantic tracing, which traces the editing target throughout the training process. Additionally, we propose Hierarchical Gaussian splatting (HGS) to achieve stabilized and fine results under stochastic generative guidance from 2D diffusion models. We also develop editing strategies for efficient object removal and integration, a challenging task for existing methods. Our comprehensive experiments demonstrate GaussianEditor's superior control, effective, and efficient performance, marking a significant advancement in 3D editing. Zilong Chen, Chi Zhang 0007, Feng Wang 0034, Yikai Wang 0001, Zhongang Cai, Lei Yang 0059, Huaping Liu 0001, Guosheng Lin |
CVPR | 9 |
| 2024 | EgoThink: Evaluating First-Person Perspective Thinking Capability of Vision-Language ModelsabstractVision-language models (VLMs) have recently shown promising results in traditional downstream tasks. Evaluation studies have emerged to assess their abilities, with the majority focusing on the third-person perspective, and only a few addressing specific tasks from the first-person per-spective. However, the capability of VLMs to “think” from a first-person perspective, a crucial attribute for advancing autonomous agents and robotics, remains largely unexplored. To bridge this research gap, we introduce EgoThink, a novel visual question-answering benchmark that encompasses six core capabilities with twelve detailed dimensions. The benchmark is constructed using selected clips from ego-centric videos, with manually annotated question-answer pairs containing first-person information. To comprehensively assess VLMs, we evaluate twenty-one popular VLMs on EgoThink. Moreover, given the open-ended format of the answers, we use GPT-4 as the automatic judge to compute single-answer grading. Experimental results indicate that although GPT-4V leads in numerous dimensions, all evaluated VLMs still possess considerable potential for improvement in first-person perspective tasks. Meanwhile, enlarging the number of trainable parameters has the most significant impact on model performance on EgoThink. In conclusion, EgoThink serves as a valuable addition to existing evaluation benchmarks for VLMs, providing an indispensable resource for future research in the realm of embodied artificial intelligence and robotics. Sijie Cheng, Zhicheng Guo, Kechen Fang, Peng Li 0030, Huaping Liu 0001, Yang Liu 0005 |
CVPR | 6 |
| 2024 | Efficient Multi-Scale Network with Learnable Discrete Wavelet Transform for Blind Motion DeblurringabstractCoarse-to-fine schemes are widely used in traditional single-image motion deblur; however, in the context of deep learning, existing multi-scale algorithms not only require the use of complex modules for feature fusion of low-scale RGB images and deep semantics, but also manually generate low-resolution pairs of images that do not have sufficient confidence. In this work, we propose a multi-scale network based on single-input and multiple-outputs(SIMO) for motion deblurring. This simplifies the complexity of algorithms based on a coarse-to-fine scheme. To alleviate restoration defects impacting detail information brought about by using a multi-scale architecture, we combine the characteristics of real-world blurring trajectories with a learnable wavelet transform module to focus on the directional continuity and frequency features of the step-by-step transitions between blurred images to sharp images. In conclusion, we propose a multi-scale network with a learnable discrete wavelet transform (MLWNet), which exhibits state-of-the-art performance on multiple real-world deblurred datasets, in terms of both subjective and objective quality as well as computational efficiency. Our code is available on https://github.com/thqiu0419/MLWNet. Xin Gao 0028, Tianheng Qiu, Xinyu Zhang 0001, Hanlin Bai, Kang Liu 0008, Hu Wei, Guoying Zhang, Huaping Liu 0001 |
CVPR | 9 |
| 2024 | Small Scale Data-Free Knowledge DistillationabstractData-free knowledge distillation is able to utilize the knowledge learned by a large teacher network to augment the training of a smaller student network without accessing the original training data, avoiding privacy, security, and proprietary risks in real applications. In this line of research, existing methods typically follow an inversion-and-distillation paradigm in which a generative adversarial network on-the-fly trained with the guidance of the pre-trained teacher network is used to synthesize a large-scale sample set for knowledge distillation. In this paper, we reexam-ine this common data-free knowledge distillation paradigm, showing that there is considerable room to improve the overall training efficiency through a lens of “small-scale inverted data for knowledge distillation”. In light of three empirical observations indicating the importance of how to balance class distributions in terms of synthetic sample di-versity and difficulty during both data inversion and distillation processes, we propose Small Scale Data-free Knowledge Distillation (SSD-KD). Informulation, SSD-KD introduces a modulating function to balance synthetic samples and a priority sampling function to select proper samples, facilitated by a dynamic replay buffer and a reinforcement learning strategy. As a result, SSD-KD can perform dis-tillation training conditioned on an extremely small scale of synthetic samples (e.g., 10x less than the original training data scale), making the overall training efficiency one or two orders of magnitude faster than many mainstream methods while retaining superior or competitive model performance, as demonstrated on popular image classification and semantic segmentation benchmarks. The code is available at https://github.com/OSVAI/SSD-KD. Yikai Wang 0001, Huaping Liu 0001, Fuchun Sun 0001, Anbang Yao |
CVPR | 3 |
| 2024 | Vision-Language Foundation Models as Effective Robot ImitatorsabstractRecent progress in vision language foundation models has shown their ability to understand multimodal data and resolve complicated vision language tasks, including robotics manipulation. We seek a straightforward way of making use of existing vision-language models (VLMs) with simple fine-tuning on robotics data.
To this end, we derive a simple and novel vision-language manipulation framework, dubbed RoboFlamingo, built upon the open-source VLMs, OpenFlamingo. Unlike prior works, RoboFlamingo utilizes pre-trained VLMs for single-step vision-language comprehension, models sequential history information with an explicit policy head, and is slightly fine-tuned by imitation learning only on language-conditioned manipulation datasets. Such a decomposition provides RoboFlamingo the flexibility for open-loop control and deployment on low-performance platforms. By exceeding the state-of-the-art performance with a large margin on the tested benchmark, we show RoboFlamingo can be an effective and competitive alternative to adapt VLMs to robot control.
Our extensive experimental results also reveal several interesting conclusions regarding the behavior of different pre-trained VLMs on manipulation tasks. We believe RoboFlamingo has the potential to be a cost-effective and easy-to-use solution for robotics manipulation, empowering everyone with the ability to fine-tune their own robotics policy. Our code will be made public upon acceptance. Xinghang Li, Minghuan Liu, Hanbo Zhang, Cunjun Yu, Chilam Cheang, Ya Jing, Weinan Zhang 0001, Huaping Liu 0001, Tao Kong |
ICLR | 10 |
| 2024 | Stimulate the Potential of Robots via CompetitionabstractIt is common for us to feel pressure in a competition environment, which arises from the desire to obtain success comparing with other individuals or opponents. Although we might get anxious under the pressure, it could also be a drive for us to stimulate our potentials to the best in order to keep up with others. Inspired by this, we propose a competitive learning framework which is able to help individual robot to acquire knowledge from the competition, fully stimulating its dynamics potential in the race. Specifically, the competition information among competitors is introduced as the additional auxiliary signal to learn advantaged actions. We further build a Multiagent-Race environment, and extensive experiments are conducted, demonstrating that robots trained in competitive environments outperform ones that are trained with SoTA algorithms in single robot environment. Kangyao Huang, Di Guo 0002, Xinyu Zhang 0001, Xiangyang Ji, Huaping Liu 0001 |
ICRA | 5 |
| 2024 | Smooth Computation without Input Delay: Robust Tube-Based Model Predictive Control for Robot Manipulator PlanningabstractModel Predictive Control (MPC) has exhibited remarkable capabilities in optimizing objectives and meeting constraints. However, the substantial computational burden associated with solving the Optimal Control Problem (OCP) at each triggering instant introduces significant delays between state sampling and control application. These delays limit the practicality of MPC in resource-constrained systems when engaging in complex tasks. The intuition to address this issue in this paper is that by predicting the successor state, the controller can solve the OCP one time step ahead of time thus avoiding the delay of the next action. To this end, we compute deviations between real and nominal system states, predicting forthcoming real states as initial conditions for the imminent OCP solution. Anticipatory computation stores optimal control based on current nominal states, thus mitigating the delay effects. Additionally, we establish an upper bound for linearization error, effectively linearizing the nonlinear system, reducing OCP complexity, and enhancing response speed. We provide empirical validation through two numerical simulations and corresponding real-world robot tasks, demonstrating significant performance improvements and augmented response speed (up to 90%) resulting from the seamless integration of our proposed approach compared to conventional time-triggered MPC strategies. Yu Luo 0021, Qie Sima, Tianying Ji, Fuchun Sun 0001, Huaping Liu 0001, Jianwei Zhang 0001 |
ICRA | 5 |
| 2024 | Bionic Soft Fingers with Hybrid Variable Stiffness Mechanisms for Multimode GraspingabstractThis paper presents a novel Bionic Soft Finger (BSF) that aims to overcome the limitations of conventional rigid manipulators in terms of adaptability and safety, as well as the challenges faced by soft hands regarding carrying capacity and stability. The BSF design uses a hybrid variable stiffness mechanism combining memory alloy actuators with particle jamming to achieve the desired bending angle and actuator stiffness. Our innovative approach utilizes a bionic finger design that incorporates a memory alloy skeleton and a water-cooled recirculation system, leading to a substantial reduction in the time required for each operation. Through the integration of particle jamming, we have enhanced the overall stiffness and performance of the manipulator, enabling load capacities of up to 3N per finger and more than twice the stiffness of a normal condition. Additionally, our design enables multimode grasping and incorporates a liquid metal strain sensor (METT) for real-time monitoring of finger bending angles. Comparative analyses demonstrate that our design exhibits superior stiffness and enables five-mode grasping in comparison to pneumatic actuators. We believe that bionic soft fingers present a promising solution for enhancing adaptability, safety, and performance in human-robot interaction applications. Xiangbo Wang, Hongze Yu, Zhenwei Wen, Lide Fang, Huaping Liu 0001, Fuchun Sun 0001, Lixue Tang, Bin Fang 0003 |
ICRA | 6 |
| 2024 | A Large-area Tactile Sensor for Distributed Force Sensing Using Highly Sensitive Piezoresistive SpongeabstractTactile sensing plays a critical role in enabling robots to interact safely with target objects in dynamic and unstructured environments. While various tactile sensors based on different sensing principles or different sensitive materials have been proposed, the development of flexible large-area tactile sensors for robots is still challenging. In this paper, a novel highly sensitive piezoresistive sponge based on multi-walled carbon nanotubes (MWCNTs) and polyurethane (PU) sponge is fabricated for pressure sensing. The sensing behavior of the piezoresistive sponge was experimentally evaluated, showing high sensitivity and fast response. Based on the piezoresistive sponge, a flexible large-area tactile sensor is designed for distributed force detection with electrical resistance tomography technology. The sensing performance of the sensor is validated by touch location, sensitivity analysis, real-time touch discrimination, and touch modality recognition. The experimental results indicate that the sensor performs well in detecting the position and force of contact in a large area. The sensor’s performance shows promise in embodied tactile sensing and human–robot interaction. Wendong Zheng, Di Guo 0002, Wuqiang Yang, Huaping Liu 0001 |
ICRA | 6 |
| 2024 | CompetEvo: Towards Morphological Evolution from Competition
Kangyao Huang, Di Guo 0002, Xinyu Zhang 0001, Xiangyang Ji, Huaping Liu 0001 |
IJCAI | 5 |
| 2024 | Skill enhancement learning with knowledge distillation
Naijun Liu, Fuchun Sun 0001, Bin Fang 0003, Huaping Liu 0001 |
Sci. China Inf. Sci. | 4 |
| 2024 | Robust tube-based MPC with smooth computation for dexterous robot manipulation
Yu Luo 0021, Tianying Ji, Fuchun Sun 0001, Qie Sima, Huaping Liu 0001, Mingxuan Jing, Jianwei Zhang 0001 |
Sci. China Inf. Sci. | 5 |
| 2024 | Data-driven electrical resistance tomography for robotic large-area tactile sensing
Wendong Zheng, Huaping Liu 0001, Fuchun Sun 0001 |
Sci. China Inf. Sci. | 2 |
| 2024 | Multimodal information bottleneck for deep reinforcement learning with multiple sensors
Bang You, Huaping Liu 0001 |
Neural Networks | 2 |
| 2024 | Learning Semantic Alignment Using Global Features and Multi-Scale ConfidenceabstractSemantic alignment aims to establish pixel correspondences between images based on semantic consistency. It can serve as a fundamental component for various downstream computer vision tasks, such as style transfer and exemplar-based colorization, etc. Many existing methods use local features and their cosine similarities to infer semantic alignment. However, they struggle with significant intra-class variation of objects, such as appearance, size, etc. In other words, contents with the same semantics tend to be significantly different in vision. To address this issue, we propose a novel deep neural network of which the core lies in global feature enhancement and adaptive multi-scale inference. Specifically, two modules are proposed: an enhancement transformer for enhancing semantic features with global awareness; a probabilistic correlation module for adaptively fusing multi-scale information based on the learned confidence scores. We use the unified network architecture to achieve two types of semantic alignment, namely, cross-object semantic alignment and cross-domain semantic alignment. Experimental results demonstrate that our method achieves competitive performance on five standard cross-object semantic alignment benchmarks, and outperforms the state of the arts in cross-domain semantic alignment. Huaiyuan Xu, Jing Liao 0001, Huaping Liu 0001, Yuxiang Sun 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Digital-Twin-Assisted Skill Learning for 3C Assembly TasksabstractThe utilization of robots in computer, communication, and consumer electronics (3C) assembly has the potential to significantly reduce labor costs and enhance assembly efficiency. However, many typical scenarios in 3C assembly, such as the assembly of flexible printed circuits (FPCs), involve complex manipulations with long-horizon steps and high-precision requirements that cannot be effectively accomplished through manual programming or conventional skill-learning methods. To address this challenge, this article proposes a learning-based framework for the acquisition of complex 3C assembly skills assisted by a multimodal digital-twin environment. First, we construct a fully equivalent digital-twin environment based on the real-world counterpart, equipped with visual, tactile force, and proprioception information, and then collect multimodal demonstration data using virtual reality (VR) devices. Next, we construct a skill knowledge base through multimodal skill parsing of demonstration data, resulting in primitive policy sequences for achieving 3C assembly tasks. Finally, we train primitive policies via a combination of curriculum learning, residual reinforcement learning, and domain randomization methods and transfer the learned skill from the digital-twin environment to the real-world environment. The experiments are conducted to verify the effectiveness of our proposed method. Fuchun Sun 0001, Naijun Liu, Ruize Sun, Shengyi Miao, Zengxin Kang, Bin Fang 0003, Huaping Liu 0001, Yongjia Zhao, Haiming Huang |
IEEE Trans. Cybern. | 8 |
| 2024 | Continual Residual Reservoir Computing for Remaining Useful Life PredictionabstractIn practical engineering applications, it is inevitable that production is often faced with different working conditions. Therefore, it is necessary to have a continual learning system, which can adapt to a sequence of tasks and keep learning over time. In this work, we propose a continual deep residual reservoir computing framework for the practical remaining useful life prediction task. Specifically, we propose a novel deep echo state network structure with residual blocks to effectively mitigate the performance degradation of the deep reservoir computing framework and reduce the difficulty of model training. Furthermore, the proposed framework is trained with the elastic weight consolidation method to alleviate the impact of catastrophic forgetting in continual learning systems. Extensive experiments are conducted with the FEMTO-ST bearing and high-intensity-radiated fields battery dataset. And the proposed framework is proven to be effective in multiple continual learning tasks compared with other state-of-the-art methods. Huaping Liu 0001, Di Guo 0002, Xi-Ming Sun |
IEEE Trans. Ind. Informatics | 2 |
| 2024 | Rich Action-Semantic Consistent Knowledge for Early Action PredictionabstractEarly action prediction (EAP) aims to recognize human actions from a part of action execution in ongoing videos, which is an important task for many practical applications. Most prior works treat partial or full videos as a whole, ignoring rich action knowledge hidden in videos, i.e., semantic consistencies among different partial videos. In contrast, we partition original partial or full videos to form a new series of partial videos and mine the Action-Semantic Consistent Knowledge (ASCK) among these new partial videos evolving in arbitrary progress levels. Moreover, a novel Rich Action-semantic Consistent Knowledge network (RACK) under the teacher-student framework is proposed for EAP. Firstly, we use a two-stream pre-trained model to extract features of videos. Secondly, we treat the RGB or flow features of the partial videos as nodes and their action semantic consistencies as edges. Next, we build a bi-directional semantic graph for the teacher network and a single-directional semantic graph for the student network to model rich ASCK among partial videos. The MSE and MMD losses are incorporated as our distillation loss to enrich the ASCK of partial videos from the teacher to the student network. Finally, we obtain the final prediction by summering the logits of different subnetworks and applying a softmax layer. Extensive experiments and ablative studies have been conducted, demonstrating the effectiveness of modeling rich ASCK for EAP. With the proposed RACK, we have achieved state-of-the-art performance on three benchmarks. The code is available at https://github.com/lily2lab/RACK.git. Jianqin Yin, Di Guo 0002, Huaping Liu 0001 |
IEEE Trans. Image Process. | 4 |
| 2024 | Auto-Points: Automatic Learning for Point Cloud Analysis With Neural Architecture SearchabstractPure point-based neural networks have recently shown tremendous promise for point cloud tasks, including 3D object classification, 3D object part segmentation, 3D semantic segmentation, and 3D object detection. Nevertheless, it is a laborious process to construct a network for each task due to the artificial parameters and hyperparameters involved, e.g., the depths and widths of the network and the number of sampled points at each stage. In this work, we propose Auto-Points, a novel one-shot search framework that automatically seeks the optimal architecture configuration for point cloud tasks. Technically, we introduce a set abstraction mixer (SAM) layer that is capable of scaling up flexibly along the depth and width of the network. Each SAM layer consists of numerous child candidates, which simplifies architecture search and enables us to discover the optimum design for each point cloud task pursuant to resource constraint from an enormous search space. To fully optimize the child candidates, we develop a weight-entwinement neural architecture search (NAS) technique that entwines the weights of different candidates in the same layer during supernet training such that all candidates can be extremely optimized. Benefiting from the proposed techniques, the trained supernet allows the searched subnets to be exceptionally well-optimized without further retraining or finetuning. In particular, the searched models deliver superior performances on multiple extensively employed benchmarks, 93.9% overall accuracy (OA) on ModelNet40, 89.1% OA on ScanObjectNN, 87.1% instance average IoU on ShapeNetPart, 69.1% mIoU on S3DIS, 70.4% [email protected] on ScanNet V2, and 64.4% [email protected] on SUN RGB-D. Li Wang 0092, Tao Xie 0010, Xinyu Zhang 0001, Linqi Yang, Yilong Ren, Haiyang Yu 0002, Jun Li 0082, Huaping Liu 0001 |
IEEE Trans. Multim. | 11 |
| 2024 | FARP-Net: Local-Global Feature Aggregation and Relation-Aware Proposals for 3D Object DetectionabstractIn this work, we introduce FARP-Net, an adaptive local-global feature aggregation and relation-aware proposal network for high-quality 3D object detection from pure point clouds. Our key insight is that learning adaptive local-global feature aggregation from an irregular yet sparse point cloud and generating superb proposals are both pivotal for detection. Technically, we propose a novel local-global feature aggregation layer (LGFAL) that fully exploits the complementary correlation between local features and global features, and fuses their strengths adaptively via an attention-based fusion module. Furthermore, we incorporate a lightweight feature affine module (LFAM) into LGFAL to map the local features into a normal distribution, thus acquiring fine-grained features of each local region in a weight-sharing manner. During object proposal generation, we propose a weighted relation-aware proposal module (WRPM) that uses an objectness-aware formalism to weigh the relation importance among object candidates for a clear and principal context, thereby facilitating the generation of high-quality proposals. The WRPM challenges the traditional practice of extracting contextual information among all object candidates, which is inefficient as object candidates are always noisy and redundant. Experimentally, FARP-Net delivers superior performance on two widely used benchmarks with fewer parameters, 64.0% [email protected] on the SUN RGB-D dataset and 70.9% [email protected] on the ScanNet V2 dataset. We further validate that the proposed LGFAL and WRPM can be integrated into both indoor and outdoor detectors to boost performance. Tao Xie 0010, Li Wang 0092, Ke Wang 0028, Ruifeng Li 0001, Xinyu Zhang 0001, Linqi Yang, Huaping Liu 0001, Jun Li 0082 |
IEEE Trans. Multim. | 8 |
| 2024 | Informative Data Selection With Uncertainty for Multimodal Object DetectionabstractNoise has always been nonnegligible trouble in object detection by creating confusion in model reasoning, thereby reducing the informativeness of the data. It can lead to inaccurate recognition due to the shift in the observed pattern, that requires a robust generalization of the models. To implement a general vision model, we need to develop deep learning models that can adaptively select valid information from multimodal data. This is mainly based on two reasons. Multimodal learning can break through the inherent defects of single-modal data, and adaptive information selection can reduce chaos in multimodal data. To tackle this problem, we propose a universal uncertainty-aware multimodal fusion model. It adopts a multipipeline loosely coupled architecture to combine the features and results from point clouds and images. To quantify the correlation in multimodal information, we model the uncertainty, as the inverse of data information, in different modalities and embed it in the bounding box generation. In this way, our model reduces the randomness in fusion and generates reliable output. Moreover, we conducted a completed investigation on the KITTI 2-D object detection dataset and its derived dirty data. Our fusion model is proven to resist severe noise interference like Gaussian, motion blur, and frost, with only slight degradation. The experiment results demonstrate the benefits of our adaptive fusion. Our analysis on the robustness of multimodal fusion will provide further insights for future research. Xinyu Zhang 0001, Zhiwei Li 0011, Zhenhong Zou, Xin Gao 0028, Yijin Xiong, Dafeng Jin, Jun Li 0082, Huaping Liu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 8 |
| 2023 | Mixed Neural Voxels for Fast Multi-view Video SynthesisabstractSynthesizing high-fidelity videos from real-world multi-view input is challenging due to the complexities of real-world environments and high-dynamic movements. Previous works based on neural radiance fields have demonstrated high-quality reconstructions of dynamic scenes. However, training such models on real-world scenes is time-consuming, usually taking days or weeks. In this paper, we present a novel method named MixVoxels to efficiently represent dynamic scenes, enabling fast training and rendering speed. The proposed MixVoxels represents the 4D dynamic scenes as a mixture of static and dynamic voxels and processes them with different networks. In this way, the computation of the required modalities for static voxels can be processed by a lightweight model, which essentially reduces the amount of computation as many daily dynamic scenes are dominated by static backgrounds. To distinguish the two kinds of voxels, we propose a novel variation field to estimate the temporal variance of each voxel. For the dynamic representations, we design an inner product time query method to efficiently query multiple time steps, which is essential to recover the high-dynamic movements. As a result, with 15 minutes of training for dynamic scenes with inputs of 300-frame videos, MixVoxels achieves better PSNR than previous methods. For rendering, MixVoxels can render a novel view video with 1K resolution at 37 fps. Codes and trained models are available at https://github.com/fengres/mixvoxels. Feng Wang 0034, Sinan Tan, Xinghang Li, Zeyue Tian, Huaping Liu 0001 |
ICCV | 6 |
| 2023 | Embodied Referring Expression for Manipulation Question Answering in Interactive EnvironmentabstractEmbodied agents are expected to perform more complicated tasks in an interactive environment, with the progress of Embodied AI in recent years. Existing embodied tasks including Embodied Referring Expression (ERE) and other QA-form tasks mainly focuses on interaction in term of linguistic instruction. Therefore, enabling the agent to manipulate objects in the environment for exploration actively has become a challenging problem for the community. To solve this problem, We introduce a new embodied task: Remote Embodied Manipulation Question Answering (REMQA) to combine ERE with manipulation tasks. In REMQA task, the agent needs to navigate to a remote position and perform manipulation with the target object to answer the question. We build a benchmark dataset for the REMQA task in AI2-THOR simulator. To this end, a framework with 3D semantic reconstruction and modular network paradigms is proposed. The evaluation of the proposed framework on REMQA dataset is presented to validate its effectiveness. Qie Sima, Sinan Tan, Huaping Liu 0001, Fuchun Sun 0001 |
ICRA | 3 |
| 2023 | Natural Language Instruction Understanding for Robotic Manipulation: a Multisensory Perception ApproachabstractIt has always been expected that the robot can understand the natural language instruction and thus a more natural human-robot interaction is achieved. Currently, the robot usually interprets the instruction by visually grounding the textual information to its surroundings, while it may be not enough for some complex situations with only visual perception. So it is reasonable for the robot to leverage its multisensory perception ability to better understand the instruction. In this paper, we propose a multisensory perception approach to tackle the task of natural language instruction understanding for robotic manipulation, in which the robot coordinates its visual, tactile and auditory perception to fully understand the instruction and then executes the manipulation task. Extensive experiments have been conducted demonstrating the superiority of the multisensory perception compared with single sensory perception for instruction understanding. Moreover, we establish a user-friendly human-robot interaction interface where the human sends instruction to the robot via a mobile APP. Yanzhi Dong, Di Guo 0002, Huaping Liu 0001 |
ICRA | 6 |
| 2023 | Adaptive Optimal Electrical Resistance Tomography for Large-Area Tactile SensingabstractIt is critical to perceive physical contact for intelligent robots to safely interact in dynamic, unstructured environments. As physical contacts can occur at any location, a well-performing tactile sensing system should be able to deploy a large area on robotic surface. Some researchers have implemented large-area tactile sensors by using sensing arrays, but it is challenging to deploy many sensing elements. Electrical resistance tomography (ERT) has recently been introduced into tactile sensing to overcome some of the limitations with conventional tactile sensing arrays, and good results have been achieved for some robotic applications. However, a particular challenge is that spatial resolution is low. Although various attempts have been made to improve the performance of ERT-based tactile sensors, the intrinsic resolution issue remains unsolved. In this paper, we propose a novel adaptive optimal drive strategy for efficient ERT-based large-area tactile sensing for robotic applications, which can adaptively select the current injection and voltage measurement pattern for optimal tactile stimulus. In particular, regions of tactile contacts are preliminarily detected and localized by a base scanning pattern with only a few measurement data. According to this detected region, the adaptive strategy can select the optimal current injection and voltage measurement pattern to improve the sensing performance by maximizing the current density. To verify the effectiveness of the proposed strategy, the proposed method is comprehensively evaluated by simulation and experiments. The results revealed that the optimal strategy can effectively improve both spatial and temporal resolution. Wendong Zheng, Huaping Liu 0001, Di Guo 0002, Wuqiang Yang |
ICRA | 2 |
| 2023 | Deep Inversion Method for Attacking Lifelong Learning Neural NetworksabstractArtificial neural networks suffer from catastrophic forgetting when knowledge needs to be learned from multi-batch or streaming data. In response to this problem, researchers have proposed a variety of lifelong learning methods to avoid catastrophic forgetting. However, current methods usually do not consider the possibility of malicious attacks. Meanwhile, in real lifelong learning scenarios, batch data or streaming data usually come from an incompletely trusted environment. Attackers can easily manipulate data or inject malicious samples into the training data set. As a result, the reliability of neural networks decreases. Recently, researches of lifelong learning attacks need to obtain real samples of the attacked classes, whether using backdoor attacks or data poisoning attacks. In this paper, we focus on an attack setting that is more suitable for lifelong learning scenario. This setting has two main features. The first is the setting does not require real samples of the attacked classes, and the second is it allows attacks to be performed on tasks that exclude the attacked classes. For this scenario, we propose a lifelong learning attack model based on deep inversion. In the scenario where EWC is used as the benchmark lifelong learning model, our experiments show that 1) in the data poisoning attack, the target accuracy can be significantly decreased by adding 0.5% of poisoned samples; 2) The backdoor attack with high accuracy can be achieved by adding 1% of backdoor samples. Boyuan Du, Yuanlong Yu 0001, Huaping Liu 0001 |
IJCNN | 3 |
| 2023 | Masked Space-Time Hash Encoding for Efficient Dynamic Scene ReconstructionabstractIn this paper, we propose the Masked Space-Time Hash encoding (MSTH), a novel method for efficiently reconstructing dynamic 3D scenes from multi-view or monocular videos. Based on the observation that dynamic scenes often contain substantial static areas that result in redundancy in storage and computations, MSTH represents a dynamic scene as a weighted combination of a 3D hash encoding and a 4D hash encoding. The weights for the two components are represented by a learnable mask which is guided by an uncertainty-based objective to reflect the spatial and temporal importance of each 3D position. With this design, our method can reduce the hash collision rate by avoiding redundant queries and modifications on static areas, making it feasible to represent a large number of space-time voxels by hash tables with small size.Besides, without the requirements to fit the large numbers of temporally redundant features independently, our method is easier to optimize and converge rapidly with only twenty minutes of training for a 300-frame dynamic scene. We evaluate our method on extensive dynamic scenes. As a result, MSTH obtains consistently better results than previous state-of-the-art methods with only 20 minutes of training time and 130 MB of memory storage. Feng Wang 0034, Zilong Chen, Guokang Wang, Huaping Liu 0001 |
NeurIPS | 5 |
| 2023 | Heterogeneous graph attention network with motif clique
Chenxu Wang 0003, Minnan Luo, Zhen Peng 0005, Yixiang Dong, Huaping Liu 0001 |
Neurocomputing | 5 |
| 2023 | SAT-GCN: Self-attention graph convolutional network-based 3D object detection for autonomous driving
Li Wang 0092, Ziying Song, Xinyu Zhang 0001, Jun Li 0082, Huaping Liu 0001 |
Knowl. Based Syst. | 8 |
| 2023 | Knowledge-Based Embodied Question AnsweringabstractIn this paper, we propose a novel Knowledge-based Embodied Question Answering (K-EQA) task, in which the agent intelligently explores the environment to answer various questions with the knowledge. Different from explicitly specifying the target object in the question as existing EQA work, the agent can resort to external knowledge to understand more complicated question such as "Please tell me what are objects used to cut food in the room?", in which the agent must know the knowledge such as "knife is used for cutting food". To address this K-EQA problem, a novel framework based on neural program synthesis reasoning is proposed, where the joint reasoning of the external knowledge and 3D scene graph is performed to realize navigation and question answering. Especially, the 3D scene graph can provide the memory to store the visual information of visited scenes, which significantly improves the efficiency for the multi-turn question answering. Experimental results have demonstrated that the proposed framework is capable of answering more complicated and realistic questions in the embodied environment. The proposed method is also applicable to multi-agent scenarios. Sinan Tan, Mengmeng Ge 0002, Di Guo 0002, Huaping Liu 0001, Fuchun Sun 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Autonomous Oropharyngeal-Swab Robot System for COVID-19 PandemicabstractThe outbreak of COVID-19 has led to the shortage of medical personnel and the increasing need for nucleic acid testing. Manual oropharyngeal sampling is susceptible to inconsistency caused by fatigue and close contact could also cause healthcare personnel exposure and cross infection. The innate deficiency calls for a safer and more consistent way to collect the oropharyngeal samples. Therefore a fully autonomous oropharyngeal-swab robot system is proposed in this paper. The system is installed in a negative pressure chamber and carrying out a standardized sampling process to minimize individual sampling differences. A hierarchical throat detection algorithm is presented and multiple modality sensory information are fused to safely and accurately localize the optimum sampling location. Also, a force/position hybrid control method is adopted to ensure both accurate sampling and subject comfort. The robot system described in this paper can safely and efficiently collect the oropharyngeal sample, providing a scalable solution for large-scale Polymerase Chain Reaction (PCR) Molecular sample collection for various respiratory diseases. Note to Practitioners—During the COVID-19 pandemic, pre-diagnostic is essential for both prevention and treatment. Existing approaches, including nasal swab and oropharyngeal-swab, require extensive medical worker training and increase the chance of cross-infection. The robot system introduced in this paper can take oropharyngeal-swab samples from subjects with minimum human intervention, reducing medical worker exposure, alleviating the work pressure of medical staff, and speed up large quantity of sampling plan. The robot will first guide the subject into position with vocal commands, and automatically detect the optimum sampling location with a real-time machine learning algorithm. A dedicated control strategy aiming at minimizing discomfort and uniforming sample quantity is then applied to safely collect nucleic samples from the throat. Eventually, while the swab is being stored in the culture medium, a disinfection process is carried out simultaneously to prepare the robot for the next subject. Preliminary clinical trials show that our robot system can safely and accurately collect samples from subjects. Fuchun Sun 0001, Huaping Liu 0001, Bin Fang 0003 |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2023 | Mix-Teaching: A Simple, Unified and Effective Semi-Supervised Learning Framework for Monocular 3D Object DetectionabstractSemi-supervised learning (SSL) has promising potential for improving model performance using both labelled and unlabelled data. Since recovering 3D information from 2D images is an ill-posed problem, the current state-of-the-art methods of monocular 3D object detection (Mono3D) have relatively low precision and recall, making semi-supervised learning for Mono3D tasks challenging and understudied. In this work, we propose a unified and effective semi-supervised learning framework called Mix-Teaching that can be applied to most monocular 3D object detectors. Based on the idea of decomposition and recombination, unlabelled samples are firstly decomposed into collections of image patches with high-quality predictions and collections of background images containing no objects. The student model is then trained on the mixed images containing dense instances with high-quality pseudo-labels generated by the recombination operation. In addition, we propose an uncertainty-based filter to distinguish high-quality pseudo-labels from noisy predictions during the decomposition process. As results in KITTI and nuScenes benchmarks, Mix-Teaching consistently improves MonoFlex and GUPNet by significant margins under various labeling ratios. Our method achieves around +6.34%$AP_{3D}$improvement against the GUPNet on the validation set when using only 10% labelled data. Using the full training set and the additional 38K raw images from KITTI, it can further improve the MonoFlex by +4.65% absolute improvement on$AP_{3D}$for car detection, reaching 18.54%$AP_{3D}$, which ranks the 1st place among all monocular based methods on the KITTI test leaderboard. Lei Yang 0060, Xinyu Zhang 0001, Jun Li 0082, Li Wang 0092, Minghan Zhu, Huaping Liu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2023 | Self-Supervised Learning by Estimating Twin Class DistributionabstractWe present Twist, a simple and theoretically explainable self-supervised representation learning method by classifying large-scale unlabeled datasets in an end-to-end way. We employ a siamese network terminated by a softmax operation to produce twin class distributions of two augmented images. Without supervision, we enforce the class distributions of different augmentations to be consistent. However, simply minimizing the divergence between augmentations will generate collapsed solutions, i.e., outputting the same class distribution for all images. In this case, little information about the input images is preserved. To solve this problem, we propose to maximize the mutual information between the input image and the output class predictions. Specifically, we minimize the entropy of the distribution for each sample to make the class prediction assertive, and maximize the entropy of the mean distribution to make the predictions of different samples diverse. In this way, Twist can naturally avoid the collapsed solutions without specific designs such as asymmetric network, stop-gradient operation, or momentum encoder. As a result, Twist outperforms previous state-of-the-art methods on a wide range of tasks. Specifically on the semi-supervised classification task, Twist achieves 61.2% top-1 accuracy with 1% ImageNet labels using a ResNet-50 as backbone, surpassing previous best results by an improvement of 6.2%. Codes and pre-trained models are available at https://github.com/bytedance/TWIST. Feng Wang 0034, Tao Kong, Rufeng Zhang, Huaping Liu 0001 |
IEEE Trans. Image Process. | 4 |
| 2023 | CAMO-MOT: Combined Appearance-Motion Optimization for 3D Multi-Object Tracking With Camera-LiDAR Fusionabstract3D Multi-object tracking (MOT) ensures consistency during continuous dynamic detection, conducive to subsequent motion planning and navigation tasks in autonomous driving. However, camera-based methods suffer in the case of occlusions and it can be challenging to track the irregular motion of objects for LiDAR-based methods accurately. Some fusion methods work well but do not consider the untrustworthy issue of appearance features under occlusion. At the same time, the false detection problem also significantly affects tracking. As such, we propose a novel camera-LiDAR fusion 3D MOT framework based on Combined Appearance-Motion Optimization (CAMO-MOT), which uses both camera and LiDAR data and significantly reduces tracking failures caused by occlusion and false detection. For occlusion problems, we are the first to propose an occlusion head to select the best object appearance features multiple times effectively, reducing the influence of occlusions. To decrease the impact of false detection in tracking, we design a motion cost matrix based on confidence scores which improve the positioning and object prediction accuracy in 3D space. As existing multi-object tracking methods always evaluate each category separately and do not consider the mismatch between objects of different categories, we also propose to build a multi-category cost to implement multi-object tracking in multi-category scenes. A series of validation experiments are conducted on the KITTI and nuScenes tracking benchmarks. Our proposed method achieves state-of-the-art performance with 79.99% HOTA and the lowest identity switches (IDS) value (23 for Car and 137 for Pedestrian) among all multi-modal MOT methods on the KITTI test dataset. And our method achieves state-of-the-art performance among all algorithms on the nuScenes test dataset with 75.3% AMOTA. Li Wang 0092, Xinyu Zhang 0001, Wenyuan Qin, Jinghan Gao, Lei Yang 0060, Zhiwei Li 0011, Jun Li 0082, Hong Wang 0014, Huaping Liu 0001 |
IEEE Trans. Intell. Transp. Syst. | 11 |
| 2023 | Hybrid Robotic Grasping With a Soft Multimodal Gripper and a Deep Multistage Learning SchemeabstractGrasping has long been considered an important and practical task in robotic manipulation. Yet achieving robust and efficient grasps of diverse objects is challenging, since it involves gripper design, perception, control, and learning, etc. Recent learning-based approaches have shown excellent performance in grasping a variety of novel objects. However, these methods either are typically limited to one single grasping mode or else more end effectors are needed to grasp various objects. In addition, gripper design and learning methods are commonly developed separately, which may not adequately explore the ability of a multimodal gripper. In this article, we present a deep reinforcement learning (DRL) framework to achieve multistage hybrid robotic grasping with a new soft multimodal gripper. A soft gripper with three grasping modes (i.e.,enveloping,sucking, andenveloping_then_sucking) can both deal with objects of different shapes and grasp more than one object simultaneously. We propose a novel hybrid grasping method integrated with the multimodal gripper to optimize the number of grasping actions. We evaluate the DRL framework under different scenarios (i.e., with different ratios of objects of two grasp types). The proposed algorithm is shown to reduce the number of grasping actions (i.e., enlarge the grasping efficiency, with maximum values of 161.0% in simulations, and 153.5% in real-world experiments) compared to single grasping modes. Fukang Liu, Fuchun Sun 0001, Bin Fang 0003, Xiang Li 0009, Songyu Sun, Huaping Liu 0001 |
IEEE Trans. Robotics | 6 |
| 2022 | Sim2Real Object-Centric Keypoint Detection and DescriptionabstractKeypoint detection and description play a central role in computer vision. Most existing methods are in the form of scene-level prediction, without returning the object classes of different keypoints. In this paper, we propose the object-centric formulation, which, beyond the conventional setting, requires further identifying which object each interest point belongs to. With such fine-grained information, our framework enables more downstream potentials, such as object-level matching and pose estimation in a clustered environment. To get around the difficulty of label collection in the real world, we develop a sim2real contrastive learning mechanism that can generalize the model trained in simulation to real-world applications. The novelties of our training method are three-fold: (i) we integrate the uncertainty into the learning framework to improve feature description of hard cases, e.g., less-textured or symmetric patches; (ii) we decouple the object descriptor into two independent branches, intra-object salience and inter-object distinctness, resulting in a better pixel-wise description; (iii) we enforce cross-view semantic consistency for enhanced robustness in representation learning. Comprehensive experiments on image matching and 6D pose estimation verify the encouraging generalization ability of our method. Particularly for 6D pose estimation, our method significantly outperforms typical unsupervised/sim2real methods, achieving a closer gap with the fully supervised counterpart. Chengliang Zhong, Chao Yang 0026, Fuchun Sun 0001, Jinshan Qi, Xiaodong Mu, Huaping Liu 0001, Wenbing Huang 0001 |
AAAI | 6 |
| 2022 | Depth-Aware Vision-and-Language Navigation using Scene Query Attention NetworkabstractVision-and-language navigation (VLN) has been an important task in the field of Robotics and Computer Vision. However, most existing vision-and-language navigation models only use features extracted from RGB observation as input, while robots can utilize depth sensors in the real world. Existing research has also shown that simply adding a depth stream to neural models could only provide a marginal improvement to the performance of the VLN task. Therefore, in our work, we develop a novel method for the VLN task using semantic map observations built from RGB-D input. We use vision-pretraining to efficiently encode the semantic map with CNN and scene query attention network by answering queries about semantic information of specific regions of a scene. The proposed method could be used with a simple model and does not require large-scale vision-language transformer pretraining, bringing a more than 10% increase in the success rate compared with a baseline model. When used together with the Speaker-Follower training technique, it achieves a success rate of 58 % on the test set for the R2R dataset in single-run setting, outperforming the previous RGB-D method and most existing RGB-only models that do not use large-scale vision-language transformers pretraining. Sinan Tan, Mengmeng Ge 0002, Di Guo 0002, Huaping Liu 0001, Fuchun Sun 0001 |
ICRA | 4 |
| 2022 | Audio-Visual Grounding Referring Expression for Robotic ManipulationabstractReferring expressions are commonly used when referring to a specific target in people's daily dialogue. In this paper, we develop a novel task of audio-visual grounding referring expression for robotic manipulation. The robot leverages both the audio and visual information to understand the referring expression in the given manipulation instruction and the corresponding manipulations are implemented. To solve the proposed task, an audio-visual framework is proposed for visual localization and sound recognition. We have also established a dataset which contains visual data, auditory data and manipulation instructions for evaluation. Finally, extensive experiments are conducted both offline and online to verify the effectiveness of the proposed audio-visual framework. And it is demonstrated that the robot performs better with the audio-visual data than with only the visual data. Yefei Wang, Di Guo 0002, Huaping Liu 0001, Fuchun Sun 0001 |
ICRA | 5 |
| 2022 | InterFusion: Interaction-based 4D Radar and LiDAR Fusion for 3D Object DetectionabstractMany recent works detect 3D objects by several sensor modalities for autonomous driving, where high-resolution cameras and high-line LiDARs are mostly used but relatively expensive. To achieve a balance between overall cost and detection accuracy, many multi-modal fusion techniques have been suggested. In recent years, the fusion of LiDAR and Radar has gained ever-increasing attention, especially 4D Radar, which can adapt to bad weather conditions due to its penetrability. Although features have been fused from multiple sensing modalities, most methods cannot learn interactions from different modalities, which does not make for their best use. Inspired by the self-attention mechanism, we present InterFusion, an interaction-based fusion framework, to fuse 16-line LiDAR with 4D Radar. It aggregates features from two modalities and identifies cross-modal relations between Radar and LiDAR features. In experimental evaluations on the Astyx HiRes 2019 dataset, our method outperformed the baseline by 4.20% mAP in 3D and 10.76% BEV mAP for the car class at the moderate level. Li Wang 0092, Xinyu Zhang 0001, Baowei Xv, Jinzhao Zhang, Haibing Ren, Pingping Lu, Jun Li 0082, Huaping Liu 0001 |
IROS | 11 |
| 2022 | Visual Affordance Guided Tactile Material Recognition for Waste RecyclingabstractBecause more and more solid waste is generated, in particular, in cities, the management of solid waste disposal has become a global challenge. A solution is to find an effective way to sort solid waste materials and recycle them into reusable products. In this article, we propose to use a material recognition method with vision-guided tactile to form a robotic system for waste sorting. The vision guidance module integrates an object detector and an affordance network together. It allows the robot to not only detect the desired containers and packaging from an assortment of the waste but also obtain a configuration to grasp the target and actively collect its tactile data. By classifying the object with the tactile data, the robot can sort containers and packaging into their respective categories according to the type of material. Our experimental results demonstrate the effectiveness of the proposed robotic waste sorting system in sorting containers and packaging various types of materials. Note to Practitioners—The management of waste has become a great challenge in environmental protection in the world. We propose a robotic waste sorting system, which utilizes a vision-guided tactile sensing approach to find target waste and sort the waste according to materials. A visual module is used to find the waste of interest. A class-specific affordance map is generated to guide a robotic hand to actively grasp waste and collect the tactile data for material recognition. With the recorded tactile data, the robot can recognize the material of the target waste and sort the waste according to the type of material. The proposed system demonstrates good performance in waste sorting, and it can be easily implemented in practical scenarios. The target material can be generalized to a broader group of objects. Di Guo 0002, Huaping Liu 0001, Bin Fang 0003, Fuchun Sun 0001, Wuqiang Yang |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2022 | Backstepping Control for a Class of Nonlinear Discrete-Time Systems Subject to Multisource Disturbances and Actuator SaturationabstractIn this article, the backstepping control scheme is designed for a class of systems with multisource disturbances, actuator saturation, and nonlinearities in the domain of discrete time. To address the multisource disturbances, we put forward a novel discrete-time hybrid observer, which can deal with both modeled and unmodeled disturbances. In virtue of the radial basis function neural networks, the unknown nonlinearities are approximated. In addition, the anti-windup technique is adopted to cope with the actuator saturation phenomenon, which is pervasive in engineering practice. Bearing all the adopted mechanisms in mind, the composite control strategy is designed in a backstepping manner. Sufficient conditions are established to guarantee that the states of the system ultimately converge to a small range with linear matrix inequalities. Finally, the effectiveness of the presented methodology is verified for the spacecraft attitude system. Yang Yu 0069, Yuan Yuan 0006, Huaping Liu 0001 |
IEEE Trans. Cybern. | 3 |
| 2022 | Nash Equilibrium Seeking for Graphic Games With Dynamic Event-Triggered MechanismabstractIn this article, a discrete-time Nash equilibrium (NE) seeking problem is studied for a class of graphic games. In order to reduce the signal transmission frequency between adjacent players, a dynamic event-triggered mechanism is designed. For the purpose of regulating the actions of players to the NE points, a discrete-time NE seeking strategy is designed by only using the local action information. Then, sufficient conditions are provided to ensure that the actions of all players converge to the NE point. Finally, a numerical example of a multisatellite communication coordination problem is given to verify the effectiveness of the proposed NE seeking method. Peng Zhang 0056, Yuan Yuan 0006, Huaping Liu 0001 |
IEEE Trans. Cybern. | 3 |
| 2022 | Fine-Grained Multilevel Fusion for Anti-Occlusion Monocular 3D Object DetectionabstractWe propose a deep fine-grained multi-level fusion architecture for monocular 3D object detection, with an additionally designed anti-occlusion optimization process. Conventional monocular 3D object detection methods usually leverage geometry constraints such as keypoints, object shape relationships, and 3D to 2D optimizations to offset the lack of accurate depth information. However, these methods still struggle against directly extracting rich information for fusion from the depth estimation. To solve the problem, we integrate the monocular 3D features with the pseudo-LiDAR filter generation network between fine-grained multi-level layers. Our network utilizes the inherent multi-scale and promotes depth and semantic information flow in different stages. The new architecture can obtain features that incorporate more reliable depth information. At the same time, the problem of occlusion among objects is prevalent in natural scenes yet remains unsolved mainly. We propose a novel loss function that aims at alleviating the problem of occlusion. Extensive experiments have proved that the framework demonstrates a competitive performance, especially for the complex scenes with occlusion. Huaping Liu 0001, Yikai Wang 0001, Fuchun Sun 0001, Wenbing Huang 0001 |
IEEE Trans. Image Process. | 2 |
| 2022 | $\hbox {PISEP}{^2}$: pseudo-image sequence evolution-based 3D pose prediction
Jianqin Yin, Huaping Liu 0001, Yilong Yin |
Vis. Comput. | 3 |
| 2021 | Understanding the Behaviour of Contrastive LossabstractUnsupervised contrastive learning has achieved out-standing success, while the mechanism of contrastive loss has been less studied. In this paper, we concentrate on the understanding of the behaviours of unsupervised contrastive loss. We will show that the contrastive loss is a hardness-aware loss function, and the temperature τ controls the strength of penalties on hard negative samples. The previous study has shown that uniformity is a key property of contrastive learning. We build relations between the uniformity and the temperature τ. We will show that uniformity helps the contrastive learning to learn separable features, however excessive pursuit to the uniformity makes the contrastive loss not tolerant to semantically similar samples, which may break the underlying semantic structure and be harmful to the formation of features useful for downstream tasks. This is caused by the inherent defect of the instance discrimination objective. Specifically, instance discrimination objective tries to push all different instances apart, ignoring the underlying relations between samples. Pushing semantically consistent samples apart has no positive effect for acquiring a prior informative to general downstream tasks. A well-designed contrastive loss should have some extents of tolerance to the closeness of semantically similar samples. Therefore, we find that the contrastive loss meets a uniformity-tolerance dilemma, and a good choice of temperature can compromise these two properties properly to both learn separable features and tolerant to semantically similar samples, improving the feature qualities and the downstream performances. Feng Wang 0034, Huaping Liu 0001 |
CVPR | 2 |
| 2021 | Adversarial Skill Learning for Robust ManipulationabstractDeep reinforcement learning has made significant progress in robotic manipulation tasks and it works well in the ideal disturbance-free environment. However, in a real-world environment, both internal and external disturbances are inevitable, thus the performance of the trained policy will dramatically drop. To improve the robustness of the policy, we introduce the adversarial training mechanism to the robotic manipulation tasks in this paper, and an adversarial skill learning algorithm based on soft actor-critic (SAC) is proposed for robust manipulation. Extensive experiments are conducted to demonstrate that the learned policy is robust to internal and external disturbances. Additionally, the proposed algorithm is evaluated in both the simulation environment and on the real robotic platform. Pingcheng Jian, Chao Yang 0026, Di Guo 0002, Huaping Liu 0001, Fuchun Sun 0001 |
ICRA | 4 |
| 2021 | Robotic Indoor Scene Captioning from Streaming VideoabstractRobots are usually equipped with cameras to explore the indoor scene and it is expected that the robot can well describe the scene with natural language. Although some great success has been achieved in image and video captioning technology, especially on many public datasets, the caption generated from indoor scene video is still not informative and coherent enough. In this paper, we propose the problem of Indoor Scene Captioning from Streaming Video, which aims at generating a more accurate and informative caption from streaming video. To solve this problem, we firstly design an algorithm to organize the visual information of the indoor scene into a scene graph, and then implement a scene graph guided captioning method, which takes the scene graph and video frames as input to generate the caption from the video streaming. The proposed framework is evaluated both on the AI2THOR dataset and a real-world robotic platform, demonstrating the effectiveness of the framework. Xinghang Li, Di Guo 0002, Huaping Liu 0001, Fuchun Sun 0001 |
ICRA | 3 |
| 2021 | Neighborhood Spatial Aggregation based Efficient Uncertainty Estimation for Point Cloud Semantic SegmentationabstractUncertainty estimation for point cloud semantic segmentation is to quantify the confidence degree for the predicted label of points, which is essential for decision-making tasks. This paper proposes a neighborhood spatial aggregation based method, NSA-MC dropout, to achieve efficient uncertainty estimation for point cloud semantic segmentation. Unlike the traditional uncertainty estimation method MC dropout de-pending on repeated inferences, our NSA-MC dropout achieves uncertainty estimation through one-time inference. Specifically, a space-dependent method is designed to sample the model many times by performing stochastic forward pass through the model just once, and it approximates the repeated inferences based sampling process in MC dropout. Besides, a neighborhood spatial aggregation module, called NSA, aggregates neighborhood probabilistic outputs for each point and works with space-dependent sampling to establish output distribution. Finally, we propose an uncertainty-aware framework NSA-MC dropout to capture the uncertainty of prediction results efficiently. Experimental results show that our method obtains comparable performance with MC dropout. More significantly, our NSA-MC dropout has little influence on the efficiency of semantic inference. It is much faster than MC dropout, and the inference time does not establish a coupling relation with the sampling times. Our code is available at https://github.com/chaoqi7/Uncertainty_Estimation_PCSS Jianqin Yin, Huaping Liu 0001, Jun Liu 0007 |
ICRA | 3 |
| 2021 | Line-based Automatic Extrinsic Calibration of LiDAR and CameraabstractReliable real-time extrinsic parameters of 3D Light Detection and Ranging (LiDAR) and camera are a key component of multi-modal perception systems. However, extrinsic transformation may drift gradually during operation, which can result in decreased accuracy of perception system. To solve this problem, we propose a line-based method that enables automatic online extrinsic calibration of LiDAR and camera in real-world scenes. Herein, the line feature is selected to constrain the extrinsic parameters for its ubiquity. Initially, the line features are extracted and filtered from point clouds and images. Afterwards, an adaptive optimization is utilized to provide accurate extrinsic parameters. We demonstrate that line features are robust geometric features that can be extracted from point clouds and images, thus contributing to the extrinsic calibration. To demonstrate the benefits of this method, we evaluate it on KITTI benchmark with ground truth value. The experiments verify the accuracy of the calibration approach. In online experiments on hundreds of frames, our approach automatically corrects miscalibration errors and achieves an accuracy of 0.2 degrees, which verifies its applicability in various scenarios. This work can provide basis for perception systems and further improve the performance of other algorithms that utilize these sensors. Xinyu Zhang 0001, Shifan Zhu, Shichun Guo, Jun Li 0082, Huaping Liu 0001 |
ICRA | 5 |
| 2021 | Lifelong Localization in Semi-Dynamic EnvironmentabstractMapping and localization in non-static environments are fundamental problems in robotics. Most of previous methods mainly focus on static and highly dynamic objects in the environment, which may suffer from localization failure in semi-dynamic scenarios without considering objects with lower dynamics, such as parked cars and stopped pedestrians. In this paper, we introduce semantic mapping and lifelong localization approaches to recognize semi-dynamic objects in non-static environments. We also propose a generic framework that can integrate mainstream object detection algorithms with mapping and localization algorithms. The mapping method combines an object detection algorithm and a SLAM algorithm to detect semi-dynamic objects and constructs a semantic map that only contains semi-dynamic objects in the environment. During navigation, the localization method can classify observation corresponding to static and non-static objects respectively and evaluate whether those semi-dynamic objects have moved, to reduce the weight of invalid observation and localization fluctuation. Real-world experiments show that the proposed method can improve the localization accuracy of mobile robots in non-static scenarios. Shifan Zhu, Xinyu Zhang 0001, Shichun Guo, Jun Li 0082, Huaping Liu 0001 |
ICRA | 5 |
| 2021 | Toward Image-to-Tactile Cross-Modal Perception for Visually Impaired PeopleabstractIt is still a great challenge for the visually impaired people to perceive their surroundings from a global perspective, which makes it difficult for them to interact with unfamiliar environments. The reason is that these conventional assisting devices only address the obstacle avoidance problem. They do not provide visually impaired people with a global perception of the surrounding environment. In this article, a new generative adversarial network (GAN) model is developed to effectively transform the ground images into the tactile signal, which can be displayed by an off-the-shelf vibration device. The algorithm module and the hardware are integrated into a portable device, which provides visually impaired people with effective surrounding perception capability. In addition, a visual-tactile cross-modal data set is constructed to train the proposed deep-learning architecture. Experimental results show that the proposed system can help visually impaired people sense the ground and bring a better traveling experience for them. Note to Practitioners-This article presents a portable device that provides tactile recognition assistance for visually impaired people. Such a technology can be extensively used in tactile mouse and white cane. The developed technology can be extensively used for various industrial applications, such as surrounding monitoring and manipulation. The proposed work demonstrates the promising ability of artificial intelligence in healthcare applications. The generated tactile signals are expected to be used in many human-centered systems, and we believe that our contribution is an important step toward the development of a more comprehensive assisting technology for visually impaired people. Huaping Liu 0001, Di Guo 0002, Xinyu Zhang 0001, Wenlin Zhu, Bin Fang 0003, Fuchun Sun 0001 |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2021 | TrajectoryCNN: A New Spatio-Temporal Feature Learning Network for Human Motion PredictionabstractHuman motion prediction is an increasingly interesting topic in computer vision and robotics. In this paper, we propose a new end-to-end feedforward network, TrajectoryCNN, to predict future poses. Compared with the most existing methods, we introduce a new trajectory space and focus on modeling motion dynamics of the input sequence with coupled spatio-temporal features, dynamic local-global features, and global temporal co-occurrence features in the new space. Specifically, the coupled spatio-temporal features describe the spatial and temporal structural information hidden in a natural human motion sequence, which can be easily mined using CNN by simultaneously covering the spatial and temporal dimensions of the sequence with the convolutional filters. The dynamic local-global features encode different correlations among joint trajectories of human motion (i.e. strong correlations among joint trajectories of one part and weak correlations among joint trajectories of different parts), which can be captured by stacking multiple residual trajectory blocks and incorporating our skeletal representation. The global temporal co-occurrence features represent different importance of different input poses to mine the motion dynamics for predicting future poses, which can be obtained automatically by learning free parameters for each pose with our TrajectoryCNN. Finally, we predict future poses with the captured motion dynamic features in a non-recursive manner. Extensive experiments show that our method achieves state-of-the-art performance on five benchmarks (e.g. Human3.6M, CMU-Mocap, 3DPW, G3D, and FNTU), which demonstrates the effectiveness of our proposed method. The code is available at https://github.com/lily2lab/TrajectoryCNN.git. Jianqin Yin, Jin Li 0002, Pengxiang Ding, Jun Liu 0007, Huaping Liu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2021 | Energy-Based Periodicity Mining With Deep Features for Action Repetition Counting in Unconstrained VideosabstractAction repetition counting is to estimate the occurrence times of the repetitive motion in one action, which is a relatively new, significant, but challenging problem. To solve this problem, we propose a new method superior to the traditional ways in two aspects, without preprocessing and applicable for arbitrary periodicity actions. Without preprocessing, the proposed model makes our scheme convenient for real applications; processing the arbitrary periodicity action makes our model more suitable for the actual circumstance. In terms of methodology, firstly, we extract action features using ConvNets and then use Principal Component Analysis algorithm to generate the intuitive periodic information from the chaotic high-dimensional features; secondly, we propose an energy-based adaptive feature mode selection scheme to adaptively select proper deep feature mode according to the background of the video; thirdly,we construct the periodic waveform of the action based on the high-energy rules by filtering the irrelevant information. Finally, we detect the peaks to obtain the times of the action repetition. Our work features two-fold: 1) We give a significant insight that features extracted by ConvNets for action recognition can well model the self-similarity periodicity of the repetitive action. 2) A high-energy based periodicity mining rule using features from ConvNets is presented, which can process arbitrary actions without preprocessing. Experimental results show that our method achieves superior or comparable performance on the three benchmark datasets, i.e. YT_Segments, QUVA, and RARV. Jianqin Yin, Yanchun Wu, Chaoran Zhu, Zijin Yin, Huaping Liu 0001, Yonghao Dang, Zhiyi Liu, Jun Liu 0007 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2021 | Model Predictive Cooperative Control With ISM for Multiagent Systems Under Stochastic Communication ProtocolabstractIn this article, the cooperative control problem is investigated for the nonlinear multiagent system (MAS). For the purpose of avoiding possible data collisions, the stochastic communication protocol (SCP) is adopted to schedule the data transmission at each time instant. To deal with the unmatched disturbances, the composite control strategy is put forward which integrates the model predictive control (MPC) and the integral sliding-mode control methods. The sufficient conditions are established to guarantee the cooperative behavior of the MAS subjected to SCP scheduling. Furthermore, the parameters of the MPC scheme are selected such that the recursive feasibility and mean-square practical stability are guaranteed. Finally, the numerical simulation on the satellites is conducted to verify the effectiveness of the proposed methodology. Yuan Yuan 0006, Lei Guo 0003, Huaping Liu 0001 |
IEEE Trans. Cybern. | 3 |
| 2021 | Road-Network-Based Fast GeolocalizationabstractIn this article, a road-network-based geolocalization method is proposed. We match roads in the onboard images to the reference road vector map, and realize successful localization over areas as large as a whole city. The road network matching problem is treated as a point cloud registration problem under the homography transformation and solved under the hypothesize-and-test framework. To tackle the point cloud registration problem, a global projective-invariant feature is proposed, which consists of two road intersections augmented with their tangents. In addition, we propose the necessary conditions for the features to match. This can reduce the candidate matching features, thus accelerating the search to a great extent. These matching candidates are first “filtered” with the model consistency check in parameter space and then tested with similarity metrics to identify the correct transformation. The experiments show that our method can localize an aerial image over an area larger than 1000 km2within several seconds on a single CPU. Our code can be found at: https://github.com/FlyAlCode/RCLGeolocalization-2.0. Yongfei Li, Hao He 0005, Jiaxing Hu, Huaping Liu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2021 | An Interactive Perception Method for Warehouse Automation in Smart CitiesabstractThe smart city is an integrated environment that heavily relies on intelligent robots, which provides the basis for the warehouse automation. However, a warehouse is a typical unstructured environment, and robotic grasp and manipulation are extremely important for the package, transfer, search, and so on. Currently, the most usual method is to detect the picking or grasping points for some specific end-effector including suction cup, gripper, or robotic hand. The manipulation performance is, therefore, strongly influenced by the visual detector. To tackle this problem, the affordance map has recently been developed. It characterizes the operation possibilities afforded by the operation scene and has been used for several grasp tasks. Nevertheless, the conventional affordance method often fails in complicated environments due to the mistake calculation results. In this article, we develop a novel framework to integrate the interactive exploration with a composite robotic hand for robotic grasping in a complicated environment. The exploration strategy is obtained by a deep reinforcement learning procedure. The developed new composite hand, which integrates the suction cup and grippers, is used to test the merits of the proposed interactive perception method. Experimental results show the proposed method significantly increases the manipulation efficiency and may bring great economic and social and benefits for smart cities. Huaping Liu 0001, Yuhong Deng, Di Guo 0002, Bin Fang 0003, Fuchun Sun 0001, Wuqiang Yang |
IEEE Trans. Ind. Informatics | 1 |
| 2021 | Active Object Discovery and Localization Using Sound-Induced AttentionabstractIndustrial intelligent devices are usually equipped with both microphones and cameras to perceive and understand the physical world. Though visual object detection technology has achieved a great success, its combination with other sensing modalities remains unsolved. In this article, we establish a novel sound-induced attention framework for the visual object detection, and develop a two-stream weakly supervised deep learning architecture to combine the visual and audio modalities for localizing the sounding object. A dataset is constructed from the Audio Set to validate the proposed method and some realistic experiments are conducted to demonstrate the effectiveness of the proposed system. Huaping Liu 0001, Feng Wang 0034, Di Guo 0002, Xinzhu Liu, Xinyu Zhang 0001, Fuchun Sun 0001 |
IEEE Trans. Ind. Informatics | 1 |
| 2021 | Lifelong Visual-Tactile Cross-Modal Learning for Robotic Material PerceptionabstractThe material attribute of an object's surface is critical to enable robots to perform dexterous manipulations or actively interact with their surrounding objects. Tactile sensing has shown great advantages in capturing material properties of an object's surface. However, the conventional classification method based on tactile information may not be suitable to estimate or infer material properties, particularly during interacting with unfamiliar objects in unstructured environments. Moreover, it is difficult to intuitively obtain material properties from tactile data as the tactile signals about material properties are typically dynamic time sequences. In this article, a visual-tactile cross-modal learning framework is proposed for robotic material perception. In particular, we address visual-tactile cross-modal learning in the lifelong learning setting, which is beneficial to incrementally improve the ability of robotic cross-modal material perception. To this end, we proposed a novel lifelong cross-modal learning model. Experimental results on the three publicly available data sets demonstrate the effectiveness of the proposed method. Wendong Zheng, Huaping Liu 0001, Fuchun Sun 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2020 | Reinforcement Learning from Imperfect Demonstrations under Soft Expert GuidanceabstractIn this paper, we study Reinforcement Learning from Demonstrations (RLfD) that improves the exploration efficiency of Reinforcement Learning (RL) by providing expert demonstrations. Most of existing RLfD methods require demonstrations to be perfect and sufficient, which yet is unrealistic to meet in practice. To work on imperfect demonstrations, we first define an imperfect expert setting for RLfD in a formal way, and then point out that previous methods suffer from two issues in terms of optimality and convergence, respectively. Upon the theoretical findings we have derived, we tackle these two issues by regarding the expert guidance as a soft constraint on regulating the policy exploration of the agent, which eventually leads to a constrained optimization problem. We further demonstrate that such problem is able to be addressed efficiently by performing a local linear search on its dual form. Considerable empirical evaluations on a comprehensive collection of benchmarks indicate our method attains consistent improvement over other RLfD counterparts. Mingxuan Jing, Xiaojian Ma 0001, Wenbing Huang 0001, Fuchun Sun 0001, Chao Yang 0026, Bin Fang 0003, Huaping Liu 0001 |
AAAI | 7 |
| 2020 | Near-duplicated Loss for Accurate Object LocalizationabstractMulti-class object detection always involves the tasks of accurate target localization which is mainly related to bounding box regression. Smooth L1 loss is the most popular bounding box regression loss used in the current state-of-the-art object detection systems. However, such loss for regressing the parameters of a bounding box can't accurately and consistently regress the bounding box to the associated ground truth well. We instead propose the near-duplicated loss, a loss that better evaluate the disparity between the bounding box and the ground truth consistently. We present an approximate algorithm associated with a kernel function that not only considers the absolute distance but also involves the relative overlap area between the two bounding boxes. The new loss doesn't need additional supervision and is easy to embed into existing networks. Our final result, by incorporating the near-duplicated loss into the state-of-the-art object detection detectors (Faster RCNN, RetinaNet), shows consistent and significant improvements on popular object detection benchmarks (MS COCO and Pascal VOC). Xiaocheng Yang, Huaping Liu 0001, Tao Kong, Fuchun Sun 0001 |
DSAA | 3 |
| 2020 | Multi-agent Embodied Question Answering in Interactive Environments
Sinan Tan, Weilai Xiang, Huaping Liu 0001, Di Guo 0002, Fuchun Sun 0001 |
ECCV (13) | 3 |
| 2020 | Self-Supervised Learning for Alignment of Objects and SoundabstractThe sound source separation problem has many useful applications in the field of robotics, such as human-robot interaction, scene understanding, etc. However, it remains a very challenging problem. In this paper, we utilize both visual and audio information of videos to perform the sound source separation task. A self-supervised learning framework is proposed to implement the object detection and sound separation modules simultaneously. Such an approach is designed to better find the alignment between the detected objects and separated sound components. Our experiments, conducted on both the synthetic and real datasets, validate this approach and demonstrate the effectiveness of the proposed model in the task of object and sound alignment. Xinzhu Liu, Di Guo 0002, Huaping Liu 0001, Fuchun Sun 0001, Haibo Min |
ICRA | 4 |
| 2020 | Unsupervised Representation Learning by Invariance PropagationabstractUnsupervised learning methods based on contrastive learning have drawn increasing attention and achieved promising results. Most of them aim to learn representations invariant to instance-level variations, which are provided by different views of the same instance. In this paper, we propose Invariance Propagation to focus on learning representations invariant to category-level variations, which are provided by different instances from the same category. Our method recursively discovers semantically consistent samples residing in the same high-density regions in representation space. We demonstrate a hard sampling strategy to concentrate on maximizing the agreement between the anchor sample and its hard positive samples, which provide more intra-class variations to help capture more abstract invariance. As a result, with a ResNet-50 as the backbone, our method achieves 71.3% top-1 accuracy on ImageNet linear classification and 78.2% top-5 accuracy fine-tuning on only 1% labels, surpassing previous results. We also achieve state-of-the-art performance on other downstream tasks, including linear classification on Places205 and Pascal VOC, and transfer learning on small scale datasets. Feng Wang 0034, Huaping Liu 0001, Di Guo 0002, Fuchun Sun 0001 |
NeurIPS | 2 |
| 2020 | Deep learning for diplomatic video analysis
Fangxin Zhou, Huaping Liu 0001 |
Multim. Tools Appl. | 3 |
| 2020 | Audiovisual cross-modal material surface retrieval
Zhuokun Liu, Huaping Liu 0001, Wenmei Huang, Bowen Wang 0006, Fuchun Sun 0001 |
Neural Comput. Appl. | 2 |
| 2020 | Weakly-paired deep dictionary learning for cross-modal retrieval
Huaping Liu 0001, Feng Wang 0034, Xinyu Zhang 0001, Fuchun Sun 0001 |
Pattern Recognit. Lett. | 1 |
| 2020 | Cross-Modal Material Perception for Novel Objects: A Deep Adversarial Learning MethodabstractTo more actively perform fine manipulation tasks in the real world, intelligent robots should be able to understand and communicate the physical attributes of the material during interaction with an object. Tactile and vision are two important sensing modalities in robotic perception system. In this article, we propose a cross-modal material perception framework for recognizing novel objects. Concretely, it first adopts an object-agnostic method to associate information from tactile and visual modalities. It then recognizes a novel object by using its tactile signal to retrieve perceptually similar surface material images through the learned cross-modal correlation. This problem exhibits a challenge because data from visual and tactile modalities are highly heterogeneous and weakly paired. Moreover, the framework should not only consider cross-modal pairwise relevance but also be discriminative and generalized for unseen objects. To this end, we propose a weakly paired cross-modal adversarial learning (WCMAL) model for the visual–tactile cross-modal retrieval, which combines the advantages of deep learning and adversarial learning. In particular, the model fully considers the weak pairing problem between the two modalities. Finally, we conduct verification experiments on a publicly available data set. The results demonstrate the effectiveness of the proposed method.Note to Practitioners—Since cross-modal perception can improve the active operation of automation systems, it is invaluable for industrial intelligence, particularly when only one sensing modality cannot be used or suitable in some applications. In this article, we provide a framework of cross-modal material perception for object recognition using the idea of the cross-modal retrieval. Concretely, we use relevant tactile data of an unknown object to retrieve perceptually similar surface images, which are used to evaluate its material properties. Different from that previous works using tactile information as a complement or alternative to visual information to recognize specific objects, our proposed framework is able to estimate and infer material properties of both seen and unseen objects, which can enhance manipulation systems intelligence and improve the quality of the interaction. In our future works, more modality information will be incorporated to further enhance the cross-modal material perception. Wendong Zheng, Huaping Liu 0001, Bowen Wang 0006, Fuchun Sun 0001 |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2020 | Bioinspired Embodiment for Intelligent Sensing and Dexterity in Fine Manipulation: A SurveyabstractRecent advances in fine manipulation have led to increased interest in both scientific research works and engineering applications. Robot manipulation at a level approaching human skills is gaining attention in both industrial and individual services. A major challenge in fine manipulation is the unavoidable uncertainties and unpredictable conditions encountered in dynamic and unstructured application environments. The employment of biologically inspired (bioinspired) embodiments in fine manipulation shows significant advantages in tackling such problems. The aim of bioinspired embodiment is to improve fine manipulation of robotic systems utilizing the knowledge gained from natural systems with biomimetic methods. Such a method includes sensing, planning, and execution. This article provides a comprehensive survey of the current state of bioinspired technologies in fine manipulation, and outlines new challenges and some potential directions. Yueyue Liu 0001, Zhijun Li 0001, Huaping Liu 0001, Zhen Kan, Bugong Xu |
IEEE Trans. Ind. Informatics | 3 |
| 2020 | FoveaBox: Beyound Anchor-Based Object DetectionabstractWe present FoveaBox, an accurate, flexible, and completely anchor-free framework for object detection. While almost all state-of-the-art object detectors utilize predefined anchors to enumerate possible locations, scales and aspect ratios for the search of the objects, their performance and generalization ability are also limited to the design of anchors. Instead, FoveaBox directly learns the object existing possibility and the bounding box coordinates without anchor reference. This is achieved by: (a) predicting category-sensitive semantic maps for the object existing possibility, and (b) producing category-agnostic bounding box for each position that potentially contains an object. The scales of target boxes are naturally associated with feature pyramid representations. In FoveaBox, an instance is assigned to adjacent feature levels to make the model more accurate.We demonstrate its effectiveness on standard benchmarks and report extensive experimental analysis. Without bells and whistles, FoveaBox achieves state-of-the-art single model performance on the standard COCO and Pascal VOC object detection benchmark. More importantly, FoveaBox avoids all computation and hyper-parameters related to anchor boxes, which are often sensitive to the final detection performance. We believe the simple and effective approach will serve as a solid baseline and help ease future research for object detection. The code has been made publicly available athttps://github.com/taokong/FoveaBox. Tao Kong, Fuchun Sun 0001, Huaping Liu 0001, Yuning Jiang 0001, Lei Li 0005, Jianbo Shi |
IEEE Trans. Image Process. | 3 |
| 2020 | Cross-Modal Zero-Shot-Learning for Tactile Object RecognitionabstractIn this paper, we address the learning problem of classifying untouched tactile instance with the help of visual modality. The proposed method is based on dictionary learning and we impose different penalty terms on coding vectors between visual and tactile modalities. Using such structured coding vectors, the visual-tactile cross-modal transfer can be achieved. A set of optimization algorithms are developed to obtain the solutions of the proposed optimization problems. After then, we can use the obtained dictionary to predict the coding vectors of the new untouched tactile samples and further determine its label. Finally, we perform extensive experimental evaluations on publicly available datasets to show the effectiveness of the proposed method. Huaping Liu 0001, Fuchun Sun 0001, Bin Fang 0003, Di Guo 0002 |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2020 | Barrier Lyapunov Functions-Based Adaptive Fault Tolerant Control for Flexible Hypersonic Flight Vehicles With Full State ConstraintsabstractOne of the key problems for space vehicles is how to deal with the contradiction between the control constraints and the disturbances from multiple resources. This paper focuses on the adaptive full state constrained controller design problem for the flexible air-breathing hypersonic vehicles with multisource uncertainties. In simultaneous presence of the aerodynamical uncertainties, the modeling errors, the external disturbances, and the contingent actuator failures, an adaptive fault-tolerant control scheme is put forward in the context of dynamic surface control. Then, the barrier Lyapunov functions are utilized to guarantee that the constraints on the velocity, the flight path angle, the altitude, and the pitch rate are never violated. In virtue of the bound estimation approach, the multisource disturbances are effectively dealt with. Finally, a number of illustrative examples are provided to demonstrate the effectiveness of the proposed methodology. Yuan Yuan 0006, Zheng Wang 0033, Lei Guo 0003, Huaping Liu 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2019 | Task Transfer by Preference-Based Cost LearningabstractThe goal of task transfer in reinforcement learning is migrating the action policy of an agent to the target task from the source task. Given their successes on robotic action planning, current methods mostly rely on two requirements: exactlyrelevant expert demonstrations or the explicitly-coded cost function on target task, both of which, however, are inconvenient to obtain in practice. In this paper, we relax these two strong conditions by developing a novel task transfer framework where the expert preference is applied as a guidance. In particular, we alternate the following two steps: Firstly, letting experts apply pre-defined preference rules to select related expert demonstrates for the target task. Secondly, based on the selection result, we learn the target cost function and trajectory distribution simultaneously via enhanced Adversarial MaxEnt IRL and generate more trajectories by the learned target distribution for the next preference selection. The theoretical analysis on the distribution learning and convergence of the proposed algorithm are provided. Extensive simulations on several benchmarks have been conducted for further verifying the effectiveness of the proposed method. Mingxuan Jing, Xiaojian Ma 0001, Wenbing Huang 0001, Fuchun Sun 0001, Huaping Liu 0001 |
AAAI | 5 |
| 2019 | Deep Point-Wise Prediction for Action Temporal Proposal
Luxuan Li, Tao Kong, Fuchun Sun 0001, Huaping Liu 0001 |
ICONIP (3) | 4 |
| 2019 | Lifelong Learning for Heterogeneous Multi-Modal TasksabstractIn this work, we investigate the lifelong learning problem from the viewpoint of heterogeneous multi-modal fusion. The main challenges come from the fact that the common representation between heterogeneous modalities should be persistently learned and the learned classifier for each multi-modal task should be persistently updated. To address this problem, we construct a multi-modal lifelong learning framework which deals with the consecutive multi-modal learning tasks and develop an efficient online dictionary learning algorithm to solve the multi-modal lifelong learning problem. Finally, we perform experimental validation on a complicated material recognition task and show the promising results. Huaping Liu 0001, Fuchun Sun 0001, Bin Fang 0003 |
ICRA | 1 |
| 2019 | Sound-Indicated Visual Object Detection for Robotic ExplorationabstractRobots are usually equipped with microphones and cameras to perceive and understand the physical world. Though visual object detection technology has achieved great success, the detection in other modalities remains unsolved. In this paper, we establish a novel robotic sound-indicated visual object detection framework, and develop a two-stream weakly-supervised deep learning architecture to connect the visual and audio modalities for localizing the sounding object. A dataset is constructed from the AudioSet to validate the proposed method and some promising applications are demonstrated on robotic platforms. Feng Wang 0034, Di Guo 0002, Huaping Liu 0001, Junfeng Zhou, Fuchun Sun 0001 |
ICRA | 3 |
| 2019 | Deep Reinforcement Learning for Robotic Pushing and Picking in Cluttered EnvironmentabstractIn this paper, a novel robotic grasping system is established to automatically pick up objects in cluttered scenes. A composite robotic hand composed of a suction cup and a gripper is designed for grasping the object stably. The suction cup is used for lifting the object from the clutter first and the gripper for grasping the object accordingly. We utilize the affordance map to provide pixel-wise lifting point candidates for the suction cup. To obtain a good affordance map, the active exploration mechanism is introduced to the system. An effective metric is designed to calculate the reward for the current affordance map, and a deep Q-Network (DQN) is employed to guide the robotic hand to actively explore the environment until the generated affordance map is suitable for grasping. Experimental results have demonstrated that the proposed robotic grasping system is able to greatly increase the success rate of the robotic grasping in cluttered scenes. Yuhong Deng, Yixuan Wei, Kai Lu 0003, Bin Fang 0003, Di Guo 0002, Huaping Liu 0001, Fuchun Sun 0001 |
IROS | 7 |
| 2019 | Imitation Learning from Observations by Minimizing Inverse Dynamics DisagreementabstractThis paper studies Learning from Observations (LfO) for imitation learning with access to state-only demonstrations. In contrast to Learning from Demonstration (LfD) that involves both action and state supervisions, LfO is more practical in leveraging previously inapplicable resources (e.g., videos), yet more challenging due to the incomplete expert guidance. In this paper, we investigate LfO and its difference with LfD in both theoretical and practical perspectives. We first prove that the gap between LfD and LfO actually lies in the disagreement of inverse dynamics models between the imitator and expert, if following the modeling approach of GAIL. More importantly, the upper bound of this gap is revealed by a negative causal entropy which can be minimized in a model-free way. We term our method as Inverse-Dynamics-Disagreement-Minimization (IDDM) which enhances the conventional LfO method through further bridging the gap to LfD. Considerable empirical results on challenging benchmarks indicate that our method attains consistent improvements over other LfO counterparts. Chao Yang 0026, Xiaojian Ma 0001, Wenbing Huang 0001, Fuchun Sun 0001, Huaping Liu 0001, Junzhou Huang, Chuang Gan 0001 |
NeurIPS | 5 |
| 2019 | A glove-based system for object recognition via visual-tactile fusion
Bin Fang 0003, Fuchun Sun 0001, Huaping Liu 0001, Chuanqi Tan, Di Guo 0002 |
Sci. China Inf. Sci. | 3 |
| 2019 | A novel multi-modal tactile sensor design using thermochromic material
Fuchun Sun 0001, Bin Fang 0003, Hongxiang Xue, Huaping Liu 0001, Haiming Huang |
Sci. China Inf. Sci. | 4 |
| 2019 | Lifelong Learning for Scene Recognition in Remote Sensing ImagesabstractThe development of visual sensing technologies has made it possible to obtain some high resolution and to gather many high-resolution satellite images. To make the best use of these images, it is essential to be able to recognize and retrieve their intrinsic scene information. The problem of scene recognition in remote sensing images has recently aroused considerable interest, mainly due to the great success achieved by deep learning methods in generic image classification. Nevertheless, such methods usually require large amounts of labeled data. By contrast, remote sensing images are relatively scarce and expensive to obtain. Moreover, data sets from different aerospace research institutions exhibit large disparities. In order to address these problems, we propose a model based on a meta-learning method with the ability of learning a classifier from just few-shot samples. With the proposed model, the knowledge learned from one data set can be easily adapted to a new data set, which, in turn, would serve in the lifelong few-shot learning. Scene-level image recognition experiments, on public high-resolution remote sensing image data sets, validate our proposed lifelong few-shot learning model. Min Zhai, Huaping Liu 0001, Fuchun Sun 0001 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2019 | Interactive video summarization with human intentions
Huaping Liu 0001, Fuchun Sun 0001, Xinyu Zhang 0001, Bin Fang 0003 |
Multim. Tools Appl. | 1 |
| 2019 | Guest Editorial Special Issue on Active Perception for Industrial IntelligenceabstractInformation technologies are permeating all aspects of manufacturing systems as well as other fields, expediting the generation of industrial big data. Traditionally, the devices collected the sensor data from various sources, and information fusion was then performed. This incurs higher burden of time and storage cost. Recently, more and more intelligent devices are equipped in the industrial environment. This provides more opportunities for better data collection and processing for industrial intelligence. Active perception technology, which performs control strategies on the data acquisition process, enables the devices to seamlessly integrate the perception and action to reach high-level goals rather than to accomplish low-level commands. It helps to select more useful information and may save the life of the sensors. However, there exist many unsolved challenging problems since the feedback is performed on complex processed sensory data, i.e., various extracted features. Huaping Liu 0001, Nathan F. Lepora, Andrea Cherubini |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2019 | Surface Material Retrieval Using Weakly Paired Cross-Modal LearningabstractIn this paper, we investigate the cross-modal material retrieval problem, which permits the user to submit a multimodal query including tactile and auditory modalities, and retrieve the image results of visual modalities. Since multiple significantly different modalities are involved in this process, we encounter more challenges compared with the existing cross-modal retrieval tasks. Our focus is to learn cross-modal representations when the modalities are significantly different and with minimal supervision. A novelty is that we establish a framework that deals with weakly paired multimodal fusion method for heterogenous tactile and auditory modalities and weakly paired cross-modal transfer for visual modality. A structured dictionary learning method with a low rank and common classifier is developed to obtain the modal-invariant representation. Finally, some cross-modal validations on publicly available data sets are performed to show the advantages of the proposed method.Note to Practitioners—Cross-modal retrieval is an important task for industrial intelligence. In this paper, we establish a framework to effectively solve the cross-modal material retrieval problem. In the developed framework, the user may submit a multimodal query including acceleration and sound about an object, and the system may return the most relevant retrieved images. Such a framework may find extensive applications in many fields, because it can be flexible to deal with a multiple-modal query and uses the minimal category label supervision without the need of strong sample pairing information between modalities. Compared with the previous material analysis systems, this paper goes beyond previously proposed surface material classification approaches as it returns an ordered list of perceptually similar surface materials for a query. Huaping Liu 0001, Feng Wang 0034, Fuchun Sun 0001, Bin Fang 0003 |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2019 | One-Shot SADI-EPE: A Visual Framework of Event Progress EstimationabstractIn many practical engineering applications, the number of actions that have been finished should be known, particularly for an untrimmed video sequence that includes an event with a series of actions, it is important to know the number of actions that have been finished. In this paper, we termed this process as visual event progress estimation (EPE). However, the research related to this problem is few in the research community. To solve this problem, a visual human action analysis-based framework, namely one-shot simultaneously action detection and identification (SADI)-EPE, is presented in this paper. The visual EPE is modeled as an online one-shot learning-based problem; sliding window and attention-based bag of key poses formulate our framework. Unlike most of the action analysis methods relying on a number of training data of some predefined classes, our method can realize SADI for any event if one sample of the event is given, which makes it feasible for practical applications. At the same time, not only SADI but also the progress estimation of the event can be realized by our algorithm. In terms of methodology, the key pose is defined by an invariant pose descriptor from skeletal data and silhouette data. Moreover, in order to extract representative and discriminative poses from one training sample, we present a new bidirectional kNN-based attention weighted key pose selection method, which can filter the unrelated actions and model different importance of various key poses. In addition, an attention-based multi-modal fusion scheme, which addresses the difficulty of high-dimensional features and few training samples, is proposed to augment the performance of our algorithm. Finally, we propose an evaluation criterion for the estimation problem. Extensive results demonstrated the efficacy of our proposed framework. Jianqin Yin, Fuchun Sun 0001, Huaping Liu 0001, Bin Wang 0045, Jun Liu 0007, Yilong Yin |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2019 | Active Object Detection With Multistep Action Prediction Using Deep Q-NetworkabstractIn recent years, great success has been achieved in visual object detection, which is one of the fundamental tasks in the field of industrial intelligence. Most of existing methods have been proposed to deal with single well-captured still images, while in practical robotic applications, due to nuisances, such as tiny scale, partial view, or occlusion, one still image may not contain enough information for object detection. However, an intelligent robot has the capability to adjust its viewpoint to get better images for detection. Therefore, active object detection becomes a very important perception strategy for intelligent robots. In this paper, by formulating active object detection as a sequential action decision process, a deep reinforcement learning framework is established to resolve it. Furthermore, a novel deep Q-learning network (DQN) with a dueling architecture is proposed, the network has two separate output channels, one predicts action type and the other predicts action range. By combining the two output channels, the action space is explored more efficiently. Several methods are extensively validated and the results show that the proposed one obtains the best results and predicts action in real time. Xiaoning Han, Huaping Liu 0001, Fuchun Sun 0001, Xinyu Zhang 0001 |
IEEE Trans. Ind. Informatics | 2 |
| 2019 | Guest Editorial Special Issue on Bioinspired Embodiment for Intelligent Sensing and Dexterity in Fine ManipulationabstractThe papers in this special section focus on robotic manipulation based on bio-inspired computing. It is the goal of this papers to present applications of human manipulation ability in robotic systems, and to outline key strategies for robotic dexterous manipulation in next generation. Zhijun Li 0001, Huaping Liu 0001, Fanny Ficuciello |
IEEE Trans. Ind. Informatics | 2 |
| 2019 | Design and Output Characteristics of Magnetostrictive Tactile Sensor for Detecting Force and Stiffness of Manipulated ObjectsabstractA novel magnetostrictive tactile sensor has been designed, based on the inverse magnetostrictive effect and bionics, which can be used to test the gripping force of a manipulator and to detect the stiffness of the manipulated objects. Based on the electromagnetics theory, the inverse magnetostrictive effect, and the Hooke's law, the force measurement model and the stiffness detection model have been established. A magnetostrictive tactile sensor is designed to measure the applied force from 0 to 5 N with a sensitivity of 114 mV/N. The stiffness classification of the manipulated objects has been tested with four different types of sample materials. The results indicate that the output voltage and slope can be applied as a criteria for classifying the stiffness of the manipulated objects. The magnetostrictive tactile sensor has a simple structure with a rapid response and can realize the precise perception of the manipulated objects. Yunkai Li, Bowen Wang 0006, Ling Weng, Wenmei Huang, Huaping Liu 0001 |
IEEE Trans. Ind. Informatics | 7 |
| 2019 | Cross-Modal Surface Material Retrieval Using Discriminant Adversarial LearningabstractThe surface properties of an object play a vital role in the tasks of robotic manipulation or interaction with its surrounding environment. Tactile sensing can provide rich information about the surface properties of an object through physical contact. Hence, how to convey and interpret the tactile information to the user is a significant problem during the human–machine interaction. To this end, a visual–tactile cross-modal retrieval framework is proposed for perceptual estimation by associating tactile information to visual information of material surfaces. Namely, we can use tactile information of an unknown material surface to retrieve perceptually similar surfaces from an available surface visual sample set. For the proposed framework, we develop a discriminant adversarial learning method, which incorporates intramodal discriminant, cross-modal correlation, and intermodal consistency into a deep learning network for common feature representation learning. Experimental results on the publicly available data set show that the proposed framework and the method are effective. Wendong Zheng, Huaping Liu 0001, Bowen Wang 0006, Fuchun Sun 0001 |
IEEE Trans. Ind. Informatics | 2 |
| 2019 | Toward Efficient Action Recognition: Principal Backpropagation for Training Two-Stream NetworksabstractIn this paper, we propose the novel principal backpropagation networks (PBNets) to revisit the backpropagation algorithms commonly used in training two-stream networks for video action recognition. We content that existing approaches always take all the frames/snippets for the backpropagation not optimal for video recognition since the desired actions only occur in a short period within a video. To remedy these drawbacks, we design a watch-and-choose mechanism. In particular, the watching stage exploits a dense snippet-wise temporal pooling strategy to discover the global characteristic for each input video, while the choosing phase only backpropagates a small number of representative snippets that are selected with two novel strategies, i.e., Max-rule and KL-rule. We prove that with the proposed selection strategies, performing the backpropagation on the selected subset is capable of decreasing the loss of the whole snippets as well. The proposed PBNets are evaluated on two standard video action recognition benchmarks UCF101 and HMDB51, where it surpasses the state of the arts consistently, but requiring less memory and computation to achieve high performance. Wenbing Huang 0001, Lijie Fan, Mehrtash Harandi, Lin Ma 0002, Huaping Liu 0001, Wei Liu 0005, Chuang Gan 0001 |
IEEE Trans. Image Process. | 5 |
| 2019 | Feature Pyramid Reconfiguration With Consistent Loss for Object DetectionabstractTaking the feature pyramids into account has become a crucial way to boost the object detection performance. While various pyramid representations have been developed, previous works are still inefficient to integrate the semantical information over different scales. Moreover, recent object detectors are suffering from accurate object location applications, mainly due to the coarse definition of the "positive" examples at training and predicting phases. In this paper, we begin by analyzing current pyramid solutions, and then propose a novel architecture by reconfiguring the feature hierarchy in a flexible yet effective way. In particular, our architecture consists of two lightweight and trainable processes: global attention and local reconfiguration. The global attention is to emphasize the global information of each feature scale, while the local reconfiguration is to capture the local correlations across different scales. Both the global attention and local reconfiguration are non-linear and thus exhibit more expressive ability. Then, we discover that the loss function for object detectors during training is the central cause of the inaccurate location problem. We propose to address this issue by reshaping the standard cross entropy loss such that it focuses more on accurate predictions. Both the feature reconfiguration and the consistent loss could be utilized in popular one-stage (SSD, RetinaNet) and two-stage (Faster R-CNN) detection frameworks. Extensive experimental evaluations on PASCAL VOC 2007, PASCAL VOC 2012 and MS COCO datasets demonstrate that, our models achieve consistent and significant boosts compared with other state-of-the-art methods. Fuchun Sun 0001, Tao Kong, Wenbing Huang 0001, Chuanqi Tan, Bin Fang 0003, Huaping Liu 0001 |
IEEE Trans. Image Process. | 6 |
| 2019 | Near-Nash Equilibrium Control Strategy for Discrete-Time Nonlinear Systems With Round-Robin ProtocolabstractIn this paper, the near-Nash equilibrium (NE) control strategies are investigated for a class of discrete-time nonlinear systems subjected to the round-robin protocol (RRP). In the studied systems, three types of complexities, namely, the additive nonlinearities, the RRP, and the output feedback form of controllers, are simultaneously taken into consideration. To tackle these complexities, an approximate dynamic programing (ADP) algorithm is first developed for NE control strategies by solving the coupled Bellman's equation. Then, a Luenberger-type observer is designed under the RRP scheduling to estimate the system states. The near-NE control strategies are implemented via the actor-critic neural networks. More importantly, the stability analysis of the closed-loop system is conducted to guarantee that the studied system with the proposed control strategies is bounded stable. Finally, simulation results are provided to demonstrate the validity of the proposed method. Peng Zhang 0056, Yuan Yuan 0006, Hongjiu Yang, Huaping Liu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2019 | Kernel Regularized Nonlinear Dictionary Learning for Sparse CodingabstractFor most sparse coding methods, data samples are first encoded as hand-crafted features, followed by another separate learning step that generates dictionary and sparse codes. However, such feature representations may not be optimally compatible with the learning process, thus producing suboptimal results. In this paper, we propose a new architecture for nonlinear dictionary learning with sparse coding, in which samples are mapped into sparse codes via carefully designed stacked auto-encoder (SAE) networks. We jointly learn a low-dimensional embedding of the data samples by means of an SAE and a dictionary in the low-dimensional space. Further, to leverage the prior knowledge, we develop a kernel regularized nonlinear dictionary learning method, which effectively incorporates the knowledge provided by the hand-crafted kernel. An iterative algorithm is developed to jointly search the solutions of the associated optimization problem and extensive experimental validations are performed to show that the proposed kernel regularized dictionary learning method achieves satisfactory performance. Huaping Liu 0001, Fuchun Sun 0001, Bin Fang 0003 |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2018 | Deep Feature Pyramid Reconfiguration for Object Detection
Tao Kong, Fuchun Sun 0001, Wenbing Huang 0001, Huaping Liu 0001 |
ECCV (5) | 4 |
| 2018 | A Dual-Modal Vision-Based Tactile Sensor for Robotic Hand GraspingabstractHumans' fingertips can perceive not only the magnitude and the direction of force but also the texture of object. When we grasp an object, the surface texture sensing of the fingertip helps us recognize the object and the force feeling that is parallel to the skin helps us grasp stably. Focusing on these points, we have developed a dual-modal vision-based tactile sensor that can measure the texture of object and a distribution of force vectors. The tactile sensor consists of a transparent elastomer, a camera, a piece of transparent acrylic board, LEDs and supporting structures. A reflective membrane and markers array are on the surface of the elastomer. An applied force on the elastic body results in movements of the markers, which are acquired by the CCD camera. In addition, the shape and texture of the object's contact surface can be reflected by the membrane deformations. The distribution of force vectors is determined by the BP neural network. The local binary pattern algorithm using captured images calculates the texture information. This paper reports experimental evaluation results concerning accuracy of determination of magnitude, direction of force, and texture recognition rate. Bin Fang 0003, Fuchun Sun 0001, Chao Yang 0026, Hongxiang Xue, Wendan Chen, Chun Zhang 0001, Di Guo 0002, Huaping Liu 0001 |
ICRA | 8 |
| 2018 | Active Object Detection Using Double DQN and Prioritized Experience ReplayabstractVisual object detection is one of the fundamental tasks in computer vision and robotics. Small scale, partial capture and occlusion often occur in robotic applications, most existing object detection algorithms perform poorly in such situations. While a robot can look at one object from different views and plan its trajectory in the next few steps, which can lead to better observations. We formulate it as a sequential action-decision process, and develop a deep reinforcement learning architecture to solve the active object detection problem. A double deep Q-learning network (DQN) is applied to predict an action at each step. Experimental validation on the Active Vision Dataset shows the efficiency of the proposed method. Xiaoning Han, Huaping Liu 0001, Fuchun Sun 0001 |
IJCNN | 2 |
| 2018 | 3D human gesture capturing and recognition by the IMMU-based data glove
Bin Fang 0003, Fuchun Sun 0001, Huaping Liu 0001, Chunfang Liu |
Neurocomputing | 3 |
| 2018 | Multi-modal local receptive field extreme learning machine for object recognition
Huaping Liu 0001, Fengxue Li, Xinying Xu, Fuchun Sun 0001 |
Neurocomputing | 1 |
| 2018 | Multi-modal Fusion
Huaping Liu 0001, Amir Hussain 0001, Shuliang Wang 0001 |
Inf. Sci. | 1 |
| 2018 | Weakly paired multimodal fusion using multilayer extreme learning machine
Xiaohong Wen, Huaping Liu 0001, Gaowei Yan, Fuchun Sun 0001 |
Soft Comput. | 2 |
| 2018 | Weakly Paired Multimodal Fusion for Object RecognitionabstractThe ever-growing development of sensor technology has led to the use of multimodal sensors to develop robotics and automation systems. It is therefore highly expected to develop methodologies capable of integrating information from multimodal sensors with the goal of improving the performance of surveillance, diagnosis, prediction, and so on. However, real multimodal data often suffer from significant weak-pairing characteristics, i.e., the full pairing between data samples may not be known, while pairing of a group of samples from one modality to a group of samples in another modality is known. In this paper, we establish a novel projective dictionary learning framework for weakly paired multimodal data fusion. By introducing a latent pairing matrix, we realize the simultaneous dictionary learning and the pairing matrix estimation, and therefore improve the fusion effect. In addition, the kernelized version and the optimization algorithms are also addressed. Extensive experimental validations on some existing data sets are performed to show the advantages of the proposed method.Note to Practitioners—In many industrial environments, we usually use multiple heterogeneous sensors, which provide multimodal information. Such multimodal data usually lead to two technical challenges. First, different sensors may provide different patterns of data. Second, the full-pairing information between modalities may not be known. In this paper, we develop a unified model to tackle such problems. This model is based on a projective dictionary learning method, which efficiently produces the representation vector for the original data by an explicit form. In addition, the latent pairing relation between samples can be learned automatically and be used to improve the classification performance. Such a method can be flexibly used for multimodal fusion with full-pairing, partial-pairing and weak-pairing cases. Huaping Liu 0001, Yupei Wu, Fuchun Sun 0001, Bin Fang 0003, Di Guo 0002 |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2018 | Extreme Trust Region Policy Optimization for Active Object RecognitionabstractIn this brief, we develop a deep reinforcement learning method to actively recognize objects by choosing a sequence of actions for an active camera that helps to discriminate between the objects. The method is realized using trust region policy optimization, in which the policy is realized by an extreme learning machine and, therefore, leads to efficient optimization algorithm. The experimental results on the publicly available data set show the advantages of the developed extreme trust region optimization method. Huaping Liu 0001, Yupei Wu, Fuchun Sun 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2017 | RON: Reverse Connection with Objectness Prior Networks for Object DetectionabstractWe present RON, an efficient and effective framework for generic object detection. Our motivation is to smartly associate the best of the region-based (e.g., Faster R-CNN) and region-free (e.g., SSD) methodologies. Under fully convolutional architecture, RON mainly focuses on two fundamental problems: (a) multi-scale object localization and (b) negative sample mining. To address (a), we design the reverse connection, which enables the network to detect objects on multi-levels of CNNs. To deal with (b), we propose the objectness prior to significantly reduce the searching space of objects. We optimize the reverse connection, objectness prior and object detector jointly by a multi-task loss function, thus RON can directly predict final detection results from all locations of various feature maps. Extensive experiments on the challenging PASCAL VOC 2007, PASCAL VOC 2012 and MS COCO benchmarks demonstrate the competitive performance of RON. Specifically, with VGG-16 and low resolution 384×384 input size, the network gets 81.3% mAP on PASCAL VOC 2007, 80.7% mAP on PASCAL VOC 2012 datasets. Its superiority increases when datasets become larger and more difficult, as demonstrated by the results on the MS COCO dataset. With 1.5G GPU memory at test phase, the speed of the network is 15 FPS, 3 times faster than the Faster R-CNN counterpart. Code will be made publicly available. Tao Kong, Fuchun Sun 0001, Anbang Yao, Huaping Liu 0001, Ming Lu 0002, Yurong Chen 0001 |
CVPR | 4 |
| 2017 | From foot to head: Active face finding using deep Q-learningabstractIn the existing work on active face detection and tracking, it is usually required that the face has to appear in the field-of-view. However, this may not be practical in some challenging scenarios. In this paper, we formulate the problem of active face finding as a Markov Decision Process and resort to the deep Q-learning to solve it in an end-to-end manner. Under the proposed framework, the agent is able to learn how to adjust the control parameters of a camera in order to find the face. Even if the captured image contains only some parts of the person, the PTZ camera can still adjust its pose until the face is found. Extensive experimental validations are performed to show the effectiveness of the developed system. Hui Zhang 0092, Huaping Liu 0001, Di Guo 0002, Fuchun Sun 0001 |
ICIP | 2 |
| 2017 | A hybrid deep architecture for robotic grasp detectionabstractThe robotic grasp detection is a great challenge in the area of robotics. Previous work mainly employs the visual approaches to solve this problem. In this paper, a hybrid deep architecture combining the visual and tactile sensing for robotic grasp detection is proposed. We have demonstrated that the visual sensing and tactile sensing are complementary to each other and important for the robotic grasping. A new THU grasp dataset has also been collected which contains the visual, tactile and grasp configuration information. The experiments conducted on a public grasp dataset and our collected dataset show that the performance of the proposed model is superior to state of the art methods. The results also indicate that the tactile data could help to enable the network to learn better visual features for the robotic grasp detection task. Di Guo 0002, Fuchun Sun 0001, Huaping Liu 0001, Tao Kong, Bin Fang 0003, Ning Xi 0001 |
ICRA | 3 |
| 2017 | Multi-label tactile property analysisabstractIn this paper, we exploit the intrinsic relation between different adjective labels and develop a novel multilabel dictionary learning and sparse coding method which is improved by introducing the structured output association information. Such a method makes use of the label correlation information and is more suitable for the multi-label tactile understanding task. In addition, we develop a globally-convergent iterative algorithms to solve the dictionary learning problem. Finally, we perform extensive experimental validations on the public available tactile sequence dataset PHAC-2 and show the advantages of the proposed method. Huaping Liu 0001, Yupei Wu, Fuchun Sun 0001, Di Guo 0002, Bin Fang 0003 |
ICRA | 1 |
| 2017 | Visual-Tactile Fusion for Object RecognitionabstractThe camera provides rich visual information regarding objects and becomes one of the most mainstream sensors in the automation community. However, it is often difficult to be applicable when the objects are not visually distinguished. On the other hand, tactile sensors can be used to capture multiple object properties, such as textures, roughness, spatial features, compliance, and friction, and therefore provide another important modality for the perception. Nevertheless, effective combination of the visual and tactile modalities is still a challenging problem. In this paper, we develop a visual–tactile fusion framework for object recognition tasks. This paper uses the multivariate-time-series model to represent the tactile sequence and the covariance descriptor to characterize the image. Further, we design a joint group kernel sparse coding (JGKSC) method to tackle the intrinsically weak pairing problem in visual–tactile data samples. Finally, we develop a visual–tactile data set, composed of 18 household objects for validation. The experimental results show that considering both visual and tactile inputs is beneficial and the proposed method indeed provides an effective strategy for fusion. Huaping Liu 0001, Yuanlong Yu 0001, Fuchun Sun 0001, Jason Gu |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2017 | An Efficient Method for Traffic Sign Recognition Based on Extreme Learning MachineabstractThis paper proposes a computationally efficient method for traffic sign recognition (TSR). This proposed method consists of two modules: 1) extraction of histogram of oriented gradient variant (HOGv) feature and 2) a single classifier trained by extreme learning machine (ELM) algorithm. The presented HOGv feature keeps a good balance between redundancy and local details such that it can represent distinctive shapes better. The classifier is a single-hidden-layer feedforward network. Based on ELM algorithm, the connection between input and hidden layers realizes the random feature mapping while only the weights between hidden and output layers are trained. As a result, layer-by-layer tuning is not required. Meanwhile, the norm of output weights is included in the cost function. Therefore, the ELM-based classifier can achieve an optimal and generalized solution for multiclass TSR. Furthermore, it can balance the recognition accuracy and computational cost. Three datasets, including the German TSR benchmark dataset, the Belgium traffic sign classification dataset and the revised mapping and assessing the state of traffic infrastructure (revised MASTIF) dataset, are used to evaluate this proposed method. Experimental results have shown that this proposed method obtains not only high recognition accuracy but also extremely high computational efficiency in both training and recognition processes in these three datasets. Zhiyong Huang 0005, Yuanlong Yu 0001, Jason Gu, Huaping Liu 0001 |
IEEE Trans. Cybern. | 4 |
| 2017 | Extreme Kernel Sparse Learning for Tactile Object RecognitionabstractTactile sensors play very important role for robot perception in the dynamic or unknown environment. However, the tactile object recognition exhibits great challenges in practical scenarios. In this paper, we address this problem by developing an extreme kernel sparse learning methodology. This method combines the advantages of extreme learning machine and kernel sparse learning by simultaneously addressing the dictionary learning and the classifier design problems. Furthermore, to tackle the intrinsic difficulties which are introduced by the representer theorem, we develop a reduced kernel dictionary learning method by introducing row-sparsity constraint. A globally convergent algorithm is developed to solve the optimization problem and the theoretical proof is provided. Finally, we perform extensive experimental validations on some public available tactile sequence datasets and show the advantages of the proposed method. Huaping Liu 0001, Fuchun Sun 0001, Di Guo 0002 |
IEEE Trans. Cybern. | 1 |
| 2017 | Structured Output-Associated Dictionary Learning for Haptic UnderstandingabstractHaptic sensing and feedback play extremely important roles for humans and robots to perceive, understand, and manipulate the world. Since many properties perceived by the haptic sensors can be characterized by adjectives, it is reasonable to develop a set of haptic adjectives for the haptic understanding. This formulates the haptic understanding as a multilabel classification problem. In this paper, we exploit the intrinsic relation between different adjective labels and develop a novel dictionary learning method which is improved by introducing the structured output association information. Such a method makes use of the label correlation information and is more suitable for the multilabel haptic understanding task. In addition, we develop two iterative algorithms to solve the dictionary learning and classifier design problems, respectively. Finally, we perform extensive experimental validations on the public available haptic sequence dataset Penn Haptic Adjective Corpus 2 and show the advantages of the proposed method. Huaping Liu 0001, Fuchun Sun 0001, Di Guo 0002, Bin Fang 0003, Zhengchun Peng |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2016 | Sparse Coding and Dictionary Learning with Linear Dynamical SystemsabstractLinear Dynamical Systems (LDSs) are the fundamental tools for encoding spatio-temporal data in various disciplines. To enhance the performance of LDSs, in this paper, we address the challenging issue of performing sparse coding on the space of LDSs, where both data and dictionary atoms are LDSs. Rather than approximate the extended observability with a finite-order matrix, we represent the space of LDSs by an infinite Grassmannian consisting of the orthonormalized extended observability subspaces. Via a homeomorphic mapping, such Grassmannian is embedded into the space of symmetric matrices, where a tractable objective function can be derived for sparse coding. Then, we propose an efficient method to learn the system parameters of the dictionary atoms explicitly, by imposing the symmetric constraint to the transition matrices of the data and dictionary systems. Moreover, we combine the state covariance into the algorithm formulation, thus further promoting the performance of the models with symmetric transition matrices. Comparative experimental evaluations reveal the superior performance of proposed methods on various tasks including video classification and tactile recognition. Wenbing Huang 0001, Fuchun Sun 0001, Le-le Cao, Deli Zhao, Huaping Liu 0001, Mehrtash Harandi |
CVPR | 5 |
| 2016 | Object discovery and grasp detection with a shared convolutional neural networkabstractGrasp an object from a stack of objects in real-time is still a challenge in robotics. This requires the robot to have the ability of both fast object discovery and grasp detection: a target object should be picked out from the stack first and then a proper grasp configuration is applied to grasp the object. In this paper, we propose a shared convolutional neural network (CNN) which can simultaneously implement these two tasks in real-time. The processing speed of the model is about 100 frames per second on a GPU which largely satisfies the requirement. Meanwhile, we also establish a labeled RGBD dataset which contains scenes of stacked objects for robotic grasping. At last, we demonstrate the implementation of our shared CNN model on a real robotic platform and show that the robot can accurately discover a target object from the stack and successfully grasp it. Di Guo 0002, Tao Kong, Fuchun Sun 0001, Huaping Liu 0001 |
ICRA | 4 |
| 2016 | Learning Stable Linear Dynamical Systems with the Weighted Least Square Method
Wenbing Huang 0001, Le-le Cao, Fuchun Sun 0001, Deli Zhao, Huaping Liu 0001 |
IJCAI | 5 |
| 2016 | Multi-Modal Local Receptive Field Extreme Learning Machine for object recognitionabstractLearning rich representations efficiently plays an important role in multi-modal recognition task, which is crucial to achieve high generalization performance. To address this problem, in this paper, we propose an effective Multi-Modal Local Receptive Field Extreme Learning Machine (MM-ELM-LRF) structure, while maintaining ELM's advantages of training efficiency. In this structure, ELM-LRF is firstly conducted for feature extraction for each modality separately. And then, the shared layer is developed by combining these features from each modality. Finally, the Extreme Learning Machine (ELM) is used as supervised feature classifier for the final decision. Experimental validation on Washington RGB-D Object Dataset illustrates that the proposed multiple modality fusion method achieves better recognition performance. Fengxue Li, Huaping Liu 0001, Xinying Xu, Fuchun Sun 0001 |
IJCNN | 2 |
| 2016 | Nonlinear non-negative matrix factorization using deep learningabstractIn this paper, we describe the deep learning method to reduce the dimension of the data samples under the framework Non-negative Matrix Factorization (NMF). That is to say, we try to find the good representation of the data samples for the task of NMF. To this end, a nonlinear NMF optimization model is constructed and the optimization algorithm is developed. The experimental results on some benchmark dataset show the nonlinear dimension reduction helps the NMF to improve the clustering performance. Hui Zhang 0092, Huaping Liu 0001, Rui Song 0002, Fuchun Sun 0001 |
IJCNN | 2 |
| 2016 | Nonlinear dictionary learning based deep neural networksabstractIn this paper, we demonstrate nonlinear features extracted by deep neural network have better results in the task of dictionary learning. A nonlinear dictionary learning model is constructed and the optimization algorithm is developed. In the learning algorithm, we use the deep neural network to convey raw samples to feature space and learn a nonlinear dictionary. The extensive experimental results of classification on some benchmark dataset show the proposed nonlinear dictionary significantly improves the classification performance. Hui Zhang 0092, Huaping Liu 0001, Rui Song 0002, Fuchun Sun 0001 |
IJCNN | 2 |
| 2016 | Sonar-based place recognition using joint sparse coding methodabstractThe problem of place recognition is central to robot navigation. The robot needs to be able to recognize or at least to be able to estimate the likelihood that it has been at a place before when it has returned to a previously visited place. We cast the place recognition problem as one of classifying among multiple linear regression models, and argue that new theory from sparse signal representation offers the key to addressing the problem. In this paper, a joint kernel sparse coding model is developed to tackle the multivariate sonar samples place recognition problem. The experimental results show that the joint sparse coding achieves better performance than 1-Nearest Neighborhood (1-NN) method. Xiangmei Zheng, Huaping Liu 0001, Fuchun Sun 0001, Meng Gao 0002, Jiakui Li |
IJCNN | 2 |
| 2016 | Extreme learning machine for time sequence classification
Huaping Liu 0001, Lianzhi Yu, Fuchun Sun 0001 |
Neurocomputing | 1 |
| 2016 | Dynamic texture video classification using extreme learning machine
Liuyang Wang, Huaping Liu 0001, Fuchun Sun 0001 |
Neurocomputing | 2 |
| 2016 | Video key-frame extraction for smart phones
Huaping Liu 0001, Fuchun Sun 0001 |
Multim. Tools Appl. | 1 |
| 2015 | Discovery of topical object in image collectionsabstractAutomatic discovery of topical objects from a set of image collections provides more strong cognitive capability of robot to understand the unstructured environment. In this paper, we propose a novel framework based on dictionary learning for such a task. Different from existing work which utilizes multiple segmentations to coarsely obtain the object regions, we adopt the most recently developed objectness operator to extract candidate objects. Such a method admits a great advantage that the interested objects can be more reliably segmented. A dictionary learning method is proposed to discover the topical objects. Such an optimization model exploits the observation that any image only includes a few topical objects and therefore sparsity is encouraged. Further, a globally convergent algorithm is developed to solve the dictionary learning problem and extensive experiments show that the proposed method outperforms the state-of-the-arts. Huaping Liu 0001, Yunhui Liu 0003, Liming Huang, Fuchun Sun 0001, Di Guo 0002 |
ICRA | 1 |
| 2015 | Scalable Gaussian Process Regression Using Deep Neural Networks
Wenbing Huang 0001, Deli Zhao, Fuchun Sun 0001, Huaping Liu 0001, Edward Y. Chang |
IJCAI | 4 |
| 2015 | Robust Kernel Dictionary Learning Using a Whole Sequence Convergent Algorithm
Huaping Liu 0001, Hong Cheng 0002, Fuchun Sun 0001 |
IJCAI | 1 |
| 2015 | Tactile sequence classification using joint kernel sparse codingabstractTactile sensors in the robotic fingertips are used to capture multiple object properties such as texture, roughness, spatial features, compliance or friction and therefore becomes a very important sense modality for intelligent robot. However, existing work neglects the intrinsic relation between different fingers which simultaneously contact the object. In this paper, a joint kernel sparse coding model is developed to tackle the multi-finger tactile sequence classification problem. In this model, the intrinsic relations between fingers are explicitly considered using the joint sparse coding which encourages different modal coding to share the same support. The experimental results show that the joint sparse coding achieves better performance than conventional sparse coding. Huaping Liu 0001, Fuchun Sun 0001, Meng Gao 0002 |
IJCNN | 2 |
| 2015 | Aerial Scene Classification with Convolutional Neural NetworksabstractA robust satellite image classification is the fundamental step for aerial image understanding. However current methods with hand-crafted features and conventional classifiers have limited performance. In this paper we introduced convolutional neural network (CNN) method into this problem. Two approaches, including using conventional classifier with CNN features and direct classification with trained CNN models, are investigated with experiments. Our method achieved 97.4% accuracy on 5-fold cross-validation test of the UCMERCED LULC dataset, which is 8% higher than state-of-the-art methods. Sibo Jia, Huaping Liu 0001, Fuchun Sun 0001 |
ISNN | 2 |
| 2015 | Representative Video Action Discovery Using Interactive Non-negative Matrix FactorizationabstractIn this paper, we develop an interactive Non-negative Matrix Factorization method for representative action video discovery. The original video is first evenly segmented into some short clips and the bag-of-words model is used to describe each clip. Then a temporal consistent Non-negative Matrix Factorization model is used for clustering and action segmentation. Since the clustering and segmentation results may not satisfy the user’s intention, two extra human operations: MERGE and ADD are developed to permit user to improve the results. The newly developed interactive Non-negative Matrix Factorization method can therefore generate personalized results. Experimental results on the public Weizman dataset demonstrate that our approach is able to improve the action discovery and segmentation results. Hui Teng, Huaping Liu 0001, Lianzhi Yu, Fuchun Sun 0001 |
ISNN | 2 |
| 2015 | RGB-D action recognition using linear coding
Huaping Liu 0001, Mingyi Yuan, Fuchun Sun 0001 |
Neurocomputing | 1 |
| 2015 | Robust Exemplar Extraction Using Structured Sparse CodingabstractRobust exemplar extraction from the noisy sample set is one of the most important problems in pattern recognition. In this brief, we propose a novel approach for exemplar extraction through structured sparse learning. The new model accounts for not only the reconstruction capability and the sparsity, but also the diversity and robustness. To solve the optimization problem, we adopt the alternating directional method of multiplier technology to design an iterative algorithm. Finally, the effectiveness of the approach is demonstrated by experiments of various examples including traffic sign sequences. Huaping Liu 0001, Yunhui Liu 0003, Fuchun Sun 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2014 | Sparse fuzzy c-regression models with application to T-S fuzzy systems identificationabstractIn this paper, the objective function of fuzzy c-regression models (FCRM) is modified to develop a novel fuzzy partition method on the basis of block structured sparse representation, namely as sparse fuzzy c-regression model. This method takes advantage of the block structured information in the objective function of FCRM and casts fuzzy partition into an optimization problem by making a tradeoff between traditional FCRM and the number of prototypes of hyper-plane with nonzero parameters. An alternating direction method of multipliers (ADMM) based algorithm is exploited to address the proposed optimization problem. Furthermore, based on sparse fuzzy c-regression models, a novel T-S fuzzy systems identification method is developed for reduction of fuzzy rules. Finally, examples on well-known benchmark data set are carried out to illustrate the effectiveness of the proposed methods. Minnan Luo, Fuchun Sun 0001, Huaping Liu 0001 |
FUZZ-IEEE | 3 |
| 2014 | Dynamie texture classification using local fuzzy codingabstractRecognition of complex dynamic texture is a challenging problem and captures the attention of the computer vision community for several decades. Essentially the dynamic texture recognition is a multi-class classification problem that has become a real challenge for computer vision and machine learning techniques. In this paper, we propose a new approach to tackle the dynamic texture recognition problem. First, we utilize the fuzzy clustering technology to design a fuzzy codebook, and then construct a soft assigned local fuzzy coding feature to represent the whole dynamic texture sequence. This new coding strategy preserves spatial and temporal characteristics of dynamic texture. Finally, by evaluating the proposed approach using with the DynTex dataset, we show the effectiveness of the proposed local fuzzy coding strategy. Liuyang Wang, Huaping Liu 0001, Fuchun Sun 0001 |
FUZZ-IEEE | 2 |
| 2014 | The intelligent grasping tactics of dexterous handabstractIn the task of grasping objects with robotic hand, power grasp, also called whole hand grasp, is usually adopted. However, power grasp may incur severe target object deformation. To address this issue, a hybrid control algorithm is proposed, which could control dexterous robotic hand at different stages. When robotic hand does not contact target object, the speed of the fingers increases sharply under the position control law, guaranteeing the rapidness of the system. During the close contact process, the features of tactile data are analyzed and matrix models that corresponds to different types of objects are set up; and then corresponding integral separation PID control law is designed meticulously. The proposed method is verified by the experiment platform, showing a promising performance. Lichong Lei, Meng Gao 0002, Huaping Liu 0001, Fuchun Sun 0001 |
ICARCV | 3 |
| 2014 | Likelihood confidence rating based multi-modal information fusion for robot fine operationabstractMulti-modal information fusion plays an important role in many robotic applications, such as target grasping, manipulation and fine operation. Traditional fusion strategies, e.g. Bayesian fusion, directly adopt each uni-modal likelihood without giving enough attention to the fact that all these likelihoods are often vulnerable to sample data and modality-specific identification algorithm, which could possibly incur inaccuracy of, say, target recognition in a practical application. To address this issue, the paper presents a likelihood confidence rating strategy to fix traditional Bayesian fusion. Due to the great importance to the modalities with more accurate likelihoods, the strategy is capable of assigning different weights to each modality meticulously. We extensively evaluate the proposed strategy on our dextrous robotic hand testbed. The results demonstrate that the proposed method can achieve significant improvement in terms of fused classification performance. Fuchun Sun 0001, Huaping Liu 0001 |
ICARCV | 4 |
| 2014 | Outlier-attenuating summarization for user-generated-videoabstractIn this paper, the key-frame extraction problem for user-generated-videos which are captured by smart phones is investigated. A collaborative sparse coding model which incorporates the 1/2, 1 and Li, 2 regularization terms are proposed to select few key-frames while attenuating the influences of the outlier frames. Further, the sensors embedded in the smart phone is used to collect the acceleration values, which can be used to improve the performance of outlier-attenuations. Finally, a real dataset is constructed to test the proposed method and the experimental validation shows promising results. Huaping Liu 0001, Yunhui Liu 0003, Fuchun Sun 0001 |
ICME | 2 |
| 2014 | Simultaneous prototype selection and outlier isolation for traffic sign recognition: A collaborative sparse optimization methodabstractVideo-based traffic sign recognition is one of the most important task for unmanned autonomous vehicle. However, there always exists unavoidable outliers in the practical scenario. Therefore, robust prototype extraction from the noisy sample set is highly expected to help traffic sign recognition in video sequence. In this paper, we propose a novel approach for simultaneous prototype extraction and outlier isolation through collaborative sparse learning. The new model accounts for not only the reconstruction capability and the sparsity, but also the robustness. To solve the optimization problem, we adopt the Alternating Directional Method of Multiplier (ADMM) technology to design an iterative algorithm. Finally, the effectiveness of the approach is demonstrated by experiments on GTSRB dataset. Huaping Liu 0001, Yuanlong Yu 0001, Fuchun Sun 0001 |
ICRA | 1 |
| 2014 | User-generated-video summarization using Sparse ModellingabstractA novel key-frame extraction method is proposed in this paper. Our method focused on user-generated-videos which were captured by smartphones or tablets or other smart devices which can record acceleration values and orientation values during video capturing. Our method use Dissimilarity-based Sparse Modeling Representative Selection(DSMRS) on orientation information to extract key-frames instead of visual features used by traditional key-frame extraction methods. Acceleration value is used in our method to exclude outliers. Huaping Liu 0001, Yunhui Liu 0003, Fuchun Sun 0001 |
IJCNN | 2 |
| 2014 | Human activity recognition using smart phone embedded sensors: A Linear Dynamical Systems methodabstractThis paper presents a novel framework of human activity recognition with time series collected from inertial sensors. We model each action sequence with a collection of Linear Dynamic Systems (LDSs), each LDS describing a small patch of the sequence. A codebook is formed by using the K-medoids clustering algorithm and a Bag-of-Systems (BoS) is developed to represent the time series. A great advantage of this method is that the complicated feature design procedure is avoided and the LDSs can well capture the dynamics of the time series. Our experiment validation on public dataset shows the promising results. Huaping Liu 0001, Lianzhi Yu, Fuchun Sun 0001 |
IJCNN | 2 |
| 2014 | Structured sparse coding method for infrared small target detection in video sequenceabstractIn this paper, the infrared small target detection in video sequence is investigated. A collaborative structured sparse coding model which incorporates the L1,2and L2,1regularization terms is proposed to detect the infrared small target in video sequence. Further, online dictionary learning is embedded into the model and temporal information is incorporated to eliminate the clutters and noises. Finally, four simulation datasets are constructed to test the proposed method and the experimental validation shows promising results. Chunwei Yang, Huaping Liu 0001, Shouyi Liao |
IJCNN | 2 |
| 2014 | Linear dynamic system method for tactile object classification
Huaping Liu 0001, Fuchun Sun 0001, Qingfen Yang, Meng Gao 0002 |
Sci. China Inf. Sci. | 2 |
| 2014 | Traffic sign recognition using group sparse coding
Huaping Liu 0001, Fuchun Sun 0001 |
Inf. Sci. | 1 |
| 2014 | Recursive depth parametrization of monocular visual navigation: Observability analysis and performance evaluation
Fuchun Sun 0001, Jinsheng Zhang, Huaping Liu 0001 |
Inf. Sci. | 5 |
| 2014 | Joint Block Structure Sparse Representation for Multi-Input-Multi-Output (MIMO) T-S Fuzzy System IdentificationabstractTakagi–Sugeno (T–S) fuzzy system identification provides a reasonable framework for modeling and approximation by decomposition of a nonlinear system into a collection of local linear models. In this paper, we extend the multidimensional output fuzzy rule from the corresponding single-output fuzzy rule with the important difference in the variable definition of output, namely the output of each fuzzy rule is represented by a multidimensional vector instead of a scalar value. In this case, the consequent of the fuzzy rule with the multidimensional output variable share a common antecedent part. Different from the traditional methods that separate the multi-input–multi-output fuzzy system into a group of multi-input-single-output fuzzy models, in this paper, we take into account the block structure information in the T–S fuzzy system and cast the problem of multidimensional output fuzzy model identification as a joint structure sparse optimization problem, where the consequent parameters are estimated with a common block structured sparsity pattern over all dimensions of the output variable. Furthermore, we exploit a joint block sparse orthogonal-match-pursuit algorithm to reduce the number of fuzzy rules in terms of all dimensions of the output variable and prove the sufficient conditions in consideration of the multidimensional output together with the block structure in the T–S fuzzy model. This method is efficient and shows good performance in well-known benchmark datasets and real-world problems. Minnan Luo, Fuchun Sun 0001, Huaping Liu 0001 |
IEEE Trans. Fuzzy Syst. | 3 |
| 2014 | Diversified Key-Frame Selection Using Structured ${L_{2, 1}}$ OptimizationabstractIn this paper, a structured L2,1optimization model, which simultaneously characterizes the reconstruction capability and diversity, is proposed to provide a semantically meaningful representation of a short video clip acquired from digital cameras or a mobile robot. In this model, a mutual inhabitation penalty term is imposed to prevent similar samples from being selected simultaneously. The proposed model is highly flexible to incorporate different mutual inhabitation terms and the temporal redundancy in video is exploited to encourage the diversity. The constructed objective function is nonconvex and an iterative algorithm is developed to solve the optimization problem. The performance is evaluated using various video clips from YouTube and also based on practical video captured by an indoor mobile robot. The results clearly indicate that the proposed strategy helps the optimization model to achieve more diversified key frames than the other existing work method. Huaping Liu 0001, Yunhui Liu 0003, Yuanlong Yu 0001, Fuchun Sun 0001 |
IEEE Trans. Ind. Informatics | 1 |
| 2014 | Spatial Neighborhood-Constrained Linear Coding for Visual Object TrackingabstractIn this paper, a new spatial neighborhood-constrained linear coding strategy which realizes sparse representation is proposed for visual object tracking. Unlike conventional sparse and locality-constrained linear coding approaches that need an extra post-processing stage to incorporate the spatial layout information, the proposed coding strategy intrinsically embeds the spatial layout information into the coding stage. The proposed coding strategy can also be used to effectively realize joint sparse representation for different feature descriptors. In addition, based on the distance to the “ideal point” in the reconstruction error space, a new multicue integration approach for robust tracking is proposed and a co-learning approach is developed to update the dictionaries. Finally, the proposed tracking algorithm is compared with other state-of-the-art trackers on some challenging video sequences and shows promising results. Huaping Liu 0001, Mingyi Yuan, Fuchun Sun 0001, Jianwei Zhang 0001 |
IEEE Trans. Ind. Informatics | 1 |
| 2013 | Traffic sign detection by ROI extraction and histogram features-based recognitionabstractWe present a traffic sign detection model consisting of two modules. The first module is for ROI (region of interest) extraction. By supervised learning, it transforms the color images to gray images such that the characteristic colors for the traffic signs are more distinguishable in the gray images. It follows shape template matching, where a set of templates for each target category of signs are designed. After that, a set of ROIs are generated. The second module is for recognition. It validates if an ROI belongs to a target category of traffic signs by supervised learning. Local shape and color features are extracted. The supervised learning methods used in the model are SVMs. The overall model is applied on the GTSDB benchmark and achieves 100%, 98.85% and 92.00% AUC (area under the precision-recall curve) for Prohibitory, Danger and Mandatory signs, respectively. The testing speed is 0.4-1.0 second per image on a mainstream PC, which demonstrates the great potential of the proposed model in real-time applications. Mingyi Yuan, Xiaolin Hu 0001, Jianmin Li 0001, Huaping Liu 0001 |
IJCNN | 5 |
| 2013 | Unsupervised multimodal feature learning for semantic image segmentationabstractIn this paper, we address the semantic segmentation problem using single-layer networks. This network is used for unsupervised feature learning for the available RGB image and the depth image. A significant contribution of the proposed approach is that the dictionary is selected from the existing samples using the L2, 1optimization. Such a dictionary can capture more meaningful representative samples and exploit intrinsic correlation between features from different modalities. The experimental results on the public NYU dataset show that this strategy dramatically improves the classification performance, compared with existing dictionary learning approach. In addition, we perform experimental verification using the practical robot platforms and show promising results. Deli Pei, Huaping Liu 0001, Fuchun Sun 0001 |
IJCNN | 2 |
| 2013 | Traffic sign detection based on convolutional neural networksabstractWe propose an approach for traffic sign detection based on Convolutional Neural Networks (CNN). We first transform the original image into the gray scale image by using support vector machines, then use convolutional neural networks with fixed and learnable layers for detection and recognition. The fixed layer can reduce the amount of interest areas to detect, and crop the boundaries very close to the borders of traffic signs. The learnable layers can increase the accuracy of detection significantly. Besides, we use bootstrap methods to improve the accuracy and avoid overfitting problem. In the German Traffic Sign Detection Benchmark, we obtained competitive results, with an area under the precision-recall curve(AUC) of 99.73% in the category “Danger”, and an AUC of 97.62% in the category “Mandatory”. Yihui Wu, Jianmin Li 0001, Huaping Liu 0001, Xiaolin Hu 0001 |
IJCNN | 4 |
| 2013 | Local Feature Coding for Action Recognition Using RGB-D Camera
Mingyi Yuan, Huaping Liu 0001, Fuchun Sun 0001 |
ISNN (1) | 2 |
| 2013 | Special issue on prediction, control and diagnosis using advanced neural computations
Fuchun Sun 0001, Ying Tan 0002, Huaping Liu 0001 |
Inf. Sci. | 3 |
| 2013 | Supervised Low-Rank Matrix Recovery for Traffic Sign Recognition in Image SequencesabstractCorrelations in image sequences can be potentially useful for recovering feature representation and subsequently prompting classification performance, which are often neglected by traditional classification approaches. In this letter, we present a supervised low-rank matrix recovery model to leverage these correlations for classification tasks by introducing a supervised penalty term to the classic low-rank matrix recovery model. This allows us to not only exploit these correlations to recover the underlying feature representation from corrupted observation, but also preserve discriminative information for classification. Our model is evaluated on both real-world data and synthetic data, and experimental results show that our model obtains highly competitive performance with state-of-the-art algorithms and is especially robust to different levels of corruptions. Deli Pei, Fuchun Sun 0001, Huaping Liu 0001 |
IEEE Signal Process. Lett. | 3 |
| 2013 | Hierarchical Structured Sparse Representation for T-S Fuzzy Systems Identificationabstract“The curse of dimensionality” has become a significant bottleneck for fuzzy system identification and approximation. In this paper, we cast the Takagi–Sugeno (T–S) fuzzy system identification into a hierarchical sparse representation problem, where our goal is to establish a T–S fuzzy system with a minimal number of fuzzy rules, which simultaneously have a minimal number of nonzero consequent parameters. The proposed method, which is called hierarchical sparse fuzzy inference systems ( H-sparseFIS), explicitly takes into account the block-structured information that exists in the T–S fuzzy model and works in an intuitive way: First, initial fuzzy rule antecedent part is extracted automatically by an iterative vector quantization clustering method; then, with block-structured sparse representation, the main important fuzzy rules are selected, and the redundant ones are eliminated for better model accuracy and generalization performance; moreover, we simplify the selected fuzzy rules consequent with sparse regularization such that more consequent parameters can approximate to zero. This algorithm is very efficient and shows good performance in well-known benchmark datasets and real-world problems. Minnan Luo, Fuchun Sun 0001, Huaping Liu 0001 |
IEEE Trans. Fuzzy Syst. | 3 |
| 2012 | Estimating viewing angles in mobile street view searchabstractRecent years have witnessed an exciting progress in mobile visual search with applications to location recognition and streaming augmented reality. Most existing works among them are deployed with reference images coming from street views in urban scene. In such scenario, an interesting yet untouched problem is how to determine the viewing angle of the visual query aside of search, which could benefit multidisciplinary applications such as purifying the visual matching and accelerating the streaming AR. In this paper, we study the viewing angle estimation by exploiting the visual appearance of the query, which might be further improved by incorporating the coarse mobile context such as gyro or compass information. Our main idea is to treat this problem as a scene classification problem, upon which the key design is an optimal visual signature to reveal diversity of different viewing angles. We introduce a novel layout based viewing angle descriptor, which is based on carefully designed spatial division as well as appearance feature like color, texture and gradient. We have validated our approach on our dataset containing 1232 street view images in the urban areas of Manhattan, New York City. We show that our proposed descriptor has outperformed several alternatives in holistic image representations, including GIST, HOG and bag-of-feature with spatial pyramid matching. Deli Pei, Rongrong Ji, Fuchun Sun 0001, Huaping Liu 0001 |
ICIP | 4 |
| 2012 | Speed Limit Sign Recognition Using Log-Polar Mapping and Visual Codebook
Huaping Liu 0001, Xiong Luo, Fuchun Sun 0001 |
ISNN (2) | 2 |
| 2012 | Fusion tracking in color and infrared images using joint sparse representation
Huaping Liu 0001, Fuchun Sun 0001 |
Sci. China Inf. Sci. | 1 |
| 2012 | Efficient visual tracking using particle filter with incremental likelihood calculation
Huaping Liu 0001, Fuchun Sun 0001 |
Inf. Sci. | 1 |
| 2012 | A new algorithm for testing diagnosability of fuzzy discrete event systems
Minnan Luo, Yongming Li 0001, Fuchun Sun 0001, Huaping Liu 0001 |
Inf. Sci. | 4 |
| 2011 | Optimal necessary conditions for general SISO Mamdani fuzzy systems as function approximators within a given accuracyabstractIn this paper, necessary conditions are investigated for a single input/single output (SISO) Mamdani fuzzy systems as function approximators of continuous functions within a given accuracy. Since general SISO Mamdani fuzzy systems are monotonic on subintervals, the optimal configuration of fuzzy systems is that the number of division points is at least the times of its monotonicity changes. Thus with the extreme of the desired continuous function, necessary conditions are obtained through generating intervals that contain division points and pruning redundant intervals. Furthermore, a dynamically constructive method is proposed to show the conditions are optimal. It has been shown that existing results concerning necessary conditions are only special cases of our results. Finally, simulation examples are given to illustrate the conclusions, the strength of the fuzzy systems as function approximators are analyzed. Fuchun Sun 0001, Minnan Luo, Huaping Liu 0001 |
FUZZ-IEEE | 4 |
| 2011 | Visual Tracking Using Iterative Sparse Approximation
Huaping Liu 0001, Fuchun Sun 0001, Meng Gao 0002 |
ISNN (2) | 1 |
| 2011 | Mutation Hopfield neural network and its applications
Laihong Hu, Fuchun Sun 0001, Hualong Xu, Huaping Liu 0001 |
Inf. Sci. | 4 |
| 2010 | Visual Tracking Using Sparsity Induced SimilarityabstractCurrently sparse signal reconstruction gains considerable interests and is applied in many fields. In this paper, we propose a new approach which utilizes the sparsity induced similarity to construct the tracking algorithm. Compared with state-of-the-art, the advantage of this approach is that the sparse representation needs to be calculated for only once and therefore the time cost is dramatically decreased. In addition, extensive experimental comparisons show that the proposed approach is more robust than some existing approaches. Huaping Liu 0001, Fuchun Sun 0001 |
ICPR | 1 |
| 2010 | Performance improvement of force feedback in bilateral teleoperation with PD controllerabstractThis paper deals with the performance improvement of force feedback in bilateral teleoperation with PD controller. In traditional PD structures, the force feedback is simply determined by the position and velocity of the master and the slave manipulators, which may induce large resistance forces to the operator even in free motion. In this paper, a novel PD bilateral controller is proposed to tackle this problem. By incorporating a distance variable in the controller, we show that the appropriate force feedback can be obtained which still guarantees the system stability. To validate the proposed algorithm, an experiment is also carried out on our single degree of freedom teleoperation system. The results indicate that this strategy is effective for safe teleoperation missions. Yuji Wang, Fuchun Sun 0001, Huaping Liu 0001, Haibo Min |
IROS | 3 |
| 2010 | A Dual-Model Jumping Fuzzy System Approach to Networked Control Systems DesignabstractA discrete-time jump fuzzy system with two hidden Markov models (HMMs) is proposed to portray the asymmetric network characteristic of a class of nonlinear networked control systems (NCSs) with random but bounded communication delays and packets dropout. The less conservative state feedback controller and the dual-model-depend guaranteed cost controller are designed base on the model. A homotopy- based iterative algorithm solving for nonlinear matrix inequality (NMI) is developed to get the control gains. Simulation examples are carried out to show the effectiveness of the proposed approaches. Fengge Wu, Fuchun Sun 0001, Huaping Liu 0001 |
Int. J. Neural Syst. | 3 |
| 2009 | Semi-supervised ensemble trackingabstractIn this paper, we propose a semi-supervised ensemble tracking approach under the framework of particle filter. The particle filter is used not only for object searching, but also for unlabelled sample generation. By adopting the semi-supervised learning technology, these unlabelled samples which are generated online are utilized to progressively modify the classifier and make the ensemble tracker to be more robust to environment changing. On the other hand, utilizing semi-supervised learning technology can avoid the drifting phenomenons which are often encountered when using supervised learning. Finally, the performance of the proposed approach is evaluated using real visual tracking examples. Huaping Liu 0001, Fuchun Sun 0001 |
ICASSP | 1 |
| 2009 | Semi-supervised particle filter for visual trackingabstractIn this paper, a semi-supervised particle filter approach is proposed for visual tracking. The combination of semi-supervised learning and particle filter is very natural since the unlabelled samples are generated by particle propagation. In addition, the proposed semi-supervised particle filter can online select different features for robust tracking. To the best knowledge of the authors, this is the first time for the semi-supervised learning technology to be incorporated into the framework of particle filter. Finally, the performance of the proposed approach is evaluated using real visual tracking examples. Huaping Liu 0001, Fuchun Sun 0001 |
ICRA | 1 |
| 2009 | Vehicle tracking based on co-learning particle filterabstractIn this paper, we propose a co-learning particle filter approach for vehicle tracking, which is very important for intelligent vehicle. The proposal distribution of the particle filter is a combination of an extra support vector machine (SVM) detector and the motion prior. Previous works focusing on how to online update the detector or the observation likelihood using the tracking results. These approaches belong to ¿self-learning¿ fashion and easily tend to drift. The major difference between the proposed approach and previous works is that the SVM detector and the likelihood function can be mutually updated in a co-learning manner. By adopting the co-learning technology, the unlabelled samples which are generated during tracking are utilized to progressively modify the SVM detector and update the observation likelihood; therefore the resulting tracker is more robust and effectively avoids the drift problem. Finally, the performance of the proposed approach is evaluated using extensive real visual tracking examples. Weilong Ye, Huaping Liu 0001, Fuchun Sun 0001, Meng Gao 0002 |
IROS | 2 |
| 2009 | Mask Particle Filter for Similar Objects Tracking
Huaping Liu 0001, Fuchun Sun 0001, Meng Gao 0002 |
ISNN (3) | 1 |
| 2008 | On-orbit long-range maneuver transfer via EDAsabstractLong-range maneuver transfer consumes the most fuel and time of the rendezvous process, and it is a multi-variable, multi-extremum optimization problem, which is difficult to solve using traditional optimization algorithms. This paper researched the mathematical model of long-range maneuver transfer of spacecraft with impulse thrust, and optimized the parameters of orbit transfer based on a class of novel stochastic optimization algorithms, estimation of distribution algorithms (EDAs) with minimum fuel-time consumption being the optimization objective, and compared with Genetic Algorithms (GAs). Simulation results showed that EDAs were effective method for solving long-range maneuver transfer. Laihong Hu, Fuchun Sun 0001, Hualong Xu, Huaping Liu 0001, Fengge Wu |
IEEE Congress on Evolutionary Computation | 4 |
| 2008 | Fusion tracking in color and infrared images using sequential belief propagationabstractIn this paper, we propose an approach to fuse the color and infrared images for visual tracking. The contribution of this paper is twofold: First, we use the covariance feature to construct the likelihood function under the framework of particle filter. This likelihood captures the spatial and statistical properties as well as their correlation within representation of covariance. Secondly, different from the existing fusion approaches, our approach automatically realizes the fusion by sequential belief propagation, which uses message passing scheme to exchange information between color and infrared image. The performance of the proposed approach is evaluated using real visual tracking examples. Huaping Liu 0001, Fuchun Sun 0001 |
ICRA | 1 |
| 2008 | Particle Filter with Improved Proposal Distribution for Vehicle Tracking
Huaping Liu 0001, Fuchun Sun 0001 |
ISNN (1) | 1 |
| 2008 | An adaptive feature fusion framework for multi-class classification based on SVM
Peipei Yin, Fuchun Sun 0001, Huaping Liu 0001 |
Soft Comput. | 4 |
| 2008 | Fuzzy Particle Filtering for Uncertain SystemsabstractIn this paper, we propose a novel fuzzy particle filtering method for online estimation of nonlinear dynamic systems with fuzzy uncertainties. This approach uses a sequential fuzzy simulation to approximate the possibilities of the state intervals in the state–space, and estimates the state by fuzzy expected value operator. To solve the degeneracy problem of the fuzzy particle filter, one corresponding resampling technique is introduced. In addition, we compare the fuzzy particle filter with ordinary particle filter in both aspects of the theoretical basis and algorithm design, and demonstrate that the proposed filter outperforms standard particle filters especially when the number of the particles is small. The numerical simulations of two continuous-state nonlinear systems and a jump Markov system are employed to show the effectiveness and robustness of the proposed fuzzy particle filter. Hao Wu 0035, Fuchun Sun 0001, Huaping Liu 0001 |
IEEE Trans. Fuzzy Syst. | 3 |
| 2007 | Symmetry-Aided Particle Filter for Vehicle TrackingabstractSymmetry is an important characteristic of vehicle and has been used for detection tasks by many researchers. However, existing results of vehicle tracking seldom used this feature. In this paper, we combine the color histogram and the symmetry measurements to design a hierarchical-like particle filter for vehicle tracking. Experimental results show that the use of symmetry information will obtain better tracking performance than the conventional color histogram-based particle filters and effectively avoid some "hijack" problems. Huaping Liu 0001, Fuchun Sun 0001, Kezhong He |
ICRA | 1 |
| 2007 | Vehicle tracking using stochastic fusion-based particle filterabstractIn this article, we propose a new observation model combination approach under particle filtering scheme, which allows robust and accurate visual tracking under typ ical circumstances of real-time visual tracking. This scheme stochastically selects single observation model to evaluate the likelihood of some particle. Since only one single observation likelihood is evaluated for any one particle, the time-cost can be reduced dramatically. To verify its performance, this particle filter is used for vehicle tracking, by stochastically selecting color histogram or edge orientation histogram. The accuracy and robustness of the stochastic fusion approach are evaluated using real sequences. Furthermore, we demonstrate through these experiments that the stochastic fusion scheme performs almost as well as the deterministic fusion approach. Huaping Liu 0001, Fuchun Sun 0001, Kezhong He |
IROS | 1 |
| 2006 | Image Filtering Using Support Vector Machine
Huaping Liu 0001, Fuchun Sun 0001, Zengqi Sun |
ISNN (2) | 1 |
| 2005 | Hinfinity control for fuzzy singularly perturbed systems
Huaping Liu 0001, Fuchun Sun 0001, Yenan Hu |
Fuzzy Sets Syst. | 1 |
| 2005 | Stability analysis and synthesis of fuzzy singularly perturbed systemsabstractIn this paper, we investigate the stability analysis and synthesis problems for both continuous-time and discrete-time fuzzy singularly perturbed systems. For continuous-time case, both the stability analysis and synthesis can be parameterized in terms of a set of linear matrix inequalities (LMIs). For discrete-time case, only the analysis problem can be cast in LMIs, while the derived stability conditions for controller design are nonlinear matrix inequalities (NMIs). Furthermore, a two-stage algorithm based on LMI and iterative LMI (ILMI) techniques is developed to solve the resulting NMIs and the stabilizing feedback controller gains can be obtained. For both continuous-time and discrete-time cases, the reduced-control law, which is only dependent on the slow variables, is also discussed. Finally, an illustrated example based on the flexible joint inverted pendulum model is given to illustrate the design procedures. Huaping Liu 0001, Fuchun Sun 0001, Zengqi Sun |
IEEE Trans. Fuzzy Syst. | 1 |
| 2004 | Comments on "Constrained controller design of discrete Takagi-Sugeno fuzzy models"
Fuchun Sun 0001, Huaping Liu 0001, Zengqi Sun |
Fuzzy Sets Syst. | 2 |
| 2003 | Fuzzy control for nonlinear singularly perturbed systems with time-delayabstractIn this paper, the time-delay fuzzy singularly perturbed model is proposed based on the Takagi-Sugeno form, the basic /spl epsiv/-dependent stability conditions are given. Then the H/sub 2/ control problem is investigated in the form of /spl epsiv/-independent linear matrix inequalities. Finally, we show that the obtained controller solve also the corresponding problem for the time-delay fuzzy descriptor system. Huaping Liu 0001, Fuchun Sun 0001, Kezhong He, Zengqi Sun |
SMC | 1 |