EDBT 2026 Demo / reviewers in the wild / expert
Ran Song 0001
dblp:10/8738
· DBLP profile ↗
68ranked-venue papers
15as first author
48since 2021 · last 2026
0000-0002-1344-4415ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 29 · 7 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 25 · 12 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 18 · 16 since 2021Systems, architecture and hardware · 13 · 10 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Dexterous Manipulation Through Imitation Learning: A SurveyabstractDexterous manipulation, which refers to the ability of a robotic hand or multi-fingered end-effector to skillfully control, reorient, and manipulate objects through precise, coordinated finger movements and adaptive force modulation, enables complex interactions similar to human hand dexterity. With recent advances in robotics and machine learning, there is a growing demand for these systems to operate in complex and unstructured environments. Traditional model-based approaches struggle to generalize across tasks and object variations due to the high dimensionality and complex contact dynamics of dexterous manipulation. Although model-free methods such as reinforcement learning (RL) show promise, they require extensive training, large-scale interaction data, and carefully designed rewards for stability and effectiveness. Imitation learning (IL) offers an alternative by allowing robots to acquire dexterous manipulation skills directly from expert demonstrations, capturing fine-grained coordination and contact dynamics while bypassing the need for explicit modeling and large-scale trial-and-error. This survey provides an overview of dexterous manipulation methods based on imitation learning, details recent advances, and addresses key challenges in the field. Additionally, it explores potential research directions to enhance IL-driven dexterous manipulation. Our goal is to offer researchers and practitioners a comprehensive introduction to this rapidly evolving domain. Shan An, Chao Tang 0001, Yuning Zhou, Tengyu Liu, Fangqiang Ding, Shufang Zhang, Yao Mu 0001, Ran Song 0001, Wei Zhang 0021, Zeng-Guang Hou, Hong Zhang 0013 |
IEEE Trans Autom. Sci. Eng. | 9 |
| 2026 | MoTL: Modality-Balanced Terrain-Aware Locomotion Learning for Quadruped Robots
Zhiheng Li 0005, Yanyun Chen, Wenhao Tan, Mingxin Zhang 0006, Ran Song 0001, Wei Zhang 0021 |
IEEE Trans Autom. Sci. Eng. | 6 |
| 2026 | PAPNet: Point-Enhanced Attention-Aware Pillar Network for 3D Object Detection in Autonomous DrivingabstractThe conversion of raw point clouds into pillar representations has been widely adopted for 3D object detection. Such conversion allows a point cloud to be discretized into structured grids, which enables more efficient spatial representation and faster processing in real-time autonomous driving systems. However, discretizing raw point clouds often leads to the misdetection of small objects such as pedestrians and cyclists. This is because the discretization inevitably results in the loss of contextual and multi-resolution information within raw point clouds. To address this issue, we propose PAPNet, a point-enhanced attention-aware pillar network mainly composed of a point-pillar cross-attention module (PCM), a pillar-wise dual attention module (PDAM), and a multi-resolution set abstraction module (MSAM). PCM integrates raw point cloud features with pillar features across different dimensions, and PDAM guides PAPNet to focus on the intrinsic characteristics of the pillars. Additionally, MSAM retains both high-resolution and low-resolution features while integrating multi-scale information. Extensive experiments on four public datasets and in real-world scenarios demonstrate the effectiveness and efficiency of PAPNet. Codes, data, and demo videos can be found at the project website https://vsislab.github.io/PAPNet/. Ruitong Li, Yuenan Zhao, Jiaming Chen 0001, Ran Song 0001, Wei Zhang 0021 |
IEEE Trans Autom. Sci. Eng. | 5 |
| 2026 | CoVA-IL: Zero-Shot Imitation Learning via Contrastive Viewpoint Alignment on Object-Centric RepresentationabstractImitation learning provides an efficient paradigm for acquiring robotic manipulation skills, yet policies trained in a single environment often generalize poorly to unseen scenes. To address this challenge, we propose CoVA-IL, a zero-shot imitation learning framework that performs contrastive viewpoint alignment on object-centric representations, enabling direct policy deployment in novel environments without retraining. CoVA-IL uses the target object’s point cloud as the visual input and learns viewpoint-invariant latent representations through contrastive learning, thereby improving robustness to background changes, viewpoint variations, and cross-environment shifts. In addition, we incorporate multi-level point-cloud augmentation into 3D visuomotor imitation learning to improve data efficiency and reduce the number of demonstrations required for training. Real-world experiments show that CoVA-IL maintains average task success rates of 82.5% and 80% under substantial background and viewpoint changes, respectively, and multi-scene evaluations further validate its effectiveness for cross-environment deployment. Moreover, CoVA-IL can learn basic manipulation skills from a small number of real demonstrations, thereby reducing demonstration collection costs. Jiangtao Luo, Chenchen Zheng, Jinqiu Fan, Ran Song 0001, Wei Zhang 0021 |
IEEE Trans Autom. Sci. Eng. | 6 |
| 2026 | RDPrompter: Reference-Defect Prompt Learning for Few-Shot Defect Segmentation Based on Visual Foundation ModelabstractFew-shot industrial defect segmentation (FIDS) is an extremely challenging task in industrial inspection, which focuses on segmenting unseen defect categories with only a few samples. Existing few-shot segmentation methods are commonly developed with constrained feature extraction capabilities, making it difficult to accurately segment diverse unseen defects from distinctive industrial scenarios. To this end, we aim to leverage strong generalization strengths of the large-scale visual foundation model, segment anything model (SAM), to handle FIDS task. However, simply incorporating SAM into FIDS is insufficient for automatically categorizing defects, as the SAM heavily relies on manual prompts for segmentation. To address this, we propose a novel reference-defect prompt learning method RDPrompter, which learns adaptive prompts from defect similarities based on SAM for guiding automated FIDS. Specifically, to obtain adaptive prompts, we introduce a multiscale feature extraction strategy to mine rich defect features by the image encoder of SAM. Then, self-correcting probability prototype and cosine similarity attention are proposed to form a mixed similarity aggregation module, which obtains multiscale feature similarities for defect localization. Based on the similarities, a prompt embedding construction module is designed to further extract fine defect location information and generate prompt embeddings for FIDS. Sufficient experiments on MVTec-Unseen, SDD, FSSD-12 and the collected CID datasets demonstrate the superiority of our method. Tiyu Fang, Lin Zhang 0041, Ran Song 0001, Xiaolei Li 0003, Wei Zhang 0021 |
IEEE Trans. Ind. Informatics | 3 |
| 2026 | B2Q-Net: Bidirectional Branch Query Network for Surgical Phase RecognitionabstractSurgical phase recognition (SPR) is essential for surgical workflow analysis and provides immediate guidance during procedures. Existing methods aggregate frame-level information into a global representation and treat the task as frame-wise classification. However, this pipeline lacks a feedback mechanism for integrating historical information into local temporal modeling. To address this limitation, we propose the Bidirectional Branch Query Network (B2Q-Net), which reformulates the SPR task as the bidirectional query between phase-level features and frame-level features. B2Q-Net incorporates historical information during the initialization of phase queries. This enables bidirectional information flow during iterative refinement of two-level feature maps between phases and frames. Furthermore, we introduce a dual-scale selector (DSS) to generate high-quality phase queries for the current video clip. These phase queries retrieve historical information from the proposed state space query (SSQ) module, which uses learnable tokens as the historical state space to preserve historical information. Extensive evaluations on three datasets demonstrate that B2Q-Net consistently outperforms state-of-the-art methods in recognition accuracy while achieving an inference speed of 106 fps. The B2Q-Net code is available at https://github.com/vsislab/B2Q-Net. Zhiheng Li 0005, Yue Bi, Xiao Jia 0005, Ran Song 0001, Wei Zhang 0021 |
IEEE Trans. Medical Imaging | 5 |
| 2025 | GROVE: A Generalized Reward for Learning Open-Vocabulary Physical SkillabstractLearning open-vocabulary physical skills for simulated agents presents a significant challenge in Artificial Intelligence (AI). Current Reinforcement Learning (RL) approaches face critical limitations: manually designed rewards lack scalability across diverse tasks, while demonstration-based methods struggle to generalize beyond their training distribution. We introduce GROVE, a generalized reward framework that enables open-vocabulary physical skill learning without manual engineering or task-specific demonstrations. Our key insight is that Large Language Models (LLMs) and Vision Language Models (VLMs) provide complementary guidance—LLMs generate precise physical constraints capturing task requirements, while VLMs evaluate motion semantics and naturalness. Through an iterative design process, VLM-based feedback continuously refines LLM-generated constraints, creating a self-improving reward system. To bridge the domain gap between simulation and natural images, we develop Pose2CLIP, a lightweight mapper that efficiently projects agent poses directly into semantic feature space without computationally expensive rendering. Extensive experiments across diverse embodiments and learning paradigms demonstrate GROVE’s effectiveness, achieving 22.2% higher motion naturalness and 25.7% better task completion scores while training 8.4× faster than previous methods. These results establish a new foundation for scalable physical skill acquisition in simulated environments. Jieming Cui, Tengyu Liu, Jiale Yu, Ran Song 0001, Wei Zhang 0021, Yixin Zhu 0001, Siyuan Huang 0001 |
CVPR | 5 |
| 2025 | DPR-Splat: Depth and Pose Refinement with Sparse-View 3D Gaussian Splatting for Novel View SynthesisabstractRecent advances in 3D Gaussian Splatting have demonstrated impressive performance in novel view synthesis, particularly with dense image sets. However, its performance degrades significantly in sparse-view scenarios, primarily due to the challenge of obtaining accurate camera poses. Also, achieving scale-consistent and detailed depth maps is crucial, while existing depth estimation methods struggle to meet both requirements, further limiting view synthesis quality in sparse settings. To address these challenges, we propose DPR-Splat, an efficient neural reconstruction framework that builds 3D Gaussian models from sparse scenes. DPR-Splat refines the coarse outputs of MASt3R by leveraging dedicated pose and depth refinement modules, resulting in precise camera poses and depth maps. With the refined outputs, it progressively expands the 3D Gaussian set to construct an accurate scene model. Extensive experiments demonstrate that DPR-Splat enhances both novel view synthesis quality and pose estimation accuracy, and significantly accelerates training and rendering. Code and demonstration video are available at https://github.com/h0xg/DPR-Splat. Lingxiang Hu, Zhiheng Li 0005, Xingfei Zhu, Dun Li, Ran Song 0001 |
IROS | 5 |
| 2025 | ALVO: Adaptive Learning with Velocity Obstacles for UGV Navigation in Dynamic ScenesabstractAutonomous navigation of unmanned ground vehicles (UGVs) in dynamic scenes is a challenging task that requires them to avoid obstacles and move toward the goal simultaneously. This paper proposes ALVO, an adaptive learning policy that leverages velocity obstacles for UGV navigation. ALVO employs an adaptive gating-based mechanism for reactive obstacle avoidance, which enables the UGV to either slow down or proactively navigate around obstacles based on the relative importance of the environmental state and the goal. A reward function based on velocity obstacles is also designed to guide the UGV to navigate toward the goal while avoiding obstacles. Extensive experiments demonstrate that ALVO outperforms the competing approaches in various dynamic environments. We also implemented our method on a real UGV and showed that it performed well in real-world scenarios. Yinduo Xie, Yuenan Zhao, Ran Song 0001, Zhiheng Li 0005, Lei Han 0001, Wei Zhang 0021 |
IROS | 3 |
| 2025 | VCADNet: Vision-based Circular Accessible Depth Prediction for UGV PerceptionabstractCircular accessible depth (CAD) provides a lightweight and robust traversability representation for autonomous navigation of unmanned ground vehicles (UGV). Aiming at the limitations of existing LiDAR-based methods in detecting low-thickness targets and executing semantic reasoning, we propose VCADNet, a vision-based neural network for circular accessible depth prediction. VCADNet comprises three core components: a geometry-based query module for multi-view bird’s eye view feature extraction, a polar coordinate transformation for CAD alignment, and a multi-scale U-Net architecture for depth prediction. In addition, we present a cross-modal contrastive learning scheme to enhance the spatial reasoning of VCADNet, which transfers knowledge from LiDAR-based encoders to vision-based counterparts. Extensive experiments demonstrate the superior performance of VCADNet in various UGV perception tasks. Yuenan Zhao, Ran Song 0001, Lei Han 0001, Wei Zhang 0021 |
IROS | 4 |
| 2025 | Learning packing-and-unpacking synergistic policy via LLM-guided DRL for robust online robotic packing
Shuai Song, Ran Song 0001, Jiyu Cheng, Yibin Li 0001, Wei Zhang 0021 |
Adv. Eng. Informatics | 3 |
| 2025 | A Generalist Agent Learning Architecture for Versatile Quadruped LocomotionabstractQuadrupeds can generate various motor behaviors with the muscle synergies activated by the central nervous system. However, versatile locomotion for quadruped robots remains challenging due to the complexity of the high-dimensional limb dynamics with many physical constraints. Current approaches typically apply a dedicated policy or controller for each motor behavior, which requires the optimization of a large number of parameters and training process is complicated. In this paper, we propose a Generalist Agent Learning Architecture (GALA) to learn diverse motor behaviors simultaneously with a single policy network for quadruped locomotion. GALA significantly decreases the number of trainable parameters while producing appropriate motor behaviors by simply reactivating the generalist policy based on different sensory feedback and commands at run time. We experimentally analyze and demonstrate the versatile locomotion delivered by GALA on both simulated and real quadruped robots in various environments. The code is available at https://github.com/vsislab/GALA. Yanyun Chen, Ran Song 0001, Jiapeng Sheng, Wenhao Tan, Yibin Li 0001, Wei Zhang 0021 |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2025 | Fast 3D Room Layout Estimation Based on Compact High-Level Representationabstract3D room layout estimation aims to reconstruct the holistic 3D structure from an indoor RGB image. For most of the deep learning-based methods, layout inference is guided by a kind of learned 2D mid-level representation such as pixel-wise surface labels. However, learning such high-resolution 2D representation might suffer from information redundancy and memory consumption, and will increase the runtime of estimation and deployment cost for practical applications. In this paper, we attempt to learn a compact high-level representation with only 29 real numbers for estimating the 3D layout using general regression networks. The learned compact high-level representation contains three components: instance-wise plane parameters, camera intrinsic parameters, and plane location indicators. With the learned representation, the inverse depth map of each plane can be calculated to reconstruct the 3D layout. We further design a set of order-agnostic loss functions to restrict the produced inverse depth maps, with which the model can be trained with either weak 2D layout labels or full 3D layout supervision. Moreover, by jointly learning the plane parameters and locations, the model is benefited from 3D reasoning. Experimental results show that our method is much faster than the existing layout estimation methods and obtains competitive performance on benchmark datasets, showing its potential for real-time applications. Weidong Zhang 0005, Yu Qiao 0001, Ying Liu 0026, Ran Song 0001, Wei Zhang 0021 |
IEEE Trans. Image Process. | 4 |
| 2025 | SLPDR: A Benchmark for Ship License Plate Detection and RecognitionabstractShip identification is a prerequisite for the intelligent management of maritime transportation, yet existing research is confined to broad ship detection and categorization, which only provides the ship’s location or type instead of its identification. Inspired by the research on the Car License Plate (CLP), we make the first attempt to propose the concept of the Ship License Plate (SLP). In addition, the limited data hinders research on ship identification. To overcome this obstacle, we construct the first large-scale Ship License Plate Detection and Recognition (SLPDR) dataset, which contains 1,472 ship identities and 88,862 images. In addition, this paper proposes an SLP detection model named YOLO-SSA and evaluates this model as well as typical detection methods on the SLPDR dataset. The experimental results demonstrate that the proposed YOLO-SSA achieves better SLP detection performance by enhancing the features where ships and SLPs are located. Furthermore, we explore the prospective applications of SLPs in intelligent maritime transportation, including ship monitoring and berth management. Project web page: https://vsislab.github.io/SLPDR/ Youmei Zhang, Ran Song 0001, Yonghuai Liu, Ardhendu Behera, Mingxin Zhang 0006, Wei Zhang 0021 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2024 | LPS-Net: Lightweight Parameter-shared Network for Point Cloud-based Place RecognitionabstractWith innovation in fields such as autonomous driving and augmented reality, point cloud-based place recognition has gained significant attention. Many methods try to address this problem by extracting and matching global descriptors in a database, but they often must balance the extraction of comprehensive contextual information and large model sizes. To overcome this challenge, we propose a lightweight parameter-shared network (LPS-Net), which includes multiple bidirectional perception units (BPUs) to extract multiscale long-range contextual information and parameter-shared NetVLADs (PS-VLADs) to aggregate descriptors. A BPU includes a parameter-shared convolution module (SharedConv) that significantly compresses the model and enhances its ability to capture informative features. In PS-VLADs, we replace half the parameters used in the original NetVLAD with trainable scalars, which further reduces the model size, and theoretically prove their equivalence. Experimental results demonstrate that LPS-Net achieves state-of-the-art performance at the task of point cloud-based place recognition while maintaining a small model size. Code and supplementary materials can be found at https://github.com/Yavinr/LPS-Net. Guiyou Chen, Ran Song 0001 |
ICRA | 3 |
| 2024 | MPP: Multiscale Path Planning for UGV Navigation in Semi-structured EnvironmentsabstractAutonomous navigation of unmanned ground vehicles (UGVs) in structured road and indoor environments has made significant progress in recent years. However, navigation in outdoor semi-structured environments remains a challenge. This paper presents the multiscale path planning (MPP) method for UGV navigation in semi-structured environments. MPP leverages global, mid-layer and local planners to obtain global path and handle local obstacles of different sizes. First, the global planner provides guidance based on road connection relationships, selecting optimal connections by evaluating the distance between road nodes. Next, the mid-layer planner perceives large-scale obstacles and constructs the costmap, generating a mid-layer path that offers a general direction for the UGV. Finally, a local trajectory planning algorithm, namely terrain-considering timed elastic band (TC-TEB), is used to obtain local trajectory. This algorithm incorporates terrain-velocity constraints into the TEB algorithm to ensure the vehicle’s vertical stability. We demonstrate the safety and effectiveness of MPP through experiments in both simulated and real-world environments. Ran Song 0001, Wei Zhang 0021 |
IROS | 3 |
| 2024 | MPGNet: Learning Move-Push-Grasping Synergy for Target-Oriented Grasping in Occluded ScenesabstractThis paper focuses on target-oriented grasping in occluded scenes, where the target object is specified by a binary mask and the goal is to grasp the target object with as few robotic manipulations as possible. Most existing methods rely on a push-grasping synergy to complete this task. To deliver a more powerful target-oriented grasping pipeline, we present MPGNet, a three-branch network for learning a synergy between moving, pushing, and grasping actions. We also propose a multi-stage training strategy to train the MPGNet which contains three policy networks corresponding to the three actions. The effectiveness of our method is demonstrated via both simulated and real-world experiments. Video of the real-world experiments is at https://youtu.be/S_QKZqkh0w8. Dayou Li, Chenkun Zhao, Ran Song 0001, Xiaolei Li 0003, Wei Zhang 0021 |
IROS | 4 |
| 2024 | Coarse-to-Fine Detection of Multiple Seams for Robotic WeldingabstractEfficiently detecting target weld seams while ensuring sub-millimeter accuracy has always been an important challenge in autonomous welding, which has significant application in industrial practice. Previous works mostly focused on recognizing and localizing welding seams one by one, leading to inferior efficiency in modeling the workpiece. This paper proposes a novel framework capable of multiple weld seams extraction using both RGB images and 3D point clouds. The RGB image is used to obtain the region of interest by approximately localizing the weld seams, and the point cloud is used to achieve the fine-edge extraction of the weld seams within the region of interest using region growth. Our method is further accelerated by using a pre-trained deep learning model to ensure both efficiency and generalization ability. The proposed method was comprehensively tested on various workpieces featuring both linear and curved weld seams, as well as in physical experiment systems. The results showcase considerable potential for real-world industrial applications, emphasizing the method’s efficiency and effectiveness. Videos of the real-world experiments can be found at https://youtu.be/pq162HSP2D4. Pengkun Wei, Dayou Li, Ran Song 0001, Wei Zhang 0021 |
IROS | 4 |
| 2024 | Exploiting Inter-Sample Affinity for Knowability-Aware Universal Domain Adaptation
Yifan Wang 0020, Lin Zhang 0041, Ran Song 0001, Hongliang Li 0001, Paul L. Rosin, Wei Zhang 0021 |
Int. J. Comput. Vis. | 3 |
| 2024 | DeTAL: Open-Vocabulary Temporal Action Localization With Decoupled NetworksabstractPre-trained visual-language (ViL) models have demonstrated good zero-shot capability in video understanding tasks, where they were usually adapted through fine-tuning or temporal modeling. However, in the task of open-vocabulary temporal action localization (OV-TAL), such adaption reduces the robustness of ViL models against different data distributions, leading to a misalignment between visual representations and text descriptions of unseen action categories. As a result, existing methods often strike a trade-off between action detection and classification. Aiming at this issue, this paper proposes DeTAL, a simple but effective two-stage approach for OV-TAL. DeTAL decouples action detection from action classification to avoid the compromise between them, and the state-of-the-art methods for close-set action localization can be handily adapted to OV-TAL, which significantly improves the performance. Meanwhile, DeTAL can easily tackle the scenario where action category annotations are unavailable in the training dataset. In the experiments, we propose a new cross-dataset setting to evaluate the zero-shot capability of different methods. And the results demonstrate that DeTAL outperforms the state-of-the-art methods for OV-TAL on both THUMOS14 and ActivityNet1.3. Zhiheng Li 0005, Ran Song 0001, Lin Ma 0002, Wei Zhang 0021 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | IGCN: A Provably Informative GCN Embedding for Semi-Supervised Learning With Extremely Limited LabelsabstractGraph Neural Networks (GNNs) have gained much more attention in the representation learning for the graph-structured data. However, the labels are always limited in the graph, which easily leads to the overfitting problem and causes the poor performance. To solve this problem, we propose a new framework called IGCN, short for Informative Graph Convolutional Network, where the objective of IGCN is designed to obtain the informative embeddings via discarding the task-irrelevant information of the graph data based on the mutual information. As the mutual information for irregular data is intractable to compute, our framework is optimized via a surrogate objective, where two terms are derived to approximate the original objective. For the former term, it demonstrates that the mutual information between the learned embeddings and the ground truth should be high, where we utilize the semi-supervised classification loss and the prototype based supervised contrastive learning loss for optimizing it. For the latter term, it requires that the mutual information between the learned node embeddings and the initial embeddings should be high and we propose to minimize the reconstruction loss between them to achieve the goal of maximizing the latter term from the feature level and the layer level, which contains the graph encoder-decoder module and a novel architecture GCN$_{Info}$. Moreover, we provably show that the designed GCN$_{Info}$can better alleviate the information loss and preserve as much useful information of the initial embeddings as possible. Experimental results show that the IGCN outperforms the state-of-the-art methods on 7 popular datasets. Lin Zhang 0041, Ran Song 0001, Wenhao Tan, Lin Ma 0002, Wei Zhang 0021 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | HRNet: 3D object detection network for point cloud with hierarchical refinement
Ran Song 0001, Yonghuai Liu |
Pattern Recognit. | 4 |
| 2024 | A Hierarchical Framework for Quadruped Omnidirectional Locomotion Based on Reinforcement LearningabstractQuadruped locomotion is challenging for many learning-based algorithms. This is because it requires tedious manual tuning to cope with different types of terrains and is difficult to deploy in reality due to the sim-to-real gap between the training and the testing scenarios. This paper proposes a quadruped robot learning system for agile locomotion which does not require any pre-training and works well in various terrains. We introduce a hierarchical framework that uses reinforcement learning as the high-level policy to adjust the low-level trajectory generator for a better adaptability to various terrains. We compact the observation and the action spaces of reinforcement learning to deploy the proposed framework on a host computer interfaced with the robot. Besides, we design an omnidirectional trajectory generator guided by robot posture, which generates omnidirectional foot trajectories to interact with the environment. Experimental results and the supplementary video demonstrate that our hierarchical framework only trained in simulation can be easily deployed in the real world, and also has the advantages of fast convergence and good terrain adaptability.Note to Practitioners—This paper presents a hierarchical framework for quadruped robots. It combines a high-level reinforcement learning controller with a posture-guided trajectory generator to adaptively generate omnidirectional motions. Our method is easy to train as it converges fast and does not need to adjust a dozen or so of rewards. The quadruped robot can be deployed in a real environment directly after being trained in simulation. With the trained hierarchical framework deployed on a remote host computer, the robot works well in a variety of real-world environments unseen in the simulation. Wenhao Tan, Wei Zhang 0021, Ran Song 0001, Yu Zheng 0001, Yibin Li 0001 |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2024 | Heuristics Integrated Deep Reinforcement Learning for Online 3D Bin PackingabstractOnline 3D Bin Packing Problem (3D-BPP) has a wide range of industrial applications and there is an emerging research interest in learning optimal bin packing policy and deploying it for real logistics applications. From the heuristic methods to the deep reinforcement learning (DRL) methods, the previous works have proposed many solutions to solve the online 3D-BPP. However, none of them have studied what and how heuristics can be modelled into DRL to build a more effective and practical bin packing pipeline. In this work, we thoroughly investigate what heuristics can be used in online 3D-BPP and how to effectively integrate the heuristics with the DRL. First, we design 3 different heuristics based on the physical rules of the real world and the experiences of the human packers, including the Physics-Heuristics, the Packing-Heuristics and the Unpacking-Heuristics. Second, we model the 3 types of heuristics into the DRL framework and propose a novel heuristic DRL method to solve the online 3D-BPP. Extensive experimental results show that our method achieves state-of-the-art bin packing performance and the resulting real-world system is able to reliably finish the bin packing task in real logistics scenarios. Supplementary video is available athttps://www.youtube.com/watch?v=x8GpmEELq18. Note to Practitioners—The rapid growth of e-commerce has significantly increased the burden of human packers in logistic warehouses, where the workers need to pick the products from a conveyor and pack them into bins (i.e. the online 3D bin packing). Thus it is of great importance to develop intelligent robotic systems to replace human labor, which is a long-standing topic in the field of control and automation science. This paper makes a substantial contribution to the related field by studying the online 3D bin packing in terms of both the theory and practice. On the one hand, the simulated experiments suggest that the presented algorithm significantly improves the space utilization of bin packing. On the other hand, the robotic system developed based on the proposed method can favourably finish the bin packing task in real logistics scenarios, demonstrating the practical use of our approach. Consequently, the approach proposed in this paper is totally applicable in logistic warehouses and is promising to drastically improve the working efficiency of the product packing in real warehouses. In the future, we will extend the presented approach to pack irregular-shaped objects and then facilitate more logistics applications. Shuai Song, Shilei Chu, Ran Song 0001, Jiyu Cheng, Yibin Li 0001, Wei Zhang 0021 |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2024 | Learning Common Semantics via Optimal Transport for Contrastive Multi-View ClusteringabstractMulti-view clustering aims to learn discriminative representations from multi-view data. Although existing methods show impressive performance by leveraging contrastive learning to tackle the representation gap between every two views, they share the common limitation of not performing semantic alignment from a global perspective, resulting in the undermining of semantic patterns in multi-view data. This paper presents CSOT, namely Common Semantics via Optimal Transport, to boost contrastive multi-view clustering via semantic learning in a common space that integrates all views. Through optimal transport, the samples in multiple views are mapped to the joint clusters which represent the multi-view semantic patterns in the common space. With the semantic assignment derived from the optimal transport plan, we design a semantic learning module where the soft assignment vector works as a global supervision to enforce the model to learn consistent semantics among all views. Moreover, we propose a semantic-aware re-weighting strategy to treat samples differently according to their semantic significance, which improves the effectiveness of cross-view contrastive representation learning. Extensive experimental results demonstrate that CSOT achieves the state-of-the-art clustering performance. Qian Zhang 0076, Lin Zhang 0041, Ran Song 0001, Runmin Cong, Yonghuai Liu, Wei Zhang 0021 |
IEEE Trans. Image Process. | 3 |
| 2024 | Exploiting Label Uncertainty for Enhanced 3D Object Detection From Point CloudsabstractAccurate detection of objects from LiDAR point clouds is crucial for autonomous driving and environment modeling. However, uncertainties in ground truth labels due to occlusions, sparsity, and truncation can hinder model training and performance. This paper introduces two strategies to address these issues: 1) Soft Regression Loss (SoRL) and 2) Discrete Quantization Sampling (DQS). SoRL utilizes Gaussian distributions for object predictions, measuring uncertainty based on the probability of ground truth labels within these distributions. This method effectively accounts for deviations in object location and orientation. Meanwhile, DQS introduces uncertainty scores for dynamic sample selection, aiming to refine the quality of positive samples for regression. Based on the proposed modules, we design a lightweight multi-stage object detection framework. Notably, these modules can enhance existing 3D object detection methods without affecting significantly inference speeds. Experiments over benchmark datasets show the effectiveness of our method, especially for cars in sparse point clouds. Yonghuai Liu, Ardhendu Behera, Ran Song 0001, Hejin Yuan |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2024 | Ship Landmark: An Informative Ship Image Annotation and Its ApplicationsabstractVisual perception of ships has been attracting increasing attention in the fields of computer vision and ocean engineering. Despite the extensive work related to landmark detection of common objects, the role of landmarks in ship perception has been overlooked. In this paper, we aim to fill this gap by focusing on ship landmarks. Specifically, we give a comprehensive analysis of both the physical structure and deep features of ships, which finds that highlighted areas in feature maps correspond with structurally significant parts of ships. By summarizing the locations of such areas in ships, we define 20 ship landmarks and build the Ship Landmark Dataset (SLAD), the first ship dataset with landmark annotations. We also provide a benchmark for ship landmark detection by evaluating state-of-the-art landmark detection methods on the newly built SLAD. Moreover, we showcased several applications of ship landmarks, including ship recognition, ship image generation, key area detection for ships, and ship detection. Project web page:https://vsislab.github.io/Ships_VSIS/. Mingxin Zhang 0006, Qian Zhang 0076, Ran Song 0001, Paul L. Rosin, Wei Zhang 0021 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2024 | Neighborhood-Aware Mutual Information Maximization for Source-Free Domain AdaptationabstractRecently, the source-free domain adaptation (SFDA) problem has attracted much attention, where the pre-trained model for the source domain is adapted to the target domain in the absence of source data. However, due to domain shift, the negative alignment usually exists between samples from the same class, which may lower intra-class feature similarity. To address this issue, we present a self-supervised representation learning strategy for SFDA, named as neighborhood-aware mutual information (NAMI), which maximizes the mutual information (MI) between the representations of target samples and their corresponding neighbors. Moreover, we theoretically demonstrate that NAMI can be decomposed into a weighted sum of local MI, which suggests that the weighted terms can better estimate NAMI. To this end, we introduce neighborhood consensus score over the set of weakly and strongly augmented views and point-wise density based on neighborhood, both of which determine the weights of local MI for NAMI by leveraging the neighborhood information of samples. The proposed method can significantly handle domain shift and adaptively reduce the noise in the neighborhood of each target sample. In combination with the consistency loss over views, NAMI leads to consistent improvement over existing state-of-the-art methods on three popular SFDA benchmarks. Lin Zhang 0041, Yifan Wang 0020, Ran Song 0001, Mingxin Zhang 0006, Xiaolei Li 0003, Wei Zhang 0021 |
IEEE Trans. Multim. | 3 |
| 2023 | Point-aware Interaction and CNN-induced Refinement Network for RGB-D Salient Object DetectionabstractBy integrating complementary information from RGB image and depth map, the ability of salient object detection (SOD) for complex and challenging scenes can be improved. In recent years, the important role of Convolutional Neural Networks (CNNs) in feature extraction and cross-modality interaction has been fully explored, but it is still insufficient in modeling global long-range dependencies of self-modality and cross-modality. To this end, we introduce CNNs-assisted Transformer architecture and propose a novel RGB-D SOD network with Point-aware Interaction and CNN-induced Refinement (PICR-Net). On the one hand, considering the prior correlation between RGB modality and depth modality, an attention-triggered cross-modality point-aware interaction (CmPI) module is designed to explore the feature interaction of different modalities with positional constraints. On the other hand, in order to alleviate the block effect and detail destruction problems brought by the Transformer naturally, we design a CNN-induced refinement (CNNR) unit for content refinement and supplementation. Extensive experiments on five RGB-D SOD datasets show that the proposed network achieves competitive results in both quantitative and qualitative comparisons. Our code is publicly available at: https://github.com/rmcong/PICR-Net_ACMMM23. Runmin Cong, Hongyu Liu 0003, Chen Zhang 0013, Wei Zhang 0021, Feng Zheng 0001, Ran Song 0001, Sam Kwong |
ACM Multimedia | 6 |
| 2023 | 3D Visual Saliency: An Independent Perceptual Measure or a Derivative of 2D Image Saliency?abstractWhile 3D visual saliency aims to predict regional importance of 3D surfaces in agreement with human visual perception and has been well researched in computer vision and graphics, latest work with eye-tracking experiments shows that state-of-the-art 3D visual saliency methods remain poor at predicting human fixations. Cues emerging prominently from these experiments suggest that 3D visual saliency might associate with 2D image saliency. This paper proposes a framework that combines a Generative Adversarial Network and a Conditional Random Field for learning visual saliency of both a single 3D object and a scene composed of multiple 3D objects with image saliency ground truth to 1) investigate whether 3D visual saliency is an independent perceptual measure or just a derivative of image saliency and 2) provide a weakly supervised method for more accurately predicting 3D visual saliency. Through extensive experiments, we not only demonstrate that our method significantly outperforms the state-of-the-art approaches, but also manage to answer the interesting and worthy question proposed within the title of this paper. Ran Song 0001, Wei Zhang 0021, Yitian Zhao, Yonghuai Liu, Paul L. Rosin |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | SMAM: Self and Mutual Adaptive Matching for Skeleton-Based Few-Shot Action RecognitionabstractThis paper focuses on skeleton-based few-shot action recognition. Since skeleton is essentially a sparse representation of human action, the feature maps extracted from it, through a standard encoder network in the few-shot condition, may not be sufficiently discriminative for some action sequences that look partially similar to each other. To address this issue, we propose a self and mutual adaptive matching (SMAM) module to convert such feature maps into more discriminative feature vectors. Our method, named as SMAM-Net, first leverages both the temporal information associated with each individual skeleton joint and the spatial relationship among them for feature extraction. Then, the SMAM module adaptively measures the similarity between labeled and query samples and further carries out feature matching within the query set to distinguish similar skeletons of various action categories. Experimental results show that the SMAM-Net outperforms other baselines on the large-scale NTU RGB + D 120 dataset in the tasks of one-shot and five-shot action recognition. We also report our results on smaller datasets including NTU RGB + D 60, SYSU and PKU-MMD to demonstrate that our method is reliable and generalises well on different datasets. Codes and the pretrained SMAM-Net will be made publicly available. Zhiheng Li 0005, Xuyuan Gong, Ran Song 0001, Peng Duan 0002, Jun Liu 0036, Wei Zhang 0021 |
IEEE Trans. Image Process. | 3 |
| 2023 | Unsupervised Maritime Vessel Re-Identification With Multi-Level Contrastive LearningabstractRe-identification (re-ID) of maritime vessels plays an important role in marine surveillance, but remains highly unexplored due to the lack of large-scale annotated datasets. In vessel re-ID, contrastive methods are supposed to learn discriminative representation from unlabeled vessel images in an unsupervised manner. However, directly introducing classical instance-level contrastive methods to maritime vessel re-ID suffers from the difficulty of finding vessel images with the same pseudo label as positive images, which potentially leads to inefficient training and unsatisfactory performance. This paper proposes a simple but effective method to solve such a hard positive problem. Our method takes all images in an intra-batch cluster as positives and excludes them from the set of negative samples when computing instance-level contrastive loss. Based on this strategy, we construct a multi-level contrastive learning (MCL) framework for vessel re-ID trained with the specifically designed intra-batch cluster-level contrastive loss along with the instance-level one. Experiments on a newly proposed dataset consisting of 1,248 vessel identities show that MCL achieves the state-of-the-art performance compared with other unsupervised methods. Qian Zhang 0076, Mingxin Zhang 0006, Jinghe Liu, Xuanyu He, Ran Song 0001, Wei Zhang 0021 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2023 | Classification of Brain Disorders in rs-fMRI via Local-to-Global Graph Neural NetworksabstractRecently, functional brain network has been used for the classification of brain disorders, such as Autism Spectrum Disorder (ASD) and Alzheimer's disease (AD). Existing methods either ignore the non-imaging information associated with the subjects and the relationship between the subjects, or cannot identify and analyze disease-related local brain regions and biomarkers, leading to inaccurate classification results. This paper proposes a local-to-global graph neural network (LG-GNN) to address this issue. A local ROI-GNN is designed to learn feature embeddings of local brain regions and identify biomarkers, and a global Subject-GNN is then established to learn the relationship between the subjects with the embeddings generated by the local ROI-GNN and the non-imaging information. The local ROI-GNN contains a self-attention based pooling module to preserve the embeddings most important for the classification. The global Subject-GNN contains an adaptive weight aggregation block to generate the multi-scale feature embedding corresponding to each subject. The proposed LG-GNN is thoroughly validated using two public datasets for ASD and AD classification. The experimental results demonstrated that it achieves the state-of-the-art performance in terms of various evaluation metrics. Hao Zhang 0113, Ran Song 0001, Lin Zhang 0041, Dawei Wang 0015, Cong Wang 0007, Wei Zhang 0021 |
IEEE Trans. Medical Imaging | 2 |
| 2023 | Unsupervised Embedding Learning With Mutual-Information Graph Convolutional NetworksabstractRecently, methods for unsupervised embedding learning have exhibited promising results for extracting desirable representations from unlabeled samples. In general, most methods learn the feature embeddings by handling each sample individually while the structural and semantic relationships between samples are not fully exploited. As a result, the learned embeddings are not sufficiently discriminative. To make use of such inter-sample information for deep embedding learning, this paper proposes an unsupervised method based on the graph convolutional network (GCN). On one hand, our method encodes structural information between the samples corresponding to the nodes in a local neighbourhood of the GCN graph. On the other hand, it leverages the mutual information between the original samples and the augmented ones to ensure that they are globally consistent with each other. Extensive experiments show that our method is not just robust to augmentation perturbations, but also learns discriminative embeddings. Consequently, it achieves the state-of-the-art performance on several challenging datasets. Lin Zhang 0041, Mingxin Zhang 0006, Ran Song 0001, Ziying Zhao, Xiaolei Li 0003 |
IEEE Trans. Multim. | 3 |
| 2023 | Circular Accessible Depth: A Robust Traversability Representation for UGV NavigationabstractIn this article, we present the circular accessible depth (CAD), a robust traversability representation for an unmanned ground vehicle (UGV) to learn traversability in various scenarios containing irregular obstacles. To predict CAD, we propose a neural network, namely CADNet, with an attention-based multiframe point cloud fusion module, stability-attention module (SAM), to encode the spatial features from point clouds captured by LiDAR. CAD is designed based on the polar coordinate system and focuses on predicting the border of traversable area. Since it encodes the spatial information of the surrounding environment, which enables a semisupervised learning for the CADNet, and thus, desirably avoids annotating a large amount of data. Extensive experiments demonstrate that CAD outperforms baselines in terms of robustness and precision. We also implement our method on a real UGV and show that it performs well in real-world scenarios. Shikuan Xie, Ran Song 0001, Yuenan Zhao, Xueqin Huang, Yibin Li 0001, Wei Zhang 0021 |
IEEE Trans. Robotics | 2 |
| 2023 | Watch and Act: Learning Robotic Manipulation From Visual DemonstrationabstractLearning from demonstration holds the promise of enabling robots to learn diverse actions from expert experience. In contrast to learning from observation-action pairs, humans learn to imitate in a more flexible and efficient manner: learning behaviors by simply “watching.” In this article, we propose a “watch-and-act” imitation learning pipeline that endows a robot with the ability of learning diverse manipulations from visual demonstrations. Specifically, we address this problem by intuitively casting it as two subtasks: 1) understanding the demonstration video and 2) learning the demonstrated manipulations. First, a captioning module based on visual change is presented to understand the demonstration by translating the demonstration video into a command sentence. Then, to execute the captioning command, a manipulation module that learns the demonstrated manipulations is built upon an instance segmentation model and a manipulation affordance prediction model. We validate the superiority of the two modules over existing methods separately via extensive experiments and demonstrate the whole robotic imitation system developed based on the two modules in diverse scenarios using a real robotic arm. Supplementary video is available athttps://vsislab.github.io/watch-and-act/. Wei Zhang 0021, Ran Song 0001, Jiyu Cheng, Hesheng Wang 0001, Yibin Li 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2022 | Unsupervised Multi-View CNN for Salient View Selection and 3D Interest Point Detection
Ran Song 0001, Wei Zhang 0021, Yitian Zhao, Yonghuai Liu |
Int. J. Comput. Vis. | 1 |
| 2022 | 3D vessel-like structure segmentation in medical images by an edge-reinforced network
Likun Xia, Hao Zhang 0113, Yufei Wu 0013, Ran Song 0001, Yuhui Ma, Lei Mou, Jiang Liu 0001, Ming Ma 0004, Yitian Zhao |
Medical Image Anal. | 4 |
| 2022 | LiTMNet: A deep CNN for efficient HDR image reconstruction from a single LDR image
Guotao Wu, Ran Song 0001, Mingxin Zhang 0006, Xiaolei Li 0003, Paul L. Rosin |
Pattern Recognit. | 2 |
| 2022 | 3D Layout Estimation via Weakly Supervised Learning of Plane Parameters From 2D SegmentationabstractThe task of 3D layout estimation in an indoor scene is to predict the holistic 3D structural information of the scene from an RGB image. It is costly to obtain the ground truth 3D layout, and this issue severely restricts the learning based 3D layout estimation approaches. In this paper, we present a novel weakly supervised learning framework that is able to learn the 3D layout effectively with 2D layout segmentation mask as supervision. We employ a deep neural network to predict the plane parameters and camera intrinsic parameters in the image. Based on the predicted plane instances, the 3D layout as well as the corresponding depth map and 2D segmentation can be generated. The key objectives for learning meaningful plane parameters are the label consistency of layout segmentation and depth consistency of border pixels from adjacent planes, with which the ground truth 2D layout segmentation is able to supervise the learning of the 3D layout. We further incorporate 3D geometric reasoning and prior knowledge in the learning process to ensure that the learned 3D layout is realistic and reasonable. Experimental results show that our method can produce accurate 3D layout estimates by weakly supervised learning. Weidong Zhang 0005, Youmei Zhang, Ran Song 0001, Ying Liu 0026, Wei Zhang 0021 |
IEEE Trans. Image Process. | 3 |
| 2022 | JoT-GAN: A Framework for Jointly Training GAN and Person Re-Identification ModelabstractTo cope with the problem caused by inadequate training data, many person re-identification (re-id) methods exploit generative adversarial networks (GAN) for data augmentation, where the training of GAN is typically independent of that of the re-id model. The coupling relation between them that probably brings in a performance gain of re-id is thus ignored. In this work, we propose a general framework, namely JoT-GAN, to jointly train GAN and the re-id model. It can simultaneously achieve the optima of both the generator and the re-id model, where the training is guided by each other through a discriminator. The re-id model is boosted for two reasons: (1) the adversarial training encourages it to fool the discriminator, and (2) the generated samples augment the training data. Extensive results on benchmark datasets show that for the re-id model trained with the identification loss as well as the triplet loss, the proposed joint training framework outperforms existing methods with separate training and achieves state-of-the-art re-id performance. Ran Song 0001, Qian Zhang 0076, Peng Duan 0002, Youmei Zhang |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2021 | Mesh Saliency: An Independent Perceptual Measure or a Derivative of Image Saliency?abstractWhile mesh saliency aims to predict regional importance of 3D surfaces in agreement with human visual perception and is well researched in computer vision and graphics, latest work with eye-tracking experiments shows that state-of-the-art mesh saliency methods remain poor at predicting human fixations. Cues emerging prominently from these experiments suggest that mesh saliency might associate with the saliency of 2D natural images. This paper proposes a novel deep neural network for learning mesh saliency using image saliency ground truth to 1) investigate whether mesh saliency is an independent perceptual measure or just a derivative of image saliency and 2) provide a weakly supervised method for more accurately predicting mesh saliency. Through extensive experiments, we not only demonstrate that our method outperforms the current state-of-the-art mesh saliency method by 116% and 21% in terms of linear correlation coefficient and AUC respectively, but also reveal that mesh saliency is intrinsically related with both image saliency and object categorical information. Codes are available at https://github.com/rsong/MIMO-GAN. Ran Song 0001, Wei Zhang 0021, Yitian Zhao, Yonghuai Liu, Paul L. Rosin |
CVPR | 1 |
| 2021 | Autonomous Multi-View Navigation via Deep Reinforcement LearningabstractIn this paper, we propose a novel deep reinforcement learning (DRL) system for the autonomous navigation of mobile robots that consists of three modules: map navigation, multi-view perception and multi-branch control. Our DRL system takes as the input a routed map provided by a global planner and three RGB images captured by a multi-camera setup to gather global and local information, respectively. In particular, we present a multi-view perception module based on an attention mechanism to filter out redundant information caused by multi-camera sensing. We also replace raw RGB images with low-dimensional representations via a specifically designed network, which benefits a more robust sim2real transfer learning. Extensive experiments in both simulated and real-world scenarios demonstrate that our system outperforms state-of-the-art approaches. Xueqin Huang, Wei Zhang 0021, Ran Song 0001, Jiyu Cheng, Yibin Li 0001 |
ICRA | 4 |
| 2021 | A Hierarchical Framework for Quadruped Locomotion Based on Reinforcement LearningabstractQuadruped locomotion is a challenging task for learning-based algorithms. It requires tedious manual tuning and is difficult to deploy in reality due to the reality gap. In this paper, we propose a quadruped robot learning system for agile locomotion which does not require any pre-training and works well in various real-world terrains. We introduce a hierarchical learning framework that uses reinforcement learning as the high-level policy to adjust the low-level trajectory generator for better adaptability to the terrain. We compact the observation and action space of the reinforcement learning to deploy it on a host computer in reality. Besides, we design a trajectory generator guided by robot posture, which can generate adaptive foot trajectory to interact with the environment. Experimental results show that our system can be easily deployed in reality while only trained in simulation, and also has the advantages of fast convergence and good terrain adaptability. The supplementary video demonstration is available at https://vsislab.github.io/hfql/. Wenhao Tan, Wei Zhang 0021, Ran Song 0001, Yu Zheng 0001, Yibin Li 0001 |
IROS | 4 |
| 2021 | PackerBot: Variable-Sized Product Packing with Heuristic Deep Reinforcement LearningabstractProduct packing is a typical application in ware-house automation that aims to pick objects from unstructured piles and place them into bins with optimized placing policy. However, it still remains a significant challenge to finish the product packing tasks in general logistics scenarios where the objects are variable-sized and the configurations are complex. In this work, we present the PackerBot, a complete robotic pipeline for performing variable-sized product packing in unstructured scenes. First, by leveraging the imperfect experience of human packer, we propose a heuristic DRL framework for learning optimal online 3D bin packing policy. Then we integrate it with a 6-DoF suction-based picking module and a product size estimation module, leading to a complete product packing system, namely the PackerBot. Extensive experimental results show that our method achieves the state-of-the-art performance in both simulated and real-world tests. The video demonstration is available at: https://vsislab.github.io/packerbot. Zifei Yang, Shuai Song, Wei Zhang 0021, Ran Song 0001, Jiyu Cheng, Yibin Li 0001 |
IROS | 5 |
| 2021 | EFNet: Enhancement-Fusion Network for Semantic Segmentation
Zhijie Wang 0010, Ran Song 0001, Peng Duan 0002, Xiaolei Li 0003 |
Pattern Recognit. | 2 |
| 2021 | PoT-GAN: Pose Transform GAN for Person Image SynthesisabstractPose-based person image synthesis aims to generate a new image containing a person with a target pose conditioned on a source image containing a person with a specified pose. It is challenging as the target pose is arbitrary and often significantly differs from the specified source pose, which leads to large appearance discrepancy between the source and the target images. This paper presents the Pose Transform Generative Adversarial Network (PoT-GAN) for person image synthesis where the generator explicitly learns the transform between the two poses by manipulating the corresponding multi-scale feature maps. By incorporating the learned pose transform information into the multi-scale feature maps of the source image in a GAN architecture, our method reliably transfers the appearance of the person in the source image to the target pose with no need for any hard-coded spatial information depicting the change of pose. According to both qualitative and quantitative results, the proposed PoT-GAN demonstrates a state-of-the-art performance on three publicly available datasets for person image synthesis. Wei Zhang 0021, Ran Song 0001, Zhiheng Li 0005, Jun Liu 0036, Xiaolei Li 0003, Shijian Lu |
IEEE Trans. Image Process. | 3 |
| 2021 | Mesh Saliency via Weakly Supervised Classification-for-Saliency CNNabstractRecently, effort has been made to apply deep learning to the detection of mesh saliency. However, one major barrier is to collect a large amount of vertex-level annotation as saliency ground truth for training the neural networks. Quite a few pilot studies showed that this task is difficult. In this work, we solve this problem by developing a novel network trained in a weakly supervised manner. The training is end-to-end and does not require any saliency ground truth but only the class membership of meshes. Our Classification-for-Saliency CNN (CfS-CNN) employs a multi-view setup and contains a newly designed two-channel structure which integrates view-based features of both classification and saliency. It essentially transfers knowledge from 3D object classification to mesh saliency. Our approach significantly outperforms the existing state-of-the-art methods according to extensive experimental results. Also, the CfS-CNN can be directly used for scene saliency. We showcase two novel applications based on scene saliency to demonstrate its utility. Ran Song 0001, Yonghuai Liu, Paul L. Rosin |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2020 | Unsupervised Multi-view CNN for Salient View Selection of 3D Objects and Scenes
Ran Song 0001, Wei Zhang 0021, Yitian Zhao, Yonghuai Liu |
ECCV (19) | 1 |
| 2020 | Grasp for Stacking via Deep Reinforcement LearningabstractIntegrated robotic arm system should contain both grasp and place actions. However, most grasping methods focus more on how to grasp objects, while ignoring the placement of the grasped objects, which limits their applications in various industrial environments. In this research, we propose a model-free deep Q-learning method to learn the grasping-stacking strategy end-to-end from scratch. Our method maps the images to the actions of the robotic arm through two deep networks: the grasping network (GNet) using the observation of the desk and the pile to infer the gripper's position and orientation for grasping, and the stacking network (SNet) using the observation of the platform to infer the optimal location when placing the grasped object. To make a long-range planning, the two observations are integrated in the grasping for stacking network (GSN). We evaluate the proposed GSN on a grasping-stacking task in both simulated and real-world scenarios. Wei Zhang 0021, Ran Song 0001, Lin Ma 0002, Yibin Li 0001 |
ICRA | 3 |
| 2020 | Learn by Observation: Imitation Learning for Drone Patrolling from Videos of A Human NavigatorabstractWe present an imitation learning method for autonomous drone patrolling based only on raw videos. Different from previous methods, we propose to let the drone learn patrolling in the air by observing and imitating how a human navigator does it on the ground. The observation process enables the automatic collection and annotation of data using inter-frame geometric consistency, resulting in less manual effort and high accuracy. Then a newly designed neural network is trained based on the annotated data to predict appropriate directions and translations for the drone to patrol in a lane-keeping manner as humans. Our method allows the drone to fly at a high altitude with a broad view and low risk. It can also detect all accessible directions at crossroads and further carry out the integration of available user instructions and autonomous patrolling control commands. Extensive experiments are conducted to demonstrate the accuracy of the proposed imitating learning process as well as the reliability of the holistic system for autonomous drone navigation. The codes, datasets as well as video demonstrations are available at https://vsislab.github.io/uavpatrol. Shilei Chu, Wei Zhang 0021, Ran Song 0001, Yibin Li 0001 |
IROS | 4 |
| 2020 | Autonomous Robot Navigation Based on Multi-Camera PerceptionabstractIn this paper, we propose an autonomous method for robot navigation based on a multi-camera setup that takes advantage of a wide field of view. A new multi-task network is designed for handling the visual information supplied by the left, central and right cameras to find the passable area, detect the intersection and infer the steering. Based on the outputs of the network, three navigation indicators are generated and then combined with the high-level control commands extracted by the proposed MapNet, which are finally fed into the driving controller. The indicators are also used through the controller for adjusting the driving velocity, which assists the robot to adjust the speed for smoothly bypassing obstacles. Experiments in real-world environments demonstrate that our method performs well in both local obstacle avoidance and global goal-directed navigation tasks. Kunyan Zhu, Wei Zhang 0021, Ran Song 0001, Yibin Li 0001 |
IROS | 4 |
| 2020 | Cerebrovascular Segmentation in MRA via Reverse Edge Attention Network
Hao Zhang 0113, Likun Xia, Ran Song 0001, Jianlong Yang, Huaying Hao, Jiang Liu 0001, Yitian Zhao |
MICCAI (6) | 3 |
| 2020 | Online Decision Based Visual Tracking via Reinforcement LearningabstractA deep visual tracker is typically based on either object detection or template matching while each of them is only suitable for a particular group of scenes. It is straightforward to consider fusing them together to pursue more reliable tracking. However, this is not wise as they follow different tracking principles. Unlike previous fusion-based methods, we propose a novel ensemble framework, named DTNet, with an online decision mechanism for visual tracking based on hierarchical reinforcement learning. The decision mechanism substantiates an intelligent switching strategy where the detection and the template trackers have to compete with each other to conduct tracking within different scenes that they are adept in. Besides, we present a novel detection tracker which avoids the common issue of incorrect proposal. Extensive results show that our DTNet achieves state-of-the-art tracking performance as well as good balance between accuracy and efficiency. The project website is available at https://vsislab.github.io/DTNet/. Ke Song 0003, Wei Zhang 0021, Ran Song 0001, Yibin Li 0001 |
NeurIPS | 3 |
| 2020 | Histogram of Fuzzy Local Spatio-Temporal Descriptors for Video Action RecognitionabstractFeature extraction plays a vital role in visual action recognition. Many existing gradient-based feature extractors, including histogram of oriented gradients, histogram of optical flow, motion boundary histograms, and histogram of motion gradients, build histograms for representing different actions over the spatio-temporal domain in a video. However, these methods require to set the number of bins for information aggregation in advance. Varying numbers of bins usually lead to inherent uncertainty within the process of pixel voting with regard to the bins in the histogram. This article proposes a novel method to handle such uncertainty by fuzzifying these feature extractors. The proposed approach has two advantages: it better represents the ambiguous boundaries between the bins and, thus, the fuzziness of the spatio-temporal visual information entailed in videos; and the contribution of each pixel is flexibly controlled by a fuzziness parameter for various scenarios. The proposed family of fuzzy descriptors and a combination of them are evaluated on two publicly available datasets, demonstrating that the proposed approach outperforms the original counterparts and other state-of-the-art methods. Zheming Zuo, Longzhi Yang, Yonghuai Liu, Fei Chao 0001, Ran Song 0001, Yanpeng Qu |
IEEE Trans. Ind. Informatics | 5 |
| 2020 | Distinction of 3D Objects and Scenes via Classification Network and Markov Random FieldabstractAn importance measure of 3D objects inspired by human perception has a range of applications since people want computers to behave like humans in many tasks. This paper revisits a well-defined measure, distinction of 3D surface mesh, which indicates how important a region of a mesh is with respect to classification. We develop a method to compute it based on a classification network and a Markov Random Field (MRF). The classification network learns view-based distinction by handling multiple views of a 3D object. Using a classification network has an advantage of avoiding the training data problem which has become a major obstacle of applying deep learning to 3D object understanding tasks. The MRF estimates the parameters of a linear model for combining the view-based distinction maps. The experiments using several publicly accessible datasets show that the distinctive regions detected by our method are not just significantly different from those detected by methods based on handcrafted features, but more consistent with human perception. We also compare it with other perceptual measures and quantitatively evaluate its performance in the context of two applications. Furthermore, due to the view-based nature of our method, we are able to easily extend mesh distinction to 3D scenes containing multiple objects. Ran Song 0001, Yonghuai Liu, Paul L. Rosin |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2019 | Multiscale Representation of 3D Surfaces via Stochastic Mesh Laplacian
Ran Song 0001 |
Comput. Aided Des. | 1 |
| 2018 | Local-to-global mesh saliency
Ran Song 0001, Yonghuai Liu, Ralph R. Martin, Karina Rodriguez-Echavarria |
Vis. Comput. | 1 |
| 2016 | Accurately estimating rigid transformations in registration using a boosting-inspired mechanism
Yonghuai Liu, Honghai Liu 0001, Ralph R. Martin, Luigi De Dominicis, Ran Song 0001, Yitian Zhao |
Pattern Recognit. | 5 |
| 2014 | Scan integration as a labelling problem
Ran Song 0001, Yonghuai Liu, Ralph R. Martin, Paul L. Rosin |
Pattern Recognit. | 1 |
| 2014 | Mesh saliency via spectral processingabstractWe propose a novel method for detecting mesh saliency, a perceptually-based measure of the importance of a local region on a 3D surface mesh. Our method incorporates global considerations by making use of spectral attributes of the mesh, unlike most existing methods which are typically based on local geometric cues. We first consider the properties of the log-Laplacian spectrum of the mesh. Those frequencies which show differences from expected behaviour capture saliency in the frequency domain. Information about these frequencies is considered in the spatial domain at multiple spatial scales to localise the salient features and give the final salient areas. The effectiveness and robustness of our approach are demonstrated by comparisons to previous approaches on a range of test models. The benefits of the proposed method are further evaluated in applications such as mesh simplification, mesh segmentation, and scan integration, where we show how incorporating mesh saliency can provide improved results. Ran Song 0001, Yonghuai Liu, Ralph R. Martin, Paul L. Rosin |
ACM Trans. Graph. | 1 |
| 2013 | 3D point of interest detection via spectral irregularity diffusion
Ran Song 0001, Yonghuai Liu, Ralph R. Martin, Paul L. Rosin |
Vis. Comput. | 1 |
| 2012 | Saliency-guided integration of multiple scansabstractWe present a novel method to integrate multiple 3D scans captured from different viewpoints. Saliency information is used to guide the integration process. The multi-scale saliency of a point is specifically designed to reflect its sensitivity to registration errors. Then scans are partitioned into salient and non-salient regions through an Markov Random Field (MRF) framework where neighbourhood consistency is incorporated to increase the robustness against potential scanning errors. We then develop different schemes to discriminatively integrate points in the two regions. For the points in salient regions which are more sensitive to registration errors, we employ the Iterative Closest Point algorithm to compensate the local registration error and find the correspondences for the integration. For the points in non-salient regions which are less sensitive to registration errors, we integrate them via an efficient and effective point-shifting scheme. A comparative study shows that the proposed method delivers improved surface integration. Ran Song 0001, Yonghuai Liu, Ralph R. Martin, Paul L. Rosin |
CVPR | 1 |
| 2012 | A saliency detection based method for 3D surface simplificationabstractTo accelerate the processing for the integration, registration, representation and recognition of point clouds, it is of growing necessity to simplify the surface of 3-D models. Simplification is an approach to vary the levels of visual details as appropriate, thereby improving on the overall performance of applications. This paper proposes a saliency detection based points sampling method for mesh simplification. By generating and enhancing the saliency map, the regions which are visually important can be located. For the mesh simplification, the local details are captured by the saliency, while for the overall shape, the approach voxelizes the model and samples points in terms of the entropy of the shape index of vertices in voxels. We present a number of results to show that the method significantly simplifies the surface without distortion and loss of local details. Yitian Zhao, Yonghuai Liu, Ran Song 0001 |
ICASSP | 3 |
| 2012 | Conditional random field-based mesh saliencyabstractWe propose a new method for detecting mesh saliency, a reflection of perception-based regional importance for 3D meshes. The basic idea is to incorporate the Conditional Random Field (CRF) framework with a saliency detection process. We first produce a multi-scale representation for a mesh. Then, a CRF is designed to robustly detect salient regions utilising neighbourhood consistency. By inferring the CRF via belief propagation algorithm, we actually make use of the global statistic information in the saliency detection process. Experimental results demonstrate the robustness and the effectiveness of the proposed method. Ran Song 0001, Yonghuai Liu, Yitian Zhao, Ralph R. Martin, Paul L. Rosin |
ICIP | 1 |
| 2012 | Extended non-local means filter for surface saliency detectionabstractMesh surface saliency detection is an important preprocessing step for many 3D applications. The salient region can be used to find the objects that are important on 3D surface, the benefits of saliency detection in the 3D domain include mesh simplification, registration, segmentation, compression, etc. This paper proposes a novel saliency detection method by diffusing the shape index field with non-local means filer, generating a random centre surround operator to yield saliency map and enhancing the saliency with the Retinex theory. The effectiveness of this method is demonstrated by simplification and registration. Experimental results demonstrate that the proposed approach has achieved competitive results. Yitian Zhao, Yonghuai Liu, Ran Song 0001 |
ICIP | 3 |
| 2011 | MRF-based automatic image ordering and its application to mosaicingabstractA fast and robust auto-sorting method for image ordering based on Markov Random Fields (MRF) is proposed. We present a specific MRF model for the ordering problem and use pairwise phase correlation for the formulation. The MRF is inferred by a modified belief propagation (BP) method. Experimental results prove that the new method can reorder a disorganised collection of images without human input, prior information or restrictions, as just the first stage of a multi stage mosaicing process, but also provides information that can be used to guide a mosaicing process in order to reduce both local mismatch and global error accumulation. Ran Song 0001, Yonghuai Liu, Yitian Zhao, Ralph R. Martin, Paul L. Rosin |
ICASSP | 1 |
| 2010 | MRF Labeling for Multi-view Range Image Integration
Ran Song 0001, Yonghuai Liu, Ralph R. Martin, Paul L. Rosin |
ACCV (2) | 1 |