EDBT 2026 Demo / reviewers in the wild / expert
Zhongxue Gan 0001
dblp:121/6289-1
· DBLP profile ↗
58ranked-venue papers
0as first author
56since 2021 · last 2026
0000-0003-1365-396XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 36 · 36 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 12 since 2021Systems, architecture and hardware · 9 · 9 since 2021Human-computer interaction and ubiquitous computing · 6 · 6 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Neuromorphic ASIC Design for Dexterous Hand Control
Hengtan Zhang, Yifu Liang, Zhongxue Gan 0001, Lirong Zheng 0001, Zhuo Zou |
ISCAS | 6 |
| 2026 | Joint-Guided Spatial and Semantic Sensitive Diffusion Policy for Robotic ManipulationabstractImitation learning has shown strong potential for enabling robots to acquire dexterous manipulation skills by integrating visual observations with proprioceptive states. However, common approaches typically use visual encoders pretrained in computer vision domains, which mainly aim to extract generic representations without emphasizing the precise spatial and semantic structures that are crucial for robotic manipulation. In this work, we propose the Joint-Guided Spatial and Semantic Sensitive Diffusion Policy (S3D), which effectively fuses structured and generic features by incorporating depth and semantic maps with RGB and proprioceptive inputs to strengthen spatial–semantic understanding in manipulation. However, naively incorporating these multimodal representations inevitably introduces additional computational overhead. Thus, we introduce a Joint-Guided Dynamic Attention module that generates joint-conditioned queries to extract behavior-specific representations with controlled complexity. Experiments across a variety of simulated and real-world robotic manipulation tasks demonstrate that S3D yields consistent performance gains over state-of-the-art methods. Hongda Zhang, Siao Liu, Yi Liu 0027, Chun Ouyang 0002, Zhongxue Gan 0001 |
ICMR | 5 |
| 2026 | Integrating channel priors and shared pattern learning for enhanced motor imagery EEG classification
Yangjie Luo, Zihua Chen, Junkongshuai Wang, Zhongxue Gan 0001, Lihua Zhang 0002, Xiaoyang Kang 0001 |
Neurocomputing | 5 |
| 2026 | Multi-head graph contrastive learning with hop augmentation for node classification
Minhao Zou, Xiaofeng Meng 0006, Zhongxue Gan 0001, Chun Guan, Siyang Leng |
Pattern Recognit. | 4 |
| 2026 | Dynamic Grouping With a Self-Aware Computational Resource Allocation for Large-Scale Multi-Objective Optimization
Yuning Chen, Ziqing Zhou, Yi Liu 0027, Linqiang Hu, Zhuo Zou, Zhongxue Gan 0001, Chun Ouyang 0002 |
IEEE Trans. Evol. Comput. | 7 |
| 2026 | ViG3D-UNet: Volumetric Vascular Connectivity-Aware Segmentation via 3D Vision Graph RepresentationabstractAccurate vascular segmentation is essential for coronary visualization and the diagnosis of coronary heart disease. This task involves the extraction of sparse tree-like vascular branches from volumetric space. However, existing methods have faced significant challenges due to discontinuous vascular segmentation and missing endpoints. To address this issue, a 3D vision graph neural network framework, named ViG3D-UNet, was introduced. This method integrates 3D graph representation and aggregation within a U-shaped architecture to facilitate continuous vascular segmentation. The ViG3D module captures volumetric vascular connectivity and topology, while the convolutional module extracts fine vascular details. These two branches are combined through channel attention to form the encoder feature. Subsequently, a paperclip-shaped offset decoder minimizes redundant computations in the sparse feature space and restores the feature map size to match the original input dimensions. To evaluate the effectiveness of the proposed approach for continuous vascular segmentation, evaluations were performed on two public datasets, ASOCA and ImageCAS. The segmentation results show that the ViG3D-UNet surpassed competing methods in maintaining vascular segmentation connectivity while achieving high segmentation accuracy. Bowen Liu 0017, Chunlei Meng, Hongda Zhang, Ziqing Zhou, Zhongxue Gan 0001, Chun Ouyang 0002 |
IEEE J. Biomed. Health Informatics | 6 |
| 2025 | Once-Tuning-Multiple-Variants: Tuning Once and Expanded as Multiple Vision-Language Model VariantsabstractVision-language model (VLM) is one of the most important models for multi-modal tasks. Real industrial applications often meet the challenge of adapting VLMs to different scenarios, such as varying hardware platforms or performance requirements. Traditional methods involve training or fine-tuning to adapt multiple unique VLMs or using model compression techniques to create multiple compact models. These approaches are complex and resource-intensive. This paper introduces a novel paradigm called Once-Tuning-Multiple-Variants (OTMV). OTMV requires only a single tuning process to inject dynamic weight expansion capacity into the original VLM structure. This tuned VLM can then be expanded into multiple variants tailored for different scenarios in inference. The tuning mechanism of OTMV is inspired by the mathematical series expansion theorem, which helps to reduce the parameter size and memory requirements while maintaining accuracy for VLM. Experiment results show that OTMV-tuned models achieve comparable accuracy to baseline VLMs across various visual-language tasks. The experiments also demonstrate the dynamic expansion capability of OTMV-tuned VLMs, outperforming traditional model compression and adaptation techniques in terms of accuracy and efficiency. Chong Yu 0001, Tao Chen 0003, Zhongxue Gan 0001 |
CVPR | 3 |
| 2025 | Think4CPP: Reinforcement Learning by Thinking with Latent World Model for Safe Coverage Path Planning
Zhentang Liao, Zhongxue Gan 0001, Lihua Zhang 0002, Zhiyan Dong |
ICIC (20) | 3 |
| 2025 | Towards Advanced Emotional Care: Embodied Emotional Care System for Humanoid RobotsabstractIn modern healthcare, emotional well-being is critical to patient recovery and overall outcomes. However, limited availability of trained professionals and time constraints often hinder the delivery of consistent emotional support. To address this gap, we propose the Embodied Emotional Care System (EECS), a comprehensive humanoid robotic framework designed to deliver personalized emotional care through an integrated, multi-layered architecture. EECS analyzes dynamic facial expressions and real-time vocal inputs to extract the patient’s emotional state and semantic information, constructs context-aware prompts processed by an LLM for reasoning, and ultimately generates empathetic dialogues synchronized with human-like facial expressions and natural body movements to address diverse emotional support needs. Experimental results show that deploying EECS on a humanoid robot significantly boosts patient engagement through real-time multimodal interaction, delivering deeper emotional support and a more human-like therapeutic experience. Furthermore, it bridges gaps in professional emotional support resources, offering a feasible pathway to improve overall healthcare quality. Yang Chang, Aoxing Li, Yuxuan Lin 0001, Lizheng Liu, Yang Liu 0246, Jing Liu 0050, Yan Wang 0068, Zhongxue Gan 0001 |
ICME | 10 |
| 2025 | HGS-Planner: Hierarchical Planning Framework for Active Scene Reconstruction Using 3D Gaussian SplattingabstractIn complex missions such as search and rescue, robots must make intelligent decisions in unknown environments, relying on their ability to perceive and understand their surroundings. High-quality and real-time reconstruction enhances situational awareness and is crucial for intelligent robotics. Traditional methods often struggle with poor scene representation or are too slow for real-time use. Inspired by the efficacy of 3D Gaussian Splatting (3DGS), we propose a hierarchical planning framework for fast and high-fidelity active reconstruction. Our method evaluates completion and quality gain to adaptively guide reconstruction, integrating global and local planning for efficiency. Experiments in simulated and realworld environments show our approach outperforms existing real-time methods. Ke Wu 0021, Zhiwei Zhang 0032, Jieru Zhao, Fei Gao 0011, Zhongxue Gan 0001, Wenchao Ding 0001 |
ICRA | 8 |
| 2025 | A Modified Resistance Model for Magnetic Honeycomb Robots to Navigate in Low Reynolds Number FluidsabstractIn recent years, magnetically controlled microrobots have garnered significant attention. This paper presents the H-robot, a self-designed microrobot featuring an innovative structure. The H-robot features a honeycomb porous spherical design specifically engineered to enhance cargo capacity. A new dynamic model for this structure has been developed for low Reynolds number fluid environments, along with a robust backstepping sliding mode control (RBSMC) strategy. Experiments were conducted in a calibrated magnetic field generated by a magnetic field generator to achieve precise motion control. The results demonstrate that the H-robot accurately tracks standard trajectories, with root mean square errors (RMSE) of$9.09 \times 10^{-4} \mathbf{~ m}$for the Number-8 path and$8.29 \times 10^{-4} \mathbf{~ m}$for the S-shaped path. Additionally, the proposed resistance model enhances tracking accuracy by 73.61% compared to traditional models, effectively adjusting the dynamic behavior of the H-robot in low Reynolds number fluids and significantly improving its motion performance. Finally, path planning experiments in a maze demonstrate the H-robot's ability to navigate and avoid obstacles. Leyao Zou, Shihao Ma, Yi Liu 0027, Xinyang Dong, Ziqing Zhou, Chun Ouyang 0002, Zhongxue Gan 0001 |
ICRA | 7 |
| 2025 | ACORN: Acyclic Coordination with Reachability Network to Reduce Communication Redundancy in Multi-Agent Systems
Ziqing Zhou, Chun Ouyang 0002, Siao Liu, Linqiang Hu, Zhongxue Gan 0001 |
AAMAS | 6 |
| 2025 | Heuristics-Assisted Experience Replay Strategy for Cooperative Multi-Agent Reinforcement Learning
Ziqing Zhou, Chun Ouyang 0002, Siao Liu, Linqiang Hu, Zhongxue Gan 0001 |
AAMAS | 6 |
| 2025 | Boost Embodied AI Models with Robust Compression BoundaryabstractThe rapid improvement of deep learning models with the integration of the physical world has dramatically improved embodied AI capabilities. Meanwhile, the powerful embodied AI models and their scales place an increasing burden on deployment efficiency. The efficiency issue is more apparent on embodied AI platforms than on data centers because they have more limited computational resources and memory bandwidth. Meanwhile, most embodied AI scenarios, like autonomous driving and robotics, are more sensitive to fast responses. Theoretically, the traditional model compression techniques can help embodied AI models with more efficient computation, lower memory and energy consumption, and reduced latency. Because the embodied AI models are expected to interact with the physical world, the corresponding compressed models are also expected to resist natural corruption caused by real-world events such as noise, blur, weather conditions, and even adversarial corruption. This paper explores the novel paradigm to boost the efficiency of the embodied AI models and the robust compression boundary. The efficacy of our method has been proven to find the optimal balance between accuracy, efficiency, and robustness in real-world conditions. Chong Yu 0001, Tao Chen 0003, Zhongxue Gan 0001 |
IJCAI | 3 |
| 2025 | SO-DETR: Leveraging Dual-Domain Features and Knowledge Distillation for Small Object DetectionabstractDetection Transformer-based methods have achieved significant advancements in general object detection. However, challenges remain in effectively detecting small objects. One key difficulty is that existing encoders struggle to efficiently fuse low-level features. Additionally, the query selection strategies are not effectively tailored for small objects. To address these challenges, this paper proposes an efficient model, Small Object Detection Transformer (SO-DETR). The model comprises three key components: a dual-domain hybrid encoder, an enhanced query selection mechanism, and a knowledge distillation strategy. The dual-domain hybrid encoder integrates spatial and frequency domains to fuse multi-scale features effectively. This approach enhances the representation of high-resolution features while maintaining relatively low computational overhead. The enhanced query selection mechanism optimizes query initialization by dynamically selecting high-scoring anchor boxes using expanded IoU, thereby improving the allocation of query resources. Furthermore, by incorporating a lightweight backbone network and implementing a knowledge distillation strategy, we develop an efficient detector for small objects. Experimental results on the VisDrone-2019-DET and UAVVaste datasets demonstrate that SO-DETR outperforms existing methods with similar computational demands. The project page is available at https://github.com/ValiantDiligent/SODETR. Huaxiang Zhang 0002, Aoran Mei, Zhongxue Gan 0001, Guoniu Zhu |
IJCNN | 4 |
| 2025 | Topology-Driven Trajectory Optimization for Modelling Controllable Interactions Within Multi-Vehicle ScenarioabstractTrajectory optimization in multi-vehicle scenarios faces challenges due to its non-linear, non-convex properties and sensitivity to initial values, making interactions between vehicles difficult to control. In this paper, inspired by topological planning, we propose a differentiable local homotopy invariant metric to model the interactions. By incorporating this topological metric as a constraint into multi-vehicle trajectory optimization, our framework is capable of generating multiple interactive trajectories from the same initial values, achieving controllable interactions as well as supporting user-designed interaction patterns. Extensive experiments demonstrate its superior optimality and efficiency over existing methods. We will release open-source code to advance relative research1. Changjia Ma, Zhongxue Gan 0001, Bingzhao Gao, Wenchao Ding 0001 |
IROS | 3 |
| 2025 | Spherical Scissor-Like Reconfigurable Palm Design in Robotic Hands: Insights from Human Hand FunctionalityabstractThe human palm demonstrates spatial reconfigurability during the gripping process and forms a spherical grasping envelope. Based on these observations, this study designs a reconfigurable spherical palm that incorporates a spatial scissor mechanism, which only requires a single actuator to reshape the palm into a range of spherical forms. We conduct a kinematic analysis and modelling of the structure, abstracting three key parameters and analysing their influence on the motion characteristics of the palm. Through multi-objective optimisation, a set of dimensional parameters is derived to balance workspace, human-like motion, and mechanical performance. The performance of the reconfigurability and the grasping capability of the proposed palm is compared to a planar folding palm by superquadrics, and the results show that the spherical design and the reconfigurable characteristics provide larger grasping arrangement and stronger grasping capability of the palm on most of the testing surfaces. Kai Chen 0026, Chang Liu 0030, Guoniu Zhu, Qiujie Lu, Zhongxue Gan 0001 |
IROS | 7 |
| 2025 | UAV-DETR: Efficient End-to-End Object Detection for Unmanned Aerial Vehicle ImageryabstractUnmanned aerial vehicle object detection (UAV-OD) has been widely used in various scenarios. However, most existing UAV-OD algorithms rely on manually designed components, which require extensive tuning. End-to-end models that do not depend on such manually designed components are mainly designed for natural images, which are less effective for UAV imagery. To address such challenges, this paper proposes an efficient detection transformer (DETR) framework tailored for UAV imagery, i.e., UAV-DETR. The framework includes a multi-scale feature fusion with frequency enhancement module, which captures both spatial and frequency information at different scales. In addition, a frequency-focused downsampling module is presented to retain critical spatial details during downsampling. A semantic alignment and calibration module is developed to align and fuse features from different fusion paths. Experimental results demonstrate the effectiveness and generalization of our approach across various UAV imagery datasets. On the VisDrone dataset, our method improves AP by 3.1% and AP50 by 4.2% over the baseline. Similar enhancements are observed on the UAVVaste dataset. The project page is available at https://github.com/ValiantDiligent/UAV-DETR. Huaxiang Zhang 0002, Zhongxue Gan 0001, Guoniu Zhu |
IROS | 4 |
| 2025 | Learning Occlusion-aware Decision-making from Agent Interaction via Active PerceptionabstractOne of the unresolved challenges for autonomous vehicles is occlusion-aware decision-making under the high uncertainty of various occlusions. Recent occlusion-aware decision-making methods encounter issues such as overly conservative behavior, high computational complexity, or scenario scalability challenges. Benefiting from automatically generating data by exploration randomization, we uncover that reinforcement learning (RL) may show promise in occlusion-aware decision-making. However, previous occlusion-aware RL faces challenges in expanding to various dynamic and static occlusion scenarios, low learning efficiency, and lack of predictive ability. To address these issues, we introduce Pad-AI, a self-reinforcing framework to learn occlusion-aware decision-making through active perception. Pad-AI utilizes vectorized representation to represent occluded environments efficiently and learns over the semantic motion primitives to focus on high-level active perception exploration. Furthermore, Pad-AI integrates prediction and RL within a unified framework to provide risk-aware learning and reliable policy optimization. Our framework was tested in challenging scenarios under both dynamic and static occlusions and demonstrated efficient perception-aware exploration performance to other strong baselines in closed-loop evaluations. Jie Jia 0002, Yiming Shu, Zhongxue Gan 0001, Wenchao Ding 0001 |
IV | 3 |
| 2025 | Non-Reciprocal Interactions Based Emergent Navigation for 3D Autonomous Drones SwarmabstractWe address a fundamental challenge in coordinating large-scale 3D drone swarms: how to achieve rapid collective response to environmental stimuli while ensuring group stability and safety. Existing swarm navigation modals often rely on sophisticated individual perception and communication capabilities, which can be computationally expensive and impractical for large swarms. In this paper, we propose the Non-reciprocal Collective Emergent Navigation model (NRCE), a decentralized approach designed for real-world drone flocking in complex environments. Unlike traditional models, our approach leverages localized non-reciprocal interactions, where boundary drones detect environmental stimuli and propagate this information throughout the swarm without directly controlling individual trajectories. Through extensive numerical simulations and physical experiments with up to 28 drones, we demonstrate how this model achieves coordinated collective motion while effectively balancing stability with responsiveness. Our findings reveal two notable insights: (1) intermediate cohesion levels (ωc) optimize collective response—a "Goldilocks zone" where individuals are neither too tightly coupled nor too independent, challenging the conventional wisdom that stronger cohesion always improves coordination; and (2) swarm queue configuration significantly affects optimal interaction parameters, with divergent trends observed between attraction- and repulsion-based coordination mechanisms as layer count increases. These discoveries provide critical design principles for cost-effective, high-density swarm systems while advancing the theoretical understanding of collective dynamics in both artificial and biological systems. Linqiang Hu, Ziqing Zhou, Yuning Chen, Hongda Zhang, Chunlei Meng, Yi Liu 0027, Zhiyan Dong, Chun Ouyang 0002, Zhongxue Gan 0001, Dunzhao Wu, Zhihua Nie |
SMC | 10 |
| 2025 | Pheromone-Focused Ant Colony Optimization algorithm for path planningabstractAnt Colony Optimization (ACO) is a prominent swarm intelligence algorithm extensively applied to path planning. However, traditional ACO methods often exhibit shortcomings, such as blind search behavior and slow convergence within complex environments. To address these challenges, this paper proposes the Pheromone-Focused Ant Colony Optimization (PFACO) algorithm, which introduces three key strategies to enhance the problem-solving ability of the ant colony. First, the initial pheromone distribution is concentrated in more promising regions based on the Euclidean distances of nodes to the start and end points, balancing the trade-off between exploration and exploitation. Second, promising solutions are reinforced during colony iterations to intensify pheromone deposition along high-quality paths, accelerating convergence while maintaining solution diversity. Third, a forward-looking mechanism is implemented to penalize redundant path turns, promoting smoother and more efficient solutions. These strategies collectively produce the focused pheromones to guide the ant colony’s search, which enhances the global optimization capabilities of the PFACO algorithm, significantly improving convergence speed and solution quality across diverse optimization problems. The experimental results demonstrate that PFACO consistently outperforms comparative ACO algorithms in terms of convergence speed and solution quality. Yi Liu 0027, Hongda Zhang, Zhongxue Gan 0001, Yuning Chen, Ziqing Zhou, Chunlei Meng, Chun Ouyang 0002 |
SMC | 3 |
| 2025 | CF-ViT: Cross-Feature Vision Transformer for Improving Feature Learning on Tiny DatasetsabstractEfficient feature learning is considered indispensable for maximizing the representation of scarce information in tiny datasets. However, existing methods are often unable to fully exploit local features and contextual dependencies when dealing with tiny datasets. To overcome this shortcoming, a Cross-Feature Vision Transformer (CF-ViT) was proposed, which decouples local feature refinement from global context modeling and leverages the complementary strengths of CNNs and Transformers. Specifically, a Cross-Scale Fusion (CSF) module was introduced to integrate features from multiple scales, ensuring that cross-scale information is globally embedded. In addition, a Feature Enhancement and Reorganization (FER) module was incorporated into CF-ViT, whereby Transformer outputs are reorganized into 2D feature maps for convolution-based detail enhancement to thoroughly exploit local information. Extensive experiments have demonstrated that CF-ViT consistently surpasses baselines across 4 tiny datasets, reaching a 96.87% (KSDD) Top-1 accuracy with only 29.19 million parameters and 2.67 billion FLOPs. Moreover, a Top-1 accuracy of 85.03% is attained on a real-world tiny dataset of wood surface defect detection, exceeding all baselines. These findings underscore the effectiveness and generalization capability of CF-ViT in capturing fine-grained local details and global context, offering a promising and deployable solution for vision tasks in tiny datasets. Chunlei Meng, Yi Liu 0027, Hongda Zhang, Yuning Chen, Bowen Liu 0017, Ziqin Zhou, Chun Ouyang 0002, Zhongxue Gan 0001, Dunzhao Wu, Zhihua Nie |
SMC | 10 |
| 2025 | Topology-preserving and structure-aware (hyper)graph contrastive learning for node classification
Minhao Zou, Zhongxue Gan 0001, Junheng Zhang, Chun Guan, Siyang Leng |
Appl. Intell. | 2 |
| 2025 | RPN: A region-to-pixel-mask-based convolutional network for lesion segmentation of fundus images
Hongda Zhang, Chun Ouyang 0002, Zhonghong Shen, Bowen Liu 0017, Yi Liu 0027, Zhongxue Gan 0001 |
Neurocomputing | 7 |
| 2025 | Taylor-Series-Expansion-Based Vision Transformer ModelsabstractTaylor-Series-Expansion (TSE) is a mathematics theorem. It proves that the expansion of the first few finite Taylor Series is a good approximation of a nonlinear function in most cases. Inspired by the TSE theorem, a brand-new TSE-based vision transformer is designed. TSE-based vision transformer uses the shared first-order TSE transformer block's weight (in analogy with the Taylor-Series first-order term), its finite multiple multiplications (in analogy with the Taylor-Series expanded high-order terms), and the corresponding learnable TSE coefficients to approximate the naive vision transformer. In this manner, the TSE-based vision model reduces the memory burden but keeps a similar accuracy as the naive counterpart. Derived from adding the Taylor skip mechanism in training, the TSE-based vision transformer has good dynamic expansion capability. Experiment results show TSE-based models can boost actual deployment latency by 1.30-1.36× on A100 GPU and 1.34-1.45× on AGX Orin with negligible accuracy degradation on ImageNet classification, COCO detection, and ADE20K segmentation benchmarking tasks. Moreover, TSE-based optimization is orthogonal to model compression. Combining with the state-of-the-art vision transformer compression method, it can boost actual deployment performance by 1.70-1.87× and 3.29-3.61× of latency and throughput on A100 GPU, and 1.67-1.74× and 2.76-2.94× improvement of latency and throughput on AGX Orin. Chong Yu 0001, Tao Chen 0003, Zhongxue Gan 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | Real-Time Scheduling Framework for Multiagent Cooperative Logistics With Dynamic Supply DemandsabstractIn logistics systems with multiagent collaboration, one of the prevailing focus lies on modeling as the dynamic multiperiod vehicle routing problem (DMPVRP). This work introduces modifications to DMPVRP to align with the requirements of real factory operations, particularly with dynamic supply demands. A self-established multiagent dynamic scheduling framework has been proposed to adapt to dynamic environmental changes and make timely adjustments, which consists of two modules: dynamic path planning and machine assignment. The first module utilizes a self-designed multioperator two-stage evolutionary algorithm to dynamically update the routes for vehicles. The second module maintains the workload balance among vehicles in real time. Experimental results demonstrate that the proposed algorithm achieves optimal outcomes compared to three state-of-the-art algorithms, surpassing others by 20% in machine output and exhibiting 5% lower transportation costs. In addition, a case study from a steel cord manufacturing factory is conducted, demonstrating its capability to promptly enhance efficiency. Yuning Chen, Yi Liu 0027, Hongda Zhang, Ziqing Zhou, Wenchao Ding 0001, Zhuo Zou, Chun Ouyang 0002, Zhongxue Gan 0001 |
IEEE Trans. Ind. Informatics | 10 |
| 2025 | RTS-ViT: Real-Time Share Vision Transformer for Image ClassificationabstractVision transformers have achieved remarkable success in image classification. The dual-branch vision transformer generates more features by taking advantage of feature fusion. Inspired by this, a dual-branch vision transformer with Real-Time Share feature was proposed during the encoding process for retinal image classification tasks. The approach processes image patches of varying sizes (base and large) through two independent branches and implements multi-stage Real-Time feature fusion via the Real-Time Share feature encoder. This encoder enables the branches to complement each other's features at each encoding stage, facilitating finer feature learning and enhancing the self-attention information passed to subsequent stages. It significantly boosts feature representation and classification performance. Additionally, a straightforward and effective feature fusion method, L-Times Attention Fusion, was proposed: vector concatenation for Real-Time Share feature in the earlier (L-1) encoding stages and element-wise addition for overall feature fusion at the L-th stage, achieving more efficient feature integration. The method was validated on a retinal image dataset. Results show that the approach outperforms the recent Cross-ViT average TOP-1 Acc by 5.61% with lower FLOPs and model parameters, without relying on pre-trained weights, highlighting stronger self-learning feature capabilities and reduced reliance on extensive pre-training data. Chunlei Meng, Bowen Liu 0017, Hongda Zhang, Zhongxue Gan 0001, Chun Ouyang 0002 |
IEEE J. Biomed. Health Informatics | 5 |
| 2025 | Primitive-Swarm: An Ultra-Lightweight and Scalable Planner for Large-Scale Aerial SwarmsabstractAchieving large-scale aerial swarms is challenging due to the inherent contradictions in balancing computational efficiency and scalability. This paper introducesPrimitive-Swarm, an ultra-lightweight and scalable planner designed specifically for large-scale autonomous aerial swarms. The proposed approach adopts a decentralized and asynchronous replanning strategy. Within it is a novel motion primitive library consisting of time-optimal and dynamically feasible trajectories. They are generated utlizing a novel time-optimial path parameterization algorithm based on reachability analysis (TOPP-RA). Then, a rapid collision checking mechanism is developed by associating the motion primitives with the discrete surrounding space according to conflicts. By considering both spatial and temporal conflicts, the mechanism handles robot-obstacle and robot-robot collisions simultaneously. Then, during a replanning process, each robot selects the safe and minimum cost trajectory from the library based on user-defined requirements. Both the time-optimal motion primitive library and the occupancy information are computed offline, turning a time-consuming optimization problem into a linear-complexity selection problem. This enables the planner to comprehensively explore the non-convex, discontinuous 3-D safe space filled with numerous obstacles and robots, effectively identifying the best hidden path. Benchmark comparisons demonstrate that our method achieves the shortest flight time and traveled distance with a computation time of less than 1 ms in dense environments. Super large-scale swarm simulations, involving up to 1000 robots, running in real-time, verify the scalability of our method. Real-world experiments validate the feasibility and robustness of our approach. The code will be released to foster community collaboration. Jialiang Hou, Xin Zhou 0015, Neng Pan, Ang Li 0042, Chao Xu 0001, Zhongxue Gan 0001, Fei Gao 0011 |
IEEE Trans. Robotics | 7 |
| 2025 | VINGS-Mono: Visual-Inertial Gaussian Splatting Monocular SLAM in Large Scenes
Ke Wu 0021, Muer Tie, Ziqing Ai, Zhongxue Gan 0001, Wenchao Ding 0001 |
IEEE Trans. Robotics | 5 |
| 2024 | Swift-Mapping: Online Neural Implicit Dense Mapping in Urban ScenesabstractOnline dense mapping of urban scenes is of paramount importance for scene understanding of autonomous navigation. Traditional online dense mapping methods fuse sensor measurements (vision, lidar, etc.) across time and space via explicit geometric correspondence. Recently, NeRF-based methods have proved the superiority of neural implicit representations by high-fidelity reconstruction of large-scale city scenes. However, it remains an open problem how to integrate powerful neural implicit representations into online dense mapping. Existing methods are restricted to constrained indoor environments and are too computationally expensive to meet online requirements. To this end, we propose Swift-Mapping, an online neural implicit dense mapping framework in urban scenes. We introduce a novel neural implicit octomap (NIO) structure that provides efficient neural representation for large and dynamic urban scenes while retaining online update capability. Based on that, we propose an online neural dense mapping framework that effectively manages and updates neural octree voxel features. Our approach achieves SOTA reconstruction accuracy while being more than 10x faster in reconstruction speed, demonstrating the superior performance of our method in both accuracy and efficiency. Ke Wu 0021, Kaizhao Zhang, Mingzhe Gao, Jieru Zhao, Zhongxue Gan 0001, Wenchao Ding 0001 |
AAAI | 5 |
| 2024 | O 2V-Mapping: Online Open-Vocabulary Mapping with Neural Implicit Representation
Muer Tie, Julong Wei, Ke Wu 0021, Zhengjun Wang, Shanshuai Yuan, Kaizhao Zhang, Jie Jia 0002, Jieru Zhao, Zhongxue Gan 0001, Wenchao Ding 0001 |
ECCV (87) | 9 |
| 2024 | OpenAnnotate3D: Open-Vocabulary Auto-Labeling System for Multi-modal 3D DataabstractIn the era of big data and large models, automatic annotating functions for multi-modal data are of great significance for real-world AI-driven applications, such as autonomous driving and embodied AI. Unlike traditional closed-set annotation, open-vocabulary annotation is essential to achieve human-level cognition capability. However, there are few open-vocabulary auto-labeling systems for multi-modal 3D data. In this paper, we introduce OpenAnnotate3D, an open-source open-vocabulary auto-labeling system that can automatically generate 2D masks, 3D masks, and 3D bounding box annotations for vision and point cloud data. Our system integrates the chain-of-thought capabilities of Large Language Models (LLMs) and the cross-modality capabilities of vision-language models (VLMs). To the best of our knowledge, OpenAnnotate3D is one of the pioneering works for open-vocabulary multi-modal 3D auto-labeling. We conduct comprehensive evaluations on both public and in-house real-world datasets, which demonstrate that the system significantly improves annotation efficiency compared to manual annotation while providing accurate open-vocabulary auto-annotating results. Likun Cai, Xianhui Cheng, Zhongxue Gan 0001, Xiangyang Xue 0001, Wenchao Ding 0001 |
ICRA | 4 |
| 2024 | Spear: Evaluate the Adversarial Robustness of Compressed Neural Models
Chong Yu 0001, Tao Chen 0003, Zhongxue Gan 0001, Jiayuan Fan 0001 |
IJCAI | 3 |
| 2024 | Robust Robot Formation Control Based on Streaming Communication and Leader-Follower ApproachabstractIn this paper, a non-visual robotic formation control method based on a streaming communication architecture and a leader-follower control model is studied. The proposed stream-based communication architecture is inspired by the flocking behavior of fish. We analogize it into a form resembling an N-ary tree for communication purposes. Communication proceeds to the next layer only when all nodes in the upper layer have completed the follower selection. We also introduce a fault-tolerance mechanism and a termination filtering mechanism to prevent multiple leaders from choosing the same follower, avoiding a scenario where the robots in the last layer enter an endless loop of follower selection. The proposed stream-based communication architecture, built upon serial and parallel tracking, can achieve more complex formations, such as rectangular formations, closely resembling real-world scenarios, significantly enhancing formation efficiency. Simulation experiments on the e-puck platform validate the effectiveness and robustness of this architecture. Zhuo Zou, Xiaoming Hu 0001, Zhongxue Gan 0001, Lizheng Liu |
IWCMC | 4 |
| 2024 | TF-Net: Triple Fusion Net for Medical Image SegmentationabstractLesion segmentation plays a crucial role in various medical image analyses, which not only improves the efficiency in clinical diagnosis but also assists in detecting early symptoms of various diseases. Most existing studies focus on directly extracting lesion information from specific types of medical images with pre-trained weights, often neglecting the underlying topological and pathological causes which lead to these lesions. Furthermore, they overlook to capture general anatomical features among lesions, which are related to the distribution of lesions, and thus the model is poorly generalized in different medical datasets. Inspired by these insights, we propose a Triple Fusion Net (TF-Net), a network structure divided into three branches: left, middle and right. The left and right branches are designed to extract lesion features and associated topological style features within various medical images, respectively. And these features are further fused and modeled in the middle branch. The proposed structure of triple branches for features fusing effectively learns multi-feature information and improves the performance of TF-Net. And our work experiments validate various feature fusion methods in the middle branch, including channel-wise concatenation, element-wise addition, attention gate, and transformer encoder block. Without using pre-trained weights in our network, the transformer encoder block performs best on some tasks of DDR and surpasses other pre-trained models. Channel concatenation exhibits performance close to other pre-trained models in both the IDRiD, Kvasir-Seg and TN3K. Attention gate fusion also shows competitive results in thyroid ultrasound segmentation. Our approach, leveraging a unique network structure and four different feature fusion methods, demonstrates remarkable generality across a spectrum of medical image segmentation tasks. Chunlei Meng, Hongda Zhang, Bowen Liu 0017, Xinyang Dong, Chun Ouyang 0002, Zhongxue Gan 0001 |
SMC | 8 |
| 2024 | Joint Optimization of Recurrence Plot Encoding and CNN Model Based on Heuristic AlgorithmsabstractPeripheral waveform analysis (PWA), which is generally used to reveal hidden health status information from peripheral pulse signals, typically involves three procedures: signal preprocessing, feature extraction, and pattern classification. With the advancement of data-driven deep neural network methodologies, feature extraction and pattern classification have progressively converged into end-to-end neural networks, where the final layer of the network is equivalent to conventional pattern classifiers. However, the performance of deep learning models heavily relies on the quality of the dataset, rendering data signal preprocessing a crucial component. This study proposes a framework that integrates signal preprocessing, feature extraction, and pattern classification into a unified learning approach using heuristic algorithms, enabling the automatic discovery of optimal data encoding methods and their corresponding models. Initially, the search space is defined based on parameters relevant to signal preprocessing, and a fitness function is constructed utilizing CNN. Subsequently, the optimal combination of data preprocessing and CNN is determined through the heuristic algorithm Particle Swarm Optimization (PSO). The proposed method was evaluated in the dataset comprising authentic clinical cases of type 2 diabetes screening involving approximately 200 volunteers. The model derived from this framework demonstrates the capability to effectively discriminate between healthy volunteers and those with diabetes, achieving the highest accuracy of 93.6%. Compared to state-of-the-art algorithms, the proposed model was shown to be competitive in both accuracy and time cost. Hongda Zhang, Zhongxue Gan 0001, Yi Liu 0027, Bowen Liu 0017, Chunlei Meng, Chun Ouyang 0002 |
SMC | 2 |
| 2024 | Heterogeneous Robot Swarms with an Attention Mechanism for Dynamic Target TrackingabstractMultirobot collaboration offers significant potential for diverse applications, including tracking and surveillance. In this paper, we introduce an attention mechanism tailored for heterogeneous robot swarms characterized by varied sensing ranges. This mechanism effectively utilizes the swarm's intrinsic characteristics, enabling rapid information transmission and ensuring consistent collective responses to external stimuli. Additionally, we introduce a pigeon-inspired navigation strategy that effectively replaces the traditional obstacle repulsion term by preventing the swarm from becoming trapped in local min-ima and reducing oscillatory behaviors. To validate the efficacy of our algorithm, we have developed an autonomously designed PlusBot swarm platform, which consists of agile vibration-driven miniature robots. Each of them is equipped with its own computing and communication system and is capable of precise closed-loop motion control. This setup meets the requirements for conducting heterogeneous swarm movement experiments in indoor environments. Through comprehensive numerical simulations and real-world experiments, our method has demonstrated exceptional precision and adaptability in tracking dynamic targets. The comparative analysis under-scores the superiority of our approach, particularly in minimizing swarm collisions and ensuring safe navigation in dynamic target-tracking scenarios involving obstacles. Ziqing Zhou, Chun Ouyang 0002, Xinyang Dong, Siao Liu, Linqiang Hu, Zhile Zhao, Zhongxue Gan 0001 |
SMC | 9 |
| 2024 | A framework for dynamical distributed flocking control in dense environments
Ziqing Zhou, Chun Ouyang 0002, Linqiang Hu, Yuning Chen, Zhongxue Gan 0001 |
Expert Syst. Appl. | 6 |
| 2024 | DiffSkill: Improving Reinforcement Learning through diffusion-based skill denoiser for robotic manipulation
Siao Liu, Yang Liu 0246, Linqiang Hu, Ziqing Zhou, Zhile Zhao, Wei Li 0055, Zhongxue Gan 0001 |
Knowl. Based Syst. | 8 |
| 2024 | UniG-Encoder: A universal feature encoder for graph and hypergraph node classification
Minhao Zou, Zhongxue Gan 0001, Junheng Zhang, Dongyan Sui, Chun Guan, Siyang Leng |
Pattern Recognit. | 2 |
| 2024 | Evaluation of Frameworks That Combine Evolution and Learning to Design Robots in Complex Morphological SpacesabstractJointly optimising both the body and brain of a robot is known to be a challenging task, especially when attempting to evolve designs in simulation that will subsequently be built in the real world. To address this, it is increasingly common to combine evolution with a learning algorithm that can either improve the inherited controllers of new offspring to fine tune them to the new body design or learn them from scratch. In this paper an approach is proposed in which a robot is specified indirectly by two compositional pattern producing networks (CPPN) encoded in a single genome, one which encodes the brain and the other the body. The body part of the genome is evolved using an evolutionary algorithm (EA), with an individual learning algorithm (also an EA) applied to the inherited controller to improve it. The goal of this paper is to determine how to utilise the results of learning process most effectively to improve task performance of the robot. Specifically, three variants are investigated: (1) evolution of the body+controller only; (2) a learning algorithm is applied to the inherited controller with the learned fitness assigned to the genome; (3) learning is applied and the genome is updated with the learned controller, as well as being assigned the learned fitness. Experiments are performed in three different scenarios chosen to favour different bodies and locomotion patterns. It is shown that better performance can be obtained using learning but only if the learned controller is inherited by the offspring. Wei Li 0055, Edgar Buchanan, Leni K. Le Goff, Emma Hart, Matthew F. Hale, Bingsheng Wei, Matteo De Carlo, Mike Angus, Robert Woolley, Zhongxue Gan 0001, Alan F. T. Winfield, Jonathan Timmis, A. E. Eiben, Andrew M. Tyrrell |
IEEE Trans. Evol. Comput. | 10 |
| 2023 | Boost Vision Transformer with GPU-Friendly Sparsity and QuantizationabstractThe transformer extends its success from the language to the vision domain. Because of the stacked self-attention and cross-attention blocks, the acceleration deployment of vision transformer on GPU hardware is challenging and also rarely studied. This paper thoroughly designs a compression scheme to maximally utilize the GPU-friendly 2:4 fine-grained structured sparsity and quantization. Specially, an original large model with dense weight parameters is first pruned into a sparse one by 2:4 structured pruning, which considers the GPU's acceleration of 2:4 structured sparse pattern with FP16 data type, then the floating-point sparse model is further quantized into a fixed-point one by sparse-distillation-aware quantization aware training, which considers GPU can provide an extra speedup of 2:4 sparse calculation with integer tensors. A mixed-strategy knowledge distillation is used during the pruning and quantization process. The proposed compression scheme is flexible to support supervised and unsupervised learning styles. Experiment results show GPUSQ-ViT scheme achieves state-of-the-art compression by reducing vision transformer models$\mathbf{6.4}-\mathbf{12.7}\times$on model size and$\mathbf{30.3}-\mathbf{62} \times$on FLOPs with negligible accuracy degradation on ImageNet classification, COCO detection and ADE20K segmentation benchmarking tasks. Moreover, GPUSQ-ViT can boost actual deployment performance by$\mathbf{1.39}-\mathbf{1.79}\times$and$\mathbf{3.22}-\mathbf{3.43}\times$of latency and throughput on A100 GPU, and$\mathbf{1.57}-\mathbf{1.69}\times$and$\mathbf{2.11}-\mathbf{2.51}\times$improvement of latency and throughput on AGX Orin. Chong Yu 0001, Tao Chen 0003, Zhongxue Gan 0001, Jiayuan Fan 0001 |
CVPR | 3 |
| 2023 | Learning-Based Neural Ant Colony OptimizationabstractIn this paper, we propose a new ant colony optimization algorithm, called learning-based neural ant colony optimization (LN-ACO), which incorporates an "intelligent ant". This intelligent ant contains a convolutional neural network pre-trained on a large set of instances which is able to predict the selection probabilities of the set of possible choices at each step of the algorithm. The intelligent ant is capable of generating a solution based on knowledge learned during training, but also guides other 'traditional' ants in improving their choices during the search. As the search progresses, the intelligent ant is also influenced by the pheromones accumulated by the colony, leading to better solutions. The key idea is that if tasks or instances share common features either in terms of their search landscape or solutions, then information learned by solving one instance can be applied to substantially accelerate the search on another. We evaluate the proposed algorithm on two public datasets and one real-world test set in the path planning domain. The results demonstrate that LN-ACO is competitive in its search capability compared to other ACO methods, with a significant improvement in convergence speed. Yi Liu 0027, Jiang Qiu, Emma Hart, Yilan Yu, Zhongxue Gan 0001, Wei Li 0055 |
GECCO | 5 |
| 2023 | Improving Generalization in Visual Reinforcement Learning via Conflict-aware Gradient Agreement AugmentationabstractLearning a policy with great generalization to unseen environments remains challenging but critical in visual reinforcement learning. Despite the success of augmentation combination in the supervised learning generalization, naively applying it to visual RL algorithms may damage the training efficiency, suffering from serve performance degradation. In this paper, we first conduct qualitative analysis and illuminate the main causes: (i) high-variance gradient magnitudes and (ii) gradient conflicts existed in various augmentation methods. To alleviate these issues, we propose a general policy gradient optimization framework, named Conflict-aware Gradient Agreement Augmentation (CG2A), and better integrate augmentation combination into visual RL algorithms to address the generalization bias. In particular, CG2A develops a Gradient Agreement Solver to adaptively balance the varying gradient magnitudes, and introduces a Soft Gradient Surgery strategy to alleviate the gradient conflicts. Extensive experiments demonstrate that CG2A significantly improves the generalization performance and sample efficiency of visual RL algorithms. Siao Liu, Zhaoyu Chen 0001, Yang Liu 0246, Dingkang Yang, Zhile Zhao, Ziqing Zhou, Xie Yi, Wei Li 0055, Zhongxue Gan 0001 |
ICCV | 11 |
| 2023 | FlowMap: Path Generation for Automated Vehicles in Open Space Using Traffic FlowabstractThere is extensive literature on perceiving road structures by fusing various sensor inputs such as lidar point clouds and camera images using deep neural nets. Leveraging the latest advance of neural architects (such as transformers) and bird-eye-view (BEV) representation, the road cognition accuracy keeps improving. However, how to cognize the “road” for automated vehicles where there is no well-defined “roads” remains an open problem. For example, how to find paths inside intersections without HD maps is hard since there is neither an explicit definition for “roads” nor explicit features such as lane markings. The idea of this paper comes from a proverb: it becomes a way when people walk on it. Although there are no “roads” from sensor readings, there are “roads” from tracks of other vehicles. In this paper, we propose FlowMap, a path generation framework for automated vehicles based on traffic flows. FlowMap is built by extending our previous work RoadMap [1], a light-weight semantic map, with an additional traffic flow layer. A path generation algorithm on traffic flow fields (TFFs) is proposed to generate human-like paths. The proposed framework is validated using real-world driving data and is amenable to generating paths for super complicated intersections without using HD maps. Wenchao Ding 0001, Jieru Zhao, Yubin Chu, Haihui Huang, Tong Qin 0001, Chunjing Xu, Zhongxue Gan 0001 |
ICRA | 8 |
| 2023 | Mechanical Intelligence for Prehensile In-Hand Manipulation of Spatial TrajectoriesabstractThe application of mechanical and other physical properties to the development of robotic systems that can easily adapt to changing external situations is known as mechanical intelligence. Following this concept, many robot hand designs can produce self-adaptive and versatile grasps with simple underactuated fingers and open-loop control, while mechanical- intelligent strategies for dexterous manipulation are still limited. This paper proposes a mechanical-intelligent technique to facilitate dexterous manipulation, in particular prehensile inhand manipulation. The proposed strategy is based on the generation of complex spatial trajectories of the hand-object system, controlled in open loop with the minimum number of actuators and using simple low-level non-position modes. This approach is exemplified by the rigorous analysis and testing of a three-fingered two-actuator underactuated robot hand, called the helical hand, which is capable of generating helical prehensile in-hand manipulation of diversiform objects under error tolerance controlled by constant speed algorithm. Qiujie Lu, Zhongxue Gan 0001, Guochao Bai, Nicolás Rojas 0002 |
ICRA | 2 |
| 2023 | Adversarial Amendment is the Only Force Capable of Transforming an Enemy into a FriendabstractAdversarial attack is commonly regarded as a huge threat to neural networks because of misleading behavior. This paper presents an opposite perspective: adversarial attacks can be harnessed to improve neural models if amended correctly. Unlike traditional adversarial defense or adversarial training schemes that aim to improve the adversarial robustness, the proposed adversarial amendment (AdvAmd) method aims to improve the original accuracy level of neural models on benign samples. We thoroughly analyze the distribution mismatch between the benign and adversarial samples. This distribution mismatch and the mutual learning mechanism with the same learning ratio applied in prior art defense strategies is the main cause leading the accuracy degradation for benign samples. The proposed AdvAmd is demonstrated to steadily heal the accuracy degradation and even leads to a certain accuracy boost of common neural models on benign classification, object detection, and segmentation tasks. The efficacy of the AdvAmd is contributed by three key components: mediate samples (to reduce the influence of distribution mismatch with a fine-grained amendment), auxiliary batch norm (to solve the mutual learning mechanism and the smoother judgment surface), and AdvAmd loss (to adjust the learning ratios according to different attack vulnerabilities) through quantitative and ablation experiments. Chong Yu 0001, Tao Chen 0003, Zhongxue Gan 0001 |
IJCAI | 3 |
| 2023 | Feedback coupling induced synchronization of neural networks
Zhihao Zuo, Ruizhi Cao, Zhongxue Gan 0001, Jiawen Hou, Chun Guan, Siyang Leng |
Neurocomputing | 3 |
| 2023 | Similarity-navigated graph neural networks for node classificationabstractGraph Neural Networks are effective in learning representations of graph-structured data. Some recent works are devoted to addressing heterophily, which exists ubiquitously in real-world networks, breaking the homophily assumption that nodes belonging to the same class are more likely to be connected and restricting the generalization of traditional methods in tasks such as node classification. However, these heterophily-oriented methods still lose efficacy in some typical heterophilic datasets. Moreover, issues on leveraging the knowledge from both node features and graph structure and investigating inherent properties of the datasets still need further consideration. In this work, we first provide insights based on similarity metrics to interpret the long-existing confusion that simple models sometimes perform better than models dedicated to heterophilic networks. Then, sticking to these insights and the classification principle of narrowing the intra-class distance and enlarging the inter-class distance of the sample's embeddings, we propose a Similarity-Navigated Graph Neural Network (SNGNN) which uses Node Similarity matrix coupled with mean aggregation operation instead of the normalized adjacency matrix in the neighborhood aggregation process. Moreover, based on SNGNN, a novel explicitly aggregating mechanism for selecting similar neighbors, named SNGNN+, is devised to preserve distinguishable features and handle the heterophilic problem. Additionally, a variant, SNGNN++, is further designed to adaptively integrate the knowledge from both node features and graph structure for improvement. Extensive experiments are conducted and demonstrate that our proposed framework outperforms the state-of-the-art methods for both small-scale and large-scale graphs regardless of their heterophilic extent. Our implementation is available online. Minhao Zou, Zhongxue Gan 0001, Ruizhi Cao, Chun Guan, Siyang Leng |
Inf. Sci. | 2 |
| 2022 | Efficient Universal Shuffle Attack for Visual Object TrackingabstractRecently, adversarial attacks have been applied in visual object tracking to deceive deep trackers by injecting imperceptible perturbations into video frames. However, previous work only generates the video-specific perturbations, which restricts its application scenarios. In addition, existing attacks are difficult to implement in reality due to the real-time of tracking and the re-initialization mechanism. To address these issues, we propose an offline universal adversarial attack called Efficient Universal Shuffle Attack. It takes only one perturbation to cause the tracker malfunction on all videos. To improve the computational efficiency and attack performance, we propose a greedy gradient strategy and a triple loss to efficiently capture and attack model-specific feature representations through the gradients. Experimental results show that EUSA can significantly reduce the performance of state-of-the-art trackers on OTB2015 and VOT2018. Siao Liu, Zhaoyu Chen 0001, Wei Li 0055, Jiwei Zhu, Zhongxue Gan 0001 |
ICASSP | 7 |
| 2022 | Imitation Learning-Based Drone Motion Planning in Dense Obstacle ScenariosabstractFor the drone motion planning problem in dense obstacle scenarios, we introduce a trajectory generation method based on imitation learning that does not require the establish-ment of a local map, which greatly increases the planning speed. This method utilizes only onboard sensors and depth camera perception. We specially made the Imitation Learning Planning-Drones (ILP-Drones) dataset for training. The kinodynamic and smoothness of the generated trajectory are improved with local nonlinear optimization. The uniform B-Spline parameterization is adopted to allocate a reasonable time interval for the generated trajectory. Ultimately, our method is able to plan high quality trajectories with excellent collision avoidance ability within mil-liseconds. This is demonstrated by comparative experiments with various advanced algorithms. At the same time, the flexibility and adaptability of our method are demonstrated by ablation experiments with different number of predicted points and different simulation environments. Ziyue Hou, Longyuan Zhang, Wei Li 0055, Zhongxue Gan 0001 |
ICTAI | 5 |
| 2022 | Deep Reinforcement Learning with Parametric Episodic MemoryabstractDeep Reinforcement Learning methods are widely acknowledged to be sample inefficient, while incorporating episodic memory significantly improves it through rapidly latching onto successful experiences to guide the action of agents. Previous episodic methods, utilizing discrete memory, cannot well accommodate the continuous control tasks and have limited generalization ability to aggregate the experience across trajectories. We propose an improved episodic memory-based RL algorithm, combining the one-step method in off-policy algorithm with Parametric Episodic Memory (PEM), which leverages the discrete memory by neural networks, and thereby enhances both sample efficiency and generalization ability. Moreover, an adaptive k-nearest-neighbors is used in determining the volume of retrieved memory, further improving its efficiency. Our algorithm, evaluated on various MuJoCo continuous control tasks, outperforms the model-free baseline methods and latest episodic memory-based RL algorithms. Kangkang Chen, Zhongxue Gan 0001, Siyang Leng, Chun Guan |
IJCNN | 2 |
| 2022 | Learning From Visual Demonstrations via Replayed Task-Contrastive Model-Agnostic Meta-LearningabstractWith the increasing application of versatile robotics, the need for end-users to teach robotic tasks via visual/video demonstrations in different environments is increasing fast. One possible method is meta-learning. However, most meta-learning methods are tailored for image classification or just focus on teaching the robot what to do, resulting in a limited ability of the robot to adapt to the real world. Thus, we propose a novel yet efficient model-agnostic meta-learning framework based on task-contrastive learning to teach the robot what to do and what not to do through positive and negative demonstrations. Our approach divides the learning procedure from visual/video demonstrations into three parts. The first part distinguishes between positive and negative demonstrations via task-contrastive learning. The second part emphasizes what the positive demo is doing, and the last part predicts what the robot needs to do. Finally, we demonstrate the effectiveness of our meta-learning approach on 1) two standard public simulated benchmarks and 2) real-world placing experiments using a UR5 robot arm, significantly outperforming current related state-of-the-art methods. Ziye Hu, Wei Li 0055, Zhongxue Gan 0001, Weikun Guo, Jiwei Zhu, James Zhiqing Wen, Decheng Zhou |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2021 | Feature Adaption with Predicted Boxes for Oriented Object Detection in Aerial Images
Minhao Zou, Ziye Hu, Zhongxue Gan 0001, Chun Guan, Siyang Leng |
PRICAI (3) | 4 |
| 2021 | Collective intelligence evolution using ant colony optimization and neural networks
Xiaoya Qi, Zhongxue Gan 0001, Xiaozhi Zhang, Wei Li 0055, Chun Ouyang 0002 |
Neural Comput. Appl. | 2 |
| 2021 | Inter-Patient Classification With Encoded Peripheral Pulse Series and Multi-Task Fusion CNN: Application in Type 2 DiabetesabstractDiabetes mellitus, a chronic disease associated with elevated accumulation of glucose in the blood, is generally diagnosed through an invasive blood test such as oral glucose tolerance test (OGTT). An effective method is proposed to test type 2 diabetes using peripheral pulse waves, which can be measured fast, simply and inexpensively by a force sensor on the wrist over the radial artery. A self-designed pulse waves collection platform includes a wristband, force sensor, cuff, air tubes, and processing module. A dataset was acquired clinically for more than one year by practitioners. A group of 127 healthy candidates and 85 patients with type 2 diabetes, all between the ages of 45 and 70, underwent assessments in both OGTT and pulse data collection at wrist arteries. After preprocessing, pulse series were encoded as images using the Gramian angular field (GAF), Markov transition field (MTF), and recurrence plots (RPs). A four-layer multi-task fusion convolutional neural network (CNN) was developed for feature recognition, the network was well-trained within 30 minutes based on our server. Compared to single-task CNN, multi-task fusion CNN was proved better in classification accuracy for nine of twelve settings with empirically selected parameters. The results show that the best accuracy reached 90.6% using an RP with threshold ϵ of 6000, which is competitive to that using state-of-the-art algorithms in diabetes classification. Chun Ouyang 0002, Zhongxue Gan 0001, Junjie Zhen, Peng Zhou 0001 |
IEEE J. Biomed. Health Informatics | 2 |
| 2020 | Multi-Scale Deep Feature Fusion for Vehicle Re-IdentificationabstractVehicle re-identification (re-id) is challenging due to the small inter-class distance. The differences between similar vehicles can be extremely subtle and only captured at particular scales and semantic levels. In this paper, we propose a novel Multi-Scale Deep Feature Fusion Network (MSDeep) to conduct both multi-scale and multi-level features for precise vehicle re-id. Based on the backbone deep CNN, MS-Deep mainly consists of two modules: 1) Multi-Scale Fusion (MSF) Block which encapsulates combination of multi-scale streams as MSF feature; 2) Multi-Level Fusion (MLF) Block which fuses MSF features of multiple levels to build the final descriptor. Importantly, in MSF, Multi-Scale Attention (MSA) is introduced to dynamically emphasize important channels of each scale, and Level-Wise Attention(LWA) is utilized in MLF to determine the different weightings for each MSF feature of different levels. As a result, experiments show that our MSDeep outperforms state-of-the-art algorithms on challenging VeRi and VehicleID benchmarks in terms of abundant and hierarchical hyper-descriptors. Yiting Cheng 0001, Chuanfa Zhang, Kangzheng Gu, Lizhe Qi, Zhongxue Gan 0001 |
ICASSP | 5 |
| 2019 | Automatic Tongue Image Segmentation For Real-Time Remote DiagnosisabstractTongue diagnosis, one of the essential diagnostic methods of Traditional Chinese Medicine (TCM), is considered an ideal candidate for remote diagnosis methods because of its convenience and noninvasiveness. However, the trade-off between accuracy and efficiency and the variation of tongue images pose great challenges in real-time tongue image segmentation. To remedy these problems, in this paper, a light weight architecture based on the encoder-decoder structure is proposed. The tongue image feature extraction (TIFE) module is designed to generate features with larger receptive fields without sacrificing spatial resolution. The context module is used to increase the performance by aggregating multi-scale contextual information. The decoder is designed as a simple yet efficient feature upsampling module to fuse different depth features and refine the segmentation results along tongue boundaries. The loss module is proposed to deal with misclassifications causing by class imbalance. A new tongue image dataset (FDU/SHUTCM) is constructed for model training and testing, which contains 5,600 tongue images and their corresponding high quality masks. We demonstrate the effectiveness of the proposed model on BioHit, PolyU/HIT, and our datasets, achieving the performance of 99.15%, 95.69%, and 99.03% IoU accuracy, respectively. Segmentation of a 513×513 image takes 165 ms on CPU. Yan Wang 0068, Lizhe Qi, Fufeng Li, Zhongxue Gan 0001 |
BIBM | 7 |