EDBT 2026 Demo / reviewers in the wild / expert
Hong-Wei Ge
dblp:70/5153 · also Hongwei Ge
· DBLP profile ↗
128ranked-venue papers
23as first author
81since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 78 · 14 first-author · 52 since 2021Graphics, computer vision, multimedia, augmented reality and games · 34 · 3 first-author · 23 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 7 first-author · 6 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Computer networks · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RHVI-FDD: A Hierarchical Decoupling Framework for Low-Light Image Enhancement
Junhao Yang, Chunguo Wu, Bo Yang 0002, Hong-Wei Ge, Yanchun Liang 0001, Heow Pueh Lee |
ICMR | 4 |
| 2026 | Heuristic-based dynamic graph actor-critic algorithm for information path planning
Ruining Zhang, Chanjuan Liu 0001, Hong-Wei Ge |
Expert Syst. Appl. | 3 |
| 2026 | Enhancing neural combinatorial optimization by progressive training paradigm
Yaoxin Wu, Yaqing Hou, Hong-Wei Ge |
Neurocomputing | 4 |
| 2026 | Semantic consistency learning across temporal scales for weakly supervised video anomaly detection
Sifan Long 0001, Hong-Wei Ge, Enxuan Gu, Zidi Li, Zhaoqin Wang |
Knowl. Based Syst. | 3 |
| 2026 | PHoM: Effective pan-sharpening via higher-order state-space model
Penglian Gao, Hong-Wei Ge, Shuzhi Su |
Neural Networks | 2 |
| 2026 | PCNet: A composite backbone for 3D point cloud representation learning
Jingkun Yan, Hong-Wei Ge, Chunguo Wu, Xinye Cai, Yi-Jia Zhang 0001 |
Pattern Recognit. | 2 |
| 2026 | Proxy-Based Classification Boundary Alignment for Few-Shot Class-Incremental Learning
Hong-Wei Ge, Guozhi Tang, Jiulin Fan |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | ModSolAgent: Automated Finite Element Code Generation for Abaqus via LLM-Based AgentabstractFinite element simulation and solution process represents a critical component in engineering analysis. While large language models (LLMs) have demonstrated remarkable capabilities in general-purpose code generation from textual descriptions, their application to generating structured and specialized finite element simulation scripts presents unique challenges. A key challenge is AI-generated hallucination, as these tasks require precise intent analysis, complex planning and reasoning, and strict adherence to logical consistency during execution. This often results in outputs that seem convincing but are ultimately erroneous. To address this challenge, we propose ModSolAgent, an LLM-based agent that generates Python scripts for Abaqus software to perform finite element modeling and solving tasks. ModSolAgent ensures the accuracy and logical coherence of code generation by leveraging a structured reasoning instruction, dynamic retrieval guidance template, and iterative verification generation. Experimental results demonstrate that LLMs augmented by ModSolAgent achieve an 83.3% success rate on real-world finite element simulation tasks, effectively meeting most Abaqus scripting requirements while significantly outperforming baseline models. To further enhance accessibility, we construct the AbqInstruct dataset by distilling knowledge from ModSolAgent to fine-tune the lightweight open-source models. Experiments show that fine-tuning on AbqInstruct leads to substantial performance improvements, with the lightweight model achieving proficiency across most finite element modeling and solving tasks. This work establishes a paradigm for integrating LLMs with specialized engineering software and providing novel insights for other structured, domain-specific code generation scenarios in industrial applications. Zidi Li, Hong-Wei Ge, Guozhi Tang, Yuxuan Liu 0015 |
IEEE Trans. Ind. Informatics | 2 |
| 2026 | Tumor Contraction-Aware Multi-Sequence MRI Framework for Accurate Post-Ablation Margin Assessment in Hepatocellular CarcinomaabstractHepatocellular carcinoma (HCC) is a major cause of cancer-related mortality, and microwave ablation (MWA) is commonly used for patients ineligible for surgical resection. A critical challenge following MWA is the assessment of the ablative margin, which is complicated by non-diffeomorphic deformations introduced by thermal effects during the procedure. This paper proposes a Multi-sequence Distance-guided Complementary Network (MDCNet) that utilizes multi-sequence MRI to quantify the extent of tumor contraction after MWA. To account for the differential contraction responses of liver parenchyma and tumor tissue, we propose a novel distance-aware mask transformation strategy. This method explicitly models the spatial attenuation of MWA energy and approximates the influence of liver parenchyma's linear elastic response on tumor shrinkage, thereby enhancing the spatial adaptiveness of feature weighting. To capture the distinct structural characteristics of liver tissue emphasized by different MRI sequences and to leverage their complementary information, a gated channel fusion module is introduced to dynamically integrate features from delayed-phase and T2-weighted images. To validate the practical effectiveness of our proposed method, we evaluate the ablative margins of 115 HCC patients using a fine-tuned TransMorph model that incorporated tumor contraction predictions generated by MDCNet, and compare the results with radiologist 2D assessments. The registration method enhanced with MDCNet improved tumor deformation accuracy and achieved a higher Youden Index in detecting incomplete ablations. Moreover, MDCNet provides interpretable predictions, thereby facilitating clinical decision support. Li-Nan Dong, Hong-Wei Ge, Jinming Hu, Shichen Yu |
IEEE J. Biomed. Health Informatics | 2 |
| 2025 | A Biogeography-based Dual-Strategy Particle Swarm Algorithm for Numerical and Engineering Design OptimizationabstractRecently, researchers developed numerous nature-inspired meta-heuristics algorithms, such as particle swarm optimization (PSO), to solve numerical and engineering optimization problems. However, the PSO still suffers from complex optimization problems, such as slow convergence speed and getting stuck in local optima. This paper proposes a Biogeography-based Dual-strategy PSO algorithm (DBPSO). Firstly, a Dual-Strategy Search (DSS) is proposed, which utilizes an improved random learning strategy for global search and a modified quasi-Newton search strategy for local search. The proportion of resources allocated to the modified quasi-Newton search strategy is associated with the inertia weight value of the improved random learning, achieving a balance between exploration and exploitation. Secondly, an improved migration operator of biogeography-based optimization is embedded in DSS. Finally, an iteration-based hybridization strategy that allows one-way information transmission is proposed to fuse the two algorithms effectively. Experiments are conducted on CEC2013 and CEC2017 test suites, and the results show that DBPSO ranks first by comparing it with multiple PSO variants and other meta-heuristics variants. In addition, experiments are conducted on two real-world engineering optimization problems to demonstrate the applicability of DBPSO in solving practical optimization problems, such as the car side impact design problem. Hong-Wei Ge, Yaqing Hou, Mengyue Wang |
CEC | 2 |
| 2025 | EFormer: An Effective Edge-based Transformer for Vehicle Routing ProblemsabstractRecent neural heuristics for the Vehicle Routing Problem (VRP) primarily rely on node coordinates as input, which may be less effective in practical scenarios where real cost metrics—such as edge-based distances—are more relevant. To address this limitation, we introduce EFormer, an Edge-based Transformer model that uses edge as the sole input for VRPs. Our approach employs a precoder module with a mixed-score attention mechanism to convert edge information into temporary node embeddings. We also present a parallel encoding strategy characterized by a graph encoder and a node encoder, each responsible for processing graph and node embeddings in distinct feature spaces, respectively. This design yields a more comprehensive representation of the global relationships among edges. In the decoding phase, parallel context embedding and multi-query integration are used to compute separate attention mechanisms over the two encoded embeddings, facilitating efficient path construction. We train EFormer using reinforcement learning in an autoregressive manner. Extensive experiments on the Traveling Salesman Problem (TSP) and Capacitated Vehicle Routing Problem (CVRP) reveal that EFormer outperforms established baselines on synthetic datasets, including large-scale and diverse distributions. Moreover, EFormer demonstrates strong generalization on real-world instances from TSPLib and CVRPLib. These findings confirm the effectiveness of EFormer’s core design in solving VRPs. Dian Meng, Zhiguang Cao, Yaoxin Wu, Yaqing Hou, Hong-Wei Ge, Qiang Zhang 0008 |
IJCAI | 5 |
| 2025 | Towards Efficient Few-shot Graph Neural Architecture Search via Partitioning Gradient ContributionabstractTo address the weight coupling problem, certain studies introduced few-shot Neural Architecture Search (NAS) methods, which partition the supernet into multiple sub-supernets. However, these methods often suffer from computational inefficiency and tend to provide suboptimal partitioning schemes. To address this problem more effectively, we analyze the weight coupling problem from a novel perspective, which primarily stems from distinct modules in succeeding layers imposing conflicting gradient directions on the preceding layer modules. Based on this perspective, we propose the Gradient Contribution (GC) method that efficiently computes the cosine similarity of gradient directions among modules by decomposing the Vector-Jacobian Product during supernet backpropagation. Subsequently, the modules with conflicting gradient directions are allocated to distinct sub-supernets while similar ones are grouped together. To assess the advantages of GC and address the limitations of existing Graph Neural Architecture Search methods, which are limited to searching a single type of Graph Neural Networks (Message Passing Neural Networks (MPNNs) or Graph Transformers (GTs)), we propose the Unified Graph Neural Architecture Search (UGAS) framework, which explores optimal combinations of MPNNs and GTs. The experimental results demonstrate that GC achieves state-of-the-art (SOTA) performance in supernet partitioning quality and time efficiency. In addition, the architectures searched by UGAS+GC outperform both the manually designed GNNs and those obtained by existing NAS methods. Finally, ablation studies further demonstrate the effectiveness of all proposed methods. Xuan Wu 0004, Bo Yang 0002, You Zhou 0008, Yubin Xiao, Yanchun Liang 0001, Hong-Wei Ge, Heow Pueh Lee, Chunguo Wu |
KDD (2) | 7 |
| 2025 | Troublemaker Learning for Low-Light Image EnhancementabstractLow-light image enhancement (LLIE) aims at restoring the color and brightness of underexposed images. Supervised methods suffer from high costs in collecting paired low-normal light images, while unsupervised approaches require intricate loss functions. To tackle these dual challenges, we propose the Trouble-Maker Learning (TML) strategy, which leverages images with normal light as training inputs. TML comprises two core components. Firstly, the Troublemaker Model (TM) generates pseudo low-light images from normal images, thereby alleviating the need for pairwise data and reducing associated costs. Secondly, the Predicting Model (PM) enhances the brightness of pseudo low-light images. Additionally, we integrate an Enhancing Model (EM) to further refine the visual quality of the PM's outputs. In LLIE tasks, it is crucial to capture global element correlations, as this allows for the extraction of more information pertaining to the same object. Convolutional Neural Networks (CNNs) and self-attention mechanisms are not well-suited to this task due to the local CNN operators, and high time complexity, respectively. To address these limitations, we propose Global Dynamic Convolution (GDC) with a time complexity of O(n). Essentially, GDC mimics the partial calculation process of self-attention to establish element-wise correlations. Building upon the GDC module, we develop the UGDC model. Finally, we explore the application of Data Fusion in the field of LLIE. Based on the Retinex theory, we conducted feature-level fusion using low-light images, illumination components and reflection components, which further enhance the performance of the LLIE system. Extensive quantitative and qualitative experiments demonstrate that UGDC, trained with TML and via data fusion, can achieve performance competitive with state-of-the-art approaches on public datasets. The source code of this paper is publicly available at https://github.com/Rainbowman0/TML_LLIE, facilitating reproducibility of the research findings. Yinghao Song, Bo Yang 0002, Yanchun Liang 0001, Hong-Wei Ge, Heow Pueh Lee, Chunguo Wu |
ICMR | 5 |
| 2025 | Dialogue-Driven Interactive Dynamic Learning for Text-to-Image Person RetrievalabstractText-to-image person retrieval aims to identify target person images using natural language descriptions. Current state-of-the-art methods predominantly rely on single-round retrieval frameworks, where retrieval accuracy heavily depends on the quality of the initial textual descriptions. However, users sometimes struggle to provide detailed and distinctive descriptions in a single attempt, resulting in generic initial queries that lack discriminative details. This fundamental limitation of the single-round retrieval framework frequently leads to the misinterpretation of user intent and suboptimal retrieval performance. To address this limitation, we propose Dialogue-driven Interactive Dynamic Learning (DIDL) for text-to-image person retrieval. Specifically, we first introduce Collaborative Query Refinement (CQR), which progressively refines retrieval conditions through multi-round dialogues. Then, we design Dynamic Context Resampling (DCR) based on a bi-granular mask strategy that enhances the model's adaptation to dialogue-style contexts and effectively balances its attention between initial descriptions and supplementary information. Based on these components, we further propose cross-modal Probabilistic Context Matching Modeling (ProCMM) that establishes effective associations between static visual features and dynamic contextual semantics. Extensive experiments demonstrate that our approach achieves state-of-the-art performance across all three benchmark datasets. Hong-Wei Ge, Yuxuan Liu 0015, Yaqing Hou |
ACM Multimedia | 2 |
| 2025 | A prior knowledge-supervised fusion network predicts survival after radiotherapy in patients with advanced gastric cancerabstractBACKGROUND AND OBJECTIVE: Predicting overall survival (OS) for advanced gastric cancer patients after radiotherapy is critical for developing an individualized treatment plan. However, existing studies have focused on gastric cancer CT images with a large amount of redundant information, neglecting the role of physicians' prior knowledge in guiding gastric cancer CT image information. We propose a multimodal fusion method based on prior knowledge to predict OS after radiotherapy in advanced gastric cancer patients to assist physicians in clinical diagnosis and treatment. METHODS: A prior knowledge supervised fusion network (PKSFnet) is proposed. Firstly, PKSFnet uses a novel sampling strategy, which enables the input model data to obtain a complete feature space by analyzing the entire patient data space. Afterwards, under the guidance of the multi-domain feature fusion module (MdFF), multimodal information of patients is adaptively fused and mined to improve the prediction performance. RESULTS: The results of the proposed model are superior to those of other unimodal and multimodal state-of-the-art methods. For the segmented survival time classification task, the AUC, specificity, sensitivity, precision of the proposed model are 0.8397, 0.875, 0.7556, and 0.875, respectively. For the survival risk regression task, the C-index and HR of the proposed model are 0.8574 and 4.658 respectively. Ablation experimental results further demonstrate the impact of each module of the proposed model. Finally, we apply the novel sampling strategy to other deep learning models and achieve significant improvement. CONCLUSION: The experimental results have demonstrated that the proposed model can effectively predict OS after radiotherapy in patients with advanced gastric cancer, which demonstrate that the proposed model can facilitate the development and application of robust clinical treatment strategies. Liang Sun 0003, Yongxin Lan, Pengfei Ji, Hong-Wei Ge, Ming Cui |
Artif. Intell. Medicine | 5 |
| 2025 | An evolutionary multitasking algorithm for multi-objective feature selection using dual-perspective reduction
Mengyue Wang, Hong-Wei Ge, Liang Sun 0003, Yaqing Hou |
Eng. Appl. Artif. Intell. | 2 |
| 2025 | Hypergraph-driven soft semantics flexible learning for visible-infrared person re-identification
Hong-Wei Ge, Yuxuan Liu 0015, Chunguo Wu, Jiulin Fan |
Eng. Appl. Artif. Intell. | 2 |
| 2025 | A secure medical image encryption algorithm for IoMT using a Quadratic-Sine chaotic map and pseudo-parallel confusion-diffusion mechanism
Ali Mansouri 0003, Pin Sun, Chengzhi Lv, Yinghua Zhu, Hong-Wei Ge, Changkai Sun |
Expert Syst. Appl. | 6 |
| 2025 | Dual-prompt complementary fusion network for RGBT trackingabstractRGBT target tracking is a significant downstream task in the field of object tracking. However, compared to visible light target tracking, RGBT target tracking faces the challenge of smaller datasets, making it difficult to achieve performance levels comparable to those achieved in visible light target tracking. To address how to effectively combine the complementary characteristics of visible and thermal modalities, as well as how to fully leverage the superior performance of models trained on visible light target tracking tasks, while also aiming for lower computational costs and higher tracking effectiveness, a dual-prompt complementary fusion strategy for an RGBT tracking network is proposed. Drawing on the concept of prompt learning, this network aims to extend the efficient performance of visible light target tracking to the RGBT target tracking domain. In its implementation, the prompt module inputs both visible and thermal modality information as dual prompts into the backbone network, where the network utilizes these prompts to generate new, enriched prompt information at each layer. Subsequently, an information enhancement fusion module enhances the acquired prompt information and refeeds it into the backbone network, aiming to improve the tracking accuracy and robustness. Experimental results on GTOT, RGBT234 and LasHeR datasets show that the tracking accuracy (PR) and success rate (SR) of the network reach 93.1%/76.8%, 84.4%/62.4% and 66.8%/53.8%, respectively, which is improved compared with the current mainstream RGBT target tracking network, which verifies the effectiveness of the network. Xihui Wu, Hong-Wei Ge, Shuzhi Su |
Intell. Data Anal. | 2 |
| 2025 | Combining decomposition and graph capsule network for multi-objective vehicle routing optimizationabstractIn order to alleviate urban congestion, improve vehicle mobility, and improve logistics delivery efficiency, this paper establishes a practical multi-objective and multi constraint logistics delivery mathematical model based on graphs, and proposes a solution algorithm framework that combines decomposition strategy and deep reinforcement learning (DRL). Firstly, taking into account the actual multiple constraints such as customer distribution, vehicle load constraints, and time windows in urban logistics distribution regions, a multi constraint and multi-objective urban logistics distribution mathematical model was established with the goal of minimizing the total length, cost, and maximum makespan of urban logistics distribution paths. Secondly, based on the decomposition strategy, a DRL framework for optimizing urban logistics delivery paths based on Graph Capsule Network (G-Caps Net) was designed. This framework takes the node information of VRP as input in the form of a 2D graph, modifies the graph attention capsule network by considering multi-layer features, edge information, and residual connections between layers in the graph structure, and replaces probability calculation with the module length of the capsule vector as output. Then, the baseline REINFORCE algorithm with rollout is used for network training, and a 2-opt local search strategy and sampling search strategy are used to improve the quality of the solution. Finally, the performance of the proposed method was evaluated on standard examples of problems of different scales. The experimental results showed that the constructed model and solution framework can improve logistics delivery efficiency. This method achieved the best comprehensive performance, surpassing the most advanced distress methods, and has great potential in practical engineering. Haifei Zhang, Hong-Wei Ge, Lujie Zhou, Shuzhi Su, Yubing Tong |
Intell. Data Anal. | 2 |
| 2025 | Spiking Trans-YOLO: A range-adaptive energy-efficient bridge between YOLO and Transformer
Yushi Huo, Hong-Wei Ge, Guozhi Tang, Shengxuan Gao |
Neurocomputing | 2 |
| 2025 | Real-time encryption of medical images in IoMT using a novel biometric and chaotic approach
Ali Mansouri 0003, Pin Sun, Chengzhi Lv, Yinghua Zhu, Hong-Wei Ge, Changkai Sun |
Knowl. Based Syst. | 6 |
| 2025 | MARIC: an efficient multi-agent real-time intention-based communication model for team cooperation
Hong-Wei Ge, Zhangang Hao, Yaqing Hou |
Neural Comput. Appl. | 2 |
| 2025 | Multi-view clustering via diversity induction and multi-layer concept factorization
Hong-Wei Ge, Shuzhi Su, Penglian Gao |
Pattern Anal. Appl. | 2 |
| 2025 | Learning to Evolve With Guiding Solutions Generated by Generative Adversarial NetworkabstractMany search strategies have been designed to generate a promising offspring population for efficiently solving large-scale multiobjective optimization problems (LSMOPs). The effectiveness of existing search strategies relies on the quality of good parent solutions. However, especially in early generations, the current population does not always include high-quality solutions. This article proposes a generative adversarial network (GAN)-guided search (G2S) strategy for learning to evolve with guiding solutions. Its main idea is to employ GAN for mapping a set of guiding points in the objective space with good convergence and diversity back to the decision space to guide evolution. Specifically, the current population is used as real data, and the guiding points consisting of nondominated solutions and reference vectors are used as virtual data. The trained GAN generates guiding solutions in the decision space to guide the population to evolve efficiently. A large-scale multiobjective evolutionary framework using G2S is also proposed which can be embedded into multiobjective evolutionary algorithms (MOEAs) to improve their ability to handle LSMOPs. Experimental studies on several benchmark problems with the highest-5000-D decision space show that the proposed G2S is competitive compared with the state-of-the-art algorithms and has impressive efficiency as the component to improve the performance of MOEAs for solving LSMOPs. Hong-Wei Ge, Yaqing Hou, Hisao Ishibuchi |
IEEE Trans. Evol. Comput. | 1 |
| 2025 | Multiscale Recovery Diffusion Model With Unsupervised Learning for Video Anomaly Detection SystemabstractThe rapid development of intelligent industry and smart city increases the number of surveillance devices, greatly enhancing the need for unsupervised automatic anomaly detection in real-time video surveillance, which uses raw data without laborious manual annotations. Existing video anomaly detection (VAD) methods encounter limitations when utilizing pretext tasks, such as reconstruction or prediction to identify abnormal events, as these tasks are not completely consistent and complementary with the essential objective of anomaly detection. Motivated by recent advances in diffusion models, we propose a multiscale recovery diffusion model, which relies on the proposed novel and effective pretext task named recovery to introduce the notion of generation speed. It utilizes critical step-by-step generation of diffusion probabilistic models in unsupervised anomaly detection scenarios. By incorporating a proposed multiscale spatial-temporal subtraction module, our model captures more detailed appearance and motion information of foreground objects without relying on other high-level pretrained models. Furthermore, an innovative push–pull loss further extends the disparity between normal and abnormal events through pseudolabels. We validate our model on five established benchmarks: UCSD Ped1, UCSD Ped2, CUHK Avenue, ShanghaiTech, and UCF-Crime, achieving frame-level area under the curves of 86.01%, 99.23%, 92.35%, 82.49%, and 74.79%, respectively, surpassing other state-of-the-art unsupervised VAD methods. Hong-Wei Ge, Yuxuan Liu 0015, Guozhi Tang |
IEEE Trans. Ind. Informatics | 2 |
| 2025 | Multi-Memory Streams: A Paradigm for Online Video Super-Resolution in Complex Exposure ScenesabstractExisting online video super-resolution methods utilize implicit memories of previous frames to provide reference information, which have a single memory stream path and are highly dependent on the continuous memory stream. However, video capture in real-world scenes is typically affected by abnormal exposures resulting in sudden changes of lightness thus interrupting the memory stream, while long-term memories suffer from memory vanishing problems during transmission. To address this problem, we propose a novel multi-memory streams based online video super-resolution paradigm that adaptively corrects for abnormal exposures and creates multi-memory streams to accurately converge long-term memories. Specifically, we first propose an exposure detection-correction module, which utilizes optical flow overfitting property and temporal lightness information to detect and correct abnormal exposures to avoid interruption of memory streams. In addition, we propose a dynamic-static decoupled alignment strategy, which can adaptively select the alignment method based on pixel displacement, thus accurately aggregating past long-term memories to create multiple memory streams. Further, we propose an adaptive memory fusion module to mine complementary information between multiple memory streams to solve the memory vanishing problem. Extensive experimental results show that our method outperforms existing video super-resolution methods on complex exposure datasets. We also conduct detailed ablation experiments to analyze and validate our contributions. Guozhi Tang, Hong-Wei Ge, Chunguo Wu |
IEEE Trans. Multim. | 2 |
| 2025 | P$^{2}$M: Progressive Perspective Mining for Referring Video Object Segmentation
Yihan Wang 0011, Baoli Sun, Xinzhu Ma, Hong-Wei Ge, Jiulin Fan |
IEEE Trans. Multim. | 4 |
| 2025 | Find Hidden Modality Divergence: Adversarial Aware Learning for Unsupervised Visible-Infrared Person Re-IdentificationabstractUnsupervised visible-infrared person re-identifi-cation (Unsupervised VI-ReID) aims to learn discriminative identity features under the large modality gap without any labeled data. Currently, the state-of-the-art methods optimize cross-modality differences by using contrastive learning as the underlying paradigm. However, they neglect the problem of modality divergence during the cross-modality optimization process. This problem means that the interclass instances between the cross-modality intraclass gaps can make cross-modality intraclass instances difficult to get closer to each other in the feature space due to the effect of contrastive learning on these interclass instances. To alleviate the negative impact of the modality divergence problem, we propose an adversarial aware learning (ADAL) framework to explore the instances that generate modal divergence and adversarially optimize these explored instances. Specifically, on the one hand, we explore the optimization directions of each cluster during the cross-modality optimization process, and the cluster centroids generating positive optimization are facilitated, while the others generating negative optimization are penalized. On the other hand, we further consider the instance-level optimization process, which increases the affinities of the positive instance pairs with large cross-modality gaps to further improve the centroid-level optimization. Extensive experiments conducted on the visible-infrared person Re-ID datasets show that the proposed method is used as a universally applicable plug-in module to add the existing unsupervised VI-ReID methods, which outperforms the existing state-of-the-art approaches. Yuxuan Liu 0015, Hong-Wei Ge, Chunguo Wu |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2025 | A Group-Based Many-Task Collaborative Optimization Framework for Evolutionary Robots DesignabstractIn evolutionary robotics (ER), the evolution of a robot’s morphology (i.e., physical structure) or controller (i.e., control algorithm or instruction sequence) often entails tackling an extensive number of tasks. The use of evolutionary multitasking (EMT) in ER, which optimizes multiple tasks simultaneously by reusing potentially useful knowledge across diverse tasks, could improve the performance of problem-solving to each task. However, existing EMT methods do not fully use intertask correlations, limiting knowledge sharing. In view of this, this study introduces a novel framework, termed adaptive group-based collaborative optimization, tailored for handling optimization problems involving a large number of tasks within the ER domain simultaneously. The proposed framework divides tasks into groups according to their similarity and then proceeds through two principal stages, namely, intergroup knowledge separation and intragroup knowledge reunion. During intergroup knowledge separation stage, an adaptive method for selecting crossover operators enables source tasks to share useful knowledge to the target task across groups. During intragroup knowledge reunion stage, an adaptive knowledge combination strategy facilitates the target task in assimilating knowledge from multiple sources intragroup. We validated the efficacy of the proposed framework in both planar manipulators and hexapod robot experiments. The results indicate that our method outperforms existing state-of-the-art algorithms (i.e., MME, MMKT) on several metrics (e.g., mean fitness and quality diversity metrics). The proposed method can effectively improve the effectiveness and diversity of solutions in solving ER problems with a large number of tasks (e.g., 5 000 or 10 000), and has broad potential in practical ER applications. Yaqing Hou, Zhaoping Yu, Wenbin Pei, Yaoxin Wu, Hong-Wei Ge, Bing Xue 0001, Mengjie Zhang 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 6 |
| 2025 | Ought to be salient and hidden: soft biological semantic-guided explicit-implicit learning for cloth-changing person re-identification
Wanda Zeng, Hong-Wei Ge, Yuxuan Liu 0015 |
Vis. Comput. | 2 |
| 2024 | CoDe: An Explicit Content Decoupling Framework for Image RestorationabstractThe performance of image restoration (IR) is highly de-pendent on the reconstruction quality of diverse contents with varying complexity. However, most IR approaches model the mapping between various complexity contents of inputs and outputs through the repeated feature calcu-lation propagation mechanism in a unified pipeline, which leads to unsatisfactory results. To address this issue, we propose an explicit Content Decoupling framework for IR, dubbed CoDe, to end-to-end model the restoration process by utilizing decoupled content components in a divide-and- conquer-like architecture. Specifically, a Content Decou-pling Module is first designed to decouple content components of inputs and outputs according to thefrequency spec-tra adaptively generatedfrom the transform domain. In ad-dition, in order to harness the divide-and-conquer strategy for reconstructing decoupled content components, we pro-pose an IR Network Container. It contains an optimized version, which is a streamlining of an arbitrary IR net-work, comprising the cascaded modulated subnets and a Reconstruction Layers Pool. Finally, a Content Consistency Loss is designed from the transform domain perspective to supervise the restoration process of each content component and further guide the feature fusion process. Exten-sive experiments on several IR tasks, such as image super-resolution, image denoising, and image blurring, covering both real and synthetic settings, demonstrate that the pro-posed paradigm can effectively take the performance of the original network to a new state-of-the-art level in multiple benchmark datasets (e.g., O.34dB@Set5 × 4 over DAT). Enxuan Gu, Hong-Wei Ge |
CVPR | 2 |
| 2024 | SAEIR: Sequentially Accumulated Entropy Intrinsic Reward for Cooperative Multi-Agent Reinforcement Learning with Sparse Reward
Hong-Wei Ge, Yaqing Hou |
IJCAI | 2 |
| 2024 | Self-Attention Guided Advice Distillation in Multi-Agent Deep Reinforcement LearningabstractAdvising is an effective method to enhance agent learning performance in multi-agent deep reinforcement learning. Existing advising methods typically rely on a teacher-student framework where a teacher agent provides student agents with action or Q-value advice. However, they share a common limitation: the advice from a teacher agent can only assist a student in making a one-time decision in the current state and cannot be internalized into the student agent’s knowledge to intrinsically change the student agent’s decision model. Consequently, the advice acts more like a one-time instruction from the teacher rather than a learning aid. If the student agent encounters the same problem again, it may still be unable to make a sound decision and need to request advice. This not only fails to rapidly enhance the agent’s decision-making ability fundamentally but also leads to a considerable waste of communication costs. Hence, we propose a multi-agent advice distillation framework through attention that allows the student agent to request advice from the experienced teacher and distill that advice into their own decision model via the self-attention mechanism. As a result, advice is fully utilized, allowing for a rapid and intrinsic improvement in the agent’s decision-making capabilities. Our empirical evaluations demonstrate that, compared to existing advising methods, our method significantly improves learning performance while reducing the communication cost. Sihan Zhou, Yaqing Hou, Liran Zhou, Hong-Wei Ge, Liang Feng 0001 |
IJCNN | 5 |
| 2024 | Robust multi-view clustering via collaborative constraints and multi-layer concept factorization
Hong-Wei Ge, Shuzhi Su, Penglian Gao |
Appl. Intell. | 2 |
| 2024 | Stochastic online decisioning hyper-heuristic for high dimensional optimization
Hong-Wei Ge, Mingde Zhao 0002, Yaqing Hou |
Appl. Intell. | 2 |
| 2024 | Three-stage multi-modal multi-objective differential evolution algorithm for vehicle routing problem with time windowsabstractIn this paper, the mathematical model of Vehicle Routing Problem with Time Windows (VRPTW) is established based on the directed graph, and a 3-stage multi-modal multi-objective differential evolution algorithm (3S-MMDEA) is proposed. In the first stage, in order to expand the range of individuals to be selected, a generalized opposition-based learning (GOBL) strategy is used to generate a reverse population. In the second stage, a search strategy of reachable distribution area is proposed, which divides the population with the selected individual as the center point to improve the convergence of the solution set. In the third stage, an improved individual variation strategy is proposed to legalize the mutant individuals, so that the individual after variation still falls within the range of the population, further improving the diversity of individuals to ensure the diversity of the solution set. Based on the synergy of the above three stages of strategies, the diversity of individuals is ensured, so as to improve the diversity of solution sets, and multiple equivalent optimal paths are obtained to meet the planning needs of different decision-makers. Finally, the performance of the proposed method is evaluated on the standard benchmark datasets of the problem. The experimental results show that the proposed 3S-MMDEA can improve the efficiency of logistics distribution and obtain multiple equivalent optimal paths. The method achieves good performance, superior to the most advanced VRPTW solution methods, and has great potential in practical projects. Haifei Zhang, Hong-Wei Ge, Shuzhi Su, Yubing Tong |
Intell. Data Anal. | 2 |
| 2024 | Occluded Person Reidentification via a Universal Framework With Difference Consistency Guidance LearningabstractOccluded person reidentification (Re-ID) aims at learning discriminative identity features to match person images under the interference of occlusion situations in video surveillance of the visual Internet of Things (VIoT). Currently, occluded person Re-ID methods have made impressive improvements in nonperson occlusion situations. However, in real-world scenarios, the target person is commonly occluded by other nontarget persons, and the fine-grained differences between the persons make the model difficult to distinguish the discriminative identity features. To this end, we propose a difference consistency guidance (DCG) learning to enlarge the fine-grained differences by the guidance of the coarse-grained differences, which can distinguish the discriminative identity features in nontarget person occlusion situations. Then, DCG reduces the identity feature representation of the nonperson occlusion instances, which further improves the ability of the model in nonperson occlusion situations and improves the guidance ability of coarse-grained difference. Moreover, DCG can enhance the robustness of the model in the unsupervised occluded person Re-ID task and further improve the universal applicability of the model. Extensive experiment results under the supervised and unsupervised settings demonstrate the DCG outperforms the state-of-the-art methods in experiments conducted on the occluded person Re-ID benchmarks. Yuxuan Liu 0015, Hong-Wei Ge, Guozhi Tang |
IEEE Internet Things J. | 2 |
| 2024 | Progressive reconstruction-decoupled face super-resolution framework with controllable knowledge guidance
Guozhi Tang, Hong-Wei Ge, Enxuan Gu, Yaqing Hou, Mingde Zhao 0002 |
Knowl. Based Syst. | 2 |
| 2024 | Localization and saturation of degradation space for weakly-supervised real-world super-resolution
Guozhi Tang, Hong-Wei Ge, Yuxuan Liu 0015, Chunguo Wu |
Knowl. Based Syst. | 2 |
| 2024 | Multi-view clustering via dual-norm and HSIC
Hong-Wei Ge, Shuzhi Su, Shuangxi Wang |
Multim. Tools Appl. | 2 |
| 2024 | Temporal prediction model with context-aware data augmentation for robust visual reinforcement learning
Xinkai Yue, Hong-Wei Ge, Yaqing Hou |
Neural Comput. Appl. | 2 |
| 2024 | Representation Robustness and Feature Expansion for Exemplar-Free Class-Incremental LearningabstractDespite deep neural networks have made outstanding achievements in many static tasks, when faced with a continuous stream of data, they suffer from catastrophic forgetting since the previous data is usually inaccessible. Stored data or generative model is commonly used for maintaining the model performance but with memory utilization and privacy safety issues. Prototype-based methods address these issues by keeping only one prototype for each class but with limitations in its ability to trade-off the model stability and plasticity. In this paper, a novel exemplar-free class-incremental learning method is proposed which improves the stability of the representation learning and the decision boundary to a great degree. First, based on the results of our exploration into the impact of the batch normalization (BN) layer on representation learning, we propose to remove the BN layer (RBNL) in the incremental training phase to improve the stability of model representation learning. Then, to further maintain the feature space, we design the prototype mixing (PM), which expands the deep features by randomly and linearly combining prototypes of the old classes to generate hybrid prototypes with composite labels for fine-tuning the fully connected layer. Experimental results on three benchmark datasets, CIFAR-100, TinyImageNet, and ImageNet, show that our proposed method can effectively balance the stability and plasticity of the model, and outperforms the state-of-the-art works. Hong-Wei Ge, Yuxuan Liu 0015, Chunguo Wu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | A Multiagent Cooperative Learning System With Evolution of Social RolesabstractRecent developments in reinforcement learning (RL) have been able to derive optimal policies for sophisticated and capable agents, and shown to achieve human-level performance on a number of challenging tasks. Unfortunately, when it comes to multiagent systems (MASs), complexities, such as nonstationarity and partial observability bring new challenges to the field. Building a flexible and efficient multiagent RL (MARL) algorithm capable of handling complex tasks has to date remained an open challenge. This article presents a multiagent learning system with the evolution of social roles (eSRMA). The main interest is placed on solving the key issues in the definition and evolution of suitable roles, and optimizing the policies accompanied by social roles in MAS efficiently. Specifically, eSRMA incorporates and cultivates role division awareness of agents to improve the ability to deal with complex cooperative tasks. Each agent is assigned a role module, which can dynamically generate roles based on the individuals’ local observations. A novel MARL algorithm is designed as the principal driving force that governs the role-policy learning process by a role-attention credit assignment mechanism. Moreover, a role evolution process is developed to help agents dynamically choose appropriate roles in decision making. Comprehensive experiments on the StarCraft II micromanagement benchmarkhave demonstrated that eSRMA exhibits superiority in achieving higher learning capability and efficiency for multiple agents compared to the state-of-the-art MARL methods. Yaqing Hou, Yifeng Zeng, Yew-Soon Ong, Yaochu Jin, Hong-Wei Ge, Qiang Zhang 0008 |
IEEE Trans. Evol. Comput. | 6 |
| 2024 | Neural Architecture Search for Text Classification With Limited Computing Resources Using Efficient Cartesian Genetic ProgrammingabstractCartesian Genetic Programming (CGP) has often been applied for Neural Architecture Search (NAS). However, the performance of CGP is less than ideal when searching for architectures with limited computing resources. To better facilitate NAS with limited computing resources, this paper proposes a crossover operator, a light-weighted age mechanism, and two adaptive mutation operators as the novel components in our Efficient Cartesian Genetic Programming (ECGP) method. To assess the performance of ECGP, we conduct extensive experiments on three text classification task datasets. The experimental results demonstrate that ECGP outperforms other NAS methods, requiring only hundreds of fitness evaluations to find architectures with competitive accuracy compared with human-designed models. Additionally, the ECGP-evolved architectures are shown as converging fast and stably, and having high-level transferability with merely a 1-2% accuracy drop. Ablation studies demonstrate the effectiveness of the proposed operators and age mechanism, and identify GRU as the most critical function in the text classification task. Finally, we summarize three design principles observed from the ECGP-evolved architectures that are in line with human-design strategies. To the best of our knowledge, this work introduces the first attention-derived NAS benchmark for the text classification task. Xuan Wu 0004, Di Wang 0004, Huanhuan Chen 0001, Lele Yan, Yubin Xiao, Chunyan Miao, Hong-Wei Ge, Dong Xu 0002, Yanchun Liang 0001, Kangping Wang, Chunguo Wu, You Zhou 0008 |
IEEE Trans. Evol. Comput. | 7 |
| 2024 | A Virtual-Sensor Construction Network Based on Physical Imaging for Image Super-ResolutionabstractImage imaging in the real world is based on physical imaging mechanisms. Existing super-resolution methods mainly focus on designing complex network structures to extract and fuse image features more effectively, but ignore the guiding role of physical imaging mechanisms for model design, and cannot mine features from a physical perspective. Inspired by the mechanism of physical imaging, we propose a novel network architecture called Virtual-Sensor Construction network (VSCNet) to simulate the sensor array inside the camera. Specifically, VSCNet first generates different splitting directions to distribute photons to construct virtual sensors, and then performs a multi-stage adaptive fine-tuning operation to fine-tune the number of photons on the virtual sensors to increase the photosensitive area and eliminate photon cross-talk, and finally converts the obtained photon distributions into RGB images. These operations can naturally be regarded as the virtual expansion of the camera's sensor array in the feature space, which makes our VSCNet bridge the physical space and feature space, and uses their complementarity to mine more effective features to improve performance. Extensive experiments on various datasets show that the proposed VSCNet achieves state-of-the-art performance with fewer parameters. Moreover, we perform experiments to validate the connection between the proposed VSCNet and the physical imaging mechanism. The implementation code is available at https://github.com/GZ-T/VSCNet. Guozhi Tang, Hong-Wei Ge, Liang Sun 0003, Yaqing Hou, Mingde Zhao 0002 |
IEEE Trans. Image Process. | 2 |
| 2024 | Discriminative Identity-Feature Exploring and Differential Aware Learning for Unsupervised Person Re-IdentificationabstractUnsupervised person re-identification (Re-ID) aims to learn discriminative representations for person retrieval from unlabeled data. Currently, state-of-the-art techniques accomplish this task by using instance contrastive learning, which contrasts the similarities of the instances in different views. However, existing contrastive methods only focus on the positive effects of inter-instance relationships, while neglecting the negative effects of intra-instance redundancy information. This redundancy information can generate invalid or spurious intra-class relationships during the instance contrasting process, which enlarges the intra-class gaps and increases the noisy pseudo-labels. To address this issue, we propose a discriminative identity-feature exploring and differential aware learning (DiDAL) framework to learn more discriminative intra-identity representations. Specifically, the DiDAL extracts intra-instance salient features by synthetic complementary attention, and further explores the discriminative identity features by modeling the relationship among these salient features based on graph neural networks. This strategy aims to reduce the intra-instance redundancy information. Moreover, DiDAL explores hard instances by leveraging the extracted intra-instance salient features, and matches an anchor with multiple hard positive instances to enhance the robustness of the model to noisy pseudo-labels. Extensive experiment results on two widely used person re-identification datasets and a vehicle re-identification dataset demonstrate the superiority of the proposed method compared with existing state-of-the-art methods. Yuxuan Liu 0015, Hong-Wei Ge, Zhen Wang 0004, Yaqing Hou, Mingde Zhao 0002 |
IEEE Trans. Multim. | 2 |
| 2024 | Clothes-Changing Person Re-Identification via Universal Framework With Association and Forgetting LearningabstractClothes-changing person re-identification (Re-ID) aims at learning identity-relevant feature representations among clothing-changed persons. Currently, the state-of-the-art methods accomplish this task by using additional assistance (e.g., silhouettes, sketches, clothes labels, etc.) to explore identity-relevant information. However, humans do not require redundant assistance information to retrieve clothing-changed persons. It is commonly known that humans can recall targets they have seen before with a simple reminder. Inspired by human perception, we propose an association and forgetting learning (AFL) framework for clothes-changing person re-identification. Specifically, on the one hand, during the association learning process, the AFL framework constructs association factors for each identity to simulate the reminders found in human perception. Then, the original instances and the explored hardest positive instances are cross-correlated by the association factors to learn identity-relevant features. On the other hand, the model is forced to forget the identity-irrelevant features by the proposed forgetting learning module, which improves the intra-class compactness. Finally, we further propose a clustering relationship exploration (CRE) module to optimize the cluster distribution of clothes-changing instances, which enables AFL to also be effectively applied in unsupervised settings, improving the universal applicability of the model. Extensive experiment results obtained on clothes-changing person Re-ID datasets under supervised and unsupervised settings demonstrate the superiority of the proposed method over the existing state-of-the-art methods. Yuxuan Liu 0015, Hong-Wei Ge, Zhen Wang 0004, Yaqing Hou, Mingde Zhao 0002 |
IEEE Trans. Multim. | 2 |
| 2023 | Surrogate-Assisted Morphology Optimization by Genetic AlgorithmsabstractDeep reinforcement learning has attracted wide interest because of its extraordinary capabilities in multiple fields. However, morphology optimization by using evolutionary computation techniques has not been intensively investigated. In this paper, we explore the use of genetic algorithms (GA) to automatically design the morphology of an agent. Evaluating the performance of an agent is very time-consuming because it needs to be trained from scratch. Moreover, it is computationally infeasible to train separate controllers for all possible different morphologies of agents to identify the optimal ones and is difficult to obtain the accurate cumulative reward of an agent to estimate the performance of the morphologies. To address these issues, we use a morphology comparator as a surrogate model to estimate the probability of one morphology being better than the other, instead of directly predicting the performance of each morphology. A set of surrogate models based on a radial basis function network are developed before evolution to make full use of the data to guide the search. Experimental results indicate that the proposed method is able to efficiently find out optimal morphologies to achieve better performance than the default morphology. Jinlin Jiang, Yongchao Chen, Wenbin Pei, Junxiang Zhang, Yaqing Hou, Hong-Wei Ge, Liang Feng 0001 |
CEC | 6 |
| 2023 | BRGR: Multi-agent cooperative reinforcement learning with bidirectional real-time gain representation
Hong-Wei Ge, Liang Sun 0003, Yaqing Hou |
Appl. Intell. | 2 |
| 2023 | Hypergraph regularized low-rank tensor multi-view subspace clustering via L1 norm constraint
Hong-Wei Ge, Shuzhi Su, Shuangxi Wang |
Appl. Intell. | 2 |
| 2023 | Research on an unsupervised person re-identification based on image quality enhancement methodabstractResearch on person re-identification(Re-ID) has important value in pedestrian detection, target tracking, criminal investigation, and other related fields. In unsupervised pedestrian recognition algorithms, the accuracy of pseudo-labels is crucial to the recognition results. However, in practical scenarios, low-quality images caused by factors such as differences in camera resolution and shooting angles can affect the extraction of pedestrian features by these algorithms, thereby negatively impacting the accuracy of the labels and the learning process of the model. To address this problem, we propose an image quality enhancement algorithm for unsupervised person Re-ID (IQE). To the best of our knowledge, this study is the first to introduce detail enhancement and the application of low-light enhancement algorithms into unsupervised person Re-ID. By improving the feature extraction quality based on these two aspects, higher-quality pseudo-labels can be constructed. This method improves the accuracy of feature extraction and clustering, thereby increasing the accuracy of pseudo-labels and reducing the interference of noisy pseudo-labels. The experimental results showed that the IQE method outperformed state-of-the-art person Re-ID methods in terms of Rank-1 accuracy and mAP. Specifically, IQE achieved an 87.9% rank-1 accuracy and a 71.2% mAP on the Market-1501 dataset; a 78.1% rank-1 accuracy and a 61.7% mAP On the DukeMTMC-reID dataset; and a 51.1% rank-1 accuracy and 24.2% mAP on the MSMT17 dataset. Zhangang Hao, Hong-Wei Ge, Jiajian Huang |
Eng. Appl. Artif. Intell. | 2 |
| 2023 | Robust multi-view subspace enhanced representation based on collaborative constraints and HSIC induction
Hong-Wei Ge, Shuzhi Su, Shuangxi Wang |
Eng. Appl. Artif. Intell. | 2 |
| 2023 | Low-rank tensor multi-view subspace clustering via cooperative regularization
Hong-Wei Ge, Shuzhi Su, Shuangxi Wang |
Multim. Tools Appl. | 2 |
| 2023 | Camera-aware progressive learning for unsupervised person re-identification
Yuxuan Liu 0015, Hong-Wei Ge, Liang Sun 0003, Yaqing Hou |
Neural Comput. Appl. | 2 |
| 2023 | Attention-guided spatial-temporal graph relation network for video-based person re-identification
Hong-Wei Ge, Wenbin Pei, Yuxuan Liu 0015, Yaqing Hou, Liang Sun 0003 |
Neural Comput. Appl. | 2 |
| 2023 | AGPN: Action Granularity Pyramid Network for Video Action RecognitionabstractVideo action recognition is a fundamental task for video understanding. Action recognition in complex spatio-temporal contexts generally requires fusing of different multi-granularity action information. However, existing works do not consider spatio-temporal information modeling and fusion from the perspective of action granularity. To address this problem, this paper proposes an Action Granularity Pyramid Network (AGPN) for action recognition, which can be flexibly integrated into 2D backbone networks. The core module is the Action Granularity Pyramid Module (AGPM), a hierarchical pyramid structure with residual connections, which is established to fuse multi-granularity action spatio-temporal information. From top to bottom level in the designed pyramid structure, the receptive field decreases and action granularity becomes more refined. To enrich temporal information of the inputs, a Multiple Frame Rate Module (MFM) is proposed to mix different frame rates at a fine-grained pixel-wise level. Moreover, a Spatio-temporal Anchor Module (SAM) is employed to fix spatio-temporal feature anchors to promote the effectiveness of feature extraction. We conduct extensive experiments on three large-scale action recognition datasets, Something-Something V1 & V2 and Kinetics-400. The results demonstrate that our proposed AGPN outperforms the state-of-the-art methods for the tasks of video action recognition. Yatong Chen 0001, Hong-Wei Ge, Yuxuan Liu 0015, Xinye Cai, Liang Sun 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Complementary Attention-Driven Contrastive Learning With Hard-Sample Exploring for Unsupervised Domain Adaptive Person Re-IDabstractUnsupervised domain adaptive (UDA) methods for person re-identification (Re-ID) aim to transfer the knowledge of the labeled source domain to the unlabeled target domain without further annotations, which is challenging due to the drift of label distribution and the missing of target domain labels. Improving the clustering accuracy of pseudo-labels can help the model fit the target domain. However, the errors of pseudo-label noise will be accumulated during training, which is harmful to the model performance. Moreover, the hard samples can lead to a large gap between intra-class features and a small gap between inter-class features. To address these problems, this paper proposes a complementary attention-driven contrastive learning with hard-sample exploring (CACHE) algorithm. In CACHE, on one hand, the complementary attention module is used to improve the discriminability of the features. The obtained discriminative features can reduce noisy pseudo-labels and improve the clustering accuracy of pseudo labels; On the other hand, we explore the hard samples based on the instance relationship and cluster relationship for contrastive learning. This way can make the cluster more compact. Extensive experiments on three large-scale person re-identification benchmarks demonstrate the effectiveness of the proposed method, which significantly outperforms state-of-the-art methods in terms of mAP and CMC. Yuxuan Liu 0015, Hong-Wei Ge, Liang Sun 0003, Yaqing Hou |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Toward Realistic 3D Human Motion Prediction With a Spatio-Temporal Cross- Transformer ApproachabstractHuman motion prediction intends to predict how humans move given a historical sequence of 3D human motions. Recent transformer-based methods have attracted increasing attentions and demonstrated their promising performance in 3D human motion prediction. However, existing methods generally decompose the input of human motion information into spatial and temporal branches in a separate way and seldom consider their inherent coherence between the two branches, hence often failing to register the dynamic spatio-temporal information during the training process. Motivated by these issues, we propose a spatio-temporal cross-transformer network (STCT) for 3D human motion predictions. Specifically, we investigate various types of interaction methods (i.e., Concatenation Interaction, Msg token interaction, and Cross-transformer) to capture the coherence of the spatial and temporal branches. According to the obtained results, the proposed cross-transformer interaction method shows its superiority over other methods. Meanwhile, considering that most existing works treat the human body as a set of 3D human joint positions, the predicted human joints are proportionally less appropriate to the realistic human body due to unreasonable bone length and non-plausible poses as time progresses. We further resort to the bone constraints of human mesh to produce more realistic human motions. By fitting a parametric body model (i.e., SMPL-X model) to the predicted human joints, a reconstruction loss function is proposed to remedy the unreasonable bone length and pose errors. Comprehensive experiments on AMASS and Human3.6M datasets have demonstrated that our method achieves superior performance over compared methods. Hua Yu 0006, Xuanzhe Fan, Yaqing Hou, Wenbin Pei, Hong-Wei Ge, Xin Yang 0011, Qiang Zhang 0008, Mengjie Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | TSINet: Efficient Breast Lesion Segmentation with Three-Stream Interactive Neural Network on Magnetic Resonance ImagesabstractWith magnetic resonance imaging (MRI), the segmentation of breast lesions is helpful for early diagnosis. Since it is subjective and time-consuming for physicians to delineate the lesions, automatic segmentation is in active demand. However, the low contrast and small size of the lesions will cause difficulties in existing segmentation methods. In this paper, we propose a three-stream interactive neural network (TSINet) for the efficient segmentation of MRI breast lesions. TSINet uses three streams for extracting boundary information, location information and predicting lesion masks, respectively. The boundary extraction stream enhances the representation of boundary pixels based on gate convolution. The internal location stream that works based on the attention mechanism identifies the location of the lesions. The lesion segmentation stream fuse the boundary and location information to accurately segment the lesion. We collect MRI data from 248 breast patients at the Second Hospital of Dalian Medical University to evaluate TSINet. The results demonstrate that the TSINet can delineate lesion boundaries close to that drawn by physicians, and outperform state-of-the-art methods. Liang Sun 0003, Yunling Zhang, Hong-Wei Ge, Yiping Zhao |
BIBM | 3 |
| 2022 | A Preliminary Study of Multi-task MAP-Elites with Knowledge Transfer for Robotic Arm DesignabstractThe structure design of robotic arms is of great importance on completing industrial tasks successfully. This is a typical multi-task optimization problem when considering different constraints as different tasks. However, mainstream methods for multi-task optimization such as evolutionary multitasking and Multi-task MAP-Elites algorithms tend to encounter problems such as high computational cost and slow convergence when solving large-scale robotic arm tasks. To this end, this paper proposes a new framework based on the MAP-Elites algorithms for solving large-scale robot arm design tasks, called Multi-task MAP-Elites with Knowledge Transfer (MMKT). Specifically, this paper designs the group-based knowledge transfer process for large-scale task optimization in which all tasks are classified into different groups according to their similarity to generate multiple knowledge transfer areas; and knowledge transfer strategies are designed to enhance the quality of solutions with low fitness value. We test the effectiveness of the MMKT framework in planar robotic arm experiments (2000, 5000, and 10,000 tasks; 10, 15-dimensional search space). The experimental results prove that the MMKT outperforms the MME, CMA-ES, and classical ES algorithms. Hua Yu 0006, Han Linghu, Yaqing Hou, Hong-Wei Ge, Qiang Zhang 0008 |
CEC | 7 |
| 2022 | Random Learning Particle Swarm Optimization with Quasi-Newton Exploitation MechanismabstractIn traditional particle swarm optimization (PSO) algorithm, each particle updates its velocity and position with a learning mechanism based on its personal historical best position and the best population position. The learning mechanism in traditional PSO is simple and easy to implement, but it suffers some potential problems, such as being easily trapped in local optimum and insufficient balance. Thus, a novel random learning PSO with improved quasi-Newton exploitation mechanism (RQ-PSO) is proposed. Firstly, to improve the global search ability, a random learning mechanism is proposed through the analysis of PSO based on many kinds of learning mechanisms. Then, the random learning mechanism is effectively integrated into PSO to obtain strong global search ability and avoid falling into local optima. Finally, to keep a better balance between exploration and exploitation, an improved quasi-Newton method with strong exploitation ability is incorporated into RL-PSO. The experimental results on the complex functions from CEC-2013 and CEC-2017 test sets show that RQ-PSO outperforms the state-of-the-art PSO variants. Hong-Wei Ge, Liang Sun 0003, Xinming Zhang 0002 |
CEC | 2 |
| 2022 | Enhancing cooperation by cognition differences and consistent representation in multi-agent reinforcement learning
Hong-Wei Ge, Zhixin Ge, Liang Sun 0003, Yuxin Wang 0001 |
Appl. Intell. | 1 |
| 2022 | Robust semi non-negative low-rank graph embedding algorithm via the L21 norm
Hong-Wei Ge, JinlongYang, Shuangxi Wang |
Appl. Intell. | 2 |
| 2022 | Transformer-based two-source motion model for multi-object tracking
Jieming Yang, Hong-Wei Ge, Shuzhi Su |
Appl. Intell. | 2 |
| 2022 | Online multi-object tracking using multi-function integration and tracking simulation training
Jieming Yang, Hong-Wei Ge, Jinlong Yang 0002, Yubing Tong, Shuzhi Su |
Appl. Intell. | 2 |
| 2022 | Improving visual multi-object tracking algorithm via integrating GM-PHD and correlation filterabstractAbstract The traditional visual multi‐object tracking methods based on the Gaussian mixture probability hypothesis density filter are generally not well adapted for tracking the targets in the complex scenarios, where there are a large number of unknowable newborn objects and occluded objects, even some missing objects cannot be associated with their previous trajectories when they are redetected. An improved visual multi‐object tracking algorithm is proposed by integrating an improved efficient convolution operator of the correlation filter and the Gaussian mixture probability hypothesis density filter. First, a similarity matrix based on the intersection‐of‐union is proposed for classifying the objects of survival objects, newborn objects, and then the improved efficient convolution operator method is employed to further identify whether the objects disappear or are missing. Moreover, the feature pyramid similarity is proposed to update the objects for enhancing the tracking accuracy. Finally, compared with some challenging methods on some challenging video sequences from publicly available MOT17 dataset, the proposed Gaussian mixture probability hypothesis density–feature pyramid similarity—efficient convolution operator* method has a good performance on detecting the newborn objects, occluded objects, blurring objects and re‐identifying the missing objects with higher multiple object tracking accuracy. Jinlong Yang 0002, Jiani Miao, Hong-Wei Ge |
IET Image Process. | 4 |
| 2022 | ICMiF: Interactive cascade microformers for cross-domain person re-identification
Jiajian Huang, Hong-Wei Ge, Liang Sun 0003, Yaqing Hou |
Inf. Sci. | 2 |
| 2022 | An effective feature extraction method via spectral-spatial filter discrimination analysis for hyperspectral image
Jianqiang Gao, Hong-Wei Ge, Haifei Zhang |
Multim. Tools Appl. | 3 |
| 2022 | Adaptive kernel selection network with attention constraint for surgical instrument classificationabstractAbstract Computer vision (CV) technologies are assisting the health care industry in many respects, i.e., disease diagnosis. However, as a pivotal procedure before and after surgery, the inventory work of surgical instruments has not been researched with the CV-powered technologies. To reduce the risk and hazard of surgical tools’ loss, we propose a study of systematic surgical instrument classification and introduce a novel attention-based deep neural network called SKA-ResNet which is mainly composed of: (a) A feature extractor with selective kernel attention module to automatically adjust the receptive fields of neurons and enhance the learnt expression and (b) A multi-scale regularizer with KL-divergence as the constraint to exploit the relationships between feature maps. Our method is easily trained end-to-end in only one stage with few additional calculation burdens. Moreover, to facilitate our study, we create a new surgical instrument dataset called SID19 (with 19 kinds of surgical tools consisting of 3800 images) for the first time. Experimental results show the superiority of SKA-ResNet for the classification of surgical tools on SID19 when compared with state-of-the-art models. The classification accuracy of our method reaches up to 97.703%, which is well supportive for the inventory and recognition study of surgical tools. Also, our method can achieve state-of-the-art performance on four challenging fine-grained visual classification datasets. Yaqing Hou, Qian Liu 0001, Hong-Wei Ge, Jun Meng, Qiang Zhang 0008, Xiaopeng Wei |
Neural Comput. Appl. | 4 |
| 2022 | An Effective Feature Extraction Approach Based on Spectral-Gabor Space Discriminant Analysis for Hyperspectral Image
Jianqiang Gao, Hong-Wei Ge, Jieming Yang |
Neural Process. Lett. | 3 |
| 2022 | Online Pedestrian Multiple-Object Tracking with Prediction Refinement and Track Classification
Jieming Yang, Hong-Wei Ge, Jinlong Yang 0002, Yubing Tong, Shuzhi Su |
Neural Process. Lett. | 2 |
| 2022 | Multi-Agent Transfer Reinforcement Learning With Multi-View Encoder for Adaptive Traffic Signal ControlabstractMulti-agent reinforcement learning (MARL) based methods for adaptive traffic signal control (ATSC) have shown promising potentials to solve the heavy traffic problems. The existing MARL methods adopt centralized or distributed strategies. The former only models the environment as an agent and suffers from the exponential growth of action and state space. The latter extends the independent reinforcement learning methods, such as DQN, to multiple interactions directly or propagates information, such as state and policy, without taking their qualities into account. In this paper, we propose a multi-agent transfer reinforcement learning method to enhance the performance of MARL for ATSC, which is termed as multi-agent transfer soft actor-critic with the multi-view encoder (MT-SAC). The MT-SAC combines centralized and distributed strategies. In MT-SAC, we propose a multi-view state encoder and a transfer learning paradigm with guidance. The encoder processes input states from multiple perspectives and uses an attention mechanism to weigh the neighborhood information. While the paradigm enables the agents to handle different conditions for improving generalization abilities by transfer learning. Experimental studies on different scale road networks show that the MT-SAC outperforms the state-of-the-art algorithms and makes the traffic signal controllers more collaborative and robust. Hong-Wei Ge, Dongwan Gao, Liang Sun 0003, Yaqing Hou, Chao Yu 0004, Yuxin Wang 0001, Guozhen Tan |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2021 | A Study on Realtime Task Selection Based on Credit Information Updating in Evolutionary Multitasking
Yumeng Cao, Yaqing Hou, Liang Feng 0001, Hong-Wei Ge, Qiang Zhang 0008, Xiaopeng Wei |
EMO | 4 |
| 2021 | Multi-Scale Attention Constraint Network for Fine-Grained Visual ClassificationabstractCapturing subtle yet discriminative features constitutes a great challenge in fine-grained visual classification due to the large intra-class and small inter-class variances. Main-stream works for this problem localize at attention mechanism and feature relationship learning. However, existing methods treat the features in isolation while neglecting the effect of attention-enhanced features on relationships between different network layers. In this paper, we propose a novel attention-based method by Multi-Scale Attention Constraint network composed of two important components: (1) a feature extractor with lightweight group-wise enhanced attention blocks that guides the generation of high representation features; and (2) a multi-scale regularizer that explores the relationships between different features. Extensive experiments show that our approach achieves state-of-the-art performance on standard benchmark datasets. Moreover, we introduce a new dataset, consisting of comprehensive surgical instrument categories based on three common surgeries, to support the classification and inventory work of surgical instruments. Yaqing Hou, Hong-Wei Ge, Qiang Zhang 0008, Xiaopeng Wei |
ICME | 4 |
| 2021 | Brain Tumor Segmentation based on Knowledge Distillation and Adversarial Trainingabstract3D MRI brain tumor segmentation is a reliable method for disease diagnosis and treatment plans in the future. Early on, the segmentation of brain tumors is mostly done manually. However, manual segmentation of 3D MRI brain tumor requires professional anatomical knowledge and may be inaccurate. In this paper, we propose a 3D MRI brain tumor segmentation architecture based on the encoder-decoder structure. Specially, we introduce knowledge distillation and adversarial training methods, which compresses and improves the accuracy and robustness of the model. Furthermore, we obtain soft targets by designing multiple teacher network training and then apply them to the student network. Finally, we evaluate our method on a challenging BraTS dataset. As a result, the performance of our proposed model is superior to state-of-the-art methods. Yaqing Hou, Tianbo Li, Qiang Zhang 0008, Hua Yu 0006, Hong-Wei Ge |
IJCNN | 5 |
| 2021 | Unpaired image to image transformation via informative coupled generative adversarial networks
Hong-Wei Ge, Wenjing Kang, Liang Sun 0003 |
Frontiers Comput. Sci. | 1 |
| 2021 | Virtual samples based robust block-diagonal dictionary learning for face recognitionabstractIt is an open question to learn an over-complete dictionary from a limited number of face samples, and the inherent attributes of the samples are underutilized. Besides, the recognition performance may be adversely affected by the noise (and outliers), and the strict binary label based linear classifier is not appropriate for face recognition. To solve above problems, we propose a virtual samples based robust block-diagonal dictionary learning for face recognition. In the proposed model, the original samples and virtual samples are combined to solve the small sample size problem, and both the structure constraint and the low rank constraint are exploited to preserve the intrinsic attributes of the samples. In addition, the fidelity term can effectively reduce negative effects of noise (and outliers), and the ε-dragging is utilized to promote the performance of the linear classifier. Finally, extensive experiments are conducted in comparison with many state-of-the-art methods on benchmark face datasets, and experimental results demonstrate the efficacy of the proposed method. Shuangxi Wang, Hong-Wei Ge, Jinlong Yang 0002, Shuzhi Su |
Intell. Data Anal. | 2 |
| 2021 | Relaxed group low rank regression model for multi-class classification
Shuangxi Wang, Hong-Wei Ge, Jinlong Yang 0002, Yubing Tong |
Multim. Tools Appl. | 2 |
| 2021 | Reciprocal kernel-based weighted collaborative-competitive representation for robust face recognition
Shuangxi Wang, Hong-Wei Ge, Jinlong Yang 0002, Yubing Tong, Shuzhi Su |
Mach. Vis. Appl. | 2 |
| 2021 | Local-aware spatio-temporal attention network with multi-stage feature fusion for human action recognitionabstractAbstract In the study of human action recognition, two-stream networks have made excellent progress recently. However, there remain challenges in distinguishing similar human actions in videos. This paper proposes a novel local-aware spatio-temporal attention network with multi-stage feature fusion based on compact bilinear pooling for human action recognition. To elaborate, taking two-stream networks as our essential backbones, the spatial network first employs multiple spatial transformer networks in a parallel manner to locate the discriminative regions related to human actions. Then, we perform feature fusion between the local and global features to enhance the human action representation. Furthermore, the output of the spatial network and the temporal information are fused at a particular layer to learn the pixel-wise correspondences. After that, we bring together three outputs to generate the global descriptors of human actions. To verify the efficacy of the proposed approach, comparison experiments are conducted with the traditional hand-engineered IDT algorithms, the classical machine learning methods (i.e., SVM) and the state-of-the-art deep learning methods (i.e., spatio-temporal multiplier networks). According to the results, our approach is reported to obtain the best performance among existing works, with the accuracy of 95.3% and 72.9% on UCF101 and HMDB51, respectively. The experimental results thus demonstrate the superiority and significance of the proposed architecture in solving the task of human action recognition. Yaqing Hou, Hua Yu 0006, Pengfei Wang 0013, Hong-Wei Ge, Jianxin Zhang 0001, Qiang Zhang 0008 |
Neural Comput. Appl. | 5 |
| 2020 | Memetic Multi-agent optimization with Problem Reformulation by Coordinate RotationabstractMemetic multi-agent system (MeMAS) is recently proposed as an enhanced version that integrates meme concept into multi-agent system (MAS) wherein all meme-inspired agents have an improvement in learning performance via meme evolution independently or social interaction. In the process of solving the black box optimization problem, the potential advantages of MeMAS have not been utilized well, which makes it a fertile area for further exploration. This paper presents a memetic multi-agent optimization paradigm through coordinate rotation (MeMAO-R) to combine MeMAS with evolutionary algorithms (EAs) to improve optimization efficiency. Based on MeMAS, the particular interest of MeMAO-R is placed on assisting original complex optimization task with new tasks generated by coordinate rotation. Further, MeMAO-R constructs the social interaction mechanism which facilitates to improve their convergence speed for solving the target optimization problem by utilizing meaningful information transferred across multiple agents with differing views of the target problem. Besides, MeMAO-R employs one or more classical EAs as the fundamental population based evolutionary solvers for multiple agents to optimize multiple tasks in a multi-agent scenario. Lastly, to testify the efficacy of the proposed MeMAO-R, comprehensive empirical studies on basic optimization problems are provided. Yaqing Hou, Qiang Zhang 0008, Hong-Wei Ge, Xin Yang 0011, Abhishek Gupta 0001, Xianneng Li |
CEC | 4 |
| 2020 | D3PG: Decomposed Deep Deterministic Policy Gradient for Continuous Control
Yinzhao Dong, Chao Yu 0004, Hong-Wei Ge |
DAI | 3 |
| 2020 | Efficient Latency Bound Analysis for Data Chains of Real-Time Tasks in Multiprocessor SystemsabstractEnd-to-end latency analysis is one of the key problems in the automotive embedded system design. In this paper, we propose an efficient worst-case end-to-end latency analysis method for data chains of periodic real-time tasks executed on multiprocessors under a partitioned fixed-priority preemptive scheduling policy. The key idea of this research is to improve the analysis efficiency by transforming the problem of bounding the worst-case latency of the data chain to a problem of bounding the releasing interval of data propagation instances for each pair of consecutive tasks in the chain. In particular, we derive an upper bound on the releasing interval of successive data propagation instances to yield the desired data chain latency bound by a simple accumulation. Based on the above idea, we present an efficient latency upper bound analysis algorithm with polynomial time complexity. Experiments with randomly generated task sets based on a generic automotive benchmark show that our proposed approach can obtain a relatively tighter data chain latency upper bound with lower computational cost. Jiankang Ren, Junlong Zhou, Hong-Wei Ge, Guozhen Tan |
DATE | 4 |
| 2020 | Two-stage Automatic Image Annotation Based on Latent Semantic Scene ClassificationabstractThe rapid growth of multimedia content makes existing automatic image annotation techniques difficult to satisfy the demands of real-world applications. In this paper, we propose a two-stage automatic image annotation algorithm (TAIA) based on latent semantic scene classification. In the offline training phase, the hidden connectivity of labels is firstly excavated by a directed-weighed graph based on label co-occurrence relation matrix, and then the latent scene categories are detected among the labels by using nonnegative matrix factorization. Further, we propose a multi-view extreme learning machine (MELM) to learn the probability that the multiple visual feature maps to the semantic scenes. In the online annotation phase, the image to be annotated is fed to the scene classifier MELM to identify its relevant scenes. Then k-nearest neighbor based annotator is conducted on the relevant scenes to predict labels for the unannotated images. The TAIA is formulated in such a framework so that the relationship between labels and semantic scenes is fully considered, and the hard classification problem is solved. The experimental results on multiple datasets have demonstrated that the proposed framework TAIA is both effective and efficient. Hong-Wei Ge, Kai Zhang 0050, Yaqing Hou, Chao Yu 0004, Mingde Zhao 0002, Zhen Wang 0004, Liang Sun 0003 |
IJCNN | 1 |
| 2020 | A Preliminary Study of Fusion ARTs with Adaptively Information Intensity Attenuation ControllingabstractFusion ART is an enhanced version of Adaptive Resonance Theory (ART) which is derived from a biologically-plausible theory of human cognitive information processing. Due to its well-established ability of learning associative mappings across multimodal pattern channels in an online and incremental manner, fusion ART has been widely applied in many real world learning problems. In this paper, we take a Fusion Architecture for Learning, Cognition, and Navigation (FALCON) as the specification and essential backbone of fusion ART and introduce an intensity attenuation controller δ for adaptively adjusting the intensity of information captured from the environment, by taking inspiration from Broadbent-Treisman Filter-Attenuation's perceptual model of environmental attention. Particularly, we propose both an adaptive δ detection algorithm as well as a δ-based pruning algorithm to enhance the learning performance of FALCON while reduce the redundant memory storage incurred by the "detrimental δ". To verify the effectiveness and efficiency of our proposed method, comprehensive experimental studies are carried out on a classical minefield navigation task. Wenxuan Zhu, Yaqing Hou, Qiang Zhang 0008, Hong-Wei Ge, Xin Yang 0011, Liang Feng 0001, Xinghua Qu |
IJCNN | 4 |
| 2020 | Image compact-resolution and reconstruction using reversible networkabstractThe dual problem of image super‐resolution (SR), which is referred to as compact‐resolution (CR), and the corresponding image reconstruction are studied. These two problems have been studied independently by the researchers. In this study, a novel model for image CR and the corresponding reconstruction using the reversible network has been proposed. The reversible network has two properties, the first property, lossless information forwarding, which makes the compact‐resolved image retain more information from the original HR image. The second property, bidirectional mapping, by which the forward and reverse propagation of a reversible network can be utilised to implement image CR and reconstruction, respectively, i.e. using the reverse process of image CR to guide the reconstruction. In addition, the utilisation of a reversible network may reduce the size of the model. The superiority of the proposed model was demonstrated by comparing its performance with the state‐of‐the‐art methods on four well‐known benchmark datasets. Jieming Yang, Hong-Wei Ge, Jinlong Yang 0002, Yubing Tong |
IET Image Process. | 2 |
| 2020 | A Novel Geometric Mean Feature Space Discriminant Analysis Method for Hyperspectral Image Feature Extraction
Hong-Wei Ge, Jianqiang Gao, Yubing Tong, Jun Sun 0008 |
Neural Process. Lett. | 2 |
| 2020 | Distributed Multiagent Coordinated Learning for Autonomous Driving in Highways Based on Dynamic Coordination GraphsabstractAutonomous driving is one of the most important AI applications and has attracted extensive interest in recent years. A large number of studies have successfully applied reinforcement learning techniques in various aspects of autonomous driving, ranging from low-level control of driving maneuvers to higher level of strategic decision-making. However, comparatively less progress has been made in investigating how co-existing autonomous vehicles would interact with each other in a common environment and how reinforcement learning can be helpful in such situations by applying multiagent reinforcement learning techniques in the high-level strategic decision-making of the following or overtaking for a group of autonomous vehicles in highway scenarios. Learning to achieve coordination among vehicles in such situations is challenging due to the unique feature of vehicular mobility, which renders it infeasible to directly apply the existing coordinated learning approaches. To solve this problem, we propose using dynamic coordination graph to model the continuously changing topology during vehicles' interactions and come up with two basic learning approaches to coordinate the driving maneuvers for a group of vehicles. Several extension mechanisms are then presented to make these approaches workable in a more complex and realistic setting with any number of vehicles. The experimental evaluation has verified the benefits of the proposed coordinated learning approaches, compared with other approaches that learn without coordination or rely on some traditional mobility models based on some expert driving rules. Chao Yu 0004, Xin Wang 0077, Xin Xu 0001, Minjie Zhang 0001, Hong-Wei Ge, Jiankang Ren, Liang Sun 0003, Bingcai Chen, Guozhen Tan |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2019 | Memetic Multi-agent Optimization in High Dimensions using Random EmbeddingsabstractIn this paper, we propose a memetic multi-agent optimization (MeMAO) paradigm to enhance the search efficacy of classical EAs (i.e., Differential Evolution (DE)) in solving the complex optimization problems. The essential backbone of MeMAO is a recently proposed memetic multi-agent learning system wherein agents acquire increasing learning capabilities by interacting with the environment mainly in a reinforcement learning manner. Differing from MeMAS, the particular interest of MeMAO is placed on addressing the specific challenges when applying classical EAs to optimize the high dimensional optimization problems with a "low effective dimensionality". To achieve this, the target optimization problem is firstly re-formulated into multiple low dimensional tasks via random embedding methods. Further, MeMAO employs DE as the fundamental population based evolutionary solver for multiple agents to optimize multiple low dimensional tasks in a multi-agent scenario. Importantly, MeMAO constructs the social interaction mechanisms among multiple agents, hence improves their convergence speed for solving the target optimization problem by sharing the beneficial information across multiple agents. Lastly, to testify the efficacy of the proposed MeMAO, comprehensive empirical studies on 8 synthetic optimization problems with a dimensionality of 2,000 are provided. Yaqing Hou, Hong-Wei Ge, Qiang Zhang 0008, Xinghua Qu, L. Feng, Abhishek Gupta 0001 |
CEC | 3 |
| 2019 | Exploring Overall Contextual Information for Image Captioning in Human-Like Cognitive StyleabstractImage captioning is a research hotspot where encoder-decoder models combining convolutional neural network (CNN) and long short-term memory (LSTM) achieve promising results. Despite significant progress, these models generate sentences differently from human cognitive styles. Existing models often generate a complete sentence from the first word to the end, without considering the influence of the following words on the whole sentence generation. In this paper, we explore the utilization of a human-like cognitive style, i.e., building overall cognition for the image to be described and the sentence to be constructed, for enhancing computer image understanding. This paper first proposes a Mutual-aid network structure with Bidirectional LSTMs (MaBi-LSTMs) for acquiring overall contextual information. In the training process, the forward and backward LSTMs encode the succeeding and preceding words into their respective hidden states by simultaneously constructing the whole sentence in a complementary manner. In the captioning process, the LSTM implicitly utilizes the subsequent semantic information contained in its hidden states. In fact, MaBi-LSTMs can generate two sentences in forward and backward directions. To bridge the gap between cross-domain models and generate a sentence with higher quality, we further develop a cross-modal attention mechanism to retouch the two sentences by fusing their salient parts as well as the salient areas of the image. Experimental results on the Microsoft COCO dataset show that the proposed model improves the performance of encoder-decoder models and achieves state-of-the-art results. Hong-Wei Ge, Zehang Yan, Kai Zhang 0050, Mingde Zhao 0002, Liang Sun 0003 |
ICCV | 1 |
| 2019 | Multi-Grained Cascade AdaBoost Extreme Learning Machine for Feature RepresentationabstractExtreme learning machine (ELM) has been well recognized for characteristics such as less training parameters, fast training speed and strong generalization ability. Due to its high efficiency, researchers have embedded ELMs into deep learning frameworks to address the problems of high time-consumption and computational complexities that are encountered in the traditional deep neural networks. However, existing ELM-based deep learning algorithms usually neglect the spatial relationship of original data. In this paper, we propose a multi-grained cascade AdaBoost based weighted ELM algorithm (gcAWELM) for feature representation. We use AdaBoost based weighted ELM as a basic module to construct cascade structure for feature learning. Different ensemble ELMs trend to extract varied features. Moreover, multi-grained scanning is employed to exploit the spatial structure of the original data. The gcAWELM can determine the number of cascade levels adaptively and has simpler structure and fewer parameters compared with the traditional deep models. The results on image datasets with different scales show that the gcAWELM can achieve competitive performance for different learning tasks even with the same parameter settings. Hong-Wei Ge, Weiting Sun, Mingde Zhao 0002, Kai Zhang 0050, Liang Sun 0003, Chao Yu 0004 |
IJCNN | 1 |
| 2019 | Strategy Selection in Complex Game Environments Based on Transfer Reinforcement LearningabstractBoosting the learning process in the new task by making use of previously obtained knowledge has been a challenging task in many fields of industrial engineering and scientific. In this paper, we propose a transfer reinforcement learning model with knowledge Inheritance and decision-making Assistance (trIA). In the stage of knowledge inheritance, trIA adopts a model that employs a simultaneous multi-task and multi-instance learning strategy to compress acquired experts knowledge from distinct task into a global multi-task agent. In the stage of decision-making assistance, trIA adopts a dual-column progressive neural network framework to fully utilize the previous knowledge in the global multi-task agent and the acquired knowledge in the new task. The experimental results on the Atari domain demonstrate that the proposed knowledge inheritance model can performed at nearly the same level as the experts on the distinct source task environments. The results also demonstrate that the decision-making assistance model can transfer knowledge from the source tasks to the target tasks effectively. Moreover, the comparative results with the state-ofthe-art algorithms validate the effectiveness of the proposed trIA for strategy selection in complex game environments. Hong-Wei Ge, Mingde Zhao 0002, Kai Zhang 0050, Liang Sun 0003 |
IJCNN | 1 |
| 2019 | Non-negative matrix factorization based modeling and training algorithm for multi-label learning
Liang Sun 0003, Hong-Wei Ge, Wenjing Kang |
Frontiers Comput. Sci. | 2 |
| 2019 | An attention mechanism based convolutional LSTM network for video action recognition
Hong-Wei Ge, Zehang Yan, Wenhao Yu 0005, Liang Sun 0003 |
Multim. Tools Appl. | 1 |
| 2019 | Hyperspectral Image Feature Extraction Using Maclaurin Series Function Curve Fitting
Hong-Wei Ge, Jianqiang Gao |
Neural Process. Lett. | 2 |
| 2019 | A Many-Objective Evolutionary Algorithm With Two Interacting Processes: Cascade Clustering and Reference Point Incremental LearningabstractResearches have shown difficulties in obtaining proximity while maintaining diversity for many-objective optimization problems. Complexities of the true Pareto front pose challenges for the reference vector-based algorithms for their insufficient adaptability to the diverse characteristics with no priori. This paper proposes a many-objective optimization algorithm with two interacting processes: cascade clustering and reference point incremental learning (CLIA). In the population selection process based on cascade clustering (CC), using the reference vectors provided by the process based on incremental learning, the nondominated and the dominated individuals are clustered and sorted with different manners in a cascade style and are selected by round-robin for better proximity and diversity. In the reference vector adaptation process based on reference point incremental learning, using the feedbacks from the process based on CC, proper distribution of reference points is gradually obtained by incremental learning. Experimental studies on several benchmark problems show that CLIA is competitive compared with the state-of-the-art algorithms and has impressive efficiency and versatility using only the interactions between the two processes without incurring extra evaluations. Hong-Wei Ge, Mingde Zhao 0002, Liang Sun 0003, Zhen Wang 0004, Guozhen Tan, Qiang Zhang 0008, C. L. Philip Chen |
IEEE Trans. Evol. Comput. | 1 |
| 2018 | Unsupervised Transformation Network Based on GANs for Target-Domain Oriented Multi-domain Image Translation
Hong-Wei Ge, Liang Sun 0003 |
ACCV (2) | 1 |
| 2018 | A Selective Ensemble Learning Framework for ECG-Based Heartbeat Classification with Imbalanced Data
Hong-Wei Ge, Keyi Sun, Liang Sun 0003, Mingde Zhao 0002, Chunguo Wu |
BIBM | 1 |
| 2018 | A Many-Objective Evolutionary Algorithm with Fast Clustering and Reference Point RedistributionabstractThe design of effective evolutionary many-objective optimization algorithms is challenging for the difficulties in obtaining proximity while maintaining diversity. In this paper, a fast Clustering based Algorithm with reference point Redistribution (fastCAR) is proposed. In the clustering process, a fast Pareto dominance based clustering mechanism is proposed to increase the evolution selection pressure that acts as a selection operator. In the redistribution process, the reference vectors are periodically redistributed by using an SVM classifier to maintain the diversity of the population without extra burden on fitness evaluations. The experimental results show that the proposed algorithm fastCAR obtains competitive results on box-constrained CEC'2018 many-objective benchmark functions in comparison with 8 state-of-the-art algorithms. Mingde Zhao 0002, Hong-Wei Ge, Hongyan Han, Liang Sun 0003 |
CEC | 2 |
| 2018 | Decentralized Multiagent Reinforcement Learning for Efficient Robotic Control by Coordination Graphs
Chao Yu 0004, Jiankang Ren, Hong-Wei Ge, Liang Sun 0003 |
PRICAI (1) | 4 |
| 2018 | Adaptively Shaping Reinforcement Learning Agents via Human Reward
Chao Yu 0004, Tianpei Yang, Wenxuan Zhu, Yuchen Li 0006, Hong-Wei Ge, Jiankang Ren |
PRICAI (1) | 6 |
| 2018 | Device Clustering Algorithm Based on Multimodal Data Correlation in Cognitive Internet of ThingsabstractWith the development of information network, the popularity of Internet of Things (IoT) is an irreversible trend, and the intelligent demands for IoT is becoming more and more urgent. How to improve the cognitive ability of IoT is a new challenge and therefore has given rise to the emergence of cognitive IoT (CIoT). In this paper, a device-level multimodal data correlation mining model is first designed based on the canonical correlation analysis to transform the data feature into a subspace and analyze the data correlation. The correlation of the device is obtained based on the comprehensive of data correlation and the location information of the device. Then a heterogeneous clustering model (heterogeneous device clustering) is proposed by using the result of the correlation analysis to classify the device. Finally, we propose a device clustering algorithm based on multimodal data correlation for CIoT, which combines the functions of multimodal data correlation analyze with device clustering. Extensive simulations are carried out and our results show that the proposed algorithm can effectively improve the quality of data transmission and the intelligent service. Di Wang 0004, Fuzhen Xia, Hong-Wei Ge |
IEEE Internet Things J. | 4 |
| 2018 | Low-density noise removal based on lambda multi-diagonal matrix filter for binary image
Hong-Wei Ge, Jianqiang Gao |
Neural Comput. Appl. | 2 |
| 2018 | Face Recognition Using Gabor-Based Feature Extraction and Feature Space Transformation Fusion Method for Single Image per Person Problem
Hong-Wei Ge, Yubing Tong |
Neural Process. Lett. | 2 |
| 2018 | Modified Gaussian inverse Wishart PHD filter for tracking multiple non-ellipsoidal extended targets
Peng Li 0076, Hong-Wei Ge, Jinlong Yang 0002 |
Signal Process. | 2 |
| 2018 | Multi-graph embedding discriminative correlation feature learning for image recognition
Shuzhi Su, Hong-Wei Ge, Yubing Tong |
Signal Process. Image Commun. | 2 |
| 2017 | An automatic motif recognition algorithm in DNA sequences based on particle swarm optimization and random projectionabstractMotif recognition plays a significant role for understanding the biological meanings of sequences. In this paper, a particle swarm optimization and random projection based algorithm (PSORPS) is proposed for recognizing DNA motifs. First, noise subsequences are filtered by a random projection strategy for constructing objective space. Then the PSO is used to automatically search the motifs of DNA sequences in the constructed l-mer objective space. Finally, the operators of independent drift and associated drift are conducted on the optimization results to alleviate base deviation. The experiments are carried out on real-world biological datasets, and the results validate the effectiveness of the proposed algorithm. Hong-Wei Ge, Liang Sun 0003, Jinghong Yu |
BIBM | 1 |
| 2017 | Fast batch searching for protein homology based on compression and clusteringabstractBACKGROUND: In bioinformatics community, many tasks associate with matching a set of protein query sequences in large sequence datasets. To conduct multiple queries in the database, a common used method is to run BLAST on each original querey or on the concatenated queries. It is inefficient since it doesn't exploit the common subsequences shared by queries. RESULTS: We propose a compression and cluster based BLASTP (C2-BLASTP) algorithm to further exploit the joint information among the query sequences and the database. Firstly, the queries and database are compressed in turn by procedures of redundancy analysis, redundancy removal and distinction record. Secondly, the database is clustered according to Hamming distance among the subsequences. To improve the sensitivity and selectivity of sequence alignments, ten groups of reduced amino acid alphabets are used. Following this, the hits finding operator is implemented on the clustered database. Furthermore, an execution database is constructed based on the found potential hits, with the objective of mitigating the effect of increasing scale of the sequence database. Finally, the homology search is performed in the execution database. Experiments on NCBI NR database demonstrate the effectiveness of the proposed C2-BLASTP for batch searching of homology in sequence database. The results are evaluated in terms of homology accuracy, search speed and memory usage. CONCLUSIONS: It can be seen that the C2-BLASTP achieves competitive results as compared with some state-of-the-art methods. Hong-Wei Ge, Liang Sun 0003, Jinghong Yu |
BMC Bioinform. | 1 |
| 2017 | A label embedding kernel method for multi-view canonical correlation analysis
Shuzhi Su, Hong-Wei Ge, Yun-Hao Yuan 0001 |
Multim. Tools Appl. | 2 |
| 2017 | Cooperative Hierarchical PSO With Two Stage Variable Interaction Reconstruction for Large Scale OptimizationabstractLarge scale optimization problems arise in diverse fields. Decomposing the large scale problem into small scale subproblems regarding the variable interactions and optimizing them cooperatively are critical steps in an optimization algorithm. To explore the variable interactions and perform the problem decomposition tasks, we develop a two stage variable interaction reconstruction algorithm. A learning model is proposed to explore part of the variable interactions as prior knowledge. A marginalized denoising model is proposed to construct the overall variable interactions using the prior knowledge, with which the problem is decomposed into small scale modules. To optimize the subproblems and relieve premature convergence, we propose a cooperative hierarchical particle swarm optimization framework, where the operators of contingency leadership, interactional cognition, and self-directed exploitation are designed. Finally, we conduct theoretical analysis for further understanding of the proposed algorithm. The analysis shows that the proposed algorithm can guarantee converging to the global optimal solutions if the problems are correctly decomposed. Experiments are conducted on the CEC2008 and CEC2010 benchmarks. The results demonstrate the effectiveness, convergence, and usefulness of the proposed algorithm. Hong-Wei Ge, Liang Sun 0003, Guozhen Tan, C. L. Philip Chen |
IEEE Trans. Cybern. | 1 |
| 2016 | Enhancing protein homology batch search algorithm with sequence compression and clusteringabstractHomology search is a tremendous application of bioinformatics in the field of molecular biology, protein function analysis and drug development. To perform batch search in the growing database, the basic approach is to run Blast on each of the original queries or concatenate queries by grouping them together. This paper proposes an enhanced protein homology batch search algorithm with sequence compression and clustering (C2-BLASTP), which takes advantage of the joint information among the query sequences as well as the database. In C2-BLASTP, the queries and database are firstly compressed by redundancy analysis. And then the database is clustered according to subsequence similarity. Following this, hits finding can be implemented in the clustered database. Furthermore, a final execution database is reconstructed based on potential hits to mitigate the increasing scale of the sequence database. Finally, homology batch search is performed in execution database. Experiments on NCBI NR database demonstrate the effectiveness of the C2-BLASTP for homology batch search in terms of homology accuracy, search speed and memory usage. Hong-Wei Ge, Liang Sun 0003, Jinghong Yu |
BIBM | 1 |
| 2016 | Intelligent scheduling in flexible job shop environments based on artificial fish swarm algorithm with estimation of distributionabstractEfficient scheduling strategy is crucial to a manufacturing system in flexible job shop environments. The flexible job shop scheduling problem (FJSP) is a complex combinatorial optimization problem due to the consideration of both machine assignment and operation sequence. In this paper, an efficient artificial fish swarm model with estimation of distribution (AFSA-ED) is proposed for obtaining intelligent scheduling strategies. In AFSA-ED, an integrated initialization algorithm is proposed for machine assignment and operation sequence initialization, and then the population is divided into two sub-populations and evolved respectively. Moreover, the designed pre-principle and post-principle arranging mechanism are applied to the different sub-populations for enhancing the diversity. Following this, an artificial fish swarm algorithm with estimation of distribution is proposed to explore the search space for promising solutions. Besides, an attracting behavior is designed to improve the global exploration ability and a public factor based critical path search strategy is presented to enhance the local exploitation ability. Simulated experiments are carried on BRdata, BCdata and HUdata benchmark sets. The statistical computation results validate the performance of the proposed algorithm in solving the FJSP, as compared with some other state of the art algorithms. Hong-Wei Ge, Liang Sun 0003 |
CEC | 1 |
| 2016 | Multi-locality correlation feature learning for image recognitionabstractLocality-based feature learning has drawn more and more attentions recently. However, most of locality-based feature learning methods only consider a kind of local neighbor information, and such the locality-based methods are difficult to well reveal intrinsic geometrical structure of raw high-dimensional data. In this paper, we propose a novel multi-locality correlation feature learning algorithm for multi-view data, called multi-locality discrimination canonical correlation analysis (MLDCCA), which can learn nonlinear correlation features with strong discriminative power. Different from the locality-based methods, our algorithm not only employs multiple local patches of each raw data to well capture the intrinsic geometrical structure information, but also fully considers intraclass scatter information for further enhancing the class separability of the learned correlation features. Extensive experimental results on several real-word image datasets have demonstrated the effectiveness of our algorithm. Shuzhi Su, Hong-Wei Ge, Yun-Hao Yuan 0001 |
ISCC | 2 |
| 2016 | Shape selection partitioning algorithm for Gaussian inverse Wishart probability hypothesis density filter for extended target trackingabstractThe Gaussian inverse Wishart probability hypothesis density (GIW‐PHD) filter is a promising approach for tracking an unknown number of extended targets. However, it does not achieve satisfactory performance if targets in different sizes are spatially close and manoeuvring because the partitioning methods are sensitive to manoeuvres. To solve this problem, the authors propose the shape selection partitioning (SSP) measurement partitioning algorithm. The proposed algorithm first calculates potential centres and shapes of targets. It then combines each centre with different shapes to divide measurements into subcells. Accordingly, some candidate partitions can be obtained. Finally, it selects the most likely candidate partition and outputs the corresponding subcells. Simulation results show that the application of SSP to the GIW‐PHD filter can achieve better performance when targets are spatially close and manoeuvring, which leads to a lower optimal subpattern assignment distance and a higher accuracy of the sum of weights. Peng Li 0076, Hong-Wei Ge, Jinlong Yang 0002, Huanqing Zhang |
IET Signal Process. | 2 |
| 2016 | Kernel propagation strategy: A novel out-of-sample propagation projection for subspace learning
Shuzhi Su, Hong-Wei Ge, Yun-Hao Yuan 0001 |
J. Vis. Commun. Image Represent. | 2 |
| 2016 | Multi-patch embedding canonical correlation analysis for multi-view feature learning
Shuzhi Su, Hong-Wei Ge, Yun-Hao Yuan 0001 |
J. Vis. Commun. Image Represent. | 2 |
| 2016 | A GM-PHD algorithm for multiple target tracking based on false alarm detection with irregular window
Huanqing Zhang, Hong-Wei Ge, Jinlong Yang 0002, Yun-Hao Yuan 0001 |
Signal Process. | 2 |
| 2014 | Support vector description of clusters for content-based image annotation
Liang Sun 0003, Hong-Wei Ge, Shinichi Yoshida, Yanchun Liang 0001, Guozhen Tan |
Pattern Recognit. | 2 |
| 2014 | Fractional-order embedding canonical correlation analysis and its applications to multi-view dimensionality reduction and recognition
Yun-Hao Yuan 0001, Quan-Sen Sun, Hong-Wei Ge |
Pattern Recognit. | 3 |
| 2013 | An improved multi-target tracking algorithm based on CBMeMBer filter and variational Bayesian approximation
Jinlong Yang 0002, Hong-Wei Ge |
Signal Process. | 2 |
| 2011 | Study of dependency between the input noise and the parameter in fuzzy linear regression modelabstractWhen noise exists in data, it is a very meaningful topic to reveal the dependency between the parameter h (i.e. the threshold value used to measure degree of fit) in Fuzzy linear regression (FLR) model and the input noise. In this paper, the FLR model is first extended to its regularized version, i.e. regularized fuzzy linear regression (RFLR) model, so as to enhance its generalization capability; then RFLR model is explained as the corresponding equivalent maximum a posteriori MAP problem; finally, the approximately inverse proportional dependency relationships that the parameter h with Laplacian noisy input and Uniform noisy input should follow are derived, respectively. Our experimental results also confirm this theoretical claim. We believe that this conclusion provides an important reference for us to determine h in FLR model with noisy input. Hong-Wei Ge, Shitong Wang 0001 |
FUZZ-IEEE | 1 |
| 2009 | An equilibrium multi-hop cluster hierarchy for wireless sensor networksabstractTo overcome the inefficiency of the cluster structure construction in the existing wireless sensor networks, we propose an equilibrium multi-hop cluster hierarchy (EMCH). Both the remaining energy of cluster head nodes and the cluster distribution are considered during the cluster construction. According to the signal intensity of communication between sensor nodes, the cluster structure is formed in a reasonable distribution. The experiment shows EMCH can balance the energy consumption and prolong the lifetime of WSNs. Keqiu Li, Hong-Wei Ge, Yanming Shen |
IWCMC | 3 |
| 2009 | Identification and control of nonlinear systems by a time-delay recurrent neural network
Hong-Wei Ge, Wenli Du, Feng Qian 0004, Yanchun Liang 0001 |
Neurocomputing | 1 |
| 2008 | Theoretical Choice of the Optimal Threshold for Possibilistic Linear Model With Noisy InputabstractBased on possibility concepts, various possibilistic linear models (PLMs) have been proposed, and their pivotal role in fuzzy modeling and associated applications has been established. When adopting PLMs, one has to adopt an appropriate threshold (lambda) value. However, choosing such a value is by no means trivial, and is still an open theoretical issue. In this paper, we propose a solution by first extending the PLM to its regularized version, i.e., a regularized PLM (RPLM), such that its generalization capability can be enhanced. The RPLM is then formulated as a maximum a posteriori (MAP) framework, which facilitates the determination of the theoretically optimal threshold value for the RPLM with noisy input. Our mathematical derivations reveal the approximately inversely proportional relationship between the threshold and the standard deviation of Gaussian noisy input. This is also confirmed by the simulation results. This finding is very helpful for the practical applications of both PLMs and RPLMs. Hong-Wei Ge, Korris Fu-Lai Chung, Shitong Wang 0001 |
IEEE Trans. Fuzzy Syst. | 1 |
| 2008 | An Effective PSO and AIS-Based Hybrid Intelligent Algorithm for Job-Shop SchedulingabstractThe optimization of job-shop scheduling is very important because of its theoretical and practical significance. In this paper, a computationally effective algorithm of combining PSO with AIS for solving the minimum makespan problem of job-shop scheduling is proposed. In the particle swarm system, a novel concept for the distance and velocity of a particle is presented to pave the way for the job-shop scheduling problem. In the artificial immune system, the models of vaccination and receptor editing are designed to improve the immune performance. The proposed algorithm effectively exploits the capabilities of distributed and parallel computing of swarm intelligence approaches. The algorithm is examined by using a set of benchmark instances with various sizes and levels of hardness and is compared with other approaches reported in some existing literature works. The computational results validate the effectiveness of the proposed approach. Hong-Wei Ge, Liang Sun 0003, Yanchun Liang 0001, Feng Qian 0004 |
IEEE Trans. Syst. Man Cybern. Part A | 1 |
| 2007 | Dependency between degree of fit and input noise in fuzzy linear regression using non-symmetric fuzzy triangular coefficients
Hong-Wei Ge, Shitong Wang 0001 |
Fuzzy Sets Syst. | 1 |
| 2006 | Mechanism Design and Analysis of Genetic Operations in Solving Traveling Salesman Problems
Hong-Wei Ge, Yanchun Liang 0001, Maurizio Marchese |
ICIC (1) | 1 |