VLDB 2026 Research / reviewers in the wild / expert
Qiang Zhang 0008
dblp:72/3527-8
· DBLP profile ↗
208ranked-venue papers
6as first author
156since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 93 · 5 first-author · 64 since 2021Graphics, computer vision, multimedia, augmented reality and games · 58 · 39 since 2021Computer networks · 31 · 31 since 2021Applied, interdisciplinary, general and emerging computing · 28 · 1 first-author · 19 since 2021Systems, architecture and hardware · 11 · 10 since 2021Databases, data management, data science and information retrieval · 9 · 5 since 2021Human-computer interaction and ubiquitous computing · 7 · 7 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Learning Structurally Stabilized Representations for Lossless DNA StorageabstractThis paper presents Reed-Solomon coded single-stranded representation learning (RSRL), a novel end-to-end model for learning representations for lossless DNA data storage. In contrast to existing learning-based methods, RSRL is inspired by both error-correction codec and structural biology. Specifically, RSRL first learns the representations for the subsequent storage from the binary data transformed by the Reed-Solomon codec (RS code). Then, the representations are masked by an RS-code-informed mask to focus on correcting the burst errors occurring in the learning process. The synergy of RS masks and graph attention enables active error localization, breaking through the limitations of traditional passive error correction. With the decoded representations with error corrections, a novel biologically stabilized loss is formulated to regularize the data representations to possess stable single-stranded structures. By incorporating these novel strategies, RSRL can learn highly durable, dense, and lossless representations for subsequent storage tasks in DNA sequences. The proposed RSRL has been compared with a number of baselines in real-world tasks of multi-type data storage. The experimental results obtained demonstrate that RSRL can store diverse types of data with much higher information density and durability, but much lower error rates. Ben Cao, Xue Li 0019, Tiantian He 0001, Bin Wang 0005, Shihua Zhou, Qiang Zhang 0008 |
AAAI | 7 |
| 2026 | Distillation-Guided Structural Transfer for Continual Learning Beyond Sparse Distributed MemoryabstractSparse neural systems are gaining traction for efficient continual learning due to their modularity and low interference. Architectures like Sparse Distributed Memory Multi-Layer Perceptrons (SDMLP) construct task-specific subnetworks via Top-K activation and have shown resilience against catastrophic forgetting. However, their rigid modularity poses two fundamental challenges: (1) the isolation of sparse subnetworks severely limits cross-task knowledge reuse; and (2) increased sparsity reduces interference but often degrades performance due to constrained feature sharing.We propose Selective Subnetwork Distillation (SSD), a structurally guided continual learning framework that treats distillation not as a regularizer, but as a topology-aligned information conduit. By identifying neurons with high activation frequency, SSD selectively distills knowledge within previous Top-K subnetworks and output logits—without requiring replay or task labels—preserving both sparsity and functional specialization.Unlike conventional distillation, SSD operates under hard modular constraints and enables structural realignment without altering the sparse architecture.While our method is validated on SDMLP, its structure-aligned mechanism has the potential to generalize to other sparse networks as a plug-in module for promoting representation sharing.Comprehensive experiments on Split CIFAR-10, CIFAR-100, and MNIST demonstrate that SSD improves accuracy, retention, and manifold coverage, offering a structurally grounded solution to sparse continual learning. Huiyan Xue, Xuming Ran, Qi Xu 0008, Enhui Li, Yi Xu 0008, Qiang Zhang 0008 |
AAAI | 7 |
| 2026 | Explaining Synergistic Effects in Social RecommendationsabstractIn social recommenders, the inherent nonlinearity and opacity of synergistic effects across multiple social networks hinders users from understanding how diverse information is leveraged for recommendations, consequently diminishing explainability. However, existing explainers can only identify the topological information in social networks that significantly influences recommendations, failing to further explain the synergistic effects among this information. Inspired by existing findings that synergistic effects enhance mutual information between inputs and predictions to generate information gain, we extend this discovery to graph data. We quantify graph information gain to identify subgraphs embodying synergistic effects. Based on the theoretical insights, we propose SemExplainer, which explains synergistic effects by identifying subgraphs that embody them. SemExplainer first extracts explanatory subgraphs from multi-view social networks to generate preliminary importance explanations for recommendations. A conditional entropy optimization strategy to maximize information gain is developed, thereby further identifying subgraphs that embody synergistic effects from explanatory subgraphs. Finally, SemExplainer searches for paths from users to recommended items within the synergistic subgraphs to generate explanations for the recommendations. Extensive experiments on three datasets demonstrate the superiority of SemExplainer over baseline methods, providing superior explanations of synergistic effects. The implementation is available at https://github.com/yushuowiki/SemExplainer. Yicong Li 0006, Shan Jin 0003, Shuo Wang 0040, Jiaying Liu 0006, Shuo Yu 0001, Qiang Zhang 0008, Kuanjiu Zhou, Feng Xia 0001 |
WWW | 7 |
| 2026 | Missingness-aware Federated Contrastive Learning on Semantic GraphsabstractSemantic graphs are fundamental to the Web, enabling applications such as semantic search, recommendation, and knowledge-intensive reasoning. In decentralized Web environments, however, these graphs are distributed across organizations and constrained by strict privacy policies, making centralized training infeasible. Federated learning provides a promising solution, yet its effectiveness is severely limited by the dual incompleteness of real-world semantic graphs: missing node attributes and incomplete relational structures. Such dual missingness, often heterogeneous and unobserved across clients, causes substantial degradation in model performance. We present FedCL, a missingness-aware federated contrastive learning framework for dual-incomplete semantic graphs. FedCL introduces two key components: a topology estimation module, grounded in rate–distortion theory, that privately quantifies structural incompleteness across clients, and a federated reconstruction module that leverages these estimations to generate plausible relations without inferring sensitive attributes. To further improve robustness, FedCL integrates graph contrastive learning across reconstructed subgraphs, ensuring semantic consistency across heterogeneous and incomplete client graphs. Experiments on benchmark citation and Web datasets demonstrate that FedCL consistently outperforms state-of-the-art baselines in accuracy and robustness under heterogeneous missingness, while preserving strong privacy guarantees. These results highlight FedCL as a scalable and trustworthy approach for federated learning on incomplete semantic graphs, advancing privacy-preserving knowledge sharing on the Web. Shuo Yu 0001, Zhuoyang Han, Guoqing Han, Tao Tang 0007, Feng Ding 0004, Qiang Zhang 0008 |
WWW | 6 |
| 2026 | An adaptive multimodal semantic knowledge enhanced framework for sarcasm detection
Jing Dong 0009, Yu Sui, Qiang Zhang 0008, Hui Fang 0003, Gerald Schaefer, Rui Liu 0015, Xiaoyong Fang |
Expert Syst. Appl. | 3 |
| 2026 | Fine-grained face personalisation using a text-guided multi-attribute embedded diffusion model
Jing Dong 0009, Qiang Zhang 0008, Hui Fang 0003, Gerald Schaefer, Rui Liu 0015, Xiaoyong Fang |
Expert Syst. Appl. | 3 |
| 2026 | Adaptive patch-Gumbel filtering for long-horizon wind power forecasting
Xiaoming Ren, Chao Che, Qiang Zhang 0008 |
Expert Syst. Appl. | 5 |
| 2026 | Cluster-Driven Block-Flip Transformer With Dynamic Positional Learning for IoT-Based Wind Power ForecastingabstractIn Internet of Things (IoT)-enabled energy systems, accurate wind power forecasting is essential for intelligent scheduling and grid stability. However, meteorological and power time series often exhibit time-varying distributions and irregular regime shifts. This complexity complicates stable pattern learning and may cause models to overemphasize repeated local patterns. To address this issue, we propose a Cluster-Driven Block-Flip Transformer with Dynamic Positional Learning (CBFT) for wind power forecasting in IoT-based energy applications. CBFT includes two components. First, a block-based clustering and flipping module splits the input into fixed-length blocks and clusters them by statistical similarity. It then applies controlled randomized flipping to selected blocks as structured regularization to reduce order-sensitive dependencies. Second, a dynamic positional learning module uses multi-layer perceptron (MLP) embeddings to learn adaptive relative positional representations among blocks. This design remains effective under block-level order perturbations. Comprehensive evaluations on multiple real-world wind power and electricity datasets show that CBFT achieves lower MAE and MSE than state-of-the-art Transformer-based models. Chao Che, Qiang Zhang 0008 |
IEEE Internet Things J. | 4 |
| 2026 | UTA-Sign: Unsupervised thermal video augmentation via event-assisted traffic signage sketching
Yuqi Han, Songqian Zhang, Weijian Su, Jin-Li Suo, Qiang Zhang 0008 |
Pattern Recognit. | 7 |
| 2026 | Seq-IF: Sequentially Consistent Infrared-Visible Video Fusion Under Time-Varying Illumination for Perception EnhancementabstractInfrared–visible image fusion leverages the complementary strengths of both modalities to enhance visual perception in challenging environments. While image-level fusion has achieved promising results, extending it to video remains challenging due to temporal illumination variations, brightness flickering, and visual inconsistency caused by motion under non-uniform illumination. To address these issues, we propose Seq-IF, a sequential fusion framework that ensures visual consistency and structural fidelity when processing video sequences. The framework comprises a static–dynamic decoupling module for robust foreground–background separation. For the background, fusion is performed by selecting the frame with the highest contrast to ensure clarity and stability. For the dynamic objects, the pixel intensity is adaptively adjusted across consecutive frames by a lightweight MLP-based illumination-consistency fine-tuning module that performs online adaptation and dynamically optimizes brightness in response to scene changes. Later we introduce a spatial–frequency fusion module integrating multi-scale encoder and edge-guided decoder to ensure structural consistency. Extensive experiments demonstrate that the fusion results produced by Seq-IF outperform baseline methods in terms of clarity, detail preservation, and illumination stability, achieving smoother temporal transitions and enhanced perceptual quality. Furthermore, the effectiveness of Seq-IF is validated on downstream tasks, including object detection and optical flow estimation, highlighting its applicability to real-world scenarios. Yuqi Han, Zhihui Zheng, Weijian Su, Mingkai Wei, Liang Zhang 0031, Jin-Li Suo, Qiang Zhang 0008 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2026 | Evolutionary Content Generation via Multimodal LLM-Based Fitness EvaluationabstractEvolutionary algorithms (EAs) have gained prominence as a powerful optimization tool inspired by biological evolution, excelling in various complex domains. In the context of Generative Artificial Intelligence (GAI), EAs have shown promise in generating diverse, high-quality solutions. However, traditional EAs heavily rely on human-designed fitness functions, which may often lack flexibility and comprehensiveness for different GAI scenarios. Recently, the emergence of Large Language Models (LLMs) has opened new avenues for enhancing the evolutionary process. This paper proposes a formal framework named ECG-LFit, which utilizes an LLM (e.g., GPT-4-Turbo) for multimodal fitness evaluations. We validate the effectiveness of our framework using EAs (CMA-ES, MAP-Elites, and CMA-ME) in our case study on Super Mario game level generation. The results show that our framework improves the quality and playability of the generated levels. Additionally, user studies indicate that participants prefer the levels generated by the ECG-LFit framework, particularly regarding aesthetics, challenge, and playability. Furthermore, to enhance evaluation efficiency, we design a distilled model to simulate the scoring process of the LLM, enabling rapid and effective content evaluation in resource-constrained environments. Yaqing Hou, Zhaoping Yu, David M. Bossens, Qiang Zhang 0008, Yew-Soon Ong |
IEEE Trans. Evol. Comput. | 6 |
| 2026 | Inverse Feature Consistency Federated Unlearning for Vision-Language ModelabstractVision-Language Models (VLMs), with their advantages in vision and language processing, exhibit immense potential in mobile intelligent systems. Integrating federated learning with parameter-efficient fine-tuning of VLMs helps address data heterogeneity challenges. However, existing methods mainly focus on task-specific patterns, neglecting the impact of general features, such as background information and low-quality data, which weakens the model's ability to generalize when handling data from different sources and dealing with fluctuations in quality. To tackle these challenges, we propose Inverse Feature Consistency Federated Unlearning (IFCFU) for VLM, comprising three components: 1) Feature Consistency Federated Learning (FCFL) aligns fine-tuned features with pre-trained features through constraints to ensure the preservation of general features; 2) Pseudo-label Low-quality Data Detection (PLDD) identifies potential low-quality data through model quality assessment and pseudo-label generation; 3) Inverse Feature Consistency Unlearning (IFCU) distances low-quality data features from optimal model features to eliminate the negative impact and restores training with pseudo-labels. Evaluations on StanfordCars show that FCFL increased accuracy by 4.97% and 29.88% under normal data and low-quality data configurations, respectively. PLDD identified over 90.00% of low-quality data, while IFCU improved the global model's accuracy by 4.43% with 80% low-quality data. Zhenwei Wang 0005, Pengfei Wang 0013, Guangjie Han, Jianxin Zhang 0001, Muhammed Ameen, Qiang Zhang 0008 |
IEEE Trans. Mob. Comput. | 7 |
| 2026 | Generic-to-Personalised Learning for Multimodal Image Synthesis With Bidirectional Variational GANabstractMultimodal image synthesis, which predicts target-modality images from source-modality images, has garnered considerable attention in the field of clinical diagnosis. Both unidirectional and bidirectional multimodal image synthesis methods have been explored in the medical domain, however, unidirectional models heavily rely on paired images, while current bidirectional models typically overlook local image details due to their unsupervised training patterns. In this work, we propose a Bidirectional Variational Generative Adversarial Network (BVGAN) for multimodal image synthesis, which achieves high-quality bidirectional translations between any two modalities using only a limited number paired images. Firstly, BVGAN's generator incorporates a variational structure (VAS) to regularise the latent space for noise reduction. This regularisation imposes smoothness to the latent space, enabling BVGAN to produce high-quality, noise-free images. Secondly, a novel generic-to-personalised (GTP) learning strategy is introduced to train BVGAN and reduce its reliance on a large sets of paired images. GTP initially leverages an unsupervised learning model to capture the global mapping between two modalities using unpaired images from generic patients. It then applies a supervised learning model to refine the mapping for individual patient, enhancing image details. Finally, the GTP learning strategy along with VAS enables BVGAN to achieve state-of-the-art performance on two multi-modality medical datasets: Brain CTMRI and BRATS. Long Chen 0019, Xirui Dong, Jiangrong Shen, Lu Zhang 0053, Qi Xu 0008, Gang Pan 0001, Qiang Zhang 0008 |
IEEE Trans. Multim. | 7 |
| 2025 | EProtoSeg: An Explainable Prototype-Based Network with Multi-Scale Context for Brain Tumor SegmentationabstractAutomatic segmentation of brain tumors in magnetic resonance imaging (MRI) is essential for clinical decision support, yet it remains a challenging task due to the pronounced heterogeneity of gliomas and the often indistinct boundaries between their subregions. Moreover, the opaque, “black-box” nature of conventional deep learning models limits their adoption in clinical workflows, where interpretability is critical. To address these challenges, we propose EProtoSeg, a novel 3D segmentation network that synergistically integrates explainable prototypebased feature learning with adaptive multi-scale context aggregation. Specifically, EProtoSeg incorporates an Explainable Prototype Fusion (EPF) module in the decoder, which learns class-specific prototypes to guide voxel-wise classification and improve boundary precision. In parallel, an Adaptive Multi-Scale Context (AMSC) module is embedded in the skip connections to dynamically fuse fine spatial details and high-level semantic information across scales. A deep supervision strategy is further employed to enhance discriminability in ambiguous regions and ensure stable optimization. Extensive experiments on the BraTS 2020 and 2021 benchmarks demonstrate that EProtoSeg achieves state-of-the-art segmentation performance while offering interpretable feature representations, thereby improving both accuracy and clinical trust. The code is available at https://github.com/chenbn266/EProtoSeg Bonian Chen, Qiule Sun, Jianxin Zhang 0001, Bin Liu 0040, Qiang Zhang 0008 |
BIBM | 6 |
| 2025 | Dual-Prompt Learning with Cross-Modal Decoders for Few-Shot Whole Slide Image ClassificationabstractFew-shot learning offers a promising solution for computational pathology by alleviating the reliance on large an-notated datasets, but faces challenges from the high redundancy in whole slide images and underutilized cross-modal knowledge. Existing methods typically use foundation models only for pre-liminary feature extraction while employing fixed or single-level prompts that lack multi-scale pathological representation. To address these limitations, we propose a Hierarchical Vision-Text Prompt (H- VTP) framework that enables multi-level cross-modal interaction through GPT-4 generated Local Instance Prompts for patch-level morphological details and Global Semantic Prompts for slide-level diagnostic context. A dual-branch decoding mech-anism with Text-Guided-Patch Decoder and Patch-Augmented-Text Decoder facilitates closed-loop vision-text fusion, while a parameter-efficient adaptation strategy trains only lightweight prompts and adapters. Extensive experiments on three cancer subtype datasets demonstrate the superiority of H-VTP in few-shot WSI classification, confirming its effectiveness for clinical applications. Bingbing Zhang 0001, Wen Zhu, Bin Liu 0040, Jianxin Zhang 0001, Qiang Zhang 0008 |
BIBM | 7 |
| 2025 | Volumetric Axial-Shift Mamba U-Net for 3D Brain Tumor SegmentationabstractAccurate MRI brain tumor segmentation is significant for disease diagnosis and treatment. U-Nets are widely explored in automatic segmentation due to high accuracy and efficiency. Additionally, global semantic information establishes effectiveness in improving MRI brain tumor segmentation accuracy, and corresponding models have gained increasing attention. Inspired by state-space models excelling in long-range dependency modeling, we propose the Volumetric Axial Shift Mamba network (VASMambaU-Net) - a novel brain tumor segmentation method more suitable for 3D images. It integrates an improved Mamba module with large kernel convolution in U-Net to capture global tumor features. In VASMambaU-Net, we propose VASMamba that using a volumetric axial shift mechanism as the bottleneck to capture long-range dependencies of 3D brain tumor images. In addition, along with the conventional convolutional encoder, a large kernel convolutional encoder is supplemented to enhance the image receptive field, further improving the capability to capture global information, thereby enhancing the brain tumor segmentation performance. VASMambaU-Net achieves mean DSC values of$84.37 \%, 84.31 \%$and 91.07 % in the BraTS2019-2021 datasets, respectively. The corresponding mean HD95 values obtained in these three datasets are 4.67 mm, 13.51 mm, and 5.77 mm, respectively. These results demonstrate the competitiveness and effectiveness of VASMambaU-Net compared to state-of-the-art methods. Muqing Zhang, Bonian Chen, Yutong Han, Bin Liu 0040, Jianxin Zhang 0001, Qiang Zhang 0008 |
BIBM | 6 |
| 2025 | Deterministic-to-Stochastic Diverse Latent Feature Mapping for Human Motion SynthesisabstractHuman motion synthesis aims to generate plausible human motion sequences, which has raised widespread attention in computer animation. Recent score-based generative models (SGMs) have demonstrated impressive results on this task. However, their training process involves complex curvature trajectories, leading to unstable training process. In this paper, we propose a Deterministic-to-Stochastic Diverse Latent Feature Mapping (DSDFM) method for human motion synthesis. DSDFM consists of two stages. The first human motion reconstruction stage aims to learn the latent space distribution of human motions. The second diverse motion generation stage aims to build connections between the Gaussian distribution and the latent space distribution of human motions, thereby enhancing the diversity and accuracy of the generated human motions. This stage is achieved by the designed deterministic feature mapping procedure with DerODE and stochastic diverse output generation procedure with DivSDE. DSDFM is easy to train compared to previous SGMs-based methods and can enhance diversity without introducing additional training parameters. Through qualitative and quantitative experiments, DSDFM achieves state-of-the-art results surpassing the latest methods, validating its superiority in human motion synthesis. Hua Yu 0006, Weiming Liu 0005, Xu Gui 0001, Yaqing Hou, Yew-Soon Ong, Qiang Zhang 0008 |
CVPR | 6 |
| 2025 | Temporal Separation with Entropy Regularization for Knowledge Distillation in Spiking Neural NetworksabstractSpiking Neural Networks (SNNs), inspired by the human brain, offer significant computational efficiency through discrete spike-based information transfer. Despite their potential to reduce inference energy consumption, a performance gap persists between SNNs and Artificial Neural Networks (ANNs), primarily due to current training methods and inherent model limitations. While recent research has aimed to enhance SNN learning by employing knowledge distillation (KD) from ANN teacher networks, traditional distillation techniques often overlook the distinctive spatiotemporal properties of SNNs, thus failing to fully leverage their advantages. To overcome these challenge, we propose a novel logit distillation method characterized by temporal separation and entropy regularization. This approach improves existing SNN distillation techniques by performing distillation learning on logits across different time steps, rather than merely on aggregated output features. Furthermore, the integration of entropy regularization stabilizes model optimization and further boosts the performance. Extensive experimental results indicate that our method surpasses prior SNN distillation strategies, whether based on logit distillation, feature distillation, or a combination of both. Our project is available at https://github.com/yukairong/TSER. Kairong Yu, Chengting Yu, Tianqing Zhang, Xiaochen Zhao, Hongwei Wang 0001, Qiang Zhang 0008, Qi Xu 0008 |
CVPR | 7 |
| 2025 | STAA-SNN: Spatial-Temporal Attention Aggregator for Spiking Neural NetworksabstractSpiking Neural Networks (SNNs) have gained significant attention due to their biological plausibility and energy efficiency, making them promising alternatives to Artificial Neural Networks (ANNs). However, the performance gap between SNNs and ANNs remains a substantial challenge hindering the widespread adoption of SNNs. In this paper, we propose a Spatial-Temporal Attention Aggregator SNN (STAA-SNN) framework, which dynamically focuses on and captures both spatial and temporal dependencies. First, we introduce a spike-driven self-attention mechanism specifically designed for SNNs. Additionally, we pioneeringly incorporate position encoding to integrate latent temporal relationships into the incoming features. For spatial-temporal information aggregation, we employ step attention to selectively amplify relevant features to variant steps. Finally, we implement a time-step random dropout strategy to avoid local optima. The framework demonstrates exceptional performance across diverse datasets and exhibits strong generalization capabilities. Notably, STAA-SNN achieves state-of-the-art results on neuromorphic datasets CIFAR10-DVS of 82.10% and with performances of 97.14%, 82.05% and 70.40% on the static datasets CIFAR-10, CIFAR-100 and ImageNet, respectively. Furthermore, this model exhibits improved performance ranging from 0.33% to 2.80% with fewer time steps. Tianqing Zhang, Kairong Yu, Xian Zhong, Hongwei Wang 0001, Qi Xu 0008, Qiang Zhang 0008 |
CVPR | 6 |
| 2025 | Factor Graph-based Interpretable Neural NetworksabstractComprehensible neural network explanations are foundations for a better understanding of decisions, especially when the input data are infused with malicious perturbations. Existing solutions generally mitigate the impact of perturbations through adversarial training, yet they fail to generate comprehensible explanations under unknown perturbations. To address this challenge, we propose AGAIN, a factor graph-based interpretable neural network, which is capable of generating comprehensible explanations under unknown perturbations. Instead of retraining like previous solutions, the proposed AGAIN directly integrates logical rules by which logical errors in explanations are identified and rectified during inference. Specifically, we construct the factor graph to express logical rules between explanations and categories. By treating logical rules as exogenous knowledge, AGAIN can identify incomprehensible explanations that violate real-world logic. Furthermore, we propose an interactive intervention switch strategy rectifying explanations based on the logical guidance from the factor graph without learning perturbations, which overcomes the inherent limitation of adversarial training-based methods in defending only against known perturbations. Additionally, we theoretically demonstrate the effectiveness of employing factor graph by proving that the comprehensibility of explanations is strongly correlated with factor graph. Extensive experiments are conducted on three datasets and experimental results illustrate the superior performance of AGAIN compared to state-of-the-art baselines. Yicong Li 0006, Kuanjiu Zhou, Shuo Yu 0001, Qiang Zhang 0008, Renqiang Luo, Xiaodong Li 0001, Feng Xia 0001 |
ICLR | 4 |
| 2025 | Self-cross Feature based Spiking Neural Networks for Efficient Few-shot LearningabstractDeep neural networks (DNNs) excel in computer vision tasks, especially, few-shot learning (FSL), which is increasingly important for generalizing from limited examples. However, DNNs are computationally expensive with scalability issues in real world. Spiking Neural Networks (SNNs), with their event-driven nature and low energy consumption, are particularly efficient in processing sparse and dynamic data, though they still encounter difficulties in capturing complex spatiotemporal features and performing accurate cross-class comparisons. To further enhance the performance and efficiency of SNNs in few-shot learning, we propose a few-shot learning framework based on SNNs, which combines a self-feature extractor module and a cross-feature contrastive module to refine feature representation and reduce power consumption. We apply the combination of temporal efficient training loss and InfoNCE loss to optimize the temporal dynamics of spike trains and enhance the discriminative power. Experimental results show that the proposed FSL-SNN significantly improves the classification performance on the neuromorphic dataset N-Omniglot, and also achieves competitive performance to ANNs on static datasets such as CUB and miniImageNet with low power consumption. Qi Xu 0008, Junyang Zhu, Dongdong Zhou, Jiangrong Shen, Qiang Zhang 0008 |
ICML | 7 |
| 2025 | EFormer: An Effective Edge-based Transformer for Vehicle Routing ProblemsabstractRecent neural heuristics for the Vehicle Routing Problem (VRP) primarily rely on node coordinates as input, which may be less effective in practical scenarios where real cost metrics—such as edge-based distances—are more relevant. To address this limitation, we introduce EFormer, an Edge-based Transformer model that uses edge as the sole input for VRPs. Our approach employs a precoder module with a mixed-score attention mechanism to convert edge information into temporary node embeddings. We also present a parallel encoding strategy characterized by a graph encoder and a node encoder, each responsible for processing graph and node embeddings in distinct feature spaces, respectively. This design yields a more comprehensive representation of the global relationships among edges. In the decoding phase, parallel context embedding and multi-query integration are used to compute separate attention mechanisms over the two encoded embeddings, facilitating efficient path construction. We train EFormer using reinforcement learning in an autoregressive manner. Extensive experiments on the Traveling Salesman Problem (TSP) and Capacitated Vehicle Routing Problem (CVRP) reveal that EFormer outperforms established baselines on synthetic datasets, including large-scale and diverse distributions. Moreover, EFormer demonstrates strong generalization on real-world instances from TSPLib and CVRPLib. These findings confirm the effectiveness of EFormer’s core design in solving VRPs. Dian Meng, Zhiguang Cao, Yaoxin Wu, Yaqing Hou, Hong-Wei Ge, Qiang Zhang 0008 |
IJCAI | 6 |
| 2025 | Preference-based Deep Reinforcement Learning for Historical Route EstimationabstractRecent Deep Reinforcement Learning (DRL) techniques have advanced solutions to Vehicle Routing Problems (VRPs). However, many of these methods focus exclusively on optimizing distance-oriented objectives (i.e., minimizing route length), often overlooking the implicit drivers' preferences for routes. These preferences, which are crucial in practice, are challenging to model using traditional DRL approaches. To address this gap, we propose a preference-based DRL method characterized by its reward design and optimization objective, which is specialized to learn historical route preferences. Our experiments demonstrate that the method aligns generated solutions more closely with human preferences. Moreover, it exhibits strong generalization performance across a variety of instances, offering a robust solution for different VRP scenarios. Boshen Pan, Yaoxin Wu, Zhiguang Cao, Yaqing Hou, Guangyu Zou, Qiang Zhang 0008 |
IJCAI | 6 |
| 2025 | SFPA-Net: Second-order Feature Enhancement Prototype Aggregation Network for MRI Brain Tumor SegmentationabstractAccurate segmentation of brain tumors faces challenges of difficult edge voxel recognition and lack of interpretability. Prototype network demonstrates significant potential for brain tumor segmentation owing to its high interpretability. However, accurately extracting typical tumor features using the sample mean as a class prototype is challenging, and directly extracting prototypes from backbone-processed features ignores feature dependencies. To address these challenges, we propose the Second-order Feature Prototype Aggregation Network (SFPANet) for brain tumor segmentation. SFPA-Net introduces Covariance Feature Enhancement Module (CFEM) and Weighted Prototype Integration Module (WPIM). CFEM enhances feature representation by computing the covariance (second-order) matrix and applying power normalization. WPIM leverages enhanced features to extract tumor category characteristics and aggregates these features to represent tumor prototypes. Additionally, to enhance gradient propagation, SFPA-Net introduces deep supervision signal to optimize the training process. The Dice Similarity Coefficients of SFPA-Net on the BraTS2020 and BraTS2021 training datasets for the whole tumor, tumor core, and enhanced tumor are 91.62% / 93.50%, 89.58% / 92.55%, and 77.76% / 86.26%, respectively. Furthermore, the 95% Hausdorff Distances are 6.81 / 4.98, 5.98 / 5.86, and 13.1 / 10.68, respectively, surpassing existing brain tumor segmentation methods. Zhenwei Wang 0005, Jianxin Zhang 0001, Qiang Zhang 0008 |
IJCNN | 6 |
| 2025 | Eliminating Poor-Quality Data Impacts from Multiple Participants with Federated UnlearningabstractFederated unlearning (FUL) facilitates targeted unlearning of data-specific artifacts from trained federated learning (FL) models. Existing methodologies primarily emphasize individual client demands at a time, such as enforcing the “right to be forgotten” or mitigating poor-quality data contributions from a single client. These strategies frequently rely on client-side computational engagement during unlearning. However, these approaches often neglect situations in which it is crucial to unlearn the contributions of multiple participants at a time from a global model. Therefore, we introduce a novel FUL method named “FULClean” to parallelly erase the adverse impacts of poor-quality data from multiple participants without using client resources and maintaining global model performance. FULClean employs a novel Contrastive Threshold-based Contribution Rectification (CTCR) mechanism, that (1) identifies and classifies the local models into unaffected and potentially affected local models, (2) computes the dynamic contribution threshold for each potentially affected model based on layer-wise parameter divergence, and (3) selectively replaces or remove contributions based on their affected ratio. FULClean can categorize and swiftly eliminate the impacts of poor-quality data from multiple participants on the global model. This can be accomplished without communication and utilization of the client's resources. Extensive experiments are also conducted on four distinct datasets with two different models to showcase the effectiveness and efficiency of FULClean. Pengfei Wang 0013, Muhammed Ameen, Mingshu Zhao, Pai Liu, Qiang Zhang 0008 |
IWQoS | 5 |
| 2025 | Divide and Conquer: The First Step Towards Adaptable Internet of ModelsabstractEnhancing model performance on low-capacity devices remains a significant challenge in the Internet of Models (IoM). Inspired by the efficiency of disentangled representation learning, this study proposes a model decomposition approach to improve response time on low-capacity devices while maintaining accuracy, through collaborative learning with models deployed on higher-capacity devices. First, we introduce a divide-and-conquer strategy that decomposes and learns knowledge within each data sample. This enables lightweight sub-models, tailored to specific knowledge components, to respond efficiently on low-capacity devices in the IoM system. Second, we design a progressive crossfusion mechanism to promote mutual enhancement among these knowledge-specific sub-models. Third, we optimize the learning process by aligning the updates of these sub-models with those of the models on higher-capacity devices. This collaborative optimization is guided via knowledge distillation during model aggregation. Experimental results on image classification tasks demonstrate that our approach reduces response time by 9.87 % to 60.31 % without compromising accuracy. Pengfei Wang 0013, Feiye Ye, Junxiang Zhang, Pai Liu, Mingshu Zhao, Liang Zhang 0031, Qiang Zhang 0008 |
IWQoS | 8 |
| 2025 | Cross Attention Guided Multimodal Network for Video Action Recognition
Bingbing Zhang 0001, Yongqi Li 0014, Jianxin Zhang 0001, Qiang Zhang 0008 |
PRCV (7) | 5 |
| 2025 | High-Order Multimodal Multi-task Video Action Recognition
Bingbing Zhang 0001, Yongqi Li 0014, Jianxin Zhang 0001, Qiang Zhang 0008 |
PRCV (7) | 5 |
| 2025 | Few-Shot Action Recognition Based on Visual-Language Prototype Hierarchical Temporal Enhancement
Bingbing Zhang 0001, Yuanchen Ma, Jianxin Zhang 0001, Qiang Zhang 0008 |
PRCV (7) | 5 |
| 2025 | HeRIF: A Mixture-of-Experts Framework for Infrared and Visible Image Fusion with Heterogeneous Resolutions
Songqian Zhang, Weijian Su, Yuqi Han, Jin-Li Suo, Qiang Zhang 0008 |
PRCV (5) | 6 |
| 2025 | GAM: A Generative Autoencoder for Diverse Human Motion Prediction
Jiapeng Bai, Hua Yu 0006, Yaqing Hou, Qiang Zhang 0008 |
PRICAI | 4 |
| 2025 | MSPL: Multimodal Statistical Prompt Learning for New Energy Equipment Defect RecognitionabstractMonitoring renewable energy devices is crucial for the timely detection of faults and the improvement of system stability, making it a key component of the Internet of Things (IoT) ecosystem. However, existing intelligent algorithms and IoT methods rely on edge devices with limited computational capacity, which requires balancing real-time performance and accuracy. Additionally, most methods require initial training on large-scale data using cloud platforms or high-performance servers, leading to high initial costs. To address these challenges, we propose multimodal statistical prompt learning (MSPL) for new energy equipment defect recognition on IoT edge devices. This method enables rapid learning with only a few samples, avoiding the need for large-scale centralized training. Specifically, MSPL uses the pretrained contrastive language-image pretraining model as its backbone, leveraging text-based conceptual information to enhance the understanding of visual inputs. A statistical query module is implemented at the end of the backbone to extract distinctive features from the outputs, integrating these features with soft prompts to customize them for the defect recognition task. Since the learnable parameters are limited to soft prompts added at the end of the backbone, MSPL restricts gradient backpropagation to this point. This improves parameter and memory efficiency, making it more suitable for scenarios with limited computational capacity on IoT edge devices. Experimental results of MSPL on two renewable energy equipment defect datasets and edge devices indicate that it meets real-time processing requirements while maintaining high accuracy, outperforming other methods. Zhenwei Wang 0005, Pengfei Wang 0013, Guangjie Han, Jianxin Zhang 0001, Guangjie Fan, Qiang Zhang 0008 |
IEEE Internet Things J. | 8 |
| 2025 | Few-Shot Defect Recognition for New Energy Equipment via Multimodal HarmonyabstractInternet of Things (IoT) technologies have been applied to fault detection in new energy equipment, which is crucial for ensuring the stable operation of energy systems. However, existing approaches typically rely on large amounts of labeled data to train intelligent algorithms, which are difficult to obtain in real-world. To address this challenge, this paper proposes a novel multi-modal few-shot defect recognition framework for new energy equipment, enabling data-efficient defect recognition in real-world scenarios through multi-modal harmony. Unlike previous multi-modal few-shot methods, it eliminates the necessity of pairwise similarity calculations, thereby simplifying the training and inference processes. Our approach trains a shared classifier through text and visual modality features integration. Initially, we extract feature vectors using text and visual encoders and then map them to common feature space. Subsequently, the vectors of these two modalities are used to train a shared classifier, which aids in the simultaneous learning of corresponding visual representations and conceptual information. Furthermore, to enhance interaction between modalities, the feature vectors of the text prompts are used to initialize the classifier weights, thereby promoting cross-modal consistency. Experiments on datasets of wind turbine blades and solar cells show that combining the two modalities improves recognition accuracy by up to 19.13% and 28.52% compared to other multi-modal few-shot methods. Despite its reliance on prompt quality, the approach provides an effective and scalable solution for defect recognition. Zhenwei Wang 0005, Pengfei Wang 0013, Mohammad S. Obaidat, Tianbao Yang, Jintao Zheng, Jianxin Zhang 0001, Qiang Zhang 0008 |
IEEE Internet Things J. | 8 |
| 2025 | Efficient hypergraph collective influence maximization in cascading processes based on general threshold model
Xilong Qu, Qiang Zhang 0008, Yinchao Yang, Xirong Xu, Wenbin Pei, Renquan Zhang |
Inf. Sci. | 2 |
| 2025 | Meta doubly robust: Debiasing CVR prediction via meta-learning with a small amount of unbiased data
Pengkun Li, Xiangrong Tong, Qiang Zhang 0008 |
Knowl. Based Syst. | 4 |
| 2025 | A sequential mixing fusion network for enhanced feature representations in multimodal sentiment analysis
Qiang Zhang 0008, Jing Dong 0009, Hui Fang 0003, Gerald Schaefer, Rui Liu 0015 |
Knowl. Based Syst. | 2 |
| 2025 | DNA sequence design model for multi-scene fusion
Yanfen Zheng, Yaqing Hou, Qiang Zhang 0008, Xiaopeng Wei |
Neural Comput. Appl. | 5 |
| 2025 | GTIGNet: Global Topology Interaction Graphormer Network for 3D hand pose estimation
Wanshu Fan, Cong Wang 0018, Shixi Wen, Xin Yang 0011, Qiang Zhang 0008, Xiaopeng Wei |
Neural Networks | 6 |
| 2025 | Stable DNA Storage Encoding Scheme Based on Repeating Substring TreeabstractDNA storage is considered to be a promising storage media in the current era of data explosion. DNA encoding is the beginning of the DNA storage process and lays the foundation for subsequent processes. However, many encoding methods suffer from low encoding rate, do not satisfy important constraints, or have insufficient sequence stability. To address these issues and improved sequences stability, this paper proposes a novel approach called the Repeating Substring Tree Encoding (RSTE) method. The method begins by applying the Longest Substring Backtracking Method (LSBM) to identify the longest repeated substrings within the binary file. These substrings are then encoded into compact DNA motifs using Huffman encoding. In contrast to the ideal coding density of 2 bits per nucleotide (2 bit/nt) targeted by previous studies, RSTE enhances the encoding rate by 13% through efficient utilization of repeated substrings. Furthermore, the DNA sequences generated by the RSTE method successfully meet three biological constraints: run-length limitation, GC content balance and end constraints. The experimental results of minimum free energy and melting temperature indicate that the stability of the sequences encoded by RSTE is also greatly improved. A series of experiments showed that the sequences encoded by RSTE have a higher coding rate, satisfy constraints, and are more stable. Jieqiong Wu, Penghao Wang 0001, Yanfen Zheng, Bin Wang 0005, Qiang Zhang 0008, Pan Zheng 0001 |
IEEE Trans. Comput. Biol. Bioinform. | 5 |
| 2025 | A Spatio-Temporal Continuous Network for Stochastic 3D Human Motion PredictionabstractStochastic Human Motion Prediction (HMP) has received increasing attention due to its wide applications. Despite the rapid progress in generative fields, existing methods often face challenges in learning continuous temporal dynamics and predicting stochastic motion sequences. They tend to overlook the flexibility inherent in complex human motions and are prone to mode collapse. To alleviate these issues, we propose a novel method called STCN, for stochastic and continuous human motion prediction, which consists of two stages. Specifically, in the first stage, we propose a spatio-temporal continuous network to generate smoother human motion sequences. In addition, the anchor set is innovatively introduced into the stochastic HMP task to prevent mode collapse, which refers to the potential human motion patterns. In the second stage, STCN endeavors to acquire the Gaussian mixture distribution (GMM) of observed motion sequences with the aid of the anchor set. It also focuses on the probability associated with each anchor, and employs the strategy of sampling multiple sequences from each anchor to alleviate intra-class differences in human motions. Experimental results on two widely-used datasets (Human3.6M and HumanEva-I) demonstrate that our model obtains competitive performance on both diversity and accuracy. Hua Yu 0006, Yaqing Hou, Xu Gui 0001, Shanshan Feng 0001, Qiang Zhang 0008 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | Cooperative Multiagent Learning and Exploration With Min-Max Intrinsic MotivationabstractIn the field of multiagent reinforcement learning (MARL), the ability to effectively explore unknown environments and collect information and experiences that are most beneficial for policy learning represents a critical research area. However, existing work often encounters difficulties in addressing the uncertainties caused by state changes and the inconsistencies between agents' local observations and global information, which presents significant challenges to coordinated exploration among multiple agents. To address this issue, this article proposes a novel MARL exploration method with Min-Max intrinsic motivation (E2M) that promotes the learning of joint policies of agents by introducing surprise minimization and social influence maximization. Since the agent is subject to unstable state changes in the environment, we introduce surprise minimization by computing state entropy to encourage the agents to cope with more stable and familiar situations. This method enables surprise estimation based on the low-dimensional representation of states obtained from random encoders. Furthermore, to prevent surprise minimization from leading to conservative policies, we introduce mutual information between agents' behaviors as social influence. By maximizing social influence, the agents are encouraged to interact to facilitate the emergence of cooperative behavior. The performance of our proposed E2M is testified across a range of popular StarCraft II and Multiagent MuJoCo tasks. Comprehensive results demonstrate its effectiveness in enhancing the cooperative capability of the multiple agents. Yaqing Hou, Haiyin Piao, Yifeng Zeng, Yew-Soon Ong, Yaochu Jin, Qiang Zhang 0008 |
IEEE Trans. Cybern. | 7 |
| 2025 | DG-SMOTE: A Distance-Angle-Based Genetic Synthetic Minority Over-Sampling Technique for Unbalanced Data LearningabstractMany real-world applications often generate unbalanced data. Learning from such data may lead to biased classifiers that perform poorly on the class of interest. Oversampling methods have been shown to be effective in rebalancing unbalanced data to help classifiers avoid performance bias. However, many existing oversampling methods rely on a predesigned linear model structure and the neighborhood information of an original instance. This may lead to the generation of noisy instances when the original data has noise. In this study, we develop a novel oversampling method in which genetic programming is introduced to automatically select good-quality instances and evolve a model structure that combines the selected instances to create a new instance. In the proposed oversampling method, an individual is used to represent a generated instance, which is evaluated by the fitness function designed based on the Euclidean distance and the cosine theorem. In the experiments, we examine the effectiveness of the proposed oversampling method in assisting different types of classifiers to solve the issue of class imbalance, and compare it with popular sampling methods in unbalanced classification. The results have been analyzed comprehensively, indicating that the new method successfully addressed the class imbalance issue by generating a group of good-quality instances for the minority class and outperformed the compared sampling methods in almost all cases. Wenbin Pei, Yuyang Cui, Bing Xue 0001, Mengjie Zhang 0001, Jiqing Zhang, Yaqing Hou, Guangyu Zou, Qiang Zhang 0008 |
IEEE Trans. Evol. Comput. | 8 |
| 2025 | MSTF: enhancing long-term forecasting with multi-scale temporal fusion in time series forecasting
Yunjiong Liu, Chao Che, Qiang Zhang 0008 |
J. Supercomput. | 5 |
| 2025 | MLFuse: Multi-Scenario Feature Joint Learning for Multi-Modality Image FusionabstractMulti-modality image fusion (MMIF) entails synthesizing images with detailed textures and prominent objects. Existing methods tend to use general feature extraction to handle different fusion tasks. However, these methods have difficulty breaking fusion barriers across various modalities owing to the lack of targeted learning routes. In this work, we propose a multi-scenario feature joint learning architecture, MLFuse, that employs the commonalities of multi-modality images to deconstruct the fusion progress. Specifically, we construct a cross-modal knowledge reinforcing network that adopts a multipath calibration strategy to promote information communication between different images. In addition, two professional networks are developed to maintain the salient and textural information of fusion results. The spatial-spectral domain optimizing network can learn the vital relationship of the source image context with the help of spatial attention and spectral attention. The edge-guided learning network utilizes the convolution operations of various receptive fields to capture image texture information. The desired fusion results are obtained by aggregating the outputs from the three networks. Extensive experiments demonstrate the superiority of MLFuse for infrared-visible image fusion and medical image fusion. The excellent results of downstream tasks (i.e., object detection and semantic segmentation) further verify the high-quality fusion performance of our method. Jia Lei 0001, Jiawei Li 0016, Jinyuan Liu 0001, Bin Wang 0005, Shihua Zhou, Qiang Zhang 0008, Xiaopeng Wei, Nikola K. Kasabov |
IEEE Trans. Multim. | 6 |
| 2025 | DivDiff: A Conditional Diffusion Model for Diverse Human Motion PredictionabstractDiverse human motion prediction (HMP) aims to predict multiple plausible future motions given an observed human motion sequence. It is a challenging task due to the diversity of potential human motions while ensuring an accurate description of future human motions. Current solutions are either low-diversity or limited in expressiveness. Recent denoising diffusion probabilistic models (DDPM) demonstrate promising performance in various generative tasks. However, introducing DDPM directly into diverse HMP incurs some issues. While DDPM can enhance the diversity of potential human motion patterns, the predicted human motions gradually become implausible over time due to significant noise disturbances in the forward process of DDPM. This phenomenon leads to the predicted human motions being unrealistic, seriously impacting the quality of predicted motions and restricting their practical applicability in real-world scenarios. To alleviate this, we propose a novel conditional diffusion-based generative model, called DivDiff, to predict more diverse and realistic human motions. Specifically, the DivDiff employs DDPM as our backbone and incorporates Discrete Cosine Transform (DCT) and Transformer mechanisms to encode the observed human motion sequence as a condition to instruct the reverse process of DDPM. More importantly, we design a diversified reinforcement sampling function (DRSF) to enforce human skeletal constraints on the predicted human motions. DRSF utilizes the acquired information from human skeletal as prior knowledge, thereby reducing significant disturbances introduced during the forward process. Extensive results received in the experiments on two widely-used datasets (Human3.6M and HumanEva-I) demonstrate that our model obtains competitive performance on both diversity and accuracy. Hua Yu 0006, Yaqing Hou, Wenbin Pei, Yew-Soon Ong, Qiang Zhang 0008 |
IEEE Trans. Multim. | 5 |
| 2025 | Robust Sensory Information Reconstruction and Classification With Augmented SpikesabstractSensory information recognition is primarily processed through the ventral and dorsal visual pathways in the primate brain visual system, which exhibits layered feature representations bearing a strong resemblance to convolutional neural networks (CNNs), encompassing reconstruction and classification. However, existing studies often treat these pathways as distinct entities, focusing individually on pattern reconstruction or classification tasks, overlooking a key feature of biological neurons, the fundamental units for neural computation of visual sensory information. Addressing these limitations, we introduce a unified framework for sensory information recognition with augmented spikes. By integrating pattern reconstruction and classification within a single framework, our approach not only accurately reconstructs multimodal sensory information but also provides precise classification through definitive labeling. Experimental evaluations conducted on various datasets including video scenes, static images, dynamic auditory scenes, and functional magnetic resonance imaging (fMRI) brain activities demonstrate that our framework delivers state-of-the-art pattern reconstruction quality and classification accuracy. The proposed framework enhances the biological realism of multimodal pattern recognition models, offering insights into how the primate brain visual system effectively accomplishes the reconstruction and classification tasks through the integration of ventral and dorsal pathways. Qi Xu 0008, Sibo Liu, Xuming Ran, Jiangrong Shen, Huajin Tang, Jian K. Liu, Gang Pan 0001, Qiang Zhang 0008 |
IEEE Trans. Neural Networks Learn. Syst. | 9 |
| 2025 | Formulating and Representing Multiagent Systems With HypergraphsabstractGraph-learning methods, especially graph neural networks (GNNs), have shown remarkable effectiveness in handling non-Euclidean data and have achieved great success in various scenarios. Existing GNNs are primarily based on message-passing schemes, that is, aggregating information from neighboring nodes. However, the diversity and complexity of complex systems from real-world circumstances are not sufficiently taken into account. In these cases, the individual should be treated as an agent, with the ability to perceive their surroundings and interact with other individuals, rather than just be viewed as nodes in existing graph approaches. Additionally, the pairwise interactions used in existing methods also lack the expressiveness for the higher-order complex relations among multiple agents, thus limiting the performance in various tasks. In this work, we propose a Multiagent Hypergraph Force-learning method dubbed MHGForce. First, we formalize the multiagent system (MAS) and illustrate its connection to graph learning. Then, we propose a generalized multiagent hypergraph-learning framework. In this framework, we integrate message-passing and force-based interactions to devise a pluggable method. The method empowers graph approaches to excel in downstream tasks while effectively maintaining structural information in the representations. Experimental results on the Cora, Citeseer, Cora-CA, Zoo, and NTU2012 datasets in node classification demonstrate the effectiveness and generality of our proposed method. We also discuss the characteristics of the MHGForce and explore its role through parametric analysis and visualization. Finally, we give a discussion, conclude our work, and propose future directions. Shuo Yu 0001, Huafei Huang 0001, Yanming Shen, Pengfei Wang 0013, Qiang Zhang 0008, Ke Sun 0011, Honglong Chen |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2025 | CaGE: A Causality-inspired Graph Neural Network Explainer for Recommender SystemsabstractGenerating post hoc causal explanations for graph neural network-based recommender systems is vital for enhancing the credibility and interpretability of recommendations. Existing model-agnostic explainers primarily capture statistical correlations between topological information and recommendation outcomes. However, they often fail to identify true causal relationships due to their model-agnostic design and the challenges posed by heterogeneous graph structures. To address these limitations, we propose a causality-inspired graph neural network explainer for recommender systems, namely CaGE, which generates explanations reflecting causality in recommendation scenarios without accessing the internal parameters of the recommender system. Unlike previous explainers that rely on correlation-based learning, CaGE leverages heterogeneous interventional distributions to eliminate backdoor paths of non-causal variables in the structural causal model of the recommendation task, ensuring causation is accurately captured. Specifically, CaGE incorporates backdoor adjustment based on heterogeneous interventional distributions and causal contrastive learning to optimize a set of heterogeneous soft masks that disentangle causation from non-causation. Additionally, a causality-inspired meta-path search strategy is employed to represent causation as paths between users and recommended items, further enhancing explanation readability. Extensive experiments are conducted on three recommendation datasets, and the experimental results illustrate the superior fidelity of CaGE as compared to state-of-the-art baselines. Shuo Yu 0001, Yicong Li 0006, Shuo Wang 0040, Tao Tang 0007, Qiang Zhang 0008, Ivan Lee 0001, Feng Xia 0001 |
ACM Trans. Inf. Syst. | 5 |
| 2025 | Fast-Response Edge Caching Scheme for Graph DataabstractBy deploying distributed storage space on edge servers, mobile edge networks significantly enhance computation and transmission efficiency for wireless tasks. The selection of an appropriate caching policy not only optimizes bandwidth utilization but also alleviates network congestion. Given the intricate connectivity and vast data volume in edge computing, coupled with users’ demand for rapid response times, proposing a cache solution that closely matches data attributes becomes imperative to enhance overall efficiency. In this paper, we introduce RECG, a high-speed edge caching scheme designed specifically for graph data, leveraging the intricate data connectivity. RECG generates query graphs from edge servers to ensure swift and accurate identification of popular nodes. Additionally, we introduce a rapid hot-spot propagation partitioning technique to optimize the partitioning process, increasing the hit rate of partitioned subgraphs in the cache while reducing runtime. Our experimental evaluation, conducted on real-world datasets, compares RECG with specialized edge caching graph partitioning algorithms such as LGPE and other baseline algorithms. The thorough experimental results demonstrate the advantages of the proposed algorithm in terms of cache hit rate and processing delay. Pengfei Wang 0013, Yuqi Han, Feiye Ye, Qiang Zhang 0008 |
IEEE Trans. Netw. | 5 |
| 2025 | Speed up Federated Unlearning With Temporary Local ModelsabstractFederated unlearning (FUL) is a solution aimed at addressing the problem of removing data contributions from trained federated learning (FL) models. Existing FUL methods only focus on iterative unlearning of clients’ contributions and fail to perform unlearning in scenarios where multiple clients request to remove their data at a time. Additionally, FUL still needs to address issues, including convergence speed, maintaining the global model’s performance, and parallel unlearning to expedite the unlearning process. To fill this gap, we introduce Federated Clients Forgetting (FedCF), a fast and accurate FUL method that can eliminate single client contributions similar to existing methods, eliminate multiple clients’ contributions on the global model parallelly, ensure the performance of the unlearned global model, and reduce the unlearning time. The key idea is to construct a temporary model by extracting knowledge from the remaining clients’ updates and adding it to the corresponding parameters of the initial global model and then leverage a temporary model to reconstruct the unlearned global model. Extensive experiments on three benchmark datasets, FedCF demonstrates its efficiency and effectiveness for single client contribution unlearning, achieving an average time efficiency of 8.3x, 6.5x, and 4.1x over existing methods FedRetrain, FedEraser, and FUL with knowledge distillation, respectively. Additionally, FedCF showcases the time efficiency and performance guarantee after unlearning the contributions of multiple clients in parallel. Muhammed Ameen, Pengfei Wang 0013, Weijian Su, Xiaopeng Wei, Qiang Zhang 0008 |
IEEE Trans. Sustain. Comput. | 5 |
| 2024 | Exploiting Polarized Material Cues for Robust Car DetectionabstractCar detection is an important task that serves as a crucial prerequisite for many automated driving functions. The large variations in lighting/weather conditions and vehicle densities of the scenes pose significant challenges to existing car detection algorithms to meet the highly accurate perception demand for safety, due to the unstable/limited color information, which impedes the extraction of meaningful/discriminative features of cars. In this work, we present a novel learning-based car detection method that leverages trichromatic linear polarization as an additional cue to disambiguate such challenging cases. A key observation is that polarization, characteristic of the light wave, can robustly describe intrinsic physical properties of the scene objects in various imaging conditions and is strongly linked to the nature of materials for cars (e.g., metal and glass) and their surrounding environment (e.g., soil and trees), thereby providing reliable and discriminative features for robust car detection in challenging scenes. To exploit polarization cues, we first construct a pixel-aligned RGB-Polarization car detection dataset, which we subsequently employ to train a novel multimodal fusion network. Our car detection network dynamically integrates RGB and polarization features in a request-and-complement manner and can explore the intrinsic material properties of cars across all learning samples. We extensively validate our method and demonstrate that it outperforms state-of-the-art detection methods. Experimental results show that polarization is a powerful cue for car detection. Our code is available at https://github.com/wind1117/AAAI24-PCDNet. Wen Dong 0008, Haiyang Mei, Ziqi Wei 0001, Ao Jin, Sen Qiu, Qiang Zhang 0008, Xin Yang 0011 |
AAAI | 6 |
| 2024 | PM2: A New Prompting Multi-modal Model Paradigm for Few-shot Medical Image ClassificationabstractFew-shot learning has become a key technical solution for addressing the challenges of limited data and difficult annotation acquisition in medical image classification. However, relying solely on a single image modality proves inadequate for capture conceptual categories. This paper proposes a novel medical image classification paradigm based on a multi-modal foundation model, called PM2. In addition to the image modality, PM2introduces supplementary text input (prompt) to further describe images or conceptual categories and facilitate cross-modal few-shot learning. We empirically studied five different prompting schemes under this new paradigm. Furthermore, linear probing in multi-modal models only takes class token as input, ignoring the rich statistical data contained in high-level visual tokens. Therefore, we alternately perform linear classification on the feature distributions of visual tokens and class token. To effectively extract statistical information, we use global covariance pool with efficient matrix power normalization to aggregate the visual tokens. We then combine two classification heads: one for handling image class token and prompt representations encoded by the text encoder, and the other for classifying the feature distributions of visual tokens. Experiments on two medical datasets demonstrate that regardless of the prompting scheme, our method PM2outperforms its counterparts, achieving state-of-the-art performance. Zhenwei Wang 0005, Qiule Sun, Bingbing Zhang 0001, Weijian Su, Pengfei Wang 0013, Jianxin Zhang 0001, Qiang Zhang 0008 |
BIBM | 7 |
| 2024 | Application of Static Virus Spread Algorithm in Base-Balanced DNA Fragment OptimizationabstractDNA has found applications in a diverse array of fields such as computing, medical diagnosis, and circuits. Designing DNA fragments that meet specific requirements is crucial for ensuring the smooth execution of tasks in these domains. However, the conventional approach of combining constraints with evolutionary algorithms often encounters challenges like orthogonality and thermodynamic instability. This poses risks to the control of reaction processes. To tackle these challenges, we have undertaken two tasks in this study: expanding constraint sets and innovating evolutionary algorithms. Firstly, we propose a base balance strategy aimed at enhancing the thermodynamic properties within DNA fragment groups by diversifying neighboring base combinations. The addition of this strategy achieves a minimum variance of 0.03 for the melting temperature while ensuring orthogonality. It represents a significant improvement compared to previous results and reduces the complexity of controlling reactions. Secondly, our static virus spread algorithm optimizes target DNA fragments by base mutations and virus amplification. Simultaneously,it demonstrates good performance across 23 benchmark functions, highlighting its optimization potential. This work is expected to further refine theoretical shortcomings and offer a convenient tool for DNA fragment optimization. Yanfen Zheng, Xin Liu 0120, Bin Wang 0005, Qiang Zhang 0008 |
BIBM | 7 |
| 2024 | Similar Locality Based Transfer Evolutionary Optimization for Minimalistic AttacksabstractDeep neural networks are powerful and popular learning models; however, recent studies have shown that deep neural network-based policies are susceptible to deception by adversarial attacks. A minimalistic attack is a specialized form of adversarial attack that aims to accomplish successful attacks at the lowest possible cost. Recently, transfer optimization algorithms have been applied to deceive previously trained policies by acquiring knowledge from previously solved tasks. Experiments indicate that the transfer optimization algorithms perform well compared to traditional optimization algorithms. However, current transfer algorithms for addressing minimalistic attacks not only select a single source task for knowledge transfer but also tend to overly rely on identified appropriate source tasks. To address this issue, this paper introduces a similar locality based transfer evolutionary optimization algorithm. It can adaptively select multiple source tasks and extract valuable knowledge from these source tasks. Moreover, by leveraging the concept of similar locality, the algorithm alleviates its excessive dependence on familiar tasks, thereby providing fresh knowledge for the optimization of the target task. On this basis, the algorithm can mine more valuable knowledge from the large source task space to achieve a successful attack in a shorter period. The algorithm is tested on three Atari games-BeamRider, Qbert, and Seaquest-demonstrating its ability and potential to outperform other transfer optimization algorithms currently available in solving this problem. Wenqiang Ma, Yaqing Hou, Hua Yu 0006, Xiangrong Tong, Zexuan Zhu 0001, Qiang Zhang 0008 |
CEC | 6 |
| 2024 | Enhancing Imbalanced Classification with Support Vector Machines via Evolutionary Oversampling AlgorithmsabstractSupport Vector Machines (SVMs), as well-known algorithms, have been successfully applied to classification problems. However, when dealing with imbalanced data, the classification performance of SVMs could be significantly compromised. One approach to tackle the class imbalance is oversampling the minority class, exemplified by methods like SMOTE and its variants. These methods generate new samples by interpolation between existing ones and determine the weights based on the ratio of samples from different classes, leading to inaccurate weight assignment, limited generation scope, and indiscriminate sample generation. To address these limitations, we propose novel evolutionary oversampling algorithms based on Support Vector Machine (SVM) and Evolutionary Algorithms (EAs) called SEOA. SEOA leverages the inherent capability of SVM to identify the samples that have a critical influence on the decision boundary and assign them appropriate weights, thereby eliminating the reliance on human experience. Furthermore, SEOA utilizes a novel approach for sample generation and emphasizes the significance of margin for classification, introducing a mechanism that employs margin as the metric to evaluate the quality of generated samples. To assess the performance of SEOA, we conducted a comprehensive comparison against various oversampling methods across 19 real-world datasets. The results underscore SEOA's superiority, showcasing its distinct strengths in addressing the challenges posed by imbalanced classification. Yongchao Chen, Yaqing Hou, Xiangrong Tong, Qiang Zhang 0008 |
CEC | 6 |
| 2024 | Adaptive deep spiking neural network with global-local learning via balanced excitatory and inhibitory mechanismabstractThe training method of Spiking Neural Networks (SNNs) is an essential problem, and how to integrate local and global learning is a worthy research interest. However, the current integration methods do not consider the network conditions suitable for local and global learning, and thus fail to balance their advantages. In this paper, we propose an Excitation-Inhibition Mechanism-assisted Hybrid Learning(EIHL) algorithm that adjusts the network connectivity by using the excitation-inhibition mechanism and then switches between local and global learning according to the network connectivity. The experimental results on CIFAR10/100 and DVS-CIFAR10 demonstrate that the EIHL not only has better accuracy performance than other methods but also has excellent sparsity advantage. Especially, the Spiking VGG11 is trained by EIHL, STBP, and STDP on DVS_CIFAR10, respectively. The accuracy of the Spiking VGG11 model on EIHL is 62.45%, which is 4.35% higher than STBP and 11.40% higher than STDP, and the sparsity is 18.74%, which is 18.74% higher than the other two methods. Moreover, the excitation-inhibition mechanism used in our method also offers a new perspective on the field of SNN learning. Qi Xu 0008, Xuming Ran, Jiangrong Shen, Pan Lv, Qiang Zhang 0008, Gang Pan 0001 |
ICLR | 6 |
| 2024 | SCAT: A Time Series Forecasting with Spectral Central Alternating Transformers
Chao Che, Pengfei Wang 0013, Qiang Zhang 0008 |
IJCAI | 4 |
| 2024 | Safety-guided Deep Reinforcement Learning for Path Planning of Autonomous Mobile RobotsabstractThe past decade has witnessed the blooming of autonomous mobile robot (AMR) applications with an increasing trend from closed to open working environments with unexpected obstacles. In this context, safety becomes a critical concern in the path planning of AMRs. However, it is a challenging task to guarantee the safety while maintaining the efficiency of path planning algorithm due to the uncertainty of open environments. Therefore, we propose in this paper to categorize the working environment of AMRs into three safety levels, i.e. low-risk, medium-risk and high-risk areas, according to the real-time distance between the AMR and the nearest obstacle. In particular, we incorporate safety levels into the famous reinforcement learning algorithm, deep deterministic policy gradient (DDPG), and develop a safety-guided DDPG algorithm for the path planning of AMRs. In low-risk areas, we adopt the conventional DDPG path planning algorithm directly to guarantee the efficiency (since the safety is generally not an issue in this case). In the medium-risk areas, we design a velocity threshold adjustment method, an OU noise with bias and propose a new reward function, aiming at encouraging AMRs to perform flexible actions as early as possible in order to avoid collisions. In the high-risk areas, we re-design the reward function based on the potential collision risk and shield misleading rewards that may cause local optimum problem, so as to ensure the safety of AMRs in dangerous situations. Simulation results confirm the satisfactory performance of the proposed scheme. Zhuoru Yu, Yaqing Hou, Qiang Zhang 0008, Qian Liu 0001 |
IJCNN | 3 |
| 2024 | Reversing Structural Pattern Learning with Biologically Inspired Knowledge Distillation for Spiking Neural NetworksabstractSpiking neural networks (SNNs) have superb characteristics in sensory information recognition tasks due to their biological plausibility. However, the performance of some current spiking-based models is limited by their structures which means either fully connected or too-deep structures bring too much redundancy. This redundancy from both connection and neurons is one of the key factors hindering the practical application of SNNs. Although Some pruning methods were proposed to tackle this problem, they normally ignored the fact the neural topology in the human brain could be adjusted dynamically. Inspired by this, this paper proposed an evolutionary-based structure construction method for constructing more reasonable SNNs. By integrating the knowledge distillation and connection pruning method, the synaptic connections in SNNs can be optimized dynamically to reach an optimal state. As a result, the structure of SNNs could not only absorb knowledge from the teacher model but also search for deep but sparse network topology. Experimental results on CIFAR100, Tiny-imagenet and DVS-Gesture show that the proposed structure learning method can get pretty well performance while reducing the connection redundancy. The proposed method explores a novel dynamical way for structure learning from scratch in SNNs which could build a bridge to close the gap between deep learning and bio-inspired neural dynamics. Qi Xu 0008, Xuanye Fang, Jiangrong Shen, Qiang Zhang 0008, Gang Pan 0001 |
ACM Multimedia | 5 |
| 2024 | Towards Efficient and Diverse Generative Model for Unconditional Human Motion SynthesisabstractRecent generative methods have revolutionized the way of human motion synthesis, such as Variational Autoencoders (VAEs), Generative Adversarial Networks (GANs), and Denoising Diffusion Probabilistic Models (DMs). These methods have gained significant attention in human motion fields. However, there are still challenges in unconditionally generating highly diverse human motions from a given distribution. To enhance the diversity of synthesized human motions, previous methods usually employ deep neural networks (DNNs) to train a transport map that transforms Gaussian noise distribution into real human motion distribution. According to Figalli's regularity theory, the optimal transport map computed by DNNs frequently exhibits discontinuities. This is due to the inherent limitation of DNNs in representing only continuous maps. Consequently, the generated human motions tend to heavily concentrate on densely populated regions of the data distribution, resulting in mode collapse or mode mixture. To address the issues, we propose an efficient method called MOOT for unconditional human motion synthesis. First, we utilize a reconstruction network based on GRU and transformer to map human motions to latent space. Next, we employ convex optimization to match the noise distribution with the latent space distribution of human motions through the Optimal Transport (OT) map. Then, we combine the extended OT map with the generator of reconstruction network to generate new human motions. Thereby overcoming the issues of mode collapse and mode mixture. MOOT generates a latent code distribution that is well-behaved and highly structured, providing a strong motion prior for various applications in the field of human motion. Through qualitative and quantitative experiments, MOOT achieves state-of-the-art results surpassing the latest methods, validating its superiority in unconditional human motion generation. Hua Yu 0006, Weiming Liu 0005, Jiapeng Bai, Xu Gui 0001, Yaqing Hou, Yew-Soon Ong, Qiang Zhang 0008 |
ACM Multimedia | 7 |
| 2024 | Dynamic Temporal Shift Feature Enhancement for Few-Shot Action Recognition
Bingbing Zhang 0001, Yuanchen Ma, Jianxin Zhang 0001, Qiang Zhang 0008 |
PRCV (10) | 6 |
| 2024 | HyperTMO: a trusted multi-omics integration framework based on hypergraph convolutional network for patient classificationabstractMOTIVATION: The rapid development of high-throughput biomedical technologies can provide researchers with detailed multi-omics data. The multi-omics integrated analysis approach based on machine learning contributes a more comprehensive perspective to human disease research. However, there are still significant challenges in representing single-omics data and integrating multi-omics information. RESULTS: This article presents HyperTMO, a Trusted Multi-Omics integration framework based on Hypergraph convolutional network for patient classification. HyperTMO constructs hypergraph structures to represent the association between samples in single-omics data, then evidence extraction is performed by hypergraph convolutional network, and multi-omics information is integrated at an evidence level. Last, we experimentally demonstrate that HyperTMO outperforms other state-of-the-art methods in breast cancer subtype classification and Alzheimer's disease classification tasks using multi-omics data from TCGA (BRCA) and ROSMAP datasets. Importantly, HyperTMO is the first attempt to integrate hypergraph structure, evidence theory, and multi-omics integration for patient classification. Its accurate and robust properties bring great potential for applications in clinical diagnosis. AVAILABILITY AND IMPLEMENTATION: HyperTMO and datasets are publicly available at https://github.com/ippousyuga/HyperTMO. Qiang Zhang 0008, Xinyu Song 0002, Jue Wu, Chenghui Zhao, Kunlun He |
Bioinform. | 3 |
| 2024 | Labeled graph partitioning scheme for distributed edge caching
Pengfei Wang 0013, Geng Sun 0001, Changjun Zhou, Chengxi Gao, Sen Qiu, Tiwei Tao, Qiang Zhang 0008 |
Future Gener. Comput. Syst. | 8 |
| 2024 | Federated Unlearning With Momentum DegradationabstractData privacy is becoming increasingly important as data becomes more valuable, as evidenced by the enactment of right-to-be-forgotten laws and regulations. However, in a federated learning (FL) system, simply deleting data from the database when a user requests data revocation is not sufficient, as the training data is already implicitly contained in the parameter distribution of the models trained with it. Furthermore, the global model in the FL system is vulnerable to data poisoning attacks by malicious nodes. Exploring a reliable data poisoning reversal method can effectively counter such attacks. In this article, we analyze the necessity of decoupling the processes of unlearning and training and propose a training-agnostic and efficient method that can effectively perform two types of unlearning tasks: 1) client revocation and 2) category removal. Specifically, we decompose the unlearning process into two steps: 1) knowledge erasure and 2) memory guidance. We first propose a novel knowledge erasure strategy called momentum degradation (MoDe) which realizes the erasure of implicit knowledge in the model and ensures that the model can move smoothly to the early state of the retrained model. To mitigate the performance degradation caused by the first step, the memory guidance strategy implements guided fine-tuning of the model on different data points, which can effectively restore the discriminability of the model on the remaining data points. Extensive experiments demonstrate that our method outperforms the existing task-specific algorithms and matches the performance of retraining, accelerating the execution time by 5–20 times compared to retraining on different data sets. Yian Zhao, Pengfei Wang 0013, Heng Qi, Jianguo Huang, Zongzheng Wei, Qiang Zhang 0008 |
IEEE Internet Things J. | 6 |
| 2024 | Mitigating Poor Data Quality Impact with Federated Unlearning for Human-Centric MetaverseabstractFederated Learning (FL), which has been employed to train machine learning models on the data with a distributed manner, could enhance the immersive user experience for the human-centric metaverse. However, it’s challenging to train machine learning models accurately and promptly with FL for the human-centric metaverse due to massive data communication and user unreliability. User experience could be negatively affected by using low-quality machine learning models for human-centric metaverse, e.g., it cannot scrutinize and arrive at decisions accurately and timely. To resolve this pressing issue, we propose MetaFul a federated unlearning solution which reduces the negative influences of low-quality data with no data transmission by removing low-quality training models at the server side. To be specific, MetaFul includes three main components. (i) Low-throughput federated learning (LT-FL) addresses the issue of large model transmission in FL by decreasing the dimension and the number of transmitted model parameters. (ii) Loss-based model quality assessment (LM-QA) utilizes the model loss generated in LT-FL to estimate user data quality. (iii) Non-communicative federated unlearning (NC-FUL) revokes the low-quality data impact on the FL model with careful designed federated unlearning at the server side. Both LM-QA and NC-FUL have no communications with clients. Finally, extensive evaluations are conducted to show MetaFul could improve the model accuracy by at least 2.5% and decrease the user perception time by at least 19.3% in human-centric metaverse compared to benchmarks. Pengfei Wang 0013, Zongzheng Wei, Heng Qi, Shaohua Wan 0001, Yunming Xiao, Geng Sun 0001, Qiang Zhang 0008 |
IEEE J. Sel. Areas Commun. | 7 |
| 2024 | A graph-based approach for traffic prediction using similarity and causal relations between nodes
Alkilane Khaled, Alfateh M. Tag Elsir, Pengfei Wang 0013, Yanming Shen, Qiang Zhang 0008 |
Knowl. Based Syst. | 5 |
| 2024 | Multi-behavior contrastive learning with graph neural networks for recommendation
Xiangrong Tong, Qiang Zhang 0008 |
Knowl. Based Syst. | 4 |
| 2024 | Human-Object Interaction detection via Global Context and Pairwise-level Fusion Features Integration
Haozhong Wang, Hua Yu 0006, Qiang Zhang 0008 |
Neural Networks | 3 |
| 2024 | Visual-guided hierarchical iterative fusion for multi-modal video action recognition
Bingbing Zhang 0001, Jianxin Zhang 0001, Qiule Sun, Qiang Zhang 0008 |
Pattern Recognit. Lett. | 6 |
| 2024 | Semantic-Aware Detail Search and Feature Constraint for Cross-Resolution Person Re-IdentificationabstractCross-resolution person re-identification (CRReID) task devotes to identifying the same person from cross-resolution and cross-camera images. Existing CRReID methods learn the identity features of persons by jointly training the super-resolution (SR) and recognition models. These methods achieve sub-optimal results because the design of SR techniques is mostly oriented towards the visual quality of images rather than recognition tasks. To address this deficiency, we propose a Semantic-Aware detail Search and feature Constraint Network (SASC-Net). Specifically, we propose the semantic-aware detail search (SDS) module that is used to customize an SR module by perceiving identity-related semantic information. Then, we devise an Intra-Scale and Inter-Scale Feature Constraint loss function. It ensures that the affinity relationships of the semantic features of repaired images are close to that of high-resolution (HR) images at the scale level, reducing the solution space of the SDS module and promoting the identification module to focus on more discriminative pedestrian features. The effectiveness of our proposed method is validated by the experimental results on five cross-resolution person datasets. Tiantian Yan, Xin Yang 0011, Qiang Zhang 0008 |
IEEE Signal Process. Lett. | 4 |
| 2024 | A Multiagent Cooperative Learning System With Evolution of Social RolesabstractRecent developments in reinforcement learning (RL) have been able to derive optimal policies for sophisticated and capable agents, and shown to achieve human-level performance on a number of challenging tasks. Unfortunately, when it comes to multiagent systems (MASs), complexities, such as nonstationarity and partial observability bring new challenges to the field. Building a flexible and efficient multiagent RL (MARL) algorithm capable of handling complex tasks has to date remained an open challenge. This article presents a multiagent learning system with the evolution of social roles (eSRMA). The main interest is placed on solving the key issues in the definition and evolution of suitable roles, and optimizing the policies accompanied by social roles in MAS efficiently. Specifically, eSRMA incorporates and cultivates role division awareness of agents to improve the ability to deal with complex cooperative tasks. Each agent is assigned a role module, which can dynamically generate roles based on the individuals’ local observations. A novel MARL algorithm is designed as the principal driving force that governs the role-policy learning process by a role-attention credit assignment mechanism. Moreover, a role evolution process is developed to help agents dynamically choose appropriate roles in decision making. Comprehensive experiments on the StarCraft II micromanagement benchmarkhave demonstrated that eSRMA exhibits superiority in achieving higher learning capability and efficiency for multiple agents compared to the state-of-the-art MARL methods. Yaqing Hou, Yifeng Zeng, Yew-Soon Ong, Yaochu Jin, Hong-Wei Ge, Qiang Zhang 0008 |
IEEE Trans. Evol. Comput. | 7 |
| 2024 | A Survey on Unbalanced Classification: How Can Evolutionary Computation Help?abstractUnbalanced classification is an essential machine learning task, which has attracted widespread attention from both the academic and industrial communities due mainly to its broad applications. Evolutionary computation (EC) has contributed greatly to unbalanced classification. However, to the best of our knowledge, there have not been any comprehensive investigations on the strengths and weaknesses of alternative EC methods in addressing various challenging problems in unbalanced classification. This article reviews the literature which utilize EC techniques for unbalanced classification, with the aim of revealing the contributions of EC to unbalanced classification, providing an overview of recent advances, and identifying limitations of existing works. In addition, we present a series of real-world applications, and identify open challenges as well as possible research directions for the future. Wenbin Pei, Bing Xue 0001, Mengjie Zhang 0001, Lin Shang 0001, Xin Yao 0001, Qiang Zhang 0008 |
IEEE Trans. Evol. Comput. | 6 |
| 2024 | Fuzzy Multiview Graph Learning on Sparse Electronic Health RecordsabstractExtracting latent disease patterns from electronic health records (EHRs) is a crucial solution for disease analysis, significantly facilitating healthcare decision-making. Multiview learning presents itself as a promising approach that offers a comprehensive exploration of both structured and unstructured EHRs. However, the intrinsic uncertainty among disease features presents a significant challenge for multiview feature alignment. Besides, the sparsity of real-world EHRs also exacerbates the difficulty of feature alignment. To address these challenges, we introduce a novel fuzzy multiview graph learning framework named FuzzyMVG, which is designed for mitigating the impacts of uncertainty in disease features derived from sparse EHRs. First, we utilize auxiliary information from sparse EHRs to construct a multiview EHR graph using the structured and unstructured records. Then, for efficient feature alignment, we specially design the fuzzy logic-enhanced graph convolutional networks to obtain the fuzzy representations of time-invariant node features. Thereby, we implement a random walk strategy and long short-term memory networks to capture the distinct features of static and dynamic nodes, respectively. Extensive experiments have been conducted on the real-world MIMIC III dataset to validate the effectiveness of FuzzyMVG. Results in the diagnosis prediction task demonstrate that FuzzyMVG outperforms other state-of-the-art baselines. Tao Tang 0007, Zhuoyang Han, Shuo Yu 0001, Adil M. Bagirov, Qiang Zhang 0008 |
IEEE Trans. Fuzzy Syst. | 5 |
| 2024 | Eye Gaze Guided Cross-Modal Alignment Network for Radiology Report GenerationabstractThe potential benefits of automatic radiology report generation, such as reducing misdiagnosis rates and enhancing clinical diagnosis efficiency, are significant. However, existing data-driven methods lack essential medical prior knowledge, which hampers their performance. Moreover, establishing global correspondences between radiology images and related reports, while achieving local alignments between images correlated with prior knowledge and text, remains a challenging task. To address these shortcomings, we introduce a novel Eye Gaze Guided Cross-modal Alignment Network (EGGCA-Net) for generating accurate medical reports. Our approach incorporates prior knowledge from radiologists' Eye Gaze Region (EGR) to refine the fidelity and comprehensibility of report generation. Specifically, we design a Dual Fine-Grained Branch (DFGB) and a Multi-Task Branch (MTB) to collaboratively ensure the alignment of visual and textual semantics across multiple levels. To establish fine-grained alignment between EGR-related images and sentences, we introduce the Sentence Fine-grained Prototype Module (SFPM) within DFGB to capture cross-modal information at different levels. Additionally, to learn the alignment of EGR-related image topics, we introduce the Multi-task Feature Fusion Module (MFFM) within MTB to refine the encoder output information. Finally, a specifically designed label matching mechanism is designed to generate reports that are consistent with the anticipated disease states. The experimental outcomes indicate that the introduced methodology surpasses previous advanced techniques, yielding enhanced performance on two extensively used benchmark datasets: Open-i and MIMIC-CXR. Peixi Peng, Wanshu Fan, Wenfei Liu, Xin Yang 0011, Qiang Zhang 0008, Xiaopeng Wei |
IEEE J. Biomed. Health Informatics | 6 |
| 2024 | A Handwriting Recognition System With WiFiabstractHandwriting recognition systems are a convenient and alternative way of writing in the air with fingers rather than typing on keyboards. However, existing recognition systems are limited by their low accuracy and the requirement to wear dedicated devices. To address these issues, we propose WiWrite, an accurate contactless handwriting recognition system that allows users to write in the air without wearing any wearable devices. Specifically, we employ a novelCSI division schemeto process the noisy raw WiFi channel state information (CSI), which stabilizes the CSI phase and reduces noise in CSI amplitude. To automatically retain low noise data for identification in the LOS scenario, we propose a self-paced dense convolutional network (SPDCN), which is a self-paced loss function based on a modified convolutional neural network coupled with a dense convolutional network. Furthermore, to achieve accurate handwriting recognition in the NLOS scenario, we combine ADOA and PCA algorithms to remove location-induced interference and extract action features. Comprehensive experiments show the merits of WiWrite, revealing that the recognition accuracy for the same-size input and different-size input are 93.6% and 89.0%, respectively. Moreover, WiWrite can achieve accurate recognition regardless of environment and target diversity in LOS and NLOS scenarios. Chi Lin 0001, Asfandeyar Ahmad, Rongsheng Qu, Yi Wang 0037, Lei Wang 0005, Guowei Wu 0001, Qiang Zhang 0008 |
IEEE Trans. Mob. Comput. | 8 |
| 2024 | Maximizing Charging Efficiency With Fresnel ZonesabstractBenefitting from the discovery of wireless power transfer (WPT) technology, the wireless rechargeable sensor network (WRSN) has become a promising way for lifetime extension for wireless sensor networks. In practical WRSN scenarios, obstacles can be found almost everywhere. Most state-of-the-art researches believe that obstacles will always degrade signal strength, and omit the influence of obstacles for simplifying the computation process. However, overlooking the positive impacts of obstacles on signal propagation is inconsistent with the intrinsic features of electromagnetic waves. To address this issue, in this paper, we explore the wireless signal propagation process and provide a theoretical charging model to enhance the charging efficiency by leveraging obstacles. Through utilizing the concept of the Fresnel Zone model, we re-formalize the wireless charging model and discretize the charging area and charging time to determine the best charging locations as well as charging duration. We model the chargingEfficiencyMaximization withObstacles (EMO) problem as a submodular function maximization problem and propose a cost-efficient algorithm to solve it. Finally, test-bed experiments and extensive simulations are both conducted to verify that our schemes outperform baseline algorithms by$33.46\%$on average in charging efficiency improvement. Chi Lin 0001, Shibo Hao, Haipeng Dai 0001, Wei Yang 0039, Lei Wang 0005, Guowei Wu 0001, Qiang Zhang 0008 |
IEEE Trans. Mob. Comput. | 7 |
| 2024 | Wi-Rotate: An Instantaneous Angular Speed Measurement System Using WiFi SignalsabstractWe propose the design, implementation, and evaluation of an instantaneous angular speed (IAS) measurement system, namely Wi-Rotate, using commercial-off-the-shelf (COTS) WiFi hardware. Wi-Rotate exploits the Channel State Information (CSI) of WiFi signals to extract the physical characteristics of the rotation object to achieve accurate contact-free IAS measurements. Wi-Rotate contains three main components: Wi-Fresnel model, Wi-Phase model, and a combination model. Wi-Fresnel model explores the signal amplitude variation features when the rotating object cuts the Fresnel zone boundary to track target rotation. Wi-Phase model leverages signal phase variation and formalizes the problem of determining IAS as a linear programming problem. The combination model combines the IAS values obtained by Wi-Fresnel and Wi-Phase and utilizes a clustering method to further improve measurement accuracy. Comprehensive experiments are conducted to demonstrate the advantages of Wi-Rotate in terms of accuracy, sensing range, and system latency. Wi-Rotate is able to achieve real-time rotation measurements at an accuracy higher than 99% when the target is within 2 meters. Even when the target is 3 meters away, Wi-Rotate can still achieve an accuracy of 94%, demonstrating the long-range tracking capability which is critical for industrial applications. Chi Lin 0001, Chuanying Ji, Jie Xiong 0001, Chaocan Xiang, Lei Wang 0005, Guowei Wu 0001, Qiang Zhang 0008 |
IEEE Trans. Mob. Comput. | 7 |
| 2024 | Maximizing Charging Utility With Fresnel Diffraction ModelabstractBenefitting from the recent breakthrough of wireless power transfer technology, Wireless Rechargeable Sensor Networks (WRSNs) have become an important research topic. Most prior arts focus on system performance enhancement in the ideal environment that ignores the impact of obstacles. This contradicts the practical applications in which obstacles can be found almost anywhere and have dramatic impacts on energy transmission. In this paper, we concentrate on the problem of charging a practical WRSN in the presence of obstacles to maximize the charging utility under specific energy constraints. First, we propose a new theoretical charging model with obstacles based on the Fresnel diffraction model and conduct experiments to verify its effectiveness. Then, we propose a spatial discretization scheme to obtain a finite feasible charging position set for mobile charger (MC), which largely reduces computation overhead. Afterwards, we re-formalize charging utility maximization with energy constraints as a submodular function maximization problem and propose a cost-efficient algorithm with an approximation guarantee to solve it. In addition, we present a theoretical analysis and a relevant mathematical proof of our algorithm. Finally, we demonstrate that our scheme outperforms other competing algorithms by 20.5% on average in terms of charging utility through test-bed experiments and extensive simulations. Chi Lin 0001, Wei Yang 0039, Haipeng Dai 0001, Mohammad S. Obaidat, Lei Wang 0005, Guowei Wu 0001, Qiang Zhang 0008 |
IEEE Trans. Mob. Comput. | 7 |
| 2024 | F2UL: Fairness-Aware Federated Unlearning for Data TradingabstractFederated learning (FL) offers a credible solution for distributed data trading since it could train machine learning models in a distributed manner thereby enhancing data privacy without sharing local data. However, it is still challenging to trade data through FL due to unfair model allocation issues arising from the unreliability of user-provided data. To tackle it, we propose F2UL, a Fairness-aware Federated UnLearning solution that distributes models to users commensurate with their data quality. F2UL is trained in a two-stage (TST) way and contains three main components: 1) Label-free model quality assessment (LMQA) promotes fairness by evaluating models without user-specific data, ensuring uniform assessment standards. 2) Fair model distribution (FMD) addresses the issue of unfair model distribution by allocating models with feature mapping deviation, ensuring that users who contribute low-quality models do not receive enhanced models. 3) User data federated unlearning (UDFU) ensures fairness in model distribution by employing rapid recovery federated unlearning, safeguarding regular users from the adverse effects of low-quality data on model performance. In experiments on CIFAR10, F2UL reduces the low-quality data user accuracy to 9.09% and increases regular users’ accuracy by 3.02%, thereby demonstrating F2UL's capacity to ensure fairness in data trading. Weijian Su, Pengfei Wang 0013, Muhammed Ameen, Tiwei Tao, Xiangrong Tong, Qiang Zhang 0008 |
IEEE Trans. Mob. Comput. | 7 |
| 2024 | Decentralized Navigation With Heterogeneous Federated Reinforcement Learning for UAV-Enabled Mobile Edge ComputingabstractUnmanned Aerial Vehicle (UAV)-enabled mobile edge computing has been proposed as an efficient task-offloading solution for user equipments (UEs). Nevertheless, the presence of heterogeneous UAVs makes centralized navigation policies impractical. Decentralized navigation policies also face significant challenges in knowledge sharing among heterogeneous UAVs. To address this, we present the soft hierarchical deep reinforcement learning network (SHDRLN) and dual-end federated reinforcement learning (DFRL) as a decentralized navigation policy solution. It enhances overall task-offloading energy efficiency for UAVs while facilitating knowledge sharing. Specifically, SHDRLN, a hierarchical DRL network based on maximum entropy learning, reduces policy differences among UAVs by abstracting atomic actions into generic skills. Simultaneously, it maximizes the average efficiency of all UAVs, optimizing coverage for UEs and minimizing task-offloading waiting time. DFRL, a federated learning (FL) algorithm, aggregates policy knowledge at the cloud server and filters it at the UAV end, enabling adaptive learning of navigation policy knowledge suitable for the UAV's performance parameters. Extensive simulations demonstrate that the proposed solution not only outperforms other baseline algorithms in overall energy efficiency but also achieves more stable navigation policy learning under different levels of heterogeneity of different UAV performance parameters. Pengfei Wang 0013, Guangjie Han, Ruiyun Yu, Leyou Yang, Geng Sun 0001, Heng Qi, Xiaopeng Wei, Qiang Zhang 0008 |
IEEE Trans. Mob. Comput. | 9 |
| 2024 | GeSeNet: A General Semantic-Guided Network With Couple Mask Ensemble for Medical Image FusionabstractAt present, multimodal medical image fusion technology has become an essential means for researchers and doctors to predict diseases and study pathology. Nevertheless, how to reserve more unique features from different modal source images on the premise of ensuring time efficiency is a tricky problem. To handle this issue, we propose a flexible semantic-guided architecture with a mask-optimized framework in an end-to-end manner, termed as GeSeNet. Specifically, a region mask module is devised to deepen the learning of important information while pruning redundant computation for reducing the runtime. An edge enhancement module and a global refinement module are presented to modify the extracted features for boosting the edge textures and adjusting overall visual performance. In addition, we introduce a semantic module that is cascaded with the proposed fusion network to deliver semantic information into our generated results. Sufficient qualitative and quantitative comparative experiments (i.e., MRI-CT, MRI-PET, and MRI-SPECT) are deployed between our proposed method and ten state-of-the-art methods, which shows our generated images lead the way. Moreover, we also conduct operational efficiency comparisons and ablation experiments to prove that our proposed method can perform excellently in the field of multimodal medical image fusion. The code is available at https://github.com/lok-18/GeSeNet. Jiawei Li 0016, Jinyuan Liu 0001, Shihua Zhou, Qiang Zhang 0008, Nikola K. Kasabov |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | MCFNet: Multi-Attentional Class Feature Augmentation Network for Real-Time Scene ParsingabstractFor real-time scene parsing tasks, capturing multi-scale semantic features and performing effective feature fusion is crucial. However, many existing solutions ignore stripe-shaped things like poles, traffic lights and are so computationally expensive that cannot meet the high real-time requirements. This article presents a novel model, the Multi-Attention Class Feature Augmentation Network (MCFNet) to address this challenge. MCFNet is designed to capture long-range dependencies across different scales with low computational cost and to perform a weighted fusion of feature maps. It features the BAM (Strip Matrix Based Attention Module) for extracting strip objects in images. The BAM module replaces the conventional self-attention method using square matrices with strip matrices, which allows it to focus more on strip objects while reducing computation. Additionally, MCFNet has a parallel branch that focuses on global information based on self-attention to avoid wasting computation. The two branches are merged to enhance the performance of traditional self-attention modules. Experimental results on two mainstream datasets demonstrate the effectiveness of MCFNet. On the Camvid and Cityscapes test sets, MCFNet achieved 207.5 FPS/73.5% mIoU and 136.1 FPS/71.63% mIoU, respectively. The experiments show that MCFNet outperforms other models on the Camvid dataset and can significantly improve the performance of real-time scene parsing tasks. Xizhong Wang, Rui Liu 0015, Xin Yang 0011, Qiang Zhang 0008 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2024 | Precise Wireless Charging in Complicated EnvironmentsabstractWireless Rechargeable Sensor Networks (WRSNs) have become an important research issue as they can overcome the energy bottleneck problem of wireless sensor networks. However, inaccurate discretization methods and imprecise charging models yield a huge gap between theoretical results and practical applications, making it difficult for wide adoptions. In this paper, we focus on designing a precise charging method for maximizing charging utility when line-of-sight (LOS) and none-line-of-sight (NLOS) charging cases exist in complicated environments. First, we design discretization methods for charging area and charging orientation for precisely constructing the charging model. Then, we develop a novel electromagnetic wave reflection model to describe the signal propagation model in the presence of obstacles. We formalize the mobile charging problem into a submodular function maximization problem which can be solved by a proposed algorithm with an approximation guarantee. Finally, extensive experiments and simulations demonstrate that our schemes outperform comparison algorithms by 32.5% on average in charging utility in complicated environments. Wei Yang 0039, Chi Lin 0001, Haipeng Dai 0001, Jiankang Ren, Lei Wang 0005, Guowei Wu 0001, Qiang Zhang 0008 |
IEEE/ACM Trans. Netw. | 7 |
| 2024 | Server-Initiated Federated Unlearning to Eliminate Impacts of Low-Quality DataabstractFederated unlearning (FUL) is an emerging distributed machine learning paradigm which enables the removal or unlearning of specific training data effects from trained Federated Learning (FL) models. While current studies mostly focus on client-side FUL to address the “right to be forgotten”, and ignore the server's right to remove local models from the global model, particularly when clients are trained with low-quality data. In this paper, we introduce the Server-Initiated Federated Unlearning (SIFU) algorithm, devised to eliminate low-quality data from the global model. SIFU consists of two main components: (i) Identifying low-quality data: we develop a category-based method for quantifying low-quality data for each client and filter out clients containing such data. Datasets are then divided accordingly. (ii) Unlearning low-quality data: we employ gradient ascent training to counteract the adverse effects of low-quality data on local models. To minimize any bias introduced, we concurrently perform several batches of boosting training with good-quality data. SIFU could identify and promptly eliminate the impact of low-quality data on the FL global model while still preserving the benefits of good-quality data. Finally, extensive evaluations are conducted to verify the performance of SIFU with four different kinds of datasets and models. Results show that, compared to retraining from scratch, SIFU accelerates the speed of unlearning by 15× for small datasets (i.e., MNIST and FMNIST) and 20× for large datasets (i.e., CIFAR-10 and CelebA) without any degradation in accuracy, which also outperforms the state of the arts. Pengfei Wang 0013, Heng Qi, Changjun Zhou, Fuliang Li, Yong Wang 0046, Peng Sun 0003, Qiang Zhang 0008 |
IEEE Trans. Serv. Comput. | 8 |
| 2023 | Parallel Dense Vision Transformer and Augmentation Network for Occluded Person Re-identification
Chuxia Yang, Wanshu Fan, Ziqi Wei 0001, Xin Yang 0011, Qiang Zhang 0008 |
CAD/Graphics | 5 |
| 2023 | A Joint Optimization Scheme in Heterogeneous UAV-Assisted MEC
Pengfei Wang 0013, Qiang Zhang 0008 |
ICA3PP (4) | 3 |
| 2023 | Anomalous Behavior Identification with Visual Federated Learning in Multi-UAVs SystemsabstractAnomaly detection aims to identify data or behav-iors that are different from the usual patterns. In traditional anomaly detection settings, edge devices collect the data and send it to a centralized server for model training, which faces two critical issues: (1) it risks data exposure during transmission; (2) it demands a large amount of network bandwidth for data transfer. To tackle these problems, we propose a Visual Federated Learning algorithm (VFLA) for anomalous behavior identification in the multi-UAVs system. To the best of our knowledge, we are the first to merge federated learning with video-based anomaly detection. VFLA consists of two phases: The initial phase is training a pseudo-label generator. UAVs collect a dataset and manually annotate it. This labeled data is then used to train the pseudo-label generator on the server, which is subsequently distributed back to the UAVs. The second phase is the federated learning-based anomaly detection model training. UAVs leverage the pseudo-label generator to automatically annotate the collected video footage. These annotated videos are fed into an anomaly detection network for training. Once the local training is completed, UAVs upload their local models to a server for federated aggregation. The global model is then redistributed to the UAVs for additional training rounds, until reach the target accuracy. Finally, we simulate the federated learning anomaly detection algorithm on the Shanghai-tech dataset, it demonstrates an average accuracy boost of 5.6% compared to baselines. Pengfei Wang 0013, Xinrui Yu, Yefei Ye, Heng Qi, Shuo Yu 0001, Leyou Yang, Qiang Zhang 0008 |
ICPADS | 7 |
| 2023 | Causal Deep Reinforcement Learning Using Observational DataabstractDeep reinforcement learning (DRL) requires the collection of interventional data, which is sometimes expensive and even unethical in the real world, such as in the autonomous driving and the medical field. Offline reinforcement learning promises to alleviate this issue by exploiting the vast amount of observational data available in the real world. However, observational data may mislead the learning agent to undesirable outcomes if the behavior policy that generates the data depends on unobserved random variables (i.e., confounders). In this paper, we propose two deconfounding methods in DRL to address this problem. The methods first calculate the importance degree of different samples based on the causal inference technique, and then adjust the impact of different samples on the loss function by reweighting or resampling the offline dataset to ensure its unbiasedness. These deconfounding methods can be flexibly combined with existing model-free DRL algorithms such as soft actor-critic and deep Q-learning, provided that a weak condition can be satisfied by the loss functions of these algorithms. We prove the effectiveness of our deconfounding methods and validate them experimentally. Wenxuan Zhu, Chao Yu 0004, Qiang Zhang 0008 |
IJCAI | 3 |
| 2023 | Charging Dynamic Sensors through Online LearningabstractAs a novel solution for IoT applications, wireless rechargeable sensor networks (WRSNs) have achieved widespread deployment in recent years. Existing WRSN scheduling methods have focused extensively on maximizing the network charging utility in the fixed node case. However, when sensor nodes are deployed in dynamic environments (e.g., maritime environments) where sensors move randomly over time, existing approaches are likely to incur significant performance loss or even fail to execute normally. In this work, we focus on serving dynamic nodes whose locations vary randomly and formalize the dynamic WRSN charging utility maximization problem (termed MATA problem). Through discretizing candidate charging locations and modeling the dynamic charging process, we propose a near-optimal algorithm for maximizing charging utility. Moreover, we point out the long-short-term conflict of dynamic sensors that their location distributions in the short-term usually deviate from the long-term expectations. To tackle this issue, we further design an online learning algorithm based on the combinatorial multi-armed bandit (CMAB) model. It iteratively adjusts the charging strategy and adapts well to nodes’ short-term location deviations. Extensive experiments and simulations demonstrate that the proposed scheme can effectively charge dynamic sensors and achieve a higher charging utility compared to baseline algorithms in both long-term and short-term. Yu Sun 0077, Chi Lin 0001, Wei Yang 0039, Jiankang Ren, Lei Wang 0005, Guowei Wu 0001, Qiang Zhang 0008 |
INFOCOM | 7 |
| 2023 | A Heuristic Framework for Personalized Route Recommendation Based on Convolutional Neural Networks
Ruining Zhang, Chanjuan Liu 0001, Qiang Zhang 0008, Xiaopeng Wei |
PRICAI (3) | 3 |
| 2023 | Diformer: A dynamic self-differential transformer for new energy power autoregressive prediction
Chao Che, Pengfei Wang 0013, Qiang Zhang 0008 |
Knowl. Based Syst. | 4 |
| 2023 | SCAU-net: 3D self-calibrated attention U-Net for brain tumor segmentation
Ning Sheng, Yutong Han, Yaqing Hou, Bin Liu 0040, Jianxin Zhang 0001, Qiang Zhang 0008 |
Neural Comput. Appl. | 7 |
| 2023 | GSoANet: Group Second-Order Aggregation Network for Video Action Recognition
Zhenwei Wang 0005, Bingbing Zhang 0001, Jianxin Zhang 0001, Bin Liu 0040, Qiang Zhang 0008 |
Neural Process. Lett. | 7 |
| 2023 | Large-Field Contextual Feature Learning for Glass DetectionabstractGlass is very common in our daily life. Existing computer vision systems neglect it and thus may have severe consequences, e.g., a robot may crash into a glass wall. However, sensing the presence of glass is not straightforward. The key challenge is that arbitrary objects/scenes can appear behind the glass. In this paper, we propose an important problem of detecting glass surfaces from a single RGB image. To address this problem, we construct the first large-scale glass detection dataset (GDD) and propose a novel glass detection network, called GDNet-B, which explores abundant contextual cues in a large field-of-view via a novel large-field contextual feature integration (LCFI) module and integrates both high-level and low-level boundary features with a boundary feature enhancement (BFE) module. Extensive experiments demonstrate that our GDNet-B achieves satisfying glass detection results on the images within and beyond the GDD testing set. We further validate the effectiveness and generalization capability of our proposed GDNet-B by applying it to other vision tasks, including mirror segmentation and salient object detection. Finally, we show the potential applications of glass detection and discuss possible future research directions. Haiyang Mei, Xin Yang 0011, Letian Yu, Qiang Zhang 0008, Xiaopeng Wei, Rynson W. H. Lau |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | High Net Information Density DNA Data Storage by the MOPE Encoding AlgorithmabstractDNA has recently been recognized as an attractive storage medium due to its high reliability, capacity, and durability. However, encoding algorithms that simply map binary data to DNA sequences have the disadvantages of low net information density and high synthesis cost. Therefore, this paper proposes an efficient, feasible, and highly robust encoding algorithm called MOPE (Modified Barnacles Mating Optimizer and Payload Encoding). The Modified Barnacles Mating Optimizer (MBMO) algorithm is used to construct the non-payload coding set, and the Payload Encoding (PE) algorithm is used to encode the payload. The results show that the lower bound of the non-payload coding set constructed by the MBMO algorithm is 3%-18% higher than the optimal result of previous work, and theoretical analysis shows that the designed PE algorithm has a net information density of 1.90 bits/nt, which is close to the ideal information capacity of 2 bits per nucleotide. The proposed MOPE encoding algorithm with high net information density and satisfying constraints can not only effectively reduce the cost of DNA synthesis and sequencing but also reduce the occurrence of errors during DNA storage. Yanfen Zheng, Ben Cao, Jieqiong Wu, Bin Wang 0005, Qiang Zhang 0008 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2023 | Learning a Coordinated Network for Detail-Refinement Multiexposure Image FusionabstractNowadays, deep learning has made rapid progress in the field of multi-exposure image fusion. However, it is still challenging to extract available features while retaining texture details and color. To address this difficult issue, in this paper, we propose a coordinated learning network for detail-refinement in an end-to-end manner. Firstly, we obtain shallow feature maps from extreme over/under-exposed source images by a collaborative extraction module. Secondly, smooth attention weight maps are generated under the guidance of a self-attention module, which can draw a global connection to correlate patches in different locations. With the cooperation of the two aforementioned used modules, our proposed network can obtain a coarse fused image. Moreover, by assisting with an edge revision module, edge details of fused results are refined and noise is suppressed effectively. We conduct subjective qualitative and objective quantitative comparisons between the proposed method and twelve state-of-the-art methods on two available public datasets, respectively. The results show that our fused images significantly outperform others in visual effects and evaluation metrics. In addition, we also perform ablation experiments to verify the function and effectiveness of each module in our proposed method. The source code can be achieved athttps://github.com/lok-18/LCNDR. Jiawei Li 0016, Jinyuan Liu 0001, Shihua Zhou, Qiang Zhang 0008, Nikola K. Kasabov |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | Toward Realistic 3D Human Motion Prediction With a Spatio-Temporal Cross- Transformer ApproachabstractHuman motion prediction intends to predict how humans move given a historical sequence of 3D human motions. Recent transformer-based methods have attracted increasing attentions and demonstrated their promising performance in 3D human motion prediction. However, existing methods generally decompose the input of human motion information into spatial and temporal branches in a separate way and seldom consider their inherent coherence between the two branches, hence often failing to register the dynamic spatio-temporal information during the training process. Motivated by these issues, we propose a spatio-temporal cross-transformer network (STCT) for 3D human motion predictions. Specifically, we investigate various types of interaction methods (i.e., Concatenation Interaction, Msg token interaction, and Cross-transformer) to capture the coherence of the spatial and temporal branches. According to the obtained results, the proposed cross-transformer interaction method shows its superiority over other methods. Meanwhile, considering that most existing works treat the human body as a set of 3D human joint positions, the predicted human joints are proportionally less appropriate to the realistic human body due to unreasonable bone length and non-plausible poses as time progresses. We further resort to the bone constraints of human mesh to produce more realistic human motions. By fitting a parametric body model (i.e., SMPL-X model) to the predicted human joints, a reconstruction loss function is proposed to remedy the unreasonable bone length and pose errors. Comprehensive experiments on AMASS and Human3.6M datasets have demonstrated that our method achieves superior performance over compared methods. Hua Yu 0006, Xuanzhe Fan, Yaqing Hou, Wenbin Pei, Hong-Wei Ge, Xin Yang 0011, Qiang Zhang 0008, Mengjie Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 8 |
| 2023 | Graph Optimized Data Offloading for Crowd-AI Hybrid Urban Tracking in Intelligent Transportation SystemsabstractUrban tracking plays a vital role for people’s urban life in intelligent transportation systems, e.g., public safety, case investigation, finding missing items, etc. However, the current tracking methods consume a large amount of communication and computing resources since they mainly offload all related sensing data, i.e., videos, generated by widely deployed cameras to the cloud where data are stored, processed, and analyzed. In this paper, we propose a graph optimized data offloading algorithm leveraging a crowd-AI hybrid method to minimize the data offloading cost and ensure the reliable urban tracking result. To be specific, we first formulate a crowd-AI hybrid urban tracking scenario, and prove the proposed data offloading problem in this scenario is NP-hard. Then, we solve it by decomposing the problem into two parts, i.e., trajectory prediction and task allocation. The trajectory prediction algorithm, leveraging the state graph, computes possible tracking areas of the target object, and the task allocation algorithm, using the dependency graph, chooses the optimal set of crowds and cameras to cover the tracking area while minimizing the data offloading cost separately. Finally, the extensive simulations with large real world data set are conducted showing that the proposed algorithm outperforms benchmarks in reducing data offloading cost while ensuring the tracking success rate in intelligent transportation systems. Pengfei Wang 0013, Yuzhu Pan, Chi Lin 0001, Heng Qi, Jiankang Ren, Ning Wang 0018, Qiang Zhang 0008 |
IEEE Trans. Intell. Transp. Syst. | 9 |
| 2023 | Near Optimal Charging Schedule for 3-D Wireless Rechargeable Sensor NetworksabstractWireless rechargeable sensor networks (WRSNs) have become a hot research issue owing to the breakthrough of wireless power transfer (WPT) technology. Previous theoretical schemes are mostly designed for 2-D networks, and few of them are tailored for 3-D scenarios, making them not suitable for wide adoptions in practical applications. In this paper, we address the issue of how to serve a 3-D WRSN with an unmanned aerial vehicle (UAV). Our main concern is to maximize the charged energy for sensors supplied by the UAV, which has energy constraints. We respectively develop a spatial discretization scheme to construct a finite feasible set of charging spots for the UAV in a 3-D environment and a temporal discretization scheme to determine the appropriate charging duration for each charging spot. Then, we reduce the problem into a submodular maximization problem with routing constraints and present a cost-efficient algorithm (CEA) with a provable approximation ratio to solve it. Finally, test-bed experiments are conducted to show the feasibility of our schemes in practical scenarios. Extensive simulations are taken to verify the superior performance of our algorithm in charged energy and robustness. The charged energy of our scheme outperforms other competing methods by at least$18.2\%$. Chi Lin 0001, Wei Yang 0039, Haipeng Dai 0001, Teng Li 0003, Yi Wang 0037, Lei Wang 0005, Guowei Wu 0001, Qiang Zhang 0008 |
IEEE Trans. Mob. Comput. | 8 |
| 2023 | Maximizing Energy Efficiency of Period-Area Coverage With a UAV for Wireless Rechargeable Sensor NetworksabstractWireless Rechargeable Sensor Networks (WRSNs) with perpetual network lifetime have been used in many Internet of Things (IoT) applications, like oceanic monitoring and precision agriculture. Rechargeable sensors, together with an Unmanned Aerial Vehicle (UAV), are collaboratively employed for fulfilling periodic coverage missions. However, traditional coverage solutions are normally based on static deployment of sensors and not suitable for such novel coverage requirements. In this paper, we propose the concept of Period-Area Coverage (PAC) problem, which requires the data of the overall area must be collected/monitored periodically. To solve the PAC problem, we employ a UAV that simultaneously acts as a mobile charger and sensor. It is responsible for charging nearly exhausted sensors and sensing vacant regions to realize complete event monitoring. To maximize the energy efficiency of the UAV, we propose a heuristic hexagon-based scheduling algorithm (HSA) which can also balance energy consumption. Furthermore, we develop an emergent node charging scheduling method to prevent node exhaustion, and introduce a grid-based boustrophedon scheduling algorithm (GBSA) to reduce the complexity. Finally, we present a charging re-allocation mechanism to further enhance energy efficiency. Extensive simulations demonstrate that the proposed schemes can solve the PAC problem and enhance energy efficiency by at least 18.2% compared to prior arts. Test-bed experiments conducted both in agriculture and oceanic monitoring applications validate the applicability of the proposed scheme in practical scenarios. Chi Lin 0001, Shibo Hao, Wei Yang 0039, Pengfei Wang 0013, Lei Wang 0005, Guowei Wu 0001, Qiang Zhang 0008 |
IEEE/ACM Trans. Netw. | 7 |
| 2023 | Robust Wireless Rechargeable Sensor NetworksabstractWireless rechargeable sensor networks have become a hot research issue as it can overcome the limited energy bottleneck of wireless sensor networks owing to the recent breakthrough of wireless power transfer technology. Though network lifetime is prolonged and sensor nodes can sustain immortally, the issue of network robustness is overlooked, yielding most theoretical work unsuitable for practical applications when confronting with unpredictable packet loss. In this paper, we address the network robustness issue by maximizing the charging utility in a risk-averse view. First, we build a risk-averse model based on the concept of CVaR (Conditional Value at Risk), which trades-off charging utility and risk aversion for quantifying robustness. Then, we propose a spatial discretization scheme to construct a charging route for mobile charger, which can reduce computational overhead. Afterwards, a path optimization scheme is designed to further improve the charging utility. We convert the original problem into the submodular function maximization problem and propose a method with a performance guarantee while maximizing the system robustness. Finally, testbed experiments and simulations are conducted, and the results demonstrate that our schemes outperform comparison algorithms by at least 22.4% in effective energy in the presence of risks to guarantee system robustness. Wei Yang 0039, Chi Lin 0001, Haipeng Dai 0001, Pengfei Wang 0013, Jiankang Ren, Lei Wang 0005, Guowei Wu 0001, Qiang Zhang 0008 |
IEEE/ACM Trans. Netw. | 8 |
| 2023 | A Contactless Authentication System Based on WiFi CSIabstractThe ubiquitous and fine-grained features of WiFi signals make it promising for realizing contactless authentication. Existing methods, though yielding reasonably good performance in certain cases, are suffering from two major drawbacks: sensitivity to environmental dynamics and over-dependence on certain activities. Thus, the challenge of solving such issues is how to validate human identities under different environments, even with different activities. Toward this goal, in this article, we develop WiTL, a transfer learning–based contactless authentication system, which works by simultaneously detecting unique human features and removing the environment dynamics contained in the signal data under different environments. To correctly detect human features (i.e., human heights used in this article), we design a Height EStimation (HES) algorithm based on Angle of Arrival (AoA). Furthermore, a transfer learning technology combined with the Residual Network (ResNet) and the adversarial network is devised to extract activity features and learn environmental independent representations. Finally, experiments through multi-activities and under multi-scenes are conducted to validate the performance of WiTL. Compared with the state-of-the-art contactless authentication systems, WiTL achieves a great accuracy over 93% and 97% in multi-scenes and multi-activities identity recognition, respectively. Chi Lin 0001, Pengfei Wang 0013, Chuanying Ji, Mohammad S. Obaidat, Lei Wang 0005, Guowei Wu 0001, Qiang Zhang 0008 |
ACM Trans. Sens. Networks | 7 |
| 2023 | Explore Contextual Information for 3D Scene Graph Generationabstract3D scene graph generation (SGG) has been of high interest in computer vision. Although the accuracy of 3D SGG on coarse classification and single relation label has been gradually improved, the performance of existing works is still far from being perfect for fine-grained and multi-label situations. In this article, we propose a framework fully exploring contextual information for the 3D SGG task, which attempts to satisfy the requirements of fine-grained entity class, multiple relation labels, and high accuracy simultaneously. Our proposed approach is composed of a Graph Feature Extraction module and a Graph Contextual Reasoning module, achieving appropriate information-redundancy feature extraction, structured organization, and hierarchical inferring. Our approach achieves superior or competitive performance over previous methods on the 3DSSG dataset, especially on the relationship prediction sub-task. Chengjiang Long, Zhaoxuan Zhang, Bokai Liu, Qiang Zhang 0008, Xin Yang 0011 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2023 | Lightweight single-image super-resolution via multi-scale feature fusion CNN and multiple attention block
Wanshu Fan, Xin Yang 0011, Qiang Zhang 0008 |
Vis. Comput. | 4 |
| 2022 | Attention multiple instance learning with Transformer aggregation for breast cancer whole slide image classificationabstractRecently, attention-based multiple instance learning (MIL) methods have received more concentration in histopathology whole slide image (WSI) applications. However, existing attention-based MIL methods rarely consider the cross-channel information interaction of pathology images when identifying discriminant patches. Additionally, they also have limitations on capturing the correlation between different discriminant instances for the bag-level classification. To address these challenges, we present a novel attention-based MIL model (AMIL-Trans) for breast cancer WSI classification. AMIL-Trans first embeds the efficient channel attention to realize the cross-channel interaction of pathology images, thus computing more robust features for instance selection without introducing too much computation cost. Then, it leverages vision Transformer encoder to directly aggregate selected instance features for better bag-level prediction, which effectively considers the correlation between different discriminant instances. Experiment results illustrate that AMIL-Trans respectively achieves its optimal AUC of 94.27% and 84.22% on the Camelyon-16 dataset and MSK external validation dataset, demonstrating the competitive performance compared with state-of-the-art MIL methods on breast cancer WSI classification task. The code will be available at https://github.con CunqiaoHou/AMIL-Trans. Jianxin Zhang 0001, Cunqiao Hou, Wen Zhu, Ying Zou 0015, Qiang Zhang 0008 |
BIBM | 7 |
| 2022 | A Preliminary Study of Multi-task MAP-Elites with Knowledge Transfer for Robotic Arm DesignabstractThe structure design of robotic arms is of great importance on completing industrial tasks successfully. This is a typical multi-task optimization problem when considering different constraints as different tasks. However, mainstream methods for multi-task optimization such as evolutionary multitasking and Multi-task MAP-Elites algorithms tend to encounter problems such as high computational cost and slow convergence when solving large-scale robotic arm tasks. To this end, this paper proposes a new framework based on the MAP-Elites algorithms for solving large-scale robot arm design tasks, called Multi-task MAP-Elites with Knowledge Transfer (MMKT). Specifically, this paper designs the group-based knowledge transfer process for large-scale task optimization in which all tasks are classified into different groups according to their similarity to generate multiple knowledge transfer areas; and knowledge transfer strategies are designed to enhance the quality of solutions with low fitness value. We test the effectiveness of the MMKT framework in planar robotic arm experiments (2000, 5000, and 10,000 tasks; 10, 15-dimensional search space). The experimental results prove that the MMKT outperforms the MME, CMA-ES, and classical ES algorithms. Hua Yu 0006, Han Linghu, Yaqing Hou, Hong-Wei Ge, Qiang Zhang 0008 |
CEC | 8 |
| 2022 | Wider and Higher: Intensive Integration and Global Foreground Perception for Image Matting
Yu Qiao 0001, Ziqi Wei 0001, Yuhao Liu 0001, Yuxin Wang 0001, Qiang Zhang 0008, Xin Yang 0011 |
CGI | 6 |
| 2022 | Research on Depth-Adaptive Dual-Arm Collaborative Grasping Method
Rui Liu 0015, Jing Dong 0009, Qiang Zhang 0008 |
CollaborateCom (2) | 5 |
| 2022 | Are You Really Charging Me?abstractWireless rechargeable sensor networks (WRSNs), which benefit from recent breakthroughs in Wireless Power Transfer (WPT) technology, emerge as very promising for network lifetime extension. Traditional methods concentrate on system performance improvement while little attention has been paid to security, making them vulnerable to novel attacks. In this paper, we develop a novel Charging Spoofing Attack (CSA), in which a mobile charger (MC) is charging a node intuitively. Nevertheless, it is launching an attack based on the nonlinear superposition principle of electromagnetic waves, causing the target node to be unable to receive any energy and finally exhausted in vain. First, we explain and model the nonlinear superposition effect through experiments, which points out the potential of launching such a novel attack. Second, we formalize the attacking problem as a charging uTility optImization problem with key noDe timE window constraints (TIDE). Then, we propose an approximation algorithm termed CSA to solve the TIDE problem with a bounded performance guarantee. Theoretical analyses are presented to exploit the feature of CSA. Finally, to demonstrate the outperformed features of our scheme, extensive simulations and test-bed experiments are conducted, revealing that CSA can exhaust at least 80% of key nodes without being detected. Chi Lin 0001, Ziwei Yang 0004, Jiankang Ren, Lei Wang 0005, Guowei Wu 0001, Qiang Zhang 0008 |
ICDCS | 7 |
| 2022 | Precise Wireless Charging in Complicated EnvironmentsabstractWireless Rechargeable Sensor Networks (WRSNs) have become an important research issue as it can overcome the energy bottleneck problem of wireless sensor networks. However, inaccurate discretization methods and imprecise charging models yield a huge gap between theoretical results and practical applications, making it difficult for wide adoptions. In this paper, we focus on designing a precise charging method for maximizing charging utility when line-of-sight (LOS) and none-line-of-sight (NLOS) charging cases exist in complicated environments. First, we design discretization methods for charging area and charging orientation for precisely constructing the charging model. Then, we develop a novel electromagnetic wave reflection model to describe the signal propagation model in the presence of obstacles. We formalize the mobile charging problem into a submodular function maximization problem which can be solved by a proposed algorithm with an approximation guarantee. Finally, extensive experiments and simulations demonstrate that our schemes outperform comparison algorithms by 31.45% on average in charging utility in complicated environments. Wei Yang 0039, Chi Lin 0001, Haipeng Dai 0001, Jiankang Ren, Pengfei Wang 0013, Lei Wang 0005, Guowei Wu 0001, Qiang Zhang 0008 |
ICDCS | 8 |
| 2022 | A New Perspective on the Effects of Spectrum in Graph Neural NetworksabstractMany improvements on GNNs can be deemed as operations on the spectrum of the underlying graph matrix, which motivates us to directly study the characteristics of the spectrum and their effects on GNN performance. By generalizing most existing GNN architectures, we show that the correlation issue caused by the unsmooth spectrum becomes the obstacle to leveraging more powerful graph filters as well as developing deep architectures, which therefore restricts GNNs’ performance. Inspired by this, we propose the correlation-free architecture which naturally removes the correlation issue among different channels, making it possible to utilize more sophisticated filters within each channel. The final correlation-free architecture with more powerful filters consistently boosts the performance of learning graph representations. Code is available at https://github.com/qslim/gnn-spectrum. Mingqi Yang, Yanming Shen, Rui Li 0086, Heng Qi, Qiang Zhang 0008 |
ICML | 5 |
| 2022 | Towards Efficient 3D Human Motion Prediction using Deformable Transformer-based Adversarial NetworkabstractHuman motion prediction is a crucial step for achieving human-robot interactions. While recent transformer-based methods have shown great potentials in 3D human motion prediction, they still suffer from mode collapse to non-plausible poses and quadratically computational complexity with respect to the increasing length of input sequences. In this paper, we propose a novel spatio-temporal deformable transformer-based adversarial network (STDTA) for 3D human motion prediction. First, we design a spatio-temporal deformable transformer module to capture the correlations between human joints while reducing the computational costs. Second, we introduce the adversarial training mechanism and design fidelity and continuity discriminators to maintain smoothness and stability for the long-term prediction. Finally, extensive experiments on Human 3.6M and AMASS benchmarks demonstrate that the proposed STDTA achieves state-of-the-art performance. Hua Yu 0006, Xuanzhe Fan, Yaqing Hou, Cai Kang, Qiang Zhang 0008 |
ICRA | 7 |
| 2022 | Multi-Omics Data Integration Patient Classification Method Based on Deep Dense Residual Shrinkage NetworkabstractRepresented by high-throughput sequencing technology, rapid advances in omics technologies have allowed scholars to understand human diseases in a more profound and comprehensive manner. However, omics data are difficult to analyze accurately on account of their characters: massive redundancy, full of noise and high dimension. Meanwhile, it is crucial to improve the performance of omics analysis from multiple modalities using data from different omics. To this end, we propose a novel multi-omics integration model for patient classification. Firstly, anti-noisy deep dense residual shrinkage neural networks (DDRSNN) are trained and utilized to make preliminary predictions on various omics data of patients. Then confidence scores are calculated for each sample through the confidence distance to evaluate the reliability of the preliminary prediction result. Finally, referring to uncertain evidence fusion theory, the preliminary prediction results from different omics are integrated based on confidence scores and the final patient classification is obtained. The algorithm combines multiple omics, treating each omics as evidence of a separate modality, and improves the accuracy and reliability of patient classification by integrating evidence from multiple modalities. Yaqing Hou, Qiang Zhang 0008 |
ICTAI | 4 |
| 2022 | SAMKR: Bottom-up Keypoint Regression Pose Estimation Method Based On Subspace Attention ModuleabstractAs a hot research issue in computer vision, 2D human pose estimation plays an important role in human-computer interaction, intelligent monitoring, 3D human pose estimation and so on. Aiming at the problem of human scale inconsistency in the situation of multi-person, a keypoint regression method based on subspace attention module (SAMKR) is proposed in this paper for the 2D human pose estimation. Firstly, each keypoint is divided into independent regression branches, and then the feature mapping in each keypoint regression branch is evenly divided into a specified number of feature mapping subspaces, and different attention mappings are derived for each feature mapping subspace. By learning different attention maps in each feature subspace, multi-scale feature representation can be effectively improved. The experimental results show that SAMKR achieved 74.4 AP score on the CrowdPose test set, which may lead to an improvement of + 7.1AP, and reached 70.4 AP score on the COCO test-dev data set, which was 0.4AP higher than the baseline. Rui Liu 0015, Qiang Zhang 0008 |
IJCNN | 4 |
| 2022 | High-order Correlation Network for Video RecognitionabstractHow to model global video representation is an important research content of video recognition. Among current convolutional neural network(CNN) based methods, only using first-order representations (i.e. global average pooling) has limitations in capturing spatiotemporal features of videos. Recent studies have shown that high-order statistics are more suitable to model complex feature distributions. To better characterize the spatiotemporal structure for video recognition, we propose a novel High-order Correlation Network (HoCNet) in this work, the core of which is to explore high-order video representations through correlation computation and covariance pooling. HoC-Net leverages the correlation module to obtain complex temporal dynamic information of frames via computing dot product of features in the fixed sliding window of two adjacent frames. As an approximate high-order calculation, the correlation module can be inserted into any stage of the deep network to model high-order representations in various spatial resolutions. Additionally, a robust high-order pooling module, i.e., iterative matrix square root normalization of covariance pooling (iSQRT-COV), is also introduced at the end of the network, and this further boosts modeling complex spatiotemporal distributions of video features. Experiments conducted on four widely used video benchmarks demonstrate the effectiveness of HoCNet, which achieves the comparable performance with the state-of-the-art models. Zhenwei Wang 0005, Bingbing Zhang 0001, Jianxin Zhang 0001, Qiang Zhang 0008 |
IJCNN | 5 |
| 2022 | A Novel Movement-supported HRI Framework for Humanoid RobotsabstractCurrent research related to human-robot interaction (HRI) of bipedal humanoid robots often assumes that the robot is in a standing stationary state, i.e., the relative position of the robot does not change, and rarely considers the effect of lower limb movement on interaction. However, HRI in the real world does not assume a moving or stationary state of the robot, and the equilibrium perturbations caused by movement can prevent HRI from functioning properly. In this paper, we propose a movement supported humanoid robot interaction method that empowers the robot to move stably while achieving HRI. First, a reinforcement learning-based neural network is run offline to generate interaction actions that satisfy the equilibrium constraint and support movement, and then an intention recognition network is introduced to run the movement-supported HRI framework online. It is demonstrated that the training method proposed in this paper can enable a bipedal robot to achieve a variety of interactive actions while moving stably. Jing Dong 0009, Rui Liu 0015, Xiaopeng Wei, Qiang Zhang 0008 |
IJCNN | 7 |
| 2022 | MDoC: Compromising WRSNs through Denial of Charge by Mobile ChargerabstractThe discovery of wireless power transfer technology enables power transferred between transceivers in a wireless manner, thus generating the concept of wireless rechargeable sensor networks (WRSNs). Previous arts paid little attention to network security issues, making them prone to novel attacks. In this work, we focus on developing a denial of charge attack for WRSNs, which aims at corrupting network functionalities by manipulating the malicious mobile charger. We formalize the maximization of destructiveness problem (MAD) and propose a denial of charge attacking method, termed MDoC, with a performance guarantee to solve it. MDoC is composed of two attacking rounds, which first triggers sensors to send requests to create a request explosion phenomenon and then figures out the longest charging route to yield nodes starving to death as much as possible. Finally, extensive testbed experiments and simulations are conducted to verify the performance of MDoC. The results reveal that MDoC attack is able to exhaust at least 20% additional nodes without being noticed. Chi Lin 0001, Pengfei Wang 0013, Qiang Zhang 0008, Hao Wang 0023, Lei Wang 0005, Guowei Wu 0001 |
INFOCOM | 3 |
| 2022 | Subset Selection for Hybrid Task Scheduling with General Cost ConstraintsabstractSubset selection problem for task scheduling with general cost constraints exists widely in IoT applications. Its objective is to select several profitable tasks to execute under routing and cost constraints such that the total profit is maximized. Most prior arts only focus on either online tasks or offline tasks, which are usually inapplicable in practical applications where online tasks and offline tasks co-exist. In this paper, we study the subset selection problem for HybrId Task Scheduling with general cost constraints (HITS), in which both online and offline tasks are scheduled to maximize the overall profit. We first divide the HITS problem into online and offline subproblems and propose two algorithms to solve them with bounded approximation ratios. Furthermore, we propose an approximation algorithm for the hybrid scenario where both online and offline tasks are considered. Extensive simulations show that our proposed algorithm outperforms baseline algorithms by 21.5% averagely in profit and also performs well in pure online/offline scenarios. We further demonstrate the feasibility of our algorithm through test-bed experiments in a realistic scene. Yu Sun 0077, Chi Lin 0001, Jiankang Ren, Pengfei Wang 0013, Lei Wang 0005, Guowei Wu 0001, Qiang Zhang 0008 |
INFOCOM | 7 |
| 2022 | Tactile Pattern Super Resolution with Taxel-based SensorsabstractIn contrast to sophisticated means of visual su-per resolution (SR), not much work has been done in the tactile SR field. Existing tactile SR algorithms for taxel-based sensors mainly focus on enhancing the localization accuracy, and generally associate with a specific type of hardware, sometimes not applicable to generic taxel-based tactile sensors. Inspired by image SR, we investigate the tactile pattern SR in this paper, and present how to transform successful image SR schemes, e.g. Convolutional Neural Network (CNN) and Generative Adversarial Network (GAN) to serve the tactile SR. We propose two tactile SR models, i.e. TactileSRCNN and TactileSRGAN, and establish a new tactile pattern SR dataset for model learning. The ground truth of high resolution (HR) tactile patterns in the dataset is obtained via multi-sampling (i.e. overlapping reception) and registration of low resolution (LR) sensor. One key contribution of this research lies in achieving ×100 (from 3×4×4 to 40×40) times tactile pattern SR with a one-time tapping of 3-axis taxel-based sensor. Different from existing tactile SR algorithms which improves the localization accuracy of a single contact point, the proposed scheme can provide multi-point contact detection to robotic applications. Qian Liu 0001, Qiang Zhang 0008 |
IROS | 3 |
| 2022 | Joint Convolutional and Self-Attention Network for Occluded Person Re-IdentificationabstractOccluded person Re-Identification (Re-ID) is built on cross views, which aims to retrieve a target person in occlusion scenes. Under the condition that occlusion leads to the interference of other objects and the loss of personal information, the efficient extraction of personal feature representation is crucial to the recognition accuracy of the system. Most of the existing methods solve this problem by designing various deep networks, which are called convolutional neural networks (CNN)-based methods. Although these methods have the powerful ability to mine local features, they may fail to capture features containing global information due to the limitation of the gaussian distribution property of convolution operation. Recently, methods based on Vision Transformer (ViT) have been successfully employed to person Re-ID task and achieved good performance. However, since ViT-based methods lack the capability of extracting local information from person images, the generated results may severely lose local details. To address these deficiencies, we design a convolution and self-attention aggregation network (CSNet) by combining the advantages of both CNN and ViT. The proposed CSNet consists of three parts. First, to better capture personal information, we adopt Dual-Branch Encoder (DBE) to encode person images. Then, we also embed a Local Information Aggregation Module (LIAM) in the feature map, which effectively leverages the useful information in the local feature map. Finally, a Multi-Head Global-to-Local Attention (MHGLA) module is designed to transmit global information to local features. Experimental results demonstrate the superiority of the proposed method compared with the state-of-the-art (SOTA) methods on both the occluded person Re-ID datasets and the holistic person Re-ID datasets. Chuxia Yang, Wanshu Fan, Qiang Zhang 0008 |
MSN | 4 |
| 2022 | High-order local connection network for 3D human pose estimation based on GCN
Qiang Zhang 0008, Jing Dong 0009, Xiaopeng Wei |
Appl. Intell. | 3 |
| 2022 | Using entropy-driven amplifier circuit response to build nonlinear model under the influence of Lévy jumpabstractBACKGROUND: Bioinformatics is a subject produced by the combination of life science and computer science. It mainly uses computer technology to study the laws of biological systems. The design and realization of DNA circuit reaction is one of the important contents of bioinformatics. RESULTS: In this paper, nonlinear dynamic system model with Lévy jump based on entropy-driven amplifier (EDA) circuit response is studied. Firstly, nonlinear biochemical reaction system model is established based on EDA circuit response. Considering the influence of disturbance factors on the system, nonlinear biochemical reaction system with Lévy jump is built. Secondly, in order to prove that the constructed system conforms to the actual meaning, the existence and uniqueness of the system solution is analyzed. Next, the sufficient conditions for the end and continuation of EDA circuit reaction are certified. Finally, the correctness of the theoretical results is proved by numerical simulation, and the reactivity of THTSignal in EDA circuit under different noise intensity is verified. CONCLUSIONS: In EDA circuit reaction, the intensity of external noise has a significant impact on the system. The end of EDA circuit reaction is closely related to the intensity of Lévy noise, and Lévy jump has a significant impact on the nature of biochemical reaction system. Qiang Zhang 0008 |
BMC Bioinform. | 3 |
| 2022 | A meta-inspired termite queen algorithm for global optimization and engineering design problems
Shihua Zhou, Qiang Zhang 0008, Nikola K. Kasabov |
Eng. Appl. Artif. Intell. | 3 |
| 2022 | A2E2: Aerial-assisted energy-efficient edge sensing in intelligent public transportation systems
Pengfei Wang 0013, Zhaohong Yan, Guangjie Han, Yian Zhao, Chi Lin 0001, Ning Wang 0002, Qiang Zhang 0008 |
J. Syst. Archit. | 8 |
| 2022 | Blockchain-Enhanced Federated Learning Market With Social Internet of ThingsabstractThe machine learning performance usually could be improved by training with massive data. However, requesters can only select a subset of devices with limited training data to execute federated learning (FL) tasks as a result of their limited budgets in today’s IoT scenario. To resolve this pressing issue, we devise a blockchain-enhanced FL market (BFL) to$(i)$make data in computationally bounded devices available for training with social Internet of things,$(ii)$maximize the amount of training data with given budgets for an FL task, and$(iii)$decentralize the FL market with blockchain. To achieve these goals, we firstly propose a trust-enhanced collaborative learning strategy (TCL) and a quality-oriented task allocation algorithm (QTA), where TCL enables training data sharing among trusted devices with social Internet of things, and QTA allocates suitable devices to execute FL tasks while maximizing the training quality with fixed budgets. Then, we devise an encrypted model training scheme (EMT) based on a simple but countervailable differential privacy methodology to prevent attacks from malicious devices. In addition, we also propose a contribution-driven delegated proof of stake (DPoS) consensus mechanism to guarantee the fairness of reward distribution in the block generation process. Finally, extensive evaluations are conducted to verify the proposed BFL could improve the total utility of requesters and average accuracy of FL models significantly. Pengfei Wang 0013, Yian Zhao, Mohammad S. Obaidat, Zongzheng Wei, Heng Qi, Chi Lin 0001, Yunming Xiao, Qiang Zhang 0008 |
IEEE J. Sel. Areas Commun. | 8 |
| 2022 | Harmonized system code prediction of import and export commodities based on Hybrid Convolutional Neural Network with Auxiliary Network
Chao Che, Xi Sheryl Zhang, Qiang Zhang 0008 |
Knowl. Based Syst. | 4 |
| 2022 | Adaptive kernel selection network with attention constraint for surgical instrument classificationabstractAbstract Computer vision (CV) technologies are assisting the health care industry in many respects, i.e., disease diagnosis. However, as a pivotal procedure before and after surgery, the inventory work of surgical instruments has not been researched with the CV-powered technologies. To reduce the risk and hazard of surgical tools’ loss, we propose a study of systematic surgical instrument classification and introduce a novel attention-based deep neural network called SKA-ResNet which is mainly composed of: (a) A feature extractor with selective kernel attention module to automatically adjust the receptive fields of neurons and enhance the learnt expression and (b) A multi-scale regularizer with KL-divergence as the constraint to exploit the relationships between feature maps. Our method is easily trained end-to-end in only one stage with few additional calculation burdens. Moreover, to facilitate our study, we create a new surgical instrument dataset called SID19 (with 19 kinds of surgical tools consisting of 3800 images) for the first time. Experimental results show the superiority of SKA-ResNet for the classification of surgical tools on SID19 when compared with state-of-the-art models. The classification accuracy of our method reaches up to 97.703%, which is well supportive for the inventory and recognition study of surgical tools. Also, our method can achieve state-of-the-art performance on four challenging fine-grained visual classification datasets. Yaqing Hou, Qian Liu 0001, Hong-Wei Ge, Jun Meng, Qiang Zhang 0008, Xiaopeng Wei |
Neural Comput. Appl. | 6 |
| 2022 | Perception-oriented Single Image Super-Resolution Network with Receptive Field Block
Yaqing Hou, Wanshu Fan, Xin Yang 0011, Qiang Zhang 0008, Xiaopeng Wei |
Neural Comput. Appl. | 6 |
| 2022 | Designing Uncorrelated Address Constrain for DNA Storage by DMVO AlgorithmabstractAt present, huge amounts of data are being produced every second, a situation that will gradually overwhelm current storage technology. DNA is a storage medium that features high storage density and long-term stability and is now considered to be a feasible storage solution. Errors are easily made during the sequencing and synthesis of DNA, however. In order to reduce the error rate, novel uncorrelated address constrain are reported, and a Damping Multi-Verse Optimizer (DMVO)algorithm is proposed to construct a set of DNA coding, which is used as the non-payload. The DMVO algorithm exchanges objects through black/white holes in order to achieve a stable state and adds damping factors as disturbances. Compared with previous work, the coding set obtained by the DMVO algorithm is larger in size and of higher quality. The results of this study reveal that the size of the DNA storage coding set obtained by the DMVO algorithm increased by 4-16 percent, and the variance of the melting temperature decreased by 3-18 percent. Ben Cao, Xue Ii, Bin Wang 0005, Qiang Zhang 0008, Xiaopeng Wei |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2022 | Design of Constraint Coding Sets for Archive DNA StorageabstractWith the advent of the era of massive data, the increase of storage demand has far exceeded current storage capacity. DNA molecules provide a reliable solution for big data storage by virtue of their large capacity, high density, and long-term stability. To reduce errors in storing procedures, constructing a sufficient set of constraint encoding is critical for achieving DNA storage. A new version of the Marine Predator algorithm (called QRSS-MPA) is proposed in this paper to increase the lower bound of the coding set while satisfying the specific combination of constraints. In order to demonstrate the effectiveness of the improvement, the classical CEC-05 test function is used to test and compare the mean, variance, scalability, and significance. In terms of storage, the lower bound of construction is compared with previous works, and the result is found to be significantly improved. In order to prevent the emergence of a secondary structure that leads to sequencing failure, we give a more stringent lower bound for the constraint coding set, which is of great significance for reducing the error rate of DNA storage amidst its rapid development. Qiang Yin 0006, Yanfen Zheng, Bin Wang 0005, Qiang Zhang 0008 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2022 | A Novel Adaptive Linear Neuron Based on DNA Strand Displacement Reaction NetworkabstractAnalog DNA strand displacement circuits can be used to build artificial neural network due to the continuity of dynamic behavior. In this study, DNA implementations of novel catalysis, novel degradation and adjustment reaction modules are designed and used to build an analog DNA strand displacement reaction network. A novel adaptive linear neuron (ADALINE) is constructed by the ordinary differential equations of an ideal formal chemical reaction network, which is built by reaction modules. When reaction network approaches equilibrium, the weights of the ADALINE are updated without learning algorithm. Simulation results indicate that, ADALINE based on the analog DNA strand displacement circuit has ability to implement the learning function of the ADALINE based on the ideal formal chemical reaction networks, and fit a class of linear function. Xiaopeng Wei, Qiang Zhang 0008, Changjun Zhou |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2022 | Synchronization of Hyper-Lorenz System Based on DNA Strand DisplacementabstractLorenz system is depicted by chemical reaction equations of an ideal formal chemical reaction network, and a series of reversible reactions are added into chemical reaction network in order to construct a cluster of hyper-Lorenz system. DNA as a universal substrate for chemical dynamics can approximate arbitrary dynamical characteristics of ideal formal chemical reaction network through auxiliary DNA strands and displacement reactions. Based on Lyapunov's stableness theory, a novel synchronization strategy is proposed. A 6-dimensional hyper-Lorenz system is taken as examples for simulation and shows that DNA strands displacement reactions can implement the synchronization of ideal formal chemical reaction networks. Numerical simulations indicate that synchronization based on DNA strand displacement is robust to the detection of DNA strand concentration, control of reaction rate, and noise. Qiang Zhang 0008, Xiaopeng Wei |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2022 | Exploring Dense Context for Salient Object DetectionabstractContexts play an important role in salient object detection (SOD). High-level contexts describe the relations between different parts/objects and thus are helpful for discovering the specific locations of salient objects while low-level contexts could provide the fine detail information for delineating the boundary of the salient objects. However, the way of perceiving/leveraging rich contexts has not been fully investigated by existing SOD works. The common context extraction strategies (e.g., leveraging convolutions with large kernels or atrous convolutions with large dilation rates) do not consider the effectiveness and efficiency simultaneously and may cause sub-optimal solutions. In this paper, we devote to exploring an effective and efficient way to learn rich contexts for accurate SOD. Specifically, we first build a dense context exploration (DCE) module to capture dense multi-scale contexts and further leverage the learned contexts to enhance the features discriminability. Then, we embed multiple DCE modules in an encoder-decoder architecture to harvest dense contexts of different levels. Furthermore, we propose an attentive skip-connection to transmit useful features from the encoder part to the decoder part for better dense context exploration. Finally, extensive experiments demonstrate that the proposed method achieves more superior detection results on the six benchmark datasets than 18 state-of-the-art SOD methods. Haiyang Mei, Ziqi Wei 0001, Xiaopeng Wei, Qiang Zhang 0008, Xin Yang 0011 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2022 | From Pixels to Semantics: Self-Supervised Video Object Segmentation With Multiperspective Feature MiningabstractExisting self-supervised methods pose one-shot video object segmentation (O-VOS) as pixel-level matching to enable segmentation mask propagation across frames. However, the two tasks are not fully equivalent since O-VOS is more reliant on semantic correspondence rather than accurate pixel matching. To remedy this issue, we explore a new self-supervised framework that integrates pixel-level correspondence learning with semantic-level adaptation. The pixel-level correspondence learning is performed through photometric reconstruction of adjacent RGB frames during offline training, while semantic-level adaption operates at test-time by enforcing a bi-directional agreement of the predicted segmentation masks. In addition, we further propose a new network architecture with multi-perspective feature mining mechanism which can not only enhance reliable features but also suppress noisy ones to facilitate more robust image matching. By training the network using the proposed self-supervised framework, we achieve state-of-the-art performance on widely adopted datasets, further closing up the gap between self-supervised learning methods and their fully supervised counterparts. Ruoqi Li, Yifan Wang 0004, Lijun Wang 0001, Huchuan Lu, Xiaopeng Wei, Qiang Zhang 0008 |
IEEE Trans. Image Process. | 6 |
| 2022 | Trading off Charging and Sensing for Stochastic Events Monitoring in WRSNsabstractAs an epoch-making technology, wireless power transfer incredibly achieves energy transmission wirelessly, enabling reliable energy supplement for Wireless Rechargeable Sensor Networks (WRSNs). Existing methods mainly concentrate on performance improvement theoretically, neglecting the fact that most Commercial Off-The-Shelf (COTS) rechargeable sensors (e.g., WISP and Powercast) are not allowed to conduct sensing and energy harvesting tasks simultaneously, termedcharging exclusivity. Therefore, their schemes are not feasible for practical applications. In this paper, we focus on the charging exclusivity issue in stochastic events monitoring while improving network performance. In specific, we pay close attention to trading off charging and sensing tasks and formulate a combinatorial optimization problem with routing constraints. We introduce novel discretization techniques and investigate the routing problem to reformulate the original problem into maximization of a submodular function. With a slightly relaxed budget, the output of our proposed algorithm is better than$(1-1/e)/2$of the optimal solution to the original problem with a smaller charging radius$(1-\xi)D_{c}$. Through extensive simulations, numerical results show that in terms of charging utility, our algorithm outperforms baseline algorithms by 21.3% on average. Moreover, we conduct test-bed experiments to demonstrate the feasibility of our scheme in real scenarios. Yu Sun 0077, Chi Lin 0001, Haipeng Dai 0001, Pengfei Wang 0013, Lei Wang 0005, Guowei Wu 0001, Qiang Zhang 0008 |
IEEE/ACM Trans. Netw. | 7 |
| 2021 | A Novel Gaze-Point-Driven HRI Framework for Single-Person
Qiang Zhang 0008, Xiaopeng Wei, Rui Liu 0015, Jing Dong 0009 |
CollaborateCom (1) | 4 |
| 2021 | A Novel and Efficient Distance Detection Based on Monocular Images for Grasp and Handover
Dianwen Liu, Qiang Zhang 0008, Xiaopeng Wei, Rui Liu 0015, Jing Dong 0009 |
CollaborateCom (1) | 4 |
| 2021 | Human-robot Interaction Method Combining Human Pose Estimation and Motion Intention RecognitionabstractAlthough human pose estimation technology based on RGB images is becoming more and more mature, most of the current mainstream methods rely on depth camera to obtain human joints information. These interaction frameworks are affected by the infrared detection distance so that they cannot well adapt to the interaction scene of different distance. Therefore, the purpose of this paper is to build a modular interactive framework based on RGB images, which aims to alleviate the problem of high dependence on depth camera and low adaptability to distance in the current human-robot interaction (HRI) framework based on human body by using advanced human pose estimation technology. To enhance the adaptability of the HRI framework to different distances, we adopt optical cameras instead of depth cameras as acquisition equipment. Firstly, the human joints information is extracted by a human pose estimation network. Then, a joints sequence filter is designed in the intermediate stage to reduce the influence of unreasonable skeletons on the interaction results. Finally, a human intention recognition model is built to recognize the human intention from reasonable joints information, and drive the robot to respond according to the predicted intention. The experimental results show that our interactive framework is more robust in the distance than the framework based on depth camera and is able to achieve effective interaction under different distances, illuminations, costumes, customers, and scenes. Yalin Cheng, Rui Liu 0015, Jing Dong 0009, Qiang Zhang 0008 |
CSCWD | 6 |
| 2021 | Asymmetric Anomaly Detection for Human-Robot InteractionabstractSecurity in human-robot interaction is the focus of research in this field. Rapid detection of abnormal events that may cause danger in the interaction process can effectively reduce the probability of occurrence of danger. In general anomaly detection methods, 2D or 3D convolutional autoencoders are widely used for anomaly detection. Among them, 2D convolutional autoencoders are with good real-time performance and lower detection accuracy, while 3D convolutional autoencoders are with higher detection accuracy and insufficient real-time performance. In order to ensure realtime performance and obtain higher accuracy, an end-to-end asymmetric convolutional autoencoder network (ACANet) using both 2D and 3D convolutions is designed. Specifically, 3D convolution is used to build the encoder to learn comprehensive information in continuous input frames, and 2D convolution is used to build the decoder to model the information fast, a dimensional alignment module is constructed to connect the encoder and the decoder while avoiding a large number of calculations in the latent space of the 3D features output by the encoder, and the skip connections module is used to obtain accurate predictions. Anomaly detection can then be completed by evaluating the differences between results predicted by the ACANet and real frames. The experimental results show that our method achieves competitive accuracy on mainstream datasets and at the same time obtains the fastest speed. Compared with mainstream methods, this method is more suitable for anomaly detection tasks in human-robot interaction. Rui Liu 0015, Yingkun Hou, Qiang Zhang 0008, Xiaopeng Wei |
CSCWD | 6 |
| 2021 | Deep Reinforcement Learning Visual Navigation Model Integrating Memory-prediction MechanismabstractDeep reinforcement learning (DRL) has been widely used in the field of visual navigation. However, due to the lack of adaptability of DRL to the new tasks, the generalization ability of current visual navigation model using DRL is not desired. In order to improve this deficiency, we introduce the memory-prediction mechanism. By enhancing the memory of the scene, and combining the past experience of navigation to predict the next state, a more reasonable action can be obtained. First, we pass the image features extracted during the navigation process to an LSTM, and use LSTM to memorize the scene information in the image features. Then, we combine all the information (including state, target, and action) of each time step in the navigation process, and pass the historical information of multiple time steps to another LSTM to predict the next state. The action performed by the robot is determined by the predicted state. We use the AI2-THOR framework to carry out experiments. The results show that the proposed method can improve the navigation performance of the DRL visual navigation model and improve its adaptability to new tasks. Rui Liu 0015, Jing Dong 0009, Qiang Zhang 0008 |
CSCWD | 6 |
| 2021 | Depth-Aware Mirror SegmentationabstractWe present a novel mirror segmentation method that leverages depth estimates from ToF-based cameras as an additional cue to disambiguate challenging cases where the contrast or relation in RGB colors between the mirror reflection and the surrounding scene is subtle. A key observation is that ToF depth estimates do not report the true depth of the mirror surface, but instead return the total length of the reflected light paths, thereby creating obvious depth dis-continuities at the mirror boundaries. To exploit depth information in mirror segmentation, we first construct a large-scale RGB-D mirror segmentation dataset, which we subsequently employ to train a novel depth-aware mirror segmentation framework. Our mirror segmentation framework first locates the mirrors based on color and depth discontinuities and correlations. Next, our model further refines the mirror boundaries through contextual contrast taking into account both color and depth information. We extensively validate our depth-aware mirror segmentation method and demonstrate that our model outperforms state-of-the-art RGB and RGB-D based methods for mirror segmentation. Experimental results also show that depth is a powerful cue for mirror segmentation. Haiyang Mei, Bo Dong 0004, Wen Dong 0008, Pieter Peers, Xin Yang 0011, Qiang Zhang 0008, Xiaopeng Wei |
CVPR | 6 |
| 2021 | A Study on Realtime Task Selection Based on Credit Information Updating in Evolutionary Multitasking
Yumeng Cao, Yaqing Hou, Liang Feng 0001, Hong-Wei Ge, Qiang Zhang 0008, Xiaopeng Wei |
EMO | 5 |
| 2021 | Attention-Guided Second-Order Pooling Convolutional NetworksabstractRecently, channel attention-guided convolutional networks (ConvNets) have shown great advance on visual recognition tasks. However, they mainly exploit coarse first-order statistics to characterize holistic image and rarely focus on long-range feature dependencies, which limits the representation power in a certain. To handle above limitations, this paper proposes a novel attention-guided second-order pooling convolutional network (ASP-Net). ASP-Net introduces bilinear pooling that captures pairwise feature interactions to model second-order statistics. Meanwhile, it explicitly collects long-range dependencies via non-local operations, thus providing a global view in lower layers. Then, the second-order statistics and non-local context features are fused to obtain the enhanced representation for predicting channel-wise attention map and scaling convolution features. Experiment results on three commonly used datasets illuminate that ASP-Net outperforms its counterparts and achieves competitive performance. The source code is available at https://github.com/ShannanChen/ASPNet. Shannan Chen, Qiule Sun, Cunhua Li, Jianxin Zhang 0001, Qiang Zhang 0008 |
ICASSP | 5 |
| 2021 | Multi-Scale Attention Constraint Network for Fine-Grained Visual ClassificationabstractCapturing subtle yet discriminative features constitutes a great challenge in fine-grained visual classification due to the large intra-class and small inter-class variances. Main-stream works for this problem localize at attention mechanism and feature relationship learning. However, existing methods treat the features in isolation while neglecting the effect of attention-enhanced features on relationships between different network layers. In this paper, we propose a novel attention-based method by Multi-Scale Attention Constraint network composed of two important components: (1) a feature extractor with lightweight group-wise enhanced attention blocks that guides the generation of high representation features; and (2) a multi-scale regularizer that explores the relationships between different features. Extensive experiments show that our approach achieves state-of-the-art performance on standard benchmark datasets. Moreover, we introduce a new dataset, consisting of comprehensive surgical instrument categories based on three common surgeries, to support the classification and inventory work of surgical instruments. Yaqing Hou, Hong-Wei Ge, Qiang Zhang 0008, Xiaopeng Wei |
ICME | 5 |
| 2021 | Weakly Supervised Gleason Grading of Prostate Cancer Slides using Graph Neural Network
Yaqing Hou, Pengfei Wang 0013, Jianxin Zhang 0001, Qiang Zhang 0008 |
ICPRAM | 6 |
| 2021 | Brain Tumor Segmentation based on Knowledge Distillation and Adversarial Trainingabstract3D MRI brain tumor segmentation is a reliable method for disease diagnosis and treatment plans in the future. Early on, the segmentation of brain tumors is mostly done manually. However, manual segmentation of 3D MRI brain tumor requires professional anatomical knowledge and may be inaccurate. In this paper, we propose a 3D MRI brain tumor segmentation architecture based on the encoder-decoder structure. Specially, we introduce knowledge distillation and adversarial training methods, which compresses and improves the accuracy and robustness of the model. Furthermore, we obtain soft targets by designing multiple teacher network training and then apply them to the student network. Finally, we evaluate our method on a challenging BraTS dataset. As a result, the performance of our proposed model is superior to state-of-the-art methods. Yaqing Hou, Tianbo Li, Qiang Zhang 0008, Hua Yu 0006, Hong-Wei Ge |
IJCNN | 3 |
| 2021 | A Contextual Attention Network for Multimodal Emotion Recognition in ConversationabstractEmotion recognition in conversation (ERC) is a challenging task due to the complexity of emotions and dynamics in dialogues. Current studies for emotion recognition mostly focus on the modeling of a single utterance in dialogue, which neglects self and inter-speaker influence. This paper presents a contextual attention neural network based on the multimodal framework that leverages the conversational information from both target and the other speaker for utterance-level emotion detection. Specifically, we utilize recurrent neural networks based on contextual attention for modeling the transaction and dependence between speakers. Further, the feature fusion is proposed to unite the important modal information extracted from multiple modalities, including audio, text and video, hence providing more useful and comprehensive knowledge for emotion recognition. The proposed approach shows its superiority in extracting contexts for self and inter-speaker influence and synthesizing them as global features that are beneficial to detect individual emotion state. Experiment result on the IEMOCAP corpus reports an accuracy of 64.6%, demonstrating the superiority of the proposed method in emotion recognition comparing to the state-of-the-arts. Tana Wang, Yaqing Hou, Qiang Zhang 0008 |
IJCNN | 4 |
| 2021 | Exploring the effects of computational costs in extensive games via modeling and simulationabstractGame theory has become a standard tool for depicting and demonstrating various game-like phenomena by providing appropriate mathematical models and for analyzing and predicting agents' behaviors and their decisions by formalizing solution concepts. The conventional game model mainly concerns ideal systems that would always guarantee optimal responses, which appears unrealistic for practical game scenarios since decision-making usually entails resource costs. Therefore, this study considers players' decision-making in extensive games when the computational cost of searching the strategy space is limited. We start with a new mathematical model of extensive games that features a bound on computational resources during players' decision-making process such that they can only foresee a part of the available alternatives in the future. This model is more appropriate in predicting players' strategies than the conventional model, under which we investigate the effects of computational costs on players' strategies as well as the computational complexity. Furthermore, a simulation experiment is performed to seek the connection between the amount of resources and the goodness of the outcomes. This study is expected to provide a foundation for players' rational decision-making with computational costs. Chanjuan Liu 0001, Enqiang Zhu, Qiang Zhang 0008, Xiaopeng Wei |
Int. J. Intell. Syst. | 3 |
| 2021 | DNA Computing Model for Satisfiability Problem Based on Hybridization Chain ReactionabstractSatisfiability problem is a famous nondeterministic polynomial-time complete (NP-complete) problem, which has always been a hotspot in artificial intelligence. In this paper, by combining the advantages of DNA origami with hybridization chain reaction, a computing model was proposed to solve the satisfiability problem. For each clause in the given formula, a DNA origami device was devised. The device corresponding to the clause was capable of searching for assignments that satisfied the clause. When all devices completed the search in parallel, the intersection of these satisfying assignments found must satisfy all the clauses. Therefore, whether the given formula is satisfiable or not was decided. The simulation results demonstrated that the proposed computing model was feasible. Our work showed the capability of DNA origami in architecting automatic computing device. The paper proposed a novel method for designing functional nanoscale devices based on DNA origami. Jing Yang 0037, Qiang Zhang 0008, Zhongtuan Zheng |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2021 | AI-Driven Collaborative Resource Allocation for Task Execution in 6G-Enabled Massive IoTabstractIn the foreseeable future, the rapid growth of devices in the Internet of Things (IoT) will make it difficult for 5G networks to ensure sufficient network resources. 6G technology has attracted increasing attention, bringing new design concepts to the dynamic real-time resource allocation. The resource requirements of devices are usually variable, so a dynamic resource allocation method is needed to ensure the smooth execution of tasks. Therefore, this article first designs a 6G-enabled massive IoT architecture that supports dynamic resource allocation. Then, a dynamic nested neural network is constructed, which adjusts the nested learning model structure online to meet training requirements of dynamic resource allocation. An AI-driven collaborative dynamic resource allocation (ACDRA) algorithm is proposed based on the nested neural network combined with Markov decision process training for 6G-enabled massive IoT. Extensive simulations have been carried out to evaluate ACDRA in terms of several performance criteria, including resource hit rate and decision delay time. The results validated that ACDRA improves the average resource hit rate by about 8% and reduces the average decision delay time by about 7% compared with three reference existing algorithms. Qiang Zhang 0008, Giancarlo Fortino |
IEEE Internet Things J. | 3 |
| 2021 | Contact Tracing Incentive for COVID-19 and Other Pandemic Diseases From a Crowdsourcing PerspectiveabstractGovernments of the world have invested a lot of manpower and material resources to combat COVID-19 this year. At this moment, the most efficient way that could stop the epidemic is to leverage the contact tracing system to monitor people's daily contact information and isolate the close contacts of COVID-19. However, the contact tracing data usually contains people's sensitive information that they do not want to share with the contact tracing system and government. Conversely, the contact tracing system could perform better when it obtains more detailed contact tracing data. In this article, we treat the process of collecting contact tracing data from a crowdsourcing perspective in order to motivate users to contribute more contact tracing data and propose the incentive algorithm named CovidCrowd. Different from previous works where they ask users to contribute their data voluntarily, the government offers some reward to users who upload their contact tracing data to reimburse the privacy and data processing cost. We formulate the problem as a Stackelberg game and show there exists a Nash equilibrium for any user given the fixed reward value. Then, CovidCrowd computes the optimal reward value which could maximize the utility of the system. Finally, we conduct a large-scale simulation with thousands of users and evaluation with real-world data set. Both results show that CovidCrowd outperforms the benchmarks, e.g., the user participating level is improved by at least 13.2% for all evaluation scenarios. Pengfei Wang 0013, Chi Lin 0001, Mohammad S. Obaidat, Ziqi Wei 0001, Qiang Zhang 0008 |
IEEE Internet Things J. | 6 |
| 2021 | ASFNet: Adaptive multiscale segmentation fusion network for real-time semantic segmentationabstractAbstract Recently, the development of deep learning has facilitated continuous progress in the field of computer vision. Pixel‐level semantic segmentation serves as a fundamental task in computer vision. It achieves significant results by connecting wider and deeper backbone networks and building fine‐grained segmentation heads. However, applications such as self‐driving cars are more critical to the computational speed of the algorithms. The trade‐off between accuracy and real‐time performance of existing algorithms is still a challenging task. To address this challenge, this article proposes an adaptive multiscale segmentation fusion network to fuse multiscale contextual, which designs an adaptive multiscale segmentation fusion module based on an attention mechanism. Using segmentation fusion instead of feature fusion, the multiscale segmentation results are aggregated to obtain more precise segmentation results. The final results achieved 70.9% mIoU of accuracy in the Cityspace test set, processing images at 61 FPS when the input is 1024 × 2048. In addition, when adjusting the input size to 512 × 1024, the images are processed at 185 FPS. Hengfeng Zha, Rui Liu 0015, Xin Yang 0011, Qiang Zhang 0008, Xiaopeng Wei |
Comput. Animat. Virtual Worlds | 5 |
| 2021 | Local-aware spatio-temporal attention network with multi-stage feature fusion for human action recognitionabstractAbstract In the study of human action recognition, two-stream networks have made excellent progress recently. However, there remain challenges in distinguishing similar human actions in videos. This paper proposes a novel local-aware spatio-temporal attention network with multi-stage feature fusion based on compact bilinear pooling for human action recognition. To elaborate, taking two-stream networks as our essential backbones, the spatial network first employs multiple spatial transformer networks in a parallel manner to locate the discriminative regions related to human actions. Then, we perform feature fusion between the local and global features to enhance the human action representation. Furthermore, the output of the spatial network and the temporal information are fused at a particular layer to learn the pixel-wise correspondences. After that, we bring together three outputs to generate the global descriptors of human actions. To verify the efficacy of the proposed approach, comparison experiments are conducted with the traditional hand-engineered IDT algorithms, the classical machine learning methods (i.e., SVM) and the state-of-the-art deep learning methods (i.e., spatio-temporal multiplier networks). According to the results, our approach is reported to obtain the best performance among existing works, with the accuracy of 95.3% and 72.9% on UCF101 and HMDB51, respectively. The experimental results thus demonstrate the superiority and significance of the proposed architecture in solving the task of human action recognition. Yaqing Hou, Hua Yu 0006, Pengfei Wang 0013, Hong-Wei Ge, Jianxin Zhang 0001, Qiang Zhang 0008 |
Neural Comput. Appl. | 7 |
| 2021 | Multi-Scale Context-Guided Deep Network for Automated Lesion Segmentation With Endoscopy Images of Gastrointestinal TractabstractAccurate lesion segmentation based on endoscopy images is a fundamental task for the automated diagnosis of gastrointestinal tract (GI Tract) diseases. Previous studies usually use hand-crafted features for representing endoscopy images, while feature definition and lesion segmentation are treated as two standalone tasks. Due to the possible heterogeneity between features and segmentation models, these methods often result in sub-optimal performance. Several fully convolutional networks have been recently developed to jointly perform feature learning and model training for GI Tract disease diagnosis. However, they generally ignore local spatial details of endoscopy images, as down-sampling operations (e.g., pooling and convolutional striding) may result in irreversible loss of image spatial information. To this end, we propose a multi-scale context-guided deep network (MCNet) for end-to-end lesion segmentation of endoscopy images in GI Tract, where both global and local contexts are captured as guidance for model training. Specifically, one global subnetwork is designed to extract the global structure and high-level semantic context of each input image. Then we further design two cascaded local subnetworks based on output feature maps of the global subnetwork, aiming to capture both local appearance information and relatively high-level semantic information in a multi-scale manner. Those feature maps learned by three subnetworks are further fused for the subsequent task of lesion segmentation. We have evaluated the proposed MCNet on 1,310 endoscopy images from the public EndoVis-Ab and CVC-ClinicDB datasets for abnormal segmentation and polyp segmentation, respectively. Experimental results demonstrate that MCNet achieves [Formula: see text] and [Formula: see text] mean intersection over union (mIoU) on two datasets, respectively, outperforming several state-of-the-art approaches in automated lesion segmentation with endoscopy images of GI Tract. Shuai Wang 0003, Yang Cong, Hancan Zhu, Xianyi Chen, Liangqiong Qu, Huijie Fan, Qiang Zhang 0008, Mingxia Liu 0001 |
IEEE J. Biomed. Health Informatics | 7 |
| 2021 | Automatic Comic Generation with Stylistic Multi-page Layouts and Emotion-driven Text Balloon GenerationabstractIn this article, we propose a fully automatic system for generating comic books from videos without any human intervention. Given an input video along with its subtitles, our approach first extracts informative keyframes by analyzing the subtitles and stylizes keyframes into comic-style images. Then, we propose a novel automatic multi-page layout framework that can allocate the images across multiple pages and synthesize visually interesting layouts based on the rich semantics of the images (e.g., importance and inter-image relation). Finally, as opposed to using the same type of balloon as in previous works, we propose an emotion-aware balloon generation method to create different types of word balloons by analyzing the emotion of subtitles and audio. Our method is able to vary balloon shapes and word sizes in balloons in response to different emotions, leading to more enriched reading experience. Once the balloons are generated, they are placed adjacent to their corresponding speakers via speaker detection. Our results show that our method, without requiring any user inputs, can generate high-quality comic pages with visually rich layouts and balloons. Our user studies also demonstrate that users prefer our generated results over those by state-of-the-art comic generation systems. Xin Yang 0011, Zongliang Ma, Letian Yu, Ying Cao 0001, Xiaopeng Wei, Qiang Zhang 0008, Rynson W. H. Lau |
ACM Trans. Multim. Comput. Commun. Appl. | 7 |
| 2021 | Smart Scribbles for Image MattingabstractImage matting is an ill-posed problem that usually requires additional user input, such as trimaps or scribbles. Drawing a fine trimap requires a large amount of user effort, while using scribbles can hardly obtain satisfactory alpha mattes for non-professional users. Some recent deep learning–based matting networks rely on large-scale composite datasets for training to improve performance, resulting in the occasional appearance of obvious artifacts when processing natural images. In this article, we explore the intrinsic relationship between user input and alpha mattes and strike a balance between user effort and the quality of alpha mattes. In particular, we propose an interactive framework, referred to as smart scribbles, to guide users to draw few scribbles on the input images to produce high-quality alpha mattes. It first infers the most informative regions of an image for drawing scribbles to indicate different categories (foreground, background, or unknown) and then spreads these scribbles (i.e., the category labels) to the rest of the image via our well-designed two-phase propagation. Both neighboring low-level affinities and high-level semantic features are considered during the propagation process. Our method can be optimized without large-scale matting datasets and exhibits more universality in real situations. Extensive experiments demonstrate that smart scribbles can produce more accurate alpha mattes with reduced additional input, compared to the state-of-the-art matting methods. Xin Yang 0011, Yu Qiao 0001, Shaozhe Chen, Shengfeng He, Qiang Zhang 0008, Xiaopeng Wei, Rynson W. H. Lau |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2020 | Efficient Attention Calibration Network for Real-Time Semantic SegmentationabstractIn recent years, the attention mechanism has been widely used in computer vision. Semantic segmentation, as one of the fundamental tasks of computer vision, has been subject to tremendous development as a result. But because of its huge computing overhead, attention-based approaches are difficult to use for real-time applications such as self-driving. In this paper, we propose a self-calibration method baesd on self-attentiion that successfully applies the attention mechanism to real-time semantic segmentation. Specifically, a spatial attention module to adjust the edges of the coarse segmentation results which gained from the real-time semantic segmentation backbone network, and obtain more granular segmentation results. We refer to this method as the Efficient Attentional Calibration Network (EACNet). Experiments on the Cityscapes dataset validate the efficiency and performance of the method. With the high-resolution input and without any post-processing, EACNet achieved 72.4% mIoU of accuracy while running at 116.9 FPS. Compared to other state-of-the-art methods for real-time semantic segmentation, our network gained a better balance between performance and speed. Hengfeng Zha, Rui Liu 0015, Xin Yang 0011, Qiang Zhang 0008, Xiaopeng Wei |
ACML | 5 |
| 2020 | Memetic Multi-agent optimization with Problem Reformulation by Coordinate RotationabstractMemetic multi-agent system (MeMAS) is recently proposed as an enhanced version that integrates meme concept into multi-agent system (MAS) wherein all meme-inspired agents have an improvement in learning performance via meme evolution independently or social interaction. In the process of solving the black box optimization problem, the potential advantages of MeMAS have not been utilized well, which makes it a fertile area for further exploration. This paper presents a memetic multi-agent optimization paradigm through coordinate rotation (MeMAO-R) to combine MeMAS with evolutionary algorithms (EAs) to improve optimization efficiency. Based on MeMAS, the particular interest of MeMAO-R is placed on assisting original complex optimization task with new tasks generated by coordinate rotation. Further, MeMAO-R constructs the social interaction mechanism which facilitates to improve their convergence speed for solving the target optimization problem by utilizing meaningful information transferred across multiple agents with differing views of the target problem. Besides, MeMAO-R employs one or more classical EAs as the fundamental population based evolutionary solvers for multiple agents to optimize multiple tasks in a multi-agent scenario. Lastly, to testify the efficacy of the proposed MeMAO-R, comprehensive empirical studies on basic optimization problems are provided. Yaqing Hou, Qiang Zhang 0008, Hong-Wei Ge, Xin Yang 0011, Abhishek Gupta 0001, Xianneng Li |
CEC | 3 |
| 2020 | Don't Hit Me! Glass Detection in Real-World ScenesabstractGlass is very common in our daily life. Existing computer vision systems neglect it and thus may have severe consequences, e.g., a robot may crash into a glass wall. However, sensing the presence of glass is not straightforward. The key challenge is that arbitrary objects/scenes can appear behind the glass, and the content within the glass region is typically similar to those behind it. In this paper, we propose an important problem of detecting glass from a single RGB image. To address this problem, we construct a large-scale glass detection dataset (GDD) and design a glass detection network, called GDNet, which explores abundant contextual cues for robust glass detection with a novel large-field contextual feature integration (LCFI) module. Extensive experiments demonstrate that the proposed method achieves more superior glass detection results on our GDD test set than state-of-the-art methods fine-tuned for glass detection. Haiyang Mei, Xin Yang 0011, Yang Wang 0106, Shengfeng He, Qiang Zhang 0008, Xiaopeng Wei, Rynson W. H. Lau |
CVPR | 6 |
| 2020 | Attention-Guided Hierarchical Structure Aggregation for Image MattingabstractExisting deep learning based matting algorithms primarily resort to high-level semantic features to improve the overall structure of alpha mattes. However, we argue that advanced semantics extracted from CNNs contribute unequally for alpha perception and we are supposed to reconcile advanced semantic information with low-level appearance cues to refine the foreground details. In this paper, we propose an end-to-end Hierarchical Attention Matting Network (HAttMatting), which can predict the better structure of alpha mattes from single RGB images without additional input. Specifically, we employ spatial and channel-wise attention to integrate appearance cues and pyramidal features in a novel fashion. This blended attention mechanism can perceive alpha mattes from refined boundaries and adaptive semantics. We also introduce a hybrid loss function fusing Structural SIMilarity (SSIM), Mean Square Error (MSE) and Adversarial loss to guide the network to further improve the overall foreground structure. Besides, we construct a large-scale image matting dataset comprised of 59,600 training images and 1000 test images (total 646 distinct foreground alpha mattes), which can further improve the robustness of our hierarchical structure aggregation model. Extensive experiments demonstrate that the proposed HAttMatting can capture sophisticated foreground structure and achieve state-of-the-art performance with single RGB images as input. Yu Qiao 0001, Yuhao Liu 0001, Xin Yang 0011, Mingliang Xu 0001, Qiang Zhang 0008, Xiaopeng Wei |
CVPR | 6 |
| 2020 | D2D-Enabled Reliable Data Collection for Mobile Crowd SensingabstractWith increasing more powerful sensing capacities of mobile devices, the Mobile Crowd Sensing (MCS) system requires to collect larger sensing data from participants. Nevertheless, collecting such large volume of data will cost a lot for participants, base stations and MCS server. Even worse, some sensing data cannot satisfy the MCS sensing requirement due to the low quality and are filtered by the MCS server in clouds. Inspired by the D2D technique, where mobile devices can communicate directly with the help of the nearby base station, in 5G networks, we propose the Reliable Data Collection (RDC) algorithm to validate the generated sensing data at device sides in this paper. To be specific, the whole progress is formulated as a Probability problem of Discovering Reliable sensing data (PDR) at client sides, and Expectation Maximization (EM) is leveraged to devise the algorithm. Finally, the extensive simulations and real-world use case are conducted to evaluate the performance of RDC algorithm, and the result shows that RDC outperforms the other two benchmarks in estimating accuracy and saving data collection cost. Pengfei Wang 0013, Chi Lin 0001, Leyou Yang, Yaqing Hou, Qiang Zhang 0008 |
ICPADS | 6 |
| 2020 | Second-order Attention Guided Convolutional Activations for Visual RecognitionabstractRecently, modeling deep convolutional activations by the global second-order pooling has shown great advance on visual recognition tasks. However, most of the existing deep second-order statistical models mainly compute second-order statistics of activations of the last convolutional layer as image representations, and they seldom introduce second-order statistics into earlier layers to better fit network topology, thus limiting the representational ability to a certain extent. Motivated by the flexibility of attention blocks that are commonly plugged into intermediate layers of deep convolutional networks (ConvNets), this work makes an attempt to combine deep second-order statistics with attention mechanisms in ConvNets, and further proposes a novel Second-order Attention Guided Network (SoAG-Net) for visual recognition. More specifically, SoAG-Net involves several SoAG modules seemingly inserted into intermediate layers of the network, in which SoAG collects second-order statistics of convolutional activations by polynomial kernel approximation to predict channel-wise attention maps utilized for guiding the learning of convolutional activations through tensor scaling along channel dimension. SoAG improves the nonlinearity of ConvNets and enables ConvNets to fit more complicated distribution of convolutional activations. Experiment results on three commonly used datasets illuminate that SoAG-Net outperforms its counterparts and achieves competitive performance with state-of-the-art models under the same backbone. Shannan Chen, Qiule Sun, Bin Liu 0040, Jianxin Zhang 0001, Qiang Zhang 0008 |
ICPR | 6 |
| 2020 | Breast Cancer Histopathological Image Classification Based on Deep Second-order Pooling NetworkabstractWith the breakthrough performance in a variety of computer vision and medical image analysis problems, convolutional neural networks (CNNs) have been successfully introduced for the classification task of breast cancer histopathological images in recent years. Nevertheless, existing breast cancer histopathological image classification networks mainly utilize the first-order statistic information of deep features to represent histopathological images, failing to characterize the complex global feature distribution of breast cancer histopathological images. To address the problem, this work makes a first attempt to explore global second-order statistics of deep features for the above task. More specifically, we propose a novel deep second-order pooling network (DSoPN) for breast cancer histopatho-logical image classification, in which a robust global covariance pooling module based on matrix power normalization (MPN) is embedded into a simple yet effective CNN architecture. The given DSoPN model can capture richer second-order statistical information of deep convolutional features and produce more informative global representations for breast cancer histopatho-logical images. Experimental results on the public BreakHis dataset illuminate the promising performance of the second-order pooling for breast cancer histopathological image classification. Besides, our DSoPN achieves very competitive performance compared to the state-of-the-art methods. Jiasen Li, Jianxin Zhang 0001, Qiule Sun, Hengbo Zhang, Jing Dong 0009, Chao Che, Qiang Zhang 0008 |
IJCNN | 7 |
| 2020 | A Preliminary Study of Fusion ARTs with Adaptively Information Intensity Attenuation ControllingabstractFusion ART is an enhanced version of Adaptive Resonance Theory (ART) which is derived from a biologically-plausible theory of human cognitive information processing. Due to its well-established ability of learning associative mappings across multimodal pattern channels in an online and incremental manner, fusion ART has been widely applied in many real world learning problems. In this paper, we take a Fusion Architecture for Learning, Cognition, and Navigation (FALCON) as the specification and essential backbone of fusion ART and introduce an intensity attenuation controller δ for adaptively adjusting the intensity of information captured from the environment, by taking inspiration from Broadbent-Treisman Filter-Attenuation's perceptual model of environmental attention. Particularly, we propose both an adaptive δ detection algorithm as well as a δ-based pruning algorithm to enhance the learning performance of FALCON while reduce the redundant memory storage incurred by the "detrimental δ". To verify the effectiveness and efficiency of our proposed method, comprehensive experimental studies are carried out on a classical minefield navigation task. Wenxuan Zhu, Yaqing Hou, Qiang Zhang 0008, Hong-Wei Ge, Xin Yang 0011, Liang Feng 0001, Xinghua Qu |
IJCNN | 3 |
| 2020 | Multi-scale Information Assembly for Image MattingabstractAbstract Image matting is a long‐standing problem in computer graphics and vision, mostly identified as the accurate estimation of the foreground in input images. We argue that the foreground objects can be represented by different‐level information, including the central bodies, large‐grained boundaries, refined details, etc. Based on this observation, in this paper, we propose a multi‐scale information assembly framework (MSIA‐matte) to pull out high‐quality alpha mattes from single RGB images. Technically speaking, given an input image, we extract advanced semantics as our subject content and retain initial CNN features to encode different‐level foreground expression, then combine them by our well‐designed information assembly strategy. Extensive experiments can prove the effectiveness of the proposed MSIA‐matte, and we can achieve state‐of‐the‐art performance compared to most existing matting networks. Yu Qiao 0001, Yuhao Liu 0001, Xin Yang 0011, Yuxin Wang 0001, Qiang Zhang 0008, Xiaopeng Wei |
Comput. Graph. Forum | 6 |
| 2019 | Memetic Multi-agent Optimization in High Dimensions using Random EmbeddingsabstractIn this paper, we propose a memetic multi-agent optimization (MeMAO) paradigm to enhance the search efficacy of classical EAs (i.e., Differential Evolution (DE)) in solving the complex optimization problems. The essential backbone of MeMAO is a recently proposed memetic multi-agent learning system wherein agents acquire increasing learning capabilities by interacting with the environment mainly in a reinforcement learning manner. Differing from MeMAS, the particular interest of MeMAO is placed on addressing the specific challenges when applying classical EAs to optimize the high dimensional optimization problems with a "low effective dimensionality". To achieve this, the target optimization problem is firstly re-formulated into multiple low dimensional tasks via random embedding methods. Further, MeMAO employs DE as the fundamental population based evolutionary solver for multiple agents to optimize multiple low dimensional tasks in a multi-agent scenario. Importantly, MeMAO constructs the social interaction mechanisms among multiple agents, hence improves their convergence speed for solving the target optimization problem by sharing the beneficial information across multiple agents. Lastly, to testify the efficacy of the proposed MeMAO, comprehensive empirical studies on 8 synthetic optimization problems with a dimensionality of 2,000 are provided. Yaqing Hou, Hong-Wei Ge, Qiang Zhang 0008, Xinghua Qu, L. Feng, Abhishek Gupta 0001 |
CEC | 4 |
| 2019 | Spatial Attentive Single-Image Deraining With a High Quality Real Rain DatasetabstractRemoving rain streaks from a single image has been drawing considerable attention as rain streaks can severely degrade the image quality and affect the performance of existing outdoor vision tasks. While recent CNN-based derainers have reported promising performances, deraining remains an open problem for two reasons. First, existing synthesized rain datasets have only limited realism, in terms of modeling real rain characteristics such as rain shape, direction and intensity. Second, there are no public benchmarks for quantitative comparisons on real rain images, which makes the current evaluation less objective. The core challenge is that real world rain/clean image pairs cannot be captured at the same time. In this paper, we address the single image rain removal problem in two ways. First, we propose a semi-automatic method that incorporates temporal priors and human supervision to generate a high-quality clean image from each input sequence of real rain images. Using this method, we construct a large-scale dataset of ∼29.5K rain/rain-free image pairs that covers a wide range of natural rain scenes. Second, to better cover the stochastic distribution of real rain streaks, we propose a novel SPatial Attentive Network (SPANet) to remove rain streaks in a local-to-global manner. Extensive experiments demonstrate that our network performs favorably against the state-of-the-art deraining methods. Tianyu Wang 0003, Xin Yang 0011, Ke Xu 0010, Shaozhe Chen, Qiang Zhang 0008, Rynson W. H. Lau |
CVPR | 5 |
| 2019 | Deep Covariance Estimation Hashing for Image RetrievalabstractRecently, combination of advanced convolutional neural networks and efficient hashing, deep hashing have achieved impressive performance for image retrieval. However, state-of-the-art deep hashing methods mainly focus on constructing hash function, loss function and training strategies to preserve semantic similarity. For the fundamental image characteristics, they depend heavily on the first-order convolutional feature statistics, failing to take their global structure into consideration. To address this problem, we present a deep covariance estimation hashing (DCEH) method with robust covariance form to improve hash code quality. The core of DCEH involves covariance pooling as deep hashing representation performing global pairwise feature interactions. Due to convolutional features are usually high dimension and small sample size, we estimate robust covariance with matrix power normalization and then insert it into deep hashing paradigm in an end-to-end learning manner. Extensive experiments on three benchmarks show that the proposed DCEH outperforms its counterparts and achieves superior performance. Qiule Sun, Jianxin Zhang 0001, Jingdong Cheng, Bin Liu 0040, Qiang Zhang 0008 |
ICIP | 6 |
| 2019 | A Word Segmentation Method of Ancient Chinese Based on Word Alignment
Chao Che, Xiaoting Wu, Qiang Zhang 0008 |
NLPCC (1) | 5 |
| 2019 | Real-virtual consistent traffic flow interaction
Xin Yang 0011, Shuai Li 0014, Xinglin Piao, Qiang Zhang 0008, Xiaopeng Wei |
Graph. Model. | 6 |
| 2019 | Logic Operation Model of the Complementer Based on Two-domain DNA Strand DisplacementabstractDNA strand replacement technology has the advantages of simple operation which makes it becomes a common method of DNA computing. A four bit binary number Complementer based on two-domain DNA strand displacement is proposed in this paper. It implements the function of converting binary code into complement code. Simulation experiment based on Visual DSD software is carried out. The simulation results show the correctness and feasibility of the logic model of the Complementer, and it makes useful exploration for further expanding the application of molecular logic circuit. Wendan Xie, Changjun Zhou, Qiang Zhang 0008 |
Fundam. Informaticae | 4 |
| 2019 | DEMC: A Deep Dual-Encoder Network for Denoising Monte Carlo Rendering
Xin Yang 0011, Wenbo Hu 0002, Lijing Zhao, Qiang Zhang 0008, Xiaopeng Wei, Hongbo Fu 0001 |
J. Comput. Sci. Technol. | 6 |
| 2019 | Cascaded network with deep intensity manipulation for scene understandingabstractAbstract Scene understanding is essential to robotic navigation and autonomous driving as it provides semantic information to their controlling system. However, it will fail when processing low‐light images/videos captured under adverse weather or at night use state‐of‐the‐art scene understanding methods. A naive way to directly infer semantics from low‐light images is ill posed because the low‐light condition distorts pixel intensities and buries details. In order to address this problem, we propose the Deep Intensity Manipulation Network (DIMNet), which could relight the input images and recover the details, and combine the DIMNet with a scene understanding network to get a cascaded network to learn the semantics from low‐light images. Through learning pixel intensity manipulation, our method can generate images not only visually pleasing but also practical for scene understanding. Qualitative and quantitative experiments demonstrate that the proposed method is effective and robust for both synthetic and real‐world images. Xin Yang 0011, Shaozhe Chen, Xinglin Piao, Qiang Zhang 0008, Xiaopeng Wei |
Comput. Animat. Virtual Worlds | 6 |
| 2019 | A Many-Objective Evolutionary Algorithm With Two Interacting Processes: Cascade Clustering and Reference Point Incremental LearningabstractResearches have shown difficulties in obtaining proximity while maintaining diversity for many-objective optimization problems. Complexities of the true Pareto front pose challenges for the reference vector-based algorithms for their insufficient adaptability to the diverse characteristics with no priori. This paper proposes a many-objective optimization algorithm with two interacting processes: cascade clustering and reference point incremental learning (CLIA). In the population selection process based on cascade clustering (CC), using the reference vectors provided by the process based on incremental learning, the nondominated and the dominated individuals are clustered and sorted with different manners in a cascade style and are selected by round-robin for better proximity and diversity. In the reference vector adaptation process based on reference point incremental learning, using the feedbacks from the process based on CC, proper distribution of reference points is gradually obtained by incremental learning. Experimental studies on several benchmark problems show that CLIA is competitive compared with the state-of-the-art algorithms and has impressive efficiency and versatility using only the interactions between the two processes without incurring extra evaluations. Hong-Wei Ge, Mingde Zhao 0002, Liang Sun 0003, Zhen Wang 0004, Guozhen Tan, Qiang Zhang 0008, C. L. Philip Chen |
IEEE Trans. Evol. Comput. | 6 |
| 2019 | DRFN: Deep Recurrent Fusion Network for Single-Image Super-Resolution With Large FactorsabstractRecently, single-image super-resolution has made great progress due to the development of deep convolutional neural networks (CNNs). The vast majority of CNN-based models use a predefined upsampling operator, such as bicubic interpolation, to upscale input low-resolution images to the desired size and learn nonlinear mapping between the interpolated image and ground truth high-resolution (HR) image. However, interpolation processing can lead to visual artifacts as details are over smoothed, particularly when the super-resolution factor is high. In this paper, we propose a deep recurrent fusion network (DRFN), which utilizes transposed convolution instead of bicubic interpolation for upsampling and integrates different-level features extracted from recurrent residual blocks to reconstruct the final HR images. We adopt a deep recurrence learning strategy and, thus, have a larger receptive field, which is conducive to reconstructing an image more accurately. Furthermore, we show that the multilevel fusion structure is suitable for dealing with image super-resolution problems. Extensive benchmark evaluations demonstrate that the proposed DRFN performs better than most current deep learning methods in terms of accuracy and visual effects, especially for large-scale images, while using fewer parameters. Xin Yang 0011, Haiyang Mei, Jiqing Zhang, Ke Xu 0010, Qiang Zhang 0008, Xiaopeng Wei |
IEEE Trans. Multim. | 6 |
| 2018 | Image Correction via Deep Reciprocating HDR TransformationabstractImage correction aims to adjust an input image into a visually pleasing one. Existing approaches are proposed mainly from the perspective of image pixel manipulation. They are not effective to recover the details in the under/over exposed regions. In this paper, we revisit the image formation procedure and notice that the missing details in these regions exist in the corresponding high dynamic range (HDR) data. These details are well perceived by the human eyes but diminished in the low dynamic range (LDR) domain because of the tone mapping process. Therefore, we formulate the image correction task as an HDR transformation process and propose a novel approach called Deep Reciprocating HDR Transformation (DRHT). Given an input LDR image, we first reconstruct the missing details in the HDR domain. We then perform tone mapping on the predicted HDR data to generate the output LDR image with the recovered details. To this end, we propose a united framework consisting of two CNNs for HDR reconstruction and tone mapping. They are integrated end-to-end for joint training and prediction. Experiments on the standard benchmarks demonstrate that the proposed method performs favorably against state-of-the-art image correction methods. Xin Yang 0011, Ke Xu 0010, Yibing Song, Qiang Zhang 0008, Xiaopeng Wei, Rynson W. H. Lau |
CVPR | 4 |
| 2018 | Deep High-order Supervised Hashing for Image RetrievalabstractRecently, deep hashing has achieved excellent performances in large-scale image retrieval by simultaneously learning deep features and hashing function. However, state-of-the-art works have so far failed to explore the feature statistics higher than first-order. In this paper, to take a step towards addressing this problem, we propose two novel Deep High-order Supervised Hashing architectures (DHoSH), i.e., point-wise labels based DHoSH (DHoSH-PO) and pair-wise labels based DHoSH (DHoSH-PA). The core of DHoSH is that a trainable layer of bilinear pooling incorporates into deep convolutional neural networks (CNNs) for end-to-end learning. This layer captures the local feature interactions of the image by outer product, employing the autocorrelation information and cross-correlation information of deep features. Furthermore, our DHoSH method systematically exploits the high-order statistics of features of multiple layers. Extensive experiments on commonly used benchmarks illuminate that both DHoSH-PO and DHoSH-PA can obtain competitive improvements over its first-order counterparts, and achieve state-of-the-art performance for image retrieval task. Jingdong Cheng, Qiule Sun, Jianxin Zhang 0001, Xiaopeng Wei, Qiang Zhang 0008 |
ICPR | 5 |
| 2018 | Active Object Reconstruction Using a Guided View PlannerabstractInspired by the recent advance of image-based object reconstruction using deep learning, we present an active reconstruction model using a guided view planner. We aim to reconstruct a 3D model using images observed from a planned sequence of informative and discriminative views. But where are such informative and discriminative views around an object? To address this we propose a unified model for view planning and object reconstruction, which is utilized to learn a guided information acquisition model and to aggregate information from a sequence of images for reconstruction. Experiments show that our model (1) increases our reconstruction accuracy with an increasing number of views (2) and generally predicts a more informative sequence of views for object reconstruction compared to other alternative methods. Xin Yang 0011, Qiang Zhang 0008, Xiaopeng Wei, Hongbo Fu 0001 |
IJCAI | 5 |
| 2018 | The memory degradation based online sequential extreme learning machine
Quanyi Zou, Xiaojun Wang 0004, Changjun Zhou, Qiang Zhang 0008 |
Neurocomputing | 4 |
| 2018 | Passivity of Reaction-Diffusion Genetic Regulatory Networks with Time-Varying Delays
Xiaopeng Wei, Qiang Zhang 0008, Changjun Zhou |
Neural Process. Lett. | 3 |
| 2018 | Constructing DNA Barcode Sets Based on Particle Swarm OptimizationabstractFollowing the completion of the human genome project, a large amount of high-throughput bio-data was generated. To analyze these data, massively parallel sequencing, namely next-generation sequencing, was rapidly developed. DNA barcodes are used to identify the ownership between sequences and samples when they are attached at the beginning or end of sequencing reads. Constructing DNA barcode sets provides the candidate DNA barcodes for this application. To increase the accuracy of DNA barcode sets, a particle swarm optimization (PSO) algorithm has been modified and used to construct the DNA barcode sets in this paper. Compared with the extant results, some lower bounds of DNA barcode sets are improved. The results show that the proposed algorithm is effective in constructing DNA barcode sets. Bin Wang 0005, Xuedong Zheng, Shihua Zhou, Changjun Zhou, Xiaopeng Wei, Qiang Zhang 0008, Ziqi Wei 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 6 |
| 2018 | Modeling of Agent Cognition in Extensive Games via Artificial Neural NetworksabstractThe decision-making process, which is regarded as cognitive and ubiquitous, has been exploited in diverse fields, such as psychology, economics, and artificial intelligence. This paper considers the problem of modeling agent cognition in a class of game-theoretic decision-making scenarios called extensive games. We present a novel framework in which artificial neural networks are incorporated to simulate agent cognition regarding the structure of the underlying game and the goodness of the game situations therein. An algorithmic procedure is investigated to describe the process for solving games with cognition, and then, a new equilibrium concept is proposed as a refinement of the classical one-subgame perfect equilibrium-by involving players' cognitive reasoning. Moreover, a series of results concerning the computational complexity, soundness, and completeness of the algorithm, as well as the existence of an equilibrium solution, is obtained. This framework, which is shown to be general enough to model the way in which AlphaGo plays Go, may offer a means for bridging the gap between theoretical models and practical problem-solving. Chanjuan Liu 0001, Enqiang Zhu, Qiang Zhang 0008, Xiaopeng Wei |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2018 | Efficient image super-resolution integration
Ke Xu 0010, Xin Wang 0118, Xin Yang 0011, Shengfeng He, Qiang Zhang 0008, Xiaopeng Wei, Rynson W. H. Lau |
Vis. Comput. | 5 |
| 2017 | LSTM based classification model and its application for doctor-patient relationship evaluationabstractThe emergence of medical social media has made it possible for more and more patients to share their views and experiences on the medical care platform. These subjective texts contains patients' evaluation information for doctors and can be analyzed to provide rich decision-making information for patients and hospitals. Therefore, we propose a LSTM (Long Short Term Memory) based text sentiment classification method to evaluate the relationship between doctors and patients through the comment data from medical social media. The classification model of doctor-patient relationship can also be performed to evaluate the hospital by calculating the praise rate to help people choose hospitals. We perform two experiments on the patients' evaluation data from the website of Haodaifu. The experiment of doctor-patient relationship classification confirms the effectiveness of the classification model. In the experiment of hospital evaluation, we calculate the praise rate of 11 hospitals in Jinan of Shandong province based on the doctor-patient relationship classification results. The consistency between the results obtained by our method and the data of Mingyihui shows that our evaluation method for hospitals is reasonable and effective. Hongrui Kuang, Chao Che, Qiang Zhang 0008, Xiaopeng Wei |
Healthcom | 3 |
| 2017 | Classification of ECG signals based on 1D convolution neural networkabstractRecently, with the obvious increasing number of cardiovascular disease, the automatic classification research of Electrocardiogram signals (ECG) has been playing a significantly important part in the clinical diagnosis of cardiovascular disease. In this paper, a 1D convolution neural network (CNN) based method is proposed to classify ECG signals. The proposed CNN model consists of five layers in addition to the input layer and the output layer, i.e., two convolution layers, two down sampling layers and one full connection layer, extracting the effective features from the original data and classifying the features automatically. This model realizes the classification of 5 typical kinds of arrhythmia signals, i.e., normal, left bundle branch block, right bundle branch block, atrial premature contraction and ventricular premature contraction. The experimental results on the public MIT-BIH arrhythmia database show that the proposed method achieves a promising classification accuracy of 97.5%, significantly outperforming several typical ECG classification methods. Jianxin Zhang 0001, Qiang Zhang 0008, Xiaopeng Wei |
Healthcom | 3 |
| 2017 | Exploring risk factors and predicting UPDRS score based on Parkinson's speech signalsabstractThe unified Parkinson's disease rating scale (UPDRS) is the most widely employed scale for tracking Parkinson's disease (PD) symptom progression. However, conventional way to achieve UPDRS, mainly based on the physical examinations of clinic patients performed by the trained medical staffs, involves the disadvantages of inconvenience and high medical expense. Hence, in this study, we try to explore some risk factors and accurately predict the UPDRS for PD, using the speech signals of PD patients published on UCI machine-learning archive. More specifically, inspired by the idea of ensemble learning, we firstly construct a framework of ensemble feature selection (EFS) to select a suitable subset of features among numerous speech signals. Subsequently, a personalized predictive model, trained by adopting information from similar patients, is developed to be customized for an individual PD patient. Finally, we employ the personalized predictive model to predict UPDRS score combined with various classical regression algorithms. Compared to conventional models, our study has a potential to capture more relevant risk factors and produces more accurate UPDRS score for individual patient. Experimental results on real-world dataset from UCI machine-learning archive show that our personalized predictive model gets a promising performance. Jianxin Zhang 0001, Qiang Zhang 0008, Bo Jin 0001, Xiaopeng Wei |
Healthcom | 3 |
| 2017 | Logic Calculation Based on Two-Domain DNA Strand Displacement
Xiaobiao Wang, Changjun Zhou, Xuedong Zheng, Qiang Zhang 0008 |
ISNN (1) | 4 |
| 2017 | Interactive traffic simulation model with learned local parameters
Xin Yang 0011, Shuai Li 0014, Wanchao Su, Guozhen Tan, Qiang Zhang 0008, Xiaopeng Wei |
Multim. Tools Appl. | 7 |
| 2016 | A Segmented Artificial Bee Colony Algorithm Based on Synchronous Learning Factors
Jianxia Zhang, Qiang Zhang 0008 |
ACIIDS (1) | 4 |
| 2016 | A Method of Data Registration for 3D Point Clouds Combining with Motion Capture Technologies
Qiang Zhang 0008 |
ACIIDS (1) | 3 |
| 2016 | 3D Protein Structure Prediction with BSA-TS Algorithm
Changjun Zhou, Qiang Zhang 0008, Bin Wang 0005 |
IEA/AIE | 3 |
| 2016 | Minimizing Legal Exposure of High-Tech Companies through Collaborative Filtering MethodsabstractPatent litigation not only covers legal and technical issues, it is also a key consideration for managers of high-technology (high-tech) companies when making strategic decisions. Patent litigation influences the market value of high-tech companies. However, this raises unique challenges. To this end, in this paper, we develop a novel recommendation framework to solve the problem of litigation risk prediction. We will introduce a specific type of patent-related litigation, that is, Section 337 investigations, which prohibit all acts of unfair competition, or any unfair trade practices, when exporting products to the United States. To build this recommendation framework, we collect and exploit a large amount of published information related to almost all Section 337 investigation cases. This study has two aims: (1) to predict the litigation risk in a specific industry category for high-tech companies and (2) to predict the litigation risk from competitors for high-tech companies. These aims can be achieved by mining historical investigation cases and related patents. Specifically, we propose two methods to meet the needs of both aims: a proximal slope one predictor and a time-aware predictor. Several factors are considered in the proposed methods, including the litigation risk if a company wants to enter a new market and the risk that a potential competitor would file a lawsuit against the new entrant. Comparative experiments using real-world data demonstrate that the proposed methods outperform several baselines with a significant margin. Bo Jin 0001, Chao Che, Kuifei Yu, Li Guo 0008, Cuili Yao, Ruiyun Yu, Qiang Zhang 0008 |
KDD | 8 |
| 2016 | DKD: a fast k-d tree update design for dynamic scenesabstractAbstract We design dynamic k‐d (DKD) tree based on classical k‐d tree for animated scene rendering. Our method can inherit the benefit of efficient traversal of k‐d tree and minimize time cost to update DKD tree, making it well suited for animated geometry. DKD employs primitive reset, redistribution to reflect the updated positions of geometry, and leaf node incremental growing to avoid the deterioration of hierarchy quality due to refitting. Our experiments show that DKD has a significant rendering performance improvement than selected existing methods. Copyright © 2016 John Wiley & Sons, Ltd. Xin Yang 0011, Pengfei Zhang 0016, Lutong Xin, Yuxin Wang 0001, Qiang Zhang 0008, Xiaopeng Wei |
Comput. Animat. Virtual Worlds | 7 |
| 2015 | A Method of Facial Animation Retargeting Based on Motion Capture
Xiaoying Liang, Qiang Zhang 0008, Jing Dong 0009 |
ICIG (1) | 3 |
| 2014 | Smart Partitioning for Product DSM Model Based on Improved Genetic Algorithm
Yangjie Zhou 0002, Chao Che, Jianxin Zhang 0001, Qiang Zhang 0008, Xiaopeng Wei |
ADMA | 4 |
| 2014 | Forward non-rigid motion tracking for facial MoCap
Xiaoyong Fang, Xiaopeng Wei, Qiang Zhang 0008 |
Vis. Comput. | 3 |
| 2013 | On the simulation of expressional animation based on facial MoCap
Xiaoyong Fang, Xiaopeng Wei, Qiang Zhang 0008, Changjun Zhou |
Sci. China Inf. Sci. | 3 |
| 2012 | A novel color image encryption algorithm based on DNA sequence operation and hyper-chaotic system
Xiaopeng Wei, Qiang Zhang 0008, Jianxin Zhang 0001, Shiguo Lian |
J. Syst. Softw. | 3 |
| 2010 | Some novel classification and learning methods and applications for neural networks - Selected papers from the Second International Conference on Bio-Inspired Computing: Theories and Applications
Xiaopeng Wei, Qiang Zhang 0008, Guangzhao Cui |
Neurocomputing | 2 |
| 2009 | Robust pitch estimation using a wavelet variance analysis model
Xiaopeng Wei, Lasheng Zhao, Qiang Zhang 0008, Jing Dong 0009 |
Signal Process. | 3 |
| 2008 | Design of DNA Sequence Based on Improved Genetic Algorithm
Bin Wang 0005, Qiang Zhang 0008 |
ICIC (1) | 2 |
| 2006 | A Novel Global Exponential Stability Result for Discrete-time Cellular Neural Networks with Variable DelaysabstractGlobal exponential stability is considered for a class of discrete-time cellular neural networks with variable delays. By employing a discrete Halanay inequality, a new result is presented ensuring global exponential stability of the unique equilibrium point of the networks. The result extends and improves the earlier publications due to the fact that it removes some restrictions on the delay. An example is given to illustrate the effectiveness of the global exponential stability condition provided here. Qiang Zhang 0008, Xiaopeng Wei, Jin Xu 0015 |
Int. J. Neural Syst. | 1 |
| 2005 | A DNA Based Evolutionary Algorithm for the Minimal Set Cover Problem
Xiangou Zhu, Guandong Xu, Qiang Zhang 0008, Lin Gao 0006 |
ICIC (2) | 4 |
| 2005 | Global Exponential Stability of Discrete Time Hopfield Neural Networks with Delays
Qiang Zhang 0008, Wenbing Liu, Xiaopeng Wei |
ISNN (1) | 1 |
| 2005 | Global Asymptotic Stability Analysis of Neural Networks with Time-Varying Delays
Qiang Zhang 0008, Xiaopeng Wei, Jin Xu 0015 |
Neural Process. Lett. | 1 |
| 2004 | Asymptotic Stability of Nonautonomous Delayed Neural Networks
Qiang Zhang 0008, Xiaopeng Wei, Jin Xu 0015 |
ICONIP | 1 |
| 2004 | On the Asymptotic Stability of Non-autonomous Delayed Neural Networks
Qiang Zhang 0008, Xiaopeng Wei |
ISNN (1) | 1 |
| 2003 | Global stability of bidirectional associative memory neural networks with continuously distributed delays
Qiang Zhang 0008, Runnian Ma, Jin Xu 0015 |
Sci. China Ser. F Inf. Sci. | 1 |