VLDB 2026 Research / reviewers in the wild / expert
Xiaopeng Wei
dblp:12/4440
· DBLP profile ↗
103ranked-venue papers
3as first author
63since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 50 · 1 first-author · 32 since 2021Graphics, computer vision, multimedia, augmented reality and games · 38 · 1 first-author · 21 since 2021Applied, interdisciplinary, general and emerging computing · 20 · 10 since 2021Databases, data management, data science and information retrieval · 8 · 6 since 2021Computer networks · 5 · 5 since 2021Human-computer interaction and ubiquitous computing · 5 · 4 since 2021Systems, architecture and hardware · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MoCoDiff: Modality-Aware Conditional Diffusion Model for 3D Brain Tumor Segmentation
Sijie Guo, Jing Dong 0009, Rui Liu 0015, Xiaopeng Wei |
ICPR (5) | 6 |
| 2026 | DCPNet: a comprehensive framework for multimodal sarcasm detection via graph topology extraction and multi-scale feature fusion
Youjiang Fang, Xiaopeng Wei |
Frontiers Comput. Sci. | 7 |
| 2026 | Hierarchical prototype-guided representation learning for robust graph classification
Liang Zhang 0031, Kongyu Chen, Bo Jin 0001, Xiaopeng Wei |
Inf. Sci. | 4 |
| 2026 | Multi-Condition Latent Diffusion Network for Semantic-aware Knowledge Graph Completion
Liang Zhang 0031, Bo Jin 0001, Xiaopeng Wei |
Knowl. Based Syst. | 4 |
| 2026 | ICAD: Rethinking the Role of Inference and Cues for the Anomaly Detection of Time Series in IIoTabstractTime series anomaly detection in real-world Industrial Internet of Things (IIoT) systems is pivotal for identifying unsafe conditions and implementing timely preventive measures. While diffusion models are popular for capturing complex patterns, they often struggle to balance diversity and fidelity across scenarios due to limited exploration of context-window logical inference relationships and trend-pattern cues. To address these challenges, we propose ICAD, a novel method that rethinks the role of inference and cues in IIoT time series anomaly detection. ICAD defines trend patterns and uncertainties in textual form and utilizes a fine-tuned large language model to encode these descriptions as conditions for the diffusion model, thereby enhancing its generalization across diverse data distributions. Additionally, a “reasoning network for contextual window” mechanism is designed to capture temporal dependencies between adjacent windows, complemented by multi-scale and spatial feature adaptive fusion modules to further enhance the predictive performance. Empirical evaluations across four benchmark datasets and a large-scale ethylene oxide production process demonstrate that ICAD consistently outperforms state-of-the-art baselines, confirming its effectiveness and practicality in overcoming current anomaly detection model limitations. Zhichao Wu 0001, Zitao Yin, Shunqi Zhang, Xirong Xu, Xiaopeng Wei, Xin Yang 0011 |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2026 | Enhanced medical image segmentation via wavelet-deformable attention networks
Rui Liu 0015, Jing Dong 0009, Xiaopeng Wei |
Vis. Comput. | 5 |
| 2025 | Separating the Wheat from the Chaff: Spatio-Temporal Transformer with View-interweaved Attention for Photon-Efficient Depth SensingabstractTime-resolved imaging is an emerging sensing modality that has been shown to enable advanced applications, including remote sensing, fluorescence lifetime imaging, and even non-line-of-sight sensing. Single-photon avalanche diodes (SPADs) outperform relevant time-resolved imaging technologies thanks to their excellent photon sensitivity and superior temporal resolution on the order of tens of picoseconds. The capability of exceeding the sensing limits of conventional cameras for SPADs also draws attention to the photon-efficient imaging area. However, photon-efficient imaging under degraded conditions with low photon counts and low signal-to-background ratio (SBR) still remains an inevitable challenge. In this paper, we propose a spatio-temporal transformer network for photon-efficient imaging under low-flux scenarios. In particular, we introduce a view-interweaved attention mechanism (VIAM) to extract both spatial-view and temporal-view self-attention in each transformer block. We also design an adaptive-weighting scheme to dynamically adjust the weights between different views of self-attention in VIAM for different signal-to-background levels. We extensively validate and demonstrate the effectiveness of our approach on the simulated Middlebury dataset and a specially self-collected dataset with real-world-captured SPAD measurements and well-annotated ground truth depth maps. Letian Yu, Qirui Bao, Felix Heide, Xiaopeng Wei |
AAAI | 7 |
| 2025 | Touch2Shape: Touch-Conditioned 3D Diffusion for Shape Exploration and ReconstructionabstractDiffusion models have made breakthroughs in 3D generation tasks. Current 3D diffusion models focus on reconstructing target shape from images or a set of partial observations. While excelling in global context understanding, they struggle to capture the local details of complex shapes and limited to the occlusion and lighting conditions. To overcome these limitations, we utilize tactile images to capture the local 3D information and propose a Touch2Shape model, which leverages a touch-conditioned diffusion model to explore and reconstruct the target shape from touch. For shape reconstruction, we have developed a touch embedding module to condition the diffusion model in creating a compact representation and a touch shape fusion module to refine the reconstructed shape. For shape exploration, we combine the diffusion model with reinforcement learning to train a policy. This involves using the generated latent vector from the diffusion model to guide the touch exploration policy training through a novel reward design. Experiments validate the reconstruction quality thorough both qualitatively and quantitative analysis, and our touch exploration policy further boosts reconstruction performance. Zhaoxuan Zhang, Jiajin Qiu, Dilong Sun, Zhengyu Meng, Xiaopeng Wei |
CVPR | 6 |
| 2025 | Can I Trust You? Advancing GUI Task Automation with Action Trust Score
Haiyang Mei, Difei Gao, Xiaopeng Wei, Xin Yang 0011, Zheng Shou 0001 |
ACM Multimedia | 3 |
| 2025 | Self-Supervised Disentangled Representation Learning for Time Series Anomaly DetectionabstractAnomaly detection is a fundamental component of intelligent monitoring in the Internet of Things (IoT), where accuracy, efficiency, and interpretability are critical requirements. However, existing methods often overlook the unique characteristics of IoT signals such as seasonality, trends, and irregular residual components, as well as the complex interactions among them. This oversight can lead to anomaly masking, increased false positives, and reduced interpretability in anomaly identification. Motivated by the effectiveness of disentangled representation learning, we propose TRAdetector, a novel disentangled reconstruction-based framework for IoT signals anomaly detection. TRAdetector explicitly models recurrent and consistent patterns, as well as irregular variations in the latent space by leveraging variational inference strategies, thereby enhancing probabilistic guidance in learning both regular and irregular temporal representations. A sparse coding strategy is incorporated within the latent space of the residual component to directly model inconsistent temporal fluctuations. Finally, a multihead cross-attention mechanism and a gated, decomposition-aware reconstruction strategy are designed to effectively model the complex interactions among different components. Extensive experiments show that our model achieves state-of-the-art performance on multiple benchmark datasets in terms of accuracy, efficiency, and interpretability. Liang Zhang 0031, Jianping Zhu 0002, Guangjie Han, Bo Jin 0001, Pengfei Wang 0013, Xiaopeng Wei |
IEEE Internet Things J. | 6 |
| 2025 | DNA sequence design model for multi-scene fusion
Yanfen Zheng, Yaqing Hou, Qiang Zhang 0008, Xiaopeng Wei |
Neural Comput. Appl. | 6 |
| 2025 | GTIGNet: Global Topology Interaction Graphormer Network for 3D hand pose estimation
Wanshu Fan, Cong Wang 0018, Shixi Wen, Xin Yang 0011, Qiang Zhang 0008, Xiaopeng Wei |
Neural Networks | 7 |
| 2025 | FDC: Feature Dropout Consistency for unsupervised domain adaptation semantic segmentation
Chaoyu Rao, Wanshu Fan, Cong Wang 0018, Xin Yang 0011, Xiaopeng Wei |
Neural Networks | 5 |
| 2025 | scPEGEnhanced Graph Convolutional Sparse Subspace Clustering Method for scRNA-Seq DataabstractThe identification of cell types by clustering single-cell RNA sequencing (scRNA-seq) data is a fundamental step in the downstream analysis of single-cell data. However, great challenges remain owing to the inherent characteristics of scRNA-seq data, including high dimensionality, high noise, and high sparsity. In this study, we propose a proximity enhanced graph convolutional sparse subspace clustering method scPEGSSC for scRNA-seq data. Method scPEGSSC generates the similarity matrix with the self-expression matrix (SEM) learned from a graph autoencoder, and enhances it further through its square. Experiments were performed on thirteen real biological datasets. The experimental results indicate compared with eleven state-of-the-art single-cell clustering methods, method scPEGSSC have attained superior performance across most datasets. Jingli Wu, Xiaopeng Wei, Gaoshi Li, Jiafei Liu 0001 |
IEEE Trans. Comput. Biol. Bioinform. | 2 |
| 2025 | MLFuse: Multi-Scenario Feature Joint Learning for Multi-Modality Image FusionabstractMulti-modality image fusion (MMIF) entails synthesizing images with detailed textures and prominent objects. Existing methods tend to use general feature extraction to handle different fusion tasks. However, these methods have difficulty breaking fusion barriers across various modalities owing to the lack of targeted learning routes. In this work, we propose a multi-scenario feature joint learning architecture, MLFuse, that employs the commonalities of multi-modality images to deconstruct the fusion progress. Specifically, we construct a cross-modal knowledge reinforcing network that adopts a multipath calibration strategy to promote information communication between different images. In addition, two professional networks are developed to maintain the salient and textural information of fusion results. The spatial-spectral domain optimizing network can learn the vital relationship of the source image context with the help of spatial attention and spectral attention. The edge-guided learning network utilizes the convolution operations of various receptive fields to capture image texture information. The desired fusion results are obtained by aggregating the outputs from the three networks. Extensive experiments demonstrate the superiority of MLFuse for infrared-visible image fusion and medical image fusion. The excellent results of downstream tasks (i.e., object detection and semantic segmentation) further verify the high-quality fusion performance of our method. Jia Lei 0001, Jiawei Li 0016, Jinyuan Liu 0001, Bin Wang 0005, Shihua Zhou, Qiang Zhang 0008, Xiaopeng Wei, Nikola K. Kasabov |
IEEE Trans. Multim. | 7 |
| 2025 | Speed up Federated Unlearning With Temporary Local ModelsabstractFederated unlearning (FUL) is a solution aimed at addressing the problem of removing data contributions from trained federated learning (FL) models. Existing FUL methods only focus on iterative unlearning of clients’ contributions and fail to perform unlearning in scenarios where multiple clients request to remove their data at a time. Additionally, FUL still needs to address issues, including convergence speed, maintaining the global model’s performance, and parallel unlearning to expedite the unlearning process. To fill this gap, we introduce Federated Clients Forgetting (FedCF), a fast and accurate FUL method that can eliminate single client contributions similar to existing methods, eliminate multiple clients’ contributions on the global model parallelly, ensure the performance of the unlearned global model, and reduce the unlearning time. The key idea is to construct a temporary model by extracting knowledge from the remaining clients’ updates and adding it to the corresponding parameters of the initial global model and then leverage a temporary model to reconstruct the unlearned global model. Extensive experiments on three benchmark datasets, FedCF demonstrates its efficiency and effectiveness for single client contribution unlearning, achieving an average time efficiency of 8.3x, 6.5x, and 4.1x over existing methods FedRetrain, FedEraser, and FUL with knowledge distillation, respectively. Additionally, FedCF showcases the time efficiency and performance guarantee after unlearning the contributions of multiple clients in parallel. Muhammed Ameen, Pengfei Wang 0013, Weijian Su, Xiaopeng Wei, Qiang Zhang 0008 |
IEEE Trans. Sustain. Comput. | 4 |
| 2024 | FreAML: A Frequency-Domain Adaptive Meta-Learning Framework for EEG-Based Emotion RecognitionabstractEmotion recognition technology, especially methods based on electroencephalogram (EEG) signals, plays a critical role in revealing deep human emotions and enhancing human-computer interaction experiences. However, the high-frequency characteristics of EEG signals lead to significant individual differences, causing domain distribution shift issues. To address this, we propose a frequency-domain adaptive meta-learning probabilistic inference framework, named FreAML, leveraging the advantages of frequency-domain signals in preserving comprehensive views and learning global dependencies. Specifically, we first introduce an adaptive meta-task construction strategy based on a clustering-matching pattern. Then, through a dual-stream amortization network, we learn task-specific frequency-domain real and imaginary parameters from the support set. Finally, using these task-specific parameters, we achieve rapid generalization learning of the query set through a frequency-domain linear mapping. Extensive experiments on two public datasets validate the effectiveness and superiority of the proposed method. Lei Wang 0196, Jianping Zhu 0002, Le Du, Bo Jin 0001, Xiaopeng Wei |
BIBM | 5 |
| 2024 | Multi-view Time-frequency Contrastive Learning for Emotion Recognition
Lei Wang 0196, Jianping Zhu 0002, Bo Jin 0001, Xiaopeng Wei |
CogSci | 4 |
| 2024 | BRG: Bidirectional Regularization Guidance for Unsupervised Domain Adaptation Semantic SegmentationabstractUnsupervised Domain Adaptation Semantic Segmentation (UDASS) aims to harness labeled data from the source domain alongside unlabeled data from the target domain to effectively segment target domain images. While self-training has achieved tremendous success in UDASS, most existing methods heavily rely on source data and lack of sufficient learning from target data, resulting in a performance decline. To address this issue, this paper introduces a Bidirectional Regularization Guidance (BRG) method, which combines a Target Perturbation Consistency (TPC) module and a thing-class ImageNet Feature Distance Reweighting (FDR) module to provide effective regularization guidance. Specifically, the TPC module is introduced to perturb the target stream by masking at the input level and adding feature noise at the feature level. This module employs a target perturbation consistency loss to penalize inconsistencies between perturbed student predictions and teacher model predictions, thereby facilitating improved learning of contextual information by the student model from the target images. To further enhance the model’s generalization ability to the target domain, We propose an FDR module that employs a transferability map to calculate a reweighted graph, helping to effectively align source features with ImageNet features by reweighting the feature distances, especially those with lower transferability Through rigorous experimentation in standard UDASS settings—training with synthetic labeled and real unlabeled data—BRG outperforms baseline model on GTAV→Cityscapes and SYNTHIA→Cityscapes. Chaoyu Rao, Wanshu Fan, Xiaopeng Wei |
IJCNN | 4 |
| 2024 | Event-intensity Stereo with Cross-modal Fusion and ContrastabstractFor binocular stereo, traditional cameras excel in capturing fine details and texture information but are limited in terms of dynamic range and their ability to handle rapid motion. On the contrary, event cameras provide pixel-level intensity changes with low latency and a wide dynamic range, albeit at the cost of less detail in their output. It is natural to leverage the strengths of both modalities. We solve this problem by introducing a cross-modal fusion module that learns a visual representation from both sensor inputs. Additionally, we extract and compare dense event-intensity stereo pair features by contrasting “pairs of event-intensity pairs from different views and different modalities and different timestamps”. This provides the flexibility in masking hard negatives and enables networks to effectively combine event-intensity signals within a contrastive learning framework, leading to an improved matching accuracy and facilitating more accurate estimation of disparity. Experimental results validate the effectiveness of our model and the improvement of disparity estimation accuracy. Shanglai Qu, Tianyu Meng, Haiyin Piao, Xiaopeng Wei, Xin Yang 0011 |
IROS | 6 |
| 2024 | Adaptive Vision Transformer for Event-Based Human Pose EstimationabstractEvent-based human pose estimation has gained popularity due to the benefits of high temporal resolution and high dynamic range offered by event cameras. The inherent spatial sparsity of event data makes discarding less significant regions a straightforward and effective way to decrease the computation. However, implementing this operation in CNNs poses a challenge, as it disrupts the regularity of dense convolutional workload. In this paper, we propose an adaptive vision transformer, a novel efficient backbone for human pose estimation with event cameras. Specifically, we present two adaptive patch and token sampling approaches based on the characteristics of events, thereby reducing the computational load while still achieving comparable performance. Firstly, we design an adaptive patch sampling scheme to eliminate inactivity patches by assessing the entropy of the events before they are inputted into the transformer. Secondly, we further propose an adaptive token reduction strategy to selectively remove less informative tokens in transformer layers through a dynamic token pruning algorithm. To exploit event-based visual cues in human pose estimation tasks, we construct a large-scale frame-event-based dataset, dubbed Event Multi Movement HPE (EventMM HPE). The dataset provides annotation frequencies up to 240 Hz. Extensive experiments demonstrate that our proposed approach outperforms existing state-of-the-art methods in estimation accuracy. The source code and dataset are available at https://github.com/doublemanyu/Adaptive-Vision-Transformer-for-Event-Based-HPE. Nannan Yu, Jiqing Zhang, Yuji Zhang 0004, Qirui Bao, Xiaopeng Wei, Xin Yang 0011 |
ACM Multimedia | 6 |
| 2024 | A self-supervised anomaly detection algorithm with interpretability
Zhichao Wu 0001, Xin Yang 0011, Xiaopeng Wei, Peijun Yuan, Yuanping Zhang, Jianming Bai |
Expert Syst. Appl. | 3 |
| 2024 | A Universal Event-Based Plug-In Module for Visual Object Tracking in Degraded Conditions
Jiqing Zhang, Bo Dong 0004, Yingkai Fu, Yuanchen Wang, Xiaopeng Wei, Xin Yang 0011 |
Int. J. Comput. Vis. | 5 |
| 2024 | Reinforcement learning from constraints and focal entity shifting in conversational KGQA
Xirong Xu, Xiaopeng Wei |
Neural Comput. Appl. | 6 |
| 2024 | Eye Gaze Guided Cross-Modal Alignment Network for Radiology Report GenerationabstractThe potential benefits of automatic radiology report generation, such as reducing misdiagnosis rates and enhancing clinical diagnosis efficiency, are significant. However, existing data-driven methods lack essential medical prior knowledge, which hampers their performance. Moreover, establishing global correspondences between radiology images and related reports, while achieving local alignments between images correlated with prior knowledge and text, remains a challenging task. To address these shortcomings, we introduce a novel Eye Gaze Guided Cross-modal Alignment Network (EGGCA-Net) for generating accurate medical reports. Our approach incorporates prior knowledge from radiologists' Eye Gaze Region (EGR) to refine the fidelity and comprehensibility of report generation. Specifically, we design a Dual Fine-Grained Branch (DFGB) and a Multi-Task Branch (MTB) to collaboratively ensure the alignment of visual and textual semantics across multiple levels. To establish fine-grained alignment between EGR-related images and sentences, we introduce the Sentence Fine-grained Prototype Module (SFPM) within DFGB to capture cross-modal information at different levels. Additionally, to learn the alignment of EGR-related image topics, we introduce the Multi-task Feature Fusion Module (MFFM) within MTB to refine the encoder output information. Finally, a specifically designed label matching mechanism is designed to generate reports that are consistent with the anticipated disease states. The experimental outcomes indicate that the introduced methodology surpasses previous advanced techniques, yielding enhanced performance on two extensively used benchmark datasets: Open-i and MIMIC-CXR. Peixi Peng, Wanshu Fan, Wenfei Liu, Xin Yang 0011, Qiang Zhang 0008, Xiaopeng Wei |
IEEE J. Biomed. Health Informatics | 7 |
| 2024 | Decentralized Navigation With Heterogeneous Federated Reinforcement Learning for UAV-Enabled Mobile Edge ComputingabstractUnmanned Aerial Vehicle (UAV)-enabled mobile edge computing has been proposed as an efficient task-offloading solution for user equipments (UEs). Nevertheless, the presence of heterogeneous UAVs makes centralized navigation policies impractical. Decentralized navigation policies also face significant challenges in knowledge sharing among heterogeneous UAVs. To address this, we present the soft hierarchical deep reinforcement learning network (SHDRLN) and dual-end federated reinforcement learning (DFRL) as a decentralized navigation policy solution. It enhances overall task-offloading energy efficiency for UAVs while facilitating knowledge sharing. Specifically, SHDRLN, a hierarchical DRL network based on maximum entropy learning, reduces policy differences among UAVs by abstracting atomic actions into generic skills. Simultaneously, it maximizes the average efficiency of all UAVs, optimizing coverage for UEs and minimizing task-offloading waiting time. DFRL, a federated learning (FL) algorithm, aggregates policy knowledge at the cloud server and filters it at the UAV end, enabling adaptive learning of navigation policy knowledge suitable for the UAV's performance parameters. Extensive simulations demonstrate that the proposed solution not only outperforms other baseline algorithms in overall energy efficiency but also achieves more stable navigation policy learning under different levels of heterogeneity of different UAV performance parameters. Pengfei Wang 0013, Guangjie Han, Ruiyun Yu, Leyou Yang, Geng Sun 0001, Heng Qi, Xiaopeng Wei, Qiang Zhang 0008 |
IEEE Trans. Mob. Comput. | 8 |
| 2023 | Adaptive Bayesian Meta-Learning for EEG Signal ClassificationabstractAccurate classification of electroencephalogram (EEG) signals is crucial for brain activity understanding. However, EEG signals are characterized by data heterogeneity and label scarcity, which present a challenging low-data learning regime when building machine learning models. Existing methods tend to suffer from overfitting problem. To this end, we propose an adaptive Bayesian meta-learning framework for instance-specific learning and inference in EEG classification tasks. Specifically, first, a query set-driven dynamic parameter-based support set selection strategy is designed to adaptively fit the query set when constructing a meta-training task. Second, we employ an amortized variational inference network to generate task-specific adapted parameters given the support set, thereby achieving rapid model adaption for the inference of the data in the query set. Especially, a time- and frequency-aware representation learning encoder is leveraged to extract more task-relevant information guided by information bottleneck principle from time and frequency views, respectively, alleviating the low signal-to-noise ratio issue. Extensive experimental results on three public datasets demonstrate the superior effectiveness of our method. Jianping Zhu 0002, Liang Zhang 0031, Bo Jin 0001, Xiaopeng Wei |
BIBM | 5 |
| 2023 | Deep Polarization Reconstruction with PDAVIS EventsabstractThe polarization event camera PDAVIS is a novel bio-inspired neuromorphic vision sensor that reports both conventional polarization frames and asynchronous, continuously per-pixel polarization brightness changes (polarization events) with fast temporal resolution and large dynamic range. A deep neural network method (Polarization FireNet) was previously developed to reconstruct the polarization angle and degree from polarization events for bridging the gap between the polarization event camera and mainstream computer vision. However, Polarization FireNet applies a network pretrained for normal event-based frame reconstruction independently on each of four channels of polarization events from four linear polarization angles, which ignores the correlations between channels and inevitably introduces content inconsistency between the four reconstructed frames, resulting in unsatisfactory polarization reconstruction performance. In this work, we strive to train an effective, yet efficient, DNN model that directly outputs polarization from the input raw polarization events. To this end, we constructed the first large-scale event-to-polarization dataset, which we subsequently employed to train our events-to-polarization network E2P. E2P extracts rich polarization patterns from input polarization events and enhances features through cross-modality context integration. We demonstrate that E2P outperforms Polarization FireNet by a significant margin with no additional computing cost. Experimental results also show that E2P produces more accurate measurement of polarization than the PDAVIS frames in challenging fast and high dynamic range scenes. Code and data are publicly available at: https://github.com/SensorsINI/e2p. Haiyang Mei, Zuowen Wang, Xin Yang 0011, Xiaopeng Wei, Tobi Delbruck |
CVPR | 4 |
| 2023 | Multi-view Spectral Polarization Propagation for Video Glass SegmentationabstractIn this paper, we present the first polarization-guided video glass segmentation propagation solution (PGVS-Net) that can robustly and coherently propagate glass segmentation in RGB-P video sequences. By leveraging spatiotemporal polarization and color information, our method combines multi-view polarization cues and thus can alleviate the view dependence of single-input intensity variations on glass objects. We demonstrate that our model can outperform glass segmentation on RGB-only video sequences as well as produce more robust segmentation than per-frame RGB-P single-image segmentation methods. To train and validate PGVS-Net, we introduce a novel RGB-P Glass Video dataset (PGV-117) containing 117 video sequences of scenes captured with different types of camera paths, lighting conditions, dynamics, and glass types. Yu Qiao 0001, Bo Dong 0004, Ao Jin, Seung-Hwan Baek, Felix Heide, Pieter Peers, Xiaopeng Wei, Xin Yang 0011 |
ICCV | 8 |
| 2023 | A Global View-Guided Autoregressive Residual Network for Irregular Time Series Classification
Jianping Zhu 0002, Haocheng Tang, Liang Zhang 0031, Bo Jin 0001, Xiaopeng Wei |
PAKDD (4) | 6 |
| 2023 | A Heuristic Framework for Personalized Route Recommendation Based on Convolutional Neural Networks
Ruining Zhang, Chanjuan Liu 0001, Qiang Zhang 0008, Xiaopeng Wei |
PRICAI (3) | 4 |
| 2023 | Camouflaged Object Segmentation with Omni Perception
Haiyang Mei, Ke Xu 0010, Yunduo Zhou, Yang Wang 0106, Haiyin Piao, Xiaopeng Wei, Xin Yang 0011 |
Int. J. Comput. Vis. | 6 |
| 2023 | Multigranularity Pruning Model for Subject Recognition Task under Knowledge Base Question Answering When General Models FailabstractIn general knowledge base question answering (KBQA) models, subject recognition (SR) is usually a precondition of finding an answer, and it is a common way to employ a general named entity recognition (NER) model such as BERT‐CRF to recognize the subject. However, in previous researches, the difference between a NER task and a SR task is usually ignored, and a wrong entity recognized by the NER model will certainly lead to a wrong answer in the KBQA task, which is one bottleneck for KBQA performance. In this paper, a multigranularity pruning model (MGPM) is proposed to answer a question when general models fail to recognize a subject. In MGPM, the set of all possible subjects in the Knowledge Base (KB) is pruned by 4 multigranularity pruning submodels successively based on the constraint of relation (domain and tuple), string similarity, and semantic similarity. Experimental results show that our model is compatible with various KBQA models for both single‐relation and complex questions answering. The integrated MGPM model (with the BERT‐CRF model) achieves a SR accuracy of 94.4% on the SimpleQuestions dataset, 68.6% on the WebQuestionsSP dataset, and 63.7% on the WebQuestions dataset, which outperforms the original model by a margin of 3.6%, 8.6%, and 5.3%, respectively. Xirong Xu, Xiaopeng Wei, Degen Huang |
Int. J. Intell. Syst. | 5 |
| 2023 | Large-Field Contextual Feature Learning for Glass DetectionabstractGlass is very common in our daily life. Existing computer vision systems neglect it and thus may have severe consequences, e.g., a robot may crash into a glass wall. However, sensing the presence of glass is not straightforward. The key challenge is that arbitrary objects/scenes can appear behind the glass. In this paper, we propose an important problem of detecting glass surfaces from a single RGB image. To address this problem, we construct the first large-scale glass detection dataset (GDD) and propose a novel glass detection network, called GDNet-B, which explores abundant contextual cues in a large field-of-view via a novel large-field contextual feature integration (LCFI) module and integrates both high-level and low-level boundary features with a boundary feature enhancement (BFE) module. Extensive experiments demonstrate that our GDNet-B achieves satisfying glass detection results on the images within and beyond the GDD testing set. We further validate the effectiveness and generalization capability of our proposed GDNet-B by applying it to other vision tasks, including mirror segmentation and salient object detection. Finally, we show the potential applications of glass detection and discuss possible future research directions. Haiyang Mei, Xin Yang 0011, Letian Yu, Qiang Zhang 0008, Xiaopeng Wei, Rynson W. H. Lau |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2023 | Distractor-Aware Event-Based TrackingabstractEvent cameras, or dynamic vision sensors, have recently achieved success from fundamental vision tasks to high-level vision researches. Due to its ability to asynchronously capture light intensity changes, event camera has an inherent advantage to capture moving objects in challenging scenarios including objects under low light, high dynamic range, or fast moving objects. Thus event camera are natural for visual object tracking. However, the current event-based trackers derived from RGB trackers simply modify the input images to event frames and still follow conventional tracking pipeline that mainly focus on object texture for target distinction. As a result, the trackers may not be robust dealing with challenging scenarios such as moving cameras and cluttered foreground. In this paper, we propose a distractor-aware event-based tracker that introduces transformer modules into Siamese network architecture (named DANet). Specifically, our model is mainly composed of a motion-aware network and a target-aware network, which simultaneously exploits both motion cues and object contours from event data, so as to discover motion objects and identify the target object by removing dynamic distractors. Our DANet can be trained in an end-to-end manner without any post-processing and can run at over 80 FPS on a single V100. We conduct comprehensive experiments on two large event tracking datasets to validate the proposed model. We demonstrate that our tracker has superior performance against the state-of-the-art trackers in terms of both accuracy and efficiency. Yingkai Fu, Meng Li 0072, Wenxi Liu, Yuanchen Wang, Jiqing Zhang, Xiaopeng Wei, Xin Yang 0011 |
IEEE Trans. Image Process. | 7 |
| 2023 | Mirror Segmentation via Semantic-aware Contextual Contrasted Feature LearningabstractMirrors are everywhere in our daily lives. Existing computer vision systems do not consider mirrors, and hence may get confused by the reflected content inside a mirror, resulting in a severe performance degradation. However, separating the real content outside a mirror from the reflected content inside it is non-trivial. The key challenge is that mirrors typically reflect contents similar to their surroundings, making it very difficult to differentiate the two. In this article, we present a novel method to segment mirrors from a single RGB image. To the best of our knowledge, this is the first work to address the mirror segmentation problem with a computational approach. We make the following contributions: First, we propose a novel network, called MirrorNet+, for mirror segmentation, by modeling both contextual contrasts and semantic associations. Second, we construct the first large-scale mirror segmentation dataset, which consists of 4,018 pairs of images containing mirrors and their corresponding manually annotated mirror masks, covering a variety of daily-life scenes. Third, we conduct extensive experiments to evaluate the proposed method and show that it outperforms the related state-of-the-art detection and segmentation methods. Fourth, we further validate the effectiveness and generalization capability of the proposed semantic awareness contextual contrasted feature learning by applying MirrorNet+ to other vision tasks, i.e., salient object detection and shadow detection. Finally, we provide some applications of mirror segmentation and analyze possible future research directions. Project homepage: https://mhaiyang.github.io/TOMM2022-MirrorNet+/index.html . Haiyang Mei, Letian Yu, Ke Xu 0010, Yang Wang 0106, Xin Yang 0011, Xiaopeng Wei, Rynson W. H. Lau |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2022 | Glass Segmentation using Intensity and Spectral Polarization CuesabstractTransparent and semi-transparent materials pose significant challenges for existing scene understanding and segmentation algorithms due to their lack of RGB texture which impedes the extraction of meaningful features. In this work, we exploit that the light-matter interactions on glass materials provide unique intensity-polarization cues for each observed wavelength of light. We present a novel learning-based glass segmentation network that leverages both trichromatic (RGB) intensities as well as trichromatic linear polarization cues from a single photograph captured without making any assumption on the polarization state of the illumination. Our novel network architecture dynamically fuses and weights both the trichromatic color and polarization cues using a novel global-guidance and multi-scale self-attention module, and leverages global cross-domain contextual information to achieve robust segmentation. We train and extensively validate our segmentation method on a new large-scale RGB-Polarization dataset (RGBP-Glass), and demonstrate that our method outperforms state-of-the-art segmentation approaches by a significant margin. Haiyang Mei, Bo Dong 0004, Wen Dong 0008, Seung-Hwan Baek, Felix Heide, Pieter Peers, Xiaopeng Wei, Xin Yang 0011 |
CVPR | 8 |
| 2022 | SSMFRP: Semantic Similarity Model for Relation Prediction in KBQA Based on Pre-trained Models
Xirong Xu, Xinzi Li, Xiaopeng Wei, Degen Huang |
ICANN (2) | 5 |
| 2022 | A Novel Movement-supported HRI Framework for Humanoid RobotsabstractCurrent research related to human-robot interaction (HRI) of bipedal humanoid robots often assumes that the robot is in a standing stationary state, i.e., the relative position of the robot does not change, and rarely considers the effect of lower limb movement on interaction. However, HRI in the real world does not assume a moving or stationary state of the robot, and the equilibrium perturbations caused by movement can prevent HRI from functioning properly. In this paper, we propose a movement supported humanoid robot interaction method that empowers the robot to move stably while achieving HRI. First, a reinforcement learning-based neural network is run offline to generate interaction actions that satisfy the equilibrium constraint and support movement, and then an intention recognition network is introduced to run the movement-supported HRI framework online. It is demonstrated that the training method proposed in this paper can enable a bipedal robot to achieve a variety of interactive actions while moving stably. Jing Dong 0009, Rui Liu 0015, Xiaopeng Wei, Qiang Zhang 0008 |
IJCNN | 6 |
| 2022 | High-order local connection network for 3D human pose estimation based on GCN
Qiang Zhang 0008, Jing Dong 0009, Xiaopeng Wei |
Appl. Intell. | 5 |
| 2022 | Bi-graph attention network for aspect category sentiment classification
Yongxue Shan, Chao Che, Xiaopeng Wei, Yongjun Zhu 0001, Bo Jin 0001 |
Knowl. Based Syst. | 3 |
| 2022 | Adaptive kernel selection network with attention constraint for surgical instrument classificationabstractAbstract Computer vision (CV) technologies are assisting the health care industry in many respects, i.e., disease diagnosis. However, as a pivotal procedure before and after surgery, the inventory work of surgical instruments has not been researched with the CV-powered technologies. To reduce the risk and hazard of surgical tools’ loss, we propose a study of systematic surgical instrument classification and introduce a novel attention-based deep neural network called SKA-ResNet which is mainly composed of: (a) A feature extractor with selective kernel attention module to automatically adjust the receptive fields of neurons and enhance the learnt expression and (b) A multi-scale regularizer with KL-divergence as the constraint to exploit the relationships between feature maps. Our method is easily trained end-to-end in only one stage with few additional calculation burdens. Moreover, to facilitate our study, we create a new surgical instrument dataset called SID19 (with 19 kinds of surgical tools consisting of 3800 images) for the first time. Experimental results show the superiority of SKA-ResNet for the classification of surgical tools on SID19 when compared with state-of-the-art models. The classification accuracy of our method reaches up to 97.703%, which is well supportive for the inventory and recognition study of surgical tools. Also, our method can achieve state-of-the-art performance on four challenging fine-grained visual classification datasets. Yaqing Hou, Qian Liu 0001, Hong-Wei Ge, Jun Meng, Qiang Zhang 0008, Xiaopeng Wei |
Neural Comput. Appl. | 7 |
| 2022 | Perception-oriented Single Image Super-Resolution Network with Receptive Field Block
Yaqing Hou, Wanshu Fan, Xin Yang 0011, Qiang Zhang 0008, Xiaopeng Wei |
Neural Comput. Appl. | 7 |
| 2022 | Designing Uncorrelated Address Constrain for DNA Storage by DMVO AlgorithmabstractAt present, huge amounts of data are being produced every second, a situation that will gradually overwhelm current storage technology. DNA is a storage medium that features high storage density and long-term stability and is now considered to be a feasible storage solution. Errors are easily made during the sequencing and synthesis of DNA, however. In order to reduce the error rate, novel uncorrelated address constrain are reported, and a Damping Multi-Verse Optimizer (DMVO)algorithm is proposed to construct a set of DNA coding, which is used as the non-payload. The DMVO algorithm exchanges objects through black/white holes in order to achieve a stable state and adds damping factors as disturbances. Compared with previous work, the coding set obtained by the DMVO algorithm is larger in size and of higher quality. The results of this study reveal that the size of the DNA storage coding set obtained by the DMVO algorithm increased by 4-16 percent, and the variance of the melting temperature decreased by 3-18 percent. Ben Cao, Xue Ii, Bin Wang 0005, Qiang Zhang 0008, Xiaopeng Wei |
IEEE ACM Trans. Comput. Biol. Bioinform. | 6 |
| 2022 | A Novel Adaptive Linear Neuron Based on DNA Strand Displacement Reaction NetworkabstractAnalog DNA strand displacement circuits can be used to build artificial neural network due to the continuity of dynamic behavior. In this study, DNA implementations of novel catalysis, novel degradation and adjustment reaction modules are designed and used to build an analog DNA strand displacement reaction network. A novel adaptive linear neuron (ADALINE) is constructed by the ordinary differential equations of an ideal formal chemical reaction network, which is built by reaction modules. When reaction network approaches equilibrium, the weights of the ADALINE are updated without learning algorithm. Simulation results indicate that, ADALINE based on the analog DNA strand displacement circuit has ability to implement the learning function of the ADALINE based on the ideal formal chemical reaction networks, and fit a class of linear function. Xiaopeng Wei, Qiang Zhang 0008, Changjun Zhou |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2022 | Synchronization of Hyper-Lorenz System Based on DNA Strand DisplacementabstractLorenz system is depicted by chemical reaction equations of an ideal formal chemical reaction network, and a series of reversible reactions are added into chemical reaction network in order to construct a cluster of hyper-Lorenz system. DNA as a universal substrate for chemical dynamics can approximate arbitrary dynamical characteristics of ideal formal chemical reaction network through auxiliary DNA strands and displacement reactions. Based on Lyapunov's stableness theory, a novel synchronization strategy is proposed. A 6-dimensional hyper-Lorenz system is taken as examples for simulation and shows that DNA strands displacement reactions can implement the synchronization of ideal formal chemical reaction networks. Numerical simulations indicate that synchronization based on DNA strand displacement is robust to the detection of DNA strand concentration, control of reaction rate, and noise. Qiang Zhang 0008, Xiaopeng Wei |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2022 | Exploring Dense Context for Salient Object DetectionabstractContexts play an important role in salient object detection (SOD). High-level contexts describe the relations between different parts/objects and thus are helpful for discovering the specific locations of salient objects while low-level contexts could provide the fine detail information for delineating the boundary of the salient objects. However, the way of perceiving/leveraging rich contexts has not been fully investigated by existing SOD works. The common context extraction strategies (e.g., leveraging convolutions with large kernels or atrous convolutions with large dilation rates) do not consider the effectiveness and efficiency simultaneously and may cause sub-optimal solutions. In this paper, we devote to exploring an effective and efficient way to learn rich contexts for accurate SOD. Specifically, we first build a dense context exploration (DCE) module to capture dense multi-scale contexts and further leverage the learned contexts to enhance the features discriminability. Then, we embed multiple DCE modules in an encoder-decoder architecture to harvest dense contexts of different levels. Furthermore, we propose an attentive skip-connection to transmit useful features from the encoder part to the decoder part for better dense context exploration. Finally, extensive experiments demonstrate that the proposed method achieves more superior detection results on the six benchmark datasets than 18 state-of-the-art SOD methods. Haiyang Mei, Ziqi Wei 0001, Xiaopeng Wei, Qiang Zhang 0008, Xin Yang 0011 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | From Pixels to Semantics: Self-Supervised Video Object Segmentation With Multiperspective Feature MiningabstractExisting self-supervised methods pose one-shot video object segmentation (O-VOS) as pixel-level matching to enable segmentation mask propagation across frames. However, the two tasks are not fully equivalent since O-VOS is more reliant on semantic correspondence rather than accurate pixel matching. To remedy this issue, we explore a new self-supervised framework that integrates pixel-level correspondence learning with semantic-level adaptation. The pixel-level correspondence learning is performed through photometric reconstruction of adjacent RGB frames during offline training, while semantic-level adaption operates at test-time by enforcing a bi-directional agreement of the predicted segmentation masks. In addition, we further propose a new network architecture with multi-perspective feature mining mechanism which can not only enhance reliable features but also suppress noisy ones to facilitate more robust image matching. By training the network using the proposed self-supervised framework, we achieve state-of-the-art performance on widely adopted datasets, further closing up the gap between self-supervised learning methods and their fully supervised counterparts. Ruoqi Li, Yifan Wang 0004, Lijun Wang 0001, Huchuan Lu, Xiaopeng Wei, Qiang Zhang 0008 |
IEEE Trans. Image Process. | 5 |
| 2022 | Prediction of Treatment Medicines With Dual Adaptive Sequential NetworksabstractPredicting treatment medicines is a key task in many intelligent healthcare systems. Prediction of treatment medicines can assist doctors in making informed prescription decisions for patients according to their Electronic Health Records (EHRs). However, predicting treatment medicines is a challenging task due to the following reasons: (1) heterogeneous nature of EHR data that typically includes laboratory results, treatment records, disease conditions, and demographic information; (2) complex correlations among EHR sequences, including inter-correlations between sequences and temporal intra-correlations within each sequence; (3) temporal dynamics of these correlations changing with disease progression. In this paper, we predict treatment medicines for patients with dual adaptive sequential networks (DASNet). Specifically, DASNet is designed with three components. First, a decomposed adaptive long short-term memory network (DA-LSTM) is designed to capture the intra- and inter-correlations in multiple heterogeneous temporal sequences. Then, we develop an attentive meta learning network (AT-MetaNet) to learn dynamic weight parameters for DA-LSTM, thus enabling it to model various correlation structures. Finally, we employ an attentive fusion network (AT-FuNet) to incorporate historical information and collectively fuse representation embeddings of heterogeneous data to predict treatment medicines. Our results on the public MIMIC-III dataset covering 11 medical conditions demonstrate that the proposed end-to-end model can achieve the state-of-the-art prediction performance while providing clinically useful insights. Liang Zhang 0031, Leilei Sun, Bo Jin 0001, Chuanren Liu, Ruiyun Yu, Xiaopeng Wei |
IEEE Trans. Knowl. Data Eng. | 8 |
| 2021 | Aspect-Level Sentiment Classification of Chinese Patient Comments Based on Pre-trained Sentiment EmbeddingabstractWith the development of information technology, online health care service platforms have collected a large amount of patient comment information. Through fine-grained sentiment classification of this information, we can provide references for patients to seek medical treatment and help doctors understand their work. Therefore, we proposed a model that integrated pre-trained emotional information and semantic information at the word and character level for aspect-level sentiment classification. Specifically, we employed adversarial learning for training sentiment word embeddings and the two-layer bidirectional long short-term memory network to extract the sentiment embedding of the entire sentence in a specific aspect. We also combined the pre-trained sentiment feature vector with the structured semantic information by linear weighting and the multi-head self-attention mechanism, enabling the model to pay more attention to the information most relevant to a given aspect category. We performed experiments on the Chinese patient comments data set constructed by our research team and the proposed model outperformed the state-of-the-art methods, which proved the effectiveness of the proposed model for aspect-level sentiment classification of Chinese patient comments. Yongxue Shan, Zhaoqian Zhong, Chao Che, Bo Jin 0001, Xiaopeng Wei |
BIBM | 5 |
| 2021 | A Novel Gaze-Point-Driven HRI Framework for Single-Person
Qiang Zhang 0008, Xiaopeng Wei, Rui Liu 0015, Jing Dong 0009 |
CollaborateCom (1) | 5 |
| 2021 | A Novel and Efficient Distance Detection Based on Monocular Images for Grasp and Handover
Dianwen Liu, Qiang Zhang 0008, Xiaopeng Wei, Rui Liu 0015, Jing Dong 0009 |
CollaborateCom (1) | 5 |
| 2021 | Asymmetric Anomaly Detection for Human-Robot InteractionabstractSecurity in human-robot interaction is the focus of research in this field. Rapid detection of abnormal events that may cause danger in the interaction process can effectively reduce the probability of occurrence of danger. In general anomaly detection methods, 2D or 3D convolutional autoencoders are widely used for anomaly detection. Among them, 2D convolutional autoencoders are with good real-time performance and lower detection accuracy, while 3D convolutional autoencoders are with higher detection accuracy and insufficient real-time performance. In order to ensure realtime performance and obtain higher accuracy, an end-to-end asymmetric convolutional autoencoder network (ACANet) using both 2D and 3D convolutions is designed. Specifically, 3D convolution is used to build the encoder to learn comprehensive information in continuous input frames, and 2D convolution is used to build the decoder to model the information fast, a dimensional alignment module is constructed to connect the encoder and the decoder while avoiding a large number of calculations in the latent space of the 3D features output by the encoder, and the skip connections module is used to obtain accurate predictions. Anomaly detection can then be completed by evaluating the differences between results predicted by the ACANet and real frames. The experimental results show that our method achieves competitive accuracy on mainstream datasets and at the same time obtains the fastest speed. Compared with mainstream methods, this method is more suitable for anomaly detection tasks in human-robot interaction. Rui Liu 0015, Yingkun Hou, Qiang Zhang 0008, Xiaopeng Wei |
CSCWD | 7 |
| 2021 | Depth-Aware Mirror SegmentationabstractWe present a novel mirror segmentation method that leverages depth estimates from ToF-based cameras as an additional cue to disambiguate challenging cases where the contrast or relation in RGB colors between the mirror reflection and the surrounding scene is subtle. A key observation is that ToF depth estimates do not report the true depth of the mirror surface, but instead return the total length of the reflected light paths, thereby creating obvious depth dis-continuities at the mirror boundaries. To exploit depth information in mirror segmentation, we first construct a large-scale RGB-D mirror segmentation dataset, which we subsequently employ to train a novel depth-aware mirror segmentation framework. Our mirror segmentation framework first locates the mirrors based on color and depth discontinuities and correlations. Next, our model further refines the mirror boundaries through contextual contrast taking into account both color and depth information. We extensively validate our depth-aware mirror segmentation method and demonstrate that our model outperforms state-of-the-art RGB and RGB-D based methods for mirror segmentation. Experimental results also show that depth is a powerful cue for mirror segmentation. Haiyang Mei, Bo Dong 0004, Wen Dong 0008, Pieter Peers, Xin Yang 0011, Qiang Zhang 0008, Xiaopeng Wei |
CVPR | 7 |
| 2021 | Camouflaged Object Segmentation With Distraction MiningabstractCamouflaged object segmentation (COS) aims to identify objects that are "perfectly" assimilate into their surroundings, which has a wide range of valuable applications. The key challenge of COS is that there exist high intrinsic similarities between the candidate objects and noise background. In this paper, we strive to embrace challenges towards effective and efficient COS. To this end, we develop a bio-inspired framework, termed Positioning and Focus Network (PFNet), which mimics the process of predation in nature. Specifically, our PFNet contains two key modules, i.e., the positioning module (PM) and the focus module (FM). The PM is designed to mimic the detection process in predation for positioning the potential target objects from a global perspective and the FM is then used to perform the identification process in predation for progressively refining the coarse prediction via focusing on the ambiguous regions. Notably, in the FM, we develop a novel distraction mining strategy for the distraction discovery and removal, to benefit the performance of estimation. Extensive experiments demonstrate that our PFNet runs in real-time (72 FPS) and significantly outperforms 18 cutting-edge models on three challenging datasets under four standard metrics. Haiyang Mei, Ge-Peng Ji, Ziqi Wei 0001, Xin Yang 0011, Xiaopeng Wei, Deng-Ping Fan |
CVPR | 5 |
| 2021 | A Study on Realtime Task Selection Based on Credit Information Updating in Evolutionary Multitasking
Yumeng Cao, Yaqing Hou, Liang Feng 0001, Hong-Wei Ge, Qiang Zhang 0008, Xiaopeng Wei |
EMO | 6 |
| 2021 | Object Tracking by Jointly Exploiting Frame and Event DomainabstractInspired by the complementarity between conventional frame-based and bio-inspired event-based cameras, we propose a multi-modal based approach to fuse visual cues from the frame- and event-domain to enhance the single object tracking performance, especially in degraded conditions (e.g., scenes with high dynamic range, low light, and fast-motion objects). The proposed approach can effectively and adaptively combine meaningful information from both domains. Our approach’s effectiveness is enforced by a novel designed cross-domain attention schemes, which can effectively enhance features based on self- and cross-domain attention schemes; The adaptiveness is guarded by a specially designed weighting scheme, which can adaptively balance the contribution of the two domains. To exploit event-based visual cues in single-object tracking, we construct a large-scale frame-event-based dataset, which we subsequently employ to train a novel frame-event fusion based model. Extensive experiments show that the proposed approach outperforms state-of-the-art frame-based tracking methods by at least 10.4% and 11.9% in terms of representative success rate and precision rate, respectively. Besides, the effectiveness of each key component of our approach is evidenced by our thorough ablation study. Jiqing Zhang, Xin Yang 0011, Yingkai Fu, Xiaopeng Wei, Bo Dong 0004 |
ICCV | 4 |
| 2021 | Multi-Scale Attention Constraint Network for Fine-Grained Visual ClassificationabstractCapturing subtle yet discriminative features constitutes a great challenge in fine-grained visual classification due to the large intra-class and small inter-class variances. Main-stream works for this problem localize at attention mechanism and feature relationship learning. However, existing methods treat the features in isolation while neglecting the effect of attention-enhanced features on relationships between different network layers. In this paper, we propose a novel attention-based method by Multi-Scale Attention Constraint network composed of two important components: (1) a feature extractor with lightweight group-wise enhanced attention blocks that guides the generation of high representation features; and (2) a multi-scale regularizer that explores the relationships between different features. Extensive experiments show that our approach achieves state-of-the-art performance on standard benchmark datasets. Moreover, we introduce a new dataset, consisting of comprehensive surgical instrument categories based on three common surgeries, to support the classification and inventory work of surgical instruments. Yaqing Hou, Hong-Wei Ge, Qiang Zhang 0008, Xiaopeng Wei |
ICME | 6 |
| 2021 | Exploring the effects of computational costs in extensive games via modeling and simulationabstractGame theory has become a standard tool for depicting and demonstrating various game-like phenomena by providing appropriate mathematical models and for analyzing and predicting agents' behaviors and their decisions by formalizing solution concepts. The conventional game model mainly concerns ideal systems that would always guarantee optimal responses, which appears unrealistic for practical game scenarios since decision-making usually entails resource costs. Therefore, this study considers players' decision-making in extensive games when the computational cost of searching the strategy space is limited. We start with a new mathematical model of extensive games that features a bound on computational resources during players' decision-making process such that they can only foresee a part of the available alternatives in the future. This model is more appropriate in predicting players' strategies than the conventional model, under which we investigate the effects of computational costs on players' strategies as well as the computational complexity. Furthermore, a simulation experiment is performed to seek the connection between the amount of resources and the goodness of the outcomes. This study is expected to provide a foundation for players' rational decision-making with computational costs. Chanjuan Liu 0001, Enqiang Zhu, Qiang Zhang 0008, Xiaopeng Wei |
Int. J. Intell. Syst. | 4 |
| 2021 | ASFNet: Adaptive multiscale segmentation fusion network for real-time semantic segmentationabstractAbstract Recently, the development of deep learning has facilitated continuous progress in the field of computer vision. Pixel‐level semantic segmentation serves as a fundamental task in computer vision. It achieves significant results by connecting wider and deeper backbone networks and building fine‐grained segmentation heads. However, applications such as self‐driving cars are more critical to the computational speed of the algorithms. The trade‐off between accuracy and real‐time performance of existing algorithms is still a challenging task. To address this challenge, this article proposes an adaptive multiscale segmentation fusion network to fuse multiscale contextual, which designs an adaptive multiscale segmentation fusion module based on an attention mechanism. Using segmentation fusion instead of feature fusion, the multiscale segmentation results are aggregated to obtain more precise segmentation results. The final results achieved 70.9% mIoU of accuracy in the Cityspace test set, processing images at 61 FPS when the input is 1024 × 2048. In addition, when adjusting the input size to 512 × 1024, the images are processed at 185 FPS. Hengfeng Zha, Rui Liu 0015, Xin Yang 0011, Qiang Zhang 0008, Xiaopeng Wei |
Comput. Animat. Virtual Worlds | 6 |
| 2021 | MeSIN: Multilevel selective and interactive network for medication recommendation
Liang Zhang 0031, Mao You, Xueqing Tian, Bo Jin 0001, Xiaopeng Wei |
Knowl. Based Syst. | 6 |
| 2021 | Automatic Comic Generation with Stylistic Multi-page Layouts and Emotion-driven Text Balloon GenerationabstractIn this article, we propose a fully automatic system for generating comic books from videos without any human intervention. Given an input video along with its subtitles, our approach first extracts informative keyframes by analyzing the subtitles and stylizes keyframes into comic-style images. Then, we propose a novel automatic multi-page layout framework that can allocate the images across multiple pages and synthesize visually interesting layouts based on the rich semantics of the images (e.g., importance and inter-image relation). Finally, as opposed to using the same type of balloon as in previous works, we propose an emotion-aware balloon generation method to create different types of word balloons by analyzing the emotion of subtitles and audio. Our method is able to vary balloon shapes and word sizes in balloons in response to different emotions, leading to more enriched reading experience. Once the balloons are generated, they are placed adjacent to their corresponding speakers via speaker detection. Our results show that our method, without requiring any user inputs, can generate high-quality comic pages with visually rich layouts and balloons. Our user studies also demonstrate that users prefer our generated results over those by state-of-the-art comic generation systems. Xin Yang 0011, Zongliang Ma, Letian Yu, Ying Cao 0001, Xiaopeng Wei, Qiang Zhang 0008, Rynson W. H. Lau |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2021 | Smart Scribbles for Image MattingabstractImage matting is an ill-posed problem that usually requires additional user input, such as trimaps or scribbles. Drawing a fine trimap requires a large amount of user effort, while using scribbles can hardly obtain satisfactory alpha mattes for non-professional users. Some recent deep learning–based matting networks rely on large-scale composite datasets for training to improve performance, resulting in the occasional appearance of obvious artifacts when processing natural images. In this article, we explore the intrinsic relationship between user input and alpha mattes and strike a balance between user effort and the quality of alpha mattes. In particular, we propose an interactive framework, referred to as smart scribbles, to guide users to draw few scribbles on the input images to produce high-quality alpha mattes. It first infers the most informative regions of an image for drawing scribbles to indicate different categories (foreground, background, or unknown) and then spreads these scribbles (i.e., the category labels) to the rest of the image via our well-designed two-phase propagation. Both neighboring low-level affinities and high-level semantic features are considered during the propagation process. Our method can be optimized without large-scale matting datasets and exhibits more universality in real situations. Extensive experiments demonstrate that smart scribbles can produce more accurate alpha mattes with reduced additional input, compared to the state-of-the-art matting methods. Xin Yang 0011, Yu Qiao 0001, Shaozhe Chen, Shengfeng He, Qiang Zhang 0008, Xiaopeng Wei, Rynson W. H. Lau |
ACM Trans. Multim. Comput. Commun. Appl. | 7 |
| 2020 | Efficient Attention Calibration Network for Real-Time Semantic SegmentationabstractIn recent years, the attention mechanism has been widely used in computer vision. Semantic segmentation, as one of the fundamental tasks of computer vision, has been subject to tremendous development as a result. But because of its huge computing overhead, attention-based approaches are difficult to use for real-time applications such as self-driving. In this paper, we propose a self-calibration method baesd on self-attentiion that successfully applies the attention mechanism to real-time semantic segmentation. Specifically, a spatial attention module to adjust the edges of the coarse segmentation results which gained from the real-time semantic segmentation backbone network, and obtain more granular segmentation results. We refer to this method as the Efficient Attentional Calibration Network (EACNet). Experiments on the Cityscapes dataset validate the efficiency and performance of the method. With the high-resolution input and without any post-processing, EACNet achieved 72.4% mIoU of accuracy while running at 116.9 FPS. Compared to other state-of-the-art methods for real-time semantic segmentation, our network gained a better balance between performance and speed. Hengfeng Zha, Rui Liu 0015, Xin Yang 0011, Qiang Zhang 0008, Xiaopeng Wei |
ACML | 6 |
| 2020 | Don't Hit Me! Glass Detection in Real-World ScenesabstractGlass is very common in our daily life. Existing computer vision systems neglect it and thus may have severe consequences, e.g., a robot may crash into a glass wall. However, sensing the presence of glass is not straightforward. The key challenge is that arbitrary objects/scenes can appear behind the glass, and the content within the glass region is typically similar to those behind it. In this paper, we propose an important problem of detecting glass from a single RGB image. To address this problem, we construct a large-scale glass detection dataset (GDD) and design a glass detection network, called GDNet, which explores abundant contextual cues for robust glass detection with a novel large-field contextual feature integration (LCFI) module. Extensive experiments demonstrate that the proposed method achieves more superior glass detection results on our GDD test set than state-of-the-art methods fine-tuned for glass detection. Haiyang Mei, Xin Yang 0011, Yang Wang 0106, Shengfeng He, Qiang Zhang 0008, Xiaopeng Wei, Rynson W. H. Lau |
CVPR | 7 |
| 2020 | Attention-Guided Hierarchical Structure Aggregation for Image MattingabstractExisting deep learning based matting algorithms primarily resort to high-level semantic features to improve the overall structure of alpha mattes. However, we argue that advanced semantics extracted from CNNs contribute unequally for alpha perception and we are supposed to reconcile advanced semantic information with low-level appearance cues to refine the foreground details. In this paper, we propose an end-to-end Hierarchical Attention Matting Network (HAttMatting), which can predict the better structure of alpha mattes from single RGB images without additional input. Specifically, we employ spatial and channel-wise attention to integrate appearance cues and pyramidal features in a novel fashion. This blended attention mechanism can perceive alpha mattes from refined boundaries and adaptive semantics. We also introduce a hybrid loss function fusing Structural SIMilarity (SSIM), Mean Square Error (MSE) and Adversarial loss to guide the network to further improve the overall foreground structure. Besides, we construct a large-scale image matting dataset comprised of 59,600 training images and 1000 test images (total 646 distinct foreground alpha mattes), which can further improve the robustness of our hierarchical structure aggregation model. Extensive experiments demonstrate that the proposed HAttMatting can capture sophisticated foreground structure and achieve state-of-the-art performance with single RGB images as input. Yu Qiao 0001, Yuhao Liu 0001, Xin Yang 0011, Mingliang Xu 0001, Qiang Zhang 0008, Xiaopeng Wei |
CVPR | 7 |
| 2020 | Fast Sparse Connectivity Network Adaption via Meta-LearningabstractPartial correlation-based connectivity networks can describe the direct connectivity between features while avoiding spurious effects, and hence they can be implemented in diagnosing complex dynamic multivariate systems. However, existing studies mainly focus on single systems that are ill-equipped for incremental learning. Moreover, related methods estimate temporal connectivity network by imposing only sparse regularization without integrating pattern priors (e.g., inter-system shared pattern and intra-system intrinsic pattern), which have been proven effective in limiting noise interference. To this end, we develop an adaptive connectivity estimation model that incorporates prior patterns, namely Sparse Adaptive Meta-Learning Connectivity Network (SAMCN). Specifically, our model extends ideas of the gradient-based meta-learning to capture inter-system shared prior information by generating fast adaptive initialization parameters for the connectivity matrix. Then, a sparse variational autoencoder is proposed to generate a weight matrix for sparse regularization penalty in reweighted LASSO, which helps extract intra-system intrinsic patterns (local manifold structure). Experimental results on both synthetic data and real-world datasets demonstrate that our method is capable of adequately capturing the aforementioned pattern priors. Further, experiments from corresponding classification tasks validate the strength of the prior pattern-aware features connectivity network in resulting in better classification performance. Bo Jin 0001, Ke Cheng 0003, Liang Zhang 0031, Keli Xiao, Xinjiang Lu, Xiaopeng Wei |
ICDM | 7 |
| 2020 | Multi-scale Information Assembly for Image MattingabstractAbstract Image matting is a long‐standing problem in computer graphics and vision, mostly identified as the accurate estimation of the foreground in input images. We argue that the foreground objects can be represented by different‐level information, including the central bodies, large‐grained boundaries, refined details, etc. Based on this observation, in this paper, we propose a multi‐scale information assembly framework (MSIA‐matte) to pull out high‐quality alpha mattes from single RGB images. Technically speaking, given an input image, we extract advanced semantics as our subject content and retain initial CNN features to encode different‐level foreground expression, then combine them by our well‐designed information assembly strategy. Extensive experiments can prove the effectiveness of the proposed MSIA‐matte, and we can achieve state‐of‐the‐art performance compared to most existing matting networks. Yu Qiao 0001, Yuhao Liu 0001, Xin Yang 0011, Yuxin Wang 0001, Qiang Zhang 0008, Xiaopeng Wei |
Comput. Graph. Forum | 7 |
| 2020 | RAHM: Relation augmented hierarchical multi-task learning framework for reasonable medication stocking
Yakun Mao, Liang Zhang 0031, Bo Jin 0001, Keli Xiao, Xiaopeng Wei, Jun Yan 0010 |
J. Biomed. Informatics | 6 |
| 2019 | Where Is My Mirror?
Xin Yang 0011, Haiyang Mei, Ke Xu 0010, Xiaopeng Wei, Rynson W. H. Lau |
ICCV | 4 |
| 2019 | Real-virtual consistent traffic flow interaction
Xin Yang 0011, Shuai Li 0014, Xinglin Piao, Qiang Zhang 0008, Xiaopeng Wei |
Graph. Model. | 8 |
| 2019 | DEMC: A Deep Dual-Encoder Network for Denoising Monte Carlo Rendering
Xin Yang 0011, Wenbo Hu 0002, Lijing Zhao, Qiang Zhang 0008, Xiaopeng Wei, Hongbo Fu 0001 |
J. Comput. Sci. Technol. | 7 |
| 2019 | Cascaded network with deep intensity manipulation for scene understandingabstractAbstract Scene understanding is essential to robotic navigation and autonomous driving as it provides semantic information to their controlling system. However, it will fail when processing low‐light images/videos captured under adverse weather or at night use state‐of‐the‐art scene understanding methods. A naive way to directly infer semantics from low‐light images is ill posed because the low‐light condition distorts pixel intensities and buries details. In order to address this problem, we propose the Deep Intensity Manipulation Network (DIMNet), which could relight the input images and recover the details, and combine the DIMNet with a scene understanding network to get a cascaded network to learn the semantics from low‐light images. Through learning pixel intensity manipulation, our method can generate images not only visually pleasing but also practical for scene understanding. Qualitative and quantitative experiments demonstrate that the proposed method is effective and robust for both synthetic and real‐world images. Xin Yang 0011, Shaozhe Chen, Xinglin Piao, Qiang Zhang 0008, Xiaopeng Wei |
Comput. Animat. Virtual Worlds | 8 |
| 2019 | DRFN: Deep Recurrent Fusion Network for Single-Image Super-Resolution With Large FactorsabstractRecently, single-image super-resolution has made great progress due to the development of deep convolutional neural networks (CNNs). The vast majority of CNN-based models use a predefined upsampling operator, such as bicubic interpolation, to upscale input low-resolution images to the desired size and learn nonlinear mapping between the interpolated image and ground truth high-resolution (HR) image. However, interpolation processing can lead to visual artifacts as details are over smoothed, particularly when the super-resolution factor is high. In this paper, we propose a deep recurrent fusion network (DRFN), which utilizes transposed convolution instead of bicubic interpolation for upsampling and integrates different-level features extracted from recurrent residual blocks to reconstruct the final HR images. We adopt a deep recurrence learning strategy and, thus, have a larger receptive field, which is conducive to reconstructing an image more accurately. Furthermore, we show that the multilevel fusion structure is suitable for dealing with image super-resolution problems. Extensive benchmark evaluations demonstrate that the proposed DRFN performs better than most current deep learning methods in terms of accuracy and visual effects, especially for large-scale images, while using fewer parameters. Xin Yang 0011, Haiyang Mei, Jiqing Zhang, Ke Xu 0010, Qiang Zhang 0008, Xiaopeng Wei |
IEEE Trans. Multim. | 7 |
| 2018 | Image Correction via Deep Reciprocating HDR TransformationabstractImage correction aims to adjust an input image into a visually pleasing one. Existing approaches are proposed mainly from the perspective of image pixel manipulation. They are not effective to recover the details in the under/over exposed regions. In this paper, we revisit the image formation procedure and notice that the missing details in these regions exist in the corresponding high dynamic range (HDR) data. These details are well perceived by the human eyes but diminished in the low dynamic range (LDR) domain because of the tone mapping process. Therefore, we formulate the image correction task as an HDR transformation process and propose a novel approach called Deep Reciprocating HDR Transformation (DRHT). Given an input LDR image, we first reconstruct the missing details in the HDR domain. We then perform tone mapping on the predicted HDR data to generate the output LDR image with the recovered details. To this end, we propose a united framework consisting of two CNNs for HDR reconstruction and tone mapping. They are integrated end-to-end for joint training and prediction. Experiments on the standard benchmarks demonstrate that the proposed method performs favorably against state-of-the-art image correction methods. Xin Yang 0011, Ke Xu 0010, Yibing Song, Qiang Zhang 0008, Xiaopeng Wei, Rynson W. H. Lau |
CVPR | 5 |
| 2018 | Deep High-order Supervised Hashing for Image RetrievalabstractRecently, deep hashing has achieved excellent performances in large-scale image retrieval by simultaneously learning deep features and hashing function. However, state-of-the-art works have so far failed to explore the feature statistics higher than first-order. In this paper, to take a step towards addressing this problem, we propose two novel Deep High-order Supervised Hashing architectures (DHoSH), i.e., point-wise labels based DHoSH (DHoSH-PO) and pair-wise labels based DHoSH (DHoSH-PA). The core of DHoSH is that a trainable layer of bilinear pooling incorporates into deep convolutional neural networks (CNNs) for end-to-end learning. This layer captures the local feature interactions of the image by outer product, employing the autocorrelation information and cross-correlation information of deep features. Furthermore, our DHoSH method systematically exploits the high-order statistics of features of multiple layers. Extensive experiments on commonly used benchmarks illuminate that both DHoSH-PO and DHoSH-PA can obtain competitive improvements over its first-order counterparts, and achieve state-of-the-art performance for image retrieval task. Jingdong Cheng, Qiule Sun, Jianxin Zhang 0001, Xiaopeng Wei, Qiang Zhang 0008 |
ICPR | 4 |
| 2018 | Active Object Reconstruction Using a Guided View PlannerabstractInspired by the recent advance of image-based object reconstruction using deep learning, we present an active reconstruction model using a guided view planner. We aim to reconstruct a 3D model using images observed from a planned sequence of informative and discriminative views. But where are such informative and discriminative views around an object? To address this we propose a unified model for view planning and object reconstruction, which is utilized to learn a guided information acquisition model and to aggregate information from a sequence of images for reconstruction. Experiments show that our model (1) increases our reconstruction accuracy with an increasing number of views (2) and generally predicts a more informative sequence of views for object reconstruction compared to other alternative methods. Xin Yang 0011, Qiang Zhang 0008, Xiaopeng Wei, Hongbo Fu 0001 |
IJCAI | 6 |
| 2018 | Passivity of Reaction-Diffusion Genetic Regulatory Networks with Time-Varying Delays
Xiaopeng Wei, Qiang Zhang 0008, Changjun Zhou |
Neural Process. Lett. | 2 |
| 2018 | Constructing DNA Barcode Sets Based on Particle Swarm OptimizationabstractFollowing the completion of the human genome project, a large amount of high-throughput bio-data was generated. To analyze these data, massively parallel sequencing, namely next-generation sequencing, was rapidly developed. DNA barcodes are used to identify the ownership between sequences and samples when they are attached at the beginning or end of sequencing reads. Constructing DNA barcode sets provides the candidate DNA barcodes for this application. To increase the accuracy of DNA barcode sets, a particle swarm optimization (PSO) algorithm has been modified and used to construct the DNA barcode sets in this paper. Compared with the extant results, some lower bounds of DNA barcode sets are improved. The results show that the proposed algorithm is effective in constructing DNA barcode sets. Bin Wang 0005, Xuedong Zheng, Shihua Zhou, Changjun Zhou, Xiaopeng Wei, Qiang Zhang 0008, Ziqi Wei 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2018 | Modeling of Agent Cognition in Extensive Games via Artificial Neural NetworksabstractThe decision-making process, which is regarded as cognitive and ubiquitous, has been exploited in diverse fields, such as psychology, economics, and artificial intelligence. This paper considers the problem of modeling agent cognition in a class of game-theoretic decision-making scenarios called extensive games. We present a novel framework in which artificial neural networks are incorporated to simulate agent cognition regarding the structure of the underlying game and the goodness of the game situations therein. An algorithmic procedure is investigated to describe the process for solving games with cognition, and then, a new equilibrium concept is proposed as a refinement of the classical one-subgame perfect equilibrium-by involving players' cognitive reasoning. Moreover, a series of results concerning the computational complexity, soundness, and completeness of the algorithm, as well as the existence of an equilibrium solution, is obtained. This framework, which is shown to be general enough to model the way in which AlphaGo plays Go, may offer a means for bridging the gap between theoretical models and practical problem-solving. Chanjuan Liu 0001, Enqiang Zhu, Qiang Zhang 0008, Xiaopeng Wei |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2018 | Efficient image super-resolution integration
Ke Xu 0010, Xin Wang 0118, Xin Yang 0011, Shengfeng He, Qiang Zhang 0008, Xiaopeng Wei, Rynson W. H. Lau |
Vis. Comput. | 7 |
| 2017 | Multitask Dyadic Prediction and Its Application in Prediction of Adverse Drug-Drug InteractionabstractAdverse drug-drug interactions (DDIs) remain a leading cause of morbidity and mortality around the world. Identifying potential DDIs during the drug design process is critical in guiding targeted clinical drug safety testing. Although detection of adverse DDIs is conducted during Phase IV clinical trials, there are still a large number of new DDIs founded by accidents after the drugs were put on market. With the arrival of big data era, more and more pharmaceutical research and development data are becoming available, which provides an invaluable resource for digging insights that can potentially be leveraged in early prediction of DDIs. Many computational approaches have been proposed in recent years for DDI prediction. However, most of them focused on binary prediction (with or without DDI), despite the fact that each DDI is associated with a different type. Predicting the actual DDI type will help us better understand the DDI mechanism and identify proper ways to prevent it. In this paper, we formulate the DDI type prediction problem as a multitask dyadic regression problem, where the prediction of each specific DDI type is treated as a task. Compared with conventional matrix completion approaches which can only impute the missing entries in the DDI matrix, our approach can directly regress those dyadic relationships (DDIs) and thus can be extend to new drugs more easily. We developed an effective proximal gradient method to solve the problem. Evaluation on real world datasets is presented to demonstrate the effectiveness of the proposed approach. Bo Jin 0001, Cao Xiao, Ping Zhang 0016, Xiaopeng Wei, Fei Wang 0001 |
AAAI | 5 |
| 2017 | LSTM based classification model and its application for doctor-patient relationship evaluationabstractThe emergence of medical social media has made it possible for more and more patients to share their views and experiences on the medical care platform. These subjective texts contains patients' evaluation information for doctors and can be analyzed to provide rich decision-making information for patients and hospitals. Therefore, we propose a LSTM (Long Short Term Memory) based text sentiment classification method to evaluate the relationship between doctors and patients through the comment data from medical social media. The classification model of doctor-patient relationship can also be performed to evaluate the hospital by calculating the praise rate to help people choose hospitals. We perform two experiments on the patients' evaluation data from the website of Haodaifu. The experiment of doctor-patient relationship classification confirms the effectiveness of the classification model. In the experiment of hospital evaluation, we calculate the praise rate of 11 hospitals in Jinan of Shandong province based on the doctor-patient relationship classification results. The consistency between the results obtained by our method and the data of Mingyihui shows that our evaluation method for hospitals is reasonable and effective. Hongrui Kuang, Chao Che, Qiang Zhang 0008, Xiaopeng Wei |
Healthcom | 4 |
| 2017 | Classification of ECG signals based on 1D convolution neural networkabstractRecently, with the obvious increasing number of cardiovascular disease, the automatic classification research of Electrocardiogram signals (ECG) has been playing a significantly important part in the clinical diagnosis of cardiovascular disease. In this paper, a 1D convolution neural network (CNN) based method is proposed to classify ECG signals. The proposed CNN model consists of five layers in addition to the input layer and the output layer, i.e., two convolution layers, two down sampling layers and one full connection layer, extracting the effective features from the original data and classifying the features automatically. This model realizes the classification of 5 typical kinds of arrhythmia signals, i.e., normal, left bundle branch block, right bundle branch block, atrial premature contraction and ventricular premature contraction. The experimental results on the public MIT-BIH arrhythmia database show that the proposed method achieves a promising classification accuracy of 97.5%, significantly outperforming several typical ECG classification methods. Jianxin Zhang 0001, Qiang Zhang 0008, Xiaopeng Wei |
Healthcom | 4 |
| 2017 | Exploring risk factors and predicting UPDRS score based on Parkinson's speech signalsabstractThe unified Parkinson's disease rating scale (UPDRS) is the most widely employed scale for tracking Parkinson's disease (PD) symptom progression. However, conventional way to achieve UPDRS, mainly based on the physical examinations of clinic patients performed by the trained medical staffs, involves the disadvantages of inconvenience and high medical expense. Hence, in this study, we try to explore some risk factors and accurately predict the UPDRS for PD, using the speech signals of PD patients published on UCI machine-learning archive. More specifically, inspired by the idea of ensemble learning, we firstly construct a framework of ensemble feature selection (EFS) to select a suitable subset of features among numerous speech signals. Subsequently, a personalized predictive model, trained by adopting information from similar patients, is developed to be customized for an individual PD patient. Finally, we employ the personalized predictive model to predict UPDRS score combined with various classical regression algorithms. Compared to conventional models, our study has a potential to capture more relevant risk factors and produces more accurate UPDRS score for individual patient. Experimental results on real-world dataset from UCI machine-learning archive show that our personalized predictive model gets a promising performance. Jianxin Zhang 0001, Qiang Zhang 0008, Bo Jin 0001, Xiaopeng Wei |
Healthcom | 5 |
| 2017 | Interactive traffic simulation model with learned local parameters
Xin Yang 0011, Shuai Li 0014, Wanchao Su, Guozhen Tan, Qiang Zhang 0008, Xiaopeng Wei |
Multim. Tools Appl. | 9 |
| 2016 | DKD: a fast k-d tree update design for dynamic scenesabstractAbstract We design dynamic k‐d (DKD) tree based on classical k‐d tree for animated scene rendering. Our method can inherit the benefit of efficient traversal of k‐d tree and minimize time cost to update DKD tree, making it well suited for animated geometry. DKD employs primitive reset, redistribution to reflect the updated positions of geometry, and leaf node incremental growing to avoid the deterioration of hierarchy quality due to refitting. Our experiments show that DKD has a significant rendering performance improvement than selected existing methods. Copyright © 2016 John Wiley & Sons, Ltd. Xin Yang 0011, Pengfei Zhang 0016, Lutong Xin, Yuxin Wang 0001, Qiang Zhang 0008, Xiaopeng Wei |
Comput. Animat. Virtual Worlds | 8 |
| 2015 | 3D Ear Shape Matching Using Joint α-Entropy
Xiao-Peng Sun, Si-Hui Li, Xiaopeng Wei |
J. Comput. Sci. Technol. | 4 |
| 2014 | Smart Partitioning for Product DSM Model Based on Improved Genetic Algorithm
Yangjie Zhou 0002, Chao Che, Jianxin Zhang 0001, Qiang Zhang 0008, Xiaopeng Wei |
ADMA | 5 |
| 2014 | 3D ear recognition using local salience and principal manifold
Hongyan Sun, Xiaopeng Wei |
Graph. Model. | 5 |
| 2014 | Forward non-rigid motion tracking for facial MoCap
Xiaoyong Fang, Xiaopeng Wei, Qiang Zhang 0008 |
Vis. Comput. | 2 |
| 2013 | On the simulation of expressional animation based on facial MoCap
Xiaoyong Fang, Xiaopeng Wei, Qiang Zhang 0008, Changjun Zhou |
Sci. China Inf. Sci. | 2 |
| 2012 | A novel color image encryption algorithm based on DNA sequence operation and hyper-chaotic system
Xiaopeng Wei, Qiang Zhang 0008, Jianxin Zhang 0001, Shiguo Lian |
J. Syst. Softw. | 1 |
| 2010 | Some novel classification and learning methods and applications for neural networks - Selected papers from the Second International Conference on Bio-Inspired Computing: Theories and Applications
Xiaopeng Wei, Qiang Zhang 0008, Guangzhao Cui |
Neurocomputing | 1 |
| 2009 | Robust pitch estimation using a wavelet variance analysis model
Xiaopeng Wei, Lasheng Zhao, Qiang Zhang 0008, Jing Dong 0009 |
Signal Process. | 1 |
| 2008 | Multi-objective optimization design of gear reducer based on adaptive genetic algorithmabstractAn adaptive genetic algorithm (GA) is introduced to solve the multi-objective optimization design of the reducer. Firstly, according to the structure, strength, etc. in a reducer, a multi-objective optimization model of the helical gear reducer is established. And then an adaptive GA based on a fuzzy controller is introduced, aiming at the characteristic of multi-objective, multiparameter, multi-constraint conditions. Finally, a numerical example is illustrated to show the advantages of this approach and the effectiveness of an adaptive genetic algorithm used in optimized design of a reducer. Rui Li 0055, Tian Chang, Xiaopeng Wei |
CSCWD | 4 |
| 2008 | Product Schemes Evaluation Method Based on Improved BP Neural Network
Xiaopeng Wei |
ICIC (2) | 2 |
| 2008 | Research on Product Case Representation and Retrieval Based on Ontology
Xiaopeng Wei |
ICIC (2) | 2 |
| 2006 | A Novel Global Exponential Stability Result for Discrete-time Cellular Neural Networks with Variable DelaysabstractGlobal exponential stability is considered for a class of discrete-time cellular neural networks with variable delays. By employing a discrete Halanay inequality, a new result is presented ensuring global exponential stability of the unique equilibrium point of the networks. The result extends and improves the earlier publications due to the fact that it removes some restrictions on the delay. An example is given to illustrate the effectiveness of the global exponential stability condition provided here. Qiang Zhang 0008, Xiaopeng Wei, Jin Xu 0015 |
Int. J. Neural Syst. | 2 |
| 2005 | Global Exponential Stability of Discrete Time Hopfield Neural Networks with Delays
Qiang Zhang 0008, Wenbing Liu, Xiaopeng Wei |
ISNN (1) | 3 |
| 2005 | Global Asymptotic Stability Analysis of Neural Networks with Time-Varying Delays
Qiang Zhang 0008, Xiaopeng Wei, Jin Xu 0015 |
Neural Process. Lett. | 2 |
| 2004 | Asymptotic Stability of Nonautonomous Delayed Neural Networks
Qiang Zhang 0008, Xiaopeng Wei, Jin Xu 0015 |
ICONIP | 2 |
| 2004 | On the Asymptotic Stability of Non-autonomous Delayed Neural Networks
Qiang Zhang 0008, Xiaopeng Wei |
ISNN (1) | 4 |