VLDB 2026 Research / reviewers in the wild / expert
Yongxiong Wang
dblp:72/9605
· DBLP profile ↗
47ranked-venue papers
5as first author
36since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 22 · 2 first-author · 19 since 2021Artificial intelligence and machine learning · 19 · 3 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Systems, architecture and hardware · 1Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ORDMP: self-supervised monocular depth estimation via optical-flow-reconstructed directional masks and large-model teacher pseudo-labels
Shuwen Jia, Yongxiong Wang, Shuai Huang 0002 |
Pattern Anal. Appl. | 2 |
| 2026 | OpenVL: Bridging 2-D and 3-D Worlds for Open-Vocabulary 3-D Scene UnderstandingabstractOpen-world 3D scene understanding is critical for applications like autonomous driving and robotics, but it faces challenges due to the scarcity of dense 3D-text pairings. While 2D vision-language models offer rich pretraining knowledge, bridging the 2D-3D gap remains challenging. To address this issue, we propose OpenVL and utilize the framework of the Vision-Language model for Open-vocabulary understanding. First, we construct a channel that connects 2D and 3D features through the method of 3D Patches, which largely avoids the problems arising from the scarcity of 3D data. Secondly, we propose a Multi-level 3D Patches Alignment. By achieving dense alignment with 3D Patches at the point, instance, and scene levels, our method not only preserves the spatial perception ability within 3D scenes but also enables a fine-grained understanding of these scenes. Experimental results across multiple bench-marks demonstrate that our approach has achieved state-of-the-art performance. OpenVL achieves significant performance improvements on ScanNet (+6.7% mAP) and nuScenes (+1.9% mAP). More importantly, it demonstrates excellent zero-shot generalization capabilities in downstream tasks including 3DVQA and 3D Embodied Planning.Embodied Planning. Yongxiong Wang, Shuai Huang 0002 |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2025 | SSAAD: A Multi-Scale Temporal-Frequency Graph Network for Binary Auditory Attention Detection with Self-Supervised LearningabstractAuditory attention detection (AAD) from electroencephalography (EEG) signals has garnered significant interest for its potential in brain-computer interfaces and hearing aids. Nevertheless, accurate decoding remains challenging due to the high-dimensional, non-stationary, and inherently noisy characteristics of EEG signals. We introduce SSAAD, a novel self-supervised approach for binary AAD with three key innovations: (1) A multi-scale temporal-frequency graph convolutional network (MST-GCN) is developed to effectively capture both spatial and temporal EEG dynamics; (2) A hierarchical self-supervised contrastive learning method is designed to cultivate robust EEG representations without extensive labeled data; (3) A bespoke triple-stage mixup data augmentation strategy is proposed to enhance model generalization. Additionally, MST-GCN initialization is achieved via EEG autoencoder pre-training, facilitating superior feature extraction. The method's effectiveness is evaluated on the KUL and DTU datasets, demonstrating state-of-the-art performance with accuracy improvements of 1.4% and 1.7% for 0.1-second windows, respectively. Shuai Huang 0002, Yongxiong Wang, Shuwen Jia, Chendong Qin, Zhongcai He |
ICASSP | 2 |
| 2025 | SVTNet: Dual Branch of Swin Transformer and Vision Transformer for Monocular Depth EstimationabstractIn monocular depth estimation, effective acquisition of global and local information is the key to improving accuracy. We introduce a novel dual branch network called Swin Vision Transformer Net (SVTNet), where the Swin Transformer and Vision Transformer are combined to learn features with global and local information. In Swin Transformer branch, we remove the Feature Pyramid Networks which reduces the noise introduced by interpolation and maintains the original characteristics of the hierarchical features extracted from each layer of Swin Transformer. In Vision Transformer Branch, we remove the multi-scale convolutional structure and keep the same-scale feature maps extracted by Vision Transformer for reducing the consumption of memory and extracting more global features. The experimental results on the NYU-Depth v2, KITTI, and CityScapes datasets show that compared with existing advanced methods, the SVTNet has achieved significant performance improvements in multiple evaluation metrics. Shuwen Jia, Yongxiong Wang, Shuai Huang 0002 |
ICASSP | 2 |
| 2025 | MINDEV: Multi-modal Integrated Diffusion Framework for Video Reconstruction from EEG SignalsabstractDespite recent progress in decoding static images from brain activity, reconstructing dynamic visual experiences from EEG signals remains challenging due to the complex temporal dynamics involved. Current approaches primarily rely on pre-trained video generation models while failing to fully leverage the rich temporal-spatial information embedded in EEG signals for video synthesis. This paper proposes MINDEV Multi-modal Integrated Neural DEcoding and Visualization), a framework that places EEG signal processing at the core of video reconstruction. We introduce three key technical contributions: (1) a dual-branch feature extractor that captures both temporal dynamics and spatial relationships in EEG signals, (2) an EEG-driven semantic bridge that uses neural patterns to guide language model interpretation, and (3) a multi-modal video synthesis pipeline where EEG features lead the generation process while semantic guidance provides refinement. Our framework prioritizes the millisecond-level temporal resolution of EEG signals, using them to drive both visual content generation and semantic understanding. Evaluated on the SEED-DV dataset, MINDEV demonstrates superior performance with a semantic classification accuracy of 93.2% and a structural similarity index (SSIM) of 0.4777, establishing a new state-of-the-art for EEG-based video reconstruction. Our code is publicly available at https://github.com/HHarr1son/MINDEV. Shuai Huang 0002, Yongxiong Wang, Haodong Jing, Chendong Qin, Jingqun Tang |
ACM Multimedia | 2 |
| 2025 | NEED: Cross-Subject and Cross-Task Generalization for Video and Image Reconstruction from EEG SignalsabstractTranslating brain activity into meaningful visual content has long been recognized as a fundamental challenge in neuroscience and brain-computer interface research. Recent advances in EEG-based neural decoding have shown promise, yet two critical limitations remain in this area: poor generalization across subjects and constraints to specific visual tasks. We introduce NEED, the first unified framework achieving zero-shot cross-subject and cross-task generalization for EEG-based visual reconstruction. Our approach addresses three fundamental challenges: (1) cross-subject variability through an Individual Adaptation Module pretrained on multiple EEG datasets to normalize subject-specific patterns, (2) limited spatial resolution and complex temporal dynamics via a dual-pathway architecture capturing both low-level visual dynamics and high-level semantics, and (3) task specificity constraints through a unified inference mechanism adaptable to different visual domains. For video reconstruction, NEED achieves better performance than existing methods. Importantly, Our model maintains 93.7% of within-subject classification performance and 92.4% of visual reconstruction quality when generalizing to unseen subjects, while achieving an SSIM of 0.352 when transferring directly to static image reconstruction without fine-tuning, demonstrating how neural decoding can move beyond subject and task boundaries toward truly generalizable brain-computer interfaces. Shuai Huang 0002, Haodong Jing, Qixian Zhang, Litao Chang, Yating Feng, Chendong Qin, Shuwen Jia, Siyi Sun, Yongxiong Wang |
NeurIPS | 12 |
| 2025 | A novel robust data synthesis method based on feature subspace interpolation to optimize samples with unknown noise
Yukun Du, Yitao Cai, Haiyue Yu 0001, Zhilong Lou, Jiang Jiang 0001, Yongxiong Wang |
Inf. Sci. | 8 |
| 2025 | Trusted Video-Based Sewer Inspection via Support Clip-Based Pareto-Optimal Evidential NetworkabstractAn automatic vision-based sewer inspection plays a vital role of sewage system in a modern city. Existing methods have utilized evidential deep learning to construct trusted models. Although the acceptable performance has been achieved in sewer defect classification, the fine-grained information of sewer defects in videos is ignored. Meanwhile, the trade-off between multi-label classification and uncertainty estimation remains challenging. In this paper, support clip-based pareto-optimal evidential network (POEN) is proposed for trusted video-based sewer inspection. Specifically, support clip module (SCM) is designed to capture the fine-grained visual representation of defects from local scale segments. Then, evidential deep learning is introduced to quantify the uncertainty for out-of-distribution detection. Furthermore, Pareto-optimal weighting scheme (PWS) is designed to solve the common trade-off dilemma in multi-task learning. Extensive experiments are conducted on VideoPipe, in which the superiority of POEN is demonstrated compared with the state-of-the-art methods. Chenyang Zhao 0009, Chuanfei Hu, Hang Shao 0001, Fir Dunkin, Yongxiong Wang |
IEEE Signal Process. Lett. | 5 |
| 2025 | Active Domain Adaptation Based on Probabilistic Fuzzy C-Means Clustering for Pancreatic Tumor Segmentation
Chendong Qin, Yongxiong Wang, Fubin Zeng, Yangsen Cao, Xiaolan Yin, Shuai Huang 0002, Huojun Zhang, Zhiyong Ju |
IEEE Trans. Fuzzy Syst. | 2 |
| 2024 | Mitigating Optimization Conflict in Domain Adversarial Neural Network via Uncertainty-AwareabstractIn prior studies, domain adversarial neural networks (DANNs) are used to align image-level features regardless of foreground and background. However, the conventional discriminator in DANNs may leads the feature extractor to disregard cross-domain features rather than aligning them. This phenomenon negatively impact classifier performance. We propose a novel loss reweighting technique that mitigates the optimization conflict between discriminator and classifier. The classification loss is reweighted based on the prediction uncertainty that is measured by two different bottleneck layers. This reweighting approach guides the model in determining which features should be activated or aligned, resulting in significantly improved adaptation performance. Additionally, we introduce a novel construction method of bottleneck layer based on pseudo label of target domain and differentiable architecture search to support our approach. Our method is rigorously evaluated across multiple benchmark datasets and outperforms state-of-the-art (SOTA) methods. Zhiqun Pan, Yongxiong Wang, Guangpeng Wang |
ICASSP | 2 |
| 2024 | ASD: Towards Attribute Spatial Decomposition for Prior-Free Facial Attribute RecognitionabstractRepresenting the spatial properties of facial attributes is a vital challenge for facial attribute recognition (FAR). Recent advances have achieved the reliable performances for FAR, benefiting from the description of spatial properties via extra prior information. However, the extra prior information might not be always available, resulting in the restricted application scenario of the prior-based methods. Meanwhile, the spatial ambiguity of facial attributes caused by inherent spatial diversities of facial parts is ignored. To address these issues, we propose a prior-free method for attribute spatial decomposition (ASD), mitigating the spatial ambiguity of facial attributes. The attribute components could be formally described in terms of the spatial locations without any extra prior information. Experimental results demonstrate the superiority of ASD compared with state-of-the-art prior-based methods on both CelebA and LFWA. Chuanfei Hu, Hang Shao 0001, Bo Dong 0001, Zhe Wang 0027, Yongxiong Wang |
ICME | 5 |
| 2024 | Cross-domain self-supervised few-shot learning via multiple crops with teacher-student network
Guangpeng Wang, Yongxiong Wang, Zhiqun Pan |
Eng. Appl. Artif. Intell. | 2 |
| 2024 | CMLCNet: medical image segmentation network based on convolution capsule encoder and multi-scale local co-occurrence
Chendong Qin, Yongxiong Wang |
Multim. Syst. | 2 |
| 2024 | Chfnet: a coarse-to-fine hierarchical refinement model for monocular depth estimation
Yongxiong Wang |
Mach. Vis. Appl. | 2 |
| 2024 | Unsupervised anomaly detection and localization via bidirectional knowledge distillation
Yongxiong Wang, Zhiqun Pan, Guangpeng Wang |
Neural Comput. Appl. | 2 |
| 2024 | Weakly supervised instance segmentation via class double-activation maps and boundary localization
Yongxiong Wang, Zhiqun Pan |
Signal Process. Image Commun. | 2 |
| 2024 | TMFF: Trustworthy Multi-Focus Fusion Framework for Multi-Label Sewer Defect Classification in Sewer Inspection VideosabstractAn automatic vision-based sewer inspection plays a vital role of sewage system in a modern city. Recent advances focus on modeling a deep learning-based method to realize the sewer inspection system, benefiting from the capability of data-driven feature extraction. Although the acceptable performances of sewer defect classification are achieved, there is still a gap between the emerged methods and actual application scenarios. The first issue is that the multi-focus complementarity is ignored to represent the sewer defect, resulting in capturing the multi-scale information of sewer defect inefficiently. Second, the inherent uncertainty of sewer defect is not considered, while the serious unknown sewer defect categories would be missed, resulting in the untrustworthy sewer inspection. In this paper, we focus on quick-view (QV)-based sewer inspection, while a trustworthy multi-focus fusion framework (TMFF) is proposed, jointly combining multi-label classification and uncertainty estimation. Specifically, focal segment module (FSM) is designed based on optical flow to split the QV sewer video into long-focus and short-focus segments, where the multi-focus segments can be modeled to represent the multi-scale information of sewer defect. Then, evidential deep learning (EDL) is introduced to quantify the uncertainty, while joint expert scheme (JES) is designed to aggregate the expert opinions of multi-focus segments. Moreover, evidential disambiguating strategy (EDS) is proposed to alleviate the ambiguity of uncertainty estimation. Extensive experiments are conducted on VideoPipe, in which the superiority of TMFF is demonstrated compared with the state-of-the-art methods. Furthermore, we validate the potential capability of TMFF against the unknown cases of sewer defects. Chuanfei Hu, Chenyang Zhao 0009, Hang Shao 0001, Jin Deng, Yongxiong Wang |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | LDTSF: A Label-Decoupling Teacher-Student Framework for Semi-Supervised Echocardiography SegmentationabstractThe accurate segmentation of the right and left ventricles with limited labeled data is a challenging task in echocardiographic data analysis. To fully leverage the easily accessible unlabeled data, we propose a label-decoupling teacher-student framework (LDTSF) based on semi-supervised learning. Specifically, the decoupled deep network within LDTSF jointly predicts pixel-wise segmentation maps, level set-based edge regression maps, target skeleton maps and target detail maps to focus on edge pixels. Several micro-task-transformable layers are used to map multi-task representations to a unified space in order to supervise the consistency among multiple tasks using massive unlabeled data. In addition, we first train a teacher model based on semi-supervised learning strategy, and then use the pseudo-labels generated by the teacher model together with the original labels to train a student model. Experiments on our self-collected 3D echocardiographic dataset and a publicly available MRI dataset show that our method outperforms state-of-the-art semi-supervised learning methods. The code will be available at: https://github.com/SwanKnightZJP/LDTSF. Yongxiong Wang, Zhiqun Pan, Zhenhui Tang |
ICASSP | 2 |
| 2023 | Towards Trustworthy Multi-Label Sewer Defect Classification via Evidential Deep LearningabstractAn automatic vision-based sewer inspection plays a key role of sewage system in a modern city. Recent advances focus on utilizing deep learning model to realize the sewer inspection system, benefiting from the capability of data-driven feature representation. However, the inherent uncertainty of sewer defects is ignored, resulting in the missed detection of serious unknown sewer defect categories. In this paper, we propose a trustworthy multi-label sewer defect classification (TMSDC) method, which can quantify the uncertainty of sewer defect prediction via evidential deep learning. Meanwhile, a novel expert base rate assignment (EBRA) is proposed to introduce the expert knowledge for describing reliable evidences in practical situations. Experimental results demonstrate the effectiveness of TMSDC and the superior capability of uncertainty estimation is achieved on the latest public benchmark. Chenyang Zhao 0009, Chuanfei Hu, Hang Shao 0001, Zhe Wang 0027, Yongxiong Wang |
ICASSP | 5 |
| 2023 | Inter-subject cognitive workload estimation based on a cascade ensemble of multilayer autoencoders
Zhanpeng Zheng, Yongxiong Wang |
Expert Syst. Appl. | 3 |
| 2023 | Self-attention network for few-shot learning based on nearest-neighbor algorithm
Guangpeng Wang, Yongxiong Wang |
Mach. Vis. Appl. | 2 |
| 2023 | Correction to: Self-attention network for few-shot learning based on nearest-neighbor algorithm
Guangpeng Wang, Yongxiong Wang |
Mach. Vis. Appl. | 2 |
| 2023 | Dual-Branch TransV-Net for 3-D Echocardiography SegmentationabstractSegmentation of left and right ventricles from 3-D echocardiographic images is premised and key for the quantitative analysis of cardiac function, which is important for pediatric cardiac diagnosis. Compared with 2-D echocardiography, 3-D echocardiography can fully represent a ventricular structure without geometric inference. However, it usually takes experts several hours to obtain a 3-D segmentation mask of the left and right ventricles. Therefore, a fast and automatic segmentation method is highly desired. Unfortunately, 3-D echocardiography suffers from low contrast, unclear left and right ventricles borders, and blind zone. Existing segmentation methods usually have poor performance at ventricular boundaries. To deal with these problems, we propose a novel dual-branch TransV-Net (DBTV). DBTV comprises two parallel, interleaved, and relatively independent V-shaped encoder–decoder branches. The main branch acts on the original data to extract image features, and the auxiliary branch acts on edge maps to extract the additional edge features. To suppress noise and enhance the edge information, extra concatenations are added to bridge the features from the main and auxiliary branches. To reduce object missing caused by blind zone, a 3-D transformer-based module is proposed in the bottom layer of the dual-branch structure to extract the global contexts. We do experiments on a self-collected dataset with 120 3-D echocardiographic images from 60 cardiac sequences, and the dice scores of 0.913 and 0.880 are obtained in the left and right ventricle segments, respectively. Each inference takes about two-thirds of a second for a single 3-D frame. Yongxiong Wang, Sunjie Zhang, Zhiqun Pan, Zhe Wang 0027, Zhenhui Tang, Ying Guo 0019 |
IEEE Trans. Ind. Informatics | 2 |
| 2022 | An Anomaly Detection Method Based on Self-Supervised Learning with Soft Label Assignment for Defect Visual InspectionabstractRecently, local-editing-based transformations are introduced in anomaly detection for defect visual inspection, which construct a pretext task with the paradigm of self-supervised learning. However, supervised information of local-editing-based transformation may be incorrect when invalid trans-formation occurs in the pretext task. The reason is that the conventional method to generate labels ignores the differences of images between before and after the transformation. To address this issue, we propose soft label assignment (SLA) to construct soft labels via measuring the similarity between the original and transformed images. Meanwhile, a novel self-supervised learning-based anomaly detection method is proposed for defect visual inspection, which exploits local-editing-based transformation with SLA as a pretext classification task. A convolutional neural network (CNN) is trained to extract deep features of ambiguity and irregularity by the pretext classification task. In the main task, an anomaly detection is modeled via the deep representations to estimate defects regarded as anomalies. Experimental results demonstrate the effect of SLA, and the proposed method achieves superior performance than state-of-the-art methods in terms of the receiver operating characteristic curve (AUC-ROC). Chuanfei Hu, Yongxiong Wang |
ICASSP | 2 |
| 2022 | Regressive Pseudo Label for Weakly Supervised Facial Landmark DetectionabstractThe progress of the deep neural network and visual sensors promote the facial landmark detection. However, faces are easily collected from Internet of Things (IoT) devices worldwide but are hard-labeled facial landmarks in consistent style. Though many weak supervision (WS) algorithms and theories have been proposed to handle the labeling problem, most of them are tailored for classification tasks and fail in regression tasks, especially in facial landmark detection. To tackle this WS regression task, first, we propose a regressive pseudo-labeling method by analyzing the cluster assumption of facial landmark detection, where overlaps are reduced and interval areas are increased among clusters of facial parts for unlabeled faces. Moreover, auxiliary information and domain loss are utilized to adapt model to samples with different styles. Second, we design a generator–regressor network to first estimate facial boundary attentions and then to locate facial landmarks. However, when generating pseudo labels based on predictive models automatically, there are two major issues. One is that the results of the multistage network highly depend on the former-stage accuracy, and another is that jointing different annotation styles always produces ambiguity feature representation. Thus, we propose two ideas. During training, the generator and regressor are decoupled to alleviate the inner dependence of the multistage network, and the heatmap discriminator is introduced to improve the quality of the predicted facial boundary. We design a transformer structure to fuse face image and boundary attentions, so that it can further complement useful features. Based on these ideas, our methods can automatically annotate accurate pseudo facial landmarks for unlabeled faces. Extensive experiments show that our model achieves good performance on different benchmark data sets, both in accuracy and efficiency. Zhiqun Pan, Yongxiong Wang, Bin Tan 0001 |
IEEE Internet Things J. | 2 |
| 2022 | Rich global feature guided network for monocular depth estimation
Bingyuan Wu, Yongxiong Wang |
Image Vis. Comput. | 2 |
| 2022 | Progressive Erasing Network with consistency loss for fine-grained visual classification
Yongxiong Wang, Zeping Zhou |
J. Vis. Commun. Image Represent. | 2 |
| 2022 | Multiple organ-specific cancers classification from PET/CT images using deep learning
Yongxiong Wang, Zhenhui Tang, Zhe Wang 0027 |
Multim. Tools Appl. | 2 |
| 2022 | Fovea localization by blood vessel vector in abnormal fundus images
Yinghua Fu, Dongyan Pan, Yongxiong Wang, Dawei Zhang 0009 |
Pattern Recognit. | 5 |
| 2022 | Joint face detection and Facial Landmark Localization using graph match and pseudo label
Zhiqun Pan, Yongxiong Wang, Sunjie Zhang |
Signal Process. Image Commun. | 2 |
| 2021 | Feature Aggregation Network with Tri-Hybrid Loss for Instance SegmentationabstractLimited by the size of feature maps in the mask head, the operation of simply stacking convolutional blocks in previous instance segmentation methods cannot obtain comprehensive features. In contrast, we design a Feature Aggregation Net-work (FAN) where three separate modules are respectively responsible for extracting salient features from the dimensions of channel, space, and scale. Then, these salient features are further aggregated based on their affinities generated by an Affinity Computation Module (ACM) with the original features. In this way, features become more distinct, which is conducive to the subsequent mask prediction. To further improve the segmentation performance of hard examples with-out introducing extra inference overhead, we also propose a novel loss function named Tri-hybrid loss where the optimization can be simultaneously performed from the perspectives of pixel, boundary, and instance. With these improvements, our model outperforms state-of-the-arts with 39.5 mask AP on COCO test-dev2017. Zeping Zhou, Yongxiong Wang |
ICME | 2 |
| 2021 | Multi-branch Channel-wise Enhancement Network for Fine-grained Visual RecognitionabstractThe challenge in fine-grained visual classification (FGVC) is that the similarity within intra-class may be larger than inter-class, where the discriminative details require more attention than traditional classification tasks. To generate channel-wise complementary and discriminative features in beneficial details of FGVC, we propose a multi-branch channel-wise enhancement network (MCEN), which includes multi-pattern spatial disruption mechanism, inter-channel complementarity module(ICM), and novel soft target loss. The raw images are scrambled in multi-pattern and then the sub-images with different degrees of confusion are combined into three pairs as inputs, where the scrambled operation can force the channel to look for the discriminative details. And ICM can measure the complementarity between key features and overall features to restrain the redundancy of features. The soft target loss is designed for classification and the semantic relationship between the blocks is learned to judge the degree of the chaos of the image. Our designed multi-branched structure utilizes the shallow visual and deep semantic features to judge the outcome jointly, where the image pairs obtained by segmentation and rearrangement are input into the different branches to extract more complementary features from different patterns of the raw image. Our method is trained end-to-end with only class labels. Experimental results show that our model outperforms the state-of-the-art performance on three fine-grained benchmarks. Guangjun Li, Yongxiong Wang, Fengting Zhu |
ACM Multimedia | 2 |
| 2021 | Learning efficient multi-task stereo matching network with richer feature information
Sunjie Zhang, Yongxiong Wang |
Neurocomputing | 3 |
| 2021 | Generative image inpainting with salient prior and relative total variation
Hang Shao 0001, Yongxiong Wang |
J. Vis. Commun. Image Represent. | 2 |
| 2021 | A sparse focus framework for visual fine-grained classification
Yongxiong Wang, Guangjun Li |
Multim. Tools Appl. | 1 |
| 2021 | Nondestructive Defect Detection in Castings by Using Spatial Attention Bilinear Convolutional Neural NetworkabstractX-ray images of castings are widely used in manufacturing for quality assurance. This article investigates the X-ray-image-based defective detection. The main contributions in this article are twofold: first, a new full-image method is proposed to classify defective castings and nondefective ones; and second, by combining two technologies, spatial attention mechanism and bilinear pooling used in deep convolutional neural networks (CNNs), a new spatial attention bilinear CNN is proposed to enhance the representation power of CNN. To validate the above initiatives, extensive experimental studies have been carried out to show the advantages of the new method over a number of existing ones. Zhenhui Tang, Engang Tian, Yongxiong Wang, Licheng Wang 0003 |
IEEE Trans. Ind. Informatics | 3 |
| 2020 | Salient Object Detection with Boundary InformationabstractHow to distinguish the low-contrast area near boundaries is a basic challenge in salient object detection. Most of recent state-of-the-art methods can achieve a good performance but still can't work well near boundaries. In this paper, we propose a novel network based on multi-level feature fusion with boundary information to solve this problem. Our model includes two separate decoding sub-networks, one is object sub-network to detect salient objects and another is boundary sub-network which outputs error maps to get boundary information by boundary maps. Moreover, we design a connection and fusion module to exchange and fuse information of objects and boundaries. In addition, to balance the two subnetworks, the optimal weight of loss function is obtained by experiments. The experimental results show that our model can distinguish the low-contrast area near boundaries well by boundary information and achieves the state-of-the-art performance on five common datasets. Yongxiong Wang, Chuanfei Hu, Hang Shao 0001 |
ICME | 2 |
| 2020 | Locally robust EEG feature selection for individual-independent emotion recognition
Lei Liu 0035, Jianing Chen 0002, Boxi Zhao, Yongxiong Wang |
Expert Syst. Appl. | 5 |
| 2020 | Real-time salient object detection with boundary information guidance
Yongxiong Wang, Yan Song 0002 |
Neurocomputing | 1 |
| 2020 | View-based weight network for 3D object recognition
Yongxiong Wang |
Image Vis. Comput. | 2 |
| 2020 | Generative image inpainting via edge structure and color aware fusion
Hang Shao 0001, Yongxiong Wang, Yinghua Fu |
Signal Process. Image Commun. | 2 |
| 2019 | Physiological-signal-based mental workload estimation via transfer dynamical autoencoders in a deep learning framework
Mengyuan Zhao 0001, Wei Zhang 0134, Yongxiong Wang, Jianhua Zhang 0004 |
Neurocomputing | 4 |
| 2017 | A novel local feature descriptor based on energy information for human activity recognition
Yongxiong Wang, Yubo Shi, Guoliang Wei |
Neurocomputing | 1 |
| 2016 | Probabilistic framework of visual anomaly detection for unbalanced data
Yongxiong Wang, Xueming Ding |
Neurocomputing | 1 |
| 2015 | A Local Feature Descriptor Based on Energy Information for Human Activity Recognition
Yubo Shi, Yongxiong Wang |
ICIC (3) | 2 |
| 2015 | A self-adaptive weighted affinity propagation clustering for key frames extraction on human action recognition
Yongxiong Wang, Shuxin Sun, Xueming Ding |
J. Vis. Commun. Image Represent. | 1 |
| 2005 | Dynamic performance enhancement of PVDF force sensor for micromanipulationabstractSo far, in-situ PVDF (polyvinylidene fluoride) films bonded to the surface of flexible cantilever structure act as the micro-force sensors, they are mostly modelled using quasi-static relationships. However, such sensors are usually a significantly compliant and easily deformable structure in order to reach highly sensitive performance in micromanipulation. As a result, this may be reasonable to consider bandwidth measurement and high frequency response for achievement of high accuracy, and thus a dynamic analysis of such sensors become essentially necessary. In this paper, a cantilever beam based micro-force sensor was designed based on the infinite dimensional system model (distributed parameter model) using the Bernouli-Euler formulation. Furthermore, in order to enable an engineering implementation, we used a zero frequency term to effectively replace the high order modes of the dynamic sensing model in the prescribed frequency range. The corrected model can efficiently minimize the effect of removed higher order modes, and then the micro-force measurement can be obtained accurately with this corrected in-bandwidth dynamic model. Preliminary simulation and experimental results both verified the performance of the developed dynamic micro-force sensor and the effectiveness of the corrected model. Yantao Shen 0001, Ning Xi 0001, Wen Jung Li, Yongxiong Wang |
IROS | 4 |