Bingshu Wang

dblp:50/301 · DBLP profile ↗
← Back
26ranked-venue papers
9as first author
19since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 13 · 4 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 5 · 4 first-author · 5 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 CATNet: Coordinate-aware transformer for all-in-one image restoration
Junling He, Yang Zhao 0038, Ziyang Chen 0002, Bingshu Wang, Yongjun Zhang 0007
Expert Syst. Appl.6
2026 Flexible multi-view feature selection with semi-supervised label semantic alignment
Han Zhang 0012, Bingshu Wang, Feiping Nie 0001, Xuelong Li 0001
Pattern Recognit.3
2025 Look Inside for More: Internal Spatial Modality Perception for 3D Anomaly Detection
abstract
3D anomaly detection has recently become a significant focus in computer vision. Several advanced methods have achieved satisfying anomaly detection performance. However, they typically concentrate on the external structure of 3D samples and struggle to leverage the internal information embedded within samples. Inspired by the basic intuition of why not look inside for more, we observed this prototype is straightforward and effective. As a result, we introduce a newly designed mode named Internal Spatial Modality Perception (ISMP) to explore the feature representation from internal views fully. Specifically, our proposed ISMP consists of a critical perception module, Spatial Insight Engine (SIE), which abstracts complex internal information of point clouds into essential global features. Besides, to better align structural information with point data, we propose an enhanced key point feature extraction method for amplifying spatial structure feature representation. Simultaneously, a novel feature filtering module is incorporated to reduce noise and redundant features for further precise spatial structure aligning. Extensive experiments validate the efficiency of our proposed method, achieving object-level and pixel-level AUROC improvements of 4.2% and 13.1%, respectively, on the Real3D-AD benchmarks. Note that the strong generalization ability of SIE has been theoretically proven and verified in both classification and segmentation tasks. Our code will be released upon acceptance.
Hanzhe Liang, Guoyang Xie, Chengbin Hou, Bingshu Wang, Can Gao, Jinbao Wang 0001
AAAI4
2025 FBI-Net: Frequency Band Integration Network for Infrared Small Target Segmentation
abstract
Small targets in infrared imagery exhibit challenging characteristics due to their minimal semantic information and the extremely imbalanced distribution between the targets and the background. In this paper, we propose a frequency band integration network to extract salient features of infrared small targets in both the spatial and frequency domains. To excavate the high-frequency features of the small targets, we propose a frequency decoupling-fusion module. To decrease the semantic loss that occurs in deep networks, we propose a semantic injection mechanism to assist in retaining critical information from shallow layers. Experimental results show that our proposed method reaches higher prediction accuracy and robustness in the infrared small target segmentation task compared with other state-of-the-art approaches.
Biqiao Xin, Qianchen Mao, Jinbao Wang 0001, Bingshu Wang
ICASSP5
2025 SLVS: A Self-Learning Approach to Achieve Near-Second Low-Latency Video Streaming Under Highly Variable Networks
abstract
Fueled by the rapid advances in high-speed mobile networks, live video streaming has seen explosive growth in recent years and some DASH-based algorithms were specifically proposed for low-latency video delivery. We conducted a measurement study for the state-of-the-art algorithms with large-scale network traces. It reveals that these algorithms are susceptible to network condition changes due to the use of solo universal adaptation logics, resulting in the playback latency that has substantial variations across highly fluctuating networks. To tackle this challenge, this paper proposes Stateful Live Video Streaming (SLVS), which is a novel self-learning approach that learns the various network features and optimizes the adaptation logic separately for different network conditions, then dynamically tunes the logic at runtime, so that bitrate decision can better match the changing networks. Moreover, we further generalize SLVS to complement the streaming platform already in service to make it compatible with any live streaming services. Extensive evaluations based on real system prototypes show that SLVS can control playback latency down to 1 s while improving Quality-of-Experience (QoE) by 17.7% to 31.8%. Moreover, it has strong robustness to maintain near-second latency over highly fluctuating networks as well as long periods of video viewing.
Ke Liu 0004, Mengbai Xiao, Bingshu Wang, Dongxiao Yu, Xiuzhen Cheng
IEEE Trans. Mob. Comput.4
2025 DIG: Improved DINO for Graffiti Detection
abstract
Graffiti detection is essential in historic building protection and urban neighborhood management. Graffiti detection has made significant progress in recent years based on the development of deep learning. However, small-scale graffiti, interference from the background, and the false detection of word parts in graffiti make it a challenging problem. This article proposes a Transformer-based high-precision graffiti detection method, namely DIG. Precisely, it consists of three modules: 1) Spatial query selection (SQS), scale-aware IoU loss (SIL); 2) Denoising Task with binary contrastive denoising (BCDN); and 3) IoU-guided box denoising (IBD) modules. To detect small-scale graffiti, this paper proposes SIL to help the loss function to perceive small-scale graffiti and large-scale graffiti fairly. To reduce the false detection of word parts, this article presents the SQS module, which integrates spatial information into the query selection process of the Encoder to filter out falsely detected bounding boxes within the graffiti. To reduce the interference from the background, this article introduces a denoising task with BCDN and IBD modules, improving the model’s ability to distinguish graffiti from the background and accurately select appropriate bounding boxes. A large number of experimental results on the STORM dataset show that our method achieves state-of-the-art results with an$AP_{50}$of 87.9%. Moreover, DIG achieved competitive results on the FineFM dataset for mask detection. This indicates that DIG can also be conveniently transferred to detect other scenarios.
Bingshu Wang, Qianchen Mao, Aifei Liu, Long Chen 0001, C. L. Philip Chen
IEEE Trans. Syst. Man Cybern. Syst.1
2024 MoCha-Stereo: Motif Channel Attention Network for Stereo Matching
abstract
Learning-based stereo matching techniques have made significant progress. However, existing methods inevitably lose geometrical structure information during the feature channel generation process, resulting in edge detail mis-matches. In this paper, the Motif Channel Attention Stereo Matching Network (MoCha-Stereo) is designed to address this problem. We provide the Motif Channel Correlation Volume (MCCV) to determine more accurate edge matching costs. MCCV is achieved by projecting motif channels, which capture common geometric structures in feature channels, onto feature maps and cost volumes. In addition, edge variations in the reconstruction error map also affect details matching, we propose the Reconstruction Error Motif Penalty (REMP) module to further refine the full-resolution disparity estimation. REMP integrates the frequency information of typical channel features from the re-construction error. MoCha-Stereo ranks 1st on the KITTI-2015 and KITTI-2012 Reflective leaderboards. Our structure also shows excellent performance in Multi- View Stereo.
Ziyang Chen 0002, He Yao, Yongjun Zhang 0007, Bingshu Wang, Yongbin Qin
CVPR5
2024 Local Information Guided Global Integration for Infrared Small Target Detection
abstract
Infrared small targets often exhibit small scale and weak semantic features, which makes it a great challenge to their detection. To address this situation, we propose a novel network for infrared small target detection that combines local details information and global contextual information. To preserve the local and high-frequency details present in infrared images, we introduce a High-frequency Aware Encoder. To extract contextual information from multi-scale feature maps, we propose a Multi-scale Context Learning Bottleneck that incorporates contextual information repeatedly and performs cross-level fusion, which enables the recognition of small targets based on their surroundings. Finally, a lightweight Transformer Decoder is employed to restore the feature map, while placing attention on the target pixels. Experimental results on the IRSTD-1k dataset demonstrate that our method outperforms other state-of-the-art approaches.
Qianchen Mao, Jinbao Wang 0001, Wenmin Wang 0001, Bingshu Wang
ICASSP6
2024 Concrete Structural Crack Damage Classification Using Nonlinear Dimension Reduction and Broad Learning System
abstract
Concrete structural crack damage classification is of importance for road safety. This paper proposes a new method based on broad neural network for crack damage classification in concrete structures. It includes three stages. Firstly, a pre-trained deep neural network is used to extract the features from crack images. Secondly, principal component analysis is used to project the retrieved features from high dimensions to low dimensions. Thirdly, broad learning system is employed to predict the classification using the low-dimensional features. Experimental results demonstrate that this method reduces the model's training time and improves classification accuracy.
Bingshu Wang, C. L. Philip Chen
SMC1
2024 SpirDet: Toward Efficient, Accurate, and Lightweight Infrared Small-Target Detector
abstract
In recent years, the detection of infrared small targets using deep learning methods has garnered substantial attention due to notable advancements. To improve the detection capability of small targets, these methods commonly maintain a pathway that preserves high-resolution (HR) features of sparse and tiny targets. However, it can result in redundant and expensive computations. To tackle this challenge, we propose a sparse infrared-target detector (SpirDet) for the efficient detection of infrared small targets. Specifically, to cope with the computational redundancy issue, we employ a new dual-branch sparse decoder (DBSD) to restore the feature map. First, the fast branch directly predicts a sparse map indicating potential small-target locations. Second, the slow branch conducts fine-grained adjustments at the positions indicated by the sparse map. In addition, we design a lightweight DO-RepEncoder based on reparameterization with the downsampling orthogonality (DO), which can effectively reduce memory consumption and inference latency. Extensive experiments show that the proposed SpirDet significantly outperforms state-of-the-art (SOTA) models while achieving faster inference speed and fewer parameters. For example, on the NUDT-SIRST dataset, SpirDet improves mean intersection over union (MIoU) by 2.09 and has a$3.1\times $frames/s acceleration compared to the previous SOTA model. The code is available athttps://github.com/laCorse/SpirDet-Pytorch.
Qianchen Mao, Qiang Li 0055, Bingshu Wang, Yongjun Zhang 0007, Tao Dai 0001, C. L. Philip Chen
IEEE Trans. Geosci. Remote. Sens.3
2024 Hierarchical Multimodality Graph Reasoning for Remote Sensing Visual Question Answering
abstract
Remote sensing visual question answering (RSVQA) targets answering the questions about RS images in natural language form. RSVQA in real-world applications is always challenging, which may contain wide-field visual information and complicated queries. The current methods in RSVQA overlook the semantic hierarchy of visual and linguistic information and ignore the complex relations of multimodal instances. Thus, they severely suffer from vital deficiencies in comprehensively representing and associating the vision–language semantics. In this research, we design an innovative end-to-end model, namedHierarchicalMultimodalityGraphReasoning (HMGR) network, which hierarchically learns multigranular vision–language joint representations, and interactively parses the heterogeneous multimodal relationships. Specifically, we design a hierarchical vision–language encoder (HVLE), which could simultaneously represent multiscale vision features and multilevel language features. Based on the representations, the vision–language semantic graphs are built, and the parallel multimodal graph relation reasoning is posed, which could explore the complex interaction patterns and implicit semantic relations of both intramodality and intermodality instances. Moreover, we raise a distinctive vision–question (VQ) feature fusion module for the collaboration of information at different semantic levels. Extensive experiments on three public large-scale datasets (RSVQA-LR, RSVQA-HRv1, and RSVQA-HRv2) demonstrate that our work is superior to the state-of-the-art results toward a mass of vision and query types.
Han Zhang 0012, Keming Wang, Laixian Zhang, Bingshu Wang, Xuelong Li 0001
IEEE Trans. Geosci. Remote. Sens.4
2024 Cross-Modality Knowledge Calibration Network for Video Corpus Moment Retrieval
abstract
Video corpus moment retrieval has become a hot topic recently, which aims to localize a consequent video moments highly relevant to the given query language description from video corpus. Existing methods towards this challenging task are suffering from the cases when the visual information and textual information in the video are very different from each other or from the cases where the redundant video content is semantically irrelevant with the query language description, which make the model confused of figuring out the truly useful within- and cross-modality information. In this article, we propose a novel Cross-Modality Knowledge Calibration Network (CKCN) to solve the issue mentioned above. Specifically, a dual calibration transformer module with improved multi-head attention is proposed to simultaneously capture the within- and cross-modality features between the visual and textual modality of the video automatically compressing the redundant information, and then a query-dependent fusion module is designed to guide feature fusion of the video's multi-modal information using the prior knowledge of query which further refine more important modality features. At last, a query-guided calibration transformer module with a well-designed learnable cell is utilized to align the query and video, forming a single joint representation for moment localization. Meanwhile, we introduce transfer learning into the task of video corpus moment retrieval (VCMR) for the first time to solve the defect of insufficient labeled data. Extensive experiments have been conducted on both the widely used TVR dataset and DiDeMo dataset which have achieved new state-of-the-art, thus verifying the effectiveness of our proposed CKCN.
Tongbao Chen, Wenmin Wang 0001, Ruochen Li 0001, Bingshu Wang
IEEE Trans. Multim.5
2023 Shadow Removal of Text Document Images Using Background Estimation and Adaptive Text Enhancement
abstract
This paper proposes a simple yet effective method to re-move shadows from text document images. It mainly includes several parts. Firstly, we propose a text elimination-based background extraction strategy to estimate shadow map. It indicates the shadow regions accurately and helps to predict global background. Secondly, a binarization-based text ex-traction algorithm is designed to obtain texts from document image. By fusing texts and global background, a preparatory shadow-free image can be obtained. Thirdly, we propose an adaptive text contrast enhancement strategy to generate shadow-free results with comfortable visual perception across shadow and non-shadow regions. Quantitative and visual results performed on open datasets indicate that the proposed method can generate clear shadow-free images from text document images. Our code will be publicly available soon.
Bingshu Wang, Jiangbin Zheng 0001, Wenmin Wang 0001
ICASSP2
2023 An Intelligent Learning Approach to Achieve Near-Second Low-Latency Live Video Streaming under Highly Fluctuating Networks
abstract
Fueled by the rapid advances in high-speed mobile networks, live video streaming has seen explosive growth in recent years and many DASH-based bitrate adaptive streaming algorithms were specifically proposed for low-latency video delivery. However, our investigations revealed that these algorithms are susceptible to network condition changes due to the use of solo universal adaptation logics, resulting the playback latency that has substantial variations across highly-fluctuating network environments and fails to meet the service quality requirement all the time. To tackle this challenge, this paper proposes Stateful Live Video Streaming (SLVS), which is a novel learning approach that learns the various network features and optimizes the adaptation logic separately for different network conditions, then dynamically tunes the logic at runtime, so that bitrate decision can better match the changing networks. Extensive evaluations show that SLVS can control playback latency down to 1s while improving Quality-of-Experience (QoE) by 17.7% to 31.8%. Moreover, it has strong robustness to maintain near-second latency over highly-fluctuating networks as well as long-period of video viewing.
Ke Liu 0004, Mengbai Xiao, Bingshu Wang, Vaneet Aggarwal
ACM Multimedia4
2022 Joint Water-Filling Algorithm with Adaptive Chroma Adjustment for Shadow Removal From Text Document Images
abstract
With smart portable devices such as smartphones and tablets in usage and popularity, people are more willing to use these devices to scan and save digitized documents. However, when capturing document images, shadows are inevitable and influence clarity and readability. How to remove the shadows of document images is an important and meaningful task. In this paper, we propose a water-filling method using chroma adjustment for shadow removal. Firstly, a global and local jointly water-filling approach is designed to estimate the shading map. Then, we design an adaptive global brightness adjustment strategy to optimize the global luminance of the output image. Since only adjusting brightness can cause color distortion of output images, we propose an adaptive chroma adjustment strategy to ensure color consistency across all areas of output images. A series of experiments show that our method can remove shadows of digitized documents, outperforming some state-of the-art methods. Moreover, the proposed method can keep the brightness and color as consistent as possible with the non-shadow area.
Ze Wang 0001, Bingshu Wang, Jiangbin Zheng 0001, C. L. Philip Chen
SMC2
2022 Random-Positioned License Plate Recognition Using Hybrid Broad Learning System and Convolutional Networks
abstract
This paper proposes a framework combing a fully convolutional network with broad learning system for license plate recognition. The fully convolutional network, which is designed as a pixel-level two-class classification method, is proposed for random-positioned object detection by the fusion of multi-scale and hierarchical features. For character segmentation, a trained AdaBoost cascade classifier is employed to locate a key character representing for an administrative area. We design a symmetric region horizontal projection method to estimate the license plate slant angles, and an approach based on vertical projection without hyphens to solve the problem of touching characters. For character recognition, the broad learning system with stacked auto-encoder of mapped feature nodes is proposed, and two structures are explored to recognize letters and digits, respectively. Experiments conducted on Macau license plates show that the proposed method outperforms some state-of-the-art approaches. The compatibility and generality can be expected by applying the proposed method to other regions or countries.
C. L. Philip Chen, Bingshu Wang
IEEE Trans. Intell. Transp. Syst.2
2021 Two-Stage Neural Network Classifier for the Data Imbalance Problem with Application to Hotspot Detection
abstract
The data imbalance problem often occurs in nanometer VLSI applications, where normal cases far outnumber error ones. Many imbalanced data handling methods have been proposed, such as oversampling minority class samples and downsampling majority class samples. However, existing methods focus on improving the quality of minority classes while causing quality deterioration of majority ones. In this paper, we propose a two-stage classifier to handle the data imbalance problem. We first develop an iterative neural network framework to reduce false alarms. Then the oversampling method on a final classification network is applied to predict the two classes better. As a result, the data imbalance problem is well handled, and the quality deterioration of majority classes is also reduced. Since the iterative stage does not change any existing network structure, any convolutional neural network can be used in the framework. Compared with the state-of-the-art imbalanced data handling methods, experimental results on the hotspot detection problem show that our two-stage classification method achieves the best prediction accuracy and reduces false alarms significantly.
Bingshu Wang, Lanfan Jiang, Wenxing Zhu, Longkun Guo, Jianli Chen, Yao-Wen Chang
DAC1
2021 Illumination and Color Adjustment for Generating Professional Headshot
abstract
When making professional headshots, it has some standards, for example, illumination and color of face images should be adjusted well to show important features. In this paper, we propose a method to convert one image captured by smart phone into identification photo or professional headshot. It includes three stages. Firstly, face region is detected and the background is removed from image captured by some mobile device. The second stage is to enhance low illumination regions by estimating illumination map. Thirdly, skin regions are detected through the Cr channel in YCrCb color space. The skin regions with high illumination are adjusted to a specific domain, which aims to provide a comfortable visual perception. These strategies contribute to a pleasing adjustment. Experiments performed on random images indicate our method can obtain a good illumination and color adjustment for face images with a promising adaptability.
Bingshu Wang, C. L. Philip Chen
SMC1
2021 Masked Face Detection Using A Two-stage Classification Approach In the COVID-19 Era
abstract
Masked face detection is a challenging task in the surveillance applications due to complex backgrounds. In this paper, we propose a two-stage method for masked face detection: pre-detection and verification. Firstly, a masked face detector based on AdaBoost algorithm and histogram of orientation feature is exploited. It may provide sufficient candidate face regions. Secondly, a two-class classifier is trained by broad learning system, which is an incremental learning algorithm with high efficiency in training. It is used to distinguish realistic masked faces from background. Moreover, this paper proposes a masked face dataset that includes multiple masked faces captured from real-life scenes . It can be used for classifier training and evaluation. Experiments conducted on the dataset indicate the effectiveness of the proposed method with Recall 94.69% and Precision 97.72%.
Bingshu Wang, Licheng Liu, C. L. Philip Chen
SMC1
2020 Moving Cast Shadows Segmentation Using Illumination Invariant Feature
abstract
This paper presents an effective framework for removing moving cast shadows. Taking the reflection property of object surface for shadow regions under static and fixed scenes, an approximation estimation strategy of bidirectional reflectance distribution function as illumination invariant feature is proposed. It is valid for different types of shadow scenes. In this paper, we propose a new multiple ratios-based technique to justify shadow type for each frame: intensity ratio, area ratio and edge ratio of shadow regions are introduced. According to shadow types, several specified strategies are designed. For weak shadows, multiple features fusion strategy is employed, including color constancy, texture consistency and illumination invariant. For strong shadows, illumination invariant is utilized to detect the umbra and color constancy is utilized to detect the penumbra. Moreover, a suite of shadow direction features is firstly proposed to identify penumbra. The proposed approach is verified in fourteen video sequences varying from weak to strong shadows. The experimental results demonstrate the effectiveness and robustness of the proposed method for both indoor and outdoor scenes compared with some state-of-the-art approaches.
Bingshu Wang, Yong Zhao 0010, C. L. Philip Chen
IEEE Trans. Multim.1
2019 An Effective Background Estimation Method for Shadows Removal of Document Images
abstract
Shadows of document images bring about difficulties for digitization application and uncomfortable perception in vision. This paper proposes an effective method to remove shadows from the single document images, which contains two stages: shadow detection and shadow removal. For the shadow detection, an iterative neighboring information-based approach is designed to estimate pixel-wise local background color image, which can be used to generate a shadow map. For the shadow removal, a global reference background color is obtained from the local background color image. The shadow scale, which is applied to relight shadowed regions, can be estimated by local and global background color. Moreover, a tone fine-tuning process is designed to make the output colors be normal values. Experiments on document images indicate that the proposed method can produce high-quality unshad-owed document images with a comparable efficiency.
Bingshu Wang, C. L. Philip Chen
ICIP1
2018 Depth Super-Resolution Using Joint Adaptive Weighted Least Squares And Patching Gradient
abstract
This paper presents a flexible framework for the challenging task of color-guided depth upsampling. Some state-of-the-art approaches apply an aligned RGB image for depth recovery. Unfortunately, these kinds of methods may result in texture copying artifacts and edge blurring artifacts. To address these difficulties, we propose an adaptive weighted least squares framework of choosing different guidance weight for variant conditions flexibly. First of all, in the framework, we propose a joint adaptive color weighting scheme in which the depth maps and color images jointly choose a proper weight term for diverse cases. Then, a patch-based smoothness measuring approach called patching-gradient method (PGM) is proposed to distinguish the discontinuities and smooth areas. Our PGM is robust to dense noise and preserve weak edges effectively. Quantitative and qualitative experiments on noisy T of -like datasets demonstrate our frameworks effectiveness on suppressing both texture copying artifacts and edge blurring artifacts.
Bingshu Wang, Yong Zhao 0010
ICASSP3
2018 Hard Shadows Removal Using an Approximate Illumination Invariant
abstract
Hard shadows detection and removal from foreground masks is a challenging step in change detection. This paper gives a simple and effective method to address hard shadows. There are inside portion and boundary portion in hard shadows. Pixel-wise neighborhood ratio is calculated to remove the most of inside shadow points. For the boundaries of shadow regions, we take advantage of color constancy to eliminate the edges of hard shadows and obtain relative accurate objects contours. Then, morphology processing is explored to enhance the integrity of objects. The main contribution of this paper is to design an approximate estimation strategy for illumination invariant based on Lambertian reflectance model without prior knowledge. The proposed method is unsupervised and experimental results on six challenging sequences show the effectiveness and robustness of our approach.
Bingshu Wang, C. L. Philip Chen, Yong Zhao 0010
ICASSP1
2016 Background subtraction using dual-class backgrounds
abstract
This paper presents a novel approach to background subtraction which aims to extract moving objects in video stream. To this end, a novel background model is proposed by using both working backgrounds and candidate backgrounds, which can be transferred to each other according to an adaptive mechanism. The input image (video frame) is compared and evaluated with these dual-class backgrounds (DCB) to detect foreground objects. Furthermore, for robust background modeling a novel background updating scheme is proposed based on the life-value which represents the existing time of a background sample, and the access-time which represents the number of valid visits of a background sample. Experiments on a standard dataset demonstrated the effectiveness and robustness of the proposed approach by comparing it with the previous typical background subtraction techniques.
Bingshu Wang, Wenqian Zhu, Yong Zhao 0010, Wenbin Zou
ICARCV1
2015 A Foreground Extraction Method by Using Multi-Resolution Edge Aggregation Algorithm
Wenqian Zhu, Bingshu Wang, Xuefeng Hu, Yong Zhao 0010
ICIG (1)2
2013 A hybrid optimization-based recurrent neural network for real-time data prediction
Liangyu Ma, Bingshu Wang
Neurocomputing3