Binglu Wang

dblp:205/4112 · DBLP profile ↗
← Back
54ranked-venue papers
17as first author
51since 2021 · last 2026
0000-0002-9266-4685ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 26 · 8 first-author · 25 since 2021Graphics, computer vision, multimedia, augmented reality and games · 20 · 10 first-author · 20 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 2 first-author · 10 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2026 Clinically Consistent Ordinal Learning for Early Pneumoconiosis Diagnosis
Binglu Wang, Jingyi Cui
KSEM (2)1
2026 Unlocking human intent perception through multimodal large models
Zhaozhong Wang, Shutong Wang, Binglu Wang
Pattern Recognit.5
2026 VL-HTR: Learning Human-Target Representation From Vision-Language Model
abstract
Human-gaze-target prediction aims to predict the target point or object that humans are looking at in images. However, existing methods predominantly rely on vision-only features, which often struggle to capture the semantic context of small or occluded objects and lack explicit priors for precise head direction regression, leading to slow convergence and suboptimal performance. Therefore, we introduce VL-HTR, a novel vision-language learning method for human-target representation, which integrates multimodal knowledge from vision-language models (VLMs) to construct robust human-target relationships. Unlike traditional approaches, extracting multimodal features via pretrained VLMs enhances the model's grasp of human-target knowledge through the learnable target class and direction context. Then, a language-guided query alignment (LQA) module is introduced to improve the semantic-aware object representation capability through vision-language query alignment. Finally, to accelerate the gaze point regression learning process, we design a language-guided direction prediction (LDP) module to introduce multimodal human gaze direction priors, thereby facilitating the human-target relationship construction. Extensive validations across two distinct tasks, i.e., gaze object prediction (GOP) and gaze target estimation, involving five challenging benchmarks, demonstrating that VL-HTR achieves superior performance and much faster training convergence.
Binglu Wang, Jingyi Cui, Haisheng Xia, Guangyu Guo 0001, Zhijun Li 0001
IEEE Trans. Cybern.1
2026 TransGOP-R: Transformer-Based Real-World Gaze Object Prediction
abstract
The goal of gaze object prediction (GOP) is to predict human gaze objects and categories. However, existing methods require additional head priors or filter the results before evaluation, which is an obstacle for real-world applications. To this end, this paper proposes aTransformer-basedGazeObjectPrediction underReal-world setting (TransGOP-R), which does not rely on any head prior input and evaluates end-to-end. We first design a head location module to generate human head location information from a head query. Then, an error analysis demonstrates that the primary error source of the existing GOP model is in gaze estimation, which is caused by the difficulty in predicting gaze points by directly regressing heatmaps. Therefore, we introduce cone prediction into the model training stage, allowing the middle-layer features of the gaze regressor to build the relationship between the target human and objects before regressing the gaze point. An oriented gradient mechanism is proposed in this process to ensure the object detection performance is not affected by cone information. Finally, we conducted very detailed and sufficient experiments to verify the superiority of our method on the GOO-Synth and GOO-Real datasets. At the same time, we also achieve advantages compared to the human-target gaze estimation methods on the GazeFollowing, VideoAttentionTarget, and ChildPlay datasets.
Guangyu Guo 0001, Zhaozhong Wang, Binglu Wang
IEEE Trans. Multim.4
2025 FinePhys: Fine-grained Human Action Generation by Explicitly Incorporating Physical Laws for Effective Skeletal Guidance
abstract
Despite significant advances in video generation, synthesizing physically plausible human actions remains a persistent challenge, particularly in modeling fine-grained semantics and complex temporal dynamics. For instance, generating gymnastics routines such as "switch leap with 0.5 turn" poses substantial difficulties for current methods, often yielding unsatisfactory results. To bridge this gap, we propose FinePhys1, a Fine-grained human action generation framework that incorporates Physics to obtain effective skeletal guidance. Specifically, FinePhys first estimates 2D poses in an online manner and then performs 2D-to-3D dimension lifting via in-context learning. To mitigate the instability and limited interpretability of purely data-driven 3D poses, we further introduce a physics-based motion re-estimation module governed by Euler-Lagrange equations, calculating joint accelerations via bidirectional temporal updating. The physically predicted 3D poses are then fused with data-driven ones, offering multi-scale 2D heatmap guidance for the diffusion process. Evaluated on three fine-grained action subsets from FineGym (FX-JUMP, FX-TURN, and FX-SALTO), FinePhys significantly outperforms competitive baselines. Comprehensive qualitative results further demonstrate FinePhys’s ability to generate more natural and plausible fine-grained human actions.
Dian Shao, Mingfei Shi, Shengda Xu, Yongle Huang, Binglu Wang
CVPR6
2025 Contextual and orientation correction modules enhance weakly-supervised aerial object detection in remote sensing images
Le Yang 0008, Shunzhou Wang, Xuerong Wang, Shutong Wang, Binglu Wang
Eng. Appl. Artif. Intell.6
2025 Radiologist-inspired Symmetric Local-Global Multi-Supervised Learning for early diagnosis of pneumoconiosis
Meiyue Song, Deng-Ping Fan, Shaoting Zhang 0001, Juntao Yang, Jiangfeng Liu, Binglu Wang
Expert Syst. Appl.9
2025 Multi-Sensory Visual-Auditory Fusion of Wearable Navigation Assistance for People With Impaired Vision
abstract
Visually impaired individuals face limited mobility and restricted independent navigation in complex environments. Improving mobility and independent navigation for people with visual impairments is essential. In this paper, we introduce wearable electronic glasses (E-Glasses) that utilize a target detection network to fuse visual and auditory information for searching desired targets. Integrated with the electronic glasses, we present a neural path planning that combines spiking neural and convolutional neural networks. Several participants took part in experiments that showcased the remarkable capabilities of the developed system in target detection and navigation. The experimental results revealed an impressive success rate of 95.46% for the target detection network, providing participants with more accurate target information. Additionally, the neural path planning network achieved a success rate of 92.60%, demonstrating a significant speed advantage compared to the enhanced$A^{\ast}$algorithm.Note to Practitioners—This article aims to address the mobility and navigation challenges faced by visually impaired individuals in complex environments. Many solutions have been proposed to assist with wearable navigation by utilizing sensors to perceive the environment and provide path information to users. In this article, we introduce wearable electronic glasses (E-Glasses) with advanced target detection and path planning capabilities to help visually impaired individuals navigate more effectively in complex environments. We have developed a target detection network that integrates visual and auditory information to effectively search for desired targets. Additionally, we have developed a neural path planning algorithm that combines spiking neural networks and convolutional neural networks. Furthermore, experiments conducted in indoor navigation challenges demonstrate the feasibility of this approach. In future research, we will focus on further exploring the target detection network and refining the neural path planning algorithms to improve the overall performance of the system.
Zhijun Li 0001, Guoxin Li 0001, Binglu Wang, Peng Shi 0001
IEEE Trans Autom. Sci. Eng.4
2025 A Patch-Based Method for Underwater Image Enhancement With Denoising Diffusion Models
abstract
The enhancement of underwater images has emerged as a significant technological challenge in advancing marine research and exploration tasks. Due to the scattering of suspended particles and absorption of light in underwater environments, underwater images tend to present blurriness and predominantly color distortion. In this study, we propose a novel approach utilizing denoising diffusion models to improve underwater degraded images. After training the noise estimation network of the denoising diffusion models, we accelerate the deterministic sampling process with denoising diffusion implicit models. We also propose a patch-based method by implementing average sampling between overlapping image patches at each sampling step, enabling the generation of images at arbitrary resolution while preserving their natural appearance and details. Through benchmark experiments, we illustrate that our method outperforms or closely approaches state-of-the-art techniques in terms of effectiveness and performance. We demonstrate that our approach reduces the interference of underwater environments with the semantic information of the images by salient object detection experiments.
Haisheng Xia, Binglei Bao, Binglu Wang, Zhijun Li 0001
IEEE Trans. Cybern.5
2025 Perspective on Wearable Systems for Human Underwater Perceptual Enhancement
abstract
Underwater areas have harsh environments with poor light, limited visibility, and high levels of noise. Humans have a weak perception of position, surroundings, and exterior information when staying underwater, which makes it difficult for humans to carry out complex underwater tasks, such as rescue, observation, and construction. Wearable devices have shown good results in enhancing human sensory function on land, thus they could potentially play a role in enhancing human underwater perception ability. This perspective aims to analyze the state-of-the-art of underwater wearable systems for human perception enhancement. This work discusses the core technology and challenges of human underwater perceptual enhancement, including wearable underwater navigation, underwater environment reconstruction, and underwater sensorial information delivery. Future research could focus on designing waterproof flexible human-machine interfaces for sensing and feedback, exploiting advanced sensors and fusion algorithms for wearable underwater positioning, and studying multimodal information interaction strategies of wearable systems.
Haisheng Xia, Binglei Bao, Binglu Wang, Qinghua Huang, Zhijun Li 0001
IEEE Trans. Cybern.5
2025 Collaborative Multimodal Fusion Network for Multiagent Perception
abstract
With the increasing popularity of autonomous driving systems and their applications in complex transportation scenarios, collaborative perception among multiple intelligent agents has become an important research direction. Existing single-agent multimodal fusion approaches are limited by their inability to leverage additional sensory data from nearby agents. In this article, we present the collaborative multimodal fusion network (CMMFNet) for distributed perception in multiagent systems. CMMFNet first extracts modality-specific features from LiDAR point clouds and camera images for each agent using dual-stream neural networks. To overcome the ambiguity in-depth prediction, we introduce a collaborative depth supervision module that projects dense fused point clouds onto image planes to generate more accurate depth ground truths. We then present modality-aware fusion strategies to aggregate homogeneous features across agents while preserving their distinctive properties. To align heterogeneous LiDAR and camera features, we introduce a modality consistency learning method. Finally, a transformer-based fusion module dynamically captures cross-modal correlations to produce a unified representation. Comprehensive evaluations on two extensive multiagent perception datasets, OPV2V and V2XSet, affirm the superiority of CMMFNet in detection performance, establishing a new benchmark in the field.
Lei Zhang 0166, Binglu Wang, Yongqiang Zhao 0001, Yuan Yuan 0006, Tianfei Zhou, Zhijun Li 0001
IEEE Trans. Cybern.2
2025 Belief-Based Fuzzy and Imprecise Clustering for Arbitrary Data Distributions
abstract
Fuzzy clustering is still a hot topic because it can calculate the support degrees of an object belonging to different clusters to characterize uncertainty. However, it remains a challenge to detect clusters of arbitrary shapes, sizes, and dimensionality. What's worse, some objects are indistinguishable (imprecise) when they are in the overlapping regions of different clusters. To address such issues, this paper investigates a belief-based fuzzy and imprecise clustering (BFI) method, which can detect arbitrary clusters and provide the behavior (support) of objects to these clusters. Moreover, BFI can assign each imprecise object to a meta-cluster, defined as the union of specific clusters, to characterize (partial) imprecision. The proposed BFI can significantly reduce the risk of misclassification, and the effectiveness is validated in image processing (e.g., image segmentation and classification) and several benchmark datasets by comparing it with some typical methods.
Zuowei Zhang 0001, Zhunga Liu, Liang-Bo Ning 0001, Hongpeng Tian, Binglu Wang
IEEE Trans. Fuzzy Syst.5
2025 Content and Relation Fuzzy Mitigation Framework for Intent Perception
abstract
In this article, we tackle the content fuzzy and relation fuzzy in image-based intent perception. Current research primarily focuses on additional modeling mechanisms and multimodal information, but it is often difficult to deal with the relation fuzzy under inaccurate annotation in intent perception. We exploit the powerful representation capabilities of large language models and innovatively use them to solve the image-based intent perception task. Comprising two innovative modules: multidimension adapter and fuzzy harmony refiner. Multidimension adapter projects intent image features of different difficulty levels to feature dimensions of different sizes to alleviate content fuzzy of different difficulties caused by image diversity. Besides, intent category labeling is somewhat subjective, and some categories are more likely to appear collaboratively. We design fuzzy harmony refiner to mine the relationship between various types to enhance perception results. Different from traditional expert-specified rules, we learn the co-occurrence frequencies between different intent categories from the training data, generate a fuzzy harmony matrix for correcting the score, and alleviate the relationship fuzzy in intent perception. We achieved new state-of-the-art performance on the Intentonomy dataset, with macro F1 score of 42.52%, micro F1 score of 54.80%, samples F1 score of 57.57%, and average F1 score of 51.63%. Compared with the current existing methods, our method improves by 9.8%.
Wenxin Zhang 0003, Shutong Wang, Binglu Wang
IEEE Trans. Fuzzy Syst.5
2025 CM-YOLO: Context Modulated Representation Learning for Ship Detection
abstract
Ship detection is essential for both military and civilian applications. Existing ship detection methods focus on prominent offshore ships, paying less attention to complex nearshore ships, which are easily confused with the intricate background. Utilizing contextual information, such as location and shape, can enhance ship detection and classification in complex environments. In this article, we propose a context modulated representation learning-based detection method termed as CM-YOLO. It adopts the classical detector design framework, which includes the backbone, neck, and head. The input image is sequentially processed through these components to obtain the detection results. Our method specifically optimizes ship detection in complex scenarios. To achieve this, we propose a dual path context enhancement neck (DCEN) to extract contextual information for ship detection. The neck builds on the path augmentation feature pyramid network with the proposed dual path context enhancement (DCE) module, which is designed to enhance feature representations by incorporating high-level semantic information. It captures long-range dependencies across both channel and spatial dimensions while suppressing irrelevant features. Additionally, to enhance the scale-aware capability of the head for detecting multiscale ships in complex environments, we introduce the multicontext boosted (MCB) detection head. The MCB can flexibly adjust the receptive field and extracts relevant context for ships of various scales using multiple large-kernel convolutions. We conduct experiments on three commonly used ship datasets: Seaships7000, ShipRSImageNet, DIOR-ship, and HRSC2016. Experiment results demonstrate that CM-YOLO achieves excellent performance compared with other leading ship detection methods.
Lingtong Min, Feiyang Dou, Dian Shao, Binglu Wang
IEEE Trans. Geosci. Remote. Sens.6
2025 Adaptive Fusion Learning for Compositional Zero-Shot Recognition
abstract
Compositional Zero-Shot Learning (CZSL) aims to learn visual concepts (i.e., attributes and objects) from seen compositions and combine them to predict unseen compositions. Existing visual encoders in CZSL typically use traditional visual encoders (i.e., CNN and Transformer) or image encoders from Visual-Language Models (VLMs) to encode image features. However, traditional visual encoders need more multi-modal textual information, and image encoders of VLMs exhibit dependence on pre-training data, making them less effective when used independently for predicting unseen compositions. To overcome this limitation, we propose a novel approach based on the joint modeling of traditional visual encoders and VLMs visual encoders to enhance the prediction ability for uncommon and unseen compositions. Specifically, we design an adaptive fusion module that automatically adjusts the weighted parameters of similarity scores between traditional and VLMs methods during training, and these weighted parameters are inherited during the inference process. Given the significance of disentangling attributes and objects, we design a Multi-Attribute Object Module that, during the training phase, incorporates multiple pairs of attributes and objects as prior knowledge, leveraging this rich prior knowledge to facilitate the disentanglement of attributes and objects. Building upon this, we select the text encoder from VLMs to construct the Adaptive Fusion Network. We conduct extensive experiments on the Clothing16 K, UT-Zappos50 K, and C-GQA datasets, achieving excellent performance on the Clothing16 K and UT-Zappos50 K datasets.
Lingtong Min, Ziman Fan, Shunzhou Wang, Feiyang Dou, Xin Li 0042, Binglu Wang
IEEE Trans. Multim.6
2025 Multimodal Large Models are Effective Action Anticipators
abstract
The task of long-term action anticipation demands solutions that can effectively model temporal dynamics over extended periods while deeply understanding the inherent semantics of actions. Traditional approaches, which primarily rely on recurrent units or Transformer layers to capture long-term dependencies, often fall short in addressing these challenges. Large Language Models (LLMs), with their robust sequential modeling capabilities and extensive commonsense knowledge, present new opportunities for long-term action anticipation. In this work, we introduce the ActionLLM framework, a novel approach that treats video sequences as successive tokens, leveraging LLMs to anticipate future actions. Our baseline model simplifies the LLM architecture by setting future tokens, incorporating an action tuning module, and reducing the textual decoder layer to a linear layer, enabling straightforward action prediction without the need for complex instructions or redundant descriptions. To further harness the commonsense reasoning of LLMs, we predict action categories for observed frames and use sequential textual clues to guide semantic understanding. In addition, we introduce a Cross-Modality Interaction Block, designed to explore the specificity within each modality and capture interactions between vision and textual modalities, thereby enhancing multimodal tuning. Extensive experiments on benchmark datasets demonstrate the superiority of the proposed ActionLLM framework, encouraging a promising direction to explore LLMs in the context of action anticipation.
Binglu Wang, Shunzhou Wang, Le Yang 0008
IEEE Trans. Multim.1
2025 Action-to-Action Diffusion Network for Weakly Supervised Temporal Action Localization
abstract
Weakly supervised temporal action localization (WTAL) aims to identify action instances in untrimmed videos with only video-level supervision. Despite recent advances in WTAL methods, achieving accurate boundary localization remains a significant challenge. A key reason is that WTAL networks following a localization-by-classification pipeline tend to focus on the most discriminative features, neglecting some ambiguous features that may contain action instances. To make the WTAL model focus on low-discriminative features that include action instances, we propose an action-to-action diffusion (ActionDiff) network. This network leverages the smoothness of data generated by the diffusion model, using the diffusion model to output smooth and high-quality features that weaken the discriminative action features from the base branch, thereby enhancing the performance of the WTAL task. First, we develop a topk-based masking strategy to generate binary masks that serve as pseudo-labels for diffusion model learning. Then, we propose a diffusion branch to generate high-quality latent action space by iteratively removing noise guided by the designed pseudo-labels and conditional information. To enhance the diffusion branch's capability to generate human behavioral features, we design an action-related conditional strategy to obtain conditional information and use it to guide the modeling of human behavior knowledge by the diffusion branch. Our comprehensive experiments demonstrate that the proposed method achieves a promising performance on three benchmark datasets: THUMOS14, ActivityNet v1.2, and v1.3.
Yuanbing Zou, Qingjie Zhao, Prodip Kumar Sarker, Le Yang 0008, Binglu Wang
IEEE Trans. Multim.5
2025 Federated Cross-Incremental Self-Supervised Learning for Medical Image Segmentation
abstract
Federated cross learning has shown impressive performance in medical image segmentation. However, it encounters the catastrophic forgetting issue caused by data heterogeneity across different clients and is particularly pronounced when simultaneously facing pixelwise label deficiency problem. In this article, we propose a novel federated cross-incremental self-supervised learning method, coined FedCSL, which not only can enable any client in the federation incrementally yet effectively learn from others without inducing knowledge forgetting or requiring massive labeled samples, but also preserve maximum data privacy. Specifically, to overcome the catastrophic forgetting issue, a novel cross-incremental collaborative distillation (CCD) mechanism is proposed, which distills explicit knowledge learned from previous clients to subsequent clients based on secure multiparty computation (MPC). Besides, an effective retrospect mechanism is designed to rearrange the training sequence of clients per round, further releasing the power of CCD by enforcing interclient knowledge propagation. In addition, to alleviate the need of large-scale densely annotated pretraining medical datasets, we also propose a two-stage training framework, in which federated cross-incremental self-supervised pretraining paradigm first extracts robust yet general image-level patterns across multi-institutional data silos via a novel round-robin distributed masked image modeling (MIM) pipeline; then, the resulting visual concepts, e.g., semantics, are transferred to the federated cross-incremental supervised fine-tuning paradigm, favoring various cross-silo medical image segmentation tasks. The experimental results on public datasets demonstrate the effectiveness of the proposed method as well as the consistently superior performance of our method over most state-of-the-art methods quantitatively and qualitatively.
Fan Zhang 0070, Chun-Mei Feng 0001, Binglu Wang, Shanshan Wang 0002, Junyu Dong, David Zhang 0001
IEEE Trans. Neural Networks Learn. Syst.5
2024 TransGOP: Transformer-Based Gaze Object Prediction
abstract
Gaze object prediction aims to predict the location and category of the object that is watched by a human. Previous gaze object prediction works use CNN-based object detectors to predict the object's location. However, we find that Transformer-based object detectors can predict more accurate object location for dense objects in retail scenarios. Moreover, the long-distance modeling capability of the Transformer can help to build relationships between the human head and the gaze object, which is important for the GOP task. To this end, this paper introduces Transformer into the fields of gaze object prediction and proposes an end-to-end Transformer-based gaze object prediction method named TransGOP. Specifically, TransGOP uses an off-the-shelf Transformer-based object detector to detect the location of objects and designs a Transformer-based gaze autoencoder in the gaze regressor to establish long-distance gaze relationships. Moreover, to improve gaze heatmap regression, we propose an object-to-gaze cross-attention mechanism to let the queries of the gaze autoencoder learn the global-memory position knowledge from the object detector. Finally, to make the whole framework end-to-end trained, we propose a Gaze Box loss to jointly optimize the object detector and gaze regressor by enhancing the gaze heatmap energy in the box of the gaze object. Extensive experiments on the GOO-Synth and GOO-Real datasets demonstrate that our TransGOP achieves state-of-the-art performance on all tracks, i.e., object detection, gaze estimation, and gaze object prediction. Our code will be available at https://github.com/chenxi-Guo/TransGOP.git.
Binglu Wang, Haisheng Xia, Nian Liu 0002
AAAI1
2024 Boosting Gaze Object Prediction via Pixel-Level Supervision from Vision Foundation Model
Lei Zhang 0166, Shi Yan 0005, Bin Fan 0002, Binglu Wang
ECCV (69)5
2024 Vision-based Wearable Steering Assistance for People with Impaired Vision in Jogging
abstract
Outdoor sports pose a challenge for people with impaired vision. The demand for higher-speed mobility inspired us to develop a vision-based wearable steering assistance. To ensure broad applicability, we focused on a representative sports environment, the athletics track. Our efforts centered on improving the speed and accuracy of perception, enhancing planning adaptability for the real world, and providing swift and safe assistance for people with impaired vision. In perception, we engineered a lightweight multitask network capable of simultaneously detecting track lines and obstacles. Additionally, due to the limitations of existing datasets for supporting multi-task detection in athletics tracks, we diligently collected and annotated a new dataset (MAT) containing 1000 images. In planning, we integrated the methods of sampling and spline curves, addressing the planning challenges of curves. Meanwhile, we utilized the positions of the track lines and obstacles as constraints to guide people with impaired vision safely along the current track. Our system is deployed on an embedded device, Jetson Orin NX. Through outdoor experiments, it demonstrated adaptability in different sports scenarios, assisting users in achieving free movement of 400meter at an average speed of 1.34 m/s, meeting the level of normal people in jogging. Our MAT dataset is publicly available from https://github.com/snoopy-l/MAT
Binglu Wang, Zhijun Li 0001
ICRA2
2024 Chareption: Change-Aware Adaption Empowers Large Language Model for Effective Remote Sensing Image Change Captioning
abstract
Remote Sensing Image Change Captioning (RSICC) faces significant challenges in effectively identifying and articulating changes between bi-temporal images. Traditional approaches often utilize individual text decoders, which may not capture the subtleties of visual changes nor fully exploit advanced language modeling capabilities. To overcome these limitations, we propose Cha nge-Awa re Ada ption , namely Chareption , a novel framework that effectively leverages pre-trained large language models (LLMs) to enhance both the accuracy and detail of change captions. Central to Chareption is a change-aware module designed to selectively identify and utilize tokens that significantly represent changes, thus avoiding the common issue of redundancy that plagues methods relying solely on class tokens or indiscriminate use of all patch tokens. Additionally, Chareption designs a lightweight change adapter module, seamlessly integrated into both the vision backbone and the LLM, requiring minimal learnable parameters while optimally adjusting representations for the RSICC task. Our experiments on the LEVIR-CC dataset demonstrate that Chareption significantly outperforms existing methods in caption accuracy and contextual relevance, while also reducing training overhead. This establishes Chareption as a pioneering solution that sets a new direction in RSICC by harnessing the rich representational power of LLMs for improved multimodal understanding.
Changhe Wang, Ningyu He, Binglu Wang
PRCV (13)3
2024 LLMAction: Adapting Large Language Model for Long-Term Action Anticipation
Binglu Wang, Changhe Wang
PRCV (10)1
2024 DSLSM: Dual-kernel-induced statistic level set model for image segmentation
Fan Zhang 0070, Xiaojun Duan, Binglu Wang, Huafeng Li 0001, Junyu Dong, David Zhang 0001
Expert Syst. Appl.4
2024 PneumoLLM: Harnessing the power of large language model for pneumoconiosis diagnosis
Meiyue Song, Zhihua Yu, Baicun Li, Qinghua Huang, Zhijun Li 0001, Nikolaos I. Kanellakis, Jiangfeng Liu, Binglu Wang, Juntao Yang
Medical Image Anal.15
2024 Temporal Action Localization in the Deep Learning Era: A Survey
abstract
The temporal action localization research aims to discover action instances from untrimmed videos, representing a fundamental step in the field of intelligent video understanding. With the advent of deep learning, backbone networks have been instrumental in providing representative spatiotemporal features, while the end-to-end learning paradigm has enabled the development of high-quality models through data-driven training. Both supervised and weakly supervised learning approaches have contributed to the rapid progress of temporal action localization, resulting in a multitude of methods and a large body of literature, making a comprehensive survey a pressing necessity. This paper presents a thorough analysis of existing action localization works, offering a well-organized taxonomy that highlights the strengths and weaknesses of each strategy. In the realm of supervised learning, in addition to the anchor mechanism, we introduce a novel classification mechanism to categorize and summarize existing works. Similarly, for weakly supervised learning, we extend the traditional pre-classification and post-classification mechanisms by providing a fresh perspective on enhancement strategies. Furthermore, we shed light on the bottleneck of confidence estimation, a critical yet overlooked aspect of current works. By conducting detailed analyses, this survey serves as a valuable resource for researchers, providing beneficial guidance to newcomers and inspiring seasoned researchers alike.
Binglu Wang, Yongqiang Zhao 0001, Le Yang 0008, Teng Long 0001, Xuelong Li 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2024 Federated Feature Augmentation and Alignment
abstract
Federated learning is a distributed paradigm that allows multiple parties to collaboratively train deep learning models without direct exchange of raw data. Nevertheless, the inherent non-independent and identically distributed (non-i.i.d.) nature of data distribution among clients results in significant degradation of the acquired model. The primary goal of this study is to develop a robust federated learning algorithm to addressfeature shiftin clients’ samples, potentially arising from a range of factors such as acquisition discrepancies in medical imaging. To reach this goal, we first propose federated feature augmentation (FedFA$^{l}$), a novel feature augmentation technique tailored for federated learning.FedFA$^{l}$is based on a crucial insight that each client's data distribution can be characterized by first-/second-order statistics (a.k.a., mean and standard deviation) of latent features; and it is feasible to manipulate these local statisticsglobally, i.e., based on information in the entire federation, to let clients have a better sense of the global distribution across clients. Grounded on this insight, we propose to augment each local feature statistic based on a normal distribution, wherein the mean corresponds to the original statistic, and the variance defines the augmentation scope. Central toFedFA$^{l}$is the determination of a meaningful Gaussian variance, which is accomplished by taking into account not only biased data of each individual client, but also underlying feature statistics represented by all participating clients. Beyond consideration oflow-orderstatistics inFedFA$^{l}$, we propose a federated feature alignment component (FedFA$^{h}$) that exploitshigher-orderfeature statistics to gain a more detailed understanding of local feature distribution and enables explicit alignment of augmented features in different clients to promote more consistent feature learning. CombiningFedFA$^{l}$andFedFA$^{h}$yields our full approachFedFA$+$+.FedFA$+$+is non-parametric, incurs negligible additional communication costs, and can be seamlessly incorporated into popular CNN and Transformer architectures. We offer rigorous theoretical analysis, as well as extensive empirical justifications to demonstrate the effectiveness of the algorithm.
Tianfei Zhou, Ye Yuan 0001, Binglu Wang, Ender Konukoglu
IEEE Trans. Pattern Anal. Mach. Intell.3
2024 Differential Feature Awareness Network Within Antagonistic Learning for Infrared-Visible Object Detection
abstract
The combination of infrared and visible videos aims to gather more comprehensive feature information from multiple sources and reach superior results on various practical tasks, such as detection and segmentation, over that of a single modality. However, most existing dual-modality object detection algorithms ignore the modal differences and fail to consider the correlation between feature extraction and fusion, which leads to incomplete extraction and inadequate fusion of dual-modality features. Hence, there raises an issue of how to preserve each unique modal feature and fully utilize the complementary infrared and visible information. Facing the above challenges, we propose a novel Differential Feature Awareness Network (DFANet) within antagonistic learning for infrared and visible object detection. The proposed model consists of an Antagonistic Feature Extraction with Divergence (AFED) module used to extract the differential infrared and visible features with unique information, and an Attention-based Differential Feature Fusion (ADFF) module used to fully fuse the extracted differential features. We conduct performance comparisons with existing state-of-the-art models on two benchmark datasets to represent the robustness and superiority of DFANet, and numerous ablation experiments to illustrate its effectiveness.
Ruiheng Zhang 0001, Qi Zhang 0004, Jin Zhang 0021, Lixin Xu 0001, Baomin Zhang, Binglu Wang
IEEE Trans. Circuits Syst. Video Technol.7
2024 U²PNet: An Unsupervised Underwater Image-Restoration Network Using Polarization
abstract
This article presents U 2PNet, a novel unsupervised underwater image restoration network using polarization for improving signal-to-noise ratio and image quality in underwater imaging environments. Traditional methods for underwater image restoration using polarization require specific cues or pairs of underwater polarization datasets, which limit their practical applications. Our proposed method requires only one mosaicked polarized image of the scene and does not require datasets for pretraining or specific cues. We design two subnetworks (T-net and B textsubscript ∞ -net) to accurately estimate the transmission map and background light, and unique nonreference loss functions to ensure effective restoration. Our experiments are based on an indoor polarization simulated dataset and a real polarization image dataset constructed from our underwater robotic platform equipped with polarization cameras. Experiment results demonstrate that our proposed method achieves state-of-the-art performance on both simulated and real underwater polarization images. The code and datasets will be available at https://github.com/polwork/U-2Pnet.
Linghao Shen, Haisheng Xia, Yongqiang Zhao 0001, Ning Li 0038, Seong G. Kong, Binglu Wang, Zhijun Li 0001
IEEE Trans. Cybern.7
2024 Scale-Aware Backprojection Transformer for Single Remote Sensing Image Super-Resolution
abstract
Backprojection networks have achieved promising super-resolution performance for nature images but not well be explored in the remote sensing image super-resolution (RSISR) field due to the high computation costs. In this article, we propose a scale-aware backprojection Transformer termed SPT for RSISR. SPT incorporates the backprojection learning strategy into a Transformer framework. It consists of scale-aware backprojection-based self-attention layers (SPALs) for scale-aware low-resolution feature learning and scale-aware backprojection-based Transformer blocks (SPTBs) for hierarchical feature learning. A backprojection-based reconstruction module (PRM) is also introduced to enhance the hierarchical features for image reconstruction. SPT stands out by efficiently learning low-resolution features without excessive modules for high-resolution processing, resulting in lower computational resources. Experimental results on UCMerced and AID datasets demonstrate that SPT obtains state-of-the-art results compared to other leading RSISR methods.
Jinglei Hao, Wukai Li, Yongqiang Zhao 0001, Shunzhou Wang, Binglu Wang
IEEE Trans. Geosci. Remote. Sens.7
2024 Two-Stage Spatial-Frequency Joint Learning for Large-Factor Remote Sensing Image Super-Resolution
abstract
Super-resolution neural networks have recently achieved great progress in restoring high-quality remote sensing images at low zoom-in magnitude. However, these networks often struggle with challenges like shape distortion and blurring effects due to the severe absence of structure and texture details in large-factor remote sensing image super-resolution. Addressing these challenges, we propose a novel Two-Stage Spatial-Frequency Joint Learning Network (TSFNet). TSFNet innovatively merges insights from both spatial and frequency domains, enabling a progressive refinement of super-resolution results from coarse to fine. Specifically, different from existing frequency feature extraction approaches, we design a novel amplitude-guided-phase adaptive filter module to explicitly disentangle and sequentially recover both the global common image degradation and specific structural degradation in the frequency domain. Additionally, we introduce the cross-stage feature fusion design to enhance feature representation and selectively propagate useful information from stage one to stage two. Quantitative and qualitative experimental results demonstrate that our proposed method surpasses state-of-the-art techniques in large-factor remote sensing image super-resolution. Our code is available at https://github.com/likakakaka/TSFNet_RSISR.
Shunzhou Wang, Binglu Wang, Teng Long 0001
IEEE Trans. Geosci. Remote. Sens.4
2024 BEVRefiner: Improving 3D Object Detection in Bird's-Eye-View via Dual Refinement
abstract
Many multi-view camera-based 3D object detection models transform the image features into Bird’s-Eye-View (BEV) via the Lift-Splat-Shoot (LSS) mechanism, which “lifts” 2D camera-view features to the 3D voxel space based on the predicted depth distribution and then “splats” 3D features into a BEV plane for subsequent 3D object detection. However, the BEV feature in such a one-stage view transformation scheme heavily relies on the quality of the predicted depth distribution and 2D camera-view features, which further determines the final detection performance. In this paper, we propose a BEVRefiner model which performs dual refinement for both depth prediction and 2D camera-view features. On the one hand, we perform light-weight depth refinement in the depth distribution frustum space by incorporating 3D context and depth distribution prior. On the other hand, we reproject the BEV feature back to each camera view to enhance 2D image features. In this way, the original camera-view features can be enhanced by implicitly incorporating 3D contexts and multi-view contexts, which cannot be achieved in the original 2D camera view. We also propose to use dominant depth bins only for the reprojection to save computational burden. Finally, we generate the refined BEV feature using the refined depth distribution and camera-view features for more accurate 3D object detection. Our BEVRefiner can be plugged into LSS-based BEV detectors and we perform extensive experiments on the representative model BEVDet, which strongly verified the efficiency of our proposed approach under several settings.
Binglu Wang, Lei Zhang 0166, Nian Liu 0002, Rao Muhammad Anwer, Hisham Cholakkal, Yongqiang Zhao 0001, Zhijun Li 0001
IEEE Trans. Intell. Transp. Syst.1
2024 Coarse-to-Fine Nutrition Prediction
abstract
Healthy dietary intake has a broad influence on the quality of life, and nutrition prediction plays a great role in the auxiliary decision-making of diet. Given a food image, existing nutrition prediction methods directly regress the nutrition content. However, due to the complex variations in food images, such as differences in viewpoint and lighting conditions, directly regressing the nutrition content faces significant challenges. The complexity of the food image data results in a high-dimensional and feature-rich input space, which poses difficulties for traditional regression models to efficiently navigate and optimize. Consequently, the direct regression paradigm usually generates inaccurate nutrition predictions. To alleviate the ambiguity challenge in the prediction progress, we propose to narrow the searchable space for the model's predictions by decomposing the direct regression into two steps: first coarsely selecting the nutrition scope and then finely refining the prediction value, forming a coarse-to-fine nutrition prediction paradigm. Although the process of coarse prediction which selects a bin from a series of scope bins can be formulated as a standard classification problem, it exhibits a distinguishable characteristic, i.e. the closer to the ground truth bin, the less punishment in the training phase. However, most of the current methods have ignored this phenomenon, thus, we specially design the linearly smoothed label in the nutrition prediction task to reveal the relative distance to the ground truth bin, leading to extraordinary improvements. Furthermore, we conduct a pair-wise comparison among all bins by extending the 1D label into 2D space and propose the structure loss to guide the bin selection process effectively. Due to the narrowed decision space, the nutrition prediction problem can be effectively optimized, and the proposed method achieves promising results on three benchmarks ECUSTFD, VFD and Nutrition5K, demonstrating the efficiency of the coarse-to-fine paradigm equipped with the linear-smoothed structure loss.
Binglu Wang, Tianci Bu, Zaiyi Hu, Le Yang 0009, Yongqiang Zhao 0001, Xuelong Li 0001
IEEE Trans. Multim.1
2024 LCH: fast RGB-D salient object detection on CPU via lightweight convolutional network with hybrid knowledge distillation
Binglu Wang, Fan Zhang 0070, Yongqiang Zhao 0001
Vis. Comput.1
2023 Densitytoken: Weakly-Supervised Crowd Counting with Density Classification
abstract
Weakly-supervised crowd counting methods using only the count-level label have achieved great progress recently. However, the count-level label can not provide information about the distribution of the crowd in the scene, which will affect the accuracy of the final count result. Therefore, we design a density token to perceive the crowd distribution in the scenes. Based on this, we propose a Dual Supervision Transformer (DSFormer) to perform weakly-supervised crowd counting in the double supervision of the total count. Specifically, the encoded features in vision transformer are sent to the proposed locality enhanced module (LEM), and one branch of them with density tokens are sent into decoder for interaction. The crowd distribution perception is then realized through cross self-attention, where the other branch of encoded features are used as queries. Finally, the output features of the decoder are respectively fed into a count regression head and a crowd density classification head to obtain the crowd count and the crowd density classification. A series of experiments are conducted on three commonly used crowd counting datasets, both the quantitative and vi-sualization results illustrate the effectiveness of DSFormer. Our model achieve a superior result among weakly supervised crowd counting methods, and code is available at: https://github.com/ZaiyiHu/DSFormer.
Zaiyi Hu, Binglu Wang, Xuelong Li 0001
ICASSP2
2023 Core: Cooperative Reconstruction for Multi-Agent Perception
abstract
This paper presents Core, a conceptually simple, effective and communication-efficient model for multi-agent cooperative perception. It addresses the task from a novel perspective of cooperative reconstruction, based on two key insights: 1) cooperating agents together provide a more holistic observation of the environment, and 2) the holistic observation can serve as valuable supervision to explicitly guide the model learning how to reconstruct the ideal observation based on collaboration. Core instantiates the idea with three major components: a compressor for each agent to create more compact feature representation for efficient broadcasting, a lightweight attentive collaboration component for cross-agent message aggregation, and a reconstruction module to reconstruct the observation based on aggregated feature representations. This learning-to-reconstruct idea is task-agnostic, and offers clear and reasonable supervision to inspire more effective collaboration, eventually promoting perception tasks. We validate Core on two large-scale multi-agent percetion dataset, OPV2V and V2X-Sim, in two tasks, i.e., 3D object detection and semantic segmentation. Results demonstrate that Core achieves state-of-the-art performance, and is more communication-efficient.
Binglu Wang, Lei Zhang 0006, Zhaozhong Wang, Yongqiang Zhao 0001, Tianfei Zhou
ICCV1
2023 Enhanced dynamic feature representation learning framework by Fourier transform for domain generalization
Qingjie Zhao, Changchun Zhang, Binglu Wang, Lei Wang 0252, Wangwang Liu
Inf. Sci.4
2023 Learning pixel-adaptive weights for portrait photo retouching
Binglu Wang, Chengzhe Lu, Yongqiang Zhao 0001, Ning Li 0038, Xuelong Li 0001
Pattern Recognit.1
2023 Cross-Spatial Pixel Integration and Cross-Stage Feature Fusion-Based Transformer Network for Remote Sensing Image Super-Resolution
abstract
Remote sensing image super-resolution (RSISR) plays a vital role in enhancing spatial detials and improving the quality of satellite imagery. Recently, Transformer-based models have shown competitive performance in RSISR. To mitigate the quadratic computational complexity resulting from global self-attention, various methods constrain attention to a local window, enhancing its efficiency. Consequently, the receptive fields in a single attention layer are inadequate, leading to insufficient context modeling. Furthermore, while most transform-based approaches reuse shallow features through skip connections, relying solely on these connections treats shallow and deep features equally, impeding the model’s ability to characterize them. To address these issues, we propose a novel transformer architecture called Cross-Spatial Pixel Integration and Cross-Stage Feature Fusion Based Transformer Network (SPIFFNet) for RSISR. Our proposed model effectively enhances context cognition and understanding of the entire image, facilitating efficient integration of features cross-stages. The model incorporates Cross-Spatial Pixel Integration Attention (CSPIA) to introduce contextual information into a local window, while Cross-Stage Feature Fusion Attention (CSFFA) adaptively fuses features from the previous stage to improve feature expression in line with the requirements of the current stage. We conducted comprehensive experiments on multiple benchmark datasets, demonstrating the superior performance of our proposed SPIFFNet in terms of both quantitative metrics and visual quality when compared to state-of-the-art methods. Our code is available at https://github.com/Dr-Lyt/SPIFFNet.
Lingtong Min, Binglu Wang, Le Zheng, Yongqiang Zhao 0001, Le Yang 0008, Teng Long 0001
IEEE Trans. Geosci. Remote. Sens.3
2023 Hybrid Attention-Based U-Shaped Network for Remote Sensing Image Super-Resolution
abstract
Recently, remote sensing image super-resolution (RSISR) has drawn considerable attention and made great breakthroughs based on convolutional neural networks (CNNs). Due to the scale and richness of texture and structural information frequently recurring inside the same remote sensing images (RSIs) but varying greatly with different RSIs, state-of-the-art CNN-based methods have begun to explore the multiscale global features in RSIs by using attention mechanisms. However, they are still insufficient to explore significant content attention clues in RSIs. In this article, we present a new hybrid attention-based U-shaped network (HAUNet) for RSISR to effectively explore the multiscale features and enhance the global feature representation by hybrid convolution-based attention. It contains two kinds of convolutional attention-based single-scale feature extraction modules (SEM) to explore the global spatial context information and abstract content information, and a cross-scale interaction module (CIM) as the skip connection between different scale feature outputs of encoders to bridge the semantic and resolution gaps between them. Considering the existence of equipment with poor hardware facilities, we further design a lighter HAUNet-S with about 596K parameters. Experimental attribution analysis method LAM results demonstrate that our HAUNet is a more efficient way to capture meaningful content information and quantitative results can show that our HAUNet can significantly improve the performance of RSISR on four remote sensing test datasets. Meanwhile, HAUNET-S also maintains competitive performance. Our code is available athttps://github.com/likakakaka/HAUNet_RSISR.
Binglu Wang, Yongqiang Zhao 0001, Teng Long 0001
IEEE Trans. Geosci. Remote. Sens.2
2023 Joint Denoising-Demosaicking Network for Long-Wave Infrared Division-of-Focal-Plane Polarization Images With Mixed Noise Level Estimation
abstract
Denoising and demosaicking long-wave infrared (LWIR) division-of-focal-plane (DoFP) polarization images are crucial for various vision applications. However, existing methods rely on the sequential application of individual denoising and demosaicking processes, which may result in the accumulation of errors produced by each process. To address this issue, we propose a joint denoising and demosaicking method for LWIR DoFP images based on a three-stage progressive deep convolutional neural network. To ensure the generalization ability of this network, it is essential to have adequate training data that closely resembles real data. Therefore, we model the complex noise sources that affect LWIR DoFP images as mixed Poisson-Additive-Stripe noise and construct a least-squares problem based on the polarization measurement redundancy error to estimate the parameters of this model on real images. Subsequently, the estimated noise parameters are used to generate training data that enables the network to learn accurate polarization image statistics and improve its generalization ability. The experimental results demonstrate the effectiveness of the proposed method in enhancing the image restoration performance on real LWIR DoFP polarization data.
Ning Li 0038, Binglu Wang, François Goudail, Yongqiang Zhao 0001, Quan Pan 0001
IEEE Trans. Image Process.2
2023 Prototype-Based Intent Perception
abstract
Intent perception is a novel task that aims to understand the intention of images, regular classification methods usually perform unsatisfactorily on intent perception due to the semantic ambiguity problem,i.e. the intra-class variety problem in which images of the same intent class may contain objects of different semantic categories and the inter-class confusion problem in which images of different intent classes may contain objects of similar semantic categories. To address this problem, this paper introduces prototype learning into the intent perception and proposes a unified framework named PIP-Net to reduce the influence of semantic ambiguity. Specifically, for each intent class, we first filter semantic ambiguity samples which are far away from the cluster center. Then we use features of the filtered samples to generate prototypes via clustering algorithm. Besides, we enhance the diversity between prototypes of different classes to better handle the inter-class confusion problem. To update the prototypes in the training process, we introduce a global matching algorithm to holistically match each feature with class prototypes, and use the momentum update strategy to stably update prototypes. Experimental results on the Intentonomy dataset demonstrate that our method can consistently outperform the traditional classification paradigm in multiple baseline models, and verify the effectiveness of our proposed prototype learning paradigm in addressing the intent perception problem. Our proposed PIP-Net achieves a new state-of-the-art performance on Intentonomy, including Macro F1 score of 31.57% and averaging F1 score of 41.85%. Code is available athttps://github.com/CodeMonsterPHD/PIP-Net.
Binglu Wang, Yongqiang Zhao 0001, Teng Long 0001, Xuelong Li 0001
IEEE Trans. Multim.1
2022 GaTector: A Unified Framework for Gaze Object Prediction
abstract
Gaze object prediction is a newly proposed task that aims to discover the objects being stared at by humans. It is of great application significance but still lacks a unified solution framework. An intuitive solution is to incorporate an object detection branch into an existing gaze prediction method. However, previous gaze prediction methods usually use two different networks to extract features from scene image and head image, which would lead to heavy network architecture and prevent each branch from joint optimization. In this paper, we build a novel framework named GaTector to tackle the gaze object prediction problem in a unified way. Particularly, a specific-general-specific (SGS) feature extractor is firstly proposed to utilize a shared backbone to extract general features for both scene and head images. To better consider the specificity of inputs and tasks, SGS introduces two input-specific blocks before the shared backbone and three task-specific blocks after the shared backbone. Specifically, a novel Defocus layer is designed to generate object-specific features for the object detection task without losing information or requiring extra computations. Moreover, the energy aggregation loss is introduced to guide the gaze heatmap to concentrate on the stared box. In the end, we propose a novel wUoC metric that can reveal the difference between boxes even when they share no overlapping area. Extensive experiments on the GOO dataset verify the superiority of our method in all three tracks, i.e. object detection, gaze estimation, and gaze object prediction.
Binglu Wang, Tao Hu 0013, Baoshan Li
CVPR1
2022 Exploring Sub-Action Granularity for Weakly Supervised Temporal Action Localization
abstract
Modeling cross-video relationship is an important issue for the weakly supervised temporal action localization task. To this end, traditional methods operate at the action level and rely on complicated strategies to prepare triplet samples, which only mines the cross-video relationships among three videos from two categories. In this work, we observe that action instances from different categories could exhibit similar motion patterns, i.e. subaction, and propose to operate at the sub-action granularity to elaborately explore cross-video relationships. However, only given video-level category labels, the sub-actions are undefined and not annotated. To tackle this challenge, we represent video features via a group of sub-actions,i.e.the sub-action family. Specifically, the sub-action family contains multiple feature vectors, where each vector is in charge of representing a specific sub-action. The sub-action family is shared among all videos in the dataset, while all videos contribute to the learning of the sub-action family. Consequently, we can not only get rid of the complicated sampling strategy but also thoroughly mine cross-video relationships from all available videos in the dataset. To learn feature vectors within the sub-action family, we employ a bottom-up temporal action localization paradigm and introduce an extra top-down branch. The sub-action family is introduced into the top-down branch, and it learns feature vectors via representing raw video features. Moreover, we propose a consistency loss to guide the learning process and a diversity loss to mine distinct sub-actions. Extensive experiments are carried out on three benchmark datasets,i.e.THUMOS14, ActivityNet v1.2 and ActivityNet v1.3, and the proposed method builds new high performance.
Binglu Wang, Yongqiang Zhao 0001
IEEE Trans. Circuits Syst. Video Technol.1
2022 Multiple Instance Graph Learning for Weakly Supervised Remote Sensing Object Detection
abstract
Weakly supervised object detection (WSOD) has recently attracted much attention in the field of remote sensing, where only image-level labels that distinguish the existence of an object in images are required. However, existing methods frequently treat the most discriminative area of an object as the optimal solution and, meanwhile, ignore the fact that more than one instance may exist in a certain class in remote sensing images (RSIs). To address the issue, we propose a unique multiple instance graph (MIG) learning framework for WSOD in RSIs. The motivation of this work is twofold: 1) a spatial graph-based vote (SGV) mechanism is proposed to find high-quality objects by collecting the top-ranking votes with highly spatial overlap and 2) an appearance graph-based instance mining (AGIM) model is further constructed to exploit all possible instances with the same class by propagating the label information according to the apparent similarity. It is noted that the formulated MIG framework that collaborates SGV and AGIM is independent of extra hyperparameters or annotations. Experimental results reported for two well-known benchmarks, i.e., NWPU VHR-10.v2 and DIOR, testify to the superiority of the proposed framework by 55.9% and 25.11% mAPs.
Binglu Wang, Yongqiang Zhao 0001, Xuelong Li 0001
IEEE Trans. Geosci. Remote. Sens.1
2021 LF-MAGNet: Learning Mutual Attention Guidance of Sub-Aperture Images for Light Field Image Super-Resolution
Zijian Wang 0007, Yao Lu 0001, Haowei Lu 0002, Shunzhou Wang, Binglu Wang
PRCV (3)6
2021 SODA: Weakly Supervised Temporal Action Localization Based on Astute Background Response and Self-Distillation Learning
Tao Zhao 0006, Junwei Han 0001, Le Yang 0008, Binglu Wang, Dingwen Zhang
Int. J. Comput. Vis.4
2021 PFWNet: Pretraining neural network via feature jigsaw puzzle for weakly-supervised temporal action localization
Binglu Wang, Yongqiang Zhao 0001
Neurocomputing1
2021 I2Net: Mining intra-video and inter-video attention for temporal action localization
abstract
This paper focuses on two challenges for temporal action localization community, i.e., lack of long-term relationship and action pattern uncertainty. The former prevents the cooperation among multiple action instances within a video, while the latter may cause incomplete localizations or false positives. The lack of long-term relationship challenge results from the limited receptive field. Instead of stacking multiple layers or using large convolution kernels, we propose the intra-video attention mechanism to bring global receptive field to each temporal point. As for the action pattern uncertainty challenge, although it is hard to precisely depict the desired action pattern, paired videos that share the same action category can provide complementary information about action pattern. Consequently, we propose an inter-video attention mechanism to assist learning accurate action patterns. Based on the intra-video attention and inter-video attention, we propose a unified framework, namely I2Net, to tackle the challenging temporal action localization task. Given two videos containing sharing action categories, I2Net adopts the widely used one-stage action localization paradigm to dispose of them in parallel. As for two neighboring layers within the same video, the intra-video attention brings global information to each temporal point and helps to learn representative features. As for two parallel layers between two videos, the inter-video attention introduces complementary information to each video and helps to learn accurate action patterns. With the cooperation of intra-video and inter-video attention mechanisms, I2Net shows obvious performance gains over the baseline and builds new state-of-the-art on two widely-used benchmarks, i.e., THUMOS14 and ActivityNet v1.3.
Binglu Wang, Songhui Ma, Yongqiang Zhao 0001
Neurocomputing2
2021 POLO: Learning Explicit Cross-Modality Fusion for Temporal Action Localization
abstract
Temporal action localization aims at discovering action instances in untrimmed videos, where RGB and flow are two widely used feature modalities. Specifically, RGB chiefly reveals appearance and flow mainly depicts motion. Given RGB and flow features, previous methods employ the early fusion or late fusion paradigm to mine the complementarity between them. By concatenating raw RGB and flow features, the early fusion implicitly achieved complementarity by the network, but it partly discards the particularity of each modality. The late fusion independently maintains two branches to explore the particularity of each modality, but it only fuses the localization results, which is insufficient to mine the complementarity. In this work, we propose explicit cross-modality fusion (POLO) to effectively utilize the complementarity between two modalities and thoroughly explore the particularity of each modality. POLO performs cross-modality fusion via estimating the attention weight from RGB modality and employing it to flow modality (vice versa). This bridges the complementarity of one modality to supply the other. Assisted with the attention weight, POLO independently learns from RGB and flow features and explores the particularity of each modality. Extensive experiments on two benchmarks demonstrate the preferable performance of POLO.
Binglu Wang, Le Yang 0008, Yongqiang Zhao 0001
IEEE Signal Process. Lett.1
2021 Hyperspectral and Multispectral Image Fusion via Graph Laplacian-Guided Coupled Tensor Decomposition
abstract
We propose a novel graph Laplacian-guided coupled tensor decomposition (gLGCTD) model for fusion of hyperspectral image (HSI) and multispectral image (MSI) for spatial and spectral resolution enhancements. The coupled Tucker decomposition is employed to capture the global interdependencies across the different modes to fully exploit the intrinsic global spatial-spectral information. To preserve local characteristics, the complementary submanifold structures embedded in high-resolution (HR)-HSI are encoded by the graph Laplacian regularizations. The global spatial-spectral information captured by the coupled Tucker decomposition and the local submanifold structures are incorporated into a unified framework. The gLGCTD fusion framework is solved by a hybrid framework between the proximal alternating optimization (PAO) and the alternating direction method of multipliers (ADMM). Experimental results on both synthetic and real data sets demonstrate that the gLGCTD fusion method is superior to state-of-the-art fusion methods with a more accurate reconstruction of the HR-HSI.
Yuanyang Bu, Yongqiang Zhao 0001, Jize Xue, Jonathan Cheung-Wai Chan, Seong G. Kong, Jinhuan Wen, Binglu Wang
IEEE Trans. Geosci. Remote. Sens.8
2020 Considering author sequence in all-author co-citation analysis
Yi Bu 0001, Binglu Wang, Zaida Chinchilla-Rodríguez, Cassidy R. Sugimoto, Yong Huang 0008, Win-Bin Huang
Inf. Process. Manag.2
2019 Analysis of Automation Systems and Virtual Reality Applications in Smart Libraries
abstract
With the rapid development of information technology, automation systems and virtual reality technology have become important pillars of economic and social development. As a key institution to store books and information materials, develop information resources and provide knowledge support, libraries are constructing and optimizing their systems. Based on the literature review about library automation systems and virtual reality application, this paper analyzes their expectable influence to promote more concern from scholars and technical personnel, and thereby, enhance the overall management and service level of libraries, especially for smart libraries, which represent the development trend of this smart+ era.
Yang Xu 0022, Binglu Wang, Yi Bu 0001, Changjiang Ji
ICIS3
2017 Share-housing allocation information service system based on Pareto optimality
abstract
Compared to living alone, sharing housing with others is usually much cheaper. Meanwhile, share-housing life would exert mammoth impact on people's social activity and self development by help them save house rental cost and improve their life quality. This paper proposes an allocation information system for share-housing management, which takes into account preference on roommates' living habits, working performance and geography distribution. Moreover, Pareto optimality is utilized to facilitate the effect of share-housing allocation. Results of case study show that the three factors of living habits, working performance as well as geography distribution should be synthetically applied as a decision support for share-housing.
Yang Xu 0022, Shuwen Liu 0001, Binglu Wang, Zhengnan Liu
SNPD3