VLDB 2026 Research / reviewers in the wild / expert
Bing Liu 0016
dblp:l/BingLiu16
· DBLP profile ↗
51ranked-venue papers
12as first author
39since 2021 · last 2026
0000-0002-2365-6606ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 29 · 7 first-author · 27 since 2021Artificial intelligence and machine learning · 20 · 6 first-author · 11 since 2021Computer networks · 6 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Few-Shot Quaternion-valued Correlation Squeeze Network for Document Image Layout Segmentation
Rui Yao 0006, Qiwei Yu, Songhui Zhao, Yong Zhou 0003, Bing Liu 0016 |
Int. J. Document Anal. Recognit. | 6 |
| 2026 | Style-controllable adversarial example generation via image editing and prompt embedding optimization
Yong Zhou 0003, Bing Liu 0016, Rui Yao 0006 |
Neurocomputing | 3 |
| 2026 | A unified multi-stream diffusion framework for robust video camouflaged object detection
Yuyao Ke, Rui Yao 0006, Kunyang Sun, Hancheng Zhu, Jiaqi Zhao 0001, Bing Liu 0016 |
Neural Networks | 7 |
| 2026 | Retrieval-Augmented Pseudo-Image Guided Alignment and Text Domain-Aware Memory Recall for Continual Zero-Shot Captioning
Bing Liu 0016, Hao Liu 0065, Peng Liu 0013, Yong Zhou 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | Facial Action Unit Detection with Iterative Rank Reduction Adapter and Directional Attention
Zhiwen Shao, Hancheng Zhu, Rui Yao 0006, Bing Liu 0016 |
CGI (1) | 6 |
| 2025 | Residual-based Efficient Bidirectional Diffusion Model for Image Dehazing and Haze GenerationabstractCurrent deep dehazing methods only focus on removing haze from hazy images, lacking the capability to translate between hazy and haze-free images. To address this issue, we propose a residual-based efficient bidirectional diffusion model (RBDM) that can model the conditional distributions for both dehazing and haze generation. Firstly, we devise dual Markov chains that can effectively shift the residuals and facilitate bidirectional smooth transitions between them. Secondly, the RBDM perturbs the hazy and haze-free images at individual timesteps and predicts the noise in the perturbed data to simultaneously learn the conditional distributions. Finally, to enhance performance on relatively small datasets and reduce computational costs, our method introduces a unified score function learned on image patches instead of entire images. Our RBDM successfully implements size-agnostic bidirectional transitions between hazefree and hazy images with only 15 sampling steps. Extensive experiments demonstrate that the proposed method achieves superior or at least comparable performance to state-of-the-art methods on both synthetic and real-world datasets. Bing Liu 0016, Hao Liu 0065 |
ICME | 1 |
| 2025 | Modality-Guided Dynamic Graph Fusion and Temporal Diffusion for Self-Supervised RGB-T TrackingabstractTo reduce the reliance on large-scale annotations, self-supervised RGB-T tracking approaches have garnered significant attention. However, the omission of the object region by erroneous pseudo-label or the introduction of background noise affects the efficiency of modality fusion, while pseudo-label noise triggered by similar object noise can further affect the tracking performance. In this paper, we propose GDSTrack, a novel approach that introduces dynamic graph fusion and temporal diffusion to address the above challenges in self-supervised RGB-T tracking. GDSTrack dynamically fuses the modalities of neighboring frames, treats them as distractor noise, and leverages the denoising capability of a generative model. Specifically, by constructing an adjacency matrix via an Adjacency Matrix Generator (AMG), the proposed Modality-guided Dynamic Graph Fusion (MDGF) module uses a dynamic adjacency matrix to guide graph attention, focusing on and fusing the object’s coherent regions. Temporal Graph-Informed Diffusion (TGID) models MDGF features from neighboring frames as interference, and thus improving robustness against similar-object noise. Extensive experiments conducted on four public RGB-T tracking datasets demonstrate that GDSTrack outperforms the existing state-of-the-art methods. The source code is available at https://github.com/LiShenglana/GDSTrack. Shenglan Li, Rui Yao 0006, Yong Zhou 0003, Hancheng Zhu, Kunyang Sun, Bing Liu 0016, Zhiwen Shao, Jiaqi Zhao 0001 |
IJCAI | 6 |
| 2025 | Counterfactual Knowledge Maintenance for Unsupervised Domain AdaptationabstractTraditional unsupervised domain adaptation (UDA) struggles to extract rich semantics due to backbone limitations. Recent large-scale pre-trained visual-language models (VLMs) have shown strong zero-shot learning capabilities in UDA tasks. However, directly using VLMs results in a mixture of semantic and domain-specific information, complicating knowledge transfer. Complex scenes with subtle semantic differences are prone to misclassification, which in turn can result in the loss of features that are crucial for distinguishing between classes. To address these challenges, we propose a novel counterfactual knowledge maintenance UDA framework. Specifically, we employ counterfactual disentanglement to separate the representation of semantic information from domain features, thereby reducing domain bias. Furthermore, to clarify ambiguous visual information specific to classes, we maintain the discriminative knowledge of both visual and textual information. This approach synergistically leverages multimodal information to preserve modality-specific distinguishable features. We conducted extensive experimental evaluations on several public datasets to demonstrate the effectiveness of our method. The source code is available at https://github.com/LiYaolab/CMKUDA Yong Zhou 0003, Jiaqi Zhao 0001, Wen-Liang Du 0002, Rui Yao 0006, Bing Liu 0016 |
IJCAI | 6 |
| 2025 | Multi-scale Progressive Low-Light Image Enhancement Networks
Bing Liu 0016, Jiale Cao |
PRCV (9) | 1 |
| 2025 | Dual-Domain Low-Light Image Enhancement Network via Frequency Interaction and Structure-Guided Attention
Bing Liu 0016, Jiale Cao, Peng Liu 0013 |
PRCV (8) | 1 |
| 2025 | Enhanced Teacher-Student Framework with Weighted Loss and Presence Awareness for Semi-supervised Semantic Segmentation
Peng Liu 0013, Chengcheng Zhou, Hengyu Cao, Xuekui Wang, Bing Liu 0016 |
PRCV (1) | 6 |
| 2025 | Facial Action Unit Detection by Adaptively Constraining Self-Attention and Causally Deconfounding Sample
Zhiwen Shao, Hancheng Zhu, Yong Zhou 0003, Xiang Xiang 0001, Bing Liu 0016, Rui Yao 0006, Lizhuang Ma |
Int. J. Comput. Vis. | 5 |
| 2025 | L2A: Learning Affinity From Attention for Weakly Supervised Continual Semantic SegmentationabstractDespite significant advances in continual semantic segmentation (CSS), they still rely on the pixel-level annotation to train models, which is time-consuming and labor-intensive. Continual learning from image-level labels is an emerging scheme in continual semantic segmentation to reduce the annotation cost. However, the incomplete and coarse pseudo-labels are insufficient to train a model to maintain a balance between stability and plasticity. To solve these issues, we propose a novel end-to-end framework based on Transformer, called L2A, for Weakly Supervised Continual Semantic Segmentation (WSCSS). In particular, to generate reliable annotations from the image-level supervision, we introduce a semantic affinity from multi-head self-attention (SA-MHSA) module to capture the semantic relationships among adjacent image coordinates. Subsequently, this acquired semantic affinity is employed to refine the initial pseudo labels of new classes trained with the image-level annotations. Furthermore, to minimize catastrophic forgetting, we propose a semantic drift compensation (SDC) strategy to optimize the pseudo-label generation process, which can effectively improve the alignment of object boundaries across both new and old categories. Comprehensive experiments conducted on the PASCAL VOC 2012 and COCO datasets demonstrate the superiority of our framework in existing WSCSS scenarios and a newly proposed challenge protocol, as well as remains competitive compared to the pixel-level supervised CSS methods. Hao Liu 0065, Yong Zhou 0003, Bing Liu 0016, Ming Yan 0007, Joey Tianyi Zhou |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Mirror Detection via Multi-Directional Similarity Perception and Spectral Saliency EnhancementabstractMirror detection is a challenging task, due to the reflective properties of mirrors. Most existing approaches rely on exploiting the relationship between the content inside the mirror and the surrounding environment to aid in locating mirrors. A typical solution is to utilize contextual contrasted features. However, the discontinuity in content at the edges of mirrors may not always be prominent. To overcome this limitation, we propose a novel mirror detection framework called S2MD including two main modules, multi-directional similarity perception module (MSPM) and spectral saliency enhancement decoder module (SSEDM). Specifically, we employ a backbone network to extract multi-scale global information from images using a dual-path approach. Then, we feed these high-level dual-path features into MSPMs to generate direction-sensitive similarity-consistent features. MSPM utilizes active rotating filters and oriented response pooling to model the similarity relations in different orientations. Moreover, the SSEDM is utilized to enhance the spatial contextual contrasted features using feature spectral residuals and fuse the dual-path features to obtain the final predicted mirror mask. Extensive experiments demonstrate that our method achieves state-of-the-art performance on challenging MSD, PMD, and RGBD-Mirror benchmarks. The code is available at https://github.com/RuiChen-stack/M2SD. Zhiwen Shao, Xuehuai Shi, Bing Liu 0016, Canlin Li, Lizhuang Ma, Dit-Yan Yeung |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Hierarchical Relation Learning for Few-Shot Semantic Segmentation in Remote Sensing ImagesabstractFew-shot semantic segmentation (FSS) aims to segment specific semantic classes in a query image using only a few annotated support samples. While FSS has gained significant attention in natural image processing, it remains underexplored in the more challenging domain of remote sensing images (RSIs). Existing FSS approaches for RSIs primarily focus on enhancing feature representations of support or query images through hierarchical/multi-level feature fusion. However, unlike fully supervised segmentation that relies on feature extraction and optimization, FSS requires segmenting the query image based on its relations with annotated support images. To address this need, we propose the concept of Hierarchical Relation Learning (HRL) to explore the intrinsic support-query relations, allowing for the direct refinement of target object appearances in the query image. Specifically, we propose a Hierarchical Relation Network (HRNet), which performs single-scale relation extraction at each network hierarchy and multi-scale relation aggregation across hierarchies. In addition, we construct a Bidirectional Hierarchical Loss (BHLoss) to guide HRNet training, providing targeted supervision at each hierarchy in both top-down and bottom-up directions, thus facilitating robust multi-scale relation learning across hierarchies. Comprehensive experiments on the iSAID-5i, DLRSD-5i, and LoveDA-2i datasets demonstrate the superiority of the proposed HRL. The code will be available at https://github.com/XinnHe/HRL. Xin He 0024, Yun Liu 0011, Yong Zhou 0003, Henghui Ding, Jiaqi Zhao 0001, Bing Liu 0016, Xudong Jiang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | Adversarial Geometric Attacks for 3D Point Cloud Object Trackingabstract3D point cloud object tracking (3D PCOT) plays a vital role in applications such as autonomous driving and robotics. Adversarial attacks offer a promising approach to enhance the robustness and security of tracking models. However, existing adversarial attack methods for 3D PCOT seldom leverage the geometric structure of point clouds and often overlook the transferability of attack strategies. To address these limitations, this paper proposes an adversarial geometric attack method tailored for 3D PCOT, which includes a point perturbation attack module (non-isometric transformation) and a rotation attack module (isometric transformation). First, we introduce a curvature-aware point perturbation attack module that enhances local transformations by applying normal perturbations to critical points identified through geometric features such as curvature and entropy. Second, we design a Thompson sampling-based rotation attack module that applies subtle global rotations to the point cloud, introducing tracking errors while maintaining imperceptibility. Additionally, we design a fused loss function to iteratively optimize the point cloud within the search region, generating adversarially perturbed samples. The proposed method is evaluated on multiple 3D PCOT models and validated through black-box tracking experiments on benchmarks. For P2B, white-box attacks on KITTI reduce the success rate from 53.3% to 29.6% and precision from 68.4% to 37.1%. On NuScenes, the success rate drops from 39.0% to 27.6%, and precision from 39.9 to 26.8%. Black-box attacks show a transferability, with BAT showing a maximum 47.0% drop in success rate and 47.2% in precision on KITTI, and a maximum 22.5% and 27.0% on NuScenes. Rui Yao 0006, Yong Zhou 0003, Jiaqi Zhao 0001, Bing Liu 0016, Abdulmotaleb El Saddik |
IEEE Trans. Multim. | 5 |
| 2025 | Syntactic-Conditional Diffusion Networks for Controllable Image CaptioningabstractCurrent diffusion model-based image captioning methods generally focus on generating descriptions in a non-autoregressive manner. Nevertheless, it is not trivial to employ such generative models to control the generation of discrete words while pursuing the balance between diversity and accuracy. Inspired by the success of continuous diffusions in image captioning, we introduce the Part-of-Speech (POS) information and classifier-free guidance into the diffusion model, and propose a novel controllable image captioning model, namely POS-Conditional Diffusion Networks (POSCD-Net), which consists of a Diffusion-based POS Generator (DPG) and a Diffusion-based Caption Generator (DCG). The DPG is built to produce diverse syntactic structures for each input image. The diverse POS sequences are further regarded as the control signals of the DCG, which produces the output sentences in a conditional diffusion process. In the DCG, a syntactic control module (SCM) is designed to strengthen the alignment progressively between words and the corresponding POS tags in a cascaded manner. Furthermore, to improve the controllability of POSCD-Net, the classifier-free guidance with learnable parameters is exploited to jointly optimize both the DPG and DCG in a non-autoregressive manner. Extensive experiments on the MSCOCO dataset demonstrate that our proposed method outperforms the state-of-the-art non-autoregressive counterparts and achieves promising performance compared with the autoregressive models. Bing Liu 0016, Hao Liu 0065, Yong Zhou 0003, Peng Liu 0013 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2025 | High-level LoRA and hierarchical fusion for enhanced micro-expression recognition
Zhiwen Shao, Yong Zhou 0003, Xiang Xiang 0001, Jian Li 0054, Bing Liu 0016, Dit-Yan Yeung |
Vis. Comput. | 6 |
| 2024 | Remote sensing image semantic segmentation via class-guided structural interaction and boundary perceptionabstractExisting remote sensing semantic segmentation methods generally ignore the structural information of objects that is vital in the human visual recognition system. The absence of overall structural information often results in weak perceptions of subtle textures and fragmented predictions, especially for complex and variable ground object scenarios. Besides, they still suffer from the semantic ambiguity caused by the unclear object boundary features in remote sensing images. In this paper, we propose a novel remote sensing semantic segmentation framework, called CSBNet, which aims to enhance the capacity of class-guided structural interaction and boundary perception simultaneously. It consists of a class-guided structure interaction module (CSIM), a Transformer-based context aggregation module (TCAM) and a class-guided boundary supervision module (CBSM). The CSIM has the ability to progressively extract the class-specific structural features, i.e. , refining the structural information of each class by iteratively exchanging information between initial coarse class tokens and contexts. Meanwhile, the TCAM is constructed to provide CSIM with more discriminative multi-scale contexts without losing spatial features. In particular, the CBSM plays an auxiliary role, which applies the boundary information obtained from the class tokens to supervise the segmentation of boundary regions. When tested on the ISPRS dataset, LoveDA dataset, UAVid dataset, our method significantly outperforms the state-of-the-art remote sensing semantic segmentation approaches. Xin He 0024, Yong Zhou 0003, Bing Liu 0016, Jiaqi Zhao 0001, Rui Yao 0006 |
Expert Syst. Appl. | 3 |
| 2024 | Efficient convolutional neural networks and network compression methods for object detection: a survey
Yong Zhou 0003, Jiaqi Zhao 0001, Rui Yao 0006, Bing Liu 0016 |
Multim. Tools Appl. | 5 |
| 2024 | Joint facial action unit recognition and self-supervised optical flow estimation
Zhiwen Shao, Yong Zhou 0003, Feiran Li, Hancheng Zhu, Bing Liu 0016 |
Pattern Recognit. Lett. | 5 |
| 2024 | CT-Net: Arbitrary-Shaped Text Detection via Contour TransformerabstractContour based scene text detection methods have rapidly developed recently, but still suffer from inaccurate front-end contour initialization, multi-stage error accumulation, or deficient local information aggregation. To tackle these limitations, we propose a novel arbitrary-shaped scene text detection framework named CT-Net by progressive contour regression with contour transformers. Specifically, we first employ a contour initialization module that generates coarse text contours without any post-processing. Then, we adopt contour refinement modules to adaptively refine text contours in an iterative manner, which are beneficial for context information capturing and progressive global contour deformation. Besides, we propose an adaptive training strategy to enable the contour transformers to learn more potential deformation paths, and introduce a re-score mechanism that can effectively suppress false positives. Extensive experiments are conducted on four challenging datasets, which demonstrate the accuracy and efficiency of our CT-Net over state-of-the-art methods. Particularly, CT-Net achieves F-measure of 86.1 at 11.2 frames per second (FPS) and F-measure of 87.8 at 10.1 FPS for CTW1500 and Total-Text datasets, respectively. Zhiwen Shao, Yong Zhou 0003, Hancheng Zhu, Bing Liu 0016, Rui Yao 0006 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | Black-box Attack against Self-supervised Video Object Segmentation Models with Contrastive LossabstractDeep learning models have been proven to be susceptible to malicious adversarial attacks, which manipulate input images to deceive the model into making erroneous decisions. Consequently, the threat posed to these models serves as a poignant reminder of the necessity to focus on the model security of object segmentation algorithms based on deep learning. However, the current landscape of research on adversarial attacks primarily centers around static images, resulting in a dearth of studies on adversarial attacks targeting Video Object Segmentation (VOS) models. Given that a majority of self-supervised VOS models rely on affinity matrices to learn feature representations of video sequences and achieve robust pixel correspondence, our investigation has delved into the impact of adversarial attacks on self-supervised VOS models. In response, we propose an innovative black-box attack method incorporating contrastive loss. This method induces segmentation errors in the model through perturbations in the feature space and the application of a pixel-level loss function. Diverging from conventional gradient-based attack techniques, we adopt an iterative black-box attack strategy that incorporates contrastive loss across the current frame, any two consecutive frames, and multiple frames. Through extensive experimentation conducted on the DAVIS 2016 and DAVIS 2017 datasets using three self-supervised VOS models and one unsupervised VOS model, we unequivocally demonstrate the potent attack efficiency of the black-box approach. Remarkably, theJ&Fmetric value experiences a significant decline of up to 50.08% post-attack. Ying Chen 0005, Rui Yao 0006, Yong Zhou 0003, Jiaqi Zhao 0001, Bing Liu 0016, Abdulmotaleb El Saddik |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2024 | Diverse Image Captioning via Panoptic Segmentation and Sequential Conditional Variational TransformerabstractRecently, transformer-based image captioning models have achieved significant performance improvement. However, due to the limitations of region visual features and deterministic projections between image space and caption space, existing methods still suffer from disentangled visual features and rigid sentences. To address these issues, we first introduce panoptic segmentation to extract the segmentation region features, which can effectively alleviate the visual confusion caused by the widely-adopted region visual features. Then, we propose a panoptic segmentation based sequential conditional variational transformer (PS-SCVT) framework for diverse image captioning, which not only accurately extracts the image visual representations by fusing the segmentation region features and object detection features, but has the ability of learning one-to-many mappings from image space to caption space. The experimental results demonstrate that our approach achieves better interpretability and generalization performance compared with the state-of-the-art diverse image captioning models. Bing Liu 0016, Jinfu Lu, Hao Liu 0065, Yong Zhou 0003, Dongping Yang |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2024 | Diverse Image Captioning via Conditional Variational Autoencoder and Dual Contrastive LearningabstractDiverse image captioning has achieved substantial progress in recent years. However, the discriminability of generative models and the limitation of cross entropy loss are generally overlooked in the traditional diverse image captioning models, which seriously hurts both the diversity and accuracy of image captioning. In this article, aiming to improve diversity and accuracy simultaneously, we propose a novel Conditional Variational Autoencoder (DCL-CVAE) framework for diverse image captioning by seamlessly integrating sequential variational autoencoder with contrastive learning. In the encoding stage, we first build conditional variational autoencoders to separately learn the sequential latent spaces for a pair of captions. Then, we introduce contrastive learning in the sequential latent spaces to enhance the discriminability of latent representations for both image-caption pairs and mismatched pairs. In the decoding stage, we leverage the captions sampled from the pre-trained Long Short-Term Memory (LSTM), LSTM decoder as the negative examples and perform contrastive learning with the greedily sampled positive examples, which can restrain the generation of common words and phrases induced by the cross entropy loss. By virtue of dual constrastive learning, DCL-CVAE is capable of encouraging the discriminability and facilitating the diversity, while promoting the accuracy of the generated captions. Extensive experiments are conducted on the challenging MSCOCO dataset, showing that our proposed methods can achieve a better balance between accuracy and diversity compared to the state-of-the-art diverse image captioning models. Bing Liu 0016, Yong Zhou 0003, Rui Yao 0006, Zhiwen Shao |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2023 | Unsupervised RGB-T object tracking with attentional multi-modal feature fusion
Shenglan Li, Rui Yao 0006, Yong Zhou 0003, Hancheng Zhu, Bing Liu 0016, Jiaqi Zhao 0001, Zhiwen Shao |
Multim. Tools Appl. | 5 |
| 2023 | Semi-supervised transformable architecture search for feature distillation
Man Zhang 0006, Yong Zhou 0003, Bing Liu 0016, Jiaqi Zhao 0001, Rui Yao 0006, Zhiwen Shao, Hancheng Zhu |
Pattern Anal. Appl. | 3 |
| 2023 | Attention-guided Adversarial Attack for Video Object SegmentationabstractVideo Object Segmentation (VOS) methods have made many breakthroughs with the help of the continuous development and advancement of deep learning. However, the deep learning model is vulnerable to malicious adversarial attacks, which mislead the model to make wrong decisions by adding adversarial perturbation that humans cannot perceive to the input image. Threats to deep learning models remind us that video object segmentation methods are also vulnerable to attacks, thereby threatening their security. Therefore, we study adversarial attacks on the VOS task to better identify the vulnerabilities of the VOS method, which in turn provides an opportunity to improve its robustness. In this paper, we propose an attention-guided adversarial attack method, which uses spatial attention blocks to capture features with global dependencies to construct correlations between consecutive video frames, and performs multipath aggregation to effectively integrate spatial-temporal perturbation, thereby guiding the deconvolution network to generate adversarial examples with strong attack capability. Specifically, the class loss function is designed to enable the deconvolution network to better activate noise in other regions and suppress the activation related to the object class based on the enhanced feature map of the object class. At the same time, attentional feature loss is designed to enhance the transferability against attack. The experimental results on the DAVIS dataset show that the proposed attention-guided adversarial attack method can significantly reduce the segmentation accuracy of OSVOS, and the J & F mean on DAVIS 2016 can reach 73.6% drop rate. The generated adversarial examples are also highly transferable to other video object segmentation models. Rui Yao 0006, Ying Chen 0005, Yong Zhou 0003, Fuyuan Hu, Jiaqi Zhao 0001, Bing Liu 0016, Zhiwen Shao |
ACM Trans. Intell. Syst. Technol. | 6 |
| 2023 | TextDCT: Arbitrary-Shaped Text Detection via Discrete Cosine Transform MaskabstractArbitrary-shaped scene text detection is a challenging task due to the variety of text changes in font, size, color, and orientation. Most existing regression based methods resort to regress the masks or contour points of text regions to model the text instances. However, regressing the complete masks requires high training complexity, and contour points are not sufficient to capture the details of highly curved texts. To tackle the above limitations, we propose a novel light-weight anchor-free text detection framework called TextDCT, which adopts the discrete cosine transform (DCT) to encode the text masks as compact vectors. Further, considering the imbalanced number of training samples among pyramid layers, we only employ a single-level head for top-down prediction. To model the multi-scale texts in a single-level head, we introduce a novel positive sampling strategy by treating the shrunk text region as positive samples, and design a feature awareness module (FAM) for spatial-awareness and scale-awareness by fusing rich contextual information and focusing on more significant features. Moreover, we propose a segmented non-maximum suppression (S-NMS) method that can filter low-quality mask regressions. Extensive experiments are conducted on four challenging datasets, which demonstrate our TextDCT obtains competitive performance on both accuracy and efficiency. Specifically, TextDCT achieves F-measure of 85.1 at 17.2 frames per second (FPS) and F-measure of 84.9 at 15.1 FPS for CTW1500 and Total-Text datasets, respectively. Zhiwen Shao, Yong Zhou 0003, Hancheng Zhu, Bing Liu 0016, Rui Yao 0006 |
IEEE Trans. Multim. | 6 |
| 2023 | Weakly Supervised Few-Shot Semantic Segmentation via Pseudo Mask Enhancement and Meta LearningabstractFew shot semantic segmentation has been proposed to enhance the generalization ability of traditional models with limited data. Previous works mainly focus on the supervised tasks, while limited amount of work is explored for the weakly supervised tasks. Weakly supervised semantic segmentation has become an active research area because weakly supervised labels effectively reduce the annotation cost of visual tasks. To this end, we propose a weakly supervised few-shot semantic segmentation model based on the meta learning framework, which utilizes prior knowledge and adjusts itself according to new tasks. Thereupon then, the proposed network is capable of both high efficiency and generalization ability to new tasks. In the pseudo mask generation stage, we develop a WRCAM method with the channel-spatial attention mechanism to refine the coverage size of targets in pseudo masks. In the few-shot semantic segmentation stage, the optimization based meta learning method is used to realize few-shot semantic segmentation by virtue of the refined pseudo masks. The experimental results show that the proposed method not only significantly outperforms weakly supervised SOTA methods, but also could be comparative to some supervised SOTA methods. Man Zhang 0006, Yong Zhou 0003, Bing Liu 0016, Jiaqi Zhao 0001, Rui Yao 0006, Zhiwen Shao, Hancheng Zhu |
IEEE Trans. Multim. | 3 |
| 2023 | Distilled Meta-learning for Multi-Class Incremental LearningabstractMeta-learning approaches have recently achieved promising performance in multi-class incremental learning. However, meta-learners still suffer from catastrophic forgetting, i.e., they tend to forget the learned knowledge from the old tasks when they focus on rapidly adapting to the new classes of the current task. To solve this problem, we propose a novel distilled meta-learning (DML) framework for multi-class incremental learning that integrates seamlessly meta-learning with knowledge distillation in each incremental stage. Specifically, during inner-loop training, knowledge distillation is incorporated into the DML to overcome catastrophic forgetting. During outer-loop training, a meta-update rule is designed for the meta-learner to learn across tasks and quickly adapt to new tasks. By virtue of the bilevel optimization, our model is encouraged to reach a balance between the retention of old knowledge and the learning of new knowledge. Experimental results on four benchmark datasets demonstrate the effectiveness of our proposal and show that our method significantly outperforms other state-of-the-art incremental learning methods. Hao Liu 0065, Zhaoyu Yan, Bing Liu 0016, Jiaqi Zhao 0001, Yong Zhou 0003, Abdulmotaleb El Saddik |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2022 | Show, Deconfound and Tell: Image Captioning with Causal InferenceabstractThe transformer-based encoder-decoder framework has shown remarkable performance in image captioning. However, most transformer-based captioning methods ever overlook two kinds of elusive confounders: the visual confounder and the linguistic confounder, which generally lead to harmful bias, induce the spurious correlations during training, and degrade the model generalization. In this paper, we first use Structural Causal Models (SCMs) to show how two confounders damage the image captioning. Then we apply the backdoor adjustment to propose a novel causal inference based image captioning (CIIC) framework, which consists of an interventional object detector (IOD) and an interventional transformer decoder (ITD) to jointly confront both confounders. In the encoding stage, the IOD is able to disentangle the region-based visual features by deconfounding the visual confounder. In the decoding stage, the ITD introduces causal intervention into the transformer decoder and deconfounds the visual and linguistic confounders simultaneously. Two modules collaborate with each other to alleviate the spurious correlations caused by the unobserved confounders. When tested on MSCOCO, our proposal significantly outperforms the state-of-the-art encoder-decoder models on Karpathy split and online test split. Code is published in https://github.com/CUMTGG/CIIC. Bing Liu 0016, Xu Yang 0021, Yong Zhou 0003, Rui Yao 0006, Zhiwen Shao, Jiaqi Zhao 0001 |
CVPR | 1 |
| 2022 | BiTMulV: Bidirectional-Decoding Based Transformer with Multi-view Visual Representation
Qiankun Yu, Xuekui Wang, Bing Liu 0016, Peng Liu 0013 |
PRCV (1) | 5 |
| 2022 | Semantic Segmentation of Remote-Sensing Images Based on Multiscale Feature Fusion and Attention RefinementabstractIn recent years, the automatic extraction of remote-sensing image information has attracted full attention. However, the particularity of remote-sensing images and the scarcity of data sets with label information have brought new challenges to existing methods. Therefore, we develop a lightweight semantic segmentation network based onmultiscale feature fusion (MFF) and attention refinement (MFFANet). Our network relies on three crucial modules for improved performance. The multiscale attention refinement module strengthens the representation ability of feature maps extracted by the deep residual network. The MFF module aggregates the information carried by the high-level and low-level features while restoring the image resolution. Furthermore, the boundary enhancement module captures boundary details to solve the semantic ambiguity problem. We achieve 83.5% mean intersection over union (MIoU) on the Urban Semantic 3-D (US3D) data set and 69.3% MIoU on the Vaihingen data set with only 8.2M parameters. Xin He 0024, Yong Zhou 0003, Jiaqi Zhao 0001, Man Zhang 0006, Rui Yao 0006, Bing Liu 0016 |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2022 | Semisupervised Multiscale Generative Adversarial Network for Semantic Segmentation of Remote Sensing ImageabstractSemantic segmentation of remote sensing images based on deep neural networks has gained wide attention recently. Although many methods have achieved amazing performance, they need large amounts of labeled images to distinguish the differences in angle, color, size, and other aspects for small targets in remote sensing data sets. However, with a few labeled images, it is difficult to extract the key features of small targets. We propose a semisupervised multiscale generative adversarial network (GAN), which not only utilizes the multipath input and atrous spatial pyramid pooling (ASPP) module but leverages unlabeled images and semisupervised learning strategy to improve the performance of small target segmentation in semantic segmentation when labeled data amount is small. Experimental results show that our model outperforms state-of-the-art methods with insufficient labeled data. Bing Liu 0016, Yong Zhou 0003, Jiaqi Zhao 0001, Shixiong Xia, Yuancan Yang, Man Zhang 0006, Liu Ming Ming |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2022 | Clustering Matters: Sphere Feature for Fully Unsupervised Person Re-identificationabstractIn person re-identification (Re-ID) , the data annotation cost of supervised learning, is huge and it cannot adapt well to complex situations. Therefore, compared with supervised deep learning methods, unsupervised methods are more in line with actual needs. In unsupervised learning, a key to solving Re-ID is to find a standard that can effectively distinguish the difference (distance) between the features of images belonging to different pedestrian identities. However, there are some differences in the images captured by different cameras (such as brightness, angle, etc.). It is well known that the training of neural networks is mainly based on the distance between features, while in unsupervised learning, especially in unsupervised learning methods based on hierarchical clustering, the distance between features plays a more important role in the clustering phase. We improve the accuracy of a deep learning method based on hierarchical clustering under fully unsupervised conditions, starting from both feature and distance metrics. First, we propose to use spherical features, by normalizing the images in the feature space, to weaken the structural differences (length) between features, while saving the feature differences (direction) between different identities. Then, we use the sum of squared errors (SSE) as a regularization term to balance different cluster states. We evaluate our method on four large-scale Re-ID datasets, and experiments show that our method achieves better results than the state-of-the-art unsupervised methods. Yong Zhou 0003, Jiaqi Zhao 0001, Ying Chen 0005, Rui Yao 0006, Bing Liu 0016, Abdulmotaleb El Saddik |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2022 | Facial action unit detection via hybrid relational reasoning
Zhiwen Shao, Yong Zhou 0003, Bing Liu 0016, Hancheng Zhu, Wen-Liang Du 0002, Jiaqi Zhao 0001 |
Vis. Comput. | 3 |
| 2021 | Point cloud classification by dynamic graph CNN with adaptive feature fusionabstractAbstract The deep neural network has made the most advanced breakthrough in almost all 2D image tasks, so we consider the application of deep learning in 3D images. Point cloud data, as the most basic and important form of representation of 3D images, can accurately and intuitively show the real world. The authors propose a new network based on feature fusion to improve the point cloud classification and segmentation tasks. Our network mainly consists of three parts: global feature extractor, local feature extractor and adaptive feature fusion module. A multi‐scale transformation network is devised to guarantee the invariance of the transformation of the global feature, and a residual block is introduced to alleviate the problem of gradient disappearance to enhance the global feature extractor. Based on the edge convolution and multi‐layer perceptron, a local feature extractor is constructed. Finally, an adaptive feature‐fusion module is proposed to complete the fusion of global features and local features. Extensive experiments on point cloud classification and segmentation tasks are carried out to verify the effectiveness of the proposed method. The classification accuracy of the ModelNet40 is 93.6%, which is 4.4% higher than that of the PointNet. Similarly, the segmentation accuracy on the ShapeNet is 85.6%, which is higher than other methods. Yong Zhou 0003, Jiaqi Zhao 0001, Yiyun Man, Minjie Liu, Rui Yao 0006, Bing Liu 0016 |
IET Comput. Vis. | 7 |
| 2021 | A siamese pedestrian alignment network for person re-identification
Yong Zhou 0003, Jiaqi Zhao 0001, Meng Jian, Rui Yao 0006, Bing Liu 0016, Ying Chen 0005 |
Multim. Tools Appl. | 6 |
| 2020 | Multiobjective ResNet pruning by means of EMOAs for remote sensing scene classification
Xuning Liu, Yong Zhou 0003, Jiaqi Zhao 0001, Rui Yao 0006, Bing Liu 0016 |
Neurocomputing | 5 |
| 2020 | Remote sensing image captioning via Variational Autoencoder and Reinforcement Learning
Xiangqing Shen, Bing Liu 0016, Yong Zhou 0003, Jiaqi Zhao 0001 |
Knowl. Based Syst. | 2 |
| 2020 | Remote sensing image caption generation via transformer and reinforcement learning
Xiangqing Shen, Bing Liu 0016, Yong Zhou 0003, Jiaqi Zhao 0001 |
Multim. Tools Appl. | 2 |
| 2019 | A Siamese Pedestrian Alignment Network for Person Re-identification
Yong Zhou 0003, Jiaqi Zhao 0001, Meng Jian, Rui Yao 0006, Bing Liu 0016, Xuning Liu |
PRCV (1) | 6 |
| 2019 | Saliency detection via double nuclear norm maximization and ensemble manifold regularization
Bing Liu 0016, Yong Zhou 0003, Peng Liu 0013, Shuang Li 0011, Xinqiu Fang |
Knowl. Based Syst. | 1 |
| 2019 | Siamese Convolutional Neural Networks for Remote Sensing Scene ClassificationabstractThe convolutional neural networks (CNNs) have shown powerful feature representation capability, which provides novel avenues to improve scene classification of remote sensing imagery. Although we can acquire large collections of satellite images, the lack of rich label information is still a major concern in the remote sensing field. In addition, remote sensing data sets have their own limitations, such as the small scale of scene classes and lack of image diversity. To mitigate the impact of the existing problems, a Siamese CNN, which combines the identification and verification models of CNNs, is proposed in this letter. A metric learning regularization term is explicitly imposed on the features learned through CNNs, which enforce the Siamese networks to be more robust. We carried out experiments on three widely used remote sensing data sets for performance evaluation. Experimental results show that our proposed method outperforms the existing methods. Xuning Liu, Yong Zhou 0003, Jiaqi Zhao 0001, Rui Yao 0006, Bing Liu 0016 |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2018 | Spectral regression based marginal Fisher analysis dimensionality reduction algorithm
Bing Liu 0016, Yong Zhou 0003, Zhanguo Xia, Peng Liu 0013, Qiuyan Yan |
Neurocomputing | 1 |
| 2016 | Manifold regularized extreme learning machine
Bing Liu 0016, Shixiong Xia, Yong Zhou 0003 |
Neural Comput. Appl. | 1 |
| 2015 | Extreme spectral regression for efficient regularized subspace learning
Bing Liu 0016, Shixiong Xia, Yong Zhou 0003 |
Neurocomputing | 1 |
| 2015 | Semi-supervised extreme learning machine with manifold and pairwise constraints regularization
Yong Zhou 0003, Beizuo Liu, Shixiong Xia, Bing Liu 0016 |
Neurocomputing | 4 |
| 2015 | Multiple kernel dimensionality reduction via spectral regression and trace ratio maximization
Bing Liu 0016 |
Knowl. Based Syst. | 3 |
| 2013 | Unsupervised non-parametric kernel learning algorithm
Bing Liu 0016, Shixiong Xia, Yong Zhou 0003 |
Knowl. Based Syst. | 1 |