EDBT 2026 Demo / reviewers in the wild / expert
Binghao Liu
dblp:279/8039
· DBLP profile ↗
21ranked-venue papers
8as first author
21since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 3 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CCMamba: A Multi-Head Criss-Cross Mamba for Pavement Crack SegmentationabstractAutomatic and accurate detection of pavement cracks on highways and main roads is very important for intelligent pavement maintenance and traffic safety. However, due to the irregular shapes and uncertain characteristics of cracks, current methods face serious challenges on crack segmentation task. To solve the above issues, we propose a Multi-Head Criss-Cross Mamba (CCMamba) to efficiently enhance crack segmentation performance. Based on the sufficient details retained in low-level features and semantic information with strong discriminative ability in high-level features, CCMamba utilizes high-level features to extract crack information from low-level features. First, because of the feature representation capacity of latent state in State Space Model (SSM), we introduce a Multi-Head Latent State Module (Multi-Head LSM) to criss-cross study high-level local features and generate multiple Dynamic Convolution kernels. Second, these Dynamic Convolution kernels are applied to low-level features in a convolutional way, filtering crack information from massive background interferences. Third, the horizontal and vertical features output from Dynamic Convolution layers are fused with head attention mechanism, producing crack sensitive features. Finally, we use the obtained multi-scale features to predict segmentation masks. Comprehensive experiments on five public datasets, Crack500, GAPs384, CFD, CrackVision12K and CPRID, are conducted and our CCMamba achieves state-of-the-art (SOTA) performances compared to current crack segmentation methods. Meanwhile, ablation studies, visualization analysis and real scene testing also validate the effectiveness of CCMamba. Codes of this paper are public available athttps://github.com/cv516Buaa/BinghaoLiu/tree/main/CCMamba Binghao Liu, Qi Zhao 0037, Hongbo Xie, Hong Zhang 0018 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2025 | Prototype antithesis for biological few-shot class-incremental learningabstractDeep learning has become essential in the biological species recognition task. However, a significant challenge is the ability to continuously learn new or mutated species with limited annotated samples. Since species within the same family typically share similar traits, distinguishing between new and existing (old) species during incremental learning often faces the issue of species confusion. This can result in "catastrophic forgetting" of old species and poor learning of new ones. To address this issue, we propose a Prototype Antithesis (PA) method, which leverages the hierarchical structures in biological taxa to reduce confusion between new and old species. PA operates in two steps: Residual Prototype Learning (RPL) and Residual Prototype Mixing (RPM). RPL enables the model to learn unique prototypes for each species alongside residual prototypes representing shared traits within families. RPM generates synthetic samples by blending features of new species with residual prototypes of old species, encouraging the model to focus on species-unique traits and minimize species confusion. By integrating RPL and RPM, the proposed PA method mitigates "catastrophic forgetting" while improving generalization to new species. Extensive experiments on CUB200, PlantVillage, and Tree-of-Life datasets demonstrate that PA significantly reduces inter-species confusion and achieves state-of-the-art performance, highlighting its potential for deep learning in biological data analysis. Binghao Liu |
ICLR | 1 |
| 2025 | HFDNet: High-Frequency Divergence Attention Network for Underwater SegmentationabstractCurrently, most underwater operations are conducted in deep water, and there is usually insufficient illumination in these areas. At this time, the local texture features of some objects are highly similar in images, and it is difficult to distinguish the inter-class boundaries. This typically results in poor performance of the current semantic segmentation models of terrestrial images in underwater scenes. Taking advantage of the general characteristic that high-frequency regions are more likely to correspond to semantic segmentation boundaries, we introduce the high-frequency Divergence Attention Network (HFDNet), a semantic segmentation model based on transformer. HFDNet extracts its frequency distribution by analyzing the frequency domain of the feature map, and then calculates the relative spectral magnitude of each component by comparing its frequency amplitude against the average amplitude within its local neighborhood in the frequency domain. The local frequency map can be incorporated into the attention matrix as a weighting factor to realize the divergence of attention to the surrounding areas, which improves the attention to the high-frequency areas. This operation can enhance the model’s focus on the object boundary region and local neigh-borhood categories for each component. Therefore, our model can alleviate the problem of determining the object boundary caused by insufficient light in underwater image segmentation, and enhance the ability to segment objects with similar local features under low light conditions. We conduct comprehensive experiments on three underwater segmentation datasets: Caveseg, SUIM and UWS. The results show that our HFDNet achieves state-of-the-art (SOTA) performance on the testing datasets. The source code is available at https://github.com/cv516Buaa/HongboXie/tree/main/HFDNet. Hongbo Xie, Qi Zhao 0037, Binghao Liu |
IROS | 3 |
| 2025 | A Voronoi Density-Based Locally Unique Network for Fine-Grained Multi-Label ClassificationabstractMulti-label image classification aims to classify all categories in images simultaneously. When current multi-label classification methods meet fine-grained objects in a single image, the extreme inter-class similarity and over-prediction problems are two major challenges that hinder model performance. To solve the above two problems, we propose Voronoi density based Locally Unique Network (VoLUNet). First, due to high correlation between predictions of different classes, following the Kolmogorov-Arnold Network (KAN), we design the Weak Inter-class Correlation Classifier (WIC-Classifier) to replace linear weights setting in MLP architecture, promoting the potential of fine-grained discrimination. Second, we propose a Local Non-Maximum Suppression (Local-NMS) loss to multi-label classification model, predicting only one unique class with high prediction value for each local region. Third, different classes may have different pixel proportions and Local-NMS loss will be imbalanced for diverse fine-grained classes, we design the Voronoi Density based Superpixel Module (VDSM) to balance the quantities of local feature vectors with different classes. Finally, comprehensive experiments are conducted on four datasets, TreeSatAI, GeoLifeCLEF, FothemNet ShipRSImageNet, and our VoLUNet can significantly improve the classification performance compared to current state-of-the-art models. Codes of this paper are public available at https://github.com/cv516Buaa/BinghaoLiu/tree/main/VoLUNet. Binghao Liu, Qi Zhao 0037, Lijiang Chen |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | A Masked Reference Token Supervision-Based Iterative Visual-Language Framework for Robust Visual GroundingabstractVisual Grounding (VG) has become a prominent task in recent years, achieving significant advancements with the development of detection and vision transformers. However, existing VG methods struggle to handle the effects of inaccurate or irrelevant textual descriptions, tending to generate false-alarm objects. Moreover, existing methods fail to capture fine-grained features, accurate localization, and comprehensive context understanding from the whole image and textual descriptions. To address these issues, we propose an Iterative Robust Visual Grounding (IR-VG) framework with Multi-stage False-alarm Sensitive Decoder (MFSD) to prevent the generation of false-alarm objects when presented with inaccurate expressions. The framework introduces Masked Reference based Centerpoint Supervision (MRCS) and Iterative Multi-level Vision-language Fusion (IMVF) for enhancing the accuracy of localization and better visual-language alignment. To investigate the elements that affect VG robustness further, we release a robust VG benchmark with 24,000 instances and we also provide a detailed classification of false-alarm according to different parts of speech. Extensive experiments on existing state-of-the-art (SOTA) VG methods and foundation models have proven that it is difficult to handle the robustness of VG by existing models. Even foundation models, which have been pre-trained with a large amount of data, have difficulty to understand inaccurate language descriptions. Our IR-VG can handle false-alarm issues in robust VG well and achieve new SOTA results on the newly proposed robust VG datasets. Ablation studies and visualization experiments demonstrate the effectiveness of the proposed components. Moreover, the proposed framework is also verified effective on five regular VG datasets. Codes and models will be publicly athttps://github.com/cv516Buaa/IR-VG. Wenquan Feng, Shuchang Lyu, Xiangtai Li, Binghao Liu, Qi Zhao 0037 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | Self-Training and Curriculum Learning Guided Dynamic Refined Network for Remote Sensing Class-Incremental Semantic SegmentationabstractClass-incremental semantic segmentation aims to update the segmentation model with training samples containing only novel categories. Within this domain, catastrophic forgetting is a common challenge. In remote sensing scenes, images always have large discrepancies caused by a large variety of geographical objects. Therefore, besides catastrophic forgetting, there exists two additional primary challenges persist. The first one is a huge imbalance in image categories while the second one is error accumulation during multiple incremental training steps. To solve these three problems, we propose a new Self-Training and Curriculum Learning Guided Dynamic Refined Network (STCL-DRNet). Specifically, we first design a self-training-based branch to ease the tendency of catastrophic forgetting. We then design a dynamic refined loss to mitigate the uneven category distribution. Through further embedding class-balanced curriculum learning, we can alleviate the performance drop from noisy accumulation. Extensive experiments on benchmark datasets, DeepGLobe and iSAID, prove that the proposed STCL-DRNet achieves new SOTA performance. Visualization and analysis further substantiate the interpretability. Hongbo Zhao 0001, Ruimin Ren, Shuchang Lyu, Binghao Liu, Qi Zhao 0037 |
IGARSS | 4 |
| 2024 | OV-VG: A benchmark for open-vocabulary visual groundingabstractOpen-vocabulary learning has emerged as a cutting-edge research area, particularly in light of the widespread adoption of vision-based foundational models. Its primary objective is to comprehend novel concepts that are not encompassed within a predefined vocabulary. One key facet of this endeavor is Visual Grounding (VG), which entails locating a specific region within an image based on a corresponding language description. While current foundational models excel at various visual language tasks, there is a noticeable absence of models specifically tailored for open-vocabulary visual grounding (OV-VG). This research endeavor introduces novel and challenging OV tasks, namely Open-Vocabulary Visual Grounding (OV-VG) and Open-Vocabulary Phrase Localization (OV-PL). The overarching aim is to establish connections between language descriptions and the localization of novel objects. To facilitate this, we have curated a comprehensive annotated benchmark, encompassing 7272 OV-VG images (comprising 10,000 instances) and 1000 OV-PL images. In our pursuit of addressing these challenges, we delved into various baseline methodologies rooted in existing open-vocabulary object detection (OV-D), VG, and phrase localization (PL) frameworks. Surprisingly, we discovered that state-of-the-art (SOTA) methods often falter in diverse scenarios. Consequently, we developed a novel framework that integrates two critical components: Text-Image Query Selection (TIQS) and Language-Guided Feature Attention (LGFA). These modules are designed to bolster the recognition of novel categories and enhance the alignment between visual and linguistic information. Extensive experiments demonstrate the efficacy of our proposed framework, which consistently attains SOTA performance across the OV-VG task. Additionally, ablation studies provide further evidence of the effectiveness of our innovative models. Codes and datasets will be made publicly available at https://github.com/cv516Buaa/OV-VG. Wenquan Feng, Xiangtai Li, Shuchang Lyu, Binghao Liu, Lijiang Chen, Qi Zhao 0037 |
Neurocomputing | 6 |
| 2024 | Improving Scarce RS Data Classification With Independent Noise and Feature Mutual ExclusionabstractData shortage and real-time demand are two essential characteristics of remote sensing signal processing tasks. Existing works rarely focus on the scenario where all categories data are relatively scarce, which may be necessary in some occasions. To further explore the potential of light-weight deep neural network for scarce data remote sensing image classification, we attempt to improve the model performance in such task with tiny proportion of the training data. Due to the information insufficiency caused by data shortage, we propose a statistically independent Gaussian noise based feature augmentation module. A variational auto-encoder is used to provide noise vectors which come from the same input space with feature vectors for classification, but independent from them. To alleviate the inter-class feature confusion caused by the feature augmentation and enhance decision boundaries, we design a 3Σ1/2area based inter-class mutual exclusion strategy to enlarge distances between samples of different classes in feature space with contrastive loss. Extensive experiments are conducted to prove that our method can significantly improve the performance of light-weight models on scarce data remote sensing classification task. Binghao Liu, Yidi Bai |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2024 | Overcomplete-to-sparse representation learning for few-shot class-incremental learning
Mengying Fu, Binghao Liu, Tianren Ma, Qixiang Ye |
Multim. Syst. | 2 |
| 2024 | LDC-PP-YOLOE: a lightweight model for detecting and counting citrus fruit
Yibo Lv, Shenglian Lu, Jiangchuan Bao, Binghao Liu |
Pattern Anal. Appl. | 5 |
| 2023 | Enhancing Spatial Consistency and Class-Level Diversity for Segmenting Fine-Grained Objects
Qi Zhao 0037, Binghao Liu, Shuchang Lyu, Yifan Yang 0003 |
ICONIP (11) | 2 |
| 2023 | Learnable Distribution Calibration for Few-Shot Class-Incremental LearningabstractFew-shot class-incremental learning (FSCIL) faces the challenges of memorizing old class distributions and estimating new class distributions given few training samples. In this study, we propose a learnable distribution calibration (LDC) approach, to systematically solve these two challenges using a unified framework. LDC is built upon a parameterized calibration unit (PCU), which initializes biased distributions for all classes based on classifier vectors (memory-free) and a single covariance matrix. The covariance matrix is shared by all classes, so that the memory costs are fixed. During base training, PCU is endowed with the ability to calibrate biased distributions by recurrently updating sampled features under supervision of real distributions. During incremental learning, PCU recovers distributions for old classes to avoid 'forgetting', as well as estimating distributions and augmenting samples for new classes to alleviate 'over-fitting' caused by the biased distributions of few-shot samples. LDC is theoretically plausible by formatting a variational inference procedure. It improves FSCIL's flexibility as the training procedure requires no class similarity priori. Experiments on CUB200, CIFAR100, and mini-ImageNet datasets show that LDC respectively outperforms the state-of-the-arts by 4.64%, 1.98%, and 3.97%. LDC's effectiveness is also validated on few-shot learning scenarios. Binghao Liu, Boyu Yang 0002, Lingxi Xie, Qi Tian 0001, Qixiang Ye |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | Dynamic Support Network for Few-Shot Class Incremental LearningabstractFew-shot class-incremental learning (FSCIL) is challenged by catastrophically forgetting old classes and over-fitting new classes. Revealed by our analyses, the problems are caused by feature distribution crumbling, which leads to class confusion when continuously embedding few samples to a fixed feature space. In this study, we propose a Dynamic Support Network (DSN), which refers to an adaptively updating network with compressive node expansion to "support" the feature space. In each training session, DSN tentatively expands network nodes to enlarge feature representation capacity for incremental classes. It then dynamically compresses the expanded network by node self-activation to pursue compact feature representation, which alleviates over-fitting. Simultaneously, DSN selectively recalls old class distributions during incremental learning to support feature distributions and avoid confusion between classes. DSN with compressive node expansion and class distribution recalling provides a systematic solution for the problems of catastrophic forgetting and overfitting. Experiments on CUB, CIFAR-100, and miniImage datasets show that DSN significantly improves upon the baseline approach, achieving new state-of-the-arts. Boyu Yang 0002, Mingbao Lin, Binghao Liu, Xiaodan Liang, Rongrong Ji, Qixiang Ye |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Learn by Oneself: Exploiting Weight-Sharing Potential in Knowledge Distillation Guided Ensemble NetworkabstractRecent CNNs (convolutional neural networks) have become more and more compact. The elegant structure design highly improves the performance of CNNs. With the development of knowledge distillation technique, the performance of CNNs gets further improved. However, existing knowledge distillation guided methods either rely on offline pretrained high-quality large teacher models or online heavy training burden. To solve the above problems, we propose a feature-sharing and weight-sharing based ensemble network (training framework) guided by knowledge distillation (EKD-FWSNet) to make baseline models stronger in terms of representation ability with less training computation and memory cost involved. Specifically, motivated by getting rid of the dependence of offline pretrained teacher model, we design an end-to-end online training scheme to optimize EKD-FWSNet. Motivated by decreasing the online training burden, we only introduce one auxiliary classmate branch to construct multiple forward branches, which will then be integrated as ensemble teacher to guide baseline model. Compared to previous online ensemble training frameworks, EKD-FWSNet can provide diverse output predictions without relying on increasing auxiliary classmate branches. Motivated by maximizing the optimization power of EKD-FWSNet, we exploit the representation potential of weight-sharing blocks and design efficient knowledge distillation mechanism in EKD-FWSNet. Extensive comparison experiments and visualization analysis on benchmark datasets (CIFAR-10/100, tiny-ImageNet, CUB-200 and ImageNet) show that self-learned EKD-FWSNet can boost the performance of baseline models by large margin, which has obvious superiority compared to previous related methods. Extensive analysis also proves the interpretability of EKD-FWSNet. Our code is available at https://github.com/cv516Buaa/EKD-FWSNet. Qi Zhao 0037, Shuchang Lyu, Lijiang Chen, Binghao Liu, Ting-Bing Xu, Wenquan Feng |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | A Similarity Distillation Guided Feature Refinement Network for Few-Shot Semantic SegmentationabstractFew-shot semantic segmentation is a challenging task of predicting object categories in pixel-wise with only few annotated samples. Existing methods mainly have two problems, which are representation inconsistency of query and support images and semantic-level feature insufficient. To tackle the two problems, we propose a similarity distillation guided feature refinement network (SD-FRNet). Specifically, we first use support label to generate support similarity feature map and coarse prediction of query image. Then, we use this coarse prediction to generate query similarity feature. To compensate feature representation inconsistency, we conduct knowledge distillation mechanism to align similarity features of query and support images. To enrich semantic-level feature, we further design a feature refinement module, which achieves high-quality segmentation. Extensive experiments show the effectiveness of SD-FRNet. On benchmark datasets, PASCAL-5iand COCO-20i, our proposed SD-FRNet outperform the previous SOTA (state-of-the-art) results. Shuchang Lyu, Binghao Liu, Lijiang Chen, Qi Zhao 0037 |
ICIP | 2 |
| 2022 | ML-FDA: Meta-Learning via Feature Distribution Alignment for Few-Shot LearningabstractComputer vision tasks suffer from the high cost of collecting large amounts of labeled data. Few-shot Learning (FSL) is a dominant approach to solve this problem because it provides an insight to learn the knowledge of novel categories with few training samples. In FSL task, Meta-learning and metric learning have achieved impressive results. However, the performance of this task is still limited by large intra-class variance and small inter-class distance caused by limited number of few samples. To solve this problem, In this paper, we propose a new method, which integrates meta-learning and metric learning techniques. Specifically, we first propose a feature representation module (FR) to construct representative support class prototypes and query features. Then, we design bias loss to minimize the bias between support and query samples. Furthermore, we design an intra-class loss to minimize the distance between query class prototype and each query sample. We denote this model as ML-FDA and validate it on standard few-shot classification benchmark datasets (MiniImageNet, CIFAR-FS, FC100). The results show that our method improves the performance over other same paradigm methods and achieves the best performance on most benchmarks. The ablation study and visulization analysis also demonstrate the effectiveness of our method. Binghao Liu, Shuchang Lyu, Lijiang Chen, Qi Zhao 0037, Wenquan Feng |
VCIP | 2 |
| 2022 | CitrusYOLO: A Algorithm for Citrus Detection under Orchard Environment Based on YOLOv4
Wenkang Chen, Shenglian Lu, Binghao Liu, Tingting Qian |
Multim. Tools Appl. | 3 |
| 2022 | A feature consistency driven attention erasing network for fine-grained image retrieval
Qi Zhao 0037, Shuchang Lyu, Binghao Liu, Yifan Yang 0003 |
Pattern Recognit. | 4 |
| 2021 | Anti-Aliasing Semantic Reconstruction for Few-Shot Semantic SegmentationabstractEncouraging progress in few-shot semantic segmentation has been made by leveraging features learned upon base classes with sufficient training data to represent novel classes with few-shot examples. However, this feature sharing mechanism inevitably causes semantic aliasing between novel classes when they have similar compositions of semantic concepts. In this paper, we reformulate few-shot segmentation as a semantic reconstruction problem, and convert base class features into a series of basis vectors which span a class-level semantic space for novel class reconstruction. By introducing contrastive loss, we maximize the orthogonality of basis vectors while minimizing semantic aliasing between classes. Within the reconstructed representation space, we further suppress interference from other classes by projecting query features to the support vector for precise semantic activation. Our proposed approach, referred to as anti-aliasing semantic reconstruction (ASR), provides a systematic yet interpretable solution for few-shot learning problems. Extensive experiments on PASCAL VOC and MS COCO datasets show that ASR achieves strong results compared with the prior works. Code will be released at github.com/Bibkiller/ASR. Binghao Liu, Yao Ding 0006, Jianbin Jiao, Xiangyang Ji, Qixiang Ye |
CVPR | 1 |
| 2021 | Knowledge-Aware Hypergraph Neural Network for Recommender Systems
Binghao Liu, Pengpeng Zhao 0001, Fuzhen Zhuang, Xuefeng Xian, Yanchi Liu, Victor S. Sheng |
DASFAA (3) | 1 |
| 2021 | Harmonic Feature Activation for Few-Shot Semantic SegmentationabstractFew-shot semantic segmentation remains an open problem because limited support (training) images are insufficient to represent the diverse semantics within target categories. Conventional methods typically model a target category solely using information from the support image(s), resulting in incomplete semantic activation. In this paper, we propose a novel few-shot segmentation approach, termed harmonic feature activation (HFA), with the aim to implement dense support-to-query semantic transform by incorporating the features of both query and support images. HFA is formulated as a bilinear model, which takes charge of the pixel-wise dense correlation (bilinear feature activation) between query and support images in a systematic way. HFA incorporates a low-rank decomposition procedure, which speeds up bilinear feature activation with negligible performance cost. In addition, a semantic diffusion procedure is fused with HFA, which further improves the global harmony and local consistency of the feature activation. Extensive experiments on commonly used datasets (PASCAL VOC and MS COCO) show that HFA improves the state-of-the-arts with significant margins. Code is available at https://github.com/Bibikiller/HFA. Binghao Liu, Jianbin Jiao, Qixiang Ye |
IEEE Trans. Image Process. | 1 |