EDBT 2026 Demo / reviewers in the wild / expert
Huafeng Liu 0004
dblp:48/4950-4
· DBLP profile ↗
18ranked-venue papers
6as first author
14since 2021 · last 2026
0000-0001-5396-3183ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 15 · 6 first-author · 11 since 2021Artificial intelligence and machine learning · 6 · 5 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Monocular 3D lane detection with geometry-guided transformation and contextual enhancement
Chunying Song, Qiong Wang 0003, Zeren Sun, Huafeng Liu 0004 |
Pattern Recognit. Lett. | 4 |
| 2025 | UNIALIGN: Scaling Multimodal Alignment within One Unified ModelabstractWe present UniAlign, a unified model to align an arbitrary number of modalities (e.g., image, text, audio, 3D point cloud, etc.) through one encoder and a single training phase. Existing solutions typically employ distinct encoders for each modality, resulting in increased parameters as the number of modalities grows. In contrast, UniAlign proposes a modality-aware adaptation of the powerful mixture- of-experts (MoE) schema and further integrates it with Low- Rank Adaptation (LoRA), efficiently scaling the encoder to accommodate inputs in diverse modalities while maintaining a fixed computational overhead. Moreover, prior work often requires separate training for each extended modality. This leads to task-specific models and further hinders the communication between modalities. To address this, we propose a soft modality binding strategy that aligns all modalities using unpaired data samples across datasets. Two additional training objectives are introduced to distill knowledge from well-aligned anchor modalities and prior multimodal models, elevating UniAlign into a high performance multimodal foundation model. Experiments on 11 benchmarks across 6 different modalities demonstrate that UniAlign could achieve comparable performance to SOTA approaches, while using merely 7.8M trainable parameters and maintaining an identical model with the same weight across all tasks. Liulei Li, Huafeng Liu 0004, Yazhou Yao, Wenguan Wang |
CVPR | 4 |
| 2025 | DynaPlane-Lane: Dynamic Multi-Plane Geometry Learning for Robust Monocular 3D Lane DetectionabstractAccurate monocular 3D lane detection remains a fundamental challenge in autonomous driving, due to the need to infer spatial geometry from single-view images under diverse road and environmental conditions. Existing approaches often struggle with limited geometric adaptability and inconsistent structural predictions, particularly in the presence of sloped or uneven road surfaces. In this paper, we propose DynaPlane-Lane, a geometry-aware framework that explicitly models road topology through a dynamic multi-plane representation. By partitioning the scene into near, transition, and far regions with learnable geometric transformations, our method effectively captures height changes and surface irregularities. To enhance spatial reasoning, we design a Multi-Scale Adaptive Feature Aggregation module with channel and spatial attention to fuse encoder features across multiple resolutions. Additionally, a Topology-Aware Lane Optimization objective is formulated to enforce continuity and smoothness of predicted lane curves under partial visibility. Comprehensive experiments on the OpenLane and Apollo 3D Synthetic benchmarks demonstrate that our approach achieves robust performance across diverse scenarios, validating the effectiveness of incorporating adaptive geometry and structural constraints into monocular 3D lane detection. Chunying Song, Huafeng Liu 0004, Tao Chen 0012, Qiong Wang 0003 |
MMAsia | 2 |
| 2025 | UncertainBEV: Uncertainty-aware BEV fusion for roadside 3D object detection
Chunying Song, Huafeng Liu 0004, Qiong Wang 0003 |
Image Vis. Comput. | 4 |
| 2025 | Semi-Supervised Semantic Segmentation With Multi-Constraint Consistency LearningabstractConsistency regularization has prevailed in semi-supervised semantic segmentation and achieved promising performance. However, existing methods typically concentrate on enhancing the Image-augmentation based Prediction consistency and optimizing the segmentation network as a whole, resulting in insufficient utilization of potential supervisory information. In this paper, we propose a Multi-Constraint Consistency Learning (MCCL) approach to facilitate the staged enhancement of the encoder and decoder. Specifically, we first design a feature knowledge alignment (FKA) strategy to promote the feature consistency learning of the encoder from image-augmentation. Our FKA encourages the encoder to derive consistent features for strongly and weakly augmented views from the perspectives of point-to-point alignment and prototype-based intra-class compactness. Moreover, we propose a self-adaptive intervention (SAI) module to increase the discrepancy of aligned intermediate feature representations, promoting Feature-perturbation based Prediction consistency learning. Self-adaptive feature masking and noise injection are designed in an instance-specific manner to perturb the features for robust learning of the decoder. Experimental results on Pascal VOC2012 and Cityscapes datasets demonstrate that our proposed MCCL achieves new state-of-the-art performance. The source code and models are made available athttps://github.com/NUST-Machine-Intelligence-Laboratory/MCCL. Jianjian Yin, Tao Chen 0012, Gensheng Pei, Huafeng Liu 0004, Yazhou Yao, Liqiang Nie, Xian-Sheng Hua 0001 |
IEEE Trans. Multim. | 4 |
| 2024 | VideoMAC: Video Masked Autoencoders Meet ConvNetsabstractRecently, the advancement of self-supervised learning techniques, like masked autoencoders (MAE), has greatly influenced visual representation learning for images and videos. Nevertheless, it is worth noting that the predomi-nant approaches in existing masked image / video modeling rely excessively on resource-intensive vision transformers (ViTs) as the feature encoder. In this paper, we propose a new approach termed as VideoMAC, which combines video masked autoencoders with resource-friendly Con-vNets. Specifically, VideoMAC employs symmetric masking on randomly sampled pairs of video frames. To prevent the issue of mask pattern dissipation, we utilize ConvNets which are implemented with sparse convolutional operators as en-coders. Simultaneously, we present a simple yet effective masked video modeling (MVM) approach, a dual encoder architecture comprising an online encoder and an exponential moving average target encoder, aimed to facilitate inter-frame reconstruction consistency in videos. Additionally, we demonstrate that VideoMAC, empowering classical (ResNet) / modern (ConvNeXt) convolutional encoders to harness the benefits of MVM, outperforms ViT-based approaches on downstream tasks, including video object segmentation (+5.2% /6.4% J&F), body part propagation (+6.3% /3.1% mIoU), and human pose tracking (+10.2% / 11.1% [email protected]). Gensheng Pei, Tao Chen 0012, Xiruo Jiang, Huafeng Liu 0004, Zeren Sun, Yazhou Yao |
CVPR | 4 |
| 2024 | SMP-Track: SAM in Multi-Pedestrian TrackingabstractMultiple Object Tracking (MOT) plays a crucial role in security data analysis as a fundamental problem for video surveillance. Our goal is to design a robust tracker for data damaged by attacks, while also emphasizing privacy protection. The mainstream paradigm for MOT is tracking-by-detection (TBD), which involves object detection followed by target association. In the association stage, most models rely on Intersection over Union (IoU) similarity of bounding boxes for short-range matching and cosine similarity of appearance features for long-range matching. However, both of these similarities contain a lot of redundant background regions except the target. To this end, we propose a new tracker named SMP-Track that integrates Segment Anything Model (SAM) into Multi-Pedestrian tracking method. Firstly we extract the pedestrian masks based on box prompt, focusing solely on the foreground information. Then we introduce a new similarity metric that combines the advantages of motion and foreground information (i.e., box-mask similarity). Extensive experiments demonstrate that SMP-Track increases main metrics on the MOT17 validation set, and achieves comparable performance to other state-of-the-art methods on the MOT17 and MOT20 test sets. Furthermore, by incorporating pedestrian masks, we reduce reliance on raw pedestrian images or features, making the model robust to corrupted data and mitigating the risk of privacy leakage. Shiyin Wang, Huafeng Liu 0004, Qiong Wang 0003, Yazhou Yao |
DSAA | 2 |
| 2024 | Universal Organizer of Segment Anything Model for Unsupervised Semantic SegmentationabstractUnsupervised semantic segmentation (USS) aims to achieve high-quality segmentation without manual pixel-level annotations. Existing USS models provide coarse category classifi-cation for regions, but the results often have blurry and imprecise edges. Recently, a robust framework called the segment anything model (SAM) has been proven to deliver precise boundary object masks. Therefore, this paper proposes a universal organizer based on SAM, termed as UO-SAM, to enhance the mask quality of USS models. Specifically, using only the original image and the masks generated by the USS model, we extract visual features to obtain positional prompts for target objects. Then, we activate a local region optimizer that performs segmentation using SAM on a per-object basis. Finally, we employ a global region optimizer to incorporate global image information and refine the masks to obtain the final fine-grained masks. Compared to existing methods, our UO-SAM achieves state-of-the-art performance. Our codes are available at https://github.com/NUST-Machine-Intelligence-Laboratory/UO-SAM. Gensheng Pei, Xinhao Cai, Qiong Wang 0003, Huafeng Liu 0004, Yazhou Yao |
ICME | 5 |
| 2024 | Learning With Imbalanced Noisy Data by Preventing Bias in Sample SelectionabstractLearning with noisy labels has gained increasing attention because the inevitable imperfect labels in real-world scenarios can substantially hurt the deep model performance. Recent studies tend to regard low-loss samples as clean ones and discard high-loss ones to alleviate the negative impact of noisy labels. However, real-world datasets contain not only noisy labels but also class imbalance. The imbalance issue is prone to causing failure in the loss-based sample selection since the under-learning of tail classes also leans to produce high losses. To this end, we propose a simple yet effective method to address noisy labels in imbalanced datasets. Specifically, we proposeClass-Balance-based sampleSelection (CBS) to prevent the tail class samples from being neglected during training. We proposeConfidence-basedSampleAugmentation (CSA) for the chosen clean samples to enhance their reliability in the training process. To exploit selected noisy samples, we resort to prediction history to rectify labels of noisy samples. Moreover, we introduce theAverageConfidenceMargin (ACM) metric to measure the quality of corrected labels by leveraging the model's evolving training dynamics, thereby ensuring that low-quality corrected noisy samples are appropriately masked out. Lastly, consistency regularization is imposed on filtered label-corrected noisy samples to boost model performance. Comprehensive experimental results on synthetic and real-world datasets demonstrate the effectiveness and superiority of our proposed method, especially in imbalanced scenarios. The source code has been made available athttps://github.com/NUST-Machine-Intelligence-Laboratory/CBS. Huafeng Liu 0004, Mengmeng Sheng, Zeren Sun, Yazhou Yao, Xian-Sheng Hua 0001, Heng Tao Shen |
IEEE Trans. Multim. | 1 |
| 2023 | Co-mining: Mining informative samples with noisy labels
Zhenhuang Cai, Huafeng Liu 0004, Yazhou Yao, Zhenmin Tang |
Signal Process. | 2 |
| 2023 | FECANet: Boosting Few-Shot Semantic Segmentation With Feature-Enhanced Context-Aware NetworkabstractFew-shot semantic segmentation is the task of learning to locate each pixel of the novel class in the query image with only a few annotated support images. The current correlation-based methods construct pair-wise feature correlations to establish the many-to-many matching because the typical prototype-based approaches cannot learn fine-grained correspondence relations. However, the existing methods still suffer from the noise contained in naive correlations and the lack of context semantic information in correlations. To alleviate these problems mentioned above, we propose a Feature-Enhanced Context-Aware Network (FECANet). Specifically, a feature enhancement module is proposed to suppress the matching noise caused by inter-class local similarity and enhance the intra-class relevance in the naive correlation. In addition, we propose a novel correlation reconstruction module that encodes extra correspondence relations between foreground and background and multi-scale context semantic features, significantly boosting the encoder to capture a reliable matching pattern. Experiments on PASCAL-$5^{i}$and COCO-$20^{i}$datasets demonstrate that our proposed FECANet leads to remarkable improvement compared to previous state-of-the-arts, demonstrating its effectiveness. The source codes and models have been made available athttps://github.com/NUST-Machine-Intelligence-Laboratory/FECANET. Huafeng Liu 0004, Tao Chen 0012, Qiong Wang 0003, Yazhou Yao, Xian-Sheng Hua 0001 |
IEEE Trans. Multim. | 1 |
| 2022 | Exploiting Web Images for Fine-Grained Visual Recognition via Dynamic Loss Correction and Global Sample SelectionabstractTo distinguish subtle differences among fine-grained categories, a large amount of well-labeled images are typically required. However, acquiring manual annotations for fine-grained categories is an extremely difficult task as it usually has a high demand for professional knowledge. To this end, directly leveraging web images for learning fine-grained models becomes a natural choice. Nevertheless, due to the existence of label noise, this learning paradigm tends to have a poor performance. In this work, we propose an end-to-end approach by combining dynamic loss correction and global sample selection to alleviate the problem of label noise. Specifically, we leverage the network to predict all samples, record the predictions of recent several epochs, and calculate the uncertainly-based dynamic loss for global sample selection. Extensive experiments on three benchmark datasets demonstrate the effectiveness of our proposed approach. The source code of our approach has been released on the website:https://github.com/NUST-Machine-Intelligence-Laboratory/dlc. Huafeng Liu 0004, Haofeng Zhang 0001, Jianfeng Lu 0003, Zhenmin Tang |
IEEE Trans. Multim. | 1 |
| 2022 | Exploiting Web Images for Fine-Grained Visual Recognition by Eliminating Open-Set Noise and Utilizing Hard ExamplesabstractLabeling objects at a subordinate level typically requires expert knowledge, which is not always available when using random annotators. As such, learning directly from web images for fine-grained recognition has attracted broad attention. However, the presence of label noise and hard examples in web images are two obstacles for training robust fine-grained recognition models. Therefore, in this paper, we propose a novel approach for removing irrelevant samples from real-world web images during training, while employing useful hard examples to update the network. Thus, our approach can alleviate the harmful effects of irrelevant noisy web images and hard examples to achieve better performance. Extensive experiments on three commonly used fine-grained datasets demonstrate that our approach is far superior to current state-of-the-art web-supervised methods. The data and source code of this work have been made publicly available at:https://github.com/NUST-Machine-Intelligence-Laboratory/Advanced-Softly-Update-Drop. Huafeng Liu 0004, Chuanyi Zhang, Yazhou Yao, Xiu-Shen Wei, Fumin Shen, Zhenmin Tang, Jian Zhang 0002 |
IEEE Trans. Multim. | 1 |
| 2022 | Co-LDL: A Co-Training-Based Label Distribution Learning Method for Tackling Label NoiseabstractPerformances of deep neural networks are prone to be degraded by label noise due to their powerful capability in fitting training data. Deeming low-loss instances as clean data is one of the most promising strategies in tackling label noise and has been widely adopted by state-of-the-art methods. However, prior works tend to drop high-loss instances directly, neglecting their valuable information. To address this issue, we propose an end-to-end framework named Co-LDL, which incorporates the low-loss sample selection strategy with label distribution learning. Specifically, we simultaneously train two deep neural networks and let them communicate useful knowledge by selecting low-loss and high-loss samples for each other. Low-loss samples are leveraged conventionally for updating network parameters. On the contrary, high-loss samples are trained in a label distribution learning manner to update network parameters and label distributions concurrently. Moreover, we propose a self-supervised module to further boost the model performance by enhancing the learned representations. Comprehensive experiments on both synthetic and real-world noisy datasets are provided to demonstrate the superiority of our Co-LDL method over state-of-the-art approaches in learning with noisy labels. The source code and models have been made available athttps://github.com/NUST-Machine-Intelligence-Laboratory/CoLDL. Zeren Sun, Huafeng Liu 0004, Qiong Wang 0003, Tianfei Zhou, Qi Wu 0001, Zhenmin Tang |
IEEE Trans. Multim. | 2 |
| 2020 | Web-Supervised Network with Softly Update-Drop Training for Fine-Grained Visual ClassificationabstractLabeling objects at the subordinate level typically requires expert knowledge, which is not always available from a random annotator. Accordingly, learning directly from web images for fine-grained visual classification (FGVC) has attracted broad attention. However, the existence of noise in web images is a huge obstacle for training robust deep neural networks. In this paper, we propose a novel approach to remove irrelevant samples from the real-world web images during training, and only utilize useful images for updating the networks. Thus, our network can alleviate the harmful effects caused by irrelevant noisy web images to achieve better performance. Extensive experiments on three commonly used fine-grained datasets demonstrate that our approach is much superior to state-of-the-art webly supervised methods. The data and source code of this work have been made anonymously available at: https://github.com/z337-408/WSNFGVC. Chuanyi Zhang, Yazhou Yao, Huafeng Liu 0004, Guosen Xie, Xiangbo Shu, Tianfei Zhou, Zheng Zhang 0006, Fumin Shen, Zhenmin Tang |
AAAI | 3 |
| 2020 | Hsi Road: A Hyper Spectral Image Dataset For Road SegmentationabstractRoad segmentation is a challenging task in the field of self-driving research. This paper present a road dataset built by hyper spectral imaging (HSI) cameras instead of the widely-used RGB cameras. HSI image is informative in spectrums and full of potential for natural environment perception. In this article, a first-of-its-kind HSI road segmentation dataset is built with careful annotation in both urban and rural scenes. It contains 3799 scenes with RGB and NIR bands as well as their respective masks. Unlike many existing datasets that provide urban scenes in RGB images only, our dataset expands the sensing spectrum to 28 bands and includes various kinds of road surfaces, such as asphalt, cement, dirt and sand, under rural and natural scenes. We also provide benchmark performances based on the recently popular segmentation algorithms on this dataset. The dataset is released at github‡.‡https://github.com/NUST-Machine-Intelligence-Laboratory/hsi_road Jiarou Lu, Huafeng Liu 0004, Yazhou Yao, Shuyin Tao, Zhenmin Tang, Jianfeng Lu 0003 |
ICME | 2 |
| 2020 | Road segmentation with image-LiDAR data fusion in deep neural network
Huafeng Liu 0004, Yazhou Yao, Zeren Sun, Xiangrui Li, Ke Jia, Zhenmin Tang |
Multim. Tools Appl. | 1 |
| 2019 | Deep representation learning for road detection using Siamese network
Huafeng Liu 0004, Xiangrui Li, Yazhou Yao, Zhenmin Tang |
Multim. Tools Appl. | 1 |