Hongmin Liu 0001

dblp:28/7697-1 · DBLP profile ↗
← Back
69ranked-venue papers
14as first author
50since 2021 · last 2026
0000-0001-9834-4087ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 41 · 9 first-author · 28 since 2021Graphics, computer vision, multimedia, augmented reality and games · 31 · 5 first-author · 25 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 Group Orthogonal Low-Rank Adaptation for RGB-T Tracking
abstract
Parameter-efficient fine-tuning has emerged as a promising paradigm in RGB-T tracking, enabling downstream task adaptation by freezing pretrained parameters and fine-tuning only a small set of parameters. This set forms a rank space made up of multiple individual ranks, whose expressiveness directly shapes the model's adaptability. However, quantitative analysis reveals low-rank adaptation exhibits significant redundancy in the rank space, with many ranks contributing almost no practical information. This hinders the model's ability to learn more diverse knowledge to address the various challenges in RGB-T tracking. To address this issue, we propose the Group Orthogonal Low-Rank Adaptation (GOLA) framework for RGB-T tracking, which effectively leverages the rank space through structured parameter learning. Specifically, we adopt a rank decomposition partitioning strategy utilizing singular value decomposition to quantify rank importance, freeze crucial ranks to preserve the pretrained priors, and cluster the redundant ranks into groups to prepare for subsequent orthogonal constraints. We further design an inter-group orthogonal constraint strategy. This constraint enforces orthogonality between rank groups, compelling them to learn complementary features that target diverse challenges, thereby alleviating information redundancy. Experimental results demonstrate that GOLA effectively reduces parameter redundancy and enhances feature representation capabilities, significantly outperforming state-of-the-art methods across four benchmark datasets and validating its effectiveness in RGB-T tracking tasks.
Zekai Shao 0002, Yufan Hu, Bin Fan 0001, Hongmin Liu 0001
AAAI5
2026 Towards Accurate 3D Object Detection in Adverse Weather by Leveraging 4D Radar for LiDAR Geometry Enhancement
abstract
3D object detection is a critical component of autonomous driving, yet its performance degrades severely in adverse weather due to the degradation of LiDAR point clouds. While existing LiDAR-4D radar fusion methods enhance robustness by incorporating weather-robust 4D radar data, they often depend on well geometric structures from LiDAR and so struggle to effectively exploit radar data in case of degraded LiDAR data. To tackle this challenge, we propose REL, a novel 4D radar-guided LiDAR geometric enhancement framework. It utilizes 4D radar features to dynamically generate virtual LiDAR points, effectively increasing the density of degraded LiDAR data. Moreover, a Position-Guided Cross Attention (PGCA) module is proposed to enhance the feature representation of virtual points, while an Adaptive Feature Fusion (AFF) module is designed to integrate virtual and real LiDAR features. Extensive experiments on the K-Radar and Vod-Fog datasets demonstrate that REL achieves state-of-the-art 3D object detection performance under diverse adverse weather conditions. Notably, REL improves the overall AP3D by 9.3% on K-Radar and boosts the cyclist class by up to 52.9% 3D mAP under the most severe foggy condition on Vod-Fog.
Tianxu Tong, Xinrun Liu, Hongmin Liu 0001, Bin Fan 0001
AAAI3
2026 Offset-driven estimation and correction for nanosecond-scale laser-gated image restoration
Linbing He, Pengpeng Pi, Gengchen Zhang, Hongmin Liu 0001
Expert Syst. Appl.6
2026 Advancing Vision Transformer With Enhanced Spatial Priors
abstract
In recent years, the Vision Transformer (ViT) has garnered significant attention within the computer vision community. However, the core component of ViT, Self-Attention, lacks explicit spatial priors and suffers from quadratic computational complexity, limiting its applicability. To address these issues, we have proposed RMT, a robust vision backbone with explicit spatial priors for general purposes. RMT utilizes Manhattan distance decay to introduce spatial information and employs a horizontal and vertical decomposition attention method to model global information. Building on the strengths of RMT, Euclidean enhanced Vision Transformer (EVT) is an expanded version that incorporates several key improvements. Firstly, EVT uses a more reasonable Euclidean distance decay to enhance the modeling of spatial information, allowing for a more accurate representation of spatial relationships compared to the Manhattan distance used in RMT. Secondly, EVT abandons the decomposed attention mechanism featured in RMT and instead adopts a simpler spatially-independent grouping approach, providing the model with greater flexibility in controlling the number of tokens within each group. By addressing these modifications, EVT offers a more sophisticated and adaptable approach to incorporating spatial priors into the Self-Attention mechanism, thus overcoming some of the limitations associated with RMT and further enhancing its applicability in various computer vision tasks. Extensive experiments on Image Classification, Object Detection, Instance Segmentation, and Semantic Segmentation demonstrate that EVT exhibits exceptional performance. Without additional training data, EVT achieves 86.6% top1-acc on ImageNet-1 k.
Qihang Fan, Huaibo Huang, Mingrui Chen 0001, Hongmin Liu 0001, Ran He 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2026 Unifying RGB and thermal object detection in one detector
Bin Fan 0001, Wei Zhao 0019, Yongjie Chen, Hongmin Liu 0001
Pattern Recognit.5
2026 IIFNet3D: Instance-to-instance fusion with dual attention for indoor RGB-D 3D object detection
Zixin Fan, Bin Fan 0001, Hongmin Liu 0001
Pattern Recognit.4
2026 OV-GT3D: A generalizable open-vocabulary two-stage 3D detector with dual path distillation
Xiuwei Xu, Bin Fan 0001, Jiwen Lu, Hongmin Liu 0001
Pattern Recognit.5
2026 MC-MVSNet: When multi-view stereo meets monocular cues
Xincheng Tang, Mengqi Rong, Bin Fan 0001, Hongmin Liu 0001, Shuhan Shen
Pattern Recognit.4
2025 PURA: Parameter Update-Recovery Test-Time Adaption for RGB-T Tracking
abstract
Maintaining stable tracking of objects in domain shift scenarios is crucial for RGB-T tracking, prompting us to explore the use of unlabeled test sample information for effective online model adaptation. However, current Test-Time Adaptation (TTA) methods in RGB-T tracking dramatically change the model’s internal parameters during long-term adaptation. At the same time, the gradient computations involved in the optimization process impose a significant computational burden. To address these challenges, we propose a Parameter Update-Recovery Adaptation (PURA) framework based on parameter decomposition. Firstly, our fast parameter update strategy adjusts model parameters using statistical information from test samples without requiring gradient calculations, ensuring consistency between the model and test data distribution. Secondly, our parameter decomposition recovery employs orthogonal decomposition to identify the principal update direction and recover parameters in this direction, aiding in the retention of critical knowledge. Finally, we leverage the information obtained from decomposition to provide feedback on the momentum during the update phase, ensuring a stable updating process. Experimental results demonstrate that PURA outperforms current state-of-the-art methods across multiple datasets, validating its effectiveness. The project page is at https://melantech.github.io/PURA.
Zekai Shao 0002, Yufan Hu, Bin Fan 0001, Hongmin Liu 0001
CVPR4
2025 Learning Evidential Delta Denoising Scores for Video Editing
Yufan Hu, Junyu Gao 0002, Bin Fan 0001, Hongmin Liu 0001
ACM Multimedia5
2025 Uncertainty Aware Multiple View Stereo Network with Accurate Supervision
abstract
Learning-based multiple view stereo has gained significant attention recently. However, most methods rely on direct network supervision using provided ground-truth depth, which poses three inherent problems: resolution-dependent ground-truth artifacts, excessively challenging training examples (with relatively featureless textures), and use of less-viewed reference pixels for supervision, all of which hinder network optimization. To alleviate these problems, we propose an accurate network supervision paradigm that includes a ground-truth mask, an entropy mask, and a consistency mask, which provide more accurate supervision signals to aid network optimization. Furthermore, we introduce UANet, an uncertainty aware multi-view stereo network, which adaptively determines a pixel-wise search range using a dynamic range sampler (DRS) built upon estimation confidence and learned uncertainty. Experimental results on recent MVS datasets demonstrate the effectiveness of our method.
Xincheng Tang, Mengqi Rong, Bin Fan 0001, Hongmin Liu 0001, Shuhan Shen
Comput. Vis. Media4
2025 Vision Generalist Model: A Survey
Ziyi Wang 0007, Yongming Rao, Shuofeng Sun, Xinrun Liu, Yi Wei 0003, Xumin Yu, Zuyan Liu, Hongmin Liu 0001, Jie Zhou 0001, Jiwen Lu
Int. J. Comput. Vis.9
2025 ProIn: Learning to predict trajectory based on progressive interactions for autonomous driving
Yinke Dong, Haifeng Yuan, Hongkun Liu, Fangzhen Li, Hongmin Liu 0001, Bin Fan 0001
Neurocomputing6
2025 Cross-Modal Guided Visual Representation Learning for Social Image Retrieval
abstract
Social images are often associated with rich but noisy tags from community contributions. Although social tags can potentially provide valuable semantic training information for image retrieval, existing studies all fail to effectively filter noises by exploiting the cross-modal correlation between image content and tags. The current cross-modal vision-and-language representation learning methods, which selectively attend to the relevant parts of the image and text, show a promising direction. However, they are not suitable for social image retrieval since: (1) they deal with natural text sequences where the relationships between words can be easily captured by language models for cross-modal relevance estimation, while the tags are isolated and noisy; (2) they take (image, text) pair as input, and consequently cannot be employed directly for unimodal social image retrieval. This paper tackles the challenge of utilizing cross-modal interactions to learn precise representations for unimodal retrieval. The proposed framework, dubbed CGVR (Cross-modal Guided Visual Representation), extracts accurate semantic representations of images from noisy tags and transfers this ability to image-only hashing subnetwork by a carefully designed training scheme. To well capture correlated semantics and filter noises, it embeds a priori common-sense relationship among tags into attention computation for joint awareness of textual and visual context. Experiments show that CGVR achieves approximately 8.82 and 5.45 points improvement in MAP over the state-of-the-art on two widely used social image benchmarks. CGVR can serve as a new baseline for the image retrieval community. The code is provided at https://github.com/zhaowanqing/CGVR.
Ziyu Guan, Wanqing Zhao, Hongmin Liu 0001, Yuta Nakashima, Noboru Babaguchi, Xiaofei He 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2025 O2Flow: Object-aware optical flow estimation
Haixu Bi, Hongmin Liu 0001
Pattern Recognit. Lett.4
2025 Bidirectional Agent-Map Interaction Feature Learning Leveraged by Map-Related Tasks for Trajectory Prediction in Autonomous Driving
abstract
Accurate prediction of the future trajectories of surrounding agents is essential for safely autonomous vehicles. However, it is quite challenging due to the dynamics of driving environments caused by frequent agent-agent and agent-map interactions. Most existing methods mainly focus on modeling the dynamics of agents but overlook the inherent dynamics of the usable map at different driving moments. To address this limitation, this paper proposes a bidirectional interaction network (DyMap) that takes into account the bilateral constraints of the map on the agents and the impact of the agents on the map, using a unique bidirectional attention module. By employing multiple stages of bidirectional interaction modules, the features of agents and map nodes are continuously updated to adapt to changing environments. To facilitate network learning, we further incorporate map-related auxiliary tasks into the learning objectives, including traffic flow prediction, map node occupancy prediction, and road direction prediction. These tasks are designed to help the network learn map features that can reflect the dynamics of the driving environment, leading to a comprehensive feature representation of the agents. This representation encodes not only the agent’s historical motions but also the latent map information triggered by other agents. This proposed network offers a new approach to modeling environmental dynamics in driving scenarios, and achieves the state-of-the-art performance on the Argoverse 1 and Argoverse 2 trajectory prediction benchmarks. Note to Practitioners—Autonomous vehicles are of particular interest to practitioners in the field of automation science and technology. Trajectory prediction is of critical importance to the safety of autonomous driving, which aims to forecast the future positions of a given traffic participator (called the agent in this paper, such as a vehicle, pedestrian, etc.) based on the observations of itself and other agents’ historical trajectories as well as the map information. This paper proposes a trajectory prediction network based on a novel bidirectional interaction module that enables bidirectional information flow between the features of map nodes and agents so as to simultaneously learn their feature representations being aware of the timing-vary property of traffic scenes. The proposed DyMap contains three blocks of bidirectional interaction modules, which are interleaved with self-interactions of the agents and map nodes. Three different kinds of map-related tasks are developed to facilitate network training by leveraging on maximizing the map features for predicting traffic flow, occupancy, and lane direction. The superiority of the proposed method has been confirmed by the widely used benchmarks in the community. Besides modelling the complex interaction between agents and map nodes in dynamic traffic scenarios, the proposed bidirectional interaction network can be beneficial to address other problems requiring symmetric modelling of interactions between different components, such as human-robot interaction. The proposed map-related tasks can be directly used in other applications requiring to extract map features while considering the time-varying conditions, including but not limited to robot navigation and embody intelligence.
Bin Fan 0001, Haifeng Yuan, Yinke Dong, Zhengyu Zhu 0002, Hongmin Liu 0001
IEEE Trans Autom. Sci. Eng.5
2025 DMotion: Diverse Modalities Alignment Enhanced Motion Prediction for Autonomous Driving
abstract
In autonomous driving, motion prediction is vital for anticipating the behaviors of surrounding vehicles, pedestrians, and other road users, enabling the system to make accurate decisions and plan driving paths effectively. Current motion prediction models typically utilize an encoder–decoder architecture, and many methods focus on the decoder’s design because the decoder is directly responsible for generating future trajectories. However, this often leads to neglecting the encoder’s capability to represent input information, resulting in a failure to provide the decoder with accurate and semantic prior features, which impacts the overall prediction accuracy. Contrastive learning, as an effective approach for enhancing feature representation through cross-modal alignment, demonstrates strong generalization and reduced reliance on labeled data. Therefore, we propose a motion prediction network named DMotion that leverages contrastive learning to align trajectory features with numeric signals and textual descriptions, improving the model’s representational capacity and enriching contextual priors. To the best of our knowledge, this is the first approach to use textual descriptions as a modality to enhance motion prediction accuracy. We sparsify dense agent attribute labels, such as historical distance and angular variation, to enable these prior features learned by the model through contrastive learning. By investigating the impact of numeric and textual supervision signals on contrastive learning effectiveness, textual supervision achieves superior results compared with numeric signals, benefiting from richer input information and more robust extraction capabilities of the text model. We further apply the low-rank adaptation (LoRA) method to fine-tune the text encoder, improving model performance and preventing catastrophic forgetting with only 0.1M additional trainable parameters. Our experiments demonstrate that DMotion shows competitive performance on the Waymo motion prediction and interaction prediction challenges. Additionally, the contrast learning module of DMotion does not introduce additional parameters or computational overhead during inference, maintaining the efficiency of the original encoder-decoder model.
Hongkun Liu, Hongmin Liu 0001, Bin Fan 0001, Jinglin Xu
IEEE Trans. Comput. Soc. Syst.2
2025 Barely-Supervised Brain Tumor Segmentation via Employing Segment Anything Model
abstract
This work explores barely-supervised brain tumor segmentation where minimal supervision,i.e., fewer than ten labeled samples, is available. Current methods often neglect two key problems in barely-supervised segmentation: i) the insufficient labeled data may be not able to offer enough information to networks for accurately segmenting tumor areas across various cases; ii) networks might overfit to the relation of multiple modalities of the limited labeled data, thus overly depending on certain modalities while overlooking other valuable modalities during segmentation. To tackle these two problems, we propose a barely-supervised training framework, called BarelySAM. BarelySAM first employs Segment Anything Model (SAM) during training by generating pseudo labels for unlabeled data. In this manner, pre-trained knowledge exhibited in SAM can be exploited to compensate for limited knowledge in labeled data, boosting network training and thus improving performance. For the overfitting problem, Multi-modality Dependency Minimization (MDM) is designed in BarelySAM to construct various partial combinations for full-modal samples, thus enforcing networks to exploit each modality effectively. Experimental results on two benchmark datasets validate the effectiveness of the integrated SAM and the designed MDM module. In particular, our method attains a 89.92% Dice score for whole tumor segmentation on BRATS2020 with just 6 (2%) labeled samples, just 1.09% lower than the performance of a fully supervised approach. Besides, experiments on barely-supervised multi-modal brain tumor segmentation also validate that our method is inherently robust against missing modalities.
Yuhang Ding, Hongmin Liu 0001
IEEE Trans. Circuits Syst. Video Technol.2
2025 BRTAL: Boundary Refinement Temporal Action Localization via Offset-Driven Diffusion Models
abstract
Temporal Action Localization (TAL) aims to classify and localize all actions within untrimmed videos. Existing TAL methods often struggle with inaccurate boundary predictions due to the similarity of action content and the uncertainty of boundaries between adjacent frames. Many of these methods rely on fixed or global proposal learning strategies, which lack a more refined method to improve localization accuracy. In this paper, we propose BRTAL, a new Boundary Refinement framework for TAL based on an offset-driven diffusion model, specifically designed to enhance action boundary precision through a refined approach iteratively. Unlike traditional TAL methods emphasizing global target predictions, BRTAL adopts a local refinement perspective by leveraging an offset-driven strategy. Specifically, our framework employs diffusion to iteratively generate local offsets between predictions and ground truth, gradually reducing these offsets to achieve better alignment with the ground truth. This refined approach is particularly effective in addressing the challenges of ambiguous boundaries frequently encountered in TAL, enabling BRTAL to achieve more refined boundary localization than existing methods. Furthermore, we introduce a lightweight yet powerful Temporal Context Modeling (TCM) module to enhance temporal information modeling for accurate action localization. TCM features a Temporal Representation Perception (TRP) layer, which captures temporal evolution and long-term contextual dependencies through a squeeze-and-excitation design combined with large convolutional kernels, ensuring robust temporal representation learning. Extensive experiments on THUMOS14, ActivityNet-1.3, and EPIC-KITCHEN 100 datasets highlight the significant advantages of BRTAL. Notably, BRTAL achieves an average mAP of 69.6% on THUMOS14, establishing a new state-of-the-art benchmark and demonstrating its outstanding boundary refinement capability.
Hongmin Liu 0001, Xueli Li, Bin Fan 0001, Jinglin Xu
IEEE Trans. Circuits Syst. Video Technol.1
2025 Dynamic Learnable Label Assignment for Indoor 3D Object Detection
abstract
In this paper, we present a dynamic learnable label assignment (DLLA) method for indoor anchor-free one-stage 3D object detection. Existing methods principally depend on hand-crafted strategies with fixed thresholds, which fail to adapt to the inherent variability in object characteristics such as size, shape, and occlusion levels. This lack of adaptability results in suboptimal sample assignments and unstable detection performance. To address this challenge, we map the features of proposals and ground truths separately into the same embedding space, enabling a dynamic strategy of assigning appropriate positive samples to each instance. Specifically, we first interact with the features of all proposals to effectively integrate information from each proposal in the scene and capture long-range dependencies between different locations. Additionally, to extract more discriminative and generalized features for positive and negative samples, we employ a contrastive learning process to optimize the elemental relationships and distances between proposals and ground truths. Finally, we introduce a denoising task to alleviate the difficulty of the unsupervised learning process in DLLA. Experimental results show that our DLLA outperforms other methods on three popular indoor datasets (ScanNet V2, SUN RGB-D, and ScanNet200).
Xinrun Liu, Linqing Zhao, Bin Fan 0001, Jiwen Lu, Hongmin Liu 0001
IEEE Trans. Circuits Syst. Video Technol.5
2025 BFSTAL: Bidirectional Feature Splitting With Cross-Layer Fusion for Temporal Action Localization
abstract
Temporal Action Localization (TAL) aims to identify the boundaries of actions and their corresponding categories in untrimmed videos. Most existing methods simultaneously process past and future information, neglecting the inherently sequential nature of action occurrence. This confused treatment of past and future information hinders the model’s ability to understand action procedures effectively. To address these issues, we propose Bidirectional Feature Splitting with Cross-Layer Fusion for Temporal Action Localization (BFSTAL), a new bidirectional feature-splitting approach based on Mamba for the TAL task, composed of two core parts, Decomposed Bidirectionally Hybrid (DBH) and Cross-Layer Fusion Detection (CLFD), which explicitly enhances the model’s capacity to understand action procedures, especially to localize temporal boundaries of actions. Specifically, we introduce the Decomposed Bidirectionally Hybrid (DBH) component, which splits video features at a given timestamp into forward features (past information) and backward features (future information). DBH integrates three key modules: Bidirectional Multi-Head Self-Attention (Bi-MHSA), Bidirectional State Space Model (Bi-SSM), and Bidirectional Convolution (Bi-CONV). DBH effectively captures long-range dependencies by combining state-space modeling, attention mechanisms, and convolutional networks while improving spatial-temporal awareness. Furthermore, we propose Cross-Layer Fusion Detection (CLFD), which aggregates multi-scale features from different pyramid levels, enhancing contextual understanding and temporal action localization precision. Extensive experiments demonstrate that BFSTAL outperforms other methods on four widely used TAL benchmarks: THUMOS14, EPIC-KITCHENS 100, Charades, and MultiTHUMOS.
Jinglin Xu, Hongmin Liu 0001
IEEE Trans. Circuits Syst. Video Technol.4
2025 Learning Boundary Continuity-Aware Gaussian Encoder for Oriented Object Detection
abstract
Oriented object detection has been crucial for rotation-sensitive tasks and has garnered significant attention. Most existing methods generate angles as detector output vectors, but this strategy can abnormally magnify visually similar differences between two boxes in certain circumstances, termed boundary discontinuity issue. To overcome this limitation, we propose a boundary continuity-aware Gaussian encoder (BCGE). Specifically, BCGE directly predicts target Gaussian distributions for proposals and learns an oriented bounding box as an integrated 2-D matrix, effectively addressing boundary discontinuity issues. We also propose a transformation from Gaussian representation back to boxes and extend this transformation theory to the complex domain to adapt to the learning characteristics of neural networks. Furthermore, BCGE serves as a versatile plug-and-play architectural encoder, directly replacing the standard coding process in various oriented detectors with adaptability. Experimental results on five popular datasets, i.e., DOTA, UCAS-AOD, HRSC2016, SSDD, and HRSID, consistently show the effectiveness of our approach.
Hongmin Liu 0001, Chengyi Zhao, Bin Fan 0001, Yufan Hu
IEEE Trans. Cybern.1
2025 Dual-Level Modality De-Biasing for RGB-T Tracking
abstract
RGB-T tracking aims to effectively leverage the complement ability of visual (RGB) and infrared (TIR) modalities to achieve robust tracking performance in various scenarios. Existing RGB-T tracking methods typically adopt backbone networks pre-trained on large-scale RGB datasets, which can lead to a predisposition toward RGB image patterns. RGB and TIR modalities also exhibit inconsistent responses to regions with diverse properties, resulting in imbalances in tracking decisions. We refer to these issues as feature-level and decision-level biases in the TIR modality. In this paper, we propose a novel dual-level modality de-biasing framework for RGB-T tracking to eliminate the inherent feature and decision-level biases. Specifically, we propose a joint infrared-fusion adapter, comprising an infrared-aware adapter and a cross-fusion adapter, designed to adaptively mitigate feature-level biases and utilize complementary information between the two modalities. In addition to implicit feature-level adjustment, we propose a response-decoupled distillation strategy to explicitly alleviate decision-level biases, aiming to achieve consistently accurate decision-making between the RGB and TIR modalities. Extensive experiments on several popular RGB-T tracking benchmarks validate the effectiveness of our proposed method.
Yufan Hu, Zekai Shao 0002, Bin Fan 0001, Hongmin Liu 0001
IEEE Trans. Image Process.4
2025 ScalableTrack: Scalable One-Stream Tracking via Alternating Learning
abstract
Transformer-based one-stream trackers are widely used to extract features and interact information for visual object tracking. However, the current one-stream tracker has fixed computational dimensions between different stages, which limits the network's ability to learn context clues and global representations, resulting in a decrease in the ability to distinguish between targets and backgrounds. To address this issue, a new scalable one-stream tracking framework, ScalableTrack, is proposed. It unifies feature extraction and information integration by intrastage mutual guidance, leveraging the scalability of target-oriented features to enhance object sensitivity and obtain discriminative global representations. In addition, we bridge interstage contextual cues by introducing an alternating learning strategy and solve the arrangement problem of the two modules. The alternating learning strategy uses alternate stacks of feature extraction and information interaction to focus on tracked objects and prevent catastrophic forgetting of target information between different stages. Experiments on eight challenging benchmarks (TrackingNet, GOT-10k, VOT2020, UAV123, LaSOT, LaSOText, OTB100, and TC128) show that ScalableTrack outperforms state-of-the-art (SOTA) methods with better generalization and global representation ability.
Hongmin Liu 0001, Yuefeng Cai, Bin Fan 0001, Jinglin Xu
IEEE Trans. Neural Networks Learn. Syst.1
2025 A Cascaded Multimodule Image Enhancement Framework for Underwater Visual Perception
abstract
Underwater images usually exhibit severe color cast, hazy appearance, and/or dark regions because of the complex lighting absorption and scattering in water. How to increase the quality of these degraded underwater images has emerged as a key issue for various underwater application tasks. Recent efforts have been made to deal with single type degradation, however, it is still challenging to deal with multiple degradations that usually coexist in an underwater image with a general network. The degradations in underwater images can be divided into medium-agnostic (hazy or low-light which also encountered in in-air images) and medium-specific (color distortion caused by the specific light attenuation property in water) ones. According to this observation, this article proposes a cascaded multimodule underwater image enhancement (UIE) framework to address the coexisted multiple degradations. In the proposed framework, an in-air image enhancement module and a novel proposed adaptive color channel compensation network (AC3Net) are cascaded, in which the former focuses primarily on solving medium-agnostic degradations and the latter is for handling the medium-specific degradation. This framework has good flexibility by cascading different types of in-air image enhancement networks with AC3Net to achieve various UIE. The effectiveness of the proposed framework has been extensively validated on various degraded underwater images as well as different underwater visual perception tasks.
Hongmin Liu 0001, Hui Zeng 0003, Huayan Pu, Jun Luo 0003, Bin Fan 0001
IEEE Trans. Neural Networks Learn. Syst.1
2024 RMT: Retentive Networks Meet Vision Transformers
abstract
Vision Transformer (ViT) has gained increasing attention in the computer vision community in recent years. How-ever, the core component of ViT, Self-Attention, lacks ex-plicit spatial priors and bears a quadratic computational complexity, thereby constraining the applicability of ViT. To alleviate these issues, we draw inspiration from the re-cent Retentive Network (RetNet) in the field of NLP, and propose RMT, a strong vision backbone with explicit spa-tial prior for general purposes. Specifically, we extend the RetNet's temporal decay mechanism to the spatial do-main, and propose a spatial decay matrix based on the Manhattan distance to introduce the explicit spatial prior to Self-Attention. Additionally, an attention decomposition form that adeptly adapts to explicit spatial prior is proposed, aiming to reduce the computational burden of modeling global information without disrupting the spa-tial decay matrix. Based on the spatial decay matrix and the attention decomposition form, we can flexibly integrate explicit spatial prior into the vision backbone with lin-ear complexity. Extensive experiments demonstrate that RMT exhibits exceptional performance across various vision tasks. Specifically, without extra training data, RMT achieves 84.8% and 86.1% top-l acc on ImageNet-lk with 27MI4.5GFLOPs and 96M/18.2GFLOPs. For downstream tasks, RMT achieves 54.5 box AP and 47.2 mask AP on the COCO detection task, and 52.8 mloU on the ADE20K se-mantic segmentation task.
Qihang Fan, Huaibo Huang, Mingrui Chen 0001, Hongmin Liu 0001, Ran He 0001
CVPR4
2024 DSPDet3D: 3D Small Object Detection with Dynamic Spatial Pruning
Xiuwei Xu, Ziwei Wang 0010, Hongmin Liu 0001, Jie Zhou 0001, Jiwen Lu
ECCV (28)4
2024 Robust Synthetic-to-Real Ensemble Dehazing Algorithm With the Intermediate Domain
abstract
Learning-based dehazing methods using synthetic datasets cannot generalize well on real-world hazy images due to the large domain discrepancy. To tackle this issue, we propose a robust synthetic-to-real dehazing framework with the construction of an intermediate domain and ensemble learning strategy. First, by mapping all examples to the intermediate domain, the bidirectional match strategy with adversarial training and the constraint of intermediated results is proposed to suppress the rich domain-specific information, which can facilitate the adaptation and perform image dehazing simultaneously. Furthermore, an ensemble dehazing algorithm based on the intermediate domain is proposed in a semisupervised manner. The reconstruction constraint and the enhanced ground-truths are employed to keep the visual fidelity and remove the dim artifacts of unsupervised dehazing results. Finally, we propose the domain-aware residual groups to deal with the distribution discrepancy between the synthetic and real hazy images. Extensive experiments of various real-world hazy images demonstrate that the proposed method outperforms the state-of-the-art dehazing methods and significantly improves the generalization in the real world.
Yingxu Qiao, Hongmin Liu 0001, Zhanqiang Huo
IEEE Trans. Comput. Soc. Syst.3
2024 Learning Proposal-Aware Re-Ranking for Weakly-Supervised Temporal Action Localization
abstract
Weakly-supervised temporal action localization (WTAL) aims to localize and classify action instances in untrimmed videos with only video-level labels available. Despite the remarkable success of existing methods, whose generated proposals are commonly far more than the ground-truth action instances, it still makes sense to improve the ranking accuracy of the generated proposals since users in real-world scenarios usually prioritize the action proposals with the highest confidence scores. The inaccuracy of the proposal ranking mainly comes from two aspects: For one thing, the traditional proposal generation manner entirely relies on snippet-level perception, resulting in a significant yet unnoticed gap with the target of proposal-level localization. For another, existing methods commonly employ a hand-crafted proposal generation manner, a post-process that does not participate in model optimization. To address the above issues, we propose an end-to-end trained two-stage method, termed as Learning Proposal-aware Re-ranking (LPR) for WTAL. In the first stage, we design a proposal-aware feature learning module to inject the proposal-aware contextual information into each snippet, and then the enhanced features are utilized for predicting initial proposals. Furthermore, to perform effective and efficient proposal re-ranking, in the second stage, we contrast the proposals attached with high confidence scores with our constructed multi-scale foreground/background prototypes for further optimization. Evaluated by both the vanilla and Top-$k$mAP metrics, results of extensive experiments on two popular benchmarks demonstrate the effectiveness of our proposed method.
Yufan Hu, Jie Fu 0004, Junyu Gao 0002, Jianfeng Dong, Bin Fan 0001, Hongmin Liu 0001
IEEE Trans. Circuits Syst. Video Technol.7
2024 Pro2Diff: Proposal Propagation for Multi-Object Tracking via the Diffusion Model
abstract
Multi-object tracking (MOT) aims to estimate the bounding boxes and ID labels of objects in videos. The challenging issue in this task is to alleviate competitive learning between the detection and tracking subtasks, for which, two-stage Tracking-By-Detection (TBD) optimizes the two subtasks individually, and the single-stage Joint Detection and Tracking (JDT) adjusts the complex network architectures finely in an end-to-end pipeline. In this paper, we propose a new MOT method, i.e., Proposal Propagation via Diffusion Models, called Pro2Diff, which integrates a diffusion model into the proposal propagation in multi-object tracking, focusing on the model training process rather than complex network design. Specifically, using a generative approach, Pro2Diff generates a considerable number of noisy proposals for the tracking image sequence in the forward process, and subsequently, Pro2Diff learns the discrepancies between these noisy proposals and the actual bounding boxes of the tracked objects, gradually optimizing these noisy proposals to obtain the initial sequence of real tracked objects. By introducing the denoising diffusion process into multi-object tracking, we have made three further important findings: 1) Generative methods can effectively handle multi-object tracking tasks; 2) Without the need to modify the model structure, we propose self-conditional proposal propagation to enhance model performance effectively during inference; 3) By adjusting the numbers of proposals and iterations appropriately for different tracking sequences, the optimal performance of the model can be achieved. Extensive experimental results on MOT17 and DanceTrack datasets demonstrate that Pro2Diff outperforms current end-to-end multi-object tracking methods. We achieve 61.9 HOTA on DanceTrack and 57.6 HOTA on MOT17, reaching the competitive result of the JDT approach.
Hongmin Liu 0001, Canbin Zhang, Bin Fan 0001, Jinglin Xu
IEEE Trans. Image Process.1
2024 Exploring Rich Semantics for Open-Set Action Recognition
abstract
Open-set action recognition (OSAR) aims to learn a recognition framework capable of both classifying known classes and identifying unknown actions in open-set scenarios. Existing OSAR methods typically reside in a data-driven paradigm, which ignore the rich semantics in both known and unknown categories. In fact, we humans have the capability of leveraging the captured semantic information, i.e., knowledge and experience, to incisively distinguish samples from known and unknown classes. Motivated by this observation, in this paper, we propose a Unified Semantic Exploration (USE) framework for recognizing actions in open-set scenarios. Specifically, we explore the explicit knowledge semantics by simulating the unknown classes with knowledge-guided virtual classes based on an external knowledge graph, which enables the model to simulate open-set perception during model training. Besides, we propose to learn the implicit data semantics by transferring the knowledge structure of action categories to the visual prototype space for semantic structure preservation. Extensive experiments on several action recognition benchmarks validate the effectiveness of our proposed method.
Yufan Hu, Junyu Gao 0002, Jianfeng Dong, Bin Fan 0001, Hongmin Liu 0001
IEEE Trans. Multim.5
2024 Image Enhancement Guided Object Detection in Visually Degraded Scenes
abstract
Object detection accuracy degrades seriously in visually degraded scenes. A natural solution is to first enhance the degraded image and then perform object detection. However, it is suboptimal and does not necessarily lead to the improvement of object detection due to the separation of the image enhancement and object detection tasks. To solve this problem, we propose an image enhancement guided object detection method, which refines the detection network with an additional enhancement branch in an end-to-end way. Specifically, the enhancement branch and detection branch are organized in a parallel way, and a feature guided module is designed to connect the two branches, which optimizes the shallow feature of the input image in the detection branch to be as consistent as possible with that of the enhanced image. As the enhancement branch is frozen during training, such a design plays a role in using the features of enhanced images to guide the learning of object detection branch, so as to make the learned detection branch being aware of both image quality and object detection. When testing, the enhancement branch and feature guided module are removed, and so no additional computation cost is introduced for detection. Extensive experimental results, on underwater, hazy, and low-light object detection datasets, demonstrate that the proposed method can improve the detection performance of popular detection networks (YOLO v3, Faster R-CNN, DetectoRS) significantly in visually degraded scenes.
Hongmin Liu 0001, Hui Zeng 0003, Huayan Pu, Bin Fan 0001
IEEE Trans. Neural Networks Learn. Syst.1
2023 Building Vision Transformers with Hierarchy Aware Feature Aggregation
abstract
Thanks to the excellent global modeling capability of attention mechanisms, the Vision Transformer has achieved better results than ConvNet in many computer tasks. However, in generating hierarchical feature maps, the Transformer still adopts the ConvNet feature aggregation scheme. This leads to the problem that the semantic information of the grid area of image becomes confused after feature aggregation, making it difficult for attention to accurately model global relationships. To address this, we propose the Hierarchy Aware Feature Aggregation framework (HAFA). HAFA enhances the extraction of local features adaptively in shallow layers where semantic information is weak, while is able to aggregate patches with similar semantics in deep layers. The clear semantic information of the aggregated patches, enables the attention mechanism to more accurately model global information at the semantic level. Extensive experiments show that after using the HAFA framework, significant improvements have been achieved relative to the baseline models in image classification, object detection, and semantic segmentation tasks.
Yongjie Chen, Hongmin Liu 0001, Bin Fan 0001
ICCV2
2023 Seeing Through Darkness: Visual Localization at Night via Weakly Supervised Learning of Domain Invariant Features
abstract
Long term visual localization has to conquer the problem of matching images with dramatic photometric changes caused by different seasons, natural and man-made illumination changes, etc. Visual localization at night plays a vital role in many applications like autonomous driving and augmented reality, for which extracting keypoints and descriptors with robustness to day-night illumination changes has became the bottleneck. This paper proposes an adversarial learning based solution to harvest from the weakly domain labels of day and night images, along with the point level correspondences among day time images, to achieve robust local feature extraction and description across day-night images. The key idea is to learn a discriminator to distinguish whether a feature map is generated from the day or night images, and simultaneously to adjust the parameters of feature extraction network so as to fool the discriminator. After adversarial training of the discriminator and feature extraction network, the feature extraction network finally reaches a stable status so that the extracted feature maps are robust to day-night photometric changes, based on which day-night domain invariant keypoints and descriptors can be extracted. Compared to existing local feature learning methods, it only requires an additional set of easily captured night images to improve the domain invariance of learned features. Experiments on two challenging benchmarks show the effectiveness of proposed method. In addition, this paper revisits the widely used image matching metrics on HPatches and finds that recall of different methods is highly related to their relative localization performance.
Bin Fan 0001, Yuzhu Yang, Wensen Feng, Fuchao Wu, Jiwen Lu, Hongmin Liu 0001
IEEE Trans. Multim.6
2022 CARR-Net: Leveraging on Subtle Variance of Neighbors for Point Cloud Semantic Segmentation
Mingming Song, Bin Fan 0001, Hongmin Liu 0001
PRCV (3)3
2022 Robust self-supervised monocular visual odometry based on prediction-update pose estimation network
Haixin Xiu, Yiyou Liang, Hui Zeng 0003, Qing Li 0015, Hongmin Liu 0001, Bin Fan 0001
Eng. Appl. Artif. Intell.5
2022 Soft Margin Triplet-Center Loss for Multi-View 3D Shape Retrieval
abstract
Obtaining discriminative features is one of the key problems in three-dimensional (3D) shape retrieval. Recently, deep metric learning-based 3D shape retrieval methods have attracted the researchers’ attention and have achieved better performance. The triplet-center loss can learn more discriminative features than traditional classification loss, and it has been successfully used in deep metric learning-based 3D shape retrieval task. However, it has a hard margin parameter that only leverages part of the training data in each mini-batch. Moreover, the margin parameter is often determined by experience and remains unchanged during the training process. To overcome the above limitations, we propose the soft margin triplet-center loss, which replaces the margin with the nonparametric soft margin. Furthermore, we combined the proposed soft margin triplet-center loss with the softmax loss to improve the training efficiency and the retrieval performance. Extensive experimental results on two popular 3D shape retrieval datasets have validated the effectiveness of the soft margin triplet-center loss, and our proposed 3D shape retrieval method has achieved better performance than other state-of-the-art method.
Ruting Cheng, Fuzhou Wang, Tianmeng Zhao, Hongmin Liu 0001, Hui Zeng 0003
Int. J. Pattern Recognit. Artif. Intell.4
2022 Incremental Translation Averaging
abstract
Translation averaging is known to be more difficult than rotation averaging due to scale ambiguity, estimation sensitivity, and solution uncertainty. Existing approaches have exposed their limitations in terms of accuracy, robustness, simplicity, or efficiency. To tackle this tough problem, a simple yet effective translation averaging pipeline, termed as Incremental Translation Averaging (ITA), is proposed in this paper. It combines the advantages of high accuracy and robustness in incremental parameter estimation pipeline and the advantages of high simplicity and efficiency in global motion averaging approach. Unlike the traditional translation averaging methods which estimate all the absolute camera locations simultaneously and suffer from inaccuracy in parameter estimation and incompleteness in scene reconstruction, our ITA computes them novelly in an incremental way with higher accuracy and robustness. Thanks to the introduction of incremental parameter estimation thought into the translation averaging pipeline, 1) our ITA is robust to measurement outliers and accurate in parameter estimation; and 2) our ITA is simple and efficient because of its less dependency on complicated optimization, carefully-designed preprocessing, or additional information. Comprehensive evaluations on the 1DSfM dataset demonstrate the effectiveness of our ITA and its advantages over several state-of-the-art translation averaging approaches.
Xiang Gao 0009, Lingjie Zhu, Bin Fan 0001, Hongmin Liu 0001, Shuhan Shen
IEEE Trans. Circuits Syst. Video Technol.4
2022 Active Learning Based 3D Semantic Labeling From Images and Videos
abstract
3D semantic segmentation is one of the most fundamental problems for 3D scene understanding and has attracted much attention in the field of computer vision. In this paper, we propose an active learning based 3D semantic labeling method for large-scale 3D mesh model generated from images or videos. Taking as input a 3D mesh model reconstructed from the image based 3D modeling system, coupled with the calibrated images, our method outputs a fine 3D semantic mesh model in which each facet is assigned a semantic label. There are three major steps in our framework: 2D semantic segmentation, 2D-3D semantic fusion, and batch image selection. A limited annotation image set is first used to fine-tune a pre-trained semantic segmentation network for obtaining the pixel-wise semantic probability maps. Then all these maps are back-projected into 3D space and fused on the 3D mesh model using Markov Random Field optimization, thus yield a preliminary 3D semantic mesh model and a heat model showing each facet’s confidence. This 3D semantic model is used as a reliable supervisor to select the parts that are not well segmented for manual annotation to boost the performance of the 2D semantic segmentation network, as well as the 3D mesh labeling, in the next iteration. This Training-Fusion-Selection process continues until the label assignment of the 3D mesh model becomes steady. By this means, we significantly reduce the amount for annotation but not the labeling quality of 3D semantic models. Extensive experiments demonstrate the effectiveness and generalization ability of our method on a wide variety of datasets.
Mengqi Rong, Hainan Cui, Zhanyi Hu, Hanqing Jiang, Hongmin Liu 0001, Shuhan Shen
IEEE Trans. Circuits Syst. Video Technol.5
2022 Progressive Multistage Learning for Discriminative Tracking
abstract
Visual tracking is typically solved as a discriminative learning problem that usually requires high-quality samples for online model adaptation. It is a critical and challenging problem to evaluate the training samples collected from previous predictions and employ sample selection by their quality to train the model. To tackle the above problem, we propose a joint discriminative learning scheme with the progressive multistage optimization policy of sample selection for robust visual tracking. The proposed scheme presents a novel time-weighted and detection-guided self-paced learning strategy for easy-to-hard sample selection, which is capable of tolerating relatively large intraclass variations while maintaining interclass separability. Such a self-paced learning strategy is jointly optimized in conjunction with the discriminative tracking process, resulting in robust tracking results. Experiments on the benchmark datasets demonstrate the effectiveness of the proposed learning framework.
Weichao Li 0003, Xi Li 0001, Omar El Farouk Bourahla, Fuxian Huang, Fei Wu 0001, Wei Liu 0005, Hongmin Liu 0001
IEEE Trans. Cybern.7
2022 Composite Kernel of Mutual Learning on Mid-Level Features for Hyperspectral Image Classification
abstract
By training different models and averaging their predictions, the performance of the machine-learning algorithm can be improved. The performance optimization of multiple models is supposed to generalize further data well. This requires the knowledge transfer of generalization information between models. In this article, a multiple kernel mutual learning method based on transfer learning of combined mid-level features is proposed for hyperspectral classification. Three-layer homogenous superpixels are computed on the image formed by PCA, which is used for computing mid-level features. The three mid-level features include: 1) the sparse reconstructed feature; 2) combined mean feature; and 3) uniqueness. The sparse reconstruction feature is obtained by a joint sparse representation model under the constraint of three-scale superpixels' boundaries and regions. The combined mean features are computed with average values of spectra in multilayer superpixels, and the uniqueness is obtained by the superposed manifold ranking values of multilayer superpixels. Next, three kernels of samples in different feature spaces are computed for mutual learning by minimizing the divergence. Then, a combined kernel is constructed to optimize the sample distance measurement and applied by employing SVM training to build classifiers. Experiments are performed on real hyperspectral datasets, and the corresponding results demonstrated that the proposed method can perform significantly better than several state-of-the-art competitive algorithms based on MKL and deep learning.
Haifeng Sima, Jing Wang 0093, Ping Guo 0002, Junding Sun, Hongmin Liu 0001, Mingliang Xu 0001, Youfeng Zou
IEEE Trans. Cybern.5
2022 Adversarial Incomplete Multiview Subspace Clustering Networks
abstract
Multiview clustering aims to leverage information from multiple views to improve the clustering performance. Most previous works assumed that each view has complete data. However, in real-world datasets, it is often the case that a view may contain some missing data, resulting in the problem of incomplete multiview clustering (IMC). Previous approaches to this problem have at least one of the following drawbacks: 1) employing shallow models, which cannot well handle the dependence and discrepancy among different views; 2) ignoring the hidden information of the missing data; and 3) being dedicated to the two-view case. To eliminate all these drawbacks, in this work, we present the adversarial IMC (AIMC) framework. In particular, AIMC seeks the common latent representation of multiview data for reconstructing raw data and inferring missing data. The elementwise reconstruction and the generative adversarial network are integrated to evaluate the reconstruction. They aim to capture the overall structure and get a deeper semantic understanding, respectively. Moreover, the clustering loss is designed to obtain a better clustering structure. We explore two variants of AIMC, namely: 1) autoencoder-based AIMC (AAIMC) and 2) generalized AIMC (GAIMC), with different strategies to obtain the multiview common representation. Experiments conducted on six real-world datasets show that AAIMC and GAIMC perform well and outperform the baseline methods.
Hongmin Liu 0001, Ziyu Guan, Xunlian Wu, Jiale Tan, Beilei Ling
IEEE Trans. Cybern.2
2022 VidSfM: Robust and Accurate Structure-From-Motion for Monocular Videos
abstract
With the popularization of smartphones, larger collection of videos with high quality is available, which makes the scale of scene reconstruction increase dramatically. However, high-resolution video produces more match outliers, and high frame rate video brings more redundant images. To solve these problems, a tailor-made framework is proposed to realize an accurate and robust structure-from-motion based on monocular videos. The key ideas include two points: one is to use the spatial and temporal continuity of video sequences to improve the accuracy and robustness of reconstruction; the other is to use the redundancy of video sequences to improve the efficiency and scalability of system. Our technical contributions include an adaptive way to identify accurate loop matching pairs, a cluster-based camera registration algorithm, a local rotation averaging scheme to verify the pose estimate and a local images extension strategy to reboot the incremental reconstruction. In addition, our system can integrate data from different video sequences, allowing multiple videos to be simultaneously reconstructed. Extensive experiments on both indoor and outdoor monocular videos demonstrate that our method outperforms the state-of-the-art approaches in robustness, accuracy and scalability.
Hainan Cui, Diantao Tu, Fulin Tang, Pengfei Xu 0013, Hongmin Liu 0001, Shuhan Shen
IEEE Trans. Image Process.5
2022 Learning Semantic-Aware Local Features for Long Term Visual Localization
abstract
Extracting robust and discriminative local features from images plays a vital role for long term visual localization, whose challenges are mainly caused by the severe appearance differences between matching images due to the day-night illuminations, seasonal changes, and human activities. Existing solutions resort to jointly learning both keypoints and their descriptors in an end-to-end manner, leveraged on large number of annotations of point correspondence which are harvested from the structure from motion and depth estimation algorithms. While these methods show improved performance over non-deep methods or those two-stage deep methods, i.e., detection and then description, they are still struggled to conquer the problems encountered in long term visual localization. Since the intrinsic semantics are invariant to the local appearance changes, this paper proposes to learn semantic-aware local features in order to improve robustness of local feature matching for long term localization. Based on a state of the art CNN architecture for local feature learning, i.e., ASLFeat, this paper leverages on the semantic information from an off-the-shelf semantic segmentation network to learn semantic-aware feature maps. The learned correspondence-aware feature descriptors and semantic features are then merged to form the final feature descriptors, for which the improved feature matching ability has been observed in experiments. In addition, the learned semantics embedded in the features can be further used to filter out noisy keypoints, leading to additional accuracy improvement and faster matching speed. Experiments on two popular long term visual localization benchmarks (Aachen Day and Night v1.1, Robotcar Seasons) and one challenging indoor benchmark (InLoc) demonstrate encouraging improvements of the localization accuracy over its counterpart and other competitive methods.
Bin Fan 0001, Wensen Feng, Huayan Pu, Yuzhu Yang, Qingqun Kong, Fuchao Wu, Hongmin Liu 0001
IEEE Trans. Image Process.8
2021 Incremental Rotation Averaging
Xiang Gao 0009, Lingjie Zhu, Zexiao Xie, Hongmin Liu 0001, Shuhan Shen
Int. J. Comput. Vis.4
2021 Geometric attentional dynamic graph convolutional neural networks for point cloud analysis
Yiming Cui 0002, Xin Liu 0027, Hongmin Liu 0001, Jiyong Zhang 0001, Alina Zare, Bin Fan 0001
Neurocomputing3
2021 PrGCN: Probability prediction with graph convolutional network for person re-identification
Hongmin Liu 0001, Zhenzhen Xiao, Bin Fan 0001, Hui Zeng 0003, Guoquan Jiang
Neurocomputing1
2021 Urban Scene LOD Vectorized Modeling From Photogrammetry Meshes
abstract
Urban scene modeling is a challenging task for the photogrammetry and computer vision community due to its large scale, structural complexity, and topological delicacy. This paper presents an efficient multistep modeling framework for large-scale urban scenes from aerial images. It takes aerial images and a textured 3D mesh model generated by an image-based modeling system as the input and outputs compact polygon models with semantics at different levels of detail (LODs). Based on the key observation that urban buildings usually have piecewise planar rooftops and vertical walls, we propose a segment-based modeling method, which consists of three major stages: scene segmentation, roof contour extraction, and building modeling. By combining the deep neural network predictions with geometric constraints of the 3D mesh, the scene is first segmented into three classes. Then, for each building mesh, the 2D line segments are detected and used to slice the ground into polygon cells, followed by assigning each cell a roof plane via a MRF optimization. Finally, the LOD model is obtained by extruding cells to their corresponding planes. Compared with direct modeling in 3D space, we transform the mesh into a uniform 2D image grid representation and most of the modeling work is performed in 2D space, which has the advantages of low computational complexity and high robustness. In addition, our method doesn't require any global prior, such as the Manhattan or Atlanta world assumption, making it flexible to model scenes with different characteristics and complexity. Experiments on both single buildings and large-scale urban scenes demonstrate that by combining 2D photometric with 3D geometric information, the proposed algorithm is robust and efficient in urban scene LOD vectorized modeling compared with the state-of-the-art approaches.
Jiali Han, Lingjie Zhu, Xiang Gao 0009, Zhanyi Hu, Liyang Zhou, Hongmin Liu 0001, Shuhan Shen
IEEE Trans. Image Process.6
2021 Efficient Style-Corpus Constrained Learning for Photorealistic Style Transfer
abstract
Photorealistic style transfer is a challenging task, which demands the stylized image remains real. Existing methods are still suffering from unrealistic artifacts and heavy computational cost. In this paper, we propose a novel Style-Corpus Constrained Learning (SCCL) scheme to address these issues. The style-corpus with the style-specific and style-agnostic characteristics simultaneously is proposed to constrain the stylized image with the style consistency among different samples, which improves photorealism of stylization output. By using adversarial distillation learning strategy, a simple fast-to-execute network is trained to substitute previous complex feature transforms models, which reduces the computational cost significantly. Experiments demonstrate that our method produces rich-detailed photorealistic images, with 13 ~ 50 times faster than the state-of-the-art method (WCT2).
Yingxu Qiao, Jiabao Cui, Fuxian Huang, Hongmin Liu 0001, Cuizhu Bao, Xi Li 0001
IEEE Trans. Image Process.4
2021 Deep Unsupervised Binary Descriptor Learning Through Locality Consistency and Self Distinctiveness
abstract
Deep learning has been successfully applied to learn local feature descriptors in recent years. However, most of existing methods are supervised methods relying on a large number of labeled training patches, which are also proposed for learning real valued descriptors. In this paper, we propose a novel unsupervised deep learning method for binary descriptor learning. The binary descriptors are much more compact and efficient than the real valued descriptors and unsupervised leaning is highly required in many applications due to its label-free characteristic as the annotations are sometimes expensive to obtain. The core idea of our method is to explore the locality consistency in the descriptor space as well as to distinguish different patches while maintaining the ability to match a patch with its geometric transformed ones. We also give a theorical analysis about the role of batch normalization in learning effective binary descriptors. Benefited from this analysis, there is no need to append two additional losses on minimizing the quantization error and maximizing the entropy to the final learning objective like previous works did, thus simplifying our network training. Experiments on four benchmarks demonstrate that the proposed method is able to learn binary descriptors significantly outperforming previous unsupervised binary descriptors, even superior to most supervised ones. Especially, it obtains 21.2% of improvement on the UBC Phototour dataset, and 19.8%, 26.7%, 26.0% of improvements for patch verification, matching, retrieval tasks respectively on the HPatches dataset compared to the previous best unsupervised method.
Bin Fan 0001, Hongmin Liu 0001, Hui Zeng 0003, Jiyong Zhang 0001, Xin Liu 0027, Junwei Han 0001
IEEE Trans. Multim.2
2020 CoHomo: A cluster-attribute correlation aware graph clustering framework
Yaming Yang 0002, Hongmin Liu 0001, Ziyu Guan, Xiaofei He 0001, Gaoliang Liu
Neurocomputing2
2020 Efficient nearest neighbor search in high dimensional hamming space
Bin Fan 0001, Qingqun Kong, Baoqian Zhang, Hongmin Liu 0001, Chunhong Pan, Jiwen Lu
Pattern Recognit.4
2020 Depth-map completion for large indoor scene reconstruction
Hongmin Liu 0001, Xincheng Tang, Shuhan Shen
Pattern Recognit.1
2020 Improved covariant local feature detector
Zhanqiang Huo, Hongmin Liu 0001, Jing Wang 0093, Xin Liu 0027, Jiyong Zhang 0001
Pattern Recognit. Lett.3
2020 Features Combined Binary Descriptor Based on Voted Ring-Sampling Pattern
abstract
Most existing binary descriptors are only based on intensity and simply compare averaged intensities of one sample pair to obtain binary output. To address this issue, we propose a novel ring-sampling pattern based binary descriptor encoding both intensity and gradient. For intensity coding, a ring-sampling pattern is presented to define a number of sample points and their neighboring points given an interest point. The intensity difference between two sample points is obtained by performing the comparisons of the intensities of their corresponding neighboring points directly. Further, a majority based voting strategy is employed to obtain compact representation (1 bit or 3 bits) based on all these intensity difference among neighboring points. As for gradient coding, the gradient orientation histogram is computed for each sample point, and the gradient difference of two sample points is obtained by comparing the gradient magnitude on each orientation bin. The raw binary descriptor is constructed by concentrating the intensity and gradient differences of all the sample pairs, and two feature selection strategies are proposed to obtain the final compact descriptor, named as Features Combined Binary Descriptor based on Voted Ring-Sampling Pattern (BDVRP). The experimental results on the tasks of object recognition and image matching demonstrate the superiority and the effectiveness of the proposed descriptor, with comparison to the state-of-the-art hand-crafted binary descriptors.
Hongmin Liu 0001, Bin Fan 0001, Zhiheng Wang 0001, Junwei Han 0001
IEEE Trans. Circuits Syst. Video Technol.1
2020 Triplet Adversarial Domain Adaptation for Pixel-Level Classification of VHR Remote Sensing Images
abstract
Pixel-level classification for very high resolution (VHR) images is a crucial but challenging task in remote sensing. However, since the diverse ways of satellite image acquisition and the distinct structures of various regions, the distributions of the same semantic classes among different data sets are dissimilar. Therefore, the classification model trained on one data set (source domain) may collapse, when it is directly applied to another one (target domain). To solve this problem, many adversarial-based domain adaptation methods have been proposed. However, these methods only consider the source and the target domains independently in the adversarial training, where only the target domain is explicitly contributed to narrow the gap between the distributions of both domains. Unlike previous methods, we propose a triplet adversarial domain adaptation (TriADA) method that jointly considers both domains to learn a domain-invariant classifier by a novel domain similarity discriminator. Specifically, the discriminator takes a triplet of segmentation maps as input, where two segmentation maps from the same domain are to be distinguished from the two maps from the different domains during the adversarial learning. Consequently, it explicitly considers both domains' information to narrow the distribution gap across domains. To enhance the discriminability of the classifier on the target domain, a class-aware self-training strategy, which depends on the output of the discriminator, is proposed to assign pseudo-labels with high adapted confidence on target data to retrain the classifier. Extensive experiments on several VHR pixel-level classification benchmarks demonstrate the effectiveness of our method as well as its superiority to the-state of the art.
Bin Fan 0001, Hongmin Liu 0001, Chunlei Huo, Shiming Xiang, Chunhong Pan
IEEE Trans. Geosci. Remote. Sens.3
2019 Personalized Attraction Enhanced Sponsored Search with Multi-task Learning
abstract
We study a novel problem of sponsored search (SS) for E-Commerce platforms: how we can attract query users to click product advertisements (ads) by presenting them features of products that attract them. This not only benefits merchants and the platform, but also improves user experience. The problem is challenging due to the following reasons: (1) We need to carefully manipulate the ad content without affecting user search experience. (2) It is difficult to obtain users' explicit feedback of their preference in product features. (3) Nowadays, a great portion of the search traffic in E-Commerce platforms is from their mobile apps (e.g., nearly 90% in Taobao). The situation would get worse in the mobile setting due to limited space. We are focused on the mobile setting and propose to manipulate ad titles by adding a few selling point keywords (SPs) to attract query users. We model it as a personalized attractive SP prediction problem and carry out both large-scale offline evaluation and online A/B tests in Taobao. The contributions include: (1) We explore various exhibition schemes of SPs. (2) We propose a surrogate of user explicit feedback for SP preference. (3) We also explore multi-task learning and various additional features to boost the performance. A variant of our best model has already been deployed in Taobao, leading to a 2% increase in revenue per thousand impressions and an opt-out rate of merchants less than 4%.
Wei Zhao 0019, Boxuan Zhang 0002, Beidou Wang, Ziyu Guan, Wanxian Guan, Guang Qiu, Wei Ning, Jiming Chen 0001, Hongmin Liu 0001
KDD9
2018 Feature Matching Based on Triangle Guidance and Constraints
abstract
For images with distortions or repetitive patterns, the existing matching methods usually work well just on one of the two kinds of images. In this paper, we present novel triangle guidance and constraints (TGC)-based feature matching method, which can achieve good results on both kinds of images. We first extract stable matched feature points and combine these points into triangles as the initial matched triangles, and triangles combined by feature points are as the candidates to be matched. Then, triangle guidance based on the connection relationship via the shared feature point between the matched triangles and the candidates is defined to find the potential matching triangles. Triangle constraints, specially the location of a vertex relative to the inscribed circle center of the triangle, the scale represented by the ratio of corresponding side lengths of two matching triangles and the included angles between the sides of two triangles with connection relationship, are subsequently used to verify the potential matches and obtain the correct ones. Comparative experiments show that the proposed TGC can increase the number of the matched points with high accuracy under various image transformations, especially more effective on images with distortions or repetitive patterns due to the fact that the triangular structure are not only stable to image transformations but also provides more geometric constraints.
Hongmin Liu 0001, Hongya Zhang, Zhiheng Wang 0001
Int. J. Pattern Recognit. Artif. Intell.1
2018 Smart pathological brain detection by synthetic minority oversampling technique, extreme learning machine, and Jaya algorithm
Yudong Zhang 0001, Guihu Zhao, Junding Sun, Xiaosheng Wu, Zhiheng Wang 0001, Hongmin Liu 0001, Vishnuvarthanan Govindaraj, Tianming Zhan, Jianwu Li
Multim. Tools Appl.6
2017 STBD: a simple tri-bit binary descriptor for point matching
abstract
This study investigates the problem of constructing binary descriptor and develops a novel binary descriptor called simple tri‐bit binary descriptor (STBD) based on a simple sampling pattern (SSP) and a tri‐value binarisation strategy (TBS). First, an SSP is proposed, in which sample points are divided into two groups according to the distance from the pattern centre and smoothed by different circular filters. Then, to make the descriptor adaptive to the matched images, a selection strategy which directly employs detected keypoints as training data is introduced to select 256 point pairs with low correlation from initial pairs. Finally, a modified TBS method is presented to properly refine intensity comparison results. Experiments show that the proposed STBD can perform well and is robust to various transformations, except for scale change.
Hongmin Liu 0001, Zhiheng Wang 0001, Zhanqiang Huo
IET Comput. Vis.1
2016 Feature matching using guidance-constraint method
abstract
Image distortions and repetitive patterns widely exist in real images, which results in that feature matching is still a challenging problem though great progress has been made recently. This study presents a matching method, called guidance‐constraints method (GCM), which has obvious advantages in resolving problems of matching features on images with image distortions or repetitive patterns. In GCM, feature points are paired and connection compatibility is introduced to describe the relative geometric relations among features, and then potential matches are found by using the defined geometric guidance and are verified by using the defined geometric constraints. Experimental evaluation shows that the proposed GCM can significantly improve both the number of correct matches and correct ratio under various image transformations, especially more effective on images with distortions or containing repetitive patterns.
Zhiheng Wang 0001, Hongmin Liu 0001, Zhanqiang Huo, Bin Fan 0001
IET Comput. Vis.3
2015 Automatic detection of crop rows based on multi-ROIs
Guoquan Jiang, Zhiheng Wang 0001, Hongmin Liu 0001
Expert Syst. Appl.3
2015 Scale-invariant feature matching based on pairs of feature points
abstract
On the basis of feature points pairing, a scale‐invariant feature matching method is proposed in this study. The distance between two features is used to compute feature pairs' support region size, which is different from the methods using detectors to provide information to find the support region. Moreover, to achieve rotation invariance, a sub‐region division method based on intensity order is introduced. For comparison to the popular descriptors scale‐invariant feature transform and speeded‐up robust features, the authors also choose the detected points by difference of Gaussian and fast Hessain detectors as feature points to start the authors' method. Additional experiments compare the reported method with similar proposed methods, such as Tell's and Fan's. The experimental results show that the authors' proposed descriptor outperforms the popular descriptors under various image transformations, especially on images with scale and viewpoint transformations.
Zhiheng Wang 0001, Hongmin Liu 0001, Zhanqiang Huo
IET Comput. Vis.3
2014 Ternary contextualised histogram pattern for curve matching
abstract
This study presents a novel texture descriptor for curve matching, called ternary contextualised histogram pattern (TCHP), which is based on the intensity order histogram and local ternary homogeneity patterns. Local ternary homogeneity patterns are firstly adopted to represent the texture features of the curve's neighbourhood, which ensures the robustness and distinctiveness of the descriptor. The proposed TCHP is constructed by three steps: firstly, the curve support region without assigning a dominant orientation is determined and then partitioned into several ordinal bins according to the intensity permutation; then, ternary contextualised histogram (TCH) feature of each point is generated by computing the statistics of the predefined ternary homogeneity patterns; finally, TCHP is achieved by accumulating the TCHs of points in each order bin. Experiments show TCHP can effectively characterise the texture features of the curve's neighbourhood and performs robust to image rotation, viewpoint change and illumination change.
Shanshan Zhi, Hongmin Liu 0001, Zhiheng Wang 0001
IET Comput. Vis.2
2014 PLDD: Point-lines distance distribution for detection of arbitrary triangles, regular polygons and circles
Hongmin Liu 0001, Zhiheng Wang 0001
J. Vis. Commun. Image Represent.1
2013 Iocd: Intensity Order Curve Descriptor
abstract
Curve matching plays an important role in pattern recognition, computer vision and image understanding. In several past years, this problem has been studied mainly based on the curve contour, while only little progress has been made using the texture feature of the curve's neighborhood. This paper develops a novel texture-based curve matching method called IOCD, which consists of three steps: (1) Curve support region (CSR) without assigning a dominant orientation is first determined; (2) CSR is equally partitioned into several order bins according to the overall intensity order; (3) The feature vector is computed based on the local intensity order mapping. Experiments prove that the proposed IOCD performs robust to image rotation, viewpoint change, illumination change, blur, noise and JPEG compress. The application of image mosaic further identifies IOCD can achieve good matching performance.
Hongmin Liu 0001, Shanshan Zhi, Zhiheng Wang 0001
Int. J. Pattern Recognit. Artif. Intell.1
2013 Geometric property based ellipse detection method
Hongmin Liu 0001, Zhiheng Wang 0001
J. Vis. Commun. Image Represent.1
2009 Image Content Based Curve Matching Using HMCD Descriptor
Zhiheng Wang 0001, Hongmin Liu 0001, Fuchao Wu
ACCV (3)2
2009 HLD: A robust descriptor for line matching
abstract
Line matching plays an important role in many applications, such as image registration, 3D reconstruction, object recognition and video understanding. However, compared with other features (such as point, region) matching, it has made little progress in recent years. In this paper, we investigate the problem of matching line segments automatically only from their neighborhood appearance, without any constraints or priori knowledge. A novel line descriptor called HLD descriptor is proposed for this purpose, which is constructed by the following three steps: (1) Line support region (LSR) is defined, then divided into a series of overlapped sub-regions with the same size; (2) Line description matrix (DM) is formed by characterizing each sub-region into a vector; (3) HLD descriptor is built by computing mean and standard deviation of DM column vectors. Experimental results show that HLD descriptor is highly distinctive and very robust for line matching on real images.
Zhiheng Wang 0001, Hongmin Liu 0001, Fuchao Wu
CAD/Graphics2