Bo Li 0006

dblp:50/3402-6 · DBLP profile ↗
← Back
113ranked-venue papers
7as first author
31since 2021 · last 2026
0000-0001-5980-4861ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 60 · 1 first-author · 10 since 2021Artificial intelligence and machine learning · 36 · 2 first-author · 18 since 2021Applied, interdisciplinary, general and emerging computing · 20 · 3 first-author · 3 since 2021Computer networks · 3 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 A Comprehensive Survey of Knowledge Graph Reasoning: Approaches and Applications
abstract
Knowledge graph reasoning (KGR) aims to infer novel knowledge based on existing facts in knowledge graphs (KGs), playing a crucial role in various cognition intelligence systems across diverse domains. The previous review works explore KGR models from specific perspectives such as KG types and embedding spaces. In contrast, this survey provides a more comprehensive perspective of KGR from foundational approaches and their applications. Notably, some seldom-attended approaches such as negative sampling strategies, popular open-source libraries, and rule-guided KGR paradigms are carefully reviewed. Besides, we explore advanced techniques, such as large language models (LLMs) and their impact on KGR. The comparison among foundational models are analyzed to declare their strengths and limitations. More interestingly, this is the first effort to provide a taxonomy of real-world KGR applications for both horizontal and vertical domains. Furthermore, we highlight the challenges and opportunities in the field of KGR, including trustworthiness, multimodal reasoning, continual learning, uncertainty and LLM-driven approaches. This work aims to bridge the gap between theoretical advancements and practical deployment of KGR models, and outline promising future directions.
Guanglin Niu, Bo Li 0006, Yangguang Lin
IEEE Trans. Big Data2
2025 Identity-aware Feature Decoupling Learning for Clothing-change Person Re-identification
abstract
Clothing-change person re-identification (CC Re-ID) has attracted increasing attention in recent years due to its application prospect. Most existing works struggle to adequately extract the ID-related information from the original RGB images. In this paper, we propose an Identity-aware Feature Decoupling (IFD) learning framework to mine identity-related features. Particularly, IFD exploits a dual stream architecture that consists of a main stream and an attention stream. The attention stream takes the clothing-masked images as inputs and derives the identity attention weights for effectively transferring the spatial knowledge to the main stream and highlighting the regions with abundant identity-related information. To eliminate the semantic gap between the inputs of two streams, we propose a clothing bias diminishing module specific to the main stream to regularize the features of clothing-relevant regions. Extensive experimental results demonstrate that our framework outperforms other baseline models on several widely-used CC Re-ID datasets.
Bo Li 0006, Guanglin Niu
ICASSP2
2025 RGB-T Tracking With Template-Bridged Search Interaction and Target-Preserved Template Updating
abstract
The goal of RGB-Thermal (RGB-T) tracking is to utilize the synergistic and complementary strengths of RGB and TIR modalities to enhance tracking in diverse situations, with cross-modal interaction being a crucial element. Earlier methods often simply combine the features of the RGB and TIR search frames, leading to a coarse interaction that also introduced unnecessary background noise. Many other approaches sample candidate boxes from search frames and apply different fusion techniques to individual pairs of RGB and TIR boxes, which confines cross-modal interactions to local areas and results in insufficient context modeling. Additionally, mining video temporal contexts is also under-explored in RGB-T tracking. To alleviate these limitations, we propose a novel Template-Bridged Search region Interaction (TBSI) module that exploits templates as the medium to bridge the cross-modal interaction between RGB and TIR search regions by gathering and distributing target-relevant object and environment contexts. An Illumination Guided Fusion (IGF) module is designed to adaptively fuse RGB and TIR search region tokens with a global illumination factor. Furthermore, in the inference stage, we also propose an efficient Target-Preserved Template Updating (TPTU) strategy, leveraging the temporal context within video sequences to accommodate the target's appearance change. Our proposed modules are integrated into a ViT backbone for joint feature extraction, search-template matching, and cross-modal interaction. Extensive experiments on three popular RGB-T tracking benchmarks demonstrate our method achieves new state-of-the-art performances.
Bo Li 0006, Fengguang Peng, Tianrui Hui, Xiaoming Wei, Xiaolin Wei, Si Liu 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2025 Heterogeneous Feature Re-Sampling for Balanced Pedestrian Attribute Recognition
abstract
In pedestrian attribute recognition (PAR), the loose umbrella term 'attribute' ranges from human soft-biometrics to wearing accessory, and even extending to various subjective body descriptors. As a result, the vast coverage of 'attributes' implies that, instead of being over-specialized to limited attributes with exclusive characteristic, PAR should be approached from a much fundamental perspective. To this end, given that most attributes are greatly under-represented in real-world datasets, we simply distill PAR into a visual task of multi-label recognition under significant data imbalance. Accordingly, we introduce feature re-sampled detached learning (FRDL) to decouple label-balanced learning from the curse of attributes co-occurrence. Specifically, FRDL is able to balance the sampling distribution of an attribute without biasing the label prior of co-occurring others. As a complementary method, we also propose gradient-oriented augment translating (GOAT) to alleviate the feature noise and semantics imbalance aggravated in FRDL. Integrated in a highly unified framework, FRDL and GOAT substantially refresh the state-of-the-art performance on various realistic benchmarks, while maintaining a minimal computational budget. Further analytical discussion and experimental evidence corroborate the veracity of our advancement: this is the first work that establishes labels-independent and impartial balanced learning for PAR.
Bo Li 0006, Hai-Miao Hu, Hanzi Wang
IEEE Trans. Pattern Anal. Mach. Intell.2
2025 A Pluggable Common Sense-Enhanced Framework for Knowledge Graph Completion
abstract
Knowledge graph completion (KGC) tasks aim to infer missing facts in a knowledge graph (KG) for many knowledgeintensive applications. However, existing embedding-based KGC approaches primarily rely on factual triples, potentially leading to outcomes inconsistent with common sense. Besides, generating explicit common sense is often impractical or costly for a KG. To address these challenges, we propose a pluggable common sense-enhanced KGC framework that incorporates both fact and common sense for KGC. This framework is adaptable to different KGs based on their entity concept richness and has the capability to automatically generate explicit or implicit common sense from factual triples. Furthermore, we introduce common senseguided negative sampling and a coarse-to-fine inference approach for KGs with rich entity concepts. For KGs without concepts, we propose a dual scoring scheme involving a relation-aware concept embedding mechanism. Importantly, our approach can be integrated as a pluggable module for many knowledge graph embedding (KGE) models, facilitating joint common sense and fact-driven training and inference. The experiments illustrate that our framework exhibits good scalability and outperforms existing models across various KGC tasks.
Guanglin Niu, Bo Li 0006, Siling Feng
IEEE Trans. Big Data2
2025 Classification Committee for Active Deep Object Detection
abstract
In object detection, the cost of labeling is very high because it needs not only to confirm the categories of multiple objects in an image but also to determine the bounding boxes of each object accurately. Thus, integrating active learning into object detection will raise pretty positive significance. In this paper, we propose a classification committee for the active deep object detection method by introducing a discrepancy mechanism of multiple classifiers for samples' selection when training object detectors. The model contains a main detector and a classification committee. The main detector denotes the target object detector trained from a labeled pool composed of the selected informative images. The role of the classification committee is to select the most informative images according to their uncertainty values from the view of classification, which is expected to focus more on the discrepancy and representative of instances. Specifically, they compute the uncertainty for a specified instance within the image by measuring its discrepancy output by the committee pre-trained via the proposed Maximum Classifiers Discrepancy Group Loss (MCDGL). The most informative images are finally determined by selecting the ones with many high-uncertainty instances. Besides, to mitigate the impact of interference instances, we design a Focusing on Positive Instances Loss (FPIL) to provide the committee the ability to automatically focus on the representative instances as well as precisely encode their discrepancies for the same instance. Experiments are conducted on Pascal VOC and COCO datasets versus some popular object detectors. And results show that our method outperforms the state-of-the-art active learning methods, which verifies the effectiveness of the proposed method.
Lei Zhao 0025, Bo Li 0006, Jixiang Jiang, Xingxing Wei 0001
IEEE Trans. Multim.2
2025 Diverse Visible-to-Thermal Image Translation via Controllable Temperature Encoding
abstract
Translating readily available visible (VIS) images into thermal infrared (TIR) images effectively alleviates the shortage of TIR data. While current methods have yielded commendable results, they fall short in generating diverse and realistic thermal infrared images, primarily due to insufficient consideration of temperature variations. In this paper, we propose a Thermally Controlled GAN (TC-GAN) that leverages VIS images to generate diverse TIR images, with the ability to control the relative temperatures of multiple objects, particularly those with temperature variations. Firstly, we introduce the physical coding module, which employs a conditional variational autoencoder GAN to learn the distributions of relative temperature information for the objects and environmental state information. Then, the physical information can be obtained by sampling the distribution. When this information is fused with the visible image, it facilitates the generation of diverse TIR images. To ensure authenticity and strengthen the physical constraints across different regions of the image, we introduce a self-attention mechanism in the generator that prioritizes the relative temperature relationships within the image. Additionally, we utilize a local discriminator that focuses on objects with actively changing temperatures and their interactions with the surrounding environment, thereby reducing the discontinuity between the target and the background. Experiments on the Drone Vehicle and AVIID datasets show that our approach outperforms mainstream diversity generation methods in terms of authenticity and diversity.
Lei Zhao 0025, Bo Li 0006, Xingxing Wei 0001
IEEE Trans. Multim.3
2024 Room-Object Entity Prompting and Reasoning for Embodied Referring Expression
abstract
Given a high-level instruction, the task of Embodied Referring Expression (REVERIE) requires an embodied agent to localise a remote referred object via navigating in the unseen environment. Previous vision-language navigation methods utilise the provided fine-grained instruction as step-by-step navigation guidance to conduct strict instruction-following, while REVERIE aims to achieve efficient goal-oriented exploration according to the high-level command. In this work, we propose a Cross-modal Knowledge Reasoning (abbreviated as CKR+) framework, which incorporates the prior knowledge as decision guidance to learn the navigation scheme comprehensively. Specifically, we design a Room-Object Aware (ROA) mechanism to explicitly decouple the room- and object-related clues from instruction and visual observations. Moreover, we propose a Knowledge-enabled Entity Relation Reasoning (KERR+) module to leverage the structured knowledge from the knowledge graph explicitly and unstructured knowledge from pre-trained model implicitly, to learn the internal-external correlations among room- and object-entities for the agent to make proper decisions. We devise an Entity Prompter (EP) that embeds in the KERR+ module, which utilises the navigation history and visual entities as prompts to transfer knowledge from the pre-trained CLIP model. In addition, we develop a Reinforced End Decider (RED) to learn the stopping scheme specifically, which is achieved by a customised reinforcement learning strategy and knowledge enhanced matching. Two techniques are also introduced to improve navigation efficiency further. Extensive experiments conducted on the REVERIE benchmark demonstrate the effectiveness and superiority of our proposed methods, which boosts the key metrics, i.e., SPL and REVERIE-success rate, to 14.46% and 13.81% respectively.
Chen Gao 0005, Si Liu 0001, Luting Wang 0001, Qi Wu 0001, Bo Li 0006, Qi Tian 0001
IEEE Trans. Pattern Anal. Mach. Intell.6
2024 PPDM++: Parallel Point Detection and Matching for Fast and Accurate HOI Detection
abstract
Human-Object Interaction (HOI) detection aims to understand human activities by detecting interaction triplets. Previous HOI detection methods adopt a two-stage instance-driven paradigm. Unfortunately, many non-interactive human-object pairs generated by the first stage are the main obstacle impeding HOI detectors from high efficiency and promising performance. To remedy this, we propose a novel top-down interaction-driven paradigm, detecting interactions first and bridging interactive human-object pairs through interactions. We formulate HOI as a point triplet human point, interaction point, object point and design a Parallel Point Detection and Matching (PPDM) framework. We further take advantage of two-stage methods and propose a novel framework, PPDM++, that detects the interactive human-object pairs by PPDM, then extracts region features for each pair to predict actions. The core of PPDM/PPDM++ is to convert the instance-driven bottom-up paradigm to an interaction-driven top-down paradigm, thus avoiding additional computation costs from traversing a tremendous number of non-interactive pairs. Benefiting from the advanced paradigm, PPDM/PPDM++ has achieved significant performance gains with high efficiency. PPDM-DLA-34 has achieved 19.94 mAP with 42 FPS as the first real-time HOI detector, and PPDM++-SwinB achieves 30.1 mAP with 17 FPS on HICO-DET dataset. We also built an application-oriented database named HOI-A, a supplement to the existing datasets.
Yue Liao, Si Liu 0001, Yulu Gao, Aixi Zhang, Fei Wang 0032, Bo Li 0006
IEEE Trans. Pattern Anal. Mach. Intell.7
2024 Revisiting the Trade-Off Between Accuracy and Robustness via Weight Distribution of Filters
abstract
Adversarial attacks have been proven to be potential threats to Deep Neural Networks (DNNs), and many methods are proposed to defend against adversarial attacks. However, while enhancing the robustness, the accuracy for clean examples will decline to a certain extent, implying a trade-off existed between the accuracy and adversarial robustness. In this paper, to meet the trade-off problem, we theoretically explore the underlying reason for the difference of the filters' weight distribution between standard-trained and robust-trained models and then argue that this is an intrinsic property for static neural networks, thus they are difficult to fundamentally improve the accuracy and adversarial robustness at the same time. Based on this analysis, we propose a sample-wise dynamic network architecture named Adversarial Weight-Varied Network (AW-Net), which focuses on dealing with clean and adversarial examples with a "divide and rule" weight strategy. The AW-Net adaptively adjusts the network's weights based on regulation signals generated by an adversarial router, which is directly influenced by the input sample. Benefiting from the dynamic network architecture, clean and adversarial examples can be processed with different network weights, which provides the potential to enhance both accuracy and adversarial robustness. A series of experiments demonstrate that our AW-Net is architecture-friendly to handle both clean and adversarial examples and can achieve better trade-off performance than state-of-the-art robust models.
Xingxing Wei 0001, Bo Li 0006
IEEE Trans. Pattern Anal. Mach. Intell.3
2024 Multi-Person Pose Regression With Distribution-Aware Single-Stage Models
abstract
Understanding human posture is a challenging topic, which encompasses several tasks, e.g., pose estimation, body mesh recovery and pose tracking. In this article, we propose a novel Distribution-Aware Single-stage (DAS) model for the pose-related tasks. The proposed DAS model estimates human position and localizes joints simultaneously, which requires only a single pass. Meanwhile, we utilize normalizing flow to enable DAS to learn the true distribution of joint locations, rather than making simple Gaussian or Laplacian assumptions. This provides a pivotal prior and greatly boosts the accuracy of regression-based methods, thus making DAS achieve comparable performance to the volumetric-based methods. We also introduce a recursively update strategy to progressively approach the regression target, reducing the difficulty of regression and improving the regression performance. We further adapt DAS to multi-person mesh recovery and pose tracking tasks and achieve considerable performance on both tasks. Comprehensive experiments on CMU Panoptic and MuPoTS-3D demonstrate the superior efficiency of DAS, specifically 1.5 times speedup over previous best method, and its state-of-the-art accuracy for multi-person pose estimation. Extensive experiments on 3DPW and PoseTrack2018 indicate the effectiveness and efficiency of DAS for human body mesh recovery and pose tracking, respectively, which prove the generality of our proposed DAS model.
Leyan Zhu, Zitian Wang, Si Liu 0001, Xuecheng Nie, Luoqi Liu, Bo Li 0006
IEEE Trans. Pattern Anal. Mach. Intell.6
2024 Cross-Modal Contrastive Learning Network for Few-Shot Action Recognition
abstract
Few-shot action recognition aims to recognize new unseen categories with only a few labeled samples of each class. However, it still suffers from the limitation of inadequate data, which easily leads to the overfitting and low-generalization problems. Therefore, we propose a cross-modal contrastive learning network (CCLN), consisting of an adversarial branch and a contrastive branch, to perform effective few-shot action recognition. In the adversarial branch, we elaborately design a prototypical generative adversarial network (PGAN) to obtain synthesized samples for increasing training samples, which can mitigate the data scarcity problem and thereby alleviate the overfitting problem. When the training samples are limited, the obtained visual features are usually suboptimal for video understanding as they lack discriminative information. To address this issue, in the contrastive branch, we propose a cross-modal contrastive learning module (CCLM) to obtain discriminative feature representations of samples with the help of semantic information, which can enable the network to enhance the feature learning ability at the class-level. Moreover, since videos contain crucial sequences and ordering information, thus we introduce a spatial-temporal enhancement module (SEM) to model the spatial context within video frames and the temporal context across video frames. The experimental results show that the proposed CCLN outperforms the state-of-the-art few-shot action recognition methods on four challenging benchmarks, including Kinetics, UCF101, HMDB51 and SSv2.
Xiao Wang 0072, Yan Yan 0001, Hai-Miao Hu, Bo Li 0006, Hanzi Wang
IEEE Trans. Image Process.4
2024 Orientation-Aware Pedestrian Attribute Recognition Based on Graph Convolution Network
abstract
Pedestrian attribute recognition (PAR) aims to generate a structured description of pedestrians and plays an important role in surveillance. Current work focusing on 2D images can achieve decent performance when there is no variation in the captured pedestrian orientation. However, the performance of these works cannot be maintained in scenarios when the orientation of pedestrians is ignored. To mitigate this problem, this paper proposes orientation-aware pedestrian attribute recognition based on graph convolution network (GCN), which is composed of an orientation-aware spatial attention (OSA) module and an orientation-guided attribute-relation learning (OAL) module. Since some attributes can be invisible for certain orientations, OSA is proposed for orientation-aware feature extraction to enhance the learned representation of the visual attributes. Moreover, since different orientations result in different relations among attributes, OAL is proposed to achieve distinguishable and impactful attribute relations by eliminating the confusion of attribute relations in different orientations. Experiments on three challenging datasets (PETA, RAP, and PA100K) demonstrate that the proposed PAR outperforms the state-of-the-art methods by considerable margins.
Weiqing Lu, Hai-Miao Hu, Jinzuo Yu, Hanzi Wang, Bo Li 0006
IEEE Trans. Multim.6
2024 Linker: Learning Long Short-term Associations for Robust Visual Tracking
abstract
iamese and Transformer trackers have demon strated exceptional performance in visual object tracking. These methods utilize initial and potentially online templates to locate the target in subsequent frames. Despite their success, these trackers are vulnerable to changes in the target's appearance due to slow template updates and interference from similar objects, resulting from the absence of scene information. To address these issues, we introduce a reference region within our tracker. The reference region is updated rapidly, providing short-term scene information. By associating the initial template, reference region, and current search region, we enhance the tracker's ability to adapt to changes in target appearance and discriminate between the target and other objects. Additionally, we propose a novel Reference-Enhance (RE) module, which aggregates contextually relevant information from the reference region to enhance the template feature. Extensive experiments show our method achieves state-of-the-art performance on six popular visual object tracking benchmarks while running at over 40 FPS.
Zizheng Xun, Shangzhe Di, Yulu Gao, Zongheng Tang, Gang Wang 0031, Si Liu 0001, Bo Li 0006
IEEE Trans. Multim.7
2023 PHA: Patch-Wise High-Frequency Augmentation for Transformer-Based Person Re-Identification
abstract
Although recent studies empirically show that injecting Convolutional Neural Networks (CNNs) into Vision Transformers (ViTs) can improve the performance of person reidentification, the rationale behind it remains elusive. From a frequency perspective, we reveal that ViTs perform worse than CNNs in preserving key high-frequency components (e.g, clothes texture details) since high-frequency components are inevitably diluted by low-frequency ones due to the intrinsic Self-Attention within ViTs. To remedy such inadequacy of the ViT, we propose a Patch-wise High-frequency Augmentation (PHA) method with two core designs. First, to enhance the feature representation ability of high-frequency components, we split patches with high-frequency components by the Discrete Haar Wavelet Transform, then empower the ViT to take the split patches as auxiliary input. Second, to prevent high-frequency components from being diluted by low-frequency ones when taking the entire sequence as input during network optimization, we propose a novel patch-wise contrastive loss. From the view of gradient optimization, it acts as an implicit augmentation to improve the representation ability of key high-frequency components. This benefits the ViT to capture key high-frequency components to extract discriminative person representations. PHA is necessary during training and can be removed during inference, without bringing extra complexity. Extensive experiments on widely-used ReID datasets validate the effectiveness of our method.
Guiwei Zhang, Yongfei Zhang, Bo Li 0006, Shiliang Pu
CVPR4
2022 CAKE: A Scalable Commonsense-Aware Framework For Multi-View Knowledge Graph Completion
abstract
Knowledge graphs store a large number of factual triples while they are still incomplete, inevitably.The previous knowledge graph completion (KGC) models predict missing links between entities merely relying on fact-view data, ignoring the valuable commonsense knowledge.The previous knowledge graph embedding (KGE) techniques suffer from invalid negative sampling and the uncertainty of fact-view link prediction, limiting KGC's performance.To address the above challenges, we propose a novel and scalable Commonsense-Aware Knowledge Embedding (CAKE) framework to automatically extract commonsense from factual triples with entity concepts.The generated commonsense augments effective selfsupervision to facilitate both high-quality negative sampling (NS) and joint commonsense and fact-view link prediction.Experimental results 1 on the KGC task demonstrate that assembling our framework could enhance the performance of the original KGE models, and the proposed commonsense-aware NS module is superior to other NS techniques.Besides, our proposed framework could be easily adaptive to various KGE models and explain the predicted results.
Guanglin Niu, Bo Li 0006, Yongfei Zhang, Shiliang Pu
ACL (1)2
2022 Perform like an Engine: A Closed-Loop Neural-Symbolic Learning Framework for Knowledge Graph Inference
abstract
Knowledge graph (KG) inference aims to address the natural incompleteness of KGs, including rule learning-based and KG embedding (KGE) models. However, the rule learning-based models suffer from low efficiency and generalization while KGE models lack interpretability. To address these challenges, we propose a novel and effective closed-loop neural-symbolic learning framework EngineKG via incorporating our developed KGE and rule learning modules. KGE module exploits symbolic rules and paths to enhance the semantic association between entities and relations for improving KG embeddings and interpretability. A novel rule pruning mechanism is proposed in the rule learning module by leveraging paths as initial candidate rules and employing KG embeddings together with concepts for extracting more high-quality rules. Experimental results on four real-world datasets show that our model outperforms the relevant baselines on link prediction tasks, demonstrating the superiority of our KG inference model in a neural-symbolic learning fashion. The source code and datasets of this paper are available at https://github.com/ngl567/EngineKG.
Guanglin Niu, Bo Li 0006, Yongfei Zhang, Shiliang Pu
COLING2
2022 POF-Based Dynamic Control in Wireless Tactical Networks
abstract
In tactical networks, link fluctuations and node movements make operation management a difficult problem. Software-defined networking (SDN) has the ability to manage wired networks, but for tactical networks, SDN faces challenges arising from unreliable transmission and inefficient operation of wireless links. Moreover, protocols in wireless tactical networks are heterogeneous to implement different tasks, while the existing Openflow protocol cannot handle such protocol heterogeneity. In this paper, we proposed a reliable dynamic control architecture based on protocol-oblivious forwarding (POF) in the wireless tactical network, which can increase the reliability of the control plane and the communication performance between switches of the data plane. Based on the proposed architecture, we customized a routing protocol, which can make data forwarding more efficient. Finally, we build a test platform based on Raspberry Pi to validate the architecture of the proposed architecture. The results demonstrate the proposed architecture achieves desired performance in dynamic tactical environments.
Bo Li 0006, Hancheng Lu, Zhuojia Gu, Haoyue Yuan
ICC1
2022 Sparse Black-Box Video Attack with Reinforcement Learning
Xingxing Wei 0001, Huanqian Yan, Bo Li 0006
Int. J. Comput. Vis.3
2022 Joint semantics and data-driven path representation for knowledge graph reasoning
Guanglin Niu, Bo Li 0006, Yongfei Zhang, Yongpan Sheng, Chuan Shi 0001, Shiliang Pu
Neurocomputing2
2022 Cross-Modal Progressive Comprehension for Referring Segmentation
abstract
Given a natural language expression and an image/video, the goal of referring segmentation is to produce the pixel-level masks of the entities described by the subject of the expression. Previous approaches tackle this problem by implicit feature interaction and fusion between visual and linguistic modalities in a one-stage manner. However, human tends to solve the referring problem in a progressive manner based on informative words in the expression, i.e., first roughly locating candidate entities and then distinguishing the target one. In this paper, we propose a cross-modal progressive comprehension (CMPC) scheme to effectively mimic human behaviors and implement it as a CMPC-I (Image) module and a CMPC-V (Video) module to improve referring image and video segmentation models. For image data, our CMPC-I module first employs entity and attribute words to perceive all the related entities that might be considered by the expression. Then, the relational words are adopted to highlight the target entity as well as suppress other irrelevant ones by spatial graph reasoning. For video data, our CMPC-V module further exploits action words based on CMPC-I to highlight the correct entity matched with the action cues by temporal graph reasoning. In addition to the CMPC, we also introduce a simple yet effective Text-Guided Feature Exchange (TGFE) module to integrate the reasoned multimodal features corresponding to different levels in the visual backbone under the guidance of textual information. In this way, multi-level features can communicate with each other and be mutually refined based on the textual context. Combining CMPC-I or CMPC-V with TGFE can form our image or video version referring segmentation frameworks and our frameworks achieve new state-of-the-art performances on four referring image segmentation benchmarks and three referring video segmentation benchmarks respectively. Our code is available at https://github.com/spyflying/CMPC-Refseg.
Si Liu 0001, Tianrui Hui, Shaofei Huang 0001, Yunchao Wei, Bo Li 0006, Guanbin Li
IEEE Trans. Pattern Anal. Mach. Intell.5
2022 PSGAN++: Robust Detail-Preserving Makeup Transfer and Removal
abstract
In this paper, we address the makeup transfer and removal tasks simultaneously, which aim to transfer the makeup from a reference image to a source image and remove the makeup from the with-makeup image respectively. Existing methods have achieved much advancement in constrained scenarios, but it is still very challenging for them to transfer makeup between images with large pose and expression differences, or handle makeup details like blush on cheeks or highlight on the nose. In addition, they are hardly able to control the degree of makeup during transferring or to transfer a specified part in the input face. These defects limit the application of previous makeup transfer methods to real-world scenarios. In this work, we propose a Pose and expression robust Spatial-aware GAN (abbreviated as PSGAN++). PSGAN++ is capable of performing both detail-preserving makeup transfer and effective makeup removal. For makeup transfer, PSGAN++ uses a Makeup Distill Network (MDNet) to extract makeup information, which is embedded into spatial-aware makeup matrices. We also devise an Attentive Makeup Morphing (AMM) module that specifies how the makeup in the source image is morphed from the reference image, and a makeup detail loss to supervise the model within the selected makeup detail area. On the other hand, for makeup removal, PSGAN++ applies an Identity Distill Network (IDNet) to embed the identity information from with-makeup images into identity matrices. Finally, the obtained makeup/identity matrices are fed to a Style Transfer Network (STNet) that is able to edit the feature maps to achieve makeup transfer or removal. To evaluate the effectiveness of our PSGAN++, we collect a Makeup Transfer In the Wild (MT-Wild) dataset that contains images with diverse poses and expressions and a Makeup Transfer High-Resolution (MT-HR) dataset that contains high-resolution images. Experiments demonstrate that PSGAN++ not only achieves state-of-the-art results with fine makeup details even in cases of large pose/expression differences but also can perform partial or degree-controllable makeup transfer. Both the code and the newly collected datasets will be released at https://github.com/wtjiang98/PSGAN.
Si Liu 0001, Chen Gao 0005, Ran He 0001, Jiashi Feng, Bo Li 0006, Shuicheng Yan
IEEE Trans. Pattern Anal. Mach. Intell.6
2022 Fine-Grained Human-Centric Tracklet Segmentation with Single Frame Supervision
abstract
In this paper, we target at the Fine-grAined human-Centric Tracklet Segmentation (FACTS) problem, where 12 human parts, e.g., face, pants, left-leg, are segmented. To reduce the heavy and tedious labeling efforts, FACTS requires only one labeled frame per video during training. The small size of human parts and the labeling scarcity makes FACTS very challenging. Considering adjacent frames of videos are continuous and human usually do not change clothes in a short time, we explicitly consider the pixel-level and frame-level context in the proposed Temporal Context segmentation Network (TCNet). On the one hand, optical flow is on-line calculated to propagate the pixel-level segmentation results to neighboring frames. On the other hand, frame-level classification likelihood vectors are also propagated to nearby frames. By fully exploiting the pixel-level and frame-level context, TCNet indirectly uses the large amount of unlabeled frames during training and produces smooth segmentation results during inference. Experimental results on four video datasets show the superiority of TCNet over the state-of-the-arts. The newly annotated datasets can be downloaded via http://liusi-group.com/projects/FACTS for the further studies.
Si Liu 0001, Guanghui Ren, Yao Sun 0004, Jinqiao Wang, Changhu Wang, Bo Li 0006, Shuicheng Yan
IEEE Trans. Pattern Anal. Mach. Intell.6
2022 Human-Centric Relation Segmentation: Dataset and Solution
abstract
Vision and language understanding techniques have achieved remarkable progress, but currently it is still difficult to well handle problems involving very fine-grained details. For example, when the robot is told to "bring me the book in the girl's left hand", most existing methods would fail if the girl holds one book respectively in her left and right hand. In this work, we introduce a new task named human-centric relation segmentation (HRS), as a fine-grained case of HOI-det. HRS aims to predict the relations between the human and surrounding entities and identify the relation-correlated human parts, which are represented as pixel-level masks. For the above exemplar case, our HRS task produces results in the form of relation triplets 〈girl [left hand], hold, book 〉 and exacts segmentation masks of the book, with which the robot can easily accomplish the grabbing task. Correspondingly, we collect a new Person In Context (PIC) dataset for this new task, which contains 17,122 high-resolution images and densely annotated entity segmentation and relations, including 141 object categories, 23 relation categories and 25 semantic human parts. We also propose a Simultaneous Matching and Segmentation (SMS) framework as a solution to the HRS task. It contains three parallel branches for entity segmentation, subject object matching and human parsing respectively. Specifically, the entity segmentation branch obtains entity masks by dynamically-generated conditional convolutions; the subject object matching branch detects the existence of any relations, links the corresponding subjects and objects by displacement estimation and classifies the interacted human parts; and the human parsing branch generates the pixelwise human part labels. Outputs of the three branches are fused to produce the final HRS results. Extensive experiments on PIC and V-COCO datasets show that the proposed SMS method outperforms baselines with the 36 FPS inference speed. Notably, SMS outperforms the best performing baseline m-KERN with only 17.6 percent time cost. The dataset and code will be released at http://picdataset.com/challenge/index/.
Si Liu 0001, Zitian Wang, Yulu Gao, Lejian Ren, Yue Liao, Guanghui Ren, Bo Li 0006, Shuicheng Yan
IEEE Trans. Pattern Anal. Mach. Intell.7
2021 UnrealPerson: An Adaptive Pipeline Towards Costless Person Re-Identification
abstract
The main difficulty of person re-identification (ReID) lies in collecting annotated data and transferring the model across different domains. This paper presents UnrealPerson, a novel pipeline that makes full use of unreal image data to decrease the costs in both the training and deployment stages. Its fundamental part is a system that can generate synthesized images of high-quality and from controllable distributions. Instance-level annotation goes with the synthesized data and is almost free. We point out some details in image synthesis that largely impact the data quality. With 3,000 IDs and 120,000 instances, our method achieves a 38.5% rank-1 accuracy when being directly transferred to MSMT17. It almost doubles the former record using synthesized data and even surpasses previous direct transfer records using real data. This offers a good basis for unsupervised domain adaption, where our pre-trained model is easily plugged into the state-of-the-art algorithms towards higher accuracy. In addition, the data distribution can be flexibly adjusted to fit some corner ReID scenarios, which widens the application of our pipeline. We publish our data synthesis toolkit and synthesized data in https://github.com/FlyHighest/UnrealPerson.
Lingxi Xie, Longhui Wei, Zijie Zhuang, Yongfei Zhang, Bo Li 0006, Qi Tian 0001
CVPR6
2021 Black-box adversarial attacks by manipulating image attributes
Xingxing Wei 0001, Ying Guo 0008, Bo Li 0006
Inf. Sci.3
2021 Ship Detection in Spaceborne Infrared Image Based on Lightweight CNN and Multisource Feature Cascade Decision
abstract
Infrared remote-sensing images have irreplaceable value in military and civilian research, such as remote surveillance and military reconnaissance. However, under the conditions of complex scenes, infrared ship detection still faces great challenges. Most importantly, because of the limited hardware resource, the spaceborne satellites usually have weak computational processing ability, which makes the traditional convolution neural network (CNN)-based detection algorithms difficult to show their power. In terms of the above facts, this article proposes a high-performance but low-computation and storage-efficient ship detection algorithm to adapt the severe spaceborne environment. Overall, our method contains the following technical steps: 1) we first present a novel iterative precise segmentation of land and sea algorithm to preprocess the complex and diverse remote-sensing scenes; 2) the multivariate Gaussian distribution is then selected to extract the ship target candidate regions to guarantee the detection recall; 3) we adopt the optical panchromatic data to assist the limited infrared data training; and 4) the cascade decision of multisource features including global and local cues is next utilized to gradually eliminate false alarms. Due to the high efficiency, the proposed method can implement well on the hardware platform of DSP and field-programmable gate array (FPGA) architecture. We conduct a series of experiments and compare with the state-of-the-art object detection algorithms. Experimental results show that our method has fewer parameters but can achieve strong detection robustness against the noise, cloud and reef interfere, which verifies the effectiveness of the proposed method.
Nan Wang 0014, Bo Li 0006, Xingxing Wei 0001, Huanqian Yan
IEEE Trans. Geosci. Remote. Sens.2
2021 Hierarchical Reasoning Network for Pedestrian Attribute Recognition
abstract
Pedestrian attribute recognition, which can benefit other tasks such as person re-identification and pedestrian retrieval, is very important in video surveillance related tasks. In this paper, we observe that the existing methods tackle this problem from the perspective of multi-label classification without considering the hierarchical relationships among the attributes. In human cognition, the attributes can be categorized according to their semantic/abstraction levels. The high-level attributes can be predicted by reasoning from the low-level and medium-level attributes, while the recognition of the low-level and medium-level attributes can be guided by the high-level attributes. Based on this attribute categorization, we propose a novel Hierarchical Reasoning Network (HR-Net), which can hierarchically predict the attributes at different abstraction levels in different stages of the network. We also propose an attribute reasoning structure to exploit the relationships among the attributes at different semantic levels. Experimental results demonstrate that the proposed network gives superior performances compared to the state-of-the-art techniques.
Haoran An, Hai-Miao Hu, Yuanfang Guo, Qianli Zhou, Bo Li 0006
IEEE Trans. Multim.5
2021 Spectrum Characteristics Preserved Visible and Near-Infrared Image Fusion Algorithm
abstract
The visible and near-infrared images fusion aims at utilizing their spectrum characteristics to enhance visibility. However, the current visible and near-infrared fusion algorithms cannot well preserve spectrum characteristics, which results in color distortion and halo artifacts. Therefore, this paper proposes a new visible and near infrared images fusion algorithm by fully considering their different reflection and scattering characteristics. According to image degradation model, the reflection weight model and the transmission weight model are established, respectively. The reflection weight model is established by calculating the difference between the visible (red, green, and blue) spectra and the near-infrared spectrum, while maintaining the correlation of the visible spectra. The proposed reflection weight model can preserve the original reflection characteristic of objects in natural scenes. On the other hand, the transmission weight model is explicitly proposed by calculating the gradient ratio of the visible spectra to the near-infrared spectrum. The proposed transmission weight model intends to make full use of the strong transmission performance of the near-infrared spectrum, which can complement the details loss of the visible spectra caused by light scattering. Moreover, the fused image based on two models is further enhanced according to the reflection characteristics of near-infrared spectrum in case of the non-uniform illumination. The experimental results demonstrate that the proposed algorithm can not only well preserve spectrum characteristics, but also avoid color distortion while maintaining the naturalness, which outperforms the state-of-the-art.
Hai-Miao Hu, Wei Zhang 0181, Shiliang Pu, Bo Li 0006
IEEE Trans. Multim.5
2021 Naive Gabor Networks for Hyperspectral Image Classification
abstract
Recently, many convolutional neural network (CNN) methods have been designed for hyperspectral image (HSI) classification since CNNs are able to produce good representations of data, which greatly benefits from a huge number of parameters. However, solving such a high-dimensional optimization problem often requires a large number of training samples in order to avoid overfitting. In addition, it is a typical nonconvex problem affected by many local minima and flat regions. To address these problems, in this article, we introduce the naive Gabor networks or Gabor-Nets that, for the first time in the literature, design and learn CNN kernels strictly in the form of Gabor filters, aiming to reduce the number of involved parameters and constrain the solution space and, hence, improve the performances of CNNs. Specifically, we develop an innovative phase-induced Gabor kernel, which is trickily designed to perform the Gabor feature learning via a linear combination of local low-frequency and high-frequency components of data controlled by the kernel phase. With the phase-induced Gabor kernel, the proposed Gabor-Nets gains the ability to automatically adapt to the local harmonic characteristics of the HSI data and, thus, yields more representative harmonic features. Also, this kernel can fulfill the traditional complex-valued Gabor filtering in a real-valued manner, hence making Gabor-Nets easily perform in a usual CNN thread. We evaluated our newly developed Gabor-Nets on three well-known HSIs, suggesting that our proposed Gabor-Nets can significantly improve the performance of CNNs, particularly with a small training set.
Chenying Liu 0001, Jun Li 0009, Lin He 0001, Antonio Plaza, Shutao Li 0001, Bo Li 0006
IEEE Trans. Neural Networks Learn. Syst.6
2021 Scene Graph Generation With Hierarchical Context
abstract
Scene graph generation has received increasing attention in recent years. Enhancing the predicate representations is an important entry point to this task. There are various methods to fully investigate the context of representation enhancement. In this brief, we analyze the decisive factors that can significantly affect the relation detection results. Our analysis shows that spatial correlations between objects, focused regions of objects, and global hints related to the relations have strong influences in relation prediction and contradiction elimination. Based on our analysis, we propose a hierarchical context network (HCNet) to generate a scene graph. HCNet consists of three contexts, including interaction context, depression context, and global context, which integrates information from pair, object, and graph levels. The experiments show that our method outperforms the state-of-the-art methods on the Visual Genome (VG) data set.
Guanghui Ren, Lejian Ren, Yue Liao, Si Liu 0001, Bo Li 0006, Jizhong Han, Shuicheng Yan
IEEE Trans. Neural Networks Learn. Syst.5
2020 Rule-Guided Compositional Representation Learning on Knowledge Graphs
abstract
Representation learning on a knowledge graph (KG) is to embed entities and relations of a KG into low-dimensional continuous vector spaces. Early KG embedding methods only pay attention to structured information encoded in triples, which would cause limited performance due to the structure sparseness of KGs. Some recent attempts consider paths information to expand the structure of KGs but lack explainability in the process of obtaining the path representations. In this paper, we propose a novel Rule and Path-based Joint Embedding (RPJE) scheme, which takes full advantage of the explainability and accuracy of logic rules, the generalization of KG embedding as well as the supplementary semantic structure of paths. Specifically, logic rules of different lengths (the number of relations in rule body) in the form of Horn clauses are first mined from the KG and elaborately encoded for representation learning. Then, the rules of length 2 are applied to compose paths accurately while the rules of length 1 are explicitly employed to create semantic associations among relations and constrain relation embeddings. Moreover, the confidence level of each rule is also considered in optimization to guarantee the availability of applying the rule to representation learning. Extensive experimental results illustrate that RPJE outperforms other state-of-the-art baselines on KG completion task, which also demonstrate the superiority of utilizing logic rules as well as paths for improving the accuracy and explainability of representation learning.
Guanglin Niu, Yongfei Zhang, Bo Li 0006, Peng Cui 0001, Si Liu 0001, Xiaowei Zhang 0003
AAAI3
2020 Single Camera Training for Person Re-Identification
abstract
Person re-identification (ReID) aims at finding the same person in different cameras. Training such systems usually requires a large amount of cross-camera pedestrians to be annotated from surveillance videos, which is labor-consuming especially when the number of cameras is large. Differently, this paper investigates ReID in an unexplored single-camera-training (SCT) setting, where each person in the training set appears in only one camera. To the best of our knowledge, this setting was never studied before. SCT enjoys the advantage of low-cost data collection and annotation, and thus eases ReID systems to be trained in a brand new environment. However, it raises major challenges due to the lack of cross-camera person occurrences, which conventional approaches heavily rely on to extract discriminative features. The key to dealing with the challenges in the SCT setting lies in designing an effective mechanism to complement cross-camera annotation. We start with a regular deep network for feature extraction, upon which we propose a novel loss function named multi-camera negative loss (MCNL). This is a metric learning loss motivated by probability, suggesting that in a multi-camera system, one image is more likely to be closer to the most similar negative sample in other cameras than to the most similar negative sample in the same camera. In experiments, MCNL significantly boosts ReID accuracy in the SCT setting, which paves the way of fast deployment of ReID systems with good performance on new target scenes.
Lingxi Xie, Longhui Wei, Yongfei Zhang, Bo Li 0006, Qi Tian 0001
AAAI5
2020 Referring Image Segmentation via Cross-Modal Progressive Comprehension
abstract
Referring image segmentation aims at segmenting the foreground masks of the entities that can well match the description given in the natural language expression. Previous approaches tackle this problem using implicit feature interaction and fusion between visual and linguistic modalities, but usually fail to explore informative words of the expression to well align features from the two modalities for accurately identifying the referred entity. In this paper, we propose a Cross-Modal Progressive Comprehension (CMPC) module and a Text-Guided Feature Exchange (TGFE) module to effectively address the challenging task. Concretely, the CMPC module first employs entity and attribute words to perceive all the related entities that might be considered by the expression. Then, the relational words are adopted to highlight the correct entity as well as suppress other irrelevant ones by multimodal graph reasoning. In addition to the CMPC module, we further leverage a simple yet effective TGFE module to integrate the reasoned multimodal features from different levels with the guidance of textual information. In this way, features from multi-levels could communicate with each other and be refined based on the textual context. We conduct extensive experiments on four popular referring segmentation benchmarks and achieve new state-of-the-art performances. Code is available at https://github.com/spyflying/CMPC-Refseg.
Shaofei Huang 0001, Tianrui Hui, Si Liu 0001, Guanbin Li, Yunchao Wei, Jizhong Han, Luoqi Liu, Bo Li 0006
CVPR8
2020 A Real-Time Cross-Modality Correlation Filtering Method for Referring Expression Comprehension
abstract
Referring expression comprehension aims to localize the object instance described by a natural language expression. Current referring expression methods have achieved good performance. However, none of them is able to achieve real-time inference without accuracy drop. The reason for the relatively slow inference speed is that these methods artificially split the referring expression comprehension into two sequential stages including proposal generation and proposal ranking. It does not exactly conform to the habit of human cognition. To this end, we propose a novel Realtime Cross-modality Correlation Filtering method (RCCF). RCCF reformulates the referring expression comprehension as a correlation filtering process. The expression is first mapped from the language domain to the visual domain and then treated as a template (kernel) to perform correlation filtering on the image feature map. The peak value in the correlation heatmap indicates the center points of the target box. In addition, RCCF also regresses a 2-D object size and 2-D offset. The center point coordinates, object size and center point offset together to form the target bounding box. Our method runs at 40 FPS while achieving leading performance in RefClef, RefCOCO, RefCOCO+ and RefCOCOg benchmarks. In the challenging RefClef dataset, our methods almost double the state-of-the-art performance (34.70% increased to 63.79%). We hope this work can arouse more attention and studies to the new cross-modality correlation filtering framework as well as the one-stage framework for referring expression comprehension.
Yue Liao, Si Liu 0001, Guanbin Li, Fei Wang 0032, Chen Qian 0006, Bo Li 0006
CVPR7
2020 Beautify As You Like
abstract
Customizable makeup transfer, which aims to transfer the makeup from an arbitrary reference face to a source face, is widely demanded in many applications such as short video platforms and online meeting applications. However, existing methods are neither user-friendly nor sufficiently fast. In this demo, we present the first fast makeup transfer system named as Fast Pose and expression robust Spatial-Aware GAN (FPSGAN). With a novel Attentive Makeup Morphing (AMM) module, FPSGAN is robust to face pose and expression. Moreover, it can achieve shade-controllable and partial makeup, improving the system's user-friendliness. In addition, FPSGAN is light-weighted and fast. To sum up, FPSGAN is the first fast customizable makeup transfer system to enable users to beautify themselves as they like.
Si Liu 0001, Chen Gao 0005, Ran He 0001, Bo Li 0006, Shuicheng Yan
ACM Multimedia5
2020 Large margin deep embedding for aesthetic image classification
Guanjun Guo, Hanzi Wang, Yan Yan 0001, Liming Zhang 0002, Bo Li 0006
Sci. China Inf. Sci.5
2020 A fast face detection method via convolutional neural network
Guanjun Guo, Hanzi Wang, Yan Yan 0001, Bo Li 0006
Neurocomputing5
2020 Accelerated Guided Sampling for Multistructure Model Fitting
abstract
The performance of many robust model fitting techniques is largely dependent on the quality of the generated hypotheses. In this paper, we propose a novel guided sampling method, called accelerated guided sampling (AGS), to efficiently generate the accurate hypotheses for multistructure model fitting. Based on the observations that residual sorting can effectively reveal the data relationship (i.e., determine whether two data points belong to the same structure), and keypoint matching scores can be used to distinguish inliers from gross outliers, AGS effectively combines the benefits of residual sorting and keypoint matching scores to efficiently generate accurate hypotheses via information theoretic principles. Moreover, we reduce the computational cost of residual sorting in AGS by designing a new residual sorting strategy, which only sorts the top-ranked residuals of input data, rather than all input data. Experimental results demonstrate the effectiveness of the proposed method in computer vision tasks, such as homography matrix and fundamental matrix estimation.
Taotao Lai, Hanzi Wang, Yan Yan 0001, Tat-Jun Chin, Bo Li 0006
IEEE Trans. Cybern.6
2020 ORDNet: Capturing Omni-Range Dependencies for Scene Parsing
abstract
Learning to capture dependencies between spatial positions is essential to many visual tasks, especially the dense labeling problems like scene parsing. Existing methods can effectively capture long-range dependencies with self-attention mechanism while short ones by local convolution. However, there is still much gap between long-range and short-range dependencies, which largely reduces the models' flexibility in application to diverse spatial scales and relationships in complicated natural scene images. To fill such a gap, we develop a Middle-Range (MR) branch to capture middle-range dependencies by restricting self-attention into local patches. Also, we observe that the spatial regions which have large correlations with others can be emphasized to exploit long-range dependencies more accurately, and thus propose a Reweighed Long-Range (RLR) branch. Based on the proposed MR and RLR branches, we build an Omni-Range Dependencies Network (ORDNet) which can effectively capture short-, middle- and long-range dependencies. Our ORDNet is able to extract more comprehensive context information and well adapt to complex spatial variance in scene images. Extensive experiments show that our proposed ORDNet outperforms previous state-of-the-art methods on three scene parsing benchmarks including PASCAL Context, COCO Stuff and ADE20K, demonstrating the superiority of capturing omni-range dependencies in deep models for scene parsing task.
Shaofei Huang 0001, Si Liu 0001, Tianrui Hui, Jizhong Han, Bo Li 0006, Jiashi Feng, Shuicheng Yan
IEEE Trans. Image Process.5
2020 Adaptive Single Image Dehazing Using Joint Local-Global Illumination Adjustment
abstract
Haze has a serious impact on the outdoor optical imaging systems, and it will result in image blurring, color shift, and saturation reduction. Recently, many single image dehazing algorithms have been proposed for practical applications, such as surveillance. However, since the widely-used global atmospheric light in image dehazing fails to well describe the local illumination differences of images, these algorithms fail to well adapt to scenes with different haze concentrations and lighting conditions. Therefore, this paper proposes an adaptive single image dehazing algorithm using joint local-global illumination adjustment. A local illumination estimation for hazy image is proposed to replace the global atmospheric light constant in the atmospheric scattering model, and it can better adapt to the local differences of image illumination. Correspondingly, the global atmospheric light constant is proposed to be utilized to adaptively compensate the illumination intensity, which may better overcome the dark illumination problem within the dehazed image. The experimental results demonstrate that the proposed algorithm can outperform the state-of-the-art algorithms in terms of not only the dehazing effect but also the adaptability.
Hai-Miao Hu, Hongda Zhang, Bo Li 0006
IEEE Trans. Multim.4
2019 GPS: Group People Segmentation with Detailed Part Inference
abstract
Noticeable progress has been witnessed in general object detection, semantic segmentation and instance segmentation, while parsing a group of people is still a challenging task for human-centric visual understanding due to severe occlusion and various poses. In this paper, we present a new large-scale dataset named “GPS (Group People Segmentation)” to boost academical study and technology development. GPS contains 14000 elaborately annotated images with 20 fine-grained semantic category labels related to human, divided into two sub-datasets corresponding to indoor and outdoor scenes involving various poses, occlusion and background. We further propose a novel GPSNet for group people segmentation. GPSNet consists of a new “Adjusted RoI Align” module to adjust position of detected person and align RoI features, such that the network does not need to fit various positions of each person. A fusion of global and local features is also employed to refine parsing results. Compared with baseline methods, GPSNet achieves the best performance on GPS Dataset.
Yue Liao, Si Liu 0001, Tianrui Hui, Chen Gao 0005, Yao Sun 0004, Bo Li 0006
ICME7
2019 A New Spatio-Temporal Fusion Method for Remotely Sensed Data Based on Convolutional Neural Networks
abstract
In some remote sensing applications such as change detection, satellite images with both high spatial and high temporal resolution are required. However, no single satellite sensor can currently provide such images due to technical specifications. To solve this problem, spatio-temporal fusion provides a cost-effective solution. In this paper, we propose a new spatio-temporal fusion approach, based on convolutional neural networks (CNNs), for Landsat and MODIS image fusion. Specifically, the proposed approach utilizes CNNs to model the heterogeneity of fine pixels from the coarse MODIS images. Here, the heterogeneity of fine pixels is defined as the difference between the reflectance changes obtained from the two types of images. After that, two transition-predicted images can be obtained using the trained CNNs, which are then fused in order to obtain a fi-nal prediction. In our newly proposed approach, CNNs are only used to learn the heterogeneity of fine pixels rather than the whole images, thus providing a more stable and less time-consuming strategy as compared to other available approaches. We evaluated the proposed approach on a public spatio-temporal fusion dataset and the obtained results suggest that our newly developed method achieves state-of-the-art performance.
Yunfei Li 0006, Chenying Liu 0001, Lin Yan 0005, Jun Li 0009, Antonio Plaza, Bo Li 0006
IGARSS6
2019 Finding Images by Dialoguing with Image
abstract
Image retrieval in complicated scene is a challenging task that requires the comprehensive understanding of an image. In this paper, we propose a scene graph based image retrieval framework that combines the scene graph generation with image retrieval and fine tuning the searching results via a dialogue mechanism. Specifically, we proposed an image retrieval oriented scene graph generation model that takes an image and a text describing the image as inputs. The additional text input is used to control the generated scene graph. It provides information for a newly introduced attributes head to better predict the attributes and helps constructing an adjacency matrix at the same time. Graph Convolutional Network is further used to gather information among nodes for precise relation estimation. Moreover, modification on the scene graph can be done by changing the text. Our proposed approach achieves the state-of-the-art performances in both scene graph based image retrieval and scene graph generation in the Visual Genome dataset.
Lejian Ren, Si Liu 0001, Han Huang 0002, Jizhong Han, Shuicheng Yan, Bo Li 0006
ACM Multimedia6
2019 Real-time object tracking via self-adaptive appearance modeling
Ming Xin 0003, Bo Li 0006, Guanglin Niu
Neurocomputing3
2019 Structured fragment-based object tracking using discrimination, uniqueness, and validity selection
abstract
Local features have widely been used in visual tracking to improve robustness in the presence of partial occlusion, deformation, and rotation. In this paper, a local fragment-based object tracking algorithm is proposed. Unlike many existing fragment-based algorithms using all the fragments and allocating the weight to each fragment according to similarity, the proposed algorithm only selects discriminative, unique, and valid fragments for tracking. First, discrimination and uniqueness metric are defined for each local fragment, and an automatic pre-selection mechanism is proposed for all these fragments. Second, a Harris-SIFT filter is used to select the current valid fragments and exclude the occluded or highly deformed fragments. By selecting the discriminative, unique, and valid fragments, these fragments are used to construct a structured description for the object. Finally, the object tracking is performed using the selected fragments combining the displacement and similarity, as well as spatial constraint of the selected fragments. The object template can be updated by fusing feature similarity and structural consistency. The experimental results on a recent OTB 2013 tracking benchmark data set demonstrate that the proposed algorithm can achieve reliable tracking results even in the presence of significant appearance changes, partial occlusion, and similar disturbances.
Bo Li 0006, Ming Xin 0003, Gang Luo 0003
Multim. Syst.2
2019 An Adaptive Multi-Projection Metric Learning for Person Re-Identification Across Non-Overlapping Cameras
abstract
Person re-identification is one of the most important and challenging problems in video analytics systems; it aims to match people across non-overlapping camera views. For person re-identification, metric learning is introduced to improve the performance by providing a metric adapted for cross-view matching. The essence of metric learning is to search for an optimal projection matrix to project the original features into a new feature space. However, most existing metric learning methods overlook the inconsistency of feature distributions in multiple cameras. In this paper, we propose a multi-projection metric learning (MPML) method to overcome the inconsistency among multiple cameras in person re-identification. Our solution is to jointly learn multiple projection matrices using paired samples from different cameras to project features from different cameras into a common feature space. To make our method adaptive to newly added cameras without affecting the learned projection matrices, we further propose an adaptive MPML method, which can learn new camera projection matrices without having to update any of the obtained projection matrices. The proposed methods are evaluated on four major person re-identification data sets, with comprehensive experiments showing the effectiveness of the proposed methods and notable improvements over the state-of-the-art approaches.
Hai-Miao Hu, Bo Li 0006, Qi Tian 0001
IEEE Trans. Circuits Syst. Video Technol.3
2019 Haze and Thin Cloud Removal Using Elliptical Boundary Prior for Remote Sensing Image
abstract
Remote sensing images play important roles in various earth surface observation applications. However, the hazy state of surface atmosphere can visually decrease the contrast and availability of remote sensing images. In this paper, we propose a haze and thin cloud removal method for single visible remote sensing images, which aims to robustly estimate haze thickness, atmospheric light, and transmission value from a remote sensing image with dense haze or thin cloud, and finally recovers a haze-free image. An elliptical boundary prior (EBP) is proposed to transform the haze thickness in each local patch from the pixels cluster in the spectral space, which is surrounded by an ellipse. With the aim of preventing highlight objects influences, an atmospheric light estimation approach is presented. The correlation of transmission and haze thickness is reconstructed to develop the scattering model for remote sensing images. The experimental results demonstrate that the proposed method can not only significantly improve the contrast and restore textures of various kinds of hazy remote sensing images but also well preserve the spectral information of visible bands.
Qiang Guo 0013, Hai-Miao Hu, Bo Li 0006
IEEE Trans. Geosci. Remote. Sens.3
2019 Single Image Defogging Based on Illumination Decomposition for Visual Maritime Surveillance
abstract
Single image fog removal is important for surveillance applications and many defogging methods have been proposed, recently. Due to the adverse atmospheric conditions, the scattering properties of foggy images depend on not only the depth information of scene, but also the atmospheric aerosol model, which has more prominent influence on illumination in a fog scene than that in a haze scene. However, recent defogging methods confuse haze and fog, and they fail to consider fully about the scattering properties. Thus, these methods are not sufficient to remove fog effects, especially for images in maritime surveillance. Therefore, this paper proposes a single image defogging method for visual maritime surveillance. Firstly, a comprehensive scattering model is proposed to formulate a fog image in the glow-shaped environmental illumination. Then, an illumination decomposition algorithm is proposed to eliminate the glow effect on the airlight radiance and recover a fog layer, in which objects at the infinite distance have uniform luminance. Secondly, a transmission-map estimation based on the non-local haze-lines prior is utilized to constrain the transmission map into a reasonable range for the input fog image. Finally, the proposed illumination compensation algorithm enables the defogging image to preserve the natural illumination information of the input image. In addition, a fog image dataset is established for visual maritime surveillance. The experimental results based on the established dataset demonstrate that the proposed method can outperform the state-of-the-art methods in terms of both the subjective and objective evaluation criteria. Moreover, the proposed method can effectively remove fog and maintain naturalness for fog images.
Hai-Miao Hu, Qiang Guo 0013, Hanzi Wang, Bo Li 0006
IEEE Trans. Image Process.5
2019 Detail Preserved Single Image Dehazing Algorithm Based on Airlight Refinement
abstract
Single-image haze removal is important for many practical applications (e.g., surveillance). However, dehazed results of existing algorithms tend to be oversmoothed with missing fine image details. This drawback is caused by two factors: inaccurate airlight estimations and disregarding multiple scattering. In this paper, we propose a detail-preserving image dehazing algorithm based on two key priors, namely, the depth-edge aware prior and the airlight impact regularity prior. The proposed algorithm makes contributions in both the haze removal step and the postprocessing step. First, based on the depth-edge aware prior, an airlight refinement algorithm is proposed. The gradient strength of the minimum channel is employed to calculate punishment weights to smooth the dark channel. Second, based on the airlight impact regularity prior, an adaptive sharpening model that considers the refined airlight to determine the sharpening strength value is established to enhance levels of detail. Experimental results demonstrate that the proposed algorithm cannot only effectively remove haze but can also enhance levels of detail to thus outperform the state of the art on a wide variety of images.
Hai-Miao Hu, Bo Li 0006, Qiang Guo 0013, Shiliang Pu
IEEE Trans. Multim.3
2019 A Novel Projective-Consistent Plane Based Image Stitching Method
abstract
When different target surfaces, in three-dimensional space, are mapped onto an image plane, they have different projections. These projections vary with the viewpoint. These local differences have influence on the accuracy of image stitching. Most of the existing image stitching methods divide an input image into a number of fixed-size cells, and the pixels within the same cell are then warped using the same local transformation model for the alignment. These methods are based on the hypothesis that the transformation models in one cell are consistent. However, this hypothesis does not hold in general. In this paper, we propose a novel projective-consistent plane based image stitching method (termed PCPS). It divides the overlapping regions of an input image into some projective-consistent planes according to the normal vectors' orientations of local regions and the reprojection errors of aligned images. The local projective transformation model is estimated for each projective-consistent plane. And then, a hybrid warping model is estimated. For the pixels in overlapping regions, the local projective transformation models are adopted to achieve a better alignment. While for the pixels in non-overlapping regions, a global projective transformation model is estimated by using the inliers uniformly distributed in the projective-consistent planes to avoid distortion. Compared with the state-of-the-art image stitching methods, the experimental results on a number of challenging image sequences show that the projective transformation model estimated by the proposed PCPS method for each projective-consistent plane is more accurate, and the achieved stitching results have less seams and projective distortion.
Yue Wang 0029, Hanzi Wang, Bo Li 0006, Hai-Miao Hu
IEEE Trans. Multim.4
2018 Perceptual hash-based feature description for person re-identification
Hai-Miao Hu, Zihao Hu, Shengcai Liao, Bo Li 0006
Neurocomputing5
2018 Conceptual space based model fitting for multi-structure data
Guobao Xiao, Hailing Luo, Bo Li 0006, Yan Yan 0001, Hanzi Wang
Neurocomputing5
2018 An effective motion object detection method using optical flow estimation under a moving camera
Yugui Zhang, Bo Li 0006
J. Vis. Commun. Image Represent.4
2018 Ship Detection From Thermal Remote Sensing Imagery Through Region-Based Deep Forest
abstract
Ship detection from thermal remote sensing imagery is a challenging task because of cluttered scenes and variable appearances of ships. In this letter, we propose a novel detection algorithm named region-based deep forest (RDF) toward overcoming these existing issues. The RDF consists of a simple region proposal network and a deep forest ensemble. The region proposal network trained over gradient features robustly generates a small number of candidates that precisely cover ship targets in various backgrounds. The deep forest ensemble adaptively learns features from remote sensing data and discriminates real ships from region proposals efficiently. The training process of deep forest ensemble is efficient and users can control training cost according to computational resource available. Experimental results on numerous thermal satellite images demonstrate the superior performance of our method compared with state-of-the-art methods.
Feng Yang 0002, Qizhi Xu, Bo Li 0006
IEEE Geosci. Remote. Sens. Lett.3
2018 A fast and HEVC-compatible perceptual video coding scheme using a transform-domain Multi-Channel JND model
Gang Wang 0023, Yongfei Zhang, Bo Li 0006, Rui Fan 0002, Mingliang Zhou 0001
Multim. Tools Appl.3
2018 Too Far to See? Not Really! - Pedestrian Detection With Scale-Aware Localization Policy
abstract
A major bottleneck of pedestrian detection lies on the sharp performance deterioration in the presence of small-size pedestrians that are relatively far from the camera. Motivated by the observation that pedestrians of disparate spatial scales exhibit distinct visual appearances, we propose in this paper an active pedestrian detector that explicitly operates over multiple-layer neuronal representations of the input still image. More specifically, convolutional neural nets, such as ResNet and faster R-CNNs, are exploited to provide a rich and discriminative hierarchy of feature representations, as well as initial pedestrian proposals. Here each pedestrian observation of distinct size could be best characterized in terms of the ResNet feature representation at a certain layer of the hierarchy. Meanwhile, initial pedestrian proposals are attained by the faster R-CNNs techniques, i.e., region proposal network and follow-up region of interesting pooling layer employed right after the specific ResNet convolutional layer of interest, to produce joint predictions on the bounding-box proposals' locations and categories (i.e., pedestrian or not). This is engaged as an input to our active detector, where for each initial pedestrian proposal, a sequence of coordinate transformation actions is carried out to determine its proper x-y 2D location and the layer of feature representation, or eventually terminated as being background. Empirically our approach is demonstrated to produce overall lower detection errors on widely used benchmarks, and it works particularly well with far-scale pedestrians. For example, compared with 60.51% log-average miss rate of the state-of-the-art MS-CNN for far-scale pedestrians (those below 80 pixels in bounding-box height) of the Caltech benchmark, the miss rate of our approach is 41.85%, with a notable reduction of 18.66%.
Xiaowei Zhang 0003, Li Cheng 0001, Bo Li 0006, Hai-Miao Hu
IEEE Trans. Image Process.3
2018 Naturalness Preserved Nonuniform Illumination Estimation for Image Enhancement Based on Retinex
abstract
Illumination estimation is important for image enhancement based on Retinex. However since illumination estimation is an ill-posed problem it is difficult to achieve accurate illumination estimation for nonuniform illumination images. The conventional illumination estimation algorithms fail to comprehensively take all the constraints into the consideration such as spatial smoothness sharp edges on illumination boundaries and limited range of illumination. Thus these algorithms cannot effectively and efficiently estimate illumination while preserving naturalness. In this paper we present a naturalness preserved illumination estimation algorithm based on the proposed joint edge-preserving filter which exploits all the abovementioned constraints. Moreover a fast estimation is implemented based on the box filter. Experimental results demonstrate that the proposed algorithm can achieve the adaptive smoothness of illumination beyond edges and ensure the range of the estimated illumination. When compared with other state-of-the-art algorithms it can achieve better quality from both subjective and objective aspects.
Hai-Miao Hu, Bo Li 0006, Qiang Guo 0013
IEEE Trans. Multim.3
2018 Background Modeling and Referencing for Moving Cameras-Captured Surveillance Video Coding in HEVC
abstract
Surveillance video coding is crucial for improving compression efficiency in intelligent video surveillance systems and applications. Plenty of work has been done, which can be roughly divided into two categories: the former mainly focuses on low-complexity background modeling to obtain the clear background, while the latter focuses on an appropriate coding strategy to generate the high-quality background reference picture for effective background prediction. However, almost all existing works focus only on stationary camera scenes, while moving cameras-captured surveillance video coding is left untouched and is still an open problem. In this paper, a background modeling and referencing scheme for moving cameras-captured surveillance video coding in high-efficiency video coding (HEVC) is proposed. First, this paper proposes a low-complexity motion background modeling algorithm for surveillance video coding using the running average based on a global-motion-compensation method. To obtain the global motion vector, we propose a global motion detection method based on character blocks by establishing a low-rank singular value decomposition model for clustering and estimating motion vectors of background character blocks in the cameras movement circumstance. Second, we propose a background referencing coding strategy, in which the motion background coding tree units (MBCTUs) would be selected by anchoring the input video frame on the modeling background frame and coded with the optimized quantization parameter. Then, the reconstructed MBCTU will be used to update the previous coding tree unit in the global compensation location of the background reference picture. Extensive experimental results show that the proposed scheme can achieve significant bit savings of up to 26.6% and, on average, 6.7% with similar subjective quality and negligible encoding complexity, compared to HM12.0. Besides, the proposed scheme consistently outperforms two state-of-the-art surveillance video coding schemes with remarkable bitrate savings.
Gang Wang 0023, Bo Li 0006, Yongfei Zhang, Jinhui Yang
IEEE Trans. Multim.2
2017 Scale-aware hierarchical loss: A multipath RPN for multi-scale pedestrian detection
abstract
Pedestrians with different spatial scales exhibiting dramatically differences, the serious performance decline with decreasing resolution is the major bottleneck for current pedestrian detection. Considering the local feature differences for multi-scale pedestrians, a scale-aware multipath region proposal network is exploited to improve the recall rate, which is divided into several branches to generate a proper object proposal for target with specific scale range. Moreover, motivated by the visual semantic concepts of different convolutional layers, a scale-aware hierarchical loss model is introduced to minimize the error rate for pedestrians with different scales, in which the hierarchical features of higher convolutional layers are jointed to calculate a multi-task loss to learn scale-aware weighting of multipath region proposal network for each object proposal. Finally, compared to state-of-the-art methods, experimental results on the challenging ETH and Caltech benchmark show the superiority of the proposed method for large variance in instance scales.
Xiaowei Zhang 0003, Bo Li 0006, Hai-Miao Hu
VCIP2
2017 A region-based video de-noising algorithm based on temporal and spatial correlations
Hai-Miao Hu, Qiang Guo 0013, Bo Li 0006
Neurocomputing4
2017 DeMS: A hybrid scheme of task scheduling and load balancing in computing clusters
Yu Liu 0031, Changjie Zhang, Bo Li 0006, Jianwei Niu 0002
J. Netw. Comput. Appl.3
2017 Complexity-based intra frame rate control by jointing inter-frame correlation for high efficiency video coding
Mingliang Zhou 0001, Yongfei Zhang, Bo Li 0006, Hai-Miao Hu
J. Vis. Commun. Image Represent.3
2017 Ship Detection From Optical Satellite Images Based on Saliency Segmentation and Structure-LBP Feature
abstract
Automatic ship detection from optical satellite imagery is a challenging task due to cluttered scenes and variability in ship sizes. This letter proposes a detection algorithm based on saliency segmentation and the local binary pattern (LBP) descriptor combined with ship structure. First, we present a novel saliency segmentation framework with flexible integration of multiple visual cues to extract candidate regions from different sea surfaces. Then, simple shape analysis is adopted to eliminate obviously false targets. Finally, a structure-LBP feature that characterizes the inherent topology structure of ships is applied to discriminate true ship targets. Experimental results on numerous panchromatic satellite images validate that our proposed scheme outperforms other state-of-the-art methods in terms of both detection time and detection accuracy.
Feng Yang 0002, Qizhi Xu, Bo Li 0006
IEEE Geosci. Remote. Sens. Lett.3
2017 A person re-identification algorithm based on pyramid color topology feature
Hai-Miao Hu, Guodong Zeng, Zihao Hu, Bo Li 0006
Multim. Tools Appl.5
2017 Efficient PCIe transmission for Multi-Channel video using dynamic splicing and conditional prefetching
Tingshan Liu, Huiyong Li 0005, Bo Li 0006, Miyi Duan
Multim. Tools Appl.4
2017 Moving object detection algorithm based on pixel spatial sample difference consensus
Yugui Zhang, Mengxiong Han, Bo Li 0006
Multim. Tools Appl.5
2017 Multidirectional parabolic prediction-based interpolation-free sub-pixel motion estimation
Rui Fan 0002, Yongfei Zhang, Bo Li 0006, Gang Wang 0023
Signal Process. Image Commun.3
2017 Local Co-Occurrence Selection via Partial Least Squares for Pedestrian Detection
abstract
Channel feature detectors are the most popular approaches for pedestrian detection recently. However, most of these approaches train the boosted decision trees by selecting a single feature at each node, which does not effectively exploit the multi-feature cues and spatial information. To address this issue, this paper proposes to construct the co-occurrence of multiple channel features in local image neighborhoods for pedestrian detection. In our approach, a binary pattern of feature co-occurrence is represented by combining the binary variables quantized from each channel feature, and the spatial information is incorporated by selecting the neighbors to jointly represent the feature co-occurrence in a local image block. However, feature co-occurrence selection leads to many possible feature combinations, which significantly increase the computational cost at the training stage. Therefore, in order to reduce the number of candidate features and obtain the most discriminative features effectively, a partial least squares-based feature selection approach called variable importance on projection is exploited. Comprehensive experiments are conducted on several challenging pedestrian data sets, and superior performances are achieved by the proposed approach in comparison with some state-of-the-art pedestrian detection approaches.
Hanzi Wang, Yan Yan 0001, Bo Li 0006, Chang Wen Chen
IEEE Trans. Intell. Transp. Syst.4
2017 Motion Classification-Based Fast Motion Estimation for High-Efficiency Video Coding
abstract
High efficiency video coding (HEVC), the latest video coding standard, is becoming popular due to its excellent coding performance. However, the significant gain in performance is achieved at the cost of substantially higher encoding complexity than its precedent H.264/AVC, in which motion estimation (ME) is the most time-consuming module that effectively removes temporal redundancy. Test zone search (TZS) is adopted as the default fast ME method in the reference software of HEVC; however, its computational complexity is still too high for real-time applications. Several fast ME algorithms have been recently proposed to further reduce ME complexity; however, these approaches typically lead to non-negligible performance loss. To address this problem, this paper proposes a motion classification-based fast ME algorithm. By exploring the motion relationship of neighboring blocks and the coding cost characteristic, the prediction unit (PU) is first categorized into one of three classes, namely, motion-smooth PU, motion-medium PU and motion-complex PU. Then different search strategies are carefully designed for PUs of each class according to their respective motion and content characteristics. Furthermore, a fast search priority-based partial internal termination scheme is presented to rapidly skip impossible positions that speeds up cost computation during the ME process. Extensive experimental results demonstrate that the proposed algorithm achieves as much as 12.47% and 20.25% reductions in total encoder complexity when compared with TZS under low delay P and random access configuration, respectively, with negligible rate-distortion degradation; thus, it outperforms state-of-the-art fast ME algorithms in terms of both coding performance and complexity reduction.
Rui Fan 0002, Yongfei Zhang, Bo Li 0006
IEEE Trans. Multim.3
2017 An Adaptive Fusion Algorithm for Visible and Infrared Videos Based on Entropy and the Cumulative Distribution of Gray Levels
abstract
Visible videos captured under different weather conditions may exhibit different characteristics, and thermal infrared videos are easily affected by ambient temperature variations; this sensitivity to environmental conditions makes the fusion of visible and thermal infrared videos a challenge. This paper proposes an adaptive fusion algorithm for visible and infrared videos, and uses cumulative distribution of gray levels and the entropy to adaptively retain infrared-hot targets and visible textures. The original visible and infrared frames are decomposed into two layers, namely, the base layer and the detail layer. The guided filter is employed to decompose frames due to its high efficiency. Two weight maps, one for the infrared base layer and one for the visible base layer, are adaptively generated based on the cumulative distribution of gray levels and the entropy, respectively. The visible base layer and the infrared base layer are fused based on their weight maps. The final fusion result is obtained by combining the fused base layer with the visible detail layer. Experimental results demonstrate that the proposed algorithm can achieve better fusion results compared with state-of-the-art methods.
Hai-Miao Hu, Bo Li 0006, Qiang Guo 0013
IEEE Trans. Multim.3
2017 Complexity Correlation-Based CTU-Level Rate Control with Direction Selection for HEVC
abstract
Rate control is a crucial consideration in high-efficiency video coding (HEVC). The estimation of model parameters is very important for coding tree unit (CTU)-level rate control, as it will significantly affect bit allocation and thus coding performance. However, the model parameters in the CTU-level rate control sometimes fails because of inadequate consideration of the correlation between model parameters and complexity characteristic. In this study, we establish a novel complexity correlation-based CTU-level rate control for HEVC. First, we formulate the model parameter estimation scheme as a multivariable estimation problem; second, based on the complexity correlation of the neighbouring CTU, an optimal direction is selected in five directions for reference CTU set selection during model parameter estimation to further improve the prediction accuracy of the complexity of the current CTU. Third, to improve their precision, the relationship between the model parameters and the complexity of the reference CTU set in the optimal direction is established by using least square method (LS), and the model parameters are solved via the estimated complexity of the current CTU. Experimental results show that the proposed algorithm can significantly improve the accuracy of the CTU-level rate control and thus the coding performance; the proposed scheme consistently outperforms HM 16.0 and other state-of-the-art algorithms in a variety of testing configurations. More specifically, up to 8.4% and on average 6.4% BD-Rate reduction is achieved compared to HM 16.0 and up to 4.7% and an average of 3.4% BD-Rate reduction is achieved compared to other algorithms, with only a slight complexity overhead.
Mingliang Zhou 0001, Yongfei Zhang, Bo Li 0006, Xupeng Lin
ACM Trans. Multim. Comput. Commun. Appl.3
2016 Conceptual space based gross outlier removal for geometric model fitting
abstract
In this paper, we propose an efficient and robust gross outlier removal method, called the Conceptual Space based Gross Outlier Removal (CSGOR) method, to remove gross outliers for geometric model fitting. In the proposed method, each data point is mapped to a conceptual space by computing the preference of "good" model hypotheses. In the conceptual space, the distributions of inliers and gross outliers are significantly different. Specifically, inliers of each model instance are distributed in a subspace and they are far away from the origin of the conceptual space, while gross outliers are distributed near the origin. In this manner, the problem of densely gross outlier removal is formulated as a binary classification problem. The main advantage of the proposed method is that it can handle data with a large proportion of outliers and effectively remove gross outliers in data. Experimental results on both synthetic and real data have demonstrated the efficiency and effectiveness of the proposed method.
Guobao Xiao, Bo Li 0006, Yan Yan 0001, Hanzi Wang
ICARCV4
2016 Robust Object Tracking Using Valid Fragments Selection
Bo Li 0006, Gang Luo 0003
MMM (1)2
2016 Parameter-adaptive nighttime image enhancement with multi-scale decomposition
abstract
As a challenging problem, image enhancement plays an important role in computer vision applications and has been widely studied. As one of the most difficult issues of image enhancement, outdoor nighttime image enhancement suffers from noise amplification easily. To solve this problem, this study proposes a parameter‐adaptive nighttime image enhancement method with multi‐scale decomposition. The main contributions of this work are threefold. First, the authors find out that noises in different scales are various, and their method decomposes an input image into three high‐frequency layers and a background layer accordingly. Second, the authors’ method enhances each high‐frequency layer using adaptive parameters based on the characteristics of noises. Third, the proposed method maps the background layer to make it suitable to present details. Experiment results demonstrate that the proposed method can suppress noises as well as improve details effectively.
Shuhang Wang, Bo Li 0006
IET Comput. Vis.3
2016 A hierarchal BoW for image retrieval by enhancing feature salience
Hai-Miao Hu, Bo Li 0006
Neurocomputing4
2016 An improved RANSAC based on the scale variation homogeneity
Yue Wang 0029, Qizhi Xu, Bo Li 0006, Hai-Miao Hu
J. Vis. Commun. Image Represent.4
2016 Memory-efficient high-speed VLSI implementation of multi-level discrete wavelet transform
Yongfei Zhang, Haiheng Cao, Bo Li 0006
J. Vis. Commun. Image Represent.4
2016 Content-adaptive parameters estimation for multi-dimensional rate control
Mingliang Zhou 0001, Bo Li 0006, Yongfei Zhang
J. Vis. Commun. Image Represent.2
2016 Automatic Change Detection in Synthetic Aperture Radar Images Based on PCANet
abstract
This letter presents a novel change detection method for multitemporal synthetic aperture radar images based on PCANet. This method exploits representative neighborhood features from each pixel using PCA filters as convolutional filters. Thus, the proposed method is more robust to the speckle noise and can generate change maps with less noise spots. Given two multitemporal images, Gabor wavelets and fuzzy c-means are utilized to select interested pixels that have high probability of being changed or unchanged. Then, new image patches centered at interested pixels are generated and a PCANet model is trained using these patches. Finally, pixels in the multitemporal images are classified by the trained PCANet model. The PCANet classification result and the preclassification result are combined to form the final change map. The experimental results obtained on three real SAR image data sets confirm the effectiveness of the proposed method.
Feng Gao 0005, Junyu Dong, Bo Li 0006, Qizhi Xu
IEEE Geosci. Remote. Sens. Lett.3
2016 Relative Radiometric Normalization for Multitemporal Remote Sensing Images by Hierarchical Regression
abstract
The existing relative radiometric normalization methods are insufficient to define the invariant pixels automatically, and the conventional methods do not perform well when the multitemporal images contain a lot of changes. Two types of changes should be particularly considered: one is caused by significant spectral differences due to change of ground objects, and the other is the pixels in the regions of misalignment caused by displacement due to differences in acquisition view angles and geometrical distortions. To automatically extract invariant pixels and reduce the influence of the changes, a hierarchical regression method is proposed to reduce the radiation difference for multitemporal images, which consists of extraction of the pseudo-invariant features (PIFs) and optimization of normalization parameters. A weighted regression based on spectral difference is proposed to automatically extract the PIFs, which can also suppress the negative effect of the first type of changes. In addition, a robust regression with gradient dependence is performed on the extracted PIFs to build the final relationship between the target image and the reference image, which can be robust for the second type of changes. Experimental results demonstrate that the proposed method has a better performance to normalize the target image.
Qizhi Xu, Bo Li 0006
IEEE Geosci. Remote. Sens. Lett.3
2016 An efficient Markov chain-based data prefetching for motion estimation of HEVC on multi-core DSPs
Bo Li 0006
Multim. Tools Appl.3
2016 Digital image stabilization based on adaptive motion filtering with feedback correction
Bo Zhai, Bo Li 0006
Multim. Tools Appl.3
2015 Pedestrian detection based on hierarchical co-occurrence model for occlusion handling
Xiaowei Zhang 0003, Hai-Miao Hu, Bo Li 0006
Neurocomputing4
2015 Joint global-local information pedestrian detection algorithm for outdoor video surveillance
Hai-Miao Hu, Xiaowei Zhang 0003, Bo Li 0006
J. Vis. Commun. Image Represent.4
2015 Pansharpening Using Regression of Classified MS and Pan Images to Reduce Color Distortion
abstract
The synthesis of low-resolution panchromatic (Pan) image is a critical step of ratio enhancement (RE) and component substitution (CS) pansharpening methods. The two types of methods assume a linear relation between Pan and multispectral (MS) images. However, due to the nonlinear spectral response of satellite sensors, the qualified low-resolution Pan image cannot be well approximated by a weighted summation of MS bands. Therefore, in some local areas, significant gray value difference exists between a synthetic Pan image and a high-resolution Pan image. To tackle this problem, the pixels of Pan and MS images are divided into several classes by$k$-means algorithm, and then multiple regression is used to calculate summation weights on each group of pixels. Experimental results demonstrate that the proposed technique can provide significant improvements on reducing color distortion.
Qizhi Xu, Yun Zhang 0014, Bo Li 0006
IEEE Geosci. Remote. Sens. Lett.3
2014 Robust patch-based tracking using valid patch selection and feature fusion update
abstract
This paper proposes a robust patch-based object tracking algorithm. Unlike many traditional algorithms, which divide the object into multiple patches and allocate the weight values for each patches, this paper uses SIFT feature matching to select valid patches and filter out invalid patches. The invalid patches usually corresponding to the occluded or partially transformed part of the object. Thus, guided by valid patch, patch-based color histogram provides a richer description of the object. The similarity of valid patch is used in particle filter to locate the object. Moreover, since feature similarity is easy to bring into object drift, this paper updates the object template fusing feature similarity and valid patches, which is both scale adaptive and robust to partial occlusion. The experimental results show that the proposed algorithm is more accurate and robust than state-of-the-art tracking algorithms in challenging scenarios.
WenBei Mao, Bo Li 0006
ICIP3
2014 Single image haze removal using content-adaptive dark channel and post enhancement
abstract
As a challenging problem, image haze removal plays an important role in computer vision applications. The dark channel prior has been widely studied for haze removal since it is simple and effective; however, it still suffers from over‐saturation, artefacts and dark‐look. To resolve these problems, this study proposes a method of single image haze removal using content‐adaptive dark channel and post enhancement. The main contributions of this work are as follows: first, an associative filter, which can transfer the structures of a reference image and the grey levels of a coarse image to the filtering output, is employed to compute the dark channel efficiently and effectively. Secondly, the dark channel confidence is utilised to restrict the dark channel based on the content of the image. Finally, a post enhancement method is devised to map the luminance of the restored haze‐free image with the preservation of local contrast. Experimental results demonstrate that the proposed method significantly improves the visibility of the hazy image.
Bo Li 0006, Shuhang Wang, Liping Zheng
IET Comput. Vis.1
2014 Ship Detection From Optical Satellite Images Based on Sea Surface Analysis
abstract
Automatic ship detection in high-resolution optical satellite images with various sea surfaces is a challenging task. In this letter, we propose a novel detection method based on sea surface analysis to solve this problem. The proposed method first analyzes whether the sea surface is homogeneous or not by using two new features. Then, a novel linear function combining pixel and region characteristics is employed to select ship candidates. Finally, Compactness and Length-width ratio are adopted to remove false alarms. Specifically, based on the sea surface analysis, the proposed method cannot only efficiently block out no-candidate regions to reduce computational time, but also automatically assign weights for candidate selection function to optimize the detection performance. Experimental results on real panchromatic satellite images demonstrate the detection accuracy and computational efficiency of the proposed method.
Guang Yang 0024, Bo Li 0006, Shufan Ji, Feng Gao 0005, Qizhi Xu
IEEE Geosci. Remote. Sens. Lett.2
2014 Visual Distortion Sensitivity Modeling for Spatially Adaptive Quantization in Remote Sensing Image Compression
abstract
As remote sensing images are often characterized with strong randomness, weak local correlation, and multiple small targets, the commonly used coarse-granularity subband-level quantization scheme fails to make use of these characteristics; thus, the performance improvements of these methods in literature are often marginal. To address this problem, this letter presents a novel spatially adaptive quantization (SAQ) method for the compression of remote sensing images based on our proposed Visual Distortion Sensitivity (ViDiS) Model. The ViDiS model takes into consideration four ViDiS components, including image luminance, spatial frequency, spatial orientation, and visual masking, to help measure the distortion more consistent to the image quality perceived by human beings. Then, a SAQ scheme is proposed to better exploit the content characteristics of remote sensing images, in which the quantization is conducted on a finer subband block level rather than subband level, with the guidance of the ViDiS model. Experimental results show that the proposed algorithm can preserve better visual quality in low-contrast areas with small targets at a competitive computational cost, which makes it more desirable in compression applications for remote sensing images.
Yongfei Zhang, Haiheng Cao, Bo Li 0006
IEEE Geosci. Remote. Sens. Lett.4
2014 A fast image dehazing algorithm based on negative correction
Hai-Miao Hu, Shuhang Wang, Bo Li 0006
Signal Process.4
2014 Delay-Bounded Priority-Driven Resource Allocation for Video Transmission Over Multihop Networks
abstract
In this paper we consider the problem of resource allocation for video transmission over mesh networks with delay bound constraints and priority-based packet scheduling. We observe that priority-driven packet scheduling at the intermediate network routers has a direct and significant impact on the queuing behaviors and delay bound violation probabilities of video packets, as well as the overall end-to-end video distortion. Using learning methods, we develop a packet delay bound violation probability model for video transmission over multihop networks with priority-based packet scheduling. With this model, we can successfully predict the probability of packets being dropped due to violation of specified delay bounds. We also observe that the transmission distortion caused by packet drops exhibits a unique exponential behavior with priority-based packet scheduling. With these analysis results, we formulate the resource allocation for multisession video transmission over networks with priority-driven packet scheduling under delay bound constraints as a multiobjective optimization problem. Evolutionary optimization methods based on single- and multiobjective genetic algorithms are proposed to solve the problem and obtain the optimal resource allocation. Extensive experiment results demonstrate the effectiveness of the proposed resource-distortion models and optimization algorithms.
Yongfei Zhang, Shiyin Qin, Bo Li 0006, Zhihai He
IEEE Trans. Circuits Syst. Video Technol.4
2014 High-Fidelity Component Substitution Pansharpening by the Fitting of Substitution Data
abstract
Due to the difference of “mean information” between substitution component and substituted component, spectral distortion often occurs in component substitution (CS) pansharpening. In this paper, a data fitting scheme is adopted to improve spectral quality in image fusion based on well-established CS approach. A generalized CS framework that is capable of modeling any CS image fusion method is also presented. In this framework, instead of injecting detail information of panchromatic (Pan) image into substituted component, the data fitting strategy is designed to adjust the mean information of Pan image in the construction of substitution component. The data fitting scheme involves two matrix subtractions and one matrix convolution. It is fast in implementation and is effective to avoid the spectral distortion problem. Experimental results on a large number of Pan and multispectral images show that the improved CS methods have good performance on the spatial and spectral fidelity. Moreover, experiments carried out on large-size images also show an excellent running time performance of the proposed methods.
Qizhi Xu, Bo Li 0006, Yun Zhang 0014
IEEE Trans. Geosci. Remote. Sens.2
2013 A person re-identification algorithm by using region-based feature selection and feature fusion
abstract
In outdoor surveillance, person appearances captured by different cameras have obvious variations due to different poses and viewpoints, which affect the accuracy of person re-identification. In this paper, a person re-identification algorithm by using region-based feature selection and future fusion is proposed to divide one body into the upper region and the lower region. According to their different characteristics, each region adopts different kinds of features, which can efficiently reduce the negative impact from different poses and viewpoints. Moreover, since different features of one region may have different intrinsic meanings, during the feature fusion, different features of one region are separately represented instead of being comprehensively processed. The proposed feature fusion can make full use of the salience of different features. The experimental results demonstrate that the proposed algorithm improves the accuracy of person re-identification compared with the state of the art.
Yanbing Geng, Hai-Miao Hu, Bo Li 0006
ICIP4
2013 Region-classification-based rate control for flicker suppression of I-frames in HEVC
abstract
In High Efficiency Video Coding (HEVC), the coding efficiency of I-frames is lower than P-frames and B-frames, which will cause the flicker artifact, especially in low bitrates applications. We propose a region-classification-based rate control for Coding Tree Units (CTUs) in I-frames to improve the reconstructed quality of I-frames to suppress the flicker artifact. The CTUs in I-frame are classified into three regions according to their motion vectors and complexity. When the bit budget of one I-frame is used up, the target bitrates for the remaining CTUs will be adjusted according to the regions they belong to, and the pixel-based unified rate-quantization (URQ) model is then used to calculate the QPs. Experimental results demonstrate that the proposed scheme can efficiently suppress the flicker artifacts and improve both the subjective and objective video quality when compared with the original scheme in HM9.0.
Yongfei Zhang, Hai-Miao Hu, Bo Li 0006
ICIP4
2013 Rate-distortion optimized unequal loss protection for video transmission over packet erasure channels
Yongfei Zhang, Shiyin Qin, Bo Li 0006, Zhihai He
Signal Process. Image Commun.3
2013 Naturalness Preserved Enhancement Algorithm for Non-Uniform Illumination Images
abstract
Image enhancement plays an important role in image processing and analysis. Among various enhancement algorithms, Retinex-based algorithms can efficiently enhance details and have been widely adopted. Since Retinex-based algorithms regard illumination removal as a default preference and fail to limit the range of reflectance, the naturalness of non-uniform illumination images cannot be effectively preserved. However, naturalness is essential for image enhancement to achieve pleasing perceptual quality. In order to preserve naturalness while enhancing details, we propose an enhancement algorithm for non-uniform illumination images. In general, this paper makes the following three major contributions. First, a lightness-order-error measure is proposed to access naturalness preservation objectively. Second, a bright-pass filter is proposed to decompose an image into reflectance and illumination, which, respectively, determine the details and the naturalness of the image. Third, we propose a bi-log transformation, which is utilized to map the illumination to make a balance between details and naturalness. Experimental results demonstrate that the proposed algorithm can not only enhance the details but also preserve the naturalness for non-uniform illumination images.
Shuhang Wang, Hai-Miao Hu, Bo Li 0006
IEEE Trans. Image Process.4
2012 Gradient-based fast decision for intra prediction in HEVC
abstract
As the next generation standard of video coding, the High Efficiency Video Coding(HEVC) achieves significantly better coding efficiency than all existing video coding standards, which is however at the cost of a much higher computation complexity. To address this issue, this paper presents a gradient-based fast decision algorithm for intra prediction in HEVC. More specifically, the intra prediction in HEVC is divided into two stages: prediction unit(PU) size decision and mode decision. At the PU size decision process, four orientation features are extracted from the coding unit by the intensity gradient filters to decide the texture complexity and texture direction of the coding unit, and then the texture direction is used to exclude impossible prediction modes at the mode decision process. Compared to HEVC reference software, the proposed algorithm saves around 56.7% of the encoding time in intra high efficiency setting and up to 70.86% in intra low complexity setting with slight performance degradation.
Yongfei Zhang, Zhe Li 0015, Bo Li 0006
VCIP3
2012 Region-Based Rate Control for H.264/AVC for Low Bit-Rate Applications
abstract
Rate control plays an important role in video coding. However, in the conventional rate control algorithms, the number and position of macroblocks (MBs) inside one basic unit for rate control is inflexible and predetermined. The different characteristics of the MBs are not fully considered. Also, there is no overall optimization of the coding of basic units. This paper proposes a new region-based rate control scheme for H.264/advanced video coding to improve the coding efficiency. The inter-frame information is explored to objectively divide one frame into multiple regions based on their rate-distortion (R-D) behaviors. The MBs with similar characteristics are classified into the same region, and the entire region, instead of a single MB or a group of contiguous MBs, is treated as a basic unit for rate control. A linear rate-quantization stepsize model and a linear distortion-quantization stepsize model are proposed to accurately describe the R-D characteristics for the region-based basic units. Moreover, based on the above linear models, an overall optimization model is proposed to obtain suitable quantization parameters for the region-based basic units. Experimental results demonstrate that the proposed region-based rate control approach can achieve both better subjective and objective quality by performing the rate control adaptively with the content, compared to the conventional rate control approaches.
Hai-Miao Hu, Bo Li 0006, Weiyao Lin, Wei Li 0209, Ming-Ting Sun
IEEE Trans. Circuits Syst. Video Technol.2
2011 Image enhancement based on Retinex and lightness decomposition
abstract
In this paper, an efficient image enhancement method based on Retinex and lightness decomposition is proposed, which enhances details and preserves the naturalness simultaneously. The quality of an enhanced image is determined by two factors, details and naturalness. Accordingly, the lightness is proposed to be decomposed into reflex lightness and ambience illumination. The reflex lightness is utilized to extract details based on Retinex, and the ambience illumination is utilized to maintain the naturalness of scenes. Moreover, a coarse evaluation that keeps the corresponding order of ambience illumination is proposed to substitute the absolute amount of ambience illumination, which may not be necessary to preserve the naturalness. The experimental results demonstrate that the proposed method can efficiently improve the perceptual quality of images by not only enhancing the details of the images, but also keeping the naturalness of scenes.
Bo Li 0006, Shuhang Wang, Yanbing Geng
ICIP1
2011 A rate-control algorithm using inter-layer information for H.264/SVC for low-delay applications
Hai-Miao Hu, Bo Li 0006, Weiyao Lin, Ming-Ting Sun
J. Vis. Commun. Image Represent.2
2011 A region-based rate-control scheme using inter-layer information for H.264/SVC
Hai-Miao Hu, Weiyao Lin, Bo Li 0006, Ming-Ting Sun
J. Vis. Commun. Image Represent.3
2011 Multiscale Contour Extraction Using a Level Set Method in Optical Satellite Images
abstract
This letter presents a novel coarse-to-fine level set method for contour extraction in optical satellite images. To distinguish objects from a background, the undecimated wavelet transform is firstly adopted to extract image features, and a homogeneity metric is defined to measure the variation of the features inside and outside contours. In addition, the weight distribution ratio is proposed to adaptively tune the relative weight of the features. Based on the homogeneity metric and the weight distribution ratio, a novel energy functional is developed to model a contour extraction problem, and in order to reduce the computation burden, a coarse-to-fine scheme is applied to progressively extract contours in finer scale, during which a contour position constraint is introduced to limit contours evolving in a small space around the candidate contours extracted in coarser scale. Extensive experiments have been carried out on optical satellite images to validate the proposed method.
Qizhi Xu, Bo Li 0006, Zhaofeng He 0001
IEEE Geosci. Remote. Sens. Lett.2
2011 Remote-Sensing Image Compression Using Two-Dimensional Oriented Wavelet Transform
abstract
In this paper, a 2-D oriented wavelet transform (OWT) is introduced for efficient remote-sensing image compression. The proposed 2-D OWT can perform integrative oriented transform in arbitrary direction and achieve a significant transform coding gain. To maximize the transform coding gain, two separable 1-D transforms are implemented in the same direction for local areas with direction consistency. Subpixel interpolation rules are designed for rectangular subbands generation. In addition, semidirection displacement is adjusted to handle direction mismatch after the first 1-D transform. Experimental results demonstrate that the proposed 2-D OWT compression scheme outperforms JPEG2000 for remote-sensing images with high resolution, up to 0.43 dB in peak signal-to-noise ratio (PSNR), 0.0261 in the measure of structural similarity, 0.44% in Kappa coefficients, respectively, and significant subjective improvement. Meanwhile, it outperforms JPEG2000, previous adaptive directional lifting and weighted adaptive lifting methods, up to 1.98, 0.36, and 0.19 dB in PSNR for natural images. Furthermore, it is suitable for real-time remote-sensing processing for its low computational cost.
Bo Li 0006
IEEE Trans. Geosci. Remote. Sens.1
2009 An error resilient video coding and transmission solution over error-prone channels
Hai-Miao Hu, Bo Li 0006
J. Vis. Commun. Image Represent.3
2008 Embedded zerotree wavelets coding based on adaptive fuzzy clustering for image compression
Xiaoyuan Yang 0003, Bo Li 0006
Image Vis. Comput.3
2007 Fast Adaptive Wavelet for Remote Sensing Image Compression
Bo Li 0006, Runhai Jiao, Yuancheng Li 0005
J. Comput. Sci. Technol.1
2006 A Fast Selection Algorithm for Multiple Reference Frames in H.264/AVC
Qing-lei Meng, Chun-lian Yao, Bo Li 0006
ICONIP (2)3
2006 Robust abnormity detecting and tracking using correlation coefficient
abstract
Abnormity detection based on computer vision is the foundation of high-level automatic surveillance that includes objects recognition, tracking and alarm etc. In this paper, a robust abnormity detection algorithm using improved correlation coefficient is proposed. It assembles every pixel and its neighboring pixels to form a vector, and computes correlation coefficient to measure the similarity of corresponding window in the background image and current image. When the correlation coefficient is bigger than the adaptive threshold, the algorithm processes pixel level detection. Moreover, using k-means and correlation coefficient respectively, it realizes the clustering and tracking of the abnormity objects. The experimental results show the algorithm is robust against noise disturbance, illumination change, shadows and reflection effect. It can detect abnormity precisely, and improve the surveillance system's adaptability for the complicated environment greatly.
Bo Li 0006, Chun-lian Yao
MMM2
2005 SVM Regression and Its Application to Image Compression
Runhai Jiao, Yuancheng Li 0005, Bo Li 0006
ICIC (1)4
2003 A Fast Block-Matching Algorithm Using Smooth Motion Vector Field Adaptive Search Technique
Bo Li 0006, Wei Li 0209, Yaming Tu
J. Comput. Sci. Technol.1
2000 A Novel Motion Estimation Algorithm Based on Dynamic Search Window and Spiral Search
Yaming Tu, Bo Li 0006, Jianwei Niu 0002
ICMI2
1993 M: An Approximate Reasoning System
abstract
A system of multivalued logical equations and its solution algorithm are put forward in this paper. Based on this work we generalize SLD-resolution into multivalued logic and establish the corresponding truth value calculus. As a result, M, an approximate reasoning system, is built. We present the language and inference rules of M. Furthermore, we analyse inconsistency of assignments to truth degrees and give the solving strategies of M.
Qinping Zhao, Bo Li 0006
Int. J. Pattern Recognit. Artif. Intell.2