EDBT 2026 Demo / reviewers in the wild / expert
Hai-Miao Hu
dblp:39/7528
· DBLP profile ↗
64ranked-venue papers
13as first author
28since 2021 · last 2026
0000-0001-6811-9209ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 46 · 12 first-author · 17 since 2021Artificial intelligence and machine learning · 13 · 1 first-author · 8 since 2021Databases, data management, data science and information retrieval · 5 · 4 since 2021Computer networks · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Place recognition for visual assistive localization under challenging visual appearance variations
Ruiqi Cheng, Hai-Miao Hu, Chongze Wang |
Comput. Vis. Image Underst. | 2 |
| 2026 | ASRL: Correlation-robust pedestrian attribute recognition via fixed orthogonal classifier
Hai-Miao Hu |
Neural Networks | 2 |
| 2025 | PEIE: Physics Embedded Illumination Estimation for Adaptive DehazingabstractDeep learning-based methods have made significant progress in image dehazing. However, these methods often falter when applied to real-world hazy images, primarily due to the scarcity of paired real-world data and the limitations of current dehazing feature extractors. Toward these issues, we introduce a novel Physics Embedded Illumination Estimation (PEIE) method for adaptive real-world dehazing. Specifically, (1) we identify the limitations of the widely used Atmospheric Scattering Model and propose a new physical model, the Illumination-Adaptive Scattering Model (IASM), for more accurate illumination representation in hazy imaging; (2) we develop a robust data synthesis pipeline that leverages the physics embedded illumination estimation to generate realistic hazy images; and (3) we design an Illumination-Adaptive Dehazing Unit (IDU) to extract dehazing features consistent with our proposed IASM in the latent space. By integrating the IDU into a U-Net architecture to create IADNet, we achieve significant improvements in dehazing performance through end-to-end training on synthetic data. Extensive experiments validate the superior performance of our PEIE method, significantly surpassing the state-of-the-arts in real-world dehazing. Huaizhuo Liu, Hai-Miao Hu, Yonglong Jiang, Yurui Liu |
AAAI | 2 |
| 2025 | DTS-BWpredictor: Dual-scale temporal strategy based bandwidth prediction in highly dynamic links
Difeng Zhu, Hai-Miao Hu |
Comput. Networks | 4 |
| 2025 | A Solution to Co-occurrence Bias in Pedestrian Attribute Recognition: Theory, Algorithms, and Improvements
Hai-Miao Hu, Jinzuo Yu, Shiliang Pu, Hanzi Wang |
Int. J. Comput. Vis. | 2 |
| 2025 | Heterogeneous Feature Re-Sampling for Balanced Pedestrian Attribute RecognitionabstractIn pedestrian attribute recognition (PAR), the loose umbrella term 'attribute' ranges from human soft-biometrics to wearing accessory, and even extending to various subjective body descriptors. As a result, the vast coverage of 'attributes' implies that, instead of being over-specialized to limited attributes with exclusive characteristic, PAR should be approached from a much fundamental perspective. To this end, given that most attributes are greatly under-represented in real-world datasets, we simply distill PAR into a visual task of multi-label recognition under significant data imbalance. Accordingly, we introduce feature re-sampled detached learning (FRDL) to decouple label-balanced learning from the curse of attributes co-occurrence. Specifically, FRDL is able to balance the sampling distribution of an attribute without biasing the label prior of co-occurring others. As a complementary method, we also propose gradient-oriented augment translating (GOAT) to alleviate the feature noise and semantics imbalance aggravated in FRDL. Integrated in a highly unified framework, FRDL and GOAT substantially refresh the state-of-the-art performance on various realistic benchmarks, while maintaining a minimal computational budget. Further analytical discussion and experimental evidence corroborate the veracity of our advancement: this is the first work that establishes labels-independent and impartial balanced learning for PAR. Bo Li 0006, Hai-Miao Hu, Hanzi Wang |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | CGATracker: Correlation-Aware Graph Alignment for Referring Multi-Object TrackingabstractReferring multi-object tracking (RMOT) aims to identify specific targets based on sentence descriptions. To enhance multi-modal learning, previous works typically relied on a simple fusion module at early or late stages. However, those methods frequently underutilize textual semantics and struggle to model the relationships between region-level features and word-level features. To address these limitations, we propose CGATracker, a correlation-aware graph alignment method for RMOT, which facilitates precise relationship modeling through relational scoring. Specifically, we design a Language-driven Relational Alignment (LRA) module, which establishes two connection graphs to generate positive and negative samples for the visual-textual alignment. Additionally, to effectively leverage referring information, we introduce a Semantic Clarify Booster (SCBooster) module based on a semantic infusion mechanism and a bias-aware verification mechanism for interactions with different modalities. Moreover, by designing a Multi-level Cross-modal Fusion (MCF) module, our method aggregates contextual features at multiple depths to enable the creation of the enriched correlation-aware graph. Extensive experiments conducted on the Refer-KITTI and Refer-KITTI-V2 datasets demonstrate the effectiveness of CGATracker. Siping Zhuang, Qiangqiang Wu, Yang Lu 0009, Hai-Miao Hu, Hanzi Wang |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | A Multi-Category Anomaly Editing Network With Correlation Exploration and Voxel-Level Attention for Unsupervised Surface Anomaly DetectionabstractDeveloping a unified model for surface anomaly detection remains challenging due to significant variations across product categories. Recent feature editing methods, as a branch of image reconstruction, mitigate the over-generalization of auto-encoders that leads to accurate anomaly reconstruction. However, these methods are only suited for texture-category products and have significant limitations in being generalized to other categories. In this article, we propose a multi-category anomaly editing network with a dual-branch training approach: one branch processes defect-free images (normal branch), while the other handles synthetic anomaly images (anomaly branch). Specifically, the paired samples are first fed into the multi-category anomaly feature editing based auto-encoder (MCAFE-AE) to perform image reconstruction and inpainting. In the normal branch, we propose a dual-entropy constrained deep embedded clustering module (DEC-DECM) to promote a more compact and orderly distribution of normal latent features, while avoiding trivial clustering solutions. Based on the clustering results, we further design a patch-based adaptive thresholding (PAT) strategy to adaptively calculate the threshold representing the central boundary of the cluster center for each local patch, thereby enabling the model to detect anomalies. Then, in the anomaly branch, we propose a multi-category anomaly feature editing module (MCAFEM) to identify anomalies in synthetic images and apply a category-oriented feature editing strategy to transform detected anomaly features into normal ones, thereby suppressing the reconstruction of anomalies. After completing the image reconstruction and inpainting, the input images from both branches and their respective output images are concatenated and fed into the correlation exploration and voxel-level attention based prediction network (CEVA-Net) for anomaly segmentation. The network is integrated with our proposed correlation-dependency exploration and voxel-level attention refinement module (CDE-VARM) and generates precise anomaly maps under the guidance of the bidirectional-path feature fusion (BPFF) and deep supervised learning (DSL). Extensive experiments on three datasets show that our method achieves state-of-the-art performance. Ruifan Zhang, Hai-Miao Hu |
IEEE Trans. Image Process. | 2 |
| 2024 | Pedestrian Attribute Recognition as Label-balanced Multi-label LearningabstractRooting in the scarcity of most attributes, realistic pedestrian attribute datasets exhibit unduly skewed data distribution, from which two types of model failures are delivered: (1) label imbalance: model predictions lean greatly towards the side of majority labels; (2) semantics imbalance: model is easily overfitted on the under-represented attributes due to their insufficient semantic diversity. To render perfect label balancing, we propose a novel framework that successfully decouples label-balanced data re-sampling from the curse of attributes co-occurrence, i.e., we equalize the sampling prior of an attribute while not biasing that of the co-occurred others. To diversify the attributes semantics and mitigate the feature noise, we propose a Bayesian feature augmentation method to introduce true in-distribution novelty. Handling both imbalances jointly, our work achieves best accuracy on various popular benchmarks, and importantly, with minimal computational budget. Hai-Miao Hu, Yirong Xiang |
ICML | 2 |
| 2024 | Cross-Modal Contrastive Learning Network for Few-Shot Action RecognitionabstractFew-shot action recognition aims to recognize new unseen categories with only a few labeled samples of each class. However, it still suffers from the limitation of inadequate data, which easily leads to the overfitting and low-generalization problems. Therefore, we propose a cross-modal contrastive learning network (CCLN), consisting of an adversarial branch and a contrastive branch, to perform effective few-shot action recognition. In the adversarial branch, we elaborately design a prototypical generative adversarial network (PGAN) to obtain synthesized samples for increasing training samples, which can mitigate the data scarcity problem and thereby alleviate the overfitting problem. When the training samples are limited, the obtained visual features are usually suboptimal for video understanding as they lack discriminative information. To address this issue, in the contrastive branch, we propose a cross-modal contrastive learning module (CCLM) to obtain discriminative feature representations of samples with the help of semantic information, which can enable the network to enhance the feature learning ability at the class-level. Moreover, since videos contain crucial sequences and ordering information, thus we introduce a spatial-temporal enhancement module (SEM) to model the spatial context within video frames and the temporal context across video frames. The experimental results show that the proposed CCLN outperforms the state-of-the-art few-shot action recognition methods on four challenging benchmarks, including Kinetics, UCF101, HMDB51 and SSv2. Xiao Wang 0072, Yan Yan 0001, Hai-Miao Hu, Bo Li 0006, Hanzi Wang |
IEEE Trans. Image Process. | 3 |
| 2024 | Explicit State Representation Guided Video-based Pedestrian Attribute RecognitionabstractThe pedestrian attribute recognition aims to generate a structured description of pedestrians, which serves an important role in surveillance. Current works usually assume that the images and the specific pedestrian states, including pedestrian occlusion and pedestrian orientation, are given. However, we argue that the current works ignore the guidance of the pedestrian state and cannot achieve the appropriate performance since the appearance feature will become unreliable due to the variance of the pedestrian state, which is common in practice. Therefore, this paper proposes the Explicit State Representation (ExSR) Guided Pedestrian Attribute Recognition to improve the accuracy through state learning and attribute fusion among frames. Firstly, the pedestrian state is explicitly represented by concatenating the pedestrian orientation and occlusion, which can be accurately determined via analyzing the pose. Secondly, the state-aware pedestrian attribute fusion method is proposed and divided into two cases, namely the inter-state case and the intra-state case. In the intra-state case, the appearance feature will remain stable and the attribute relations are propagated to refine. The method of exploiting attribute relations within a single frame is the Graph Neural Network. In the inter-state case, the state changes, the attribute relationship propagation is prevented, and the advantages of attribute recognition in each frame are complemented to make a reliable judgment on the invisible region. The experimental results demonstrate that the ExSR outperforms the state-of-the-art methods on two public databases, benefiting from the explicit introduction of the state into the attribute recognition. Weiqing Lu, Hai-Miao Hu, Jinzuo Yu, Hanzi Wang |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2024 | From Appearance to Inherence: A Hyperspectral Image Dataset and Benchmark of Material Classification for SurveillanceabstractImage understanding and analysis primarily rely on object appearances. However, when faced with challenges such as occlusion, camouflage, and small targets in surveillance scenarios, the discriminative features of objects cannot be effectively extracted. This paper proposes the utilization of hyperspectral imaging techniques as a solution to these issues, aiming to unlock new potential in the field. Hyperspectral images have the unique capability to identify a variety of materials, thereby offering a distinct advantage in surveillance. However, existing hyperspectral image datasets are not specifically tailored for image classification tasks within surveillance scenarios. To address this issue, we introduce an innovative hyperspectral image dataset designed explicitly for real-world surveillance, with the goal of setting a new benchmark for material classification. Our aspiration extends beyond merely deploying deep learning methods for hyperspectral material classification, aiming to contribute insightful understanding of spectral patterns inherent in natural surveillance scenes. The proposed dataset is currently the largest of its kind and the first one designed specifically for surveillance scenarios. It encompasses 128 spectral bands and provides annotations for 28 common material categories. Furthermore, we introduce a novel texture-metric-based spatial and spectral fusion network, meticulously crafted to accommodate our unique scenario and dataset. This model significantly outperforms existing networks in enhancing the fusion of spatial and spectral features, achieving state-of-the-art results on both our proposed dataset and existing public hyperspectral image classification datasets. For access to this unique dataset, please visithttps://github.com/robberban/HIC-for-Surveillance.git. Likun Gao, Hai-Miao Hu, Xinhui Xue, Haoxin Hu |
IEEE Trans. Multim. | 2 |
| 2024 | Orientation-Aware Pedestrian Attribute Recognition Based on Graph Convolution NetworkabstractPedestrian attribute recognition (PAR) aims to generate a structured description of pedestrians and plays an important role in surveillance. Current work focusing on 2D images can achieve decent performance when there is no variation in the captured pedestrian orientation. However, the performance of these works cannot be maintained in scenarios when the orientation of pedestrians is ignored. To mitigate this problem, this paper proposes orientation-aware pedestrian attribute recognition based on graph convolution network (GCN), which is composed of an orientation-aware spatial attention (OSA) module and an orientation-guided attribute-relation learning (OAL) module. Since some attributes can be invisible for certain orientations, OSA is proposed for orientation-aware feature extraction to enhance the learned representation of the visual attributes. Moreover, since different orientations result in different relations among attributes, OAL is proposed to achieve distinguishable and impactful attribute relations by eliminating the confusion of attribute relations in different orientations. Experiments on three challenging datasets (PETA, RAP, and PA100K) demonstrate that the proposed PAR outperforms the state-of-the-art methods by considerable margins. Weiqing Lu, Hai-Miao Hu, Jinzuo Yu, Hanzi Wang, Bo Li 0006 |
IEEE Trans. Multim. | 2 |
| 2024 | Instance-Based Continual Learning: A Real-World Dataset and Baseline for Fresh RecognitionabstractReal-time learning on real-world data streams with temporal relations is essential for intelligent agents. However, current online Continual Learning (CL) benchmarks adopt the mini-batch setting and are composed of temporally unrelated and disjoint tasks as well as pre-set class boundaries. In this paper, we delve into a real-world CL scenario for fresh recognition where algorithms are required to recognize a huge variety of products to facilitate the checkout speed. Products mainly consists of packaged cereals, seasonal fruits, and vegetables from local farms or shipped from overseas. Since algorithms process instance streams consisting of sequential images, we name this real-world CL problem as Instance-Based Continual Learning (IBCL) . Different from the current online CL setting, algorithms are required to perform instant testing and learning upon each incoming instance. Moreover, IBCL has no task boundaries or class boundaries and allows the evolution and the forgetting of old samples within each class. To promote the researches on real CL challenges, we propose the first real-world CL dataset coined the Continual Fresh Recognition (CFR) dataset, which consists of fresh recognition data streams (766 K labelled images in total) collected from 30 supermarkets. Based on the CFR dataset, we extensively evaluate the performance of current online CL methods under various settings and find that current prominent online CL methods operate at high latency and demand significant memory consumption to cache old samples for replaying. Therefore, we make the first attempt to design an efficient and effective Instant Training-Free Learning (ITFL) framework for IBCL. ITFL consists of feature extractors trained in the metric learning manner and reformulates CL as a temporal classification problem among several most similar classes. Unlike current online CL methods that cache image samples (150 KB per image) and rely on training to learn new knowledge, our framework only caches features (2 KB per image) and is free of training in deployment. Extensive evaluations across three datasets demonstrate that our method achieves comparable recognition accuracy to current methods with lower latency and less resource consumption. Our codes and datasets will be publicly available at https://github.com/detectRecog/IBCL . Zhenbo Xu, Hai-Miao Hu, Wenming Tan |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2023 | One-Shot Neural Band Selection for Spectral RecoveryabstractBand selection has a great impact on the spectral recovery quality. To solve this ill-posed inverse problem, most band selection methods adopt hand-crafted priors or exploit clustering or sparse regularization constraints to find most prominent bands. These methods are either very slow due to the computational cost of repeatedly training with respect to different selection frequencies or different band combinations. Many traditional methods rely on the scene prior and thus are not applicable to other scenarios. In this paper, we present a novel one-shot Neural Band Selection (NBS) framework for spectral recovery. Unlike conventional searching approaches with a discrete search space and a non-differentiable search strategy, our NBS is based on the continuous relaxation of the band selection process, thus allowing efficient band search using gradient descent. To enable the compatibility for selecting any number of bands in one-shot, we further exploit the band-wise correlation matrices to progressively suppress similar adjacent bands. Extensive evaluations on the NTIRE 2022 Spectral Reconstruction Challenge demonstrate that our NBS achieves consistent performance gains over competitive baselines when examined with four different spectral recovery methods. Our code will be publicly available. Hai-Miao Hu, Zhenbo Xu, Wenshuai Xu, You Song, YiTao Zhang, Zhilin Han, Ajin Meng |
ICASSP | 1 |
| 2023 | A Solution to Co-occurence Bias: Attributes Disentanglement via Mutual Information Minimization for Pedestrian Attribute RecognitionabstractRecent studies on pedestrian attribute recognition progress with either explicit or implicit modeling of the co-occurence among attributes. Considering that this known a prior is highly variable and unforeseeable regarding the specific scenarios, we show that current methods can actually suffer in generalizing such fitted attributes interdependencies onto scenes or identities off the dataset distribution, resulting in the underlined bias of attributes co-occurence. To render models robust in realistic scenes, we propose the attributes-disentangled feature learning to ensure the recognition of an attribute not inferring on the existence of others, and which is sequentially formulated as a problem of mutual information minimization. Rooting from it, practical strategies are devised to efficiently decouple attributes, which substantially improve the baseline and establish state-of-the-art performance on realistic datasets like PETAzs and RAPzs. Hai-Miao Hu, Jinzuo Yu, Zhenbo Xu, Weiqing Lu, Yuran Cao |
IJCAI | 2 |
| 2023 | From Global to Local: An Adaptive Environmental Illumination Estimation for Non-uniform ScatteringabstractThe atmospheric scattering model is one of the most widely used model to describe the optical imaging processing of hazy images. However, the global atmospheric light used in the traditional atmospheric scattering model has limitations in describing images with varying local environmental illumination. In this paper, by extending the global atmospheric light to the local illumination, a non-uniform scattering model is proposed, which can better describe real scenes under non-uniform environmental illumination. Based on this model, an adaptive local illumination estimation for hazy image is proposed, which can adapt to the local differences of environment illumination. The experimental results demonstrate that the proposed algorithm can outperform the state-of-the-art algorithms in terms of not only the non-uniform scattering removal but also the adaptability. Huaizhuo Liu, Hai-Miao Hu |
MMAsia | 2 |
| 2023 | A Spatial-Spectral Decoupling Fusion Framework for Visible and Near-Infrared ImagesabstractThe visible and near-infrared image fusion aims to generate an image that integrates complementary information from images captured in different spectral bands. However, existing fusion methods either only focus on the fusion of spatial information, or fuse spatial and spectral information without decoupling, resulting in undesirable effects such as halo artifacts, information loss, and inferior visual quality. To address these issues, we propose a Spatial-Spectral Decoupling Fusion (SSDF) framework that can effectively fuse the spatial and spectral information of visible and near-infrared images. The SSDF framework decomposes the image pairs into two main branches: the Spatial Feature Enhancement (SFE) branch and the Spectral Characteristic Preservation (SCP) branch. The SFE branch enhances the salient details in the fused image by exploiting the contrast between spatial features and generating region-based fusion weights, while the SCP branch preserves the intrinsic spectral characteristics of the scene by fusing the reflectance characteristics of visible and near-infrared images. The final image is obtained by combining the spatial and spectral information. We conduct extensive experiments to show that our SSDF method can achieve superior fusion performance in subjective visual quality and objective metrics compared with state-of-the-art methods. Zhenglin Tang, Hai-Miao Hu |
MMAsia | 2 |
| 2023 | HR-HAR: A hierarchical relation representation for human activity recognition based on Wi-FiabstractAbstract The Wi‐Fi‐based human activity recognition shows immense potential, as it is device‐free, non‐intrusive to privacy, and low‐cost. However, current learning‐based recognition methods mostly adopt the hybrid representation without distinguished contributions of features to different activities, which will be seriously affected by environment variations and interference of other persons, and costly to extend to new activities. Therefore, this paper proposes HR‐HAR, a hierarchical relation representation for human activity recognition, to improve the performance, extensibility, and robustness by exploiting the hierarchical relation of features of activities. The hierarchical relation reflects the different contributions of features to recognize different activities and effectively distinguishes similar activities. It naturally leads to a layered structure that can be extended to new activities without re‐training the entire model. With the layered structure, HR‐HAR first detects the existence of other persons and then processes un‐interfered scene and interfered scene signals with different methods, so it is robust to the interference. The experimental results on the public dataset with 95.6% accuracy and on the self‐collected dataset with 95.4% accuracy for un‐interfered scene and 95.0% for interfered scene indicate that HR‐HAR is of reliable performance on human activity recognition and is robust to environmental changes and interference of other persons. Yanglin Pu, Yongqiang Jiang, Hai-Miao Hu |
IET Commun. | 3 |
| 2023 | Improving Color Constancy Using Chromaticity-Line PriorabstractColor constancy is the ability to remove the effect of illumination on color. Since color constancy is an ill-posed problem, many methods have been proposed based on assumptions to constraint the solution space. However, most existing assumptions require specular pixels or abundant colors, and fail to produce satisfactory results for different scenarios. According to extensive experiments, we observe that the chromaticity distribution of pixels within main color under canonical illumination, which we called canonical pixels, is linear and can also locate the position of illumination under the non-canonical illumination. Therefore, this paper proposes a chromaticity-line prior (CLP) as an additional linear constraint on the ill-posed problem of color constancy. In the calculation of CLP, the simple linear iterative clustering is firstly employed to segment an image into several super-pixel blocks. And the random sampling consensus is utilized to remove non-primary color points and fit the chromaticity-line. Based on the proposed CLP, a color constancy algorithm is implemented correspondingly. Since the main idea of the CLP is to extract the canonical pixels, which is the inherent property of image, the proposed CLP is more general and adaptive in real scenes. The experiments on two public datasets demonstrate that the proposed algorithm not only outperforms state-of-the-art learning-free algorithms, but also achieves results that are competitive to those of learning-based algorithms. Hai-Miao Hu, Hongda Zhang, Qiang Guo 0013 |
IEEE Trans. Multim. | 2 |
| 2022 | Embedding Adaptation Network with Transformer for Few-Shot Action Recognition
Rongrong Jin, Xiao Wang 0072, Guangge Wang, Yang Lu 0009, Hai-Miao Hu, Hanzi Wang |
ACML | 5 |
| 2022 | A Progressive Domain Adaptation for Object Detection via Coarse-grained Foreground GuidanceabstractDomain adaptive object detection has always been a problem widely concerned by academia and industry. The variability of application scenes and the unknown target labels in the target domain make it challenging for well-trained detectors to generalize well by supervised learning. Due to the lack of instance-level information, previous approaches often tend to achieve image-level and pixel-level alignment using the model's own knowledge, but this is of finite help in separating the object to be detected from the image background, and it is difficult for the model itself to learn the foreground characteristics of the target domain. In this paper, we propose a method for domain-adaptive detection based on coarse-grained foreground guidance. Our approach uses an unsupervised foreground extraction algorithm to estimate the coarse locations of objects in the video, which can provide knowledge for instance features and progressively improve the cross-domain detection ability of the model. Experiments on the cross-domain object detection dataset confirm the effectiveness of our approach. We have also conducted experiments in real outdoor traffic scenarios, and the results indicate that our method can operate in real application scenarios. Likun Gao, Hai-Miao Hu, Mingzhu Li |
MMSP | 2 |
| 2022 | Revealing the real-world applicable setting of online continual learningabstractThe motivation of online continual learning (CL) is training agents to learn from an infinite stream of data and quickly accommodate changes in the data distribution. However, current online CL datasets are synthesized by common classification datasets by splitting all classes into disjoint tasks where disjoint task streams have little temporal relations, resulting in a CL setting far from realistic. In this paper, we ask two questions: (i) What are the characteristics of real-world CL scenarios? (ii) How existing methods perform on real-world CL scenarios? To answer the first question, we propose the first realistic CL setting coined instance-based continual learning (IBCL). IBCL has no task or class boundaries and requires algorithms to predict and learn from instance streams simultaneously. The life cycles of classes under IBCL are dynamic and instances belonging to the same class might evolve over time. For each sequentially arrival instance, algorithms are required to give the recognition result and then perform changes based on its label. No additional training resource are available except for the instance stream in evaluation. To answer the second question, on CORe50 and mini-ImageNet, we compare current online CL methods under the IBCL setting with both the traditional ResNet18 backbone as well as the recent transformer-based backbone ViT on the IBCL setting. Three aspects including the recognition performance, the latency, and the memory usage of current methods are analyzed. Experiment results show that current online CL methods perform poorly in the real CL scenarios, and methods using the transformer-based backbone perform better than the CNN-based counterparts. Zhenbo Xu, Hai-Miao Hu |
MMSP | 2 |
| 2022 | Referring image segmentation with attention guided cross modal fusion for semantic oriented languages
Qianli Zhou, Hai-Miao Hu, Quange Tan |
Frontiers Comput. Sci. | 3 |
| 2022 | Correlation Graph Convolutional Network for Pedestrian Attribute RecognitionabstractThe pedestrian attribute recognition aims at generating the structured description of pedestrian, which plays an important role in surveillance. However, it is difficult to achieve accurate recognition results due to diverse illumination, partial body occlusion and limited resolutions. Therefore, this paper proposes a comprehensive relationship framework for comprehensively describing and utilizing relations among attributes, describing different type of relations in the same dimension, and implementing complex transfers of relations in a GCN manner. This framework is named Correlation Graph Convolutional Network (CGCN). Based on the proposed framework, the feature vectors are built to associate attributes with image features and generate different relation matrices through self-attention among different feature vectors, describing different attribute relations. Then, we conduct multi-layer transfer of attribute relations by means of graph convolution, realizing complex utilization of attribute relations. In addition, the relations among attributes are fully exploited and two types of relations, namely the explicit and implicit relations, are proposed to be integrate into the proposed comprehensive relationship framework. The experimental results on RAP and PETA demonstrate that the recognition performance of the proposed CGCN can obviously outperform the state-of-the-arts, and moreover, the CGCN can achieve a better synergy with different relations. Haonan Fan, Hai-Miao Hu, Shuailing Liu, Weiqing Lu, Shiliang Pu |
IEEE Trans. Multim. | 2 |
| 2021 | Attentive Excitation and Aggregation for Bilingual Referring Image SegmentationabstractThe goal of referring image segmentation is to identify the object matched with an input natural language expression. Previous methods only support English descriptions, whereas Chinese is also broadly used around the world, which limits the potential application of this task. Therefore, we propose to extend existing datasets with Chinese descriptions and preprocessing tools for training and evaluating bilingual referring segmentation models. In addition, previous methods also lack the ability to collaboratively learn channel-wise and spatial-wise cross-modal attention to well align visual and linguistic modalities. To tackle these limitations, we propose a Linguistic Excitation module to excite image channels guided by language information and a Linguistic Aggregation module to aggregate multimodal information based on image-language relationships. Since different levels of features from the visual backbone encode rich visual information, we also propose a Cross-Level Attentive Fusion module to fuse multilevel features gated by language information. Extensive experiments on four English and Chinese benchmarks show that our bilingual referring image segmentation model outperforms previous methods. Qianli Zhou, Tianrui Hui, Hai-Miao Hu, Si Liu 0001 |
ACM Trans. Intell. Syst. Technol. | 4 |
| 2021 | Hierarchical Reasoning Network for Pedestrian Attribute RecognitionabstractPedestrian attribute recognition, which can benefit other tasks such as person re-identification and pedestrian retrieval, is very important in video surveillance related tasks. In this paper, we observe that the existing methods tackle this problem from the perspective of multi-label classification without considering the hierarchical relationships among the attributes. In human cognition, the attributes can be categorized according to their semantic/abstraction levels. The high-level attributes can be predicted by reasoning from the low-level and medium-level attributes, while the recognition of the low-level and medium-level attributes can be guided by the high-level attributes. Based on this attribute categorization, we propose a novel Hierarchical Reasoning Network (HR-Net), which can hierarchically predict the attributes at different abstraction levels in different stages of the network. We also propose an attribute reasoning structure to exploit the relationships among the attributes at different semantic levels. Experimental results demonstrate that the proposed network gives superior performances compared to the state-of-the-art techniques. Haoran An, Hai-Miao Hu, Yuanfang Guo, Qianli Zhou, Bo Li 0006 |
IEEE Trans. Multim. | 2 |
| 2021 | Spectrum Characteristics Preserved Visible and Near-Infrared Image Fusion AlgorithmabstractThe visible and near-infrared images fusion aims at utilizing their spectrum characteristics to enhance visibility. However, the current visible and near-infrared fusion algorithms cannot well preserve spectrum characteristics, which results in color distortion and halo artifacts. Therefore, this paper proposes a new visible and near infrared images fusion algorithm by fully considering their different reflection and scattering characteristics. According to image degradation model, the reflection weight model and the transmission weight model are established, respectively. The reflection weight model is established by calculating the difference between the visible (red, green, and blue) spectra and the near-infrared spectrum, while maintaining the correlation of the visible spectra. The proposed reflection weight model can preserve the original reflection characteristic of objects in natural scenes. On the other hand, the transmission weight model is explicitly proposed by calculating the gradient ratio of the visible spectra to the near-infrared spectrum. The proposed transmission weight model intends to make full use of the strong transmission performance of the near-infrared spectrum, which can complement the details loss of the visible spectra caused by light scattering. Moreover, the fused image based on two models is further enhanced according to the reflection characteristics of near-infrared spectrum in case of the non-uniform illumination. The experimental results demonstrate that the proposed algorithm can not only well preserve spectrum characteristics, but also avoid color distortion while maintaining the naturalness, which outperforms the state-of-the-art. Hai-Miao Hu, Wei Zhang 0181, Shiliang Pu, Bo Li 0006 |
IEEE Trans. Multim. | 2 |
| 2020 | Adaptive Single Image Dehazing Using Joint Local-Global Illumination AdjustmentabstractHaze has a serious impact on the outdoor optical imaging systems, and it will result in image blurring, color shift, and saturation reduction. Recently, many single image dehazing algorithms have been proposed for practical applications, such as surveillance. However, since the widely-used global atmospheric light in image dehazing fails to well describe the local illumination differences of images, these algorithms fail to well adapt to scenes with different haze concentrations and lighting conditions. Therefore, this paper proposes an adaptive single image dehazing algorithm using joint local-global illumination adjustment. A local illumination estimation for hazy image is proposed to replace the global atmospheric light constant in the atmospheric scattering model, and it can better adapt to the local differences of image illumination. Correspondingly, the global atmospheric light constant is proposed to be utilized to adaptively compensate the illumination intensity, which may better overcome the dark illumination problem within the dehazed image. The experimental results demonstrate that the proposed algorithm can outperform the state-of-the-art algorithms in terms of not only the dehazing effect but also the adaptability. Hai-Miao Hu, Hongda Zhang, Bo Li 0006 |
IEEE Trans. Multim. | 1 |
| 2019 | WiLay: building wi-fi-based human activity recognition system through activity hierarchical relationshipabstractRecently, Wi-Fi-based human activity recognition technique has attracted attentions extensively. Due to its ease of access and low cost, Wi-Fi-based technique achieves great potential on building human activity recognition systems. However, this technique is limited because the Wi-Fi signal is less-informative and susceptible to environmental changes. To build a practical Wi-Fi-base human activity recognition system, in this paper, WiLay, a layer-structured human activity recognition system is proposed. To recognize 7 different activities, WiLay used an activity-oriented process to select and extract features according to the hierarchical relationship between different activities, and trained multiple classifiers to build its layer-structured recognition system. We collected data on several different environments and tested our system. The experimental results with 95.4% accuracy and 89.1% recall rate indicate that our system has very well performance on recognition human activities and is robust to environmental changes. Yongqiang Jiang, Hai-Miao Hu, Yanglin Pu |
MobiQuitous | 2 |
| 2019 | Part-guided Network for Pedestrian Attribute RecognitionabstractPedestrian attribute recognition, which can benefit other tasks such as person re-identification and pedestrian retrieval, is very important in video surveillance related tasks. In this paper, we observe that the existing methods tackle this problem from the perspective of multi-label classification without considering the spatial location constraints, which means that the attributes tend to be recognized at certain body parts. Based on that, we propose a novel Part-guided Network (P-Net), which guides the refined convolutional feature maps to capture different location information for the attributes related to different body parts. The part-guided attention module employs the pix-level classification to produce attention maps which can be interpreted as the probability of each pixel belonging to the 6 pre-defined body parts. Experimental results demonstrate that the proposed network gives superior performances compared to the state-of-the-art techniques. Haoran An, Haonan Fan, Kaiwen Deng, Hai-Miao Hu |
VCIP | 4 |
| 2019 | Fast Image Deblurring Based on Dual-Exposure PriorabstractImage deblurring is an important task for practical applications. However, it remains a challenge due to the serious camera shake. In this paper, we propose a prior named Dual-Exposure Prior (DEP) according to the observation that image captured with a relatively short exposure time may achieve sharp edges. Based on the DEP, the blur kernels of the blurry images can be accurately estimated. An effective optimization between blur kernel estimation and intermediate image restoration is established by using the L2-regularized DEP in the maximum a posterior framework. Each sub-problem of the optimization is convex, it can be solved in close-form with fast Fourier transforms acceleration. The experiments based on real and synthesized blurry images show that the proposed image deblurring algorithm can outperform the state-of-the-art methods in terms of subjective and objective quality. Moreover, the proposed algorithm can be implemented with significantly low computational complexity. Hai-Miao Hu |
VCIP | 2 |
| 2019 | An Adaptive Multi-Projection Metric Learning for Person Re-Identification Across Non-Overlapping CamerasabstractPerson re-identification is one of the most important and challenging problems in video analytics systems; it aims to match people across non-overlapping camera views. For person re-identification, metric learning is introduced to improve the performance by providing a metric adapted for cross-view matching. The essence of metric learning is to search for an optimal projection matrix to project the original features into a new feature space. However, most existing metric learning methods overlook the inconsistency of feature distributions in multiple cameras. In this paper, we propose a multi-projection metric learning (MPML) method to overcome the inconsistency among multiple cameras in person re-identification. Our solution is to jointly learn multiple projection matrices using paired samples from different cameras to project features from different cameras into a common feature space. To make our method adaptive to newly added cameras without affecting the learned projection matrices, we further propose an adaptive MPML method, which can learn new camera projection matrices without having to update any of the obtained projection matrices. The proposed methods are evaluated on four major person re-identification data sets, with comprehensive experiments showing the effectiveness of the proposed methods and notable improvements over the state-of-the-art approaches. Hai-Miao Hu, Bo Li 0006, Qi Tian 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2019 | Haze and Thin Cloud Removal Using Elliptical Boundary Prior for Remote Sensing ImageabstractRemote sensing images play important roles in various earth surface observation applications. However, the hazy state of surface atmosphere can visually decrease the contrast and availability of remote sensing images. In this paper, we propose a haze and thin cloud removal method for single visible remote sensing images, which aims to robustly estimate haze thickness, atmospheric light, and transmission value from a remote sensing image with dense haze or thin cloud, and finally recovers a haze-free image. An elliptical boundary prior (EBP) is proposed to transform the haze thickness in each local patch from the pixels cluster in the spectral space, which is surrounded by an ellipse. With the aim of preventing highlight objects influences, an atmospheric light estimation approach is presented. The correlation of transmission and haze thickness is reconstructed to develop the scattering model for remote sensing images. The experimental results demonstrate that the proposed method can not only significantly improve the contrast and restore textures of various kinds of hazy remote sensing images but also well preserve the spectral information of visible bands. Qiang Guo 0013, Hai-Miao Hu, Bo Li 0006 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2019 | Single Image Defogging Based on Illumination Decomposition for Visual Maritime SurveillanceabstractSingle image fog removal is important for surveillance applications and many defogging methods have been proposed, recently. Due to the adverse atmospheric conditions, the scattering properties of foggy images depend on not only the depth information of scene, but also the atmospheric aerosol model, which has more prominent influence on illumination in a fog scene than that in a haze scene. However, recent defogging methods confuse haze and fog, and they fail to consider fully about the scattering properties. Thus, these methods are not sufficient to remove fog effects, especially for images in maritime surveillance. Therefore, this paper proposes a single image defogging method for visual maritime surveillance. Firstly, a comprehensive scattering model is proposed to formulate a fog image in the glow-shaped environmental illumination. Then, an illumination decomposition algorithm is proposed to eliminate the glow effect on the airlight radiance and recover a fog layer, in which objects at the infinite distance have uniform luminance. Secondly, a transmission-map estimation based on the non-local haze-lines prior is utilized to constrain the transmission map into a reasonable range for the input fog image. Finally, the proposed illumination compensation algorithm enables the defogging image to preserve the natural illumination information of the input image. In addition, a fog image dataset is established for visual maritime surveillance. The experimental results based on the established dataset demonstrate that the proposed method can outperform the state-of-the-art methods in terms of both the subjective and objective evaluation criteria. Moreover, the proposed method can effectively remove fog and maintain naturalness for fog images. Hai-Miao Hu, Qiang Guo 0013, Hanzi Wang, Bo Li 0006 |
IEEE Trans. Image Process. | 1 |
| 2019 | Detail Preserved Single Image Dehazing Algorithm Based on Airlight RefinementabstractSingle-image haze removal is important for many practical applications (e.g., surveillance). However, dehazed results of existing algorithms tend to be oversmoothed with missing fine image details. This drawback is caused by two factors: inaccurate airlight estimations and disregarding multiple scattering. In this paper, we propose a detail-preserving image dehazing algorithm based on two key priors, namely, the depth-edge aware prior and the airlight impact regularity prior. The proposed algorithm makes contributions in both the haze removal step and the postprocessing step. First, based on the depth-edge aware prior, an airlight refinement algorithm is proposed. The gradient strength of the minimum channel is employed to calculate punishment weights to smooth the dark channel. Second, based on the airlight impact regularity prior, an adaptive sharpening model that considers the refined airlight to determine the sharpening strength value is established to enhance levels of detail. Experimental results demonstrate that the proposed algorithm cannot only effectively remove haze but can also enhance levels of detail to thus outperform the state of the art on a wide variety of images. Hai-Miao Hu, Bo Li 0006, Qiang Guo 0013, Shiliang Pu |
IEEE Trans. Multim. | 2 |
| 2019 | A Novel Projective-Consistent Plane Based Image Stitching MethodabstractWhen different target surfaces, in three-dimensional space, are mapped onto an image plane, they have different projections. These projections vary with the viewpoint. These local differences have influence on the accuracy of image stitching. Most of the existing image stitching methods divide an input image into a number of fixed-size cells, and the pixels within the same cell are then warped using the same local transformation model for the alignment. These methods are based on the hypothesis that the transformation models in one cell are consistent. However, this hypothesis does not hold in general. In this paper, we propose a novel projective-consistent plane based image stitching method (termed PCPS). It divides the overlapping regions of an input image into some projective-consistent planes according to the normal vectors' orientations of local regions and the reprojection errors of aligned images. The local projective transformation model is estimated for each projective-consistent plane. And then, a hybrid warping model is estimated. For the pixels in overlapping regions, the local projective transformation models are adopted to achieve a better alignment. While for the pixels in non-overlapping regions, a global projective transformation model is estimated by using the inliers uniformly distributed in the projective-consistent planes to avoid distortion. Compared with the state-of-the-art image stitching methods, the experimental results on a number of challenging image sequences show that the projective transformation model estimated by the proposed PCPS method for each projective-consistent plane is more accurate, and the achieved stitching results have less seams and projective distortion. Yue Wang 0029, Hanzi Wang, Bo Li 0006, Hai-Miao Hu |
IEEE Trans. Multim. | 5 |
| 2018 | Adaptive Query Re-ranking Based on ImageGraph for Image RetrievalabstractWith the exponential growth of images, the accuracy of image retrieval for improving the performance of rank layer and feature layer in content-based image retrieval (CBIR) becomes more and more attractive. Better single feature search results will further enhance the effect. Better sorting results will robust the performance. Therefore, this paper focuses on rank reordering to improve the performance of image retrieval. We propose a rank-level framework for feature reordering based on hierarchical undirected graphs. First, we calculate the K Nearest Neighbors for each image in the database. Then, we propose a method based on the combination of reciprocal nearest neighbor and K nearest neighbor to construct an undirected graph, and a method for calculating the image distance based on the shared nearest neighbor. An undirected graph is used to represent the similarity relationship of an image, where the vertices are composed of pictures and the edges are weighted according to the distance of the pictures. Finally, based on the constructed undirected graph in this paper, we perform sorting optimization. Experiments on four public data sets demonstrate the effectiveness of the method and are challenging. Haonan Fan, Hai-Miao Hu, Yugui Zhang |
IEEE BigData | 2 |
| 2018 | Perceptual hash-based feature description for person re-identification
Hai-Miao Hu, Zihao Hu, Shengcai Liao, Bo Li 0006 |
Neurocomputing | 2 |
| 2018 | Too Far to See? Not Really! - Pedestrian Detection With Scale-Aware Localization PolicyabstractA major bottleneck of pedestrian detection lies on the sharp performance deterioration in the presence of small-size pedestrians that are relatively far from the camera. Motivated by the observation that pedestrians of disparate spatial scales exhibit distinct visual appearances, we propose in this paper an active pedestrian detector that explicitly operates over multiple-layer neuronal representations of the input still image. More specifically, convolutional neural nets, such as ResNet and faster R-CNNs, are exploited to provide a rich and discriminative hierarchy of feature representations, as well as initial pedestrian proposals. Here each pedestrian observation of distinct size could be best characterized in terms of the ResNet feature representation at a certain layer of the hierarchy. Meanwhile, initial pedestrian proposals are attained by the faster R-CNNs techniques, i.e., region proposal network and follow-up region of interesting pooling layer employed right after the specific ResNet convolutional layer of interest, to produce joint predictions on the bounding-box proposals' locations and categories (i.e., pedestrian or not). This is engaged as an input to our active detector, where for each initial pedestrian proposal, a sequence of coordinate transformation actions is carried out to determine its proper x-y 2D location and the layer of feature representation, or eventually terminated as being background. Empirically our approach is demonstrated to produce overall lower detection errors on widely used benchmarks, and it works particularly well with far-scale pedestrians. For example, compared with 60.51% log-average miss rate of the state-of-the-art MS-CNN for far-scale pedestrians (those below 80 pixels in bounding-box height) of the Caltech benchmark, the miss rate of our approach is 41.85%, with a notable reduction of 18.66%. Xiaowei Zhang 0003, Li Cheng 0001, Bo Li 0006, Hai-Miao Hu |
IEEE Trans. Image Process. | 4 |
| 2018 | Naturalness Preserved Nonuniform Illumination Estimation for Image Enhancement Based on RetinexabstractIllumination estimation is important for image enhancement based on Retinex. However since illumination estimation is an ill-posed problem it is difficult to achieve accurate illumination estimation for nonuniform illumination images. The conventional illumination estimation algorithms fail to comprehensively take all the constraints into the consideration such as spatial smoothness sharp edges on illumination boundaries and limited range of illumination. Thus these algorithms cannot effectively and efficiently estimate illumination while preserving naturalness. In this paper we present a naturalness preserved illumination estimation algorithm based on the proposed joint edge-preserving filter which exploits all the abovementioned constraints. Moreover a fast estimation is implemented based on the box filter. Experimental results demonstrate that the proposed algorithm can achieve the adaptive smoothness of illumination beyond edges and ensure the range of the estimated illumination. When compared with other state-of-the-art algorithms it can achieve better quality from both subjective and objective aspects. Hai-Miao Hu, Bo Li 0006, Qiang Guo 0013 |
IEEE Trans. Multim. | 2 |
| 2017 | Scale-aware hierarchical loss: A multipath RPN for multi-scale pedestrian detectionabstractPedestrians with different spatial scales exhibiting dramatically differences, the serious performance decline with decreasing resolution is the major bottleneck for current pedestrian detection. Considering the local feature differences for multi-scale pedestrians, a scale-aware multipath region proposal network is exploited to improve the recall rate, which is divided into several branches to generate a proper object proposal for target with specific scale range. Moreover, motivated by the visual semantic concepts of different convolutional layers, a scale-aware hierarchical loss model is introduced to minimize the error rate for pedestrians with different scales, in which the hierarchical features of higher convolutional layers are jointed to calculate a multi-task loss to learn scale-aware weighting of multipath region proposal network for each object proposal. Finally, compared to state-of-the-art methods, experimental results on the challenging ETH and Caltech benchmark show the superiority of the proposed method for large variance in instance scales. Xiaowei Zhang 0003, Bo Li 0006, Hai-Miao Hu |
VCIP | 3 |
| 2017 | A region-based video de-noising algorithm based on temporal and spatial correlations
Hai-Miao Hu, Qiang Guo 0013, Bo Li 0006 |
Neurocomputing | 1 |
| 2017 | Complexity-based intra frame rate control by jointing inter-frame correlation for high efficiency video coding
Mingliang Zhou 0001, Yongfei Zhang, Bo Li 0006, Hai-Miao Hu |
J. Vis. Commun. Image Represent. | 4 |
| 2017 | An auto-encoder-based summarization algorithm for unstructured videos
Mengxiong Han, Hai-Miao Hu, Rong-Peng Tian |
Multim. Tools Appl. | 2 |
| 2017 | A person re-identification algorithm based on pyramid color topology feature
Hai-Miao Hu, Guodong Zeng, Zihao Hu, Bo Li 0006 |
Multim. Tools Appl. | 1 |
| 2017 | A region-based intra-frame rate control scheme by jointing inter-frame dependency and inter-frame correlation
Hai-Miao Hu, Mingliang Zhou 0001, Naiyu Yin |
Multim. Tools Appl. | 1 |
| 2017 | An Adaptive Fusion Algorithm for Visible and Infrared Videos Based on Entropy and the Cumulative Distribution of Gray LevelsabstractVisible videos captured under different weather conditions may exhibit different characteristics, and thermal infrared videos are easily affected by ambient temperature variations; this sensitivity to environmental conditions makes the fusion of visible and thermal infrared videos a challenge. This paper proposes an adaptive fusion algorithm for visible and infrared videos, and uses cumulative distribution of gray levels and the entropy to adaptively retain infrared-hot targets and visible textures. The original visible and infrared frames are decomposed into two layers, namely, the base layer and the detail layer. The guided filter is employed to decompose frames due to its high efficiency. Two weight maps, one for the infrared base layer and one for the visible base layer, are adaptively generated based on the cumulative distribution of gray levels and the entropy, respectively. The visible base layer and the infrared base layer are fused based on their weight maps. The final fusion result is obtained by combining the fused base layer with the visible detail layer. Experimental results demonstrate that the proposed algorithm can achieve better fusion results compared with state-of-the-art methods. Hai-Miao Hu, Bo Li 0006, Qiang Guo 0013 |
IEEE Trans. Multim. | 1 |
| 2016 | A realtime fusion algorithm of visible and infrared videos based on spectrum characteristicsabstractThe fusion of visible and infrared videos can improve the perceptual quality of videos under severe environments, which is important for video surveillance. Since the infrared and visible videos have different spectrum characteristics, it is difficult to retain the hot targets in the fused videos and enhance the textures at the same time. Moreover, a real-time fusion algorithm is highly required in video surveillance. Therefore, this paper proposes a real-time fusion algorithm of visible and infrared videos by fully utilizing their different spectrum characteristics. A hot-target-oriented fusion based on the gray distribution is firstly proposed to retain the hot targets in the fused videos. Moreover, a texture-enhanced fusion is proposed to enhance the texture information of the fused videos based on the guided filter. Furthermore, the inter-frame correlations are employed to speed up fusion process. The experimental results demonstrate that the proposed algorithm can achieve real-time video fusion and improve both subjective and objective qualities compared with the state-of-the-art. Hai-Miao Hu |
ICIP | 2 |
| 2016 | A hierarchal BoW for image retrieval by enhancing feature salience
Hai-Miao Hu, Bo Li 0006 |
Neurocomputing | 2 |
| 2016 | An improved RANSAC based on the scale variation homogeneity
Yue Wang 0029, Qizhi Xu, Bo Li 0006, Hai-Miao Hu |
J. Vis. Commun. Image Represent. | 5 |
| 2015 | Video denoising algorithm via multi-scale joint luma-chroma bilateral filterabstractVideo denoising is important for display and subsequent analysis, but remains to be a challenging problem. Key insights that limit the performance of algorithms include two main aspects. First, low-frequency scene information and the coarse-grained noise in the chroma are mixed with each other, which is different from that in the luma. Thus, denoising the chroma by using only its own information is difficult. Second, it is impossible to directly use the denoised luma edge information to guide the chroma denoising for high quality results. Therefore, in this paper, we propose a multi-scale joint luma-chroma bilateral filter to improve chroma denoising performance. Experimental results demonstrate that the proposed algorithm surpasses state-of-the-art algorithms in most cases. Hai-Miao Hu |
VCIP | 2 |
| 2015 | Pedestrian detection based on hierarchical co-occurrence model for occlusion handling
Xiaowei Zhang 0003, Hai-Miao Hu, Bo Li 0006 |
Neurocomputing | 2 |
| 2015 | A person re-identification algorithm by exploiting region-based feature salience
Yanbing Geng, Hai-Miao Hu, Guodong Zeng |
J. Vis. Commun. Image Represent. | 2 |
| 2015 | Joint global-local information pedestrian detection algorithm for outdoor video surveillance
Hai-Miao Hu, Xiaowei Zhang 0003, Bo Li 0006 |
J. Vis. Commun. Image Represent. | 1 |
| 2014 | A person re-identification algorithm based on color topologyabstractThe color feature is one of the main features in person re-identification. Due to the illumination variation and similarity interference, the color feature is not robust. For practical applications, different pedestrian may have similar statistical distribution in terms of color histogram, while the same pedestrian may have different statistical distribution. However, it can be observed that although the color statistical information is not robust, the color spatial distribution can remain stable. According to this observation, this paper proposes a new feature, namely the color topology, to describe the color spatial distribution. One image is segmented into patches and the proposed color topology is represented according to both the gradient and the value of color changes among adjacent patches. Based on the color topology, a person re-identification algorithm is implemented in this paper. The experimental results demonstrate that the proposed algorithm can significantly improve the recognition performance when compared with the state-of-the-art. Guodong Zeng, Hai-Miao Hu, Yanbing Geng |
ICIP | 2 |
| 2014 | A fast image dehazing algorithm based on negative correction
Hai-Miao Hu, Shuhang Wang, Bo Li 0006 |
Signal Process. | 2 |
| 2013 | A person re-identification algorithm by using region-based feature selection and feature fusionabstractIn outdoor surveillance, person appearances captured by different cameras have obvious variations due to different poses and viewpoints, which affect the accuracy of person re-identification. In this paper, a person re-identification algorithm by using region-based feature selection and future fusion is proposed to divide one body into the upper region and the lower region. According to their different characteristics, each region adopts different kinds of features, which can efficiently reduce the negative impact from different poses and viewpoints. Moreover, since different features of one region may have different intrinsic meanings, during the feature fusion, different features of one region are separately represented instead of being comprehensively processed. The proposed feature fusion can make full use of the salience of different features. The experimental results demonstrate that the proposed algorithm improves the accuracy of person re-identification compared with the state of the art. Yanbing Geng, Hai-Miao Hu, Bo Li 0006 |
ICIP | 2 |
| 2013 | Region-classification-based rate control for flicker suppression of I-frames in HEVCabstractIn High Efficiency Video Coding (HEVC), the coding efficiency of I-frames is lower than P-frames and B-frames, which will cause the flicker artifact, especially in low bitrates applications. We propose a region-classification-based rate control for Coding Tree Units (CTUs) in I-frames to improve the reconstructed quality of I-frames to suppress the flicker artifact. The CTUs in I-frame are classified into three regions according to their motion vectors and complexity. When the bit budget of one I-frame is used up, the target bitrates for the remaining CTUs will be adjusted according to the regions they belong to, and the pixel-based unified rate-quantization (URQ) model is then used to calculate the QPs. Experimental results demonstrate that the proposed scheme can efficiently suppress the flicker artifacts and improve both the subjective and objective video quality when compared with the original scheme in HM9.0. Yongfei Zhang, Hai-Miao Hu, Bo Li 0006 |
ICIP | 3 |
| 2013 | Naturalness Preserved Enhancement Algorithm for Non-Uniform Illumination ImagesabstractImage enhancement plays an important role in image processing and analysis. Among various enhancement algorithms, Retinex-based algorithms can efficiently enhance details and have been widely adopted. Since Retinex-based algorithms regard illumination removal as a default preference and fail to limit the range of reflectance, the naturalness of non-uniform illumination images cannot be effectively preserved. However, naturalness is essential for image enhancement to achieve pleasing perceptual quality. In order to preserve naturalness while enhancing details, we propose an enhancement algorithm for non-uniform illumination images. In general, this paper makes the following three major contributions. First, a lightness-order-error measure is proposed to access naturalness preservation objectively. Second, a bright-pass filter is proposed to decompose an image into reflectance and illumination, which, respectively, determine the details and the naturalness of the image. Third, we propose a bi-log transformation, which is utilized to map the illumination to make a balance between details and naturalness. Experimental results demonstrate that the proposed algorithm can not only enhance the details but also preserve the naturalness for non-uniform illumination images. Shuhang Wang, Hai-Miao Hu, Bo Li 0006 |
IEEE Trans. Image Process. | 3 |
| 2012 | Region-Based Rate Control for H.264/AVC for Low Bit-Rate ApplicationsabstractRate control plays an important role in video coding. However, in the conventional rate control algorithms, the number and position of macroblocks (MBs) inside one basic unit for rate control is inflexible and predetermined. The different characteristics of the MBs are not fully considered. Also, there is no overall optimization of the coding of basic units. This paper proposes a new region-based rate control scheme for H.264/advanced video coding to improve the coding efficiency. The inter-frame information is explored to objectively divide one frame into multiple regions based on their rate-distortion (R-D) behaviors. The MBs with similar characteristics are classified into the same region, and the entire region, instead of a single MB or a group of contiguous MBs, is treated as a basic unit for rate control. A linear rate-quantization stepsize model and a linear distortion-quantization stepsize model are proposed to accurately describe the R-D characteristics for the region-based basic units. Moreover, based on the above linear models, an overall optimization model is proposed to obtain suitable quantization parameters for the region-based basic units. Experimental results demonstrate that the proposed region-based rate control approach can achieve both better subjective and objective quality by performing the rate control adaptively with the content, compared to the conventional rate control approaches. Hai-Miao Hu, Bo Li 0006, Weiyao Lin, Wei Li 0209, Ming-Ting Sun |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2011 | A rate-control algorithm using inter-layer information for H.264/SVC for low-delay applications
Hai-Miao Hu, Bo Li 0006, Weiyao Lin, Ming-Ting Sun |
J. Vis. Commun. Image Represent. | 1 |
| 2011 | A region-based rate-control scheme using inter-layer information for H.264/SVC
Hai-Miao Hu, Weiyao Lin, Bo Li 0006, Ming-Ting Sun |
J. Vis. Commun. Image Represent. | 1 |
| 2009 | An error resilient video coding and transmission solution over error-prone channels
Hai-Miao Hu, Bo Li 0006 |
J. Vis. Commun. Image Represent. | 1 |