EDBT 2026 Demo / reviewers in the wild / expert
Huihui Song 0003
dblp:63/8017-3
· DBLP profile ↗
38ranked-venue papers
6as first author
18since 2021 · last 2025
0000-0002-7275-9871ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 19 · 3 first-author · 13 since 2021Artificial intelligence and machine learning · 17 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Learning Deep Frequency Degradation Prior for Remote Sensing Spatio-temporal FusionabstractExisting deep learning-based remote sensing spatiotemporal fusion (STF) relies on a data-driven paradigm without considering the degradation prior modeling from the coarseto fine-resolution images. This makes the learned model easy to overfit to the training dataset, resulting in poor domain generalization across different datasets. To this end, this paper presents a deep frequency degradation prior to STF, dubbed as DeepFDP. The DeepFDP is based on a statistical observation that the frequency feature distributions of the tokens from the coarse- and fine-resolution images in different datasets have a low intra-resolution variance and a high inter-resolution variance with compact well-separated clusters. Therefore, the DeepFDP first designs a frequency proximity module to learn the mapping function between the token frequency representations of the coarse- and fine-resolution images, which is to narrow their feature distribution difference. Since the statistical properties of the token frequency representations are independent of the land-cover classes, the DeepFDP can be learned on a training set of limited image pairs without extra-supervision signals, which has a favorable zero-shot generalization capability across different datasets. Then, to faithfully recover the land-surface details, a high-frequency feature modulation module is designed that uses the fine-resolution image as guidance to progressively learn the multi-scale residual features in a coarse-to-fine fashion, yielding the fused features with rich high-frequency details. Finally, the progressively-fused features at each stage are fed into a hybrid fusion module, yielding the fine-resolution image prediction. Extensive evaluations on LGC and CIA datasets demonstrate favorable performance of the DeepFDP over state-of-the-art methods. Especially, the DeepFDP also shows a good zero-shot generalization performance when training on LGC and testing on CIA. Yiting Bian, Shenglong Hu, Huihui Song 0003, Kaihua Zhang 0001 |
ICASSP | 3 |
| 2025 | Learning Joint Appearance and Shape Co-Representations for Co-Saliency DetectionabstractExisting leading Co-saliency Detection (CoD) framework aims to segment the co-salient objects by learning the consensus visual representation of the foreground objects. However, despite different categories, some distractors may have similar appearance to the co-salient objects, such as Apples vs. Bananas have similar color and textures. This makes it challenging to distinguish the distractors only through learning the co-salient object appearance representations. To address this issue, we propose a joint appearance and shape co-representation learner for CoD, dubbed as ASCoD. The ASCoD is composed of a Co-Appearance learning Module (CoAM) and a Co-Shape learning Module (CoSM). The CoAM first learns a co-salient object appearance embedding that encodes the global cross-image and spatial context information. Then, this embedding is set as a co-appearance prototype, which guides the model to enhance the features to highlight the co-salient object regions. Afterwards, we design the CoSM that is a cross-attention module, among which the key and the value encode the shape information from a set of salient tokens dynamically selected by a Co-Shape Prototype generation Module (CSPM). Finally, through jointly optimizing the cascaded CoAM and CoSM, the optimal appearance and shape co-representations are achieved, marrying the merits of both appearance and shape co-representations that are not only robust to co-salient objects appearance variations, but also can well discriminate the co-salient objects from the distractors with similar appearance. Extensive evaluations on three challenging benchmarks including CoCA, CoSOD3k and CoSal2015, demonstrate superiority of the ASCoD to a variety of state-of-the-art CoD methods. Guanting Guo, Shenglong Hu, Huihui Song 0003, Kaihua Zhang 0001 |
ICASSP | 3 |
| 2025 | Open-Vocabulary Saliency-Guided Progressive Refinement Network for Unsupervised Video Object SegmentationabstractExisting leading unsupervised video object segmentation (UVOS) paradigm often leverages a dual-stream architecture with motion and appearance branches, where only the motion cues from optical flow are used as a guide to locating the primary foreground objects. When suffering from challenging factors such as static scenes, fast camera shaking, severe motion blur, etc., the estimated optical flow is noisy with low quality, leading to erroneous primary foreground objects estimation. To address this issue, we propose an open-vocabulary saliency-guided progressive refinement network for UVOS, dubbed as OVSNet. It is observed that most of the primary foreground objects also demonstrate saliency characteristics in the appearance branch. Based on this, our OVSNet complements motion cues with saliency cues predicted by a series of foundation models equipped with favorable zero-shot generalization capabilities. Specifically, we first leverage the off-the-shelf contrastive vision-language pre-training (CLIP) and CLIPSeg to generate an OVS attention map as saliency cues. Then, the saliency cues together with motion cues prompt the segment anything model (SAM) to generate a location map. In the location process, we design two lightweight adapters to fine-tune SAM, which makes SAM well adapt to the downstream UVOS task. Finally, the location map generated by SAM is used to progressively guide object representation refinement in the appearance branch, ultimately achieving an accurate segmentation mask prediction. Extensive evaluations on DAVIS-16, FBMS, and YouTube-Objects demonstrate the favorable performance of our OVSNet over the state-of-the-art methods. Zhidong Han, Shenglong Hu, Huihui Song 0003, Kaihua Zhang 0001 |
ICASSP | 3 |
| 2025 | Group-wise Semantic-enhanced Interaction Network for Remote Sensing Spatio-Temporal FusionabstractRemote sensing spatio-temporal fusion (STF) aims at fusing temporally-dense coarse-resolution images and temporally-sparse fine-resolution images to reconstruct high spatio-temporal resolution images. Multi-band remote sensing images are often accepted as inputs for STF that have complementary characteristics for high-fidelity land surface reconstruction, yet the existing STF framework often treats different bands uniformly without considering the statistical correlation between different bands, resulting in unsatisfying results. To address this problem, this paper presents a group-wise semantic-enhanced interactive network for STF, dubbed as GSINet. Based on statistical observations, the feature correlation between the visible-light group and the invisible-light group is weak, while the intra-group correlation is strong. Therefore, the GSINet first separates the inputs into visible-light and invisible-light groups, which are fed into different branches with independent encoders for feature extraction and fusion. Afterwards, to address the issue of land cover changes between the prediction coarse- and reference fine-resolution images, a Semantic-Enhancement Fusion Module (SEFM) is designed to interact with the features from the same group with enhanced semantic information captured in an unsupervised learning manner. Then, the semantic-enhanced fused features from different bands are fed into an Interleaved Cross-attention Module (ICM) for further fusion. Finally, the output fusion features fully encode the intra- and inter-group information, which are fed into the decoder, reconstructing the spatio-temporal high-resolution images. Extensive experiments on CIA and LGC benchmark datasets demonstrate that the GSINet outperforms a variety of state-of-the-art methods in terms of multiple metrics. Baoluo Zhu, Shenglong Hu, Huihui Song 0003, Kaihua Zhang 0001 |
ICASSP | 3 |
| 2025 | Joint Feature Learning and Mixing via State Space Model for Remote Sensing Change DetectionabstractRemote sensing change detection (RSCD) aims to identify change areas between bi-temporal images of the same location captured at different points in time. The existing siamese framework adopts a shared-weight strategy to process each bitemporal image independently. Due to the lack of inter-image information interaction, this strategy exhibits limited capability to target change perception and discrimination when faced with small targets or ambiguous changes such as low-covering change areas and irregular morphology in real-world complex scenes. To this end, we propose a one-stream framework using the State Space (SS) Model Mamba to jointly perform feature learning and mixing for RSCD, dubbed as SSCD. Specifically, by leveraging the long-range modeling capability and linear computational complexity of the SS model Mamba, the SSCD leverages a unified approach to feature extraction and information integration through simultaneously processing the bi-temporal images. This enables the model to intensively mutually guide to extract discriminative change features. In addition, to recover more spatial details, we design a texture enhancement module that makes full use of the selective scan modeling capability of the Mamba to enhance the texture features in different directions. Without bells and whistles, our SSCD achieves the state-of-the-art performance on three benchmark datasets including SYSU-CD, LEVIR-CD, and LEVIR+-CD. Shenglong Hu, Huihui Song 0003, Kaihua Zhang 0001 |
ICME | 3 |
| 2024 | Generalizable Fourier Augmentation for Unsupervised Video Object SegmentationabstractThe performance of existing unsupervised video object segmentation methods typically suffers from severe performance degradation on test videos when tested in out-of-distribution scenarios. The primary reason is that the test data in real- world may not follow the independent and identically distribution (i.i.d.) assumption, leading to domain shift. In this paper, we propose a generalizable fourier augmentation method during training to improve the generalization ability of the model. To achieve this, we perform Fast Fourier Transform (FFT) over the intermediate spatial domain features in each layer to yield corresponding frequency representations, including amplitude components (encoding scene-aware styles such as texture, color, contrast of the scene) and phase components (encoding rich semantics). We produce a variety of style features via Gaussian sampling to augment the training data, thereby improving the generalization capability of the model. To further improve the cross-domain generalization performance of the model, we design a phase feature update strategy via exponential moving average using phase features from past frames in an online update manner, which could help the model to learn cross-domain-invariant features. Extensive experiments show that our proposed method achieves the state-of-the-art performance on popular benchmarks. Huihui Song 0003, Tiankang Su, Yuhui Zheng, Kaihua Zhang 0001, Bo Liu 0005, Dong Liu 0002 |
AAAI | 1 |
| 2024 | Segment Anything Model Guided Semantic Knowledge Learning For Remote Sensing Change DetectionabstractExisting deep learning based remote sensing change detection (RSCD) methods only rely on binary ground-truth to guide the network learning while neglecting the useful semantic guidance. As a result, the network can be readily misled by irrelevant category changes, leading to degraded performance and slow convergence of the model. To this end, we propose a novel segment anything model (SAM) guided framework, termed as SAM-CD, which mines the rich semantic knowledge from the SAM for RSCD. Specifically, we first employ a transformer encoder to extract multi-scale global features from the bi-temporal images. Meanwhile, we obtain semantic prior masks from the bi-temporal images by providing the SAM with category-relevant text prompts. Then, using the semantic prior masks as constraints, we design a masked attention module (MAM) that generates local features related to the interested categories. Finally, the local and global features are fused and fed into a multi-layer perception (MLP) decoder to obtain the change map. The whole network is trained in an end-to-end manner that can readily encode the rich semantic knowledge of the changed targets to predict an accurate change map. Extensive experiments demonstrate that the proposed SAM-CD achieves state-of-the-art performance on a variety of benchmark datasets. Zixuan Sun, Huihui Song 0003, Kaihua Zhang 0001, Gang Dong, Lingyan Liang, Yaqian Zhao |
ICASSP | 2 |
| 2024 | Language-Guided Semantic Alignment for Co-saliency DetectionabstractPrevious pure vision paradigm for co-saliency detection (COD) predominantly employs supervised training. The supervisory signals often consist of binary masks or a combination of masks and category labels. However, constrained by limited training samples, these models often suffer from overfitting issue, struggling to generalize to unseen samples. To this end, this paper presents the constrative language-image pretraining-COD (CLIP-COD), a novel language-guided semantic alignment paradigm for COD. The primary objective is to leverage CLIP for aligning concepts between language and images, where the alignment can effectively leverage the powerful language understanding capability of CLIP and transfer its knowledge to image domain, thereby enhancing the model’s zero-shot generalization ability for COD. Firstly, we propose a semantic alignment branch (SAB) that can learn rich knowledge for comprehending images globally. Meanwhile, the SAB can narrow the gap in high-dimensional feature space between the language and image features, transferring the powerful semantic knowledge from CLIP to our model. Subsequently, we devise an intra-group multi-fusion module (IMM) to capture features that integrate group knowledge as dense prompts, providing spatial localization information for subsequent fine segmentation. Finally, we input sparse language prompts and dense mask cues into the pre-trained SAM decoder to obtain the final COD results. Additionally, we further design a transfer optimization adaptor, which can reduce the model training scale, saving computing resource and cost greatly. Extensive experiments on three benchmark datasets, including CoSal2015, CoCA, and CoSOD3k, demonstrate the superior performance of our CLIP-COD to a variety of state-of-the-art methods. Chuang Ding, Huihui Song 0003, Kaihua Zhang 0001 |
ICME | 3 |
| 2023 | Co-Salient Object Detection with Uncertainty-Aware Group Exchange-MaskingabstractThe traditional definition of co-salient object detection (CoSOD) task is to segment the common salient objects in a group of relevant images. Existing CoSOD models by-default adopt the group consensus assumption. This brings about model robustness defect under the condition of irrelevant images in the testing image group, which hinders the use of CoSOD models in real-world applications. To address this issue, this paper presents a group exchange-masking (GEM) strategy for robust CoSOD model learning. With two group of image containing different types of salient object as input, the GEM first selects a set of images from each group by the proposed learning based strategy, then these images are exchanged. The proposed feature extraction module considers both the uncertainty caused by the irrelevant images and group consensus in the remaining relevant images. We design a latent variable generator branch which is made of conditional variational autoencoder to generate uncertainly-based global stochastic features. A CoSOD transformer branch is devised to capture the correlation-based local features that contain the group consistency information. At last, the output of two branches are concatenated and fed into a transformer-based decoder, producing robust co-saliency prediction. Extensive evaluations on co-saliency detection with and without irrelevant images demonstrate the superiority of our method over a variety of state-of-the-art methods. Huihui Song 0003, Bo Liu 0005, Kaihua Zhang 0001, Dong Liu 0002 |
CVPR | 2 |
| 2023 | Unsupervised Video Object Segmentation with Online Adversarial Self-TuningabstractThe existing unsupervised video object segmentation methods depend heavily on the segmentation model trained offline on a labeled training video set, and cannot well generalize to the test videos from a different domain with possible distribution shifts. We propose to perform online fine-tuning on the pre-trained segmentation model to adapt to any ad-hoc videos at the test time. To achieve this, we design an offline semi-supervised adversarial training process, which leverages the unlabeled video frames to improve the model generalizability while aligning the features of the labeled video frames with the features of the unlabeled video frames. With the trained segmentation model, we further conduct an online self-supervised adversarial finetuning, in which a teacher model and a student model are first initialized with the pre-trained segmentation model weights, and the pseudo label produced by the teacher model is used to supervise the student model in an adversarial learning framework. Through online finetuning, the student model is progressively updated according to the emerging patterns in each test video, which significantly reduces the test-time domain gap. We integrate our offline training and online fine-tuning in a unified framework for unsupervised video object segmentation and dub our method Online Adversarial Self-Tuning (OAST). The experiments show that our method outperforms the state-of-the-arts with significant gains on the popular video object segmentation datasets. Tiankang Su, Huihui Song 0003, Dong Liu 0002, Bo Liu 0005, Qingshan Liu 0001 |
ICCV | 2 |
| 2023 | Intellectual property protection for deep semantic segmentation models
Hongjia Ruan, Huihui Song 0003, Bo Liu 0005, Yong Cheng 0002, Qingshan Liu 0001 |
Frontiers Comput. Sci. | 2 |
| 2023 | Gradient-Guided Temporal Cross-Attention Transformer for High-Performance Remote Sensing Change DetectionabstractGiven a group of bi-temporal remote sensing images acquired in the same geographical area, the task of Change Detection (CD) aims to detect and segment the change regions therein. Existing leading CD methods typically use the self-attention (SA) mechanism to directly fuse the concatenated features of the bi-temporal images. Despite demonstrated success, the SA focuses primarily on modeling spatial-wise relationships, rather than channel-wise relationships. Meanwhile, since the temporal direction is along the channel direction, this makes the SA difficult to model the temporal-wise relationship between the bi-temporal features, making it fail to learn the feature correspondence from the significantly-changed regions. To this end, this letter presents a gradient-guided temporal cross-attention (Grad-TCA) mechanism for CD. First, we design a temporal cross-attention module (TCAM) that mixes the cross- and self-attention to model both temporal- and spatial-wise interactions, which fully mines the complementary cues between bi-temporal features to learn a strong feature presentation. Afterwards, to further highlight the salient features between the change regions, we design a gradient-guided module (GGM) to enhance the difference of the learned bi-temporal features through feedback gradient information. Both the TCAM and the GGM construct our Grad-TCA module, which is seamlessly integrated into a Transformer framework for end-to-end learning. Finally, to reduce the computation overhead, we design a simple change discrimination module (CDM) that outputs a score to directly filter out the unchanged features from the GGM with no need of passing the features through the decoder. Comprehensive evaluations on the two widely used benchmark datasets including LEVIR-CD and WHU-CD demonstrate our model outperforms a variety of state-of-the-art methods. Huihui Song 0003, Kaihua Zhang 0001 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2022 | Video Super-Resolution with Frame-Wise Dynamic Fusion and Self-Calibrated Deformable Alignment
Huihui Song 0003, Yutong Jin |
Neural Process. Lett. | 2 |
| 2022 | Learning interlaced sparse Sinkhorn matching network for video super-resolution
Huihui Song 0003, Yutong Jin, Yong Cheng 0002, Bo Liu 0005, Dong Liu 0002, Qingshan Liu 0001 |
Pattern Recognit. | 1 |
| 2021 | Conditional generative adversarial network with densely-connected residual learning for single image super-resolution
Jiaojiao Qiao, Huihui Song 0003, Kaihua Zhang 0001 |
Multim. Tools Appl. | 2 |
| 2021 | Video saliency prediction using enhanced spatiotemporal alignment network
Huihui Song 0003, Kaihua Zhang 0001, Bo Liu 0005, Qingshan Liu 0001 |
Pattern Recognit. | 2 |
| 2021 | Feature Alignment and Aggregation Siamese Networks for Fast Visual TrackingabstractSiamese networks have been successfully introduced into visual tracking, which match the best candidate and a target template via a couple of networks with shared parameters. However, most Siamese network-based trackers (SNTs) are tailored to best match the canonical posture of the template and the search-region images, resulting in inferior performance when the target objects have large-scale pose variations. Besides, SNTs fail to discriminate distractors well because they only leverage high-level semantic features as target representations that cannot well tell from different targets of the same category. To address these issues, this paper presents an efficient and effective SNT that is based on feature alignment and aggregation networks. Specifically, we first design an effective feature alignment network module to calibrate the search-region image. This module results in a more reliable matching response that is robust to severe target pose variations. Then, we develop an effective shallow-level and high-level feature aggregation network module to complement the feature characteristics, making the learned feature representation not only well differentiate the target from distractors, but also robust to target appearance variations. Afterwards, we employ a channel-attention mechanism to further strengthen the discriminative capability of the aggregated feature representation. Finally, both the alignment and the aggregation modules are seamlessly integrated into the Siamese networks for robust tracking. Meanwhile, we offline learn the network parameters end-to-end without time-consuming fine-tuning. Extensive evaluations on a variety of benchmarks including VOT-2017, OTB-100, UAV123 and GOT-10k demonstrate favorable performance of our tracker against state-of-the-art ones with a speed of 60 fps. Jiaqing Fan, Huihui Song 0003, Kaihua Zhang 0001, Qingshan Liu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | Multi-Stage Feature Fusion Network for Video Super-ResolutionabstractVideo super-resolution (VSR) is to restore a photo-realistic high-resolution (HR) frame from both its corresponding low-resolution (LR) frame (reference frame) and multiple neighboring frames (supporting frames). An important step in VSR is to fuse the feature of the reference frame with the features of the supporting frames. The major issue with existing VSR methods is that the fusion is conducted in a one-stage manner, and the fused feature may deviate greatly from the visual information in the original LR reference frame. In this paper, we propose an end-to-end Multi-Stage Feature Fusion Network that fuses the temporally aligned features of the supporting frames and the spatial feature of the original reference frame at different stages of a feed-forward neural network architecture. In our network, the Temporal Alignment Branch is designed as an inter-frame temporal alignment module used to mitigate the misalignment between the supporting frames and the reference frame. Specifically, we apply the multi-scale dilated deformable convolution as the basic operation to generate temporally aligned features of the supporting frames. Afterwards, the Modulative Feature Fusion Branch, the other branch of our network accepts the temporally aligned feature map as a conditional input and modulates the feature of the reference frame at different stages of the branch backbone. This enables the feature of the reference frame to be referenced at each stage of the feature fusion process, leading to an enhanced feature from LR to HR. Experimental results on several benchmark datasets demonstrate that our proposed method can achieve state-of-the-art performance on VSR task. Huihui Song 0003, Dong Liu 0002, Bo Liu 0005, Qingshan Liu 0001, Dimitris N. Metaxas |
IEEE Trans. Image Process. | 1 |
| 2020 | Top-Down Fusing Multi-level Contextual Features for Salient Object Detection
Mingyuan Pan, Huihui Song 0003, Junxia Li, Kaihua Zhang 0001, Qingshan Liu 0001 |
PRCV (3) | 2 |
| 2020 | Learning lightweight Multi-Scale Feedback Residual network for single image super-resolution
Huihui Song 0003, Kaihua Zhang 0001, Qingshan Liu 0001, Jia Liu 0034 |
Comput. Vis. Image Underst. | 2 |
| 2020 | Real-time manifold regularized context-aware correlation tracking
Jiaqing Fan, Huihui Song 0003, Kaihua Zhang 0001, Qingshan Liu 0001, Wei Lian |
Frontiers Comput. Sci. | 2 |
| 2020 | Recurrent reverse attention guided residual learning for saliency object detection
Tengpeng Li, Huihui Song 0003, Kaihua Zhang 0001, Qingshan Liu 0001 |
Neurocomputing | 2 |
| 2020 | Single image super-resolution with enhanced Laplacian pyramid network via conditional generative adversarial learning
Huihui Song 0003, Kaihua Zhang 0001, Jiaojiao Qiao, Qingshan Liu 0001 |
Neurocomputing | 2 |
| 2020 | Hierarchical attentive Siamese network for real-time visual tracking
Huihui Song 0003, Kaihua Zhang 0001, Qingshan Liu 0001 |
Neural Comput. Appl. | 2 |
| 2020 | Learning residual refinement network with semantic context representation for real-time saliency object detection
Tengpeng Li, Huihui Song 0003, Kaihua Zhang 0001, Qingshan Liu 0001 |
Pattern Recognit. | 2 |
| 2020 | Dynamically Spatiotemporal Regularized Correlation TrackingabstractRecently, due to the high performance, spatially regularized strategy has been widely applied to addressing the issue of boundary effects existed in correlation filter (CF)-based visual tracking. Specifically, it introduces a spatially regularized term to penalize the coefficients of the CFs to be learned depending on their spatial locations. However, the regularization weights are often formed as a fixed Gaussian function, and hence may cause the learned model degenerate due to the inflexible constraints on the ever-changing CFs to be learned over time during tracking. To address this issue, in this paper, we develop a dynamically spatiotemporal regularization model to constrain the CFs to be learned with the ever-changing regularization weights learned from two consecutive frames. The proposed method jointly learns the CFs along with the dynamically spatiotemporal constraint term, which can be efficiently solved in the Fourier domain by the alternative direction method. Extensive evaluations on the popular data sets OTB-100 and VOT-2016 demonstrate that the proposed tracker performs favorably against the baseline tracker and several recently proposed state-of-the-art methods. Yuhui Zheng, Huihui Song 0003, Kaihua Zhang 0001, Jiaqing Fan, Xinyan Liu 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2019 | Image super-resolution using conditional generative adversarial networkabstractRecently, extensive studies on a generative adversarial network (GAN) have made great progress in single image super‐resolution (SISR). However, there still exists a significant difference between the reconstructed high‐frequency and the real high‐frequency details. To address this issue, this study presents an SISR approach based on conditional GAN (SRCGAN). SRCGAN includes a generator network that generates super‐resolution (SR) images and a discriminator network that is trained to distinguish the SR images from ground‐truth high‐resolution (HR) ones. Specifically, the discriminator network uses the ground‐truth HR image as a conditional variable, which guides the network to distinguish the real images from the SR images, facilitating training a more stable generator model than GAN without this guidance. Furthermore, a residual‐learning module is introduced into the generator network to solve the issue of detail information loss in SR images. Finally, the network is trained in an end‐to‐end manner by optimizing a perceptual loss function. Extensive evaluations on four benchmark datasets including Set5, Set14, BSD100, and Urban100 demonstrate the superiority of the proposed SRCGAN over state‐of‐the‐art methods in terms of PSNR, SSIM, and visual effect. Jiaojiao Qiao, Huihui Song 0003, Kaihua Zhang 0001, Qingshan Liu 0001 |
IET Image Process. | 2 |
| 2019 | Low-rank weighted co-saliency detection via efficient manifold ranking
Tengpeng Li, Huihui Song 0003, Kaihua Zhang 0001, Qingshan Liu 0001, Wei Lian |
Multim. Tools Appl. | 2 |
| 2018 | Visual tracking using spatio-temporally nonlocally regularized correlation filter
Kaihua Zhang 0001, Huihui Song 0003, Qingshan Liu 0001, Wei Lian |
Pattern Recognit. | 3 |
| 2018 | Visual Tracking via Nonlocal Similarity LearningabstractEither global (e.g., intensity histograms and coefficients of sparse representation) or local (e.g., scale-invariant feature transform and histogram of oriented gradient) feature representations have been widely exploited for visual tracking. However, most of these representations describe a target appearance with a fixed spatial grid layout without considering the interactions between different grids, and hence may adversely affect their performance when the target appearance suffers from large-scale pose variations. In this paper, we learn a similarity function that considers the interactions of features in the grids not only from the same spatial positions, but also from different positions, thereby taking charge of the nonlocal information of the target appearances to effectively handle the significant appearance variations. Specifically, we explore the polynomial kernel feature map to characterize the nonlocal similarity information of all pairs of grids among the target and its background samples, and combine these feature maps as the target representations. Moveover, we learn a linear logistic regression classifier with online update to separate the target from its local background, and integrate this classifier into a particle filtering tracking framework. Extensive experimental results on the CVPR2013 tracking benchmark demonstrate the proposed approach performs favorably against some representative tracking algorithms. Qingshan Liu 0001, Jiaqing Fan, Huihui Song 0003, Kaihua Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2015 | A Variational Approach to Simultaneous Image Segmentation and Bias CorrectionabstractThis paper presents a novel variational approach for simultaneous estimation of bias field and segmentation of images with intensity inhomogeneity. We model intensity of inhomogeneous objects to be Gaussian distributed with different means and variances, and then introduce a sliding window to map the original image intensity onto another domain, where the intensity distribution of each object is still Gaussian but can be better separated. The means of the Gaussian distributions in the transformed domain can be adaptively estimated by multiplying the bias field with a piecewise constant signal within the sliding window. A maximum likelihood energy functional is then defined on each local region, which combines the bias field, the membership function of the object region, and the constant approximating the true signal from its corresponding object. The energy functional is then extended to the whole image domain by the Bayesian learning approach. An efficient iterative algorithm is proposed for energy minimization, via which the image segmentation and bias field correction are simultaneously achieved. Furthermore, the smoothness of the obtained optimal bias field is ensured by the normalized convolutions without extra cost. Experiments on real images demonstrated the superiority of the proposed algorithm to other state-of-the-art representative methods. Kaihua Zhang 0001, Qingshan Liu 0001, Huihui Song 0003, Xuelong Li 0001 |
IEEE Trans. Cybern. | 3 |
| 2015 | Improving the Spatial Resolution of Landsat TM/ETM+ Through Fusion With SPOT5 Images via Learning-Based Super-ResolutionabstractTo take advantage of the wide swath width of Landsat Thematic Mapper (TM)/Enhanced Thematic Mapper Plus (ETM+) images and the high spatial resolution of Système Pour l'Observation de la Terre 5 (SPOT5) images, we present a learning-based super-resolution method to fuse these two data types. The fused images are expected to be characterized by the swath width of TM/ETM+ images and the spatial resolution of SPOT5 images. To this end, we first model the imaging process from a SPOT image to a TM/ETM+ image at their corresponding bands, by building an image degradation model via blurring and downsampling operations. With this degradation model, we can generate a simulated Landsat image from each SPOT5 image, thereby avoiding the requirement for geometric coregistration for the two input images. Then, band by band, image fusion can be implemented in two stages: 1) learning a dictionary pair representing the high- and low-resolution details from the given SPOT5 and the simulated TM/ETM+ images; 2) super-resolving the input Landsat images based on the dictionary pair and a sparse coding algorithm. It is noteworthy that the proposed method can also deal with the conventional spatial and spectral fusion of TM/ETM+ and SPOT5 images by using the learned dictionary pairs. To examine the performance of the proposed method of fusing the swath width of TM/ETM+ and the spatial resolution of SPOT5, we illustrate the fusion results on the actual TM images and compare with several classic pansharpening methods by assuming that the corresponding SPOT5 panchromatic image exists. Furthermore, we implement the classification experiments on both actual images and fusion results to demonstrate the benefits of the proposed method for further classification applications. Huihui Song 0003, Bo Huang 0001, Qingshan Liu 0001, Kaihua Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2014 | Shadow Detection and Reconstruction in High-Resolution Satellite Images via Morphological Filtering and Example-Based LearningabstractThe shadows in high-resolution satellite images are usually caused by the constraints of imaging conditions and the existence of high-rise objects, and this is particularly so in urban areas. To alleviate the shadow effects in high-resolution images for their further applications, this paper proposes a novel shadow detection algorithm based on the morphological filtering and a novel shadow reconstruction algorithm based on the example learning method. In the shadow detection stage, an initial shadow mask is generated by the thresholding method, and then, the noise and wrong shadow regions are removed by the morphological filtering method. The shadow reconstruction stage consists of two phases: the example-based learning phase and the inference phase. During the example-based learning phase, the shadow and the corresponding nonshadow pixels are first manually sampled from the study scene, and then, these samples form a shadow library and a nonshadow library, which are correlated by a Markov random field (MRF). During the inference phase, the underlying land-cover pixels are reconstructed from the corresponding shadow pixels by adopting the Bayesian belief propagation algorithm to solve the MRF. Experimental results on QuickBird and WorldView-2 satellite images have demonstrated that the proposed shadow detection algorithm can generate accurate and continuous shadow masks and also that the estimated nonshadow regions from the proposed shadow reconstruction algorithm are highly compatible with their surrounding nonshadow regions. Finally, we examine the effects of the reconstructed image on the application of classification by comparing the classification maps of images before and after shadow reconstruction. Huihui Song 0003, Bo Huang 0001, Kaihua Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2013 | Real-time visual tracking via online weighted multiple instance learning
Kaihua Zhang 0001, Huihui Song 0003 |
Pattern Recognit. | 2 |
| 2013 | Reinitialization-Free Level Set Evolution via Reaction DiffusionabstractThis paper presents a novel reaction-diffusion (RD) method for implicit active contours that is completely free of the costly reinitialization procedure in level set evolution (LSE). A diffusion term is introduced into LSE, resulting in an RD-LSE equation, from which a piecewise constant solution can be derived. In order to obtain a stable numerical solution from the RD-based LSE, we propose a two-step splitting method to iteratively solve the RD-LSE equation, where we first iterate the LSE equation, then solve the diffusion equation. The second step regularizes the level set function obtained in the first step to ensure stability, and thus the complex and costly reinitialization procedure is completely eliminated from LSE. By successfully applying diffusion to LSE, the RD-LSE model is stable by means of the simple finite difference method, which is very easy to implement. The proposed RD method can be generalized to solve the LSE for both variational level set method and partial differential equation-based level set method. The RD-LSE method shows very good performance on boundary antileakage. The extensive and promising experimental results on synthetic and real images validate the effectiveness of the proposed RD-LSE approach. Kaihua Zhang 0001, Lei Zhang 0006, Huihui Song 0003, David Zhang 0001 |
IEEE Trans. Image Process. | 3 |
| 2010 | AN adaptive L1-L2 hybrid error model to super-resolutionabstractA hybrid error model with L1and L2norm minimization criteria is proposed in this paper for image/video super-resolution. A membership function is defined to adaptively control the tradeoff between the L1and L2norm terms. Therefore, the proposed hybrid model can have the advantages of both L1norm minimization (i.e. edge preservation) and L2norm minimization (i.e. smoothing noise). In addition, an effective convergence criterion is proposed, which is able to terminate the iterative L1and L2norm minimization process efficiently. Experimental results on images corrupted with various types of noises demonstrate the robustness of the proposed algorithm and its superiority to representative algorithms. Huihui Song 0003, Lei Zhang 0006, Peikang Wang, Kaihua Zhang 0001, Xin Li 0005 |
ICIP | 1 |
| 2010 | Active contours with selective local or global segmentation: A new formulation and level set method
Kaihua Zhang 0001, Lei Zhang 0006, Huihui Song 0003, Wengang Zhou 0001 |
Image Vis. Comput. | 3 |
| 2010 | Active contours driven by local image fitting energy
Kaihua Zhang 0001, Huihui Song 0003, Lei Zhang 0006 |
Pattern Recognit. | 2 |