VLDB 2026 Research / reviewers in the wild / expert
Kaihua Zhang 0001
dblp:84/8017-1
· DBLP profile ↗
84ranked-venue papers
23as first author
40since 2021 · last 2026
0000-0002-1613-3401ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 53 · 15 first-author · 30 since 2021Artificial intelligence and machine learning · 38 · 15 first-author · 15 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Robust Exemplar Prompt Learning via Bi-directional Visual-Semantic Alignment for Multi-Object TrackingabstractRecent multi-object tracking (MOT) approaches increasingly leverage pre-trained CLIP models to boost cross-domain generalization. A common strategy uses a predefined TrackBook—a closed-set of visual concepts—as textual prompts to guide learning of domain-invariant representations. However, these fixed prompts lack adaptive context, causing limited generalization. To address this limitation, this paper introduces a robust Exemplar Prompt Learning (EPL) framework via Bi-directional Visual-Semantic Alignment (BiVSA), termed EPL-MOT, which augments textual prompts with instance-aware contextual information derived during tracking. Specifically, an EPL module is designed to dynamically enrich textual prompts with contextual cues, enabling instance-specific adaptation without inducing category shift. Furthermore, a BiVSA module is proposed to deepen cross-modal interaction by incorporating bidirectional learnable prompts into both textual and visual branches. This facilitates progressive integration of global semantic features with local visual structures, resulting in a more effectively aligned visual-semantic space. Finally, to enhance robustness against distractors, a Category-guided Detection Query Generator (CDQG) is constructed, which incorporates base-class textual information to suppress irrelevant targets. Comprehensive evaluations on MOT17 and MOT20 demonstrate that the proposed EPL-MOT achieves competitive performance across both in-domain and cross-domain settings. Lingyan Liang, Gang Dong, Dongchao Wen, Kaihua Zhang 0001 |
ICMR | 5 |
| 2026 | Partitioned observation network for camouflaged object detection
Jinxia Zhang, Yin Yuan, Xuwen Zhu, Kaihua Zhang 0001 |
Pattern Recognit. | 5 |
| 2025 | WeatherGen: A Unified Diverse Weather Generator for LiDAR Point Clouds via Spider Mamba Diffusionabstract3D scene perception demands a large amount of adverse-weather LiDAR data, yet the cost of LiDAR data collection presents a significant scaling-up challenge. To this end, a series of LiDAR simulators have been proposed. Yet, they can only simulate a single adverse weather with a single physical model, and the fidelity of the generated data is quite limited. This paper presents WeatherGen, the first unified diverse-weather LiDAR data diffusion generation framework, significantly improving fidelity. Specifically, we first design a map-based data producer, which can provide a vast amount of high-quality diverse-weather data for training purposes. Then, we utilize the diffusion-denoising paradigm to construct a diffusion model. Among them, we propose a spider mamba generator to restore the disturbed diverse weather data gradually. The spider mamba models the feature interactions by scanning the Li-Dar beam circle or central ray, excellently maintaining the physical structure of the LiDAR data. Subsequently, following the generator to transfer real-world knowledge, we design a latent feature aligner. Afterward, we devise a contrastive learning-based controller, which equips weather control signals with compact semantic knowledge through language supervision, guiding the diffusion model to generate more discriminative data. Extensive evaluations demonstrate the high generation quality of WeatherGen. Through WeatherGen, we construct the mini-weather dataset, promoting the performance of the downstream task under adverse weather conditions. Code is available: https://github.com/wuyang98/weathergen Yun Zhu 0011, Kaihua Zhang 0001, Jianjun Qian, Jin Xie 0001, Jian Yang 0003 |
CVPR | 3 |
| 2025 | Learning Deep Frequency Degradation Prior for Remote Sensing Spatio-temporal FusionabstractExisting deep learning-based remote sensing spatiotemporal fusion (STF) relies on a data-driven paradigm without considering the degradation prior modeling from the coarseto fine-resolution images. This makes the learned model easy to overfit to the training dataset, resulting in poor domain generalization across different datasets. To this end, this paper presents a deep frequency degradation prior to STF, dubbed as DeepFDP. The DeepFDP is based on a statistical observation that the frequency feature distributions of the tokens from the coarse- and fine-resolution images in different datasets have a low intra-resolution variance and a high inter-resolution variance with compact well-separated clusters. Therefore, the DeepFDP first designs a frequency proximity module to learn the mapping function between the token frequency representations of the coarse- and fine-resolution images, which is to narrow their feature distribution difference. Since the statistical properties of the token frequency representations are independent of the land-cover classes, the DeepFDP can be learned on a training set of limited image pairs without extra-supervision signals, which has a favorable zero-shot generalization capability across different datasets. Then, to faithfully recover the land-surface details, a high-frequency feature modulation module is designed that uses the fine-resolution image as guidance to progressively learn the multi-scale residual features in a coarse-to-fine fashion, yielding the fused features with rich high-frequency details. Finally, the progressively-fused features at each stage are fed into a hybrid fusion module, yielding the fine-resolution image prediction. Extensive evaluations on LGC and CIA datasets demonstrate favorable performance of the DeepFDP over state-of-the-art methods. Especially, the DeepFDP also shows a good zero-shot generalization performance when training on LGC and testing on CIA. Yiting Bian, Shenglong Hu, Huihui Song 0003, Kaihua Zhang 0001 |
ICASSP | 4 |
| 2025 | Continuously Learning Video-level Object Tokens for Robust UAV trackingabstractDue to the dynamic changes in flight motion and viewpoint, the objects in unmanned aerial vehicle (UAV) tracking scenarios often suffer from drastic appearance variations. Existing UAV trackers often leverage a frame-level matching mechanism, which measures the appearance similarity between the object template and the search frame. The drastic object appearance variations degrade the learned model, leading to drift issue. To this end, this paper presents a video-level UAV tracking framework that focuses on Continuously Learning (CL) effective and efficient spatio-temporal object tokens for robust tracking, dubbed as CLTrack. Specifically, the CLTrack first learns a series of spatio-temporal object tokens via a dynamic filtering module (DFM), which encodes more consensus object appearance information from each frame. Afterwards, a spatio-temporal enhancement module (STEM) is designed via cascading a temporal and a spatial attention to fully interact with the selected tokens with stable long-range spatio-temporal context information of the tracked object. Finally, to ensure the learned model encodes the rich context information without catastrophic forgetting, a video-level tracking loss is designed to supervise feature learning from the whole video frames. Extensive experiments on three UAV benchmarks including UAV123, DTB70 and VisDrone2018 demonstrate that the proposed CLTrack achieves state-of-the-art performance. Shenglong Hu, Gang Dong, Lingyan Liang, Dongchao Wen, Kaihua Zhang 0001 |
ICASSP | 6 |
| 2025 | Easy-to-hard Instance-level Feature Fusion for Co-saliency DetectionabstractExisting leading deep learning-based Co-saliency Detection (CoD) methods often learn the consensus features from the input image group without considering the complexity of each image. Despite the demonstrated success, the input images may contain hard samples with high complexity, e.g., those containing distractors that have similar appearance but different semantics to the co-salient objects. This is prone to mislead the learned model to treat these distractors as co-salient objects, leading to classification ambiguity. To address this issue, this paper presents an easy-to-hard instance-level feature Fusion framework for CoD, termed E2HCoD. The E2HCoD exploits the instance-level co-salient object consensus cues from the easy samples as reliable guidance to accurately fuse the co-salient object features in the hard samples. First, we design a Feature Filtering Module (FFM) that evaluates image complexity by integrating entropy, variance, texture, and edge density cues, allowing the model to select the easy samples with relatively easy backgrounds. Then, we develop an Easy-instance Embedding Branch (EEB), which accurately segments the co-salient object masks from the easy samples as the instance-level guidance to learn the accurate co-salient object consensus cues. Then, with the consensus knowledge from the easy samples as guidance, we construct an Easy-instance guided Fusion Branch (EFB), which fully interacts with the consensus features from the hard samples via a cross-attention mechanism, yielding the refined features that highlight the co-salient objects while suppressing the distractors. Finally, the refined features are fed into the decoder, generating a high-quality CoD prediction. Extensive experiments demonstrate that the proposed E2HCoD achieves state-of-the-art performance on CoSal2015, CoCA, and CoSOD3k. Chuang Ding, Zhidong Han, Gang Dong, Lingyan Liang, Dongchao Wen, Kaihua Zhang 0001 |
ICASSP | 6 |
| 2025 | Learning Joint Appearance and Shape Co-Representations for Co-Saliency DetectionabstractExisting leading Co-saliency Detection (CoD) framework aims to segment the co-salient objects by learning the consensus visual representation of the foreground objects. However, despite different categories, some distractors may have similar appearance to the co-salient objects, such as Apples vs. Bananas have similar color and textures. This makes it challenging to distinguish the distractors only through learning the co-salient object appearance representations. To address this issue, we propose a joint appearance and shape co-representation learner for CoD, dubbed as ASCoD. The ASCoD is composed of a Co-Appearance learning Module (CoAM) and a Co-Shape learning Module (CoSM). The CoAM first learns a co-salient object appearance embedding that encodes the global cross-image and spatial context information. Then, this embedding is set as a co-appearance prototype, which guides the model to enhance the features to highlight the co-salient object regions. Afterwards, we design the CoSM that is a cross-attention module, among which the key and the value encode the shape information from a set of salient tokens dynamically selected by a Co-Shape Prototype generation Module (CSPM). Finally, through jointly optimizing the cascaded CoAM and CoSM, the optimal appearance and shape co-representations are achieved, marrying the merits of both appearance and shape co-representations that are not only robust to co-salient objects appearance variations, but also can well discriminate the co-salient objects from the distractors with similar appearance. Extensive evaluations on three challenging benchmarks including CoCA, CoSOD3k and CoSal2015, demonstrate superiority of the ASCoD to a variety of state-of-the-art CoD methods. Guanting Guo, Shenglong Hu, Huihui Song 0003, Kaihua Zhang 0001 |
ICASSP | 4 |
| 2025 | Open-Vocabulary Saliency-Guided Progressive Refinement Network for Unsupervised Video Object SegmentationabstractExisting leading unsupervised video object segmentation (UVOS) paradigm often leverages a dual-stream architecture with motion and appearance branches, where only the motion cues from optical flow are used as a guide to locating the primary foreground objects. When suffering from challenging factors such as static scenes, fast camera shaking, severe motion blur, etc., the estimated optical flow is noisy with low quality, leading to erroneous primary foreground objects estimation. To address this issue, we propose an open-vocabulary saliency-guided progressive refinement network for UVOS, dubbed as OVSNet. It is observed that most of the primary foreground objects also demonstrate saliency characteristics in the appearance branch. Based on this, our OVSNet complements motion cues with saliency cues predicted by a series of foundation models equipped with favorable zero-shot generalization capabilities. Specifically, we first leverage the off-the-shelf contrastive vision-language pre-training (CLIP) and CLIPSeg to generate an OVS attention map as saliency cues. Then, the saliency cues together with motion cues prompt the segment anything model (SAM) to generate a location map. In the location process, we design two lightweight adapters to fine-tune SAM, which makes SAM well adapt to the downstream UVOS task. Finally, the location map generated by SAM is used to progressively guide object representation refinement in the appearance branch, ultimately achieving an accurate segmentation mask prediction. Extensive evaluations on DAVIS-16, FBMS, and YouTube-Objects demonstrate the favorable performance of our OVSNet over the state-of-the-art methods. Zhidong Han, Shenglong Hu, Huihui Song 0003, Kaihua Zhang 0001 |
ICASSP | 4 |
| 2025 | Spatio-Semantic Prompt guided Adaptive Segment Anything for Remote Sensing Change DetectionabstractExisting leading remote sensing change detection (RSCD) often takes a semantic-agnostic learning paradigm, which uses a binary ground-truth mask as supervision for model training. Despite the demonstrated success, due to the intrinsic characteristic of extremely complicated scene changes in RS images, this paradigm is prone to be misled by irrelevant semantic category changes, leading to a noisy CD mask prediction. To address this issue, this paper presents a Spatio-Semantic Prompt (SSP) guided adaptive Segment Anything Model (SAM) for RSCD, dubbed as SSP-SAM. The SSP-SAM introduces sparse textual and dense mask prompts into SAM to encode the task-specific semantic knowledge for RSCD. Specifically, we first encode the powerful textual semantic knowledge using Contrastive Language-Image Pre-training (CLIP) to determine the desired change semantic category. Then, we design a spatial dense prompt module that yields an attention map as prompt features to further refine the desired changed regions. Subsequently, we fine-tune the SAM through an adaptor to integrate the spatial-semantic prompt cues, yielding a coarse CD mask prediction. Finally, guided by the coarse CD mask, a multi-scale mask attention mechanism is adopted to learn the refined semantic representations of the changed targets, predicting the accurate CD mask. Extensive experiments on a variety of benchmark datasets demonstrate that the proposed SSP-SAM achieves state-of-the-art performance. Shenglong Hu, Zhidong Han, Gang Dong, Lingyan Liang, Dongchao Wen, Kaihua Zhang 0001 |
ICASSP | 6 |
| 2025 | Group-wise Semantic-enhanced Interaction Network for Remote Sensing Spatio-Temporal FusionabstractRemote sensing spatio-temporal fusion (STF) aims at fusing temporally-dense coarse-resolution images and temporally-sparse fine-resolution images to reconstruct high spatio-temporal resolution images. Multi-band remote sensing images are often accepted as inputs for STF that have complementary characteristics for high-fidelity land surface reconstruction, yet the existing STF framework often treats different bands uniformly without considering the statistical correlation between different bands, resulting in unsatisfying results. To address this problem, this paper presents a group-wise semantic-enhanced interactive network for STF, dubbed as GSINet. Based on statistical observations, the feature correlation between the visible-light group and the invisible-light group is weak, while the intra-group correlation is strong. Therefore, the GSINet first separates the inputs into visible-light and invisible-light groups, which are fed into different branches with independent encoders for feature extraction and fusion. Afterwards, to address the issue of land cover changes between the prediction coarse- and reference fine-resolution images, a Semantic-Enhancement Fusion Module (SEFM) is designed to interact with the features from the same group with enhanced semantic information captured in an unsupervised learning manner. Then, the semantic-enhanced fused features from different bands are fed into an Interleaved Cross-attention Module (ICM) for further fusion. Finally, the output fusion features fully encode the intra- and inter-group information, which are fed into the decoder, reconstructing the spatio-temporal high-resolution images. Extensive experiments on CIA and LGC benchmark datasets demonstrate that the GSINet outperforms a variety of state-of-the-art methods in terms of multiple metrics. Baoluo Zhu, Shenglong Hu, Huihui Song 0003, Kaihua Zhang 0001 |
ICASSP | 4 |
| 2025 | Joint Feature Learning and Mixing via State Space Model for Remote Sensing Change DetectionabstractRemote sensing change detection (RSCD) aims to identify change areas between bi-temporal images of the same location captured at different points in time. The existing siamese framework adopts a shared-weight strategy to process each bitemporal image independently. Due to the lack of inter-image information interaction, this strategy exhibits limited capability to target change perception and discrimination when faced with small targets or ambiguous changes such as low-covering change areas and irregular morphology in real-world complex scenes. To this end, we propose a one-stream framework using the State Space (SS) Model Mamba to jointly perform feature learning and mixing for RSCD, dubbed as SSCD. Specifically, by leveraging the long-range modeling capability and linear computational complexity of the SS model Mamba, the SSCD leverages a unified approach to feature extraction and information integration through simultaneously processing the bi-temporal images. This enables the model to intensively mutually guide to extract discriminative change features. In addition, to recover more spatial details, we design a texture enhancement module that makes full use of the selective scan modeling capability of the Mamba to enhance the texture features in different directions. Without bells and whistles, our SSCD achieves the state-of-the-art performance on three benchmark datasets including SYSU-CD, LEVIR-CD, and LEVIR+-CD. Shenglong Hu, Huihui Song 0003, Kaihua Zhang 0001 |
ICME | 4 |
| 2025 | Learning Self-Corrective Network via Adaptive Self-Labeling and Dynamic NMS for High-Performance Long-Term TrackingabstractThis article presents a self-corrective network-based long-term tracker (SCLT) including a self-modulated tracking reliability evaluator (STRE) and a self-adjusting proposal postprocessor (SPPP). The targets in the long-term sequences often suffer from severe appearance variations. Existing long-term trackers often online update their models to adapt the variations, but the inaccurate tracking results introduce cumulative error into the updated model that may cause severe drift issue. To this end, a robust long-term tracker should have the self-corrective capability that can judge whether the tracking result is reliable or not, and then it is able to recapture the target when severe drift happens caused by serious challenges (e.g., full occlusion and out-of-view). To address the first issue, the STRE designs an effective tracking reliability classifier that is built on a modulation subnetwork. The classifier is trained using the samples with pseudo labels generated by an adaptive self-labeling strategy. The adaptive self-labeling can automatically label the hard negative samples that are often neglected in existing trackers according to the statistical characteristics of target state, and the network modulation mechanism can guide the backbone network to learn more discriminative features without extra training data. To address the second issue, after the STRE has been triggered, the SPPP follows it with a dynamic NMS to recapture the target in time and accurately. In addition, the STRE and the SPPP demonstrate good transportability ability, and their performance is improved when combined with multiple baselines. Compared to the commonly used greedy NMS, the proposed dynamic NMS leverages an adaptive strategy to effectively handle the different conditions of in view and out of view, thereby being able to select the most probable object box that is essential to accurately online update the basic tracker. Extensive evaluations on four large-scale and challenging benchmark datasets including VOT2021LT, OxUvALT, TLP, and LaSOT demonstrate superiority of the proposed SCLT to a variety of state-of-the-art long-term trackers in terms of all measures. Source codes and demos can be found at https://github.com/TJUT-CV/SCLT. Wanli Xue, Kaihua Zhang 0001, Bo Liu 0005, Chengwei Zhang 0001, Jingen Liu, Shengyong Chen |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Generalizable Fourier Augmentation for Unsupervised Video Object SegmentationabstractThe performance of existing unsupervised video object segmentation methods typically suffers from severe performance degradation on test videos when tested in out-of-distribution scenarios. The primary reason is that the test data in real- world may not follow the independent and identically distribution (i.i.d.) assumption, leading to domain shift. In this paper, we propose a generalizable fourier augmentation method during training to improve the generalization ability of the model. To achieve this, we perform Fast Fourier Transform (FFT) over the intermediate spatial domain features in each layer to yield corresponding frequency representations, including amplitude components (encoding scene-aware styles such as texture, color, contrast of the scene) and phase components (encoding rich semantics). We produce a variety of style features via Gaussian sampling to augment the training data, thereby improving the generalization capability of the model. To further improve the cross-domain generalization performance of the model, we design a phase feature update strategy via exponential moving average using phase features from past frames in an online update manner, which could help the model to learn cross-domain-invariant features. Extensive experiments show that our proposed method achieves the state-of-the-art performance on popular benchmarks. Huihui Song 0003, Tiankang Su, Yuhui Zheng, Kaihua Zhang 0001, Bo Liu 0005, Dong Liu 0002 |
AAAI | 4 |
| 2024 | Text2LiDAR: Text-Guided LiDAR Point Cloud Generation via Equirectangular Transformer
Kaihua Zhang 0001, Jianjun Qian, Jin Xie 0001, Jian Yang 0003 |
ECCV (56) | 2 |
| 2024 | Segment Anything Model Guided Semantic Knowledge Learning For Remote Sensing Change DetectionabstractExisting deep learning based remote sensing change detection (RSCD) methods only rely on binary ground-truth to guide the network learning while neglecting the useful semantic guidance. As a result, the network can be readily misled by irrelevant category changes, leading to degraded performance and slow convergence of the model. To this end, we propose a novel segment anything model (SAM) guided framework, termed as SAM-CD, which mines the rich semantic knowledge from the SAM for RSCD. Specifically, we first employ a transformer encoder to extract multi-scale global features from the bi-temporal images. Meanwhile, we obtain semantic prior masks from the bi-temporal images by providing the SAM with category-relevant text prompts. Then, using the semantic prior masks as constraints, we design a masked attention module (MAM) that generates local features related to the interested categories. Finally, the local and global features are fused and fed into a multi-layer perception (MLP) decoder to obtain the change map. The whole network is trained in an end-to-end manner that can readily encode the rich semantic knowledge of the changed targets to predict an accurate change map. Extensive experiments demonstrate that the proposed SAM-CD achieves state-of-the-art performance on a variety of benchmark datasets. Zixuan Sun, Huihui Song 0003, Kaihua Zhang 0001, Gang Dong, Lingyan Liang, Yaqian Zhao |
ICASSP | 3 |
| 2024 | Glance, Focus and Refinement Network for Remote Sensing Change DetectionabstractExisting change detection (CD) methods often directly fuse the multi-level features from bi-temporal remote sensing images without discriminatively considering each pixel's importance. Despite the demonstrated success, unselectively mixing the features degrades the model's performance to effectively capture the change targets due to the imbalance ratio between the change regions and the whole scene. To this end, this paper presents a glance, focus, and refinement network (GFRNet), which formulates CD as a continuous, step-by-step focusing process to mimic the human visual system. Specifically, the GFRNet first employs a transformer encoder to extract the global features from the bi-temporal images, where each feature takes a glance at the whole scene. Then, the GFRNet gradually pays attention to a cascade of salient regions, and ultimately progressively refines its focus on the desired areas of change. Comprehensive evaluations on two extensively utilized benchmark datasets, including LEVIR-CD and WHU-CD, demonstrate the superiority of our GFR-Net to a variety of state-of-the-art methods. Zixuan Sun, Yuhui Zheng, Kaihua Zhang 0001, Gang Dong, Lingyan Liang, Yaqian Zhao |
ICASSP | 4 |
| 2024 | Language-Guided Semantic Alignment for Co-saliency DetectionabstractPrevious pure vision paradigm for co-saliency detection (COD) predominantly employs supervised training. The supervisory signals often consist of binary masks or a combination of masks and category labels. However, constrained by limited training samples, these models often suffer from overfitting issue, struggling to generalize to unseen samples. To this end, this paper presents the constrative language-image pretraining-COD (CLIP-COD), a novel language-guided semantic alignment paradigm for COD. The primary objective is to leverage CLIP for aligning concepts between language and images, where the alignment can effectively leverage the powerful language understanding capability of CLIP and transfer its knowledge to image domain, thereby enhancing the model’s zero-shot generalization ability for COD. Firstly, we propose a semantic alignment branch (SAB) that can learn rich knowledge for comprehending images globally. Meanwhile, the SAB can narrow the gap in high-dimensional feature space between the language and image features, transferring the powerful semantic knowledge from CLIP to our model. Subsequently, we devise an intra-group multi-fusion module (IMM) to capture features that integrate group knowledge as dense prompts, providing spatial localization information for subsequent fine segmentation. Finally, we input sparse language prompts and dense mask cues into the pre-trained SAM decoder to obtain the final COD results. Additionally, we further design a transfer optimization adaptor, which can reduce the model training scale, saving computing resource and cost greatly. Extensive experiments on three benchmark datasets, including CoSal2015, CoCA, and CoSOD3k, demonstrate the superior performance of our CLIP-COD to a variety of state-of-the-art methods. Chuang Ding, Huihui Song 0003, Kaihua Zhang 0001 |
ICME | 4 |
| 2024 | Group-wise co-salient object detection via multi-view self-labeling novel class discovery
Gang Dong, Lingyan Liang, Yaqian Zhao, Kaihua Zhang 0001 |
Frontiers Comput. Sci. | 5 |
| 2024 | Dual temporal memory network with high-order spatio-temporal graph learning for video object segmentation
Jiaqing Fan, Shenglong Hu, Kaihua Zhang 0001, Bo Liu 0005 |
Image Vis. Comput. | 4 |
| 2024 | Hunt-inspired Transformer for visual object tracking
Wanli Xue, Kaihua Zhang 0001, Shengyong Chen |
Pattern Recognit. | 4 |
| 2024 | Gloss Prior Guided Visual Feature Learning for Continuous Sign Language RecognitionabstractContinuous sign language recognition (CSLR) is to recognize the glosses in a sign language video. Enhancing the generalization ability of CSLR's visual feature extractor is a worthy area of investigation. In this paper, we model glosses as priors that help to learn more generalizable visual features. Specifically, the signer-invariant gloss feature is extracted by a pre-trained gloss BERT model. Then we design a gloss prior guidance network (GPGN). It contains a novel parallel densely-connected temporal feature extraction (PDC-TFE) module for multi-resolution visual feature extraction. The PDC-TFE captures the complex temporal patterns of the glosses. The pre-trained gloss feature guides the visual feature learning through a cross-modality matching loss. We propose to formulate the cross-modality feature matching into a regularized optimal transport problem, it can be efficiently solved by a variant of the Sinkhorn algorithm. The GPGN parameters are learned by optimizing a weighted sum of the cross-modality matching loss and CTC loss. The experiment results on German and Chinese sign language benchmarks demonstrate that the proposed GPGN achieves competitive performance. The ablation study verifies the effectiveness of several critical components of the GPGN. Furthermore, the proposed pre-trained gloss BERT model and cross-modality matching can be seamlessly integrated into other RGB-cue-based CSLR methods as plug-and-play formulations to enhance the generalization ability of the visual feature extractor. Leming Guo, Wanli Xue, Bo Liu 0005, Kaihua Zhang 0001, Tiantian Yuan, Dimitris N. Metaxas |
IEEE Trans. Image Process. | 4 |
| 2024 | Learning Dynamic Compact Memory Embedding for Deformable Visual Object TrackingabstractRecently, template-based trackers have become the leading tracking algorithms with promising performance in terms of efficiency and accuracy. However, the correlation operation between query feature and the given template only achieves accurate target localization, but is prone to state estimation error, especially when the target suffers from severe deformation. To address this issue, segmentation-based trackers are proposed that use per-pixel matching to improve the tracking performance of deformable objects effectively. However, most of the existing trackers only match with the target features of the initial frame, thereby lacking the discrimination for handling a variety of challenging factors, e.g., similar distractors, background clutter, and appearance change. To this end, we propose a dynamic compact memory embedding technique to enhance the discrimination of the segmentation-based visual tracking method that can well tell the target from the background. Specifically, we initialize a memory embedding with the target features in the first frame. During the tracking process, the current target features that have certain correlation with the existing memory are updated to the memory embedding online. To further improve the tracking accuracy for deformable objects, we use a weighted point-to-global matching strategy to measure the correlation between the pixelwise query feature and the whole template, so as to capture more detailed deformation information. Extensive evaluations on six challenging tracking benchmarks including VOT2016, VOT2018, VOT2019, GOT-10K, TrackingNet, and LaSOT demonstrate the superiority of our method over recent remarkable trackers. Besides, our tracker outperforms the excellent segmentation-based trackers, i.e., D3S and SiamMask on the DAVIS2017 benchmark. The code is available at https://github.com/peace-love243/CMEDFL. Pengfei Zhu 0001, Kaihua Zhang 0001, Yu Wang 0106, Tianzhu Zhang 0001, Qinghua Hu |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | Distilling Cross-Temporal Contexts for Continuous Sign Language RecognitionabstractContinuous sign language recognition (CSLR) aims to recognize glosses in a sign language video. State-of-the-art methods typically have two modules, a spatial perception module and a temporal aggregation module, which are jointly learned end-to-end. Existing results in [9, 20, 25, 36] have indicated that, as the frontal component of the over-all model, the spatial perception module used for spatial feature extraction tends to be insufficiently trained. In this paper, we first conduct empirical studies and show that a shallow temporal aggregation module allows more thor-ough training of the spatial perception module. However, a shallow temporal aggregation module cannot well capture both local and global temporal context information in sign language. To address this dilemma, we propose a cross-temporal context aggregation (CTCA) model. Specifically, we build a dual-path network that contains two branches for perceptions of local temporal context and global temporal context. We further design a cross-context knowledge distil-lation learning objective to aggregate the two types of con-text and the linguistic prior. The knowledge distillation en-ables the resultant one-branch temporal aggregation mod-ule to perceive local-global temporal and semantic context. This shallow temporal perception module structure facili-tates spatial perception module learning. Extensive exper-iments on challenging CSLR benchmarks demonstrate that our method outperforms all state-of-the-art methods. Leming Guo, Wanli Xue, Qing Guo 0005, Bo Liu 0005, Kaihua Zhang 0001, Tiantian Yuan, Shengyong Chen |
CVPR | 5 |
| 2023 | Co-Salient Object Detection with Uncertainty-Aware Group Exchange-MaskingabstractThe traditional definition of co-salient object detection (CoSOD) task is to segment the common salient objects in a group of relevant images. Existing CoSOD models by-default adopt the group consensus assumption. This brings about model robustness defect under the condition of irrelevant images in the testing image group, which hinders the use of CoSOD models in real-world applications. To address this issue, this paper presents a group exchange-masking (GEM) strategy for robust CoSOD model learning. With two group of image containing different types of salient object as input, the GEM first selects a set of images from each group by the proposed learning based strategy, then these images are exchanged. The proposed feature extraction module considers both the uncertainty caused by the irrelevant images and group consensus in the remaining relevant images. We design a latent variable generator branch which is made of conditional variational autoencoder to generate uncertainly-based global stochastic features. A CoSOD transformer branch is devised to capture the correlation-based local features that contain the group consistency information. At last, the output of two branches are concatenated and fed into a transformer-based decoder, producing robust co-saliency prediction. Extensive evaluations on co-saliency detection with and without irrelevant images demonstrate the superiority of our method over a variety of state-of-the-art methods. Huihui Song 0003, Bo Liu 0005, Kaihua Zhang 0001, Dong Liu 0002 |
CVPR | 4 |
| 2023 | Group-Wise Co-Salient Object Detection with Siamese Transformers Via Brownian Distance Covariance MatchingabstractCo-salient object detection (CoSOD) aims to discover and segment foreground targets in a group of images with the same semantic category. Existing mainstream approaches often employ convolutional neural networks (CNNs) to learn the semantic-invariant features from a group of images. Despite demonstrated success, there exist two limitations: 1) The CNNs introduce the inductive bias of locality that are difficult to model long-range dependency, limiting their feature representation capability. 2) Their models lack discriminability to differentiate semantic differences between different groups since only one group of images with the same semantic category has been taken into account for model training. To address these issues, this paper presents a Siamese Transformer architecture for CoSOD that can fully mine the group-wise semantic contrast information for more discriminative feature learning. Specifically, the designed Siamese Transformer takes two groups of images as input for feature contrastive learning. Each group is processed by a Transformer branch with shared weights to capture the long-range interaction information. Besides, to model the complex non-linear interactions between these two branches, we further design a Brownian distance covariance (BDC) module that uses joint distribution to measure the inter- and intra-group semantic similarity. The BDC can be efficiently calculated in closed form that can fully characterize independence for effective feature contrastive learning. Extensive evaluations on the three largest and most challenging benchmark datasets (CoSal2015, CoCA, and CoSOD3k) demonstrate the superiority of our method over a variety of state-of-the-art methods. Lingyan Liang, Yaqian Zhao, Kaihua Zhang 0001 |
ICASSP | 5 |
| 2023 | Object-Aware Calibrated Depth-Guided Transformer for RGB-D Co-Salient Object DetectionabstractThe key role of RGB-D co-salient object detection is to effectively fuse the common information of RGB and depth signals. Existing works directly mix the information captured from both original depth maps and RGB images, but ignore one critical issue: due to the low contrast of the neighborhood objects in depth, the depth maps’ salient regions may correspond to the interference background regions in the RGB images, thereby leading to unsatisfying performance. To address this issue, we propose an Object-aware Calibrated Depth guided transformer (dubbed as OCDFormer) for RGB-D co-salient object detection. The OCDFormer mainly consists of two key designs: First, we design a depth calibration module via spectral clustering, which yields a group of calibrated depth maps that can highlight the co-object region while suppressing the interference regions. Second, we construct a cross-modal transformer, in which the common information from the RGB and the calibrated depth maps are fully captured by first injecting common tokens into the individual tokens, and then mixing them with an interaction-attention mechanism. Extensive evaluations demonstrate that our OCDFormer sets a new state-of-the-art on two public standard benchmarks including RGB-D CoSall5O and RGB-D CoSegl83. Lingyan Liang, Yaqian Zhao, Kaihua Zhang 0001 |
ICME | 4 |
| 2023 | 'Skimming-Perusal' Detection: A Simple Object Detection Baseline in GigaPixel-level ImagesabstractObject detection has achieved amazing performance in regular-sized images, but with the emergence of gigapixel-level images, even the most advanced object detection methods cannot be directly used to process them quickly and efficiently. Therefore, this paper proposes a simple baseline for gigapixel-level images object detection called Skimming-Perusal Detection (SPDet). The SPDet consists mainly of two parts, a skimming model and a perusal model. The skimming model is based on an efficient global-to-local search strategy to detect possible regions containing objects. Non-object regions are merged through a skimming iterative merging strategy to generate skimming patch candidates. The perusal model adaptive scales the skimming patch candidates guided by the coarse detection of the skimming model. Extensive evaluations on the PANDA dataset demonstrate that the SPDet boosts detection speed on gigapixel-level images by 6× while achieving better performance than a variety of state-of-the-art methods. The source code is released at https://github.com/TJUT-CV/SPDet. Wanli Xue, Kaihua Zhang 0001, Shengyong Chen |
ICME | 3 |
| 2023 | Temporally Efficient Gabor Transformer for Unsupervised Video Object SegmentationabstractSpatial-temporal structural details of targets in video (e.g. varying edges, textures over time) are essential to accurate Unsupervised Video Object Segmentation (UVOS). The vanilla multi-head self-attention in the Transformer-based UVOS methods usually concentrates on learning the general low-frequency information (e.g. illumination, color), while neglecting the high-frequency texture details, leading to unsatisfying segmentation results. To address this issue, this paper presents a Temporally efficient Gabor Transformer (TGFormer) for UVOS. The TGFormer jointly models the spatial dependencies and temporal coherence intra- and inter-frames, which can fully capture the rich structural details for accurate UVOS. Concretely, we first propose an effective learnable Gabor filtering Transformer to mine the structural texture details of the object for accurate UVOS. Then, to adaptively store the redundant neighboring historical information, we present an efficient dynamic neighboring frame selection module to automatically choose the useful temporal information, which simultaneously relieves the blurry frame and reduces the computation burden. Finally, we make the UVOS model be a fully Transformer architecture, meanwhile aggregating the information from space, Gabor and time domains, yielding a strong representation with rich structure details. Extensive experiments on five mainstream UVOS benchmarks (DAVIS2016, FBMS, DAVSOD, ViSal, and MCL) demonstrate the superiority of the presented solution to sate-of-the-art methods. Jiaqing Fan, Tiankang Su, Kaihua Zhang 0001, Bo Liu 0005, Qingshan Liu 0001 |
ACM Multimedia | 3 |
| 2023 | Gradient-Guided Temporal Cross-Attention Transformer for High-Performance Remote Sensing Change DetectionabstractGiven a group of bi-temporal remote sensing images acquired in the same geographical area, the task of Change Detection (CD) aims to detect and segment the change regions therein. Existing leading CD methods typically use the self-attention (SA) mechanism to directly fuse the concatenated features of the bi-temporal images. Despite demonstrated success, the SA focuses primarily on modeling spatial-wise relationships, rather than channel-wise relationships. Meanwhile, since the temporal direction is along the channel direction, this makes the SA difficult to model the temporal-wise relationship between the bi-temporal features, making it fail to learn the feature correspondence from the significantly-changed regions. To this end, this letter presents a gradient-guided temporal cross-attention (Grad-TCA) mechanism for CD. First, we design a temporal cross-attention module (TCAM) that mixes the cross- and self-attention to model both temporal- and spatial-wise interactions, which fully mines the complementary cues between bi-temporal features to learn a strong feature presentation. Afterwards, to further highlight the salient features between the change regions, we design a gradient-guided module (GGM) to enhance the difference of the learned bi-temporal features through feedback gradient information. Both the TCAM and the GGM construct our Grad-TCA module, which is seamlessly integrated into a Transformer framework for end-to-end learning. Finally, to reduce the computation overhead, we design a simple change discrimination module (CDM) that outputs a score to directly filter out the unchanged features from the GGM with no need of passing the features through the decoder. Comprehensive evaluations on the two widely used benchmark datasets including LEVIR-CD and WHU-CD demonstrate our model outperforms a variety of state-of-the-art methods. Huihui Song 0003, Kaihua Zhang 0001 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2023 | Bi-RRNet: Bi-level recurrent refinement network for camouflaged object detection
Yan Liu 0004, Kaihua Zhang 0001, Yaqian Zhao, Qingshan Liu 0001 |
Pattern Recognit. | 2 |
| 2023 | Deep Object Co-Segmentation and Co-Saliency Detection via High-Order Spatial-Semantic Network ModulationabstractObject co-segmentation (CSG) is to segment the common objects of the same category in multiple relevant images while the co-saliency detection (CSD) aims to discover the salient and common foreground objects in a group of images. To process both tasks simultaneously, this paper presents an adaptive spatially and high-order semantically modulated deep network framework. A backbone network is first adopted to extract multi-resolution image features. With the multi-resolution features of the relevant images as input, we design an adaptive spatial modulator to learn a spatial representation that can highlight the co-object regions for each image. The adaptive spatial modulator fully captures the rich correlations of all image feature descriptors via unsupervised clustering and a graph aggregation strategy. The learned representation can well localize the common foreground object while effectively suppressing the background signals. For the high-order semantic modulator, we model it as a supervised image classification task. We propose a hierarchical high-order pooling module to learn the rich semantic features for classification use. The outputs of the two modulators manipulate the multi-resolution features by a shift-and-scale operation so that the features focus on segmenting common object regions. The proposed model is trained end-to-end without any intricate post-processing. Extensive experiments on three CSG benchmark datasets (MSRC, i-Coseg, and PASCAL-VOC) and three CSD datasets (Cosal2015, CoCA, and CoSOD3k) demonstrate the superior accuracy of the proposed method compared to state-of-the-art methods on both tasks. Kaihua Zhang 0001, Mingliang Dong, Bo Liu 0005, Dong Liu 0002, Qingshan Liu 0001 |
IEEE Trans. Multim. | 1 |
| 2022 | Bidirectionally Learning Dense Spatio-temporal Feature Propagation Network for Unsupervised Video Object SegmentationabstractSpatio-temporal feature representation is essential for accurate unsupervised video object segmentation, which needs an effective feature propagation paradigm for both appearance and motion features that can fully interchange information across frames. However, existing solutions mainly focus on the forward feature propagation from the preceding frame to the current one, either using the former segmentation mask or motion propagation in a frame-by-frame manner. This ignores the bi-directional temporal feature interactions (including the backward propagation from the future to the current frame) across all frames that can help to enhance the spatiotemporal feature representation for segmentation prediction. To this end, this paper presents a novel Dense Bidirectional Spatio-temporal feature propagation Network (DBSNet) to fully integrate the forward and the backward propagations across all frames. Specifically, a dense bi-ConvLSTM module is first developed to propagate the features across all frames in a forward and backward manner. This can fully capture the multi-level spatio-temporal contextual information across all frames, producing an effective feature representation that has a strong discriminative capability to tell from noisy backgrounds. Following it, a spatio-temporal Transformer refinement module is designed to further enhance the propagated features, which can effectively capture the spatio-temporal long-range dependencies among all frames. Afterwards, a Co-operative Direction-aware Graph Attention (Co-DGA) module is designed to integrate the propagated appearancemotion cues, yielding a strong spatio-temporal feature representation for segmentation mask prediction. The Co-DGA assigns proper attentional weights to neighboring points along the coordinate axis, making the segmentation model to selectively focus on the most relevant neighbors. Extensive evaluations on four mainstream challenging benchmarks including DAVIS16, FBMS, DAVSOD, and MCL demonstrate that the proposed DBSNet achieves favorable performance against state-of-the-art methods in terms of all evaluation metrics. Jiaqing Fan, Tiankang Su, Kaihua Zhang 0001, Qingshan Liu 0001 |
ACM Multimedia | 3 |
| 2022 | Learning Self-supervised Low-Rank Network for Single-Stage Weakly and Semi-supervised Semantic Segmentation
Junwen Pan, Pengfei Zhu 0001, Kaihua Zhang 0001, Bing Cao 0002, Yu Wang 0106, Dingwen Zhang, Junwei Han 0001, Qinghua Hu |
Int. J. Comput. Vis. | 3 |
| 2022 | Semi-Supervised Video Object Segmentation via Learning Object-Aware Global-Local CorrespondenceabstractIn semi-supervised video object segmentation (VOS) task, temporal coherent object-level cues play a key role yet are hard to accurately model. To this end, this paper presents an object-aware global-local correspondence architecture, which enables to extract the inter-frame temporal coherent object-level features for accurate VOS. Specifically, we first generate a set of object masks by the ground-truth segmentation, and then we squeeze the current frame representation inside the object masks into a set of global object embeddings. Second, we compute the similarity between each embedding and the feature map, producing an object-aware weight for each pixel. The object-aware feature at each pixel is then constructed by summing the object embeddings weighted by their corresponding object-aware weights, which is able to capture rich object category information. Third, to establish the accurate correspondences between the inter-frame temporal coherent cues, we further design a novel global-local correspondence module to refine the temporal feature representations. Finally, we augment the object-aware features with the global-local aligned information to produce a strong spatio-temporal representation, which is essential to a more reliable pixel-wise segmentation prediction. Extensive evaluations are conducted on three popular VOS benchmarks containing Youtube-VOS, Davis2017 and Davis2016, demonstrating that the proposed method achieves favourable performance compared to the state-of-the-arts. Jiaqing Fan, Bo Liu 0005, Kaihua Zhang 0001, Qingshan Liu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Image Co-Saliency Detection and Instance Co-Segmentation Using Attention Graph Clustering Based Graph Convolutional NetworkabstractCo-Saliency Detection (CSD) is to explore the concurrent patterns and salient objects from a group of relevant images, while Instance Co-Segmentation (ICS) aims to identify and segment out all of these co-salient instances, generating corresponding mask for each instance. To simultaneously tackle these two tasks, we present a novel adaptive graph convolutional network with attention graph clustering (GCAGC) for CSD and ICS, termed as GCAGC-CSD and GCAGC-ICS, respectively. The GCAGC-CSD contains three key model designs: first, we develop a graph convolutional network architecture to extract multi-scale representations to characterize the intra- and inter-image consistency. Second, we propose an attention graph clustering algorithm to distinguish the salient foreground objects from common areas in an unsupervised manner. Third, we present a unified framework with encoder-decoder structure to jointly train and optimize the graph convolutional network, attention graph cluster, and CSD decoder in an end-to-end fashion. Afterwards, we design a salient instance segmentation network for GCAGC-ICS, and combine the outputs of GCAGC-CSD and the instance segmentation branch to obtain instance-aware co-segmentation masks. The proposed GCAGC-CSD and GCAGC-ICS are extensively evaluated on four CSD benchmark datasets (iCoseg, Cosal2015, COCO-SEG and CoSOD3k) and five ICS benchmark datasets (CoSOD3k, COCO-NONVOC, COCO-VOC, VOC12 and SOC), and achieve superior performance over state-of-the-arts on both tasks. Tengpeng Li, Kaihua Zhang 0001, Shiwen Shen, Bo Liu 0005, Qingshan Liu 0001, Zhu Li 0001 |
IEEE Trans. Multim. | 2 |
| 2021 | DeepACG: Co-Saliency Detection via Semantic-Aware Contrast Gromov-Wasserstein DistanceabstractThe objective of co-saliency detection is to segment the co-occurring salient objects in a group of images. To address this task, we introduce a new deep network architecture via semantic-aware contrast Gromov-Wasserstein distance (DeepACG). We first adopt the Gromov-Wasserstein (GW) distance to build dense 4D correlation volumes for all pairs of image pixels within the image group. These dense correlation volumes enable the network to accurately discover the structured pair-wise pixel similarities among the common salient objects. Second, we develop a semantic-aware co-attention module (SCAM) to enhance the foreground co-saliency through predicted categorical information. Specifically, SCAM recognizes the semantic class of the foreground co-objects, and this information is then modulated to the deep representations to localize the related pixels. Third, we design a contrast edge-enhanced module (EEM) to capture richer contexts and preserve fine-grained spatial information. We validate the effectiveness of our model using three largest and most challenging benchmark datasets (Cosal2015, CoCA, and CoSOD3k). Extensive experiments have demonstrated the substantial practical merit of each module. Compared with the existing works, DeepACG shows significant improvements and achieves state-of-the-art performance. Kaihua Zhang 0001, Mingliang Dong, Bo Liu 0005, Xiao-Tong Yuan, Qingshan Liu 0001 |
CVPR | 1 |
| 2021 | Deep Transport Network for Unsupervised Video Object SegmentationabstractThe popular unsupervised video object segmentation methods fuse the RGB frame and optical flow via a two-stream network. However, they cannot handle the distracting noises in each input modality, which may vastly deteriorate the model performance. We propose to establish the correspondence between the input modalities while suppressing the distracting signals via optimal structural matching. Given a video frame, we extract the dense local features from the RGB image and optical flow, and treat them as two complex structured representations. The Wasserstein distance is then employed to compute the global optimal flows to transport the features in one modality to the other, where the magnitude of each flow measures the extent of the alignment between two local features. To plug the structural matching into a two-stream network for end-to-end training, we factorize the input cost matrix into small spatial blocks and design a differentiable long-short Sinkhorn module consisting of a long-distant Sinkhorn layer and a short-distant Sinkhorn layer. We integrate the module into a dedicated two-stream network and dub our model TransportNet. Our experiments show that aligning motion-appearance yields the state-of-the-art results on the popular video object segmentation datasets. Kaihua Zhang 0001, Zicheng Zhao, Dong Liu 0002, Qingshan Liu 0001, Bo Liu 0005 |
ICCV | 1 |
| 2021 | Conditional generative adversarial network with densely-connected residual learning for single image super-resolution
Jiaojiao Qiao, Huihui Song 0003, Kaihua Zhang 0001 |
Multim. Tools Appl. | 3 |
| 2021 | Video saliency prediction using enhanced spatiotemporal alignment network
Huihui Song 0003, Kaihua Zhang 0001, Bo Liu 0005, Qingshan Liu 0001 |
Pattern Recognit. | 3 |
| 2021 | Feature Alignment and Aggregation Siamese Networks for Fast Visual TrackingabstractSiamese networks have been successfully introduced into visual tracking, which match the best candidate and a target template via a couple of networks with shared parameters. However, most Siamese network-based trackers (SNTs) are tailored to best match the canonical posture of the template and the search-region images, resulting in inferior performance when the target objects have large-scale pose variations. Besides, SNTs fail to discriminate distractors well because they only leverage high-level semantic features as target representations that cannot well tell from different targets of the same category. To address these issues, this paper presents an efficient and effective SNT that is based on feature alignment and aggregation networks. Specifically, we first design an effective feature alignment network module to calibrate the search-region image. This module results in a more reliable matching response that is robust to severe target pose variations. Then, we develop an effective shallow-level and high-level feature aggregation network module to complement the feature characteristics, making the learned feature representation not only well differentiate the target from distractors, but also robust to target appearance variations. Afterwards, we employ a channel-attention mechanism to further strengthen the discriminative capability of the aggregated feature representation. Finally, both the alignment and the aggregation modules are seamlessly integrated into the Siamese networks for robust tracking. Meanwhile, we offline learn the network parameters end-to-end without time-consuming fine-tuning. Extensive evaluations on a variety of benchmarks including VOT-2017, OTB-100, UAV123 and GOT-10k demonstrate favorable performance of our tracker against state-of-the-art ones with a speed of 60 fps. Jiaqing Fan, Huihui Song 0003, Kaihua Zhang 0001, Qingshan Liu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2020 | Deep Object Co-Segmentation via Spatial-Semantic Network ModulationabstractObject co-segmentation is to segment the shared objects in multiple relevant images, which has numerous applications in computer vision. This paper presents a spatial and semantic modulated deep network framework for object co-segmentation. A backbone network is adopted to extract multi-resolution image features. With the multi-resolution features of the relevant images as input, we design a spatial modulator to learn a mask for each image. The spatial modulator captures the correlations of image feature descriptors via unsupervised learning. The learned mask can roughly localize the shared foreground object while suppressing the background. For the semantic modulator, we model it as a supervised image classification task. We propose a hierarchical second-order pooling module to transform the image features for classification use. The outputs of the two modulators manipulate the multi-resolution features by a shift-and-scale operation so that the features focus on segmenting co-object regions. The proposed model is trained end-to-end without any intricate post-processing. Extensive experiments on four image co-segmentation benchmark datasets demonstrate the superior accuracy of the proposed method compared to state-of-the-art methods. The codes are available at http://kaihuazhang.net/. Kaihua Zhang 0001, Bo Liu 0005, Qingshan Liu 0001 |
AAAI | 1 |
| 2020 | Adaptive Graph Convolutional Network With Attention Graph Clustering for Co-Saliency DetectionabstractCo-saliency detection aims to discover the common and salient foregrounds from a group of relevant images. For this task, we present a novel adaptive graph convolutional network with attention graph clustering (GCAGC). Three major contributions have been made, and are experimentally shown to have substantial practical merits. First, we propose a graph convolutional network design to extract information cues to characterize the intra- and inter-image correspondence. Second, we develop an attention graph clustering algorithm to discriminate the common objects from all the salient foreground objects in an unsupervised fashion. Third, we present a unified framework with encoder-decoder structure to jointly train and optimize the graph convolutional network, attention graph cluster, and co-saliency detection decoder in an end-to-end manner. We evaluate our proposed GCAGC method on three co-saliency detection benchmark datasets (iCoseg, Cosal2015 and COCO-SEG). Our GCAGC method obtains significant improvements over the state-of-the-arts on most of them. Kaihua Zhang 0001, Tengpeng Li, Shiwen Shen, Bo Liu 0005, Qingshan Liu 0001 |
CVPR | 1 |
| 2020 | Dual Temporal Memory Network for Efficient Video Object SegmentationabstractVideo Object Segmentation (VOS) is typically formulated in a semi-supervised setting. Given the ground-truth segmentation mask on the first frame, the task of VOS is to track and segment the single or multiple objects of interests in the rest frames of the video at the pixel level. One of the fundamental challenges in VOS is how to make the most use of the temporal information to boost the performance. We present an end-to-end network which stores short- and long-term video sequence information preceding the current frame as the temporal memories to address the temporal modeling in VOS. Our network consists of two temporal sub-networks including a short-term memory sub-network and a long-term memory sub-network. The short-term memory sub-network models the fine-grained spatial-temporal interactions between local regions across neighboring frames in video via a graph-based learning framework, which can well preserve the visual consistency of local regions over time. The long-term memory sub-network models the long-range evolution of object via a Simplified-Gated Recurrent Unit (S-GRU), making the segmentation be robust against occlusions and drift errors. In our experiments, we show that our proposed method achieves a favorable and competitive performance on three frequently-used VOS datasets, including DAVIS 2016, DAVIS 2017 and Youtube-VOS in terms of both speed and accuracy. Kaihua Zhang 0001, Dong Liu 0002, Bo Liu 0005, Qingshan Liu 0001, Zhu Li 0001 |
ACM Multimedia | 1 |
| 2020 | Top-Down Fusing Multi-level Contextual Features for Salient Object Detection
Mingyuan Pan, Huihui Song 0003, Junxia Li, Kaihua Zhang 0001, Qingshan Liu 0001 |
PRCV (3) | 4 |
| 2020 | Hierarchical Representations with Discriminative Meta-filters in Dual Path Network for Tracking
Ning Wang 0020, Yuncong Yao, Wankou Yang, Kaihua Zhang 0001, Bo Liu 0005 |
PRCV (2) | 5 |
| 2020 | Learning lightweight Multi-Scale Feedback Residual network for single image super-resolution
Huihui Song 0003, Kaihua Zhang 0001, Qingshan Liu 0001, Jia Liu 0034 |
Comput. Vis. Image Underst. | 3 |
| 2020 | Real-time manifold regularized context-aware correlation tracking
Jiaqing Fan, Huihui Song 0003, Kaihua Zhang 0001, Qingshan Liu 0001, Wei Lian |
Frontiers Comput. Sci. | 3 |
| 2020 | Recurrent reverse attention guided residual learning for saliency object detection
Tengpeng Li, Huihui Song 0003, Kaihua Zhang 0001, Qingshan Liu 0001 |
Neurocomputing | 3 |
| 2020 | Single image super-resolution with enhanced Laplacian pyramid network via conditional generative adversarial learning
Huihui Song 0003, Kaihua Zhang 0001, Jiaojiao Qiao, Qingshan Liu 0001 |
Neurocomputing | 3 |
| 2020 | Hierarchical attentive Siamese network for real-time visual tracking
Huihui Song 0003, Kaihua Zhang 0001, Qingshan Liu 0001 |
Neural Comput. Appl. | 3 |
| 2020 | Learning residual refinement network with semantic context representation for real-time saliency object detection
Tengpeng Li, Huihui Song 0003, Kaihua Zhang 0001, Qingshan Liu 0001 |
Pattern Recognit. | 3 |
| 2020 | Multi-Task Deep Dual Correlation Filters for Visual TrackingabstractCorrelation filters combined with deep features have delivered impressive results in visual tracking task. However, existing approaches treat deep features produced by different network layers independently, limiting their representation power. To address this issue, this paper proposes a multi-task deep dual correlation filters (MDDCF) based method for robust visual tracking. First, a new multi-task learning scheme is designed to take full advantage of the multi-level features of deep networks, where target representation with individual features is regarded as a single task. As such, the interdependencies between different levels of features can be better explored. Second, we reformulate the objective function of the dual correlation filters and propose a new alternating optimization method, allowing joint training of the correlation filters and network parameters. Third, we design an effective object template update scheme which can well capture the target appearance variations. Extensive experimental evaluations on seven benchmark datasets show that the proposed MDDCF tracker performs favorably against state-ofthe-art methods. Yuhui Zheng, Xinyan Liu 0002, Xu Cheng 0003, Kaihua Zhang 0001, Yi Wu 0001, Shengyong Chen |
IEEE Trans. Image Process. | 4 |
| 2020 | Dynamically Spatiotemporal Regularized Correlation TrackingabstractRecently, due to the high performance, spatially regularized strategy has been widely applied to addressing the issue of boundary effects existed in correlation filter (CF)-based visual tracking. Specifically, it introduces a spatially regularized term to penalize the coefficients of the CFs to be learned depending on their spatial locations. However, the regularization weights are often formed as a fixed Gaussian function, and hence may cause the learned model degenerate due to the inflexible constraints on the ever-changing CFs to be learned over time during tracking. To address this issue, in this paper, we develop a dynamically spatiotemporal regularization model to constrain the CFs to be learned with the ever-changing regularization weights learned from two consecutive frames. The proposed method jointly learns the CFs along with the dynamically spatiotemporal constraint term, which can be efficiently solved in the Fourier domain by the alternative direction method. Extensive evaluations on the popular data sets OTB-100 and VOT-2016 demonstrate that the proposed tracker performs favorably against the baseline tracker and several recently proposed state-of-the-art methods. Yuhui Zheng, Huihui Song 0003, Kaihua Zhang 0001, Jiaqing Fan, Xinyan Liu 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2019 | Co-Saliency Detection via Mask-Guided Fully Convolutional Networks With Multi-Scale Label SmoothingabstractIn image co-saliency detection problem, one critical issue is how to model the concurrent pattern of the co-salient parts, which appears both within each image and across all the relevant images. In this paper, we propose a hierarchical image co-saliency detection framework as a coarse to fine strategy to capture this pattern. We first propose a mask-guided fully convolutional network structure to generate the initial co-saliency detection result. The mask is used for background removal and it is learned from the high-level feature response maps of the pre-trained VGG-net output. We next propose a multi-scale label smoothing model to further refine the detection result. The proposed model jointly optimizes the label smoothness of pixels and superpixels. Experiment results on three popular image co-saliency detection benchmark datasets including iCoseg, MSRC and Cosal2015 demonstrate the remarkable performance compared with the state-of-the-art methods. Kaihua Zhang 0001, Tengpeng Li, Bo Liu 0005, Qingshan Liu 0001 |
CVPR | 1 |
| 2019 | Image super-resolution using conditional generative adversarial networkabstractRecently, extensive studies on a generative adversarial network (GAN) have made great progress in single image super‐resolution (SISR). However, there still exists a significant difference between the reconstructed high‐frequency and the real high‐frequency details. To address this issue, this study presents an SISR approach based on conditional GAN (SRCGAN). SRCGAN includes a generator network that generates super‐resolution (SR) images and a discriminator network that is trained to distinguish the SR images from ground‐truth high‐resolution (HR) ones. Specifically, the discriminator network uses the ground‐truth HR image as a conditional variable, which guides the network to distinguish the real images from the SR images, facilitating training a more stable generator model than GAN without this guidance. Furthermore, a residual‐learning module is introduced into the generator network to solve the issue of detail information loss in SR images. Finally, the network is trained in an end‐to‐end manner by optimizing a perceptual loss function. Extensive evaluations on four benchmark datasets including Set5, Set14, BSD100, and Urban100 demonstrate the superiority of the proposed SRCGAN over state‐of‐the‐art methods in terms of PSNR, SSIM, and visual effect. Jiaojiao Qiao, Huihui Song 0003, Kaihua Zhang 0001, Qingshan Liu 0001 |
IET Image Process. | 3 |
| 2019 | Low-rank weighted co-saliency detection via efficient manifold ranking
Tengpeng Li, Huihui Song 0003, Kaihua Zhang 0001, Qingshan Liu 0001, Wei Lian |
Multim. Tools Appl. | 3 |
| 2019 | Parallel Attentive Correlation TrackingabstractPsychological and cognitive findings indicate that human visual perception is attentive and selective, which may process spatial and appearance selective attentions in parallel. By reflecting some aspects of these attentions, this paper presents a novel correlation filter (CF) based tracking approach, corresponding to processing a local and a semi-local background domains, respectively. In the local domain, inspired by the Gestalt principle of figure-ground segregation, we leverage an efficient Boolean map representation, which characterizes an image by a set of Boolean maps via randomly thresholding its color channels, yielding a location response map as a weighted sum of all Boolean maps. The Boolean maps capture the topological structures of target and its scene with different granularities, thereby enabling to effectively improve tracking of non-rectangular objects. Alternatively, in the semi-local domains, we introduce a novel distractor-resilient metric regularization into CF, which acts as a force to push distractors into negative space. Consequently, the unwanted boundary effects of CF can be effectively alleviated. Finally, both models associated with the local and the semi-local domains are seamlessly integrated into a Bayesian framework, and the tracked location is determined by maximizing its likelihood function. Extensive evaluations on the OTB50, OTB100, VOT2016 and VOT2017 tracking benchmarks demonstrate that the proposed method achieves favorable performance against a variety of state-of-the-art trackers with a speed of 45 fps on a single CPU. Kaihua Zhang 0001, Jiaqing Fan, Qingshan Liu 0001, Jian Yang 0003, Wei Lian |
IEEE Trans. Image Process. | 1 |
| 2018 | Visual tracking using spatio-temporally nonlocally regularized correlation filter
Kaihua Zhang 0001, Huihui Song 0003, Qingshan Liu 0001, Wei Lian |
Pattern Recognit. | 1 |
| 2018 | Visual tracking via Boolean map representations
Kaihua Zhang 0001, Qingshan Liu 0001, Jian Yang 0003, Ming-Hsuan Yang 0001 |
Pattern Recognit. | 1 |
| 2018 | Visual Tracking via Nonlocal Similarity LearningabstractEither global (e.g., intensity histograms and coefficients of sparse representation) or local (e.g., scale-invariant feature transform and histogram of oriented gradient) feature representations have been widely exploited for visual tracking. However, most of these representations describe a target appearance with a fixed spatial grid layout without considering the interactions between different grids, and hence may adversely affect their performance when the target appearance suffers from large-scale pose variations. In this paper, we learn a similarity function that considers the interactions of features in the grids not only from the same spatial positions, but also from different positions, thereby taking charge of the nonlocal information of the target appearances to effectively handle the significant appearance variations. Specifically, we explore the polynomial kernel feature map to characterize the nonlocal similarity information of all pairs of grids among the target and its background samples, and combine these feature maps as the target representations. Moveover, we learn a linear logistic regression classifier with online update to separate the target from its local background, and integrate this classifier into a particle filtering tracking framework. Extensive experimental results on the CVPR2013 tracking benchmark demonstrate the proposed approach performs favorably against some representative tracking algorithms. Qingshan Liu 0001, Jiaqing Fan, Huihui Song 0003, Kaihua Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2018 | Visual Tracking With Weighted Adaptive Local Sparse Appearance Model via Spatio-Temporal Context LearningabstractSparse representation has been widely exploited to develop an effective appearance model for object tracking due to its well discriminative capability in distinguishing the target from its surrounding background. However, most of these methods only consider either the holistic representation or the local one for each patch with equal importance, and hence may fail when the target suffers from severe occlusion or large-scale pose variation. In this paper, we propose a simple yet effective approach that exploits rich feature information from reliable patches based on weighted local sparse representation that takes into account the importance of each patch. Specifically, we design a reconstruction-error based weight function with the reconstruction error of each patch via sparse coding to measure the patch reliability. Moreover, we explore spatio-temporal context information to enhance the robustness of the appearance model, in which the global temporal context is learned via incremental subspace and sparse representation learning with a novel dynamic template update strategy to update the dictionary, while the local spatial context considers the correlation between the target and its surrounding background via measuring the similarity among their sparse coefficients. Extensive experimental evaluations on two large tracking benchmarks demonstrate favorable performance of the proposed method over some state-of-the-art trackers. Zhetao Li, Jie Zhang 0136, Kaihua Zhang 0001, Zhiyong Li 0001 |
IEEE Trans. Image Process. | 3 |
| 2018 | Efficient Correlation Tracking via Center-Biased Spatial RegularizationabstractCorrelation filters (CFs) have been applied to visual tracking with success providing excellent performance in terms of accuracy and efficiency. The underlying periodic assumption of the training samples results in their great efficiency when using the fast Fourier transform (FFT), yet it also brings unwanted boundary effects. To address this issue, the recently proposed spatially-regularized discriminative CF (SRDCF) method introduces a Gaussian weight function to regularize the learning filter, yielding favorable performances in accuracy but high computational complexity because the objective of the SRDCF cannot achieve a closed solution via the FFT. Motivated by SRDCF, we present an efficient and effective CF-based tracker using center-biased constraint weights (CBCWs), which improve simultaneously speed and accuracy. Specifically, we first construct a CBCW function by exploiting the symmetry of the Fourier transform. The values of the constraint weights are real in both time and frequency domains, so that the optimization can be directly solved in the frequency domain without any data transformation, thereby greatly reducing its computational complexity. Moreover, according to the average peak-tocorrelation energy value of the CF response, we propose an efficient and effective filter update strategy to handle occlusions during tracking. Extensive experiments on the OTB-2013, OTB- 2015, and VOT2016 benchmarks demonstrate that the proposed tracker significantly outperforms the baseline SRDCF in terms of accuracy and efficiency. Moreover, the proposed method performs favorably against 16 other representative state-of-the-art methods regarding robustness and success rate. Jianghong Han, Fan Yang 0063, Kaihua Zhang 0001, Richang Hong |
IEEE Trans. Image Process. | 4 |
| 2017 | Robust facial landmark tracking via cascade regression
Qingshan Liu 0001, Jing Yang 0038, Jiankang Deng, Kaihua Zhang 0001 |
Pattern Recognit. | 4 |
| 2017 | Adaptive Compressive Tracking via Online Vector Boosting Feature SelectionabstractRecently, the compressive tracking (CT) method has attracted much attention due to its high efficiency, but it cannot well deal with the large scale target appearance variations due to its data-independent random projection matrix that results in less discriminative features. To address this issue, in this paper, we propose an adaptive CT approach, which selects the most discriminative features to design an effective appearance model. Our method significantly improves CT in three aspects. First, the most discriminative features are selected via an online vector boosting method. Second, the object representation is updated in an effective online manner, which preserves the stable features while filtering out the noisy ones. Furthermore, a simple and effective trajectory rectification approach is adopted that can make the estimated location more accurate. Finally, a multiple scale adaptation mechanism is explored to estimate object size, which helps to relieve interference from background information. Extensive experiments on the CVPR2013 tracking benchmark and the VOT2014 challenges demonstrate the superior performance of our method. Qingshan Liu 0001, Jing Yang 0038, Kaihua Zhang 0001, Yi Wu 0001 |
IEEE Trans. Cybern. | 3 |
| 2017 | Low-Rank Latent Pattern Approximation With Applications to Robust Image ClassificationabstractThis paper develops a novel method to address the structural noise in samples for image classification. Recently, regression-related classification methods have shown promising results when facing the pixelwise noise. However, they become weak in coping with the structural noise due to ignoring of relationships between pixels of noise image. Meanwhile, most of them need to implement the iterative process for computing representation coefficients, which leads to the high time consumption. To overcome these problems, we exploit a latent pattern model called low-rank latent pattern approximation (LLPA) to reconstruct the test image having structural noise. The rank function is applied to characterize the structure of the reconstruction residual between test image and the corresponding latent pattern. Simultaneously, the error between the latent pattern and the reference image is constrained by Frobenius norm to prevent overfitting. LLPA involves a closed-form solution by the virtue of a singular value thresholding operator. The provided theoretic analysis demonstrates that LLPA indeed removes the structural noise during classification task. Additionally, LLPA is further extended to the form of matrix regression by connecting multiple training samples, and alternating direction of multipliers method with Gaussian back substitution algorithm is used to solve the extended LLPA. Experimental results on several popular data sets validate that the proposed methods are more robust to image classification with occlusion and illumination changes, as compared to some existing state-of-the-art reconstruction-based methods and one deep neural network-based method. Shuo Chen 0003, Jian Yang 0003, Lei Luo 0001, Yang Wei 0003, Kaihua Zhang 0001, Ying Tai |
IEEE Trans. Image Process. | 5 |
| 2016 | Robust object tracking by online Fisher discrimination boosting feature selection
Jing Yang 0038, Kaihua Zhang 0001, Qingshan Liu 0001 |
Comput. Vis. Image Underst. | 2 |
| 2016 | Robust visual tracking via patch based kernel correlation filters with adaptive multiple feature ensemble
Kaihua Zhang 0001, Qingshan Liu 0001 |
Neurocomputing | 2 |
| 2016 | Dual Group Structured TrackingabstractThe sparse representation (SR)-based tracking framework generally considers the testing candidates and dictionary atoms individually, thus failing to model the structured information within data. In this paper, we present a robust tracking framework by exploiting the dual group structure of both candidate samples and dictionary templates, and formulate the SR at group level. The similar samples are encoded simultaneously by a few atom groups, which induces the inter-group sparsity, and also each group enjoys different internal sparsity. In this way, not only the potential commonality shared by the related candidates is taken into account but also the individual differences between samples are reflected. Then, we provide two effective optimization methods to solve our formulation by block-coordinate gradient descent and alternating direction method of multipliers, respectively, and make a comparison between them in terms of both effectiveness and efficiency. Finally, we embed the dual group structure model into the particle filter framework for visual tracking. Extensive experimental results demonstrate that our tracker achieves favorable performance against the state-of-the-art tracking methods. Fu Li 0003, Huchuan Lu, Dong Wang 0004, Yi Wu 0001, Kaihua Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2016 | A Level Set Approach to Image Segmentation With Intensity InhomogeneityabstractIt is often a difficult task to accurately segment images with intensity inhomogeneity, because most of representative algorithms are region-based that depend on intensity homogeneity of the interested object. In this paper, we present a novel level set method for image segmentation in the presence of intensity inhomogeneity. The inhomogeneous objects are modeled as Gaussian distributions of different means and variances in which a sliding window is used to map the original image into another domain, where the intensity distribution of each object is still Gaussian but better separated. The means of the Gaussian distributions in the transformed domain can be adaptively estimated by multiplying a bias field with the original signal within the window. A maximum likelihood energy functional is then defined on the whole image region, which combines the bias field, the level set function, and the piecewise constant function approximating the true image signal. The proposed level set method can be directly applied to simultaneous segmentation and bias correction for 3 and 7T magnetic resonance images. Extensive evaluation on synthetic and real-images demonstrate the superiority of the proposed method over other representative algorithms. Kaihua Zhang 0001, Lei Zhang 0006, Kin-Man Lam 0001, David Zhang 0001 |
IEEE Trans. Cybern. | 1 |
| 2016 | Robust Visual Tracking via Convolutional Networks Without TrainingabstractDeep networks have been successfully applied to visual tracking by learning a generic representation offline from numerous training images. However, the offline training is time-consuming and the learned generic representation may be less discriminative for tracking specific objects. In this paper, we present that, even without offline training with a large amount of auxiliary data, simple two-layer convolutional networks can be powerful enough to learn robust representations for visual tracking. In the first frame, we extract a set of normalized patches from the target region as fixed filters, which integrate a series of adaptive contextual filters surrounding the target to define a set of feature maps in the subsequent frames. These maps measure similarities between each filter and useful local intensity patterns across the target, thereby encoding its local structural information. Furthermore, all the maps together form a global representation, via which the inner geometric layout of the target is also preserved. A simple soft shrinkage method that suppresses noisy values below an adaptive threshold is employed to de-noise the global representation. Our convolutional networks have a lightweight structure and perform favorably against several state-of-the-art methods on the recent tracking benchmark data set with 50 challenging videos. Kaihua Zhang 0001, Qingshan Liu 0001, Yi Wu 0001, Ming-Hsuan Yang 0001 |
IEEE Trans. Image Process. | 1 |
| 2015 | A Variational Approach to Simultaneous Image Segmentation and Bias CorrectionabstractThis paper presents a novel variational approach for simultaneous estimation of bias field and segmentation of images with intensity inhomogeneity. We model intensity of inhomogeneous objects to be Gaussian distributed with different means and variances, and then introduce a sliding window to map the original image intensity onto another domain, where the intensity distribution of each object is still Gaussian but can be better separated. The means of the Gaussian distributions in the transformed domain can be adaptively estimated by multiplying the bias field with a piecewise constant signal within the sliding window. A maximum likelihood energy functional is then defined on each local region, which combines the bias field, the membership function of the object region, and the constant approximating the true signal from its corresponding object. The energy functional is then extended to the whole image domain by the Bayesian learning approach. An efficient iterative algorithm is proposed for energy minimization, via which the image segmentation and bias field correction are simultaneously achieved. Furthermore, the smoothness of the obtained optimal bias field is ensured by the normalized convolutions without extra cost. Experiments on real images demonstrated the superiority of the proposed algorithm to other state-of-the-art representative methods. Kaihua Zhang 0001, Qingshan Liu 0001, Huihui Song 0003, Xuelong Li 0001 |
IEEE Trans. Cybern. | 1 |
| 2015 | Improving the Spatial Resolution of Landsat TM/ETM+ Through Fusion With SPOT5 Images via Learning-Based Super-ResolutionabstractTo take advantage of the wide swath width of Landsat Thematic Mapper (TM)/Enhanced Thematic Mapper Plus (ETM+) images and the high spatial resolution of Système Pour l'Observation de la Terre 5 (SPOT5) images, we present a learning-based super-resolution method to fuse these two data types. The fused images are expected to be characterized by the swath width of TM/ETM+ images and the spatial resolution of SPOT5 images. To this end, we first model the imaging process from a SPOT image to a TM/ETM+ image at their corresponding bands, by building an image degradation model via blurring and downsampling operations. With this degradation model, we can generate a simulated Landsat image from each SPOT5 image, thereby avoiding the requirement for geometric coregistration for the two input images. Then, band by band, image fusion can be implemented in two stages: 1) learning a dictionary pair representing the high- and low-resolution details from the given SPOT5 and the simulated TM/ETM+ images; 2) super-resolving the input Landsat images based on the dictionary pair and a sparse coding algorithm. It is noteworthy that the proposed method can also deal with the conventional spatial and spectral fusion of TM/ETM+ and SPOT5 images by using the learned dictionary pairs. To examine the performance of the proposed method of fusing the swath width of TM/ETM+ and the spatial resolution of SPOT5, we illustrate the fusion results on the actual TM images and compare with several classic pansharpening methods by assuming that the corresponding SPOT5 panchromatic image exists. Furthermore, we implement the classification experiments on both actual images and fusion results to demonstrate the benefits of the proposed method for further classification applications. Huihui Song 0003, Bo Huang 0001, Qingshan Liu 0001, Kaihua Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2014 | Fast Visual Tracking via Dense Spatio-temporal Context Learning
Kaihua Zhang 0001, Lei Zhang 0006, Qingshan Liu 0001, David Zhang 0001, Ming-Hsuan Yang 0001 |
ECCV (5) | 1 |
| 2014 | Fast Compressive TrackingabstractIt is a challenging task to develop effective and efficient appearance models for robust object tracking due to factors such as pose variation, illumination change, occlusion, and motion blur. Existing online tracking algorithms often update models with samples from observations in recent frames. Despite much success has been demonstrated, numerous issues remain to be addressed. First, while these adaptive appearance models are data-dependent, there does not exist sufficient amount of data for online algorithms to learn at the outset. Second, online tracking algorithms often encounter the drift problems. As a result of self-taught learning, misaligned samples are likely to be added and degrade the appearance models. In this paper, we propose a simple yet effective and efficient tracking algorithm with an appearance model based on features extracted from a multiscale image feature space with data-independent basis. The proposed appearance model employs non-adaptive random projections that preserve the structure of the image feature space of objects. A very sparse measurement matrix is constructed to efficiently extract the features for the appearance model. We compress sample images of the foreground target and the background using the same sparse measurement matrix. The tracking task is formulated as a binary classification via a naive Bayes classifier with online update in the compressed domain. A coarse-to-fine search strategy is adopted to further reduce the computational complexity in the detection procedure. The proposed compressive tracking algorithm runs in real-time and performs favorably against state-of-the-art methods on challenging sequences in terms of efficiency, accuracy and robustness. Kaihua Zhang 0001, Lei Zhang 0006, Ming-Hsuan Yang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2014 | Shadow Detection and Reconstruction in High-Resolution Satellite Images via Morphological Filtering and Example-Based LearningabstractThe shadows in high-resolution satellite images are usually caused by the constraints of imaging conditions and the existence of high-rise objects, and this is particularly so in urban areas. To alleviate the shadow effects in high-resolution images for their further applications, this paper proposes a novel shadow detection algorithm based on the morphological filtering and a novel shadow reconstruction algorithm based on the example learning method. In the shadow detection stage, an initial shadow mask is generated by the thresholding method, and then, the noise and wrong shadow regions are removed by the morphological filtering method. The shadow reconstruction stage consists of two phases: the example-based learning phase and the inference phase. During the example-based learning phase, the shadow and the corresponding nonshadow pixels are first manually sampled from the study scene, and then, these samples form a shadow library and a nonshadow library, which are correlated by a Markov random field (MRF). During the inference phase, the underlying land-cover pixels are reconstructed from the corresponding shadow pixels by adopting the Bayesian belief propagation algorithm to solve the MRF. Experimental results on QuickBird and WorldView-2 satellite images have demonstrated that the proposed shadow detection algorithm can generate accurate and continuous shadow masks and also that the estimated nonshadow regions from the proposed shadow reconstruction algorithm are highly compatible with their surrounding nonshadow regions. Finally, we examine the effects of the reconstructed image on the application of classification by comparing the classification maps of images before and after shadow reconstruction. Huihui Song 0003, Bo Huang 0001, Kaihua Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2013 | Real-time visual tracking via online weighted multiple instance learning
Kaihua Zhang 0001, Huihui Song 0003 |
Pattern Recognit. | 1 |
| 2013 | Robust Object Tracking Via Active Feature SelectionabstractAdaptive tracking by detection has been widely studied with promising results. The key idea of such trackers is how to train an online discriminative classifier, which can well separate an object from its local background. The classifier is incrementally updated using positive and negative samples extracted from the current frame around the detected object location. However, if the detection is less accurate, the samples are likely to be less accurately extracted, thereby leading to visual drift. Recently, the multiple instance learning (MIL) based tracker has been proposed to solve these problems to some degree. It puts samples into the positive and negative bags, and then selects some features with an online boosting method via maximizing the bag likelihood function. Finally, the selected features are combined for classification. However, in MIL tracker the features are selected by a likelihood function, which can be less informative to tell the target from complex background. Motivated by the active learning method, in this paper we propose an active feature selection approach that is able to select more informative features than the MIL tracker by using the Fisher information criterion to measure the uncertainty of the classification model. More specifically, we propose an online boosting feature selection approach via optimizing the Fisher information criterion, which can yield more robust and efficient real-time object tracking performance. Experimental evaluations on challenging sequences demonstrate the efficiency, accuracy, and robustness of the proposed tracker in comparison with state-of-the-art trackers. Kaihua Zhang 0001, Lei Zhang 0006, Ming-Hsuan Yang 0001, Qinghua Hu |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2013 | Real-Time Object Tracking Via Online Discriminative Feature SelectionabstractMost tracking-by-detection algorithms train discriminative classifiers to separate target objects from their surrounding background. In this setting, noisy samples are likely to be included when they are not properly sampled, thereby causing visual drift. The multiple instance learning (MIL) paradigm has been recently applied to alleviate this problem. However, important prior information of instance labels and the most correct positive instance (i.e., the tracking result in the current frame) can be exploited using a novel formulation much simpler than an MIL approach. In this paper, we show that integrating such prior information into a supervised learning algorithm can handle visual drift more effectively and efficiently than the existing MIL tracker. We present an online discriminative feature selection algorithm that optimizes the objective function in the steepest ascent direction with respect to the positive samples while in the steepest descent direction with respect to the negative ones. Therefore, the trained classifier directly couples its score with the importance of samples, leading to a more robust and efficient tracker. Numerous experimental evaluations with state-of-the-art algorithms on challenging sequences demonstrate the merits of the proposed algorithm. Kaihua Zhang 0001, Lei Zhang 0006, Ming-Hsuan Yang 0001 |
IEEE Trans. Image Process. | 1 |
| 2013 | Reinitialization-Free Level Set Evolution via Reaction DiffusionabstractThis paper presents a novel reaction-diffusion (RD) method for implicit active contours that is completely free of the costly reinitialization procedure in level set evolution (LSE). A diffusion term is introduced into LSE, resulting in an RD-LSE equation, from which a piecewise constant solution can be derived. In order to obtain a stable numerical solution from the RD-based LSE, we propose a two-step splitting method to iteratively solve the RD-LSE equation, where we first iterate the LSE equation, then solve the diffusion equation. The second step regularizes the level set function obtained in the first step to ensure stability, and thus the complex and costly reinitialization procedure is completely eliminated from LSE. By successfully applying diffusion to LSE, the RD-LSE model is stable by means of the simple finite difference method, which is very easy to implement. The proposed RD method can be generalized to solve the LSE for both variational level set method and partial differential equation-based level set method. The RD-LSE method shows very good performance on boundary antileakage. The extensive and promising experimental results on synthetic and real images validate the effectiveness of the proposed RD-LSE approach. Kaihua Zhang 0001, Lei Zhang 0006, Huihui Song 0003, David Zhang 0001 |
IEEE Trans. Image Process. | 1 |
| 2012 | Real-Time Compressive Tracking
Kaihua Zhang 0001, Lei Zhang 0006, Ming-Hsuan Yang 0001 |
ECCV (3) | 1 |
| 2010 | AN adaptive L1-L2 hybrid error model to super-resolutionabstractA hybrid error model with L1and L2norm minimization criteria is proposed in this paper for image/video super-resolution. A membership function is defined to adaptively control the tradeoff between the L1and L2norm terms. Therefore, the proposed hybrid model can have the advantages of both L1norm minimization (i.e. edge preservation) and L2norm minimization (i.e. smoothing noise). In addition, an effective convergence criterion is proposed, which is able to terminate the iterative L1and L2norm minimization process efficiently. Experimental results on images corrupted with various types of noises demonstrate the robustness of the proposed algorithm and its superiority to representative algorithms. Huihui Song 0003, Lei Zhang 0006, Peikang Wang, Kaihua Zhang 0001, Xin Li 0005 |
ICIP | 4 |
| 2010 | A variational multiphase level set approach to simultaneous segmentation and bias correctionabstractThis paper presents a novel level set approach to simultaneous tissue segmentation and bias correction of Magnetic Resonance Imaging (MRI) images. We first model the distribution of intensity belonging to each tissue as a Gaussian distribution with spatially varying mean and variance. Then a sliding window is used to transform the intensity domain to another domain, where the distribution overlap between different tissues is significantly suppressed. A maximum likelihood objective function is defined for each point in the transformed domain, which is then integrated over the entire domain to form a variational level set formulation. Tissue segmentation and bias correction are simultaneously achieved via a level set evolution process. The proposed method is robust to initialization, thereby allowing automatic applications. Experiments on images of various modalities demonstrated the superior performance of the proposed approach over state-of-the-art methods. Kaihua Zhang 0001, Lei Zhang 0006 |
ICIP | 1 |
| 2010 | Active contours with selective local or global segmentation: A new formulation and level set method
Kaihua Zhang 0001, Lei Zhang 0006, Huihui Song 0003, Wengang Zhou 0001 |
Image Vis. Comput. | 1 |
| 2010 | Active contours driven by local image fitting energy
Kaihua Zhang 0001, Huihui Song 0003, Lei Zhang 0006 |
Pattern Recognit. | 1 |