Lihe Zhang

dblp:46/10700 · DBLP profile ↗
← Back
99ranked-venue papers
14as first author
54since 2021 · last 2026
0000-0002-9241-1688ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 68 · 7 first-author · 38 since 2021Graphics, computer vision, multimedia, augmented reality and games · 58 · 7 first-author · 28 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 4 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 PulseMind: A Multi-Modal Medical Model for Real-World Clinical Diagnosis
abstract
Recent advances in medical multi-modal models focus on specialized image analysis like dermatology, pathology, or radiology. However, they do not fully capture the complexity of real-world clinical diagnostics, which involve heterogeneous inputs and require ongoing contextual understanding during patient-physician interactions. To bridge this gap, we introduce PulseMind, a new family of multi-modal diagnostic models that integrates a systematically curated dataset, a comprehensive evaluation benchmark, and a tailored training framework. Specifically, we first construct a diagnostic dataset, MediScope, which comprises 98,000 real-world multi-turn consultations and 601,500 medical images, spanning over 10 major clinical departments and more than 200 sub-specialties. Then, to better reflect the requirements of real-world clinical diagnosis, we develop the PulseMind Benchmark, a multi-turn diagnostic consultation benchmark with a four-dimensional evaluation protocol comprising proactiveness, accuracy, usefulness, and language quality. Finally, we design a training framework tailored for multi-modal clinical diagnostics, centered around a core component named Comparison-based Reinforcement Policy Optimization (CRPO). Compared to absolute score rewards, CRPO uses relative preference signals from multi-dimensional comparisons to provide stable and human-aligned training guidance. Extensive experiments demonstrate that PulseMind achieves competitive performance on both the diagnostic consultation benchmark and public medical benchmarks.
Jiangwei Lao, Qi Zhu 0010, Congyun Jin, Shinan Liu, Zhihong Lu 0002, Lihe Zhang, Jian Wang 0108
AAAI9
2026 EipFormer: Enhancing 3D instance segmentation by emphasizing instance positions
Mengnan Zhao 0001, Lihe Zhang, Yuqiu Kong
Expert Syst. Appl.2
2026 MedSegFM: A Generative Perspective for Lesion Segmentation via Flow Matching
Shijie Chang, Yilong Hu, Lihe Zhang, Zechen Liu, Huchuan Lu
Int. J. Comput. Vis.3
2026 Power Battery Detection
Xiaoqi Zhao 0003, Peiqian Cao, Chenyang Yu, Zonglei Feng, Lihe Zhang, Hanqi Liu, Jiaming Zuo, Youwei Pang, Jinsong Ouyang, Weisi Lin, Georges El Fakhri, Huchuan Lu, Xiaofeng Liu 0001
Int. J. Comput. Vis.5
2026 DecoupleMAD: Boosting sensitivity in multimodal anomaly detection via representation decoupling
Yuan Zhao 0006, Bocen Li, Lihe Zhang, Huchuan Lu
Pattern Recognit.4
2026 LiveMatte: Dynamic Scene Background Restoration and Selective Portrait Patch Enhancement
abstract
Real-time and accurate portrait matting in videos is a challenging problem in computer vision research. Recent approaches have explored incorporating prior conditions for accurate inference. Notably, some methods ask the user to provide the background image, which requires extra effort from the user to capture the background image and is limited to videos with static backgrounds only. We note that real-time video motion segmentation methods often train a background model to detect the foreground. Our insight of this work is that if we can directly restore the background content from the input video as a prior, we may be able to achieve more precise portrait matting. In addition, this approach could potentially work even with dynamic backgrounds, without requiring additional user input. However, automatically restoring the background content is not straightforward due to the difficulty in distinguishing between foreground and background. While it may seem that stationary pixel values represent the background, these values can vary across frames. To this end, we propose a novel dynamic scene background restoration (DSBR) module that learns a background model by accumulating background content from each input video frame. It restores the current background content, which serves as a matting prior for alpha prediction of the subsequent frame. DSBR is extremely lightweight and can be easily integrated into existing matting models. Based on it, we present a real-time portrait video matting framework,LiveMatte. To more efficiently process high-resolution videos, we also introduce a selective portrait patch enhancement (SPPE) module. Extensive experiments and user studies demonstrate that our method is better and faster than existing methods.
Zhanghan Ke, Lihe Zhang, Huchuan Lu, Rynson W. H. Lau
IEEE Trans. Circuits Syst. Video Technol.3
2026 Classification and Calibration: Dual-Guidance Diffusion Model for Mitigating Reconstruction Hallucinations in Multi-Class Anomaly Detection
Yuan Zhao 0006, Xiaoqi Zhao 0003, Lihe Zhang, Huchuan Lu
IEEE Trans. Circuits Syst. Video Technol.4
2025 Unified Medical Lesion Segmentation via Self-referring Indicator
abstract
The recently emerged in-context-learning-based (ICL-based) models have the potential towards the unification of medical lesion segmentation. However, due to their cross-fusion designs, existing ICL-based unified segmentation models fail to accurately localize lesions with low-matched reference sets. Considering that the query itself can be regarded as a high-matched reference, which better indicates the target, we design a self-referencing mechanism that adaptively extracts self-referring indicator vectors from the query based on coarse predictions, thus effectively overcoming the negative impact caused by low-match reference sets. To further facilitate the self-referring mechanism, we introduce reference indicator generation to efficiently extract reference information for coarse predictions instead of using cross-fusion modules, which heavily rely on reference sets. Our designs successfully address the challenges of applying ICL to unified medical lesion segmentation, forming a novel framework named SR-ICL. Our method achieves state-of-the-art results on 8 medical lesion segmentation tasks with only 4 image-mask pairs as reference. Notably, SR-ICL still accomplishes remarkable performance even when using weak reference annotations such as boxes and points, and maintains fixed and low memory consumption even if more tasks are combined. We hope that SR-ICL can provide new insights for the clinical application of medical lesion segmentation.
Shijie Chang, Xiaoqi Zhao 0003, Lihe Zhang
CVPR3
2025 Learning Dynamic Collaborative Network for Semi-supervised 3D Vessel Segmentation
abstract
In this paper, we present a new dynamic collaborative network for semi-supervised 3D vessel segmentation, termed DiCo. Conventional mean teacher (MT) methods typically employ a static approach, where the roles of the teacher and student models are fixed. However, due to the complexity of 3D vessel data, the teacher model may not always outperform the student model, leading to cognitive biases that can limit performance. To address this issue, we propose a dynamic collaborative network that allows the two models to dynamically switch their teacher-student roles. Additionally, we introduce a multi-view integration module to capture various perspectives of the inputs, mirroring the way doctors conduct medical analysis. We also incorporate adversarial supervision to constrain the shape of the segmented vessels in unlabeled data. In this process, the 3D volume is projected into 2D views to mitigate the impact of label inconsistencies. Experiments demonstrate that our DiCo method sets new state-of-the-art performance on three 3D vessel segmentation benchmarks. The code repository address is https://github.com/xujiaommcome/DiCo.
Lihe Zhang
CVPR3
2025 High-Precision Dichotomous Image Segmentation via Probing Diffusion Capacity
abstract
In the realm of high-resolution (HR), fine-grained image segmentation, the primary challenge is balancing broad contextual awareness with the precision required for detailed object delineation, capturing intricate details and the finest edges of objects. Diffusion models, trained on vast datasets comprising billions of image-text pairs, such as SD V2.1, have revolutionized text-to-image synthesis by delivering exceptional quality, fine detail resolution, and strong contextual awareness, making them an attractive solution for high-resolution image segmentation. To this end, we propose DiffDIS, a diffusion-driven segmentation model that taps into the potential of the pre-trained U-Net within diffusion models, specifically designed for high-resolution, fine-grained object segmentation. By leveraging the robust generalization capabilities and rich, versatile image representation prior of the SD models, coupled with a task-specific stable one-step denoising approach, we significantly reduce the inference time while preserving high-fidelity, detailed generation. Additionally, we introduce an auxiliary edge generation task to not only enhance the preservation of fine details of the object boundaries, but reconcile the probabilistic nature of diffusion with the deterministic demands of segmentation. With these refined strategies in place, DiffDIS serves as a rapid object mask generation model, specifically optimized for generating detailed binary maps at high resolutions, while demonstrating impressive accuracy and swift processing. Experiments on the DIS5K dataset demonstrate the superiority of DiffDIS, achieving state-of-the-art results through a streamlined inference process. The source code will be publicly available at \href{https://github.com/qianyu-dlut/DiffDIS}{DiffDIS}.
Qian Yu 0015, Peng-Tao Jiang, Hao Zhang 0063, Jinwei Chen 0003, Bo Li 0115, Lihe Zhang, Huchuan Lu
ICLR6
2025 UniSegDiff: Boosting Unified Lesion Segmentation via a Staged Diffusion Model
Yilong Hu, Shijie Chang, Lihe Zhang, Feng Tian 0001, Weibing Sun, Huchuan Lu
MICCAI (2)3
2025 Rethinking Evaluation of Infrared Small Target Detection
abstract
As an essential vision task, infrared small target detection (IRSTD) has seen significant advancements through deep learning. However, critical limitations in current evaluation protocols impede further progress. First, existing methods rely on fragmented pixel- and target-level specific metrics, which fails to provide a comprehensive view of model capabilities. Second, an excessive emphasis on overall performance scores obscures crucial error analysis, which is vital for identifying failure modes and improving real-world system performance. Third, the field predominantly adopts dataset-specific training-testing paradigms, hindering the understanding of model robustness and generalization across diverse infrared scenarios. This paper addresses these issues by introducing a hybrid-level metric incorporating pixel- and target-level performance, proposing a systematic error analysis method, and emphasizing the importance of cross-dataset evaluation. These aim to offer a more thorough and rational hierarchical analysis framework, ultimately fostering the development of more effective and robust IRSTD models. An open-source toolkit has be released to facilitate standardized benchmarking.
Youwei Pang, Xiaoqi Zhao 0003, Lihe Zhang, Huchuan Lu, Georges El Fakhri, Xiaofeng Liu 0001, Shijian Lu
NeurIPS3
2025 UniMRSeg: Unified Modality-Relax Segmentation via Hierarchical Self-Supervised Compensation
abstract
Multi-modal image segmentation faces real-world deployment challenges from incomplete/corrupted modalities degrading performance. While existing methods address training-inference modality gaps via specialized per-combination models, they introduce high deployment costs by requiring exhaustive model subsets and model-modality matching. In this work, we propose a unified modality-relax segmentation network (UniMRSeg) through hierarchical self-supervised compensation (HSSC). Our approach hierarchically bridges representation gaps between complete and incomplete modalities across input, feature and output levels. First, we adopt modality reconstruction with the hybrid shuffled-masking augmentation, encouraging the model to learn the intrinsic modality characteristics and generate meaningful representations for missing modalities through cross-modal fusion. Next, modality-invariant contrastive learning implicitly compensates the feature space distance among incomplete-complete modality pairs. Furthermore, the proposed lightweight reverse attention adapter explicitly compensates for the weak perceptual semantics in the frozen encoder. Last, UniMRSeg is fine-tuned under the hybrid consistency constraint to ensure stable prediction under all modality combinations without large performance fluctuations. Without bells and whistles, UniMRSeg significantly outperforms the state-of-the-art methods under diverse missing modality scenarios on MRI-based brain tumor segmentation, RGB-D semantic segmentation, RGB-D/T salient object segmentation. The code will be released at \url{https://github.com/Xiaoqi-Zhao-DLUT/UniMRSeg}.
Xiaoqi Zhao 0003, Youwei Pang, Chenyang Yu, Lihe Zhang, Huchuan Lu, Shijian Lu, Georges El Fakhri, Xiaofeng Liu 0001
NeurIPS4
2025 Bidirectional Spatial Semantics Correlation for Referring Image Segmentation
Lihe Zhang, Huchuan Lu
PRCV (5)2
2025 ComPtr: Toward Diverse Bi-Source Dense Prediction Tasks via a Simple Yet General Complementary Transformer
abstract
Deep learning (DL) has advanced the field of dense prediction, while gradually dissolving the inherent barriers between different tasks. However, most existing works focus on designing architectures and constructing visual cues only for the specific task, which ignores the potential uniformity introduced by the DL paradigm. In this paper, we attempt to construct a novel ComPlementary transformer, ComPtr, for diverse bi-source dense prediction tasks. Specifically, unlike existing methods that over-specialize in a single task or a subset of tasks, ComPtr starts from the more general concept of bi-source dense prediction. Based on the basic dependence on information complementarity, we propose consistency enhancement and difference awareness components with which ComPtr can evacuate and collect important visual semantic cues from different image sources for diverse tasks, respectively. ComPtr treats different inputs equally and builds an efficient dense interaction model in the form of sequence-to-sequence on top of the transformer. This task-generic design provides a smooth foundation for constructing the unified model that can simultaneously deal with various bi-source information. In extensive experiments across several representative vision tasks, i.e. remote sensing change detection, RGB-T crowd counting, RGB-D/T salient object detection, and RGB-D semantic segmentation, the proposed method consistently obtains favorable performance.
Youwei Pang, Xiaoqi Zhao 0003, Lihe Zhang, Huchuan Lu
IEEE Trans. Pattern Anal. Mach. Intell.3
2025 Beyond mask: Rethinking guidance types in few-shot segmentation
Shijie Chang, Youwei Pang, Xiaoqi Zhao 0003, Huchuan Lu, Lihe Zhang
Pattern Recognit.5
2025 FocusCLIP: Focusing on Anomaly Regions by Visual-Text Discrepancies
abstract
Few-shot anomaly detection aims to detect defects with only a limited number of normal samples for training. Recent few-shot methods typically focus on object-level features rather than subtle defects within objects, as pretrained models are generally trained on classification or image-text matching datasets. However, object-level features are often insufficient to detect defects, which are characterized by fine-grained texture variations. To address this, we propose FocusCLIP, which consists of a vision-guided branch and a language-guided branch. FocusCLIP leverages the complementary relationship between visual and text modalities to jointly emphasize discrepancies in fine-grained textures of defect regions. Specifically, we design three modules to mine these discrepancies. In the vision-guided branch, we propose the Bidirectional Self-knowledge Distillation (BSD) structure, which identifies anomaly regions through inconsistent representations and accumulates these discrepancies. Within this structure, the Anomaly Capture Module (ACM) is designed to refine features and detect more comprehensive anomalies by leveraging semantic cues from multi-head self-attention. In the language-guided branch, Multi-level Adversarial Class Activation Mapping (MACAM) utilizes foreground-invariant responses to adversarial text prompts, reducing interference from object regions and further focusing on defect regions. Our approach outperforms the state-of-the-art methods in few-shot anomaly detection. Additionally, the language-guided branch within FocusCLIP also demonstrates competitive performance in zero-shot anomaly detection, further validating the effectiveness of our proposed method.
Yuan Zhao 0006, Lihe Zhang, Huchuan Lu
IEEE Trans. Circuits Syst. Video Technol.3
2025 CollabLearn: Propelling Weakly-Supervised Referring Image Segmentation Through Collaboration Between Semantics and Details
abstract
This work presents a weakly supervised referring image segmentation method, namedCollabLearn, that segments objects described by free-form referring expression utilizing solely image-text pairs. Existing methods suffer from incorrect localization of referring expressions due to the lack of high-level semantics in cross-modal alignment or rough segmentation of referenced objects stemming from the absence of low-level details. To address these issues, we propose an innovative framework for generating cross-modal features encompassing both high-level semantics and low-level details via two fusion modules: a semantic awareness module and a detail cognition module. Each of these modules generates an activation map, and they mutually correct each other through a collaborative learning strategy. Specifically, the semantic awareness module performs in-depth cross-modal interaction and achieves accurate localization in a top-down manner. The detail cognition module facilitates the segmentation of entire objects in a bottom-up manner. A collaborative learning strategy is designed to enable interaction between these two modules, enforcing sufficient vision-language alignment. Experiments on three benchmarks demonstrate that CollabLearn consistently outperforms state-of-the-art weakly supervised methods.
Yuqiu Kong, Mengnan Zhao 0001, Lihe Zhang
IEEE Trans. Multim.4
2025 Semantics Alternating Enhancement and Bidirectional Aggregation for Referring Video Object Segmentation
abstract
Referring Video Object Segmentation (RVOS) aims at segmenting out the described object in a video clip according to given expression. The task requires methods to effectively fuse cross-modality features, communicate temporal information, and delineate referent appearance. However, existing solutions bias their focus to mainly mining one or two clues, causing their performance inferior. In this paper, we propose Semantics Alternating Enhancement (SAE) to achieve cross-modality fusion and temporal-spatial semantics mining in an alternate way that makes comprehensive exploit of three cues possible. During each update, SAE will generate a cross-modality and temporal-aware vector that guides vision feature to amplify its referent semantics while filtering out irrelevant contents. In return, the purified feature will provide the contextual soil to produce a more refined guider. Overall, cross-modality interaction and temporal communication are together interleaved into axial semantics enhancement steps. Moreover, we design a simplified SAE by dropping spatial semantics enhancement steps, and employ the variant in the early stages of vision encoder to further enhance usability. To integrate features of different scales, we propose Bidirectional Semantic Aggregation decoder (BSA) to obtain referent mask. The BSA arranges the comprehensively-enhanced features into two groups, and then employs difference awareness step to achieve intra-group feature aggregation bidirectionally and consistency constraint step to realize inter-group integration of semantics-dense and appearance-rich features. Extensive results on challenging benchmarks show that our method performs favorably against the state-of-the-art competitors.
Lihe Zhang, Huchuan Lu
IEEE Trans. Multim.2
2024 Multi-View Aggregation Network for Dichotomous Image Segmentation
abstract
Dichotomous Image Segmentation (DIS) has recently emerged towards high-precision object segmentation from high-resolution natural images. When designing an effective DIS model, the main challenge is how to balance the semantic dispersion of high-resolution targets in the small receptive field and the loss of high-precision details in the large receptive field. Existing methods rely on tedious multiple encoder-decoder streams and stages to gradually complete the global localization and local refinement. Human visual system captures regions of interest by observing them from multiple views. Inspired by it, we model DIS as a multi-view object perception problem and provide a parsi-monious multi-view aggregation network (MVANet), which unifies the feature fusion of the distant view and close-up view into a single stream with one encoder-decoder structure. With the help of the proposed multi-view complementary localization and refinement modules, our approach established long-range, profound visual interactions across multiple views, allowing the features of the detailed close-up view to focus on highly slender structures. Experiments on the popular DIS-5K dataset show that our MVANet significantly outperforms state-of-the-art methods in both accuracy and speed. The source code and datasets will be publicly available at MVANet.
Qian Yu 0015, Xiaoqi Zhao 0003, Youwei Pang, Lihe Zhang, Huchuan Lu
CVPR4
2024 Towards Automatic Power Battery Detection: New Challenge, Benchmark Dataset and Baseline
abstract
We conduct a comprehensive study on a new task named power battery detection (PBD), which aims to localize the dense cathode and anode plates endpoints from X-ray images to evaluate the quality of power batteries. Existing manufacturers usually rely on human eye observation to complete PBD, which makes it difficult to balance the accuracy and efficiency of detection. To address this issue and drive more attention into this meaningful task, we first elaborately collect a dataset, called X-ray PBD, which has 1,500 diverse X-ray images selected from thousands of power batteries of 5 manufacturers, with 7 different visual interference. Then, we propose a novel segmentation-based solution for PBD, termed multi-dimensional collaborative network (MDCNet). With the help of line and counting predictors, the representation of the point segmentation branch can be improved at both semantic and detail aspects. Besides, we design an effective distance-adaptive mask generation strategy, which can alleviate the visual challenge caused by the inconsistent distribution density of plates to provide MDCNet with stable supervision. Without any bells and whistles, our segmentation-based MDCNet consistently outperforms various other corner detection, crowd counting and general/tiny object detection-based so-lutions, making it a strong baseline that can help facilitate future research in PBD. Finally, we share some potential difficulties and works for future researches. The source code and datasets will be publicly available at X-ray PBD.
Xiaoqi Zhao 0003, Youwei Pang, Zhenyu Chen 0001, Qian Yu 0015, Lihe Zhang, Hanqi Liu, Jiaming Zuo, Huchuan Lu
CVPR5
2024 Open-Vocabulary Camouflaged Object Segmentation
Youwei Pang, Xiaoqi Zhao 0003, Jiaming Zuo, Lihe Zhang, Huchuan Lu
ECCV (47)4
2024 Catastrophic Overfitting: A Potential Blessing in Disguise
Mengnan Zhao 0001, Lihe Zhang, Yuqiu Kong
ECCV (42)2
2024 Spider: A Unified Framework for Context-dependent Concept Segmentation
abstract
Different from the context-independent (CI) concepts such as human, car, and airplane, context-dependent (CD) concepts require higher visual understanding ability, such as camouflaged object and medical lesion. Despite the rapid advance of many CD understanding tasks in respective branches, the isolated evolution leads to their limited cross-domain generalisation and repetitive technique innovation. Since there is a strong coupling relationship between foreground and background context in CD tasks, existing methods require to train separate models in their focused domains. This restricts their real-world CD concept understanding towards artificial general intelligence (AGI). We propose a unified model with a single set of parameters, Spider, which only needs to be trained once. With the help of the proposed concept filter driven by the image-mask group prompt, Spider is able to understand and distinguish diverse strong context-dependent concepts to accurately capture the Prompter's intention. Without bells and whistles, Spider significantly outperforms the state-of-the-art specialized models in 8 different context-dependent segmentation tasks, including 4 natural scenes (salient, camouflaged, and transparent objects and shadow) and 4 medical lesions (COVID-19, polyp, breast, and skin lesion with color colonoscopy, CT, ultrasound, and dermoscopy modalities). Besides, Spider shows obvious advantages in continuous learning. It can easily complete the training of new tasks by fine-tuning parameters less than 1% and bring a tolerable performance degradation of less than 5% for all old tasks. The source code will be publicly available at https://github.com/Xiaoqi-Zhao-DLUT/Spider-UniCDSeg.
Xiaoqi Zhao 0003, Youwei Pang, Wei Ji 0011, Baicheng Sheng, Jiaming Zuo, Lihe Zhang, Huchuan Lu
ICML6
2024 Neural image re-exposure
Xinyu Zhang 0017, Hefei Huang, Xu Jia 0012, Dong Wang 0004, Lihe Zhang, Bolun Zheng, Wei Zhou 0021, Huchuan Lu
Comput. Vis. Image Underst.5
2024 Adaptive Multi-Source Predictor for Zero-Shot Video Object Segmentation
Xiaoqi Zhao 0003, Shijie Chang, Youwei Pang, Lihe Zhang, Huchuan Lu
Int. J. Comput. Vis.5
2024 Towards Diverse Binary Segmentation via a Simple yet General Gated Network
Xiaoqi Zhao 0003, Youwei Pang, Lihe Zhang, Huchuan Lu, Lei Zhang 0006
Int. J. Comput. Vis.3
2024 ZoomNeXt: A Unified Collaborative Pyramid Network for Camouflaged Object Detection
abstract
Recent camouflaged object detection (COD) attempts to segment objects visually blended into their surroundings, which is extremely complex and difficult in real-world scenarios. Apart from the high intrinsic similarity between camouflaged objects and their background, objects are usually diverse in scale, fuzzy in appearance, and even severely occluded. To this end, we propose an effective unified collaborative pyramid network that mimics human behavior when observing vague images and videos, i.e., zooming in and out. Specifically, our approach employs the zooming strategy to learn discriminative mixed-scale semantics by the multi-head scale integration and rich granularity perception units, which are designed to fully explore imperceptible clues between candidate objects and background surroundings. The former's intrinsic multi-head aggregation provides more diverse visual patterns. The latter's routing mechanism can effectively propagate inter-frame differences in spatiotemporal scenarios and be adaptively deactivated and output all-zero results for static representations. They provide a solid foundation for realizing a unified architecture for static and dynamic COD. Moreover, considering the uncertainty and ambiguity derived from indistinguishable textures, we construct a simple yet effective regularization, uncertainty awareness loss, to encourage predictions with higher confidence in candidate regions. Our highly task-friendly framework consistently outperforms existing state-of-the-art methods in image and video COD benchmarks.
Youwei Pang, Xiaoqi Zhao 0003, Tian-Zhu Xiang, Lihe Zhang, Huchuan Lu
IEEE Trans. Pattern Anal. Mach. Intell.4
2024 Class correlation correction for unbiased scene graph generation
Mengnan Zhao 0001, Yuqiu Kong, Lihe Zhang
Pattern Recognit.3
2024 Adversarial Attacks on Scene Graph Generation
abstract
Scene graph generation (SGG) effectively improves semantic understanding of the visual world. However, the recent interest of researchers focuses on enhancing SGG in non-adversarial settings, which raises our curiosity about the adversarial robustness of SGG models. To bridge this gap, we perform adversarial attacks on two typical SGG tasks, Scene Graph Detection (SGDet) and Scene Graph Classification (SGCls). Specifically, we initially propose a bounding box relabeling method to reconstruct reasonable attack targets for SGCls. It solves the inconsistency between the specified bounding boxes and the scene graphs selected as attack targets. Subsequently, we introduce a two-step weighted attack by removing the predicted objects and relational triples that affect attack performance, which significantly increases the success rate of adversarial attacks on two SGG tasks. Extensive experiments demonstrate the effectiveness of our methods on five popular SGG models and four adversarial attacks. The Pytorch® implementation can be downloaded from an open-source Github project https://github.com/Dlut-lab-zmn/SGG_Attack.
Mengnan Zhao 0001, Lihe Zhang, Wei Wang 0025, Yuqiu Kong
IEEE Trans. Inf. Forensics Secur.2
2024 GroupMorph: Medical Image Registration via Grouping Network With Contextual Fusion
abstract
Pyramid-based deformation decomposition is a promising registration framework, which gradually decomposes the deformation field into multi-resolution subfields for precise registration. However, most pyramid-based methods directly produce one subfield per resolution level, which does not fully depict the spatial deformation. In this paper, we propose a novel registration model, called GroupMorph. Different from typical pyramid-based methods, we adopt the grouping-combination strategy to predict deformation field at each resolution. Specifically, we perform group-wise correlation calculation to measure the similarities of grouped features. After that, n groups of deformation subfields with different receptive fields are predicted in parallel. By composing these subfields, a deformation field with multi-receptive field ranges is formed, which can effectively identify both large and small deformations. Meanwhile, a contextual fusion module is designed to fuse the contextual features and provide the inter-group information for the field estimator of the next level. By leveraging the inter-group correspondence, the synergy among deformation subfields is enhanced. Extensive experiments on four public datasets demonstrate the effectiveness of GroupMorph. Code is available at https://github.com/TVayne/GroupMorph.
Zuopeng Tan, Lihe Zhang, Yanan Lv, Yili Ma, Huchuan Lu
IEEE Trans. Medical Imaging2
2024 Learning From Box Annotations for Referring Image Segmentation
abstract
Referring image segmentation (RIS) has obtained an impressive achievement by fully convolutional networks (FCNs). However, previous RIS methods require a large number of pixel-level annotations. In this article, we present a weakly supervised RIS method by using bounding box (BB) annotations. In the first stage, we introduce an adversarial boundary loss to extract the object contour from the BB, which is then used to select appropriate region proposals for pseudoground-truth (PGT) generation. In the second stage, we design a co-training (Co-T) strategy to purify the pseudolabels. Specifically, we train two networks and interactively guide them to pick clean labels for each other's networks, which can weaken the effect of noisy labels on model training. Experiment results on four benchmark datasets demonstrate that the proposed method can produce high-quality masks with a speed of 63 frames/s.
Lihe Zhang, Zhiwei Hu, Huchuan Lu
IEEE Trans. Neural Networks Learn. Syst.2
2024 Referring Image Segmentation With Fine-Grained Semantic Funneling Infusion
abstract
Recently, referring image segmentation has attracted wide attention given its huge potential in human-robot interaction. Networks to identify the referred region must have a deep understanding of both the image and language semantics. To do so, existing works tend to design various mechanisms to achieve cross-modality fusion, for example, tile and concatenation and vanilla nonlocal manipulation. However, the plain fusion usually is either coarse or constrained by the exorbitant computation overhead, finally causing not enough understanding of the referent. In this work, we propose a fine-grained semantic funneling infusion (FSFI) mechanism to solve the problem. The FSFI introduces a constant spatial constraint on the querying entities from different encoding stages and dynamically infuses the gleaned language semantic into the vision branch. Moreover, it decomposes the features from different modalities into more delicate components, allowing the fusion to happen in multiple low-dimensional spaces. The fusion is more effective than the one only happening in one high-dimensional space, given its ability to sink more representative information along the channel dimension. Another problem haunting the task is that the instilling of high-abstract semantic will blur the details of the referent. Targetedly, we propose a multiscale attention-enhanced decoder (MAED) to alleviate the problem. We design a detail enhancement operator (DeEh) and apply it in a multiscale and progressive way. Features from the higher level are used to generate attention guidance to enlighten the lower-level features to more attend to the detail regions. Extensive results on the challenging benchmarks show that our network performs favorably against the state-of-the-arts (SOTAs).
Lihe Zhang, Huchuan Lu
IEEE Trans. Neural Networks Learn. Syst.2
2023 Referring Image Segmentation Using Text Supervision
abstract
Existing Referring Image Segmentation (RIS) methods typically require expensive pixel-level or box-level annotations for supervision. In this paper, we observe that the referring texts used in RIS already provide sufficient information to localize the target object. Hence, we propose a novel weakly-supervised RIS framework to formulate the target localization problem as a classification process to differentiate between positive and negative text expressions. While the referring text expressions for an image are used as positive expressions, the referring text expressions from other images can be used as negative expressions for this image. Our framework has three main novelties. First, we propose a bilateral prompt method to facilitate the classification process, by harmonizing the domain discrepancy between visual and linguistic features. Second, we propose a calibration method to reduce noisy background information and improve the correctness of the response maps for target object localization. Third, we propose a positive response map selection strategy to generate high-quality pseudo-labels from the enhanced response maps, for training a segmentation network for RIS inference. For evaluation, we propose a new metric to measure localization accuracy. Experiments on four benchmarks show that our framework achieves promising performances to existing fully-supervised RIS methods while outperforming state-of-the-art weakly-supervised methods adapted from related areas. Code is available at https://github.com/fawnliu/TRIS.
Fang Liu 0033, Yuhao Liu 0001, Yuqiu Kong, Ke Xu 0010, Lihe Zhang, Gerhard P. Hancke 0002, Rynson W. H. Lau
ICCV5
2023 Adaptive Illumination Mapping for Shadow Detection in Raw Images
abstract
Shadow detection methods rely on multi-scale contrast, especially global contrast, information to locate shadows correctly. However, we observe that the camera image signal processor (ISP) tends to preserve more local contrast information by sacrificing global contrast information during the raw-to-sRGB conversion process. This often causes existing methods to fail in scenes with high global contrast but low local contrast in shadow regions. In this paper, we propose a novel method to detect shadows from raw images. Our key idea is that instead of performing a many-to-one mapping like the ISP process, we can learn a many-to-many mapping from the high dynamic range raw images to the sRGB images of different illumination, which is able to preserve multi-scale contrast for accurate shadow detection. To this end, we first construct a new shadow dataset with ~ 7000 raw images and shadow masks. We then propose a novel network, which includes a novel adaptive illumination mapping (AIM) module to project the input raw images into sRGB images of different intensity ranges and a shadow detection module to leverage the preserved multi-scale contrast information to detect shadows. To learn the shadow-aware adaptive illumination mapping process, we propose a novel feedback mechanism to guide the AIM during training. Experiments show that our method outperforms state- of-the-art shadow detectors. Code and dataset are available at https://github.com/jiayusun/SARA.
Ke Xu 0010, Youwei Pang, Lihe Zhang, Huchuan Lu, Gerhard P. Hancke 0002, Rynson W. H. Lau
ICCV4
2023 Fast Adversarial Training with Smooth Convergence
abstract
Fast adversarial training (FAT) is beneficial for improving the adversarial robustness of neural networks. However, previous FAT work has encountered a significant issue known as catastrophic overfitting when dealing with large perturbation budgets, i.e. the adversarial robustness of models declines to near zero during training. To address this, we analyze the training process of prior FAT work and observe that catastrophic overfitting is accompanied by the appearance of loss convergence outliers. Therefore, we argue a moderately smooth loss convergence process will be a stable FAT process that solves catastrophic overfitting. To obtain a smooth loss convergence process, we propose a novel oscillatory constraint (dubbed ConvergeSmooth) to limit the loss difference between adjacent epochs. The convergence stride of ConvergeSmooth is introduced to balance convergence and smoothing. Likewise, we design weight centralization without introducing additional hyperparameters other than the loss balance coefficient. Our proposed methods are attack-agnostic and thus can improve the training stability of various FAT techniques. Extensive experiments on popular datasets show that the proposed methods efficiently avoid catastrophic overfitting and outperform all previous FAT methods. Code is available at https://github.com/FAT-CS/ConvergeSmooth.
Mengnan Zhao 0001, Lihe Zhang, Yuqiu Kong
ICCV2
2023 Progressively Coupling Network for Brain MRI Registration in Few-Shot Situation
Zuopeng Tan, Feng Tian 0001, Lihe Zhang, Weibing Sun, Huchuan Lu
MICCAI (10)4
2023 Only Classification Head Is Sufficient for Medical Image Segmentation
Hongbin Wei, Zhiwei Hu, Zhilong Ji, Hongpeng Jia, Lihe Zhang, Huchuan Lu
PRCV (13)6
2023 Growth Simulation Network for Polyp Segmentation
Hongbin Wei, Xiaoqi Zhao 0003, Long Lv, Lihe Zhang, Weibing Sun, Huchuan Lu
PRCV (13)4
2023 Temporal knowledge graph reasoning triggered by memories
Mengnan Zhao 0001, Lihe Zhang, Yuqiu Kong
Appl. Intell.2
2023 Local-global coordination with transformers for referring image segmentation
Fang Liu 0033, Yuqiu Kong, Lihe Zhang
Neurocomputing3
2023 Referring Segmentation via Encoder-Fused Cross-Modal Attention Network
abstract
This paper focuses on referring segmentation, which aims to selectively segment the corresponding visual region in an image (or video) according to the referring expression. However, the existing methods usually consider the interaction between multi-modal features at the decoding end of the network. Specifically, they interact the visual features of each scale with language respectively, thus ignoring the correlation between multi-scale features. In this work, we present an encoder fusion network (EFN), which transfers the multi-modal feature learning process from the decoding end to the encoding end and realizes the gradual refinement of multi-modal features by the language. In EFN, we also adopt a co-attention mechanism to promote the mutual alignment of language and visual information in feature space. In the decoding stage, a boundary enhancement module (BEM) is proposed to enhance the network's attention to the details of the target. For video data, we introduce an asymmetric cross-frame attention module (ACFM) to effectively capture the temporal information from the video frames by computing the relationship between each pixel of the current frame and each pooled sub-region of the reference frames. Extensive experiments on referring image/video segmentation datasets show that our method outperforms the state-of-the-art performance.
Lihe Zhang, Zhiwei Hu, Huchuan Lu
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 Lane Detection with Versatile AtrousFormer and Local Semantic Guidance
Lihe Zhang, Huchuan Lu
Pattern Recognit.2
2023 CAVER: Cross-Modal View-Mixed Transformer for Bi-Modal Salient Object Detection
abstract
Most of the existing bi-modal (RGB-D and RGB-T) salient object detection methods utilize the convolution operation and construct complex interweave fusion structures to achieve cross-modal information integration. The inherent local connectivity of the convolution operation constrains the performance of the convolution-based methods to a ceiling. In this work, we rethink these tasks from the perspective of global information alignment and transformation. Specifically, the proposed cross-modal view-mixed transformer (CAVER) cascades several cross-modal integration units to construct a top-down transformer-based information propagation path. CAVER treats the multi-scale and multi-modal feature integration as a sequence-to-sequence context propagation and update process built on a novel view-mixed attention mechanism. Besides, considering the quadratic complexity w.r.t. the number of input tokens, we design a parameter-free patch-wise token re-embedding strategy to simplify operations. Extensive experimental results on RGB-D and RGB-T SOD datasets demonstrate that such a simple two-stream encoder-decoder framework can surpass recent state-of-the-art methods when it is equipped with the proposed components.
Youwei Pang, Xiaoqi Zhao 0003, Lihe Zhang, Huchuan Lu
IEEE Trans. Image Process.3
2023 Bidirectional Relationship Inferring Network for Referring Image Localization and Segmentation
abstract
Recently, referring image localization and segmentation has aroused widespread interest. However, the existing methods lack a clear description of the interdependence between language and vision. To this end, we present a bidirectional relationship inferring network (BRINet) to effectively address the challenging tasks. Specifically, we first employ a vision-guided linguistic attention module to perceive the keywords corresponding to each image region. Then, language-guided visual attention adopts the learned adaptive language to guide the update of the visual features. Together, they form a bidirectional cross-modal attention module (BCAM) to achieve the mutual guidance between language and vision. They can help the network align the cross-modal features better. Based on the vanilla language-guided visual attention, we further design an asymmetric language-guided visual attention, which significantly reduces the computational cost by modeling the relationship between each pixel and each pooled subregion. In addition, a segmentation-guided bottom-up augmentation module (SBAM) is utilized to selectively combine multilevel information flow for object localization. Experiments show that our method outperforms other state-of-the-art methods on three referring image localization datasets and four referring image segmentation datasets.
Zhiwei Hu, Lihe Zhang, Huchuan Lu
IEEE Trans. Neural Networks Learn. Syst.3
2022 Self-Supervised Pretraining for RGB-D Salient Object Detection
abstract
Existing CNNs-Based RGB-D salient object detection (SOD) networks are all required to be pretrained on the ImageNet to learn the hierarchy features which helps provide a good initialization. However, the collection and annotation of large-scale datasets are time-consuming and expensive. In this paper, we utilize self-supervised representation learning (SSL) to design two pretext tasks: the cross-modal auto-encoder and the depth-contour estimation. Our pretext tasks require only a few and unlabeled RGB-D datasets to perform pretraining, which makes the network capture rich semantic contexts and reduce the gap between two modalities, thereby providing an effective initialization for the downstream task. In addition, for the inherent problem of cross-modal fusion in RGB-D SOD, we propose a consistency-difference aggregation (CDA) module that splits a single feature fusion into multi-path fusion to achieve an adequate perception of consistent and differential information. The CDA module is general and suitable for cross-modal and cross-level feature fusion. Extensive experiments on six benchmark datasets show that our self-supervised pretrained model performs favorably against most state-of-the-art methods pretrained on ImageNet. The source code will be publicly available at https://github.com/Xiaoqi-Zhao-DLUT/SSLSOD.
Xiaoqi Zhao 0003, Youwei Pang, Lihe Zhang, Huchuan Lu, Xiang Ruan
AAAI3
2022 Semantics-Adding Flaw-Erasing Network for Semantic Human Matting
Zhanghan Ke, Ke Xu 0010, Fan Shao, Lihe Zhang, Huchuan Lu, Rynson W. H. Lau
BMVC5
2022 Zoom In and Out: A Mixed-scale Triplet Network for Camouflaged Object Detection
abstract
The recently proposed camouflaged object detection (COD) attempts to segment objects that are visually blended into their surroundings, which is extremely complex and difficult in real-world scenarios. Apart from high intrinsic similarity between the camouflaged objects and their background, the objects are usually diverse in scale, fuzzy in appearance, and even severely occluded. To deal with these problems, we propose a mixed-scale triplet network, Zoom- Net, which mimics the behavior of humans when observing vague images, i.e., zooming in and out. Specifically, our ZoomNet employs the zoom strategy to learn the discriminative mixed-scale semantics by the designed scale integration unit and hierarchical mixed-scale unit, which fully explores imperceptible clues between the candidate objects and background surroundings. Moreover, considering the uncertainty and ambiguity derived from indistinguishable textures, we construct a simple yet effective regularization constraint, uncertainty-aware loss, to promote the model to accurately produce predictions with higher confidence in candidate regions. Without bells and whistles, our proposed highly task-friendly model consistently surpasses the existing 23 state-of-the-art methods on four public datasets. Besides, the superior performance over the recent cutting-edge models on the SOD task also verifies the effectiveness and generality of our model. The code will be available at https://github.com/lartpang/ZoomNet.
Youwei Pang, Xiaoqi Zhao 0003, Tian-Zhu Xiang, Lihe Zhang, Huchuan Lu
CVPR4
2022 Learning to Detect Salient Object With Multi-Source Weak Supervision
abstract
High-cost pixel-level annotations makes it appealing to train saliency detection models with weak supervision. However, a single weak supervision source hardly contain enough information to train a well-performing model. To this end, we introduce a unified two-stage framework to learn from category labels, captions, web images and unlabeled images. In the first stage, we design a classification network (CNet) and a caption generation network (PNet), which learn to predict object categories and generate captions, respectively, meanwhile highlights the potential foreground regions. We present an attention transfer loss to transmit supervisions between two tasks and an attention coherence loss to encourage the networks to detect generally salient regions instead of task-specific regions. In the second stage, we create two complementary training datasets using CNet and PNet, i.e., natural image dataset with noisy labels for adapting saliency prediction network (SNet) to natural image input, and synthesized image dataset by pasting objects on background images for providing SNet with accurate ground-truth. During the testing phases, we only need SNet to predict saliency maps. Experiments indicate the performance of our method compares favorably against unsupervised, weakly supervised methods and even some supervised methods.
Hongshuang Zhang, Yu Zeng 0001, Huchuan Lu, Lihe Zhang, Jinqing Qi
IEEE Trans. Pattern Anal. Mach. Intell.4
2022 Encoder deep interleaved network with multi-scale aggregation for RGB-D salient object detection
Jinyu Meng, Lihe Zhang, Huchuan Lu
Pattern Recognit.3
2022 Joint Learning of Salient Object Detection, Depth Estimation and Contour Extraction
abstract
Benefiting from color independence, illumination invariance and location discrimination attributed by the depth map, it can provide important supplemental information for extracting salient objects in complex environments. However, high-quality depth sensors are expensive and can not be widely applied. While general depth sensors produce the noisy and sparse depth information, which brings the depth-based networks with irreversible interference. In this paper, we propose a novel multi-task and multi-modal filtered transformer (MMFT) network for RGB-D salient object detection (SOD). Specifically, we unify three complementary tasks: depth estimation, salient object detection and contour estimation. The multi-task mechanism promotes the model to learn the task-aware features from the auxiliary tasks. In this way, the depth information can be completed and purified. Moreover, we introduce a multi-modal filtered transformer (MFT) module, which equips with three modality-specific filters to generate the transformer-enhanced feature for each modality. The proposed model works in a depth-free style during the testing phase. Experiments show that it not only significantly surpasses the depth-based RGB-D SOD methods on multiple datasets, but also precisely predicts a high-quality depth map and salient contour at the same time. And, the resulted depth map can help existing RGB-D SOD methods obtain significant performance gain.
Xiaoqi Zhao 0003, Youwei Pang, Lihe Zhang, Huchuan Lu
IEEE Trans. Image Process.3
2021 Encoder Fusion Network With Co-Attention Embedding for Referring Image Segmentation
abstract
Recently, referring image segmentation has aroused widespread interest. Previous methods perform the multi-modal fusion between language and vision at the decoding side of the network. And, linguistic feature interacts with visual feature of each scale separately, which ignores the continuous guidance of language to multi-scale visual features. In this work, we propose an encoder fusion network (EFN), which transforms the visual encoder into a multi-modal feature learning network, and uses language to refine the multi-modal features progressively. Moreover, a co-attention mechanism is embedded in the EFN to realize the parallel update of multi-modal features, which can promote the consistent of the cross-modal information representation in the semantic space. Finally, we propose a boundary enhancement module (BEM) to make the network pay more attention to the fine structure. The experiment results on four benchmark datasets demonstrate that the proposed approach achieves the state-of-the-art performance under different evaluation metrics without any post-processing.
Zhiwei Hu, Lihe Zhang, Huchuan Lu
CVPR3
2021 Automatic Polyp Segmentation via Multi-scale Subtraction Network
Xiaoqi Zhao 0003, Lihe Zhang, Huchuan Lu
MICCAI (1)2
2021 Multi-Source Fusion and Automatic Predictor Selection for Zero-Shot Video Object Segmentation
abstract
Location and appearance are the key cues for video object segmentation. Many sources such as RGB, depth, optical flow and static saliency can provide useful information about the objects. However, existing approaches only utilize the RGB or RGB and optical flow. In this paper, we propose a novel multi-source fusion network for zero-shot video object segmentation. With the help of interoceptive spatial attention module (ISAM), spatial importance of each source is highlighted. Furthermore, we design a feature purification module (FPM) to filter the inter-source incompatible features. By the ISAM and FPM, the multi-source features are effectively fused. In addition, we put forward an automatic predictor selection network (APS) to select the better prediction of either the static saliency predictor or the moving object predictor in order to prevent over-reliance on the failed results caused by low-quality optical flow maps. Extensive experiments on three challenging public benchmarks (i.e. DAVIS$_16 $, Youtube-Objects and FBMS) show that the proposed model achieves compelling performance against the state-of-the-arts. The source code will be publicly available at https://github.com/Xiaoqi-Zhao-DLUT/Multi-Source-APS-ZVOS
Xiaoqi Zhao 0003, Youwei Pang, Lihe Zhang, Huchuan Lu
ACM Multimedia4
2020 Bi-Directional Relationship Inferring Network for Referring Image Segmentation
abstract
Most existing methods do not explicitly formulate the mutual guidance between vision and language. In this work, we propose a bi-directional relationship inferring network (BRINet) to model the dependencies of cross-modal information. In detail, the vision-guided linguistic attention is used to learn the adaptive linguistic context corresponding to each visual region. Combining with the language-guided visual attention, a bi-directional cross-modal attention module (BCAM) is built to learn the relationship between multi-modal features. Thus, the ultimate semantic context of the target object and referring expression can be represented accurately and consistently. Moreover, a gated bi-directional fusion module (GBFM) is designed to integrate the multi-level features where a gate function is used to guide the bi-directional flow of multi-level information. Extensive experiments on four benchmark datasets demonstrate that the proposed method outperforms other state-of-the-art methods under different evaluation metrics.
Zhiwei Hu, Lihe Zhang, Huchuan Lu
CVPR4
2020 Multi-Scale Interactive Network for Salient Object Detection
abstract
Deep-learning based salient object detection methods achieve great progress. However, the variable scale and unknown category of salient objects are great challenges all the time. These are closely related to the utilization of multi-level and multi-scale features. In this paper, we propose the aggregate interaction modules to integrate the features from adjacent levels, in which less noise is introduced because of only using small up-/down-sampling rates. To obtain more efficient multi-scale features from the integrated features, the self-interaction modules are embedded in each decoder unit. Besides, the class imbalance issue caused by the scale variation weakens the effect of the binary cross entropy loss and results in the spatial inconsistency of the predictions. Therefore, we exploit the consistency-enhanced loss to highlight the fore-/back-ground difference and preserve the intra-class consistency. Experimental results on five benchmark datasets demonstrate that the proposed method without any post-processing performs favorably against 23 state-of-the-art approaches. The source code will be publicly available at https://github.com/lartpang/MINet.
Youwei Pang, Xiaoqi Zhao 0003, Lihe Zhang, Huchuan Lu
CVPR3
2020 Hierarchical Dynamic Filtering Network for RGB-D Salient Object Detection
Youwei Pang, Lihe Zhang, Xiaoqi Zhao 0003, Huchuan Lu
ECCV (25)2
2020 Suppress and Balance: A Simple Gated Network for Salient Object Detection
Xiaoqi Zhao 0003, Youwei Pang, Lihe Zhang, Huchuan Lu, Lei Zhang 0006
ECCV (2)3
2020 A Single Stream Network for Robust and Real-Time RGB-D Salient Object Detection
Xiaoqi Zhao 0003, Lihe Zhang, Youwei Pang, Huchuan Lu, Lei Zhang 0006
ECCV (22)2
2020 CACNet: Salient object detection via context aggregation and contrast embedding
Hongguang Bo, Lihe Zhang, Huchuan Lu
Neurocomputing4
2020 Salient object detection via double random walks with dual restarts
Lihe Zhang, Huchuan Lu, Guohua Wei
Image Vis. Comput.3
2020 Visual Saliency Detection via Kernelized Subspace Ranking With Active Learning
abstract
Saliency detection task has witnessed a booming interest for years, due to the growth of the computer vision community. In this paper, we introduce a new saliency model that performs active learning with kernelized subspace ranker (KSR) referred to as KSR-AL. This pool-based active learning algorithm ranks the informativeness of unlabeled data by considering both uncertainty sampling and information density, thereby minimizing the cost of labeling. The informative images are selected to train the KSR iteratively and incrementally. The learning model of this algorithm is designed on object-level proposals and region-based convolutional neural network (R-CNN) features, by jointly learning a Rank-SVM classifier and a subspace projection. When the active learning process meets its stopping criteria, the saliency map of each image is generated by a weight fusion of its top-ranked proposals, whose ranking scores are graded by the learned ranker. We show that the KSR-AL achieves a reduction in annotation, as well as improvement in performance, compared with the supervised learning scheme. Besides, the proposed algorithm also outperforms the state-of-the-art methods. These improvements are demonstrated by extensive experiments on six publicly available benchmark datasets.
Lihe Zhang, Tiantian Wang 0002, Yifan Min, Huchuan Lu
IEEE Trans. Image Process.1
2020 A Multistage Refinement Network for Salient Object Detection
abstract
Deep convolutional neural networks (CNNs) have been successfully applied to a wide variety of problems in computer vision, including salient object detection. To accurately detect and segment salient objects, it is necessary to extract and combine high-level semantic features with low-level fine details simultaneously. This is challenging for CNNs because repeated subsampling operations such as pooling and convolution lead to a significant decrease in the feature resolution, which results in the loss of spatial details and finer structures. Therefore, we propose augmenting feedforward neural networks by using the multistage refinement mechanism. In the first stage, a master net is built to generate a coarse prediction map in which most detailed structures are missing. In the following stages, the refinement net with layerwise recurrent connections to the master net is equipped to progressively combine local context information across stages to refine the preceding saliency maps in a stagewise manner. Furthermore, the pyramid pooling module and channel attention module are applied to aggregate different-region-based global contexts. Extensive evaluations over six benchmark datasets show that the proposed method performs favorably against the state-of-the-art approaches.
Lihe Zhang, Jie Wu 0028, Tiantian Wang 0002, Ali Borji, Guohua Wei, Huchuan Lu
IEEE Trans. Image Process.1
2019 Multi-Source Weak Supervision for Saliency Detection
abstract
The high cost of pixel-level annotations makes it appealing to train saliency detection models with weak supervision. However, a single weak supervision source usually does not contain enough information to train a well-performing model. To this end, we propose a unified framework to train saliency detection models with diverse weak supervision sources. In this paper, we use category labels, captions, and unlabelled data for training, yet other supervision sources can also be plugged into this flexible framework. We design a classification network (CNet) and a caption generation network (PNet), which learn to predict object categories and generate captions, respectively, meanwhile highlight the most important regions for corresponding tasks. An attention transfer loss is designed to transmit supervision signal between networks, such that the network designed to be trained with one supervision source can benefit from another. An attention coherence loss is defined on unlabelled data to encourage the networks to detect generally salient regions instead of task-specific regions. We use CNet and PNet to generate pixel-level pseudo labels to train a saliency prediction network (SNet). During the testing phases, we only need SNet to predict saliency maps. Experiments demonstrate the performance of our method compares favourably against unsupervised and weakly supervised methods and even some supervised methods.
Yu Zeng 0001, Yunzhi Zhuge, Huchuan Lu, Lihe Zhang, Mingyang Qian, Yizhou Yu
CVPR4
2019 Deep Learning for Light Field Saliency Detection
abstract
Recent research in 4D saliency detection is limited by the deficiency of a large-scale 4D light field dataset. To address this, we introduce a new dataset to assist the subsequent research in 4D light field saliency detection. To the best of our knowledge, this is to date the largest light field dataset in which the dataset provides 1465 all-focus images with human-labeled ground truth masks and the corresponding focal stacks for every light field image. To verify the effectiveness of the light field data, we first introduce a fusion framework which includes two CNN streams where the focal stacks and all-focus images serve as the input. The focal stack stream utilizes a recurrent attention mechanism to adaptively learn to integrate every slice in the focal stack, which benefits from the extracted features of the good slices. Then it is incorporated with the output map generated by the all-focus stream to make the saliency prediction. In addition, we introduce adversarial examples by adding noise intentionally into images to help train the deep network, which can improve the robustness of the proposed network. The noise is designed by users, which is imperceptible but can fool the CNNs to make the wrong prediction. Extensive experiments show the effectiveness and superiority of the proposed model on the popular evaluation metrics. The proposed method performs favorably compared with the existing 2D, 3D and 4D saliency detection methods on the proposed dataset and existing LFSD light field dataset. The code and results can be found at https://github.com/OIPLab-DUT/ ICCV2019_Deeplightfield_Saliency. Moreover, to facilitate research in this field, all images we collected are shared in a ready-to-use manner.
Tiantian Wang 0002, Yongri Piao, Huchuan Lu, Lihe Zhang
ICCV5
2019 Joint Learning of Saliency Detection and Weakly Supervised Semantic Segmentation
abstract
Existing weakly supervised semantic segmentation (WSSS) methods usually utilize the results of pre-trained saliency detection (SD) models without explicitly modelling the connections between the two tasks, which is not the most efficient configuration. Here we propose a unified multi-task learning framework to jointly solve WSSS and SD using a single network, i.e. saliency and segmentation network (SSNet). SSNet consists of a segmentation network (SN) and a saliency aggregation module (SAM). For an input image, SN generates the segmentation result and, SAM predicts the saliency of each category and aggregating the segmentation masks of all categories into a saliency map. The proposed network is trained end-to-end with image-level category labels and class-agnostic pixel-level saliency labels. Experiments on PASCAL VOC 2012 segmentation dataset and four saliency benchmark datasets show the performance of our method compares favorably against state-of-the-art weakly supervised segmentation methods and fully supervised saliency detection methods.
Yu Zeng 0001, Yunzhi Zhuge, Huchuan Lu, Lihe Zhang
ICCV4
2019 Salient object detection by local and global manifold regularized SVM model
Lihe Zhang, Guohua Wei, Hongguang Bo
Neurocomputing1
2019 Language-aware weak supervision for salient object detection
Mingyang Qian, Jinqing Qi, Lihe Zhang, Mengyang Feng, Huchuan Lu
Pattern Recognit.3
2019 Edge-Aware Convolution Neural Network Based Salient Object Detection
abstract
Salient object detection has received great amount of attention in recent years. In this letter, we propose a novel salient object detection algorithm, which combines the global contextual information along with the low-level edge features. First, we train an edge detection stream based on the state-of-the-art holistically-nested edge detection (HED) model and extract hierarchical boundary information from each VGG block. Then, the edge contours are served as the complementary edge-aware information and integrated with the saliency detection stream to depict continuous boundary for salient objects. Finally, we combine pyramid pooling modules with auxiliary side output supervision to form the multi-scale pyramid-based supervision module, providing multi-scale global contextual information for the saliency detection network. Compared with the previous methods, the proposed network contains more explicit edge-aware features and exploit the multi-scale global information more effectively. Experiments demonstrate the effectiveness of the proposed method, which achieves the state-of-the-art performance on five popular benchmarks.
Wenlong Guan, Tiantian Wang 0002, Jinqing Qi, Lihe Zhang, Huchuan Lu
IEEE Signal Process. Lett.4
2018 Detect Globally, Refine Locally: A Novel Approach to Saliency Detection
abstract
Effective integration of contextual information is crucial for salient object detection. To achieve this, most existing methods based on 'skip' architecture mainly focus on how to integrate hierarchical features of Convolutional Neural Networks (CNNs). They simply apply concatenation or element-wise operation to incorporate high-level semantic cues and low-level detailed information. However, this can degrade the quality of predictions because cluttered and noisy information can also be passed through. To address this problem, we proposes a global Recurrent Localization Network (RLN) which exploits contextual information by the weighted response map in order to localize salient objects more accurately. Particularly, a recurrent module is employed to progressively refine the inner structure of the CNN over multiple time steps. Moreover, to effectively recover object boundaries, we propose a local Boundary Refinement Network (BRN) to adaptively learn the local contextual information for each spatial position. The learned propagation coefficients can be used to optimally capture relations between each pixel and its neighbors. Experiments on five challenging datasets show that our approach performs favorably against all existing methods in terms of the popular evaluation metrics.
Tiantian Wang 0002, Lihe Zhang, Huchuan Lu, Gang Yang 0002, Xiang Ruan, Ali Borji
CVPR2
2018 Learning to Promote Saliency Detectors
abstract
The categories and appearance of salient objects vary from image to image, therefore, saliency detection is an image-specific task. Due to lack of large-scale saliency training data, using deep neural networks (DNNs) with pretraining is difficult to precisely capture the image-specific saliency cues. To solve this issue, we formulate a zero-shot learning problem to promote existing saliency detectors. Concretely, a DNN is trained as an embedding function to map pixels and the attributes of the salient/background regions of an image into the same metric space, in which an image-specific classifier is learned to classify the pixels. Since the image-specific task is performed by the classifier, the DNN embedding effectively plays the role of a general feature extractor. Compared with transferring the learning to a new recognition task using limited data, this formulation makes the DNN learn more effectively from small data. Extensive experiments on five data sets show that our method significantly improves accuracy of existing methods and compares favorably against state-of-the-art approaches.
Yu Zeng 0001, Huchuan Lu, Lihe Zhang, Mengyang Feng, Ali Borji
CVPR3
2018 Deep multi-level networks with multi-task learning for saliency detection
Lihe Zhang, Hongguang Bo, Tiantian Wang 0002, Huchuan Lu
Neurocomputing1
2018 Salient object detection via proposal selection
Lihe Zhang
Neurocomputing1
2018 Saliency Detection via Absorbing Markov Chain With Learnt Transition Probability
abstract
In this paper, we propose a bottom-up saliency model based on absorbing Markov chain (AMC). First, a sparsely connected graph is constructed to capture the local context information of each node. All image boundary nodes and other nodes are, respectively, treated as the absorbing nodes and transient nodes in the absorbing Markov chain. Then, the expected number of times from each transient node to all other transient nodes can be used to represent the saliency value of this node. The absorbed time depends on the weights on the path and their spatial coordinates, which are completely encoded in the transition probability matrix. Considering the importance of this matrix, we adopt different hierarchies of deep features extracted from fully convolutional networks and learn a transition probability matrix, which is called learnt transition probability matrix. Although the performance is significantly promoted, salient objects are not uniformly highlighted very well. To solve this problem, an angular embedding technique is investigated to refine the saliency results. Based on pairwise local orderings, which are produced by the saliency maps of AMC and boundary maps, we rearrange the global orderings (saliency value) of all nodes. Extensive experiments demonstrate that the proposed algorithm outperforms the state-of-the-art methods on six publicly available benchmark data sets.
Lihe Zhang, Jianwu Ai, Huchuan Lu, Xiukui Li
IEEE Trans. Image Process.1
2017 A Stagewise Refinement Model for Detecting Salient Objects in Images
Tiantian Wang 0002, Ali Borji, Lihe Zhang, Huchuan Lu
ICCV3
2017 Ranking Saliency
abstract
Most existing bottom-up algorithms measure the foreground saliency of a pixel or region based on its contrast within a local context or the entire image, whereas a few methods focus on segmenting out background regions and thereby salient objects. Instead of only considering the contrast between salient objects and their surrounding regions, we consider both foreground and background cues in this work. We rank the similarity of image elements with foreground or background cues via graph-based manifold ranking. The saliency of image elements is defined based on their relevances to the given seeds or queries. We represent an image as a multi-scale graph with fine superpixels and coarse regions as nodes. These nodes are ranked based on the similarity to background and foreground queries using affinity matrices. Saliency detection is carried out in a cascade scheme to extract background regions and foreground salient objects efficiently. Experimental results demonstrate the proposed method performs well against the state-of-the-art methods in terms of accuracy and speed. We also propose a new benchmark dataset containing 5,168 images for large-scale performance evaluation of saliency detection methods.
Lihe Zhang, Huchuan Lu, Xiang Ruan, Ming-Hsuan Yang 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2017 Salient Object Detection via Multiple Instance Learning
abstract
Object proposals are a series of candidate segments containing objects of interest, which are taken as preprocessing and widely applied in various vision tasks. However, most of existing saliency approaches only utilize the proposals to compute a location prior. In this paper, we naturally take the proposals as the bags of instances of multiple instance learning (MIL), where the instances are the superpixels contained in the proposals, and formulate saliency detection problem as a MIL task (i.e., predict the labels of instances using the classifier in the MIL framework). This method allows some flexibility in finding a decision boundary based on the bag-level representations and can identify salient superpixels from ambiguous proposals. In addition, we introduce the MIL to an optimization mechanism, which iteratively updates training bags from easy to complex ones to learn a strong model. The significant improvement can be consistently achieved when applying the optimization model to existing saliency approaches. Extensive experiments demonstrate that the proposed algorithms perform favorably against the stateof- art saliency detection methods on several benchmark datasets.
Jinqing Qi, Huchuan Lu, Lihe Zhang, Xiang Ruan
IEEE Trans. Image Process.4
2016 Kernelized Subspace Ranking for Saliency Detection
Tiantian Wang 0002, Lihe Zhang, Huchuan Lu, Jinqing Qi
ECCV (8)2
2016 Combining motion and appearance cues for anomaly detection
Ying Zhang 0021, Huchuan Lu, Lihe Zhang, Xiang Ruan
Pattern Recognit.3
2016 Video anomaly detection based on locality sensitive hashing filters
Ying Zhang 0021, Huchuan Lu, Lihe Zhang, Xiang Ruan, Shun Sakai
Pattern Recognit.3
2016 Salient object detection via point-to-set metric learning
Lihe Zhang, Jinqing Qi, Huchuan Lu
Pattern Recognit. Lett.2
2016 Discriminative Hash Tracking With Group Sparsity
abstract
In this paper, we propose a novel tracking framework based on discriminative supervised hashing algorithm. Different from previous methods, we treat tracking as a problem of object matching in a binary space. Using the hash functions, all target templates and candidates are mapped into compact binary codes, with which the target matching is conducted effectively. To be specific, we make full use of the label information to assign a compact and discriminative binary code for each sample. And to deal with out-of-sample case, multiple hash functions are trained to describe the learned binary codes, and group sparsity is introduced to the hash projection matrix to select the representative and discriminative features dynamically, which is crucial for the tracker to adapt to target appearance variations. The whole training problem is formulated as an optimization function where the hash codes and hash function are learned jointly. Extensive experiments on various challenging image sequences demonstrate the effectiveness and robustness of the proposed tracker.
Dandan Du, Lihe Zhang, Huchuan Lu, Xue Mei, Xiaoli Li 0011
IEEE Trans. Cybern.2
2016 Dense and Sparse Reconstruction Error Based Saliency Descriptor
abstract
In this paper, we propose a visual saliency detection algorithm from the perspective of reconstruction error. The image boundaries are first extracted via superpixels as likely cues for background templates, from which dense and sparse appearance models are constructed. First, we compute dense and sparse reconstruction errors on the background templates for each image region. Second, the reconstruction errors are propagated based on the contexts obtained from K -means clustering. Third, the pixel-level reconstruction error is computed by the integration of multi-scale reconstruction errors. Both the pixel-level dense and sparse reconstruction errors are then weighted by image compactness, which could more accurately detect saliency. In addition, we introduce a novel Bayesian integration method to combine saliency maps, which is applied to integrate the two saliency measures based on dense and sparse reconstruction errors. Experimental results show that the proposed algorithm performs favorably against 24 state-of-the-art methods in terms of precision, recall, and F-measure on three public standard salient object detection databases.
Huchuan Lu, Xiaohui Li 0005, Lihe Zhang, Xiang Ruan, Ming-Hsuan Yang 0001
IEEE Trans. Image Process.3
2016 Sparse Hashing Tracking
abstract
In this paper, we propose a novel tracking framework based on a sparse and discriminative hashing method. Different from the previous work, we treat object tracking as an approximate nearest neighbor searching process in a binary space. Using the hash functions, the target templates and the candidates can be projected into the Hamming space, facilitating the distance calculation and tracking efficiency. First, we integrate both the inter-class and intra-class information to train multiple hash functions for better classification, while most classifiers in previous tracking methods usually neglect the inter-class correlation, which may cause the inaccuracy. Then, we introduce sparsity into the hash coefficient vectors for dynamic feature selection, which is crucial to select the discriminative and stable features to adapt to visual variations during the tracking process. Extensive experiments on various challenging sequences show that the proposed algorithm performs favorably against the state-of-the-art methods.
Lihe Zhang, Huchuan Lu, Dandan Du, Luning Liu
IEEE Trans. Image Process.1
2015 Visual tracking via guided filter
abstract
In this paper, we propose a novel tracking algorithm based on an explicit image filter - guided filter. The guided filter utilizes the structure in the guidance image and performs as an edge-preserving smoothing operator. In our work, we treat the target as the guidance and the incoming candidates are filtered depending on the similarity between the guidance image and each input. The edge-preserving smoothing property depending on the guidance image is a critical advantage for object tracking. First, the guided filter can help to pick out the valuable candidates and make the inaccurate ones blurry so that the tracker can distinguish the target from numerous bad candidates easily. Besides, the filtering process can recover the content of the target being occluded according to the guidance image, which can help to alleviate the drifting problem effectively. Eventually, to generate a robust tracker, we take advantage of the combination of positive and negative templates to conduct effective sparse representation. Experimental results show that our algorithm outperforms relative trackers.
Dandan Du, Huchuan Lu, Lihe Zhang, Fu Li 0003
ICIP3
2015 Visual tracking with astructured local model
abstract
In this paper, we propose a novel visual tracking algorithm combining appearance feature and spatial information. These two aspects extract the constant elements of the target object. To represent information of an object, we utilize a block-local model and a structured visual histogram learned from an object's spatial information, which can adapt to changes in the appearance of the target and environment of the background. During tracking, the proposed algorithm introduces an occlusion model to prevent the environment change from affecting the target's posterior probability. Besides, the online updating is based on incremental algorithms for principal component analysis and the renewal of a spatial histogram queue. Experimental results show that our method outperforms relative trackers.
Minghao Sun, Dandan Du, Huchuan Lu, Lihe Zhang
ICIP4
2015 Saliency detection via sparse reconstruction and joint label inference in multiple features
Lihe Zhang, Shoufeng Zhao, Wei Liu 0055, Huchuan Lu
Neurocomputing1
2015 Salient Object Detection with Higher Order Potentials and Learning Affinity
abstract
In this paper, we propose a novel graph-based salient object detection algorithm which exploits higher order potential to capture the cross-scale grouping cues instead of using multi-scale graph model or naive multi-scale fusion (i.e. individually compute a saliency result for each scale and then combine them). And, we investigate the importance of graph affinities in graph labeling. We take both local (spatial distribution) and nonlocal (feature distribution) priors into account and learn the pairwise similarity values in a semi-supervised manner, thereby obtaining a faithful graph affinity model. With the guidance of foreground and background seeds, salient object detection is formulated as a labeling inference problem. Extensive experiments on two large benchmark datasets demonstrate the proposed method performs well when against the state-of-the-art methods in terms of accuracy.
Lihe Zhang
IEEE Signal Process. Lett.1
2014 Low-rank decomposition and Laplacian group sparse coding for image classification
Lihe Zhang, Chen Ma 0002
Neurocomputing1
2014 Saliency Detection with Multi-Scale Superpixels
abstract
We propose a salient object detection algorithm via multi-scale analysis on superpixels. First, multi-scale segmentations of an input image are computed and represented by superpixels. In contrast to prior work, we utilize various Gaussian smoothing parameters to generate coarse or fine results, thereby facilitating the analysis of salient regions. At each scale, three essential cues from local contrast, integrity and center bias are considered within the Bayesian framework. Next, we compute saliency maps by weighted summation and normalization. The final saliency map is optimized by a guided filter which further improves the detection results. Extensive experiments on two large benchmark datasets demonstrate the proposed algorithm performs favorably against state-of-the-art methods. The proposed method achieves the highest precision value of 97.39% when evaluated on one of the most popular datasets, the ASD dataset.
Na Tong, Huchuan Lu, Lihe Zhang, Xiang Ruan
IEEE Signal Process. Lett.3
2013 Saliency Detection via Graph-Based Manifold Ranking
abstract
Most existing bottom-up methods measure the foreground saliency of a pixel or region based on its contrast within a local context or the entire image, whereas a few methods focus on segmenting out background regions and thereby salient objects. Instead of considering the contrast between the salient objects and their surrounding regions, we consider both foreground and background cues in a different way. We rank the similarity of the image elements (pixels or regions) with foreground cues or background cues via graph-based manifold ranking. The saliency of the image elements is defined based on their relevances to the given seeds or queries. We represent the image as a close-loop graph with super pixels as nodes. These nodes are ranked based on the similarity to background and foreground queries, based on affinity matrices. Saliency detection is carried out in a two-stage scheme to extract background regions and foreground salient objects efficiently. Experimental results on two large benchmark databases demonstrate the proposed method performs well when against the state-of-the-art methods in terms of accuracy and speed. We also create a more difficult benchmark database containing 5,172 images to test the proposed saliency model and make this database publicly available with this paper for further studies in the saliency field.
Lihe Zhang, Huchuan Lu, Xiang Ruan, Ming-Hsuan Yang 0001
CVPR2
2013 Saliency Detection via Absorbing Markov Chain
abstract
In this paper, we formulate saliency detection via absorbing Markov chain on an image graph model. We jointly consider the appearance divergence and spatial distribution of salient objects and the background. The virtual boundary nodes are chosen as the absorbing nodes in a Markov chain and the absorbed time from each transient node to boundary absorbing nodes is computed. The absorbed time of transient node measures its global similarity with all absorbing nodes, and thus salient objects can be consistently separated from the background when the absorbed time is used as a metric. Since the time from transient node to absorbing nodes relies on the weights on the path and their spatial distance, the background region on the center of image may be salient. We further exploit the equilibrium distribution in an ergodic Markov chain to reduce the absorbed time in the long-range smooth background regions. Extensive experiments on four benchmark datasets demonstrate robustness and efficiency of the proposed method against the state-of-the-art methods.
Lihe Zhang, Huchuan Lu, Ming-Hsuan Yang 0001
ICCV2
2013 Saliency Detection via Dense and Sparse Reconstruction
abstract
In this paper, we propose a visual saliency detection algorithm from the perspective of reconstruction errors. The image boundaries are first extracted via super pixels as likely cues for background templates, from which dense and sparse appearance models are constructed. For each image region, we first compute dense and sparse reconstruction errors. Second, the reconstruction errors are propagated based on the contexts obtained from K-means clustering. Third, pixel-level saliency is computed by an integration of multi-scale reconstruction errors and refined by an object-biased Gaussian model. We apply the Bayes formula to integrate saliency measures based on dense and sparse reconstruction errors. Experimental results show that the proposed algorithm performs favorably against seventeen state-of-the-art methods in terms of precision and recall. In addition, the proposed algorithm is demonstrated to be more effective in highlighting salient objects uniformly and robust to background noise.
Xiaohui Li 0005, Huchuan Lu, Lihe Zhang, Xiang Ruan, Ming-Hsuan Yang 0001
ICCV3
2013 Superpixel level object recognition under local learning framework
Huchuan Lu, Xuejiao Feng, Xiaohui Li 0005, Lihe Zhang
Neurocomputing4
2013 Graph-Regularized Saliency Detection With Convex-Hull-Based Center Prior
abstract
Object level saliency detection is useful for many content-based computer vision tasks. In this letter, we present a novel bottom-up salient object detection approach by exploiting contrast, center and smoothness priors. First, we compute an initial saliency map using contrast and center priors. Unlike most existing center prior based methods, we apply the convex hull of interest points to estimate the center of the salient object rather than directly use the image center. This strategy makes the saliency result more robust to the location of objects. Second, we refine the initial saliency map through minimizing a continuous pairwise saliency energy function with graph regularization which encourages adjacent pixels or segments to take the similar saliency value (i.e., smoothness prior). The smoothness prior enables the proposed method to uniformly highlight the salient object and simultaneously suppress the background effectively. Extensive experiments on a large dataset demonstrate that the proposed method performs favorably against the state-of-the-art methods in terms of accuracy and efficiency.
Lihe Zhang, Huchuan Lu
IEEE Signal Process. Lett.2
2012 Superpixel level object recognition under local learning framework
abstract
In this paper, we propose a simple yet efficient method for superpixel level object recognition on the bag-of-feature framework. Instead of using general classifiers for the superpixel categorization, we introduce local learning classifiers into our framework, so as to tackle the intraclass variation problem brought by superpixel based representations of objects. In addition, context information is used to make better performance by combining each superpixel with its most similar neighbors. We test our proposed method on Graz-02 datasets, and get results comparable to the state-of-the-art.
Xuejiao Feng, Huchuan Lu, Lihe Zhang
ICIP3
2012 Low-rank, sparse matrix decomposition and group sparse coding for image classification
abstract
This paper presents a novel image classification framework (referred to as LR-GSC) by leveraging the low-rank, sparse matrix decomposition and group sparse coding. First, motivated by the observation that local features (such as SIFT) extracted from neighboring patches in an image usually contain correlated (or common) items and specific (or noisy) items, we decompose the local features matrix of an image into a low-rank matrix and a sparse matrix. Second, we train the group sparse dictionaries on the low-rank parts and sparse parts respectively. And then, the dictionaries of the two parts are jointed to encode the original SIFT features by group coding. Finally, linear SVM classifier is used for the classification. The method is tested on the Caltech-101 dataset and UIUC-sports dataset, and achieves competitive or better results than the state-of-the-art methods.
Lihe Zhang, Chen Ma 0002
ICIP1
2011 Online sparse learning utilizing multi-feature combination for image classification
abstract
Bag-of-features has become very popular in Image classification. Offline codebook learning has to limit the number of training sample concerned with memory, and it influences classification accuracy to some extent. We propose an online sparse learning algorithm, which utilizes the reconstruction error to update the current codebook. It can capture salient properties of images in real-time. Most of image representation approaches in Gabor domain merely utilize magnitude information, and some important phase information is missing. Taking both magnitude and phase response into account, a Local Gabor Magnitude Weighted Phase (LGMWP) descriptor is proposed in this paper. The technique works by dividing the image into local patches, extracting SIFT and LGMWP features to online learn the codebook respectively, implementing spatial pyramid matching (SPM) and binary SVM classifier. The experiment results demonstrate our approach outperforms offline learning with a single type of descriptors.
Lihe Zhang, Kunyu Zhang, Xiaoli Dong
ICIP1
2007 A Video Watermarking Scheme Resistant to Synchronization Attacks Based on Statistics and Shot Segmentation
abstract
One of the challenges of blind watermark detection is synchronization. In this paper, a new video watermarking procedure for resistant synchronization is proposed. We partition watermark into several segments, and embed every segment into different scenes of video sequence. Firstly, motion vectors are grouped according to their magnitude and each group is further partitioned into two subsets. We use element number ratio in two subsets of the same group to denote watermark bit 0 or bit 1. Thus, watermark is associated with motion vectors' statistical characteristics, and those motion vectors carrying the identical watermark bits are uniformly distributed to the whole video sequence, they are spatio-temporal indiscerptibly. It is shown that this kind of watermark is more resilient against temporal synchronization attacks. Experimental results from an implementation of the algorithm are presented.
Lihe Zhang, Ji-jun Zhou, Lin-jie Shen
ISDA1