VLDB 2026 Research / reviewers in the wild / expert
Ping Hu 0001
dblp:53/5490-1
· DBLP profile ↗
50ranked-venue papers
14as first author
39since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 38 · 8 first-author · 32 since 2021Graphics, computer vision, multimedia, augmented reality and games · 32 · 10 first-author · 23 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Graph Smoothing for Enhanced Local Geometry Learning in Point Cloud AnalysisabstractGraph-based methods have proven to be effective in capturing relationships among points for 3D point cloud analysis. However, these methods often suffer from suboptimal graph structures, particularly due to sparse connections at boundary points and noisy connections in junction areas. To address these challenges, we propose a novel method that integrates a graph smoothing module with an enhanced local geometry learning module. Specifically, we identify the limitations of conventional graph structures, particularly in handling boundary points and junction areas. In response, we introduce a graph smoothing module designed to optimize the graph structure and minimize the negative impact of unreliable sparse and noisy connections. Based on the optimized graph structure, we improve the feature extract function with local geometry information. These include shape features derived from adaptive geometric descriptors based on eigenvectors and distribution features obtained through cylindrical coordinate transformation. Experimental results on real-world datasets validate the effectiveness of our method in various point cloud learning tasks, i.e., classification, part segmentation, and semantic segmentation. Shangbo Yuan, Jie Xu 0044, Ping Hu 0001, Xiaofeng Zhu 0001, Na Zhao 0004 |
AAAI | 3 |
| 2026 | Generalized prompt-driven zero-shot domain adaptive segmentation with feature rectification and semantic modulation
Jinyi Li, Longyu Yang, Donghyun Kim 0006, Kuniaki Saito, Kate Saenko, Stan Sclaroff, Xiaofeng Zhu 0001, Ping Hu 0001 |
Comput. Vis. Image Underst. | 8 |
| 2026 | Explicit Geometry-Reflectance Domain Shift Modeling for Robust LiDAR Segmentation in Adverse Weather
Longyu Yang, Shangbo Yuan, Lu Zhang 0053, Jun Liu 0036, Heng Tao Shen, Xiaofeng Zhu 0001, Ping Hu 0001 |
Int. J. Comput. Vis. | 7 |
| 2026 | Adaptive memory refinement and perception enhancement for exo-to-ego video generation
Weipeng Hu, Jiun Tian Hoe, Ping Hu 0001, Xudong Jiang 0001, Yap-Peng Tan |
Neurocomputing | 5 |
| 2026 | Unleashing the Power of Text-to-Image Diffusion Models for Category-Agnostic Pose EstimationabstractCategory-Agnostic Pose Estimation (CAPE) aims to detect keypoints of unseen object categories in a few-shot setting, where the scarcity of labeled data poses significant challenges to generalization. In this work, we propose Prompt Pose Matching (PPM), a novel framework that unleashes the power of off-the-shelf text-to-image diffusion models for CAPE. PPM learns pseudo prompts from few-shot examples via the text-to-image diffusion model. These learned pseudo prompts capture semantic information of keypoints, which can then be used to locate the same type of keypoints from images. To provide prompts with representative initialization, we introduce a category-agnostic pre-training strategy to capture the foreground prior shared across categories and keypoints. To support the reliable prompt pre-training, we propose a Foreground-Aware Region Aggregation (FARA) module to provide robust and consistent supervision signal. Based on the foreground prior, a Foreground-Guided Attention Refinement (FGAR) module is further proposed to reinforce cross-attention responses for accurate keypoint localization. For efficiency, a Prompt Ensemble Inference (PEI) scheme enables joint keypoint prediction. Unlike previous methods that highly rely on base-category annotated data, our PPM framework can operate in a base-category-free setting while retaining strong performance. Code will be available at: https://github.com/DuoPeng-CVer/Prompt-Pose-Matching. Duo Peng, Zhengbo Zhang, Ping Hu 0001, Qiuhong Ke, De Wen Soh, Mohammed Bennamoun, Jun Liu 0036 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2026 | DiMuS: Disentangled Multi-Signal Learning for Weakly Supervised Point-Based 3D Object DetectionabstractWeakly supervised 3D object detection has emerged as a promising paradigm to reduce the reliance on costly 3D annotations. Existing methods often rely on 2D projection constraints or heuristic priors to supervise 3D box regression with inexpensive 2D labels. However, they still suffer from projection ambiguity and geometry inconsistency due to the entangled optimization of 3D parameters. In this paper, we propose DiMuS, a Disentangled Multi- $\boldsymbol {S}$ ignal learning framework that integrates complementary supervision from 2D boxes, LLM-derived semantic prior, and 3D geometric alignment to enhance distinct 3D properties of position, dimension, and orientation, respectively. Specifically, DiMuS incorporates three key components: (i) a Centerness-enhanced Projection Constraint (CPC) that improves position estimation through a centerness weighting strategy, (ii) a Semantic Prior Anchoring (SPA) module that leverages LLM-derived category-specific priors for robust dimension decoding, and (iii) a Rotation-aware Consistency Regularization (RCR) mechanism that enforces orientation consistency through synthetic rotations and self-supervised invariance learning. Additionally, an Adversarial Geometric Alignment (AGA) module is proposed to build attraction/repulsion forces between LiDAR points and box edges for dynamic boundary refinement. Extensive experiments on the KITTI dataset demonstrate that DiMuS outperforms previous weakly supervised methods, achieving 96.82% of fully supervised performance on car detection while maintaining robustness across different categories. Yunzhi Zhuge, Lu Zhang 0053, Ping Hu 0001, Dong Wang 0004, Huchuan Lu |
IEEE Trans. Image Process. | 4 |
| 2025 | Bootstraping Clustering of Gaussians for View-consistent 3D Scene UnderstandingabstractInjecting semantics into 3D Gaussian Splatting (3DGS) has recently garnered significant attention. While current approaches typically distill 3D semantic features from 2D foundational models (e.g., CLIP and SAM) to facilitate novel view segmentation and semantic understanding, their heavy reliance on 2D supervision can undermine cross-view semantic consistency and necessitate complex data preparation processes, therefore hindering view-consistent scene understanding. In this work, we present FreeGS, an unsupervised semantic-embedded 3DGS framework that achieves view-consistent 3D scene understanding without the need for 2D labels. Instead of directly learning semantic features, we introduce the IDentity-coupled Semantic Field (IDSF) into 3DGS, which captures both semantic representations and view-consistent instance indices for each Gaussian. We optimize IDSF with a two-step alternating strategy: semantics help to extract coherent instances in 3D space, while the resulting instances regularize the injection of stable semantics from 2D space. Additionally, we adopt a 2D-3D joint contrastive loss to enhance the complementarity between view-consistent 3D geometry and rich semantics during the bootstrapping process, enabling FreeGS to uniformly perform tasks such as novel-view semantic segmentation, object selection, and 3D object detection. Extensive experiments on LERF-Mask, 3D-OVS, and ScanNet datasets demonstrate that FreeGS performs comparably to state-of-the-art methods while avoiding the complex data preprocessing workload. Lu Zhang 0053, Ping Hu 0001, Liqian Ma, Yunzhi Zhuge, Huchuan Lu |
AAAI | 3 |
| 2025 | Towards Explicit Geometry-Reflectance Collaboration for Generalized LiDAR Segmentation in Adverse WeatherabstractExisting LiDAR semantic segmentation models often suffer from decreased accuracy when exposed to adverse weather conditions. Recent methods addressing this issue focus on enhancing training data through weather simulation or universal augmentation techniques. However, few works have studied the negative impacts caused by the heterogeneous domain shifts in the geometric structure and reflectance intensity of point clouds. In this paper, we delve into this challenge and address it with a novel Geometry-Reflectance Collaboration (GRC) framework that explicitly separates feature extraction for geometry and reflectance. Specifically, GRC employs a dual-branch architecture designed to independently process geometric and reflectance features initially, thereby capitalizing on their distinct characteristic. Then, GRC adopts a robust multi-level feature collaboration module to suppress redundant and unreliable information from both branches. Consequently, without complex simulation or augmentation, our method effectively extracts intrinsic information about the scene while suppressing interference, thus achieving better robustness and generalization in adverse weather conditions. We demonstrate the effectiveness of GRC through comprehensive experiments on challenging benchmarks, showing that our method out-performs previous approaches and establishes new state-of-the-art results. Longyu Yang, Ping Hu 0001, Shangbo Yuan, Lu Zhang 0053, Jun Liu 0036, Heng Tao Shen, Xiaofeng Zhu 0001 |
CVPR | 2 |
| 2025 | Boundary Probing for Input Privacy Protection when Using LMM Services
Xiaofei Hui, Haoxuan Qu, Ping Hu 0001, Hossein Rahmani 0001, Jun Liu 0036 |
ICCV | 3 |
| 2025 | TSTMotion: Training-free Scene-aware Text-to-motion GenerationabstractText-to-motion generation has recently garnered significant research interest, primarily focusing on generating human motion sequences in blank backgrounds. However, human motions commonly occur within diverse 3D scenes, which has prompted exploration into scene-aware text-to-motion generation methods. Yet, existing scene-aware methods often rely on large-scale ground-truth motion sequences in diverse 3D scenes, which poses practical challenges due to the expensive cost. To mitigate this challenge, we are the first to propose a Training-free Scene-aware Text-to-Motion framework, dubbed as TSTMotion, that efficiently empowers pre-trained blank-background motion generators with the scene-aware capability. Specifically, conditioned on the given 3D scene and text description, we adopt foundation models together to reason, predict and validate a scene-aware motion guidance. Then, the motion guidance is incorporated into the blank-background motion generators with two modifications, resulting in scene-aware text-driven motion sequences. Extensive experiments demonstrate the efficacy and generalizability of our proposed framework. We release our code in Project Page. Ziyan Guo, Haoxuan Qu, Hossein Rahmani 0001, De Wen Soh, Ping Hu 0001, Qiuhong Ke, Jun Liu 0036 |
ICME | 5 |
| 2025 | Seeking Proxy Point via Stable Feature Space for Noisy Correspondence LearningabstractTo meet the growing demand for cross-modal training data, directly collecting multimodal data from the Internet has become prevalent. However, such data inevitably suffer from Noisy Correspondence. Previous works focused on recasting soft labels to mitigate noise's negative impact. We explore a novel perspective to solve this problem: pursuing proxy representation for noisy data to enable reliable feature learning. To this end, we propose a novel framework: Seeking Proxy Point via Stable Feature Space (SPS). This framework employs a fine-grained partitioning strategy to obtain a high-confidence reliable set. By imposing intermodal cross-transformation consistency constraints and intramodal metric consistency constraints, a stable feature space is constructed. Building on this foundation, SPS seeks proxy points for noisy data, enabling even noisy data to be accurately embedded into appropriate positions within the feature space. Combined with partial alignment for partially matched data pairs, SPS ultimately achieves robust learning under Noisy Correspondence. Experiments on three widely used cross-modal datasets demonstrate that SPS significantly outperforms previous methods. Our code is available at https://github.com/C-TeaRanger/SPS. Yucheng Xie, Songyue Cai, Tao Tong, Ping Hu 0001, Xiaofeng Zhu 0001 |
IJCAI | 4 |
| 2025 | FineRS: Fine-grained Reasoning and Segmentation of Small Objects with Reinforcement LearningabstractMulti-modal Large Language Models (MLLMs) have shown remarkable capabilities across a wide range of vision-language tasks. However, due to the restricted input resolutions, MLLMs face significant challenges in precisely understanding and localizing visual details in high-resolution images---particularly when dealing with extra-small objects embedded in cluttered contexts.
To address this issue, we propose FineRS, a two-stage MLLM-based reinforcement learning framework for jointly reasoning and segmenting extremely small objects within high-resolution scenes. FineRS adopts a coarse-to-fine pipeline comprising Global Semantic Exploration (GSE) and Localized Perceptual Refinement (LPR). Specifically, GSE performs instruction-guided reasoning to generate a textural response and a coarse target region, while LPR refines this region to produce an accurate bounding box and segmentation mask. To couple the two stages, we introduce a locate-informed retrospective reward, where LPR's outputs are used to optimize GSE for more robust coarse region exploration. Additionally, we present FineRS-4k, a new dataset for evaluating MLLMs on attribute-level reasoning and pixel-level segmentation on subtle, small-scale targets in complex high-resolution scenes. Experimental results on FineRS-4k and public datasets demonstrate that our method consistently outperforms state-of-the-art MLLM-based approaches on both instruction-guided segmentation and visual reasoning tasks. Lu Zhang 0053, Jiazuo Yu 0001, Haomiao Xiong, Ping Hu 0001, Yunzhi Zhuge, Huchuan Lu, You He 0002 |
NeurIPS | 4 |
| 2025 | On Efficient Variants of Segment Anything Model: A Survey
Xiaorui Sun, Jun Liu 0036, Heng Tao Shen, Xiaofeng Zhu 0001, Ping Hu 0001 |
Int. J. Comput. Vis. | 5 |
| 2025 | Unified Prompt Attack Against Text-to-Image Generation ModelsabstractText-to-Image (T2I) models have advanced significantly, but their growing popularity raises security concerns due to their potential to generate harmful images. To address these issues, we propose UPAM, a novel framework to evaluate the robustness of T2I models from an attack perspective. Unlike prior methods that focus solely on textual defenses, UPAM unifies the attack on both textual and visual defenses. Additionally, it enables gradient-based optimization, overcoming reliance on enumeration for improved efficiency and effectiveness. To handle cases where T2I models block image outputs due to defenses, we introduce Sphere-Probing Learning (SPL) to enable optimization even without image results. Following SPL, our model bypasses defenses, inducing the generation of harmful content. To ensure semantic alignment with attacker intent, we propose Semantic-Enhancing Learning (SEL) for precise semantic control. UPAM also prioritizes the naturalness of adversarial prompts using In-context Naturalness Enhancement (INE), making them harder for human examiners to detect. Additionally, we address the issue of iterative queries-common in prior methods and easily detectable by API defenders-by introducing Transferable Attack Learning (TAL), allowing effective attacks with minimal queries. Extensive experiments validate UPAM's superiority in effectiveness, efficiency, naturalness, and low query detection rates. Duo Peng, Qiuhong Ke, Mark He Huang, Ping Hu 0001, Jun Liu 0036 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | MoE-Adapters++: Toward More Efficient Continual Learning of Vision-Language Models Via Dynamic Mixture-of-Experts AdaptersabstractIn this paper, we first propose MoE-Adapters, a parameter-efficient training framework to alleviate long-term forgetting issues in incremental learning with Vision-Language Models (VLM). Our MoE-Adapters leverages incrementally added routers to activate and integrate exclusive expert adapters from a pre-defined static expert set, enabling the pre-trained CLIP to efficiently adapt to new tasks. To preserve the zero-shot capability of VLM, a Distribution Discriminative Auto-Selector (DDAS) is introduced that automatically routes in-distribution and out-of-distribution inputs to the MoE-Adapters and the original CLIP, respectively. However, relying on a static expert set and a separate distribution selector can lead to parameter redundancy and increased training complexity. In response, we further extend an MoE-Adapters++ framework by introducing dynamic MoE-adapters, which allows experts to be adaptively involved during the continual learning process. Additionally, a Latent Embedding Auto-Selector (LEAS) is proposed that incorporates distribution selection within CLIP to create a more unified architecture. Extensive experiments across diverse settings demonstrate that the proposed method consistently surpasses previous state-of-the-art approaches while concurrently improving training efficiency. Jiazuo Yu 0001, Zichen Huang 0004, Yunzhi Zhuge, Lu Zhang 0053, Ping Hu 0001, Dong Wang 0004, Huchuan Lu, You He 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | Unsupervised multiplex graph representation learning via maximizing coding rate reduction
Xin Wang 0148, Rongyao Hu, Ping Hu 0001, Xiaofeng Zhu 0001 |
Pattern Recognit. | 4 |
| 2025 | RVMamba: Selective Text-Vision Mamba for Referring Video Object SegmentationabstractExisting RVOS methods typically employ Transformers to model global cross-modal, temporal-spatial correspondences, but their quadratic complexity limits deployment on resource-constrained devices. To overcome this limitation, Mamba offers a sequence modeling framework with linear computational complexity. We introduceRVMamba, which utilizes weight modulation to selectively update hidden states across text-frame sequences, enabling effective linguistic context propagation, and a learning-based scanning strategy to efficiently capture spatio-temporal dependencies with linear memory consumption. Extensive experiments demonstrate thatRVMambaachieves state-of-the-art performance on public benchmarks, with significantly reduced memory growth, offering an efficient and scalable solution for long video processing. Zhenyu Chen 0001, Jiawen Zhu 0003, Lu Zhang 0053, Ping Hu 0001, Yunzhi Zhuge, Huchuan Lu, You He 0002 |
IEEE Signal Process. Lett. | 4 |
| 2025 | MaskTrack: Auto-Labeling and Stable Tracking for Video Object SegmentationabstractVideo object segmentation (VOS) has witnessed notable progress due to the establishment of video training datasets and the introduction of diverse, innovative network architectures. However, video mask annotation is a highly intricate and labor-intensive task, as meticulous frame-by-frame comparisons are needed to ascertain the positions and identities of targets in the subsequent frames. Current VOS benchmarks often annotate only a few instances in each video to save costs, which, however, hinders the model's understanding of the complete context of the video scenes. To simplify video annotation and achieve efficient dense labeling, we introduce a zero-shot auto-labeling strategy based on the segment anything model (SAM), enabling it to densely annotate video instances without access to any manual annotations. Moreover, although existing VOS methods demonstrate improving performance, segmenting long-term and complex video scenes remains challenging due to the difficulties in stably discriminating and tracking instance identities. To this end, we further introduce a new framework, MaskTrack, which excels in long-term VOS and also exhibits significant performance advantages in distinguishing instances in complex videos with densely packed similar objects. We conduct extensive experiments to demonstrate the effectiveness of the proposed method and show that without introducing image datasets for pretraining, it achieves excellent performance on both short-term (86.2% in YouTube-VOS val) and long-term (68.2% in LVOS val) VOS benchmarks. Our method also surprisingly demonstrates strong generalization ability and performs well in visual object tracking (VOT) (65.6% in VOTS2023) and referring VOS (RVOS) (65.2% in Ref YouTube VOS) challenges. Zhenyu Chen 0001, Lu Zhang 0053, Ping Hu 0001, Huchuan Lu, You He 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Koala: Key Frame-Conditioned Long Video-LLMabstractLong video question answering is a challenging task that involves recognizing short-term activities and reasoning about their fine-grained relationships. State-of-the-art video Large Language Models (vLLMs) hold promise as a viable solution due to their demonstrated emergent capabilities on new tasks. However, despite being trained on millions of short seconds-long videos, vLLMs are unable to understand minutes-long videos and accurately answer questions about them. To address this limitation, we propose a lightweight and self-supervised approach, Key frame-conditioned long video-LLM (Koala), that introduces learnable spatiotemporal queries to adapt pretrained vLLMs for generalizing to longer videos. Our approach introduces two new tokenizers that condition on visual tokens computed from sparse video key frames for understanding short and long video moments. We train our proposed approach on HowTo100M and demonstrate its effectiveness on zero-shot long video understanding benchmarks, where it outperforms state-of-the-art large models by 3 - 6% in absolute accuracy across all tasks. Surprisingly, we also empirically show that our approach not only helps a pretrained vLLM to understand long videos but also improves its accuracy on short-term action recognition. Reuben Tan, Ximeng Sun, Ping Hu 0001, Jui-Hsien Wang, Hanieh Deilamsalehy, Bryan A. Plummer, Bryan C. Russell, Kate Saenko |
CVPR | 3 |
| 2024 | Boosting Continual Learning of Vision-Language Models via Mixture-of-Experts AdaptersabstractContinual learning can empower vision-language models to continuously acquire new knowledge, without the need for access to the entire historical dataset. However, mitigating the performance degradation in large-scale models is non-trivial due to (i) parameter shifts throughout life-long learning and (ii) significant computational burdens associated with full-model tuning. In this work, we present a parameter-efficient continual learning framework to alleviate long-term forgetting in incremental learning with vision-language models. Our approach involves the dynamic expansion of a pre-trained CLIP model, through the integration of Mixture-of-Experts (MoE) adapters in response to new tasks. To preserve the zero-shot recognition capability of vision-language models, we further introduce a Distribution Discriminative Auto-Selector (DDAS) that automatically routes in-distribution and out-of-distribution inputs to the MoE Adapter and the original CLIP, respectively. Through extensive experiments across various settings, our proposed method consistently outperforms previous state-of-the-art approaches while concurrently reducing parameter training burdens by 60%. Our code locates at https://github.com/JiazuoYu/MoE-Adapters4CL Jiazuo Yu 0001, Yunzhi Zhuge, Lu Zhang 0053, Ping Hu 0001, Dong Wang 0004, Huchuan Lu, You He 0002 |
CVPR | 4 |
| 2024 | Harnessing Text-to-Image Diffusion Models for Category-Agnostic Pose Estimation
Duo Peng, Zhengbo Zhang, Ping Hu 0001, Qiuhong Ke, David K. Y. Yau, Jun Liu 0036 |
ECCV (13) | 3 |
| 2024 | Self-Supervised Heterogeneous Graph Learning: a Homophily and Heterogeneity ViewabstractSelf-supervised heterogeneous graph learning has achieved promising results in various real applications, but it still suffers from the following issues: (i) meta-paths can be employed to capture the homophily in the heterogeneous graph, but meta-paths are human-defined, requiring substantial expert knowledge and computational costs; and (ii) the heterogeneity in the heterogeneous graph is usually underutilized, leading to the loss of task-related information. To solve these issues, this paper proposes to capture both homophily and heterogeneity in the heterogeneous graph without pre-defined meta-paths. Specifically, we propose to learn a self-expressive matrix to capture the homophily from the subspace and nearby neighbors. Meanwhile, we propose to capture the heterogeneity by aggregating the information of nodes from different types. We further design a consistency loss and a specificity loss, respectively, to extract the consistent information between homophily and heterogeneity and to preserve their specific task-related information. We theoretically analyze that the learned homophilous representations exhibit the grouping effect to capture the homophily, and considering both homophily and heterogeneity introduces more task-related information. Extensive experimental results verify the superiority of the proposed method on different downstream tasks. Yujie Mo, Feiping Nie 0001, Ping Hu 0001, Heng Tao Shen, Zheng Zhang 0006, Xinchao Wang, Xiaofeng Zhu 0001 |
ICLR | 3 |
| 2024 | GATrack: Group-Aware features for multiple object trackingabstractCurrent multiple object tracking methods typically associate two detected objects from consecutive frames via discriminative appearance features or motion modeling at the object level. However, in scenarios of dense crowds and prolonged occlusions, the extracted object-level features lack reliability, resulting in less effective target association. To tackle this challenge, we introduce GAT, a novel Group-Aware Transformer that learns to automatically group objects and complement targets’ appearance features with multi-level contextual information. In response to the prolonged occlusions issue, we further introduce an effective Trajectory Merging Mechanism (TMM), which relink the failed trajectories in prolonged occlusion based on motion patterns inferred from historical information. We demonstrate the effectiveness of our method with ablative experiments and exhibit outstanding tracking performance on the three popular multiple object tracking benchmarks MOT17, MOT20, and DanceTrack. Ping Hu 0001, Rongyao Hu, Xiaofeng Zhu 0001 |
ICME | 2 |
| 2024 | Exploring the Role of Node Diversity in Directed Graph Representation Learning
Jincheng Huang 0005, Yujie Mo, Ping Hu 0001, Xiaoshuang Shi, Shangbo Yuan, Xiaofeng Zhu 0001 |
IJCAI | 3 |
| 2024 | Towards Dynamic-Prompting Collaboration for Source-Free Domain Adaptation
Mengmeng Zhan, Zongqian Wu, Rongyao Hu, Ping Hu 0001, Heng Tao Shen, Xiaofeng Zhu 0001 |
IJCAI | 4 |
| 2024 | Adaptive Multi-Modality Prompt LearningabstractAlthough current prompt learning methods have successfully been designed to effectively reuse the large pre-trained models without fine-tuning their large number of parameters, they still have limitations to be addressed, i.e., without considering the adverse impact of meaningless patches in every image and without simultaneously considering in-sample generalization and out-of-sample generalization. In this paper, we propose an adaptive multi-modality prompt learning to address the above issues. To do this, we employ previous text prompt learning and propose a new image prompt learning. The image prompt learning achieves in-sample and out-of-sample generalization, by first masking meaningless patches and then padding them with the learnable parameters and the information from texts. Moreover, each of the prompts provides auxiliary information to each other, further strengthening these two kinds of generalization. Experimental results on real datasets demonstrate that our method outperforms SOTA methods, in terms of different downstream tasks. Zongqian Wu, Mengmeng Zhan, Ping Hu 0001, Xiaofeng Zhu 0001 |
ACM Multimedia | 4 |
| 2024 | Learning depth-aware decomposition for single image dehazing
Yumeng Kang, Lu Zhang 0053, Ping Hu 0001, Yu Liu 0005, Huchuan Lu, You He 0002 |
Comput. Vis. Image Underst. | 3 |
| 2024 | Video Frame Interpolation With Many-to-Many Splatting and Spatial Selective RefinementabstractIn this work, we first propose a fully differentiable Many-to-Many (M2M) splatting framework to interpolate frames efficiently. Given a frame pair, we estimate multiple bidirectional flows to directly forward warp the pixels to the desired time step before fusing any overlapping pixels. In doing so, each source pixel renders multiple target pixels and each target pixel can be synthesized from a larger area of visual context, establishing a many-to-many splatting scheme with robustness to undesirable artifacts. For each input frame pair, M2M has a minuscule computational overhead when interpolating an arbitrary number of in-between frames, hence achieving fast multi-frame interpolation. However, directly warping and fusing pixels in the intensity domain is sensitive to the quality of motion estimation and may suffer from less effective representation capacity. To improve interpolation accuracy, we further extend an M2M++ framework by introducing a flexible Spatial Selective Refinement (SSR) component, which allows for trading computational efficiency for interpolation quality and vice versa. Instead of refining the entire interpolated frame, SSR only processes difficult regions selected under the guidance of an estimated error map, thereby avoiding redundant computation. Evaluation on multiple benchmark datasets shows that our method is able to improve the efficiency while maintaining competitive video interpolation quality, and it can be adjusted to use more or less compute as needed. Ping Hu 0001, Simon Niklaus, Lu Zhang 0053, Stan Sclaroff, Kate Saenko |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | DualCoOp++: Fast and Effective Adaptation to Multi-Label Recognition With Limited AnnotationsabstractMulti-label image recognition in the low-label regime is a task of great challenge and practical significance. Previous works have focused on learning the alignment between textual and visual spaces to compensate for limited image labels, yet may suffer from reduced accuracy due to the scarcity of high-quality multi-label annotations. In this research, we leverage the powerful alignment between textual and visual features pretrained with millions of auxiliary image-text pairs. We introduce an efficient and effective framework calledEvidence-guided Dual Context Optimization(DualCoOp++), which serves as a unified approach for addressing partial-label and zero-shot multi-label recognition. InDualCoOp++we separately encode evidential, positive, and negative contexts for target classes as parametric components of the linguistic input (i.e., prompts). The evidential context aims to discover all the related visual content for the target class, and serves as guidance to aggregate positive and negative contexts from the spatial domain of the image, enabling better distinguishment between similar categories. Additionally, we introduce a Winner-Take-All module that promotes inter-class interaction during training, while avoiding the need for extra parameters and costs. AsDualCoOp++imposes minimal additional learnable overhead on the pretrained vision-language framework, it enables rapid adaptation to multi-label recognition tasks with limited annotations and even unseen classes. Experiments on standard multi-label recognition benchmarks across two challenging low-label settings demonstrate the superior performance of our approach compared to state-of-the-art methods. Ping Hu 0001, Ximeng Sun, Stan Sclaroff, Kate Saenko |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | Token Boosting for Robust Self-Supervised Visual Transformer Pre-trainingabstractLearning with large-scale unlabeled data has become a powerful tool for pre-training Visual Transformers (VTs). However, prior works tend to overlook that, in real-world scenarios, the input data may be corrupted and unreliable. Pre-training VTs on such corrupted data can be challenging, especially when we pre-train via the masked autoencoding approach, where both the inputs and masked “ground truth” targets can potentially be unreliable in this case. To address this limitation, we introduce the Token Boosting Module (TBM) as a plug-and-play component for VTs that effectively allows the VT to learn to extract clean and robust features during masked autoencoding pre-training. We provide theoretical analysis to show how TBM improves model pre-training with more robust and generalizable representations, thus benefiting down stream tasks. We conduct extensive experiments to analyze TBM's effectiveness, and results on four corrupted datasets demonstrate that TBM consistently improves performance on downstream tasks. Lin Geng Foo, Ping Hu 0001, Xindi Shang, Hossein Rahmani 0001, Zehuan Yuan, Jun Liu 0036 |
CVPR | 3 |
| 2023 | Diffusion-based Image Translation with Label Guidance for Domain Adaptive Semantic SegmentationabstractTranslating images from a source domain to a target domain for learning target models is one of the most common strategies in domain adaptive semantic segmentation (DASS). However, existing methods still struggle to preserve semantically-consistent local details between the original and translated images. In this work, we present an innovative approach that addresses this challenge by using sourcedomain labels as explicit guidance during image translation. Concretely, we formulate cross-domain image translation as a denoising diffusion process and utilize a novel Semantic Gradient Guidance (SGG) method to constrain the translation process, conditioning it on the pixel-wise source labels. Additionally, a Progressive Translation Learning (PTL) strategy is devised to enable the SGG method to work reliably across domains with large gaps. Extensive experiments demonstrate the superiority of our approach over state-of-the-art methods. Duo Peng, Ping Hu 0001, Qiuhong Ke, Jun Liu 0036 |
ICCV | 2 |
| 2023 | Joint Attribute and Model Generalization Learning for Privacy-Preserving Action RecognitionabstractPrivacy-Preserving Action Recognition (PPAR) aims to transform raw videos into anonymous ones to prevent privacy leakage while maintaining action clues, which is an increasingly important problem in intelligent vision applications. Despite recent efforts in this task, it is still challenging to deal with novel privacy attributes and novel privacy attack models that are unavailable during the training phase. In this paper, from the perspective of meta-learning (learning to learn), we propose a novel Meta Privacy-Preserving Action Recognition (MPPAR) framework to improve both generalization abilities above (i.e., generalize to *novel privacy attributes* and *novel privacy attack models*) in a unified manner. Concretely, we simulate train/test task shifts by constructing disjoint support/query sets w.r.t. privacy attributes or attack models. Then, a virtual training and testing scheme is applied based on support/query sets to provide feedback to optimize the model's learning toward better generalization. Extensive experiments demonstrate the effectiveness and generalization of the proposed framework compared to state-of-the-arts. Duo Peng, Qiuhong Ke, Ping Hu 0001, Jun Liu 0036 |
NeurIPS | 4 |
| 2023 | Semantic Consistent Embedding for Domain Adaptive Zero-Shot LearningabstractUnsupervised domain adaptation has limitations when encountering label discrepancy between the source and target domains. While open-set domain adaptation approaches can address situations when the target domain has additional categories, these methods can only detect them but not further classify them. In this paper, we focus on a more challenging setting dubbed Domain Adaptive Zero-Shot Learning (DAZSL), which uses semantic embeddings of class tags as the bridge between seen and unseen classes to learn the classifier for recognizing all categories in the target domain when only the supervision of seen categories in the source domain is available. The main challenge of DAZSL is to perform knowledge transfer across categories and domain styles simultaneously. To this end, we propose a novel end-to-end learning mechanism dubbed Three-way Semantic Consistent Embedding (TSCE) to embed the source domain, target domain, and semantic space into a shared space. Specifically, TSCE learns domain-irrelevant categorical prototypes from the semantic embedding of class tags and uses them as the pivots of the shared space. The source domain features are aligned with the prototypes via their supervised information. On the other hand, the mutual information maximization mechanism is introduced to push the target domain features and prototypes towards each other. By this way, our approach can align domain differences between source and target images, as well as promote knowledge transfer towards unseen classes. Moreover, as there is no supervision in the target domain, the shared space may suffer from the catastrophic forgetting problem. Hence, we further propose a ranking-based embedding alignment mechanism to maintain the consistency between the semantic space and the shared space. Experimental results on both I2AwA and I2WebV clearly validate the effectiveness of our method. Code is available at https://github.com/tiggers23/TSCE-Domain-Adaptive-Zero-Shot-Learning. Jianyang Zhang, Guowu Yang, Ping Hu 0001, Guosheng Lin, Fengmao Lv |
IEEE Trans. Image Process. | 3 |
| 2022 | Video Object Segmentation via Structural Feature Reconfiguration
Zhenyu Chen 0001, Ping Hu 0001, Lu Zhang 0053, Huchuan Lu, You He 0002, Maodi Hu |
ACCV (7) | 2 |
| 2022 | ZeroWaste Dataset: Towards Deformable Object Segmentation in Cluttered ScenesabstractLess than 35% of recyclable waste is being actually recycled in the US [2], which leads to increased soil and sea pollution and is one of the major concerns of environmental researchers as well as the common public. At the heart of the problem are the inefficiencies of the waste sorting process (separating paper, plastic, metal, glass, etc.) due to the extremely complex and cluttered nature of the waste stream. Recyclable waste detection poses a unique computer vision challenge as it requires detection of highly deformable and often translucent objects in cluttered scenes without the kind of context information usually present in human-centric datasets. This challenging computer vision task currently lacks suitable datasets or methods in the available literature. In this paper, we take a step towards computer-aided waste detection and present the first in-the-wild industrial-grade waste detection and segmentation dataset, ZeroWaste. We believe that ZeroWaste will catalyze research in object detection and semantic segmentation in extreme clutter as well as applications in the recycling domain. Our project page can be found at http://ai.bu.edu/zerowaste/ Dina Bashkirova, Mohamed Abdelfattah, Ziliang Zhu, James Akl, Fadi M. Alladkani, Ping Hu 0001, Vitaly Ablavsky, Berk Çalli, Sarah Adel Bargal, Kate Saenko |
CVPR | 6 |
| 2022 | Many-to-many Splatting for Efficient Video Frame InterpolationabstractMotion-based video frame interpolation commonly relies on optical flow to warp pixels from the inputs to the desired interpolation instant. Yet due to the inherent challenges of motion estimation (e.g. occlusions and discontinuities), most state-of-the-art interpolation approaches require subsequent refinement of the warped result to generate satisfying outputs, which drastically decreases the efficiency for multi-frame interpolation. In this work, we propose a fully differentiable Many-to-Many (M2M) splatting framework to interpolate frames efficiently. Specifically, given a frame pair, we estimate multiple bidirectional flows to directly forward warp the pixels to the desired time step, and then fuse any overlapping pixels. In doing so, each source pixel renders multiple target pixels and each target pixel can be synthesized from a larger area of visual context. This establishes a many-to-many splatting scheme with robustness to artifacts like holes. Moreover, for each input frame pair, M2M only performs motion estimation once and has a minuscule computational overhead when interpolating an arbitrary number of in-between frames, hence achieving fast multi-frame interpolation. We conducted extensive experiments to analyze M2M, and found that it significantly improves the efficiency while maintaining high effectiveness. Ping Hu 0001, Simon Niklaus, Stan Sclaroff, Kate Saenko |
CVPR | 1 |
| 2022 | Learning to Detect Every Thing in an Open World
Kuniaki Saito, Ping Hu 0001, Trevor Darrell, Kate Saenko |
ECCV (24) | 2 |
| 2022 | DualCoOp: Fast Adaptation to Multi-Label Recognition with Limited AnnotationsabstractSolving multi-label recognition (MLR) for images in the low-label regime is a challenging task with many real-world applications. Recent work learns an alignment between textual and visual spaces to compensate for insufficient image labels, but loses accuracy because of the limited amount of available MLR annotations. In this work, we utilize the strong alignment of textual and visual features pretrained with millions of auxiliary image-text pairs and propose \textit{Dual Context Optimization} (DualCoOp) as a unified framework for partial-label MLR and zero-shot MLR. \ours encodes positive and negative contexts with class names as part of the linguistic input (i.e. prompts). Since \ours only introduces a very light learnable overhead upon the pretrained vision-language framework, it can quickly adapt to multi-label recognition tasks that have limited annotations and even unseen classes. Experiments on standard multi-label recognition benchmarks across two challenging low-label settings demonstrate the advantages of our approach over state-of-the-art methods. Our code will be publicly available.Project page: https://cs-people.bu.edu/sunxm/DualCoOp/project.html Ximeng Sun, Ping Hu 0001, Kate Saenko |
NeurIPS | 2 |
| 2022 | Leveraging Geometric Structure for Label-Efficient Semi-Supervised Scene SegmentationabstractLabel-efficient scene segmentation aims to achieve effective per-pixel classification with reduced labeling effort. Recent approaches for this task focus on leveraging unlabelled images by formulating consistency regularization or pseudo labels for individual pixels. Yet most of these methods ignore the 3D geometric structures naturally conveyed by image scenes, which is free for enhancing training segmentation models with better discrimination of image details. In this work, we present a novel Geometric Structure Refinement (GSR) framework to explicitly exploit the geometric structures of image scenes to enhance the semi-supervised training of segmentation models. In the training phase, we generate initial dense pseudo labels based on fast and coarse annotations, and then utilize the free unsupervised 3D reconstruction of the image scene to calibrate the dense pseudo labels with more reliable details. With the calibrated pseudo groundtruth, we are able to conveniently train any existing image segmentation models without increasing the costs of annotations or modifying the models' architectures. Moreover, we explore different strategies for allocating labeling effort in semi-supervised scene segmentation, and find that a combination of finely-labeled samples and coarsely-labeled samples performs better than the traditional dense-fine only annotations. Extensive experiments on datasets including Cityscapes and KITTI are conducted to evaluate our proposed methods. The results demonstrate that GSR can be easily applied to boost the performance of existing models like PSPNet, DeepLabv3+, etc with reduced annotations. With half of the annotation effort, GSR achieves 99% of the accuracy of its fully supervised state-of-the-art counterparts. Ping Hu 0001, Stan Sclaroff, Kate Saenko |
IEEE Trans. Image Process. | 1 |
| 2020 | Temporally Distributed Networks for Fast Video Semantic SegmentationabstractWe present TDNet, a temporally distributed network designed for fast and accurate video semantic segmentation. We observe that features extracted from a certain high-level layer of a deep CNN can be approximated by composing features extracted from several shallower sub-networks. Leveraging the inherent temporal continuity in videos, we distribute these sub-networks over sequential frames. Therefore, at each time step, we only need to perform a lightweight computation to extract a sub-features group from a single sub-network. The full features used for segmentation are then recomposed by application of a novel attention propagation module that compensates for geometry deformation between frames. A grouped knowledge distillation loss is also introduced to further improve the representation power at both full and sub-feature levels. Experiments on Cityscapes, CamVid, and NYUD-v2 demonstrate that our method achieves state-of-the-art accuracy with significantly faster speed and lower latency. Ping Hu 0001, Fabian Caba Heilbron, Oliver Wang, Zhe Lin 0001, Stan Sclaroff, Federico Perazzi |
CVPR | 1 |
| 2020 | Uncertainty-Aware Learning for Zero-Shot Semantic SegmentationabstractZero-shot semantic segmentation (ZSS) aims to classify pixels of novel classes without training examples available. Recently, most ZSS methods focus on learning the visual-semantic correspondence to transfer knowledge from seen classes to unseen classes at the pixel level. Yet, few works study the adverse effects caused by the noisy and outlying training samples in the seen classes. In this paper, we identify this challenge and address it with a novel framework that learns to discriminate noisy samples based on Bayesian uncertainty estimation. Specifically, we model the network outputs with Gaussian and Laplacian distributions, with the variances accounting for the observation noise as well as the uncertainty of input samples. Learning objectives are then derived with the estimated variances playing as adaptive attenuation for individual samples in training. Consequently, our model learns more attentively from representative samples of seen classes while suffering less from noisy and outlying ones, thus providing better reliability and generalization toward unseen categories. We demonstrate the effectiveness of our framework through comprehensive experiments on multiple challenging benchmarks, and show that our method achieves significant accuracy improvement over previous approaches for large open-set segmentation. Ping Hu 0001, Stan Sclaroff, Kate Saenko |
NeurIPS | 1 |
| 2020 | DIPNet: Dynamic Identity Propagation Network for Video Object SegmentationabstractMany recent methods for semi-supervised Video Object Segmentation (VOS) have achieved good performance by exploiting the annotated first frame via one-shot fine-tuning or mask propagation. However, heavily relying on the first frame may weaken the robustness for VOS, since video objects can show large variations through time. In this work, we propose a Dynamic Identity Propagation Network (DIPNet) that adaptively propagates and accurately segments the video objects over time. To achieve this, DIPNet factors the VOS task at each time step into a dynamic propagation phase and a spatial segmentation phase. The former utilizes a novel identity representation to adaptively propagate objects’ reference information over time, which enhances the robustness to videos’ temporal variations. The segmentation phase uses the propagated information to tackle the object segmentation as an easier static image problem that can be optimized via light-weight fine-tuning on the first frame, thus reducing the computational cost. As a result, by optimizing these two components to complement each other, we can achieve a robust system for VOS. Evaluations on four benchmark datasets show that DIPNet provides state-of-the-art performance with time efficiency. Ping Hu 0001, Jun Liu 0036, Gang Wang 0012, Vitaly Ablavsky, Kate Saenko, Stan Sclaroff |
WACV | 1 |
| 2020 | Motion-Guided Cascaded Refinement Network for Video Object SegmentationabstractIn this work, we propose a motion-guided cascaded refinement network for video object segmentation. By assuming the foreground objects show different motion patterns from the background, for each video frame we apply an active contour model on optical flow to coarsely segment the foreground. The proposed Cascaded Refinement Network (CRN) then takes as guidance the coarse segmentation to generate an accurate segmentation in full resolution. In this way, the motion information and the deep CNNs can complement each other well to accurately segment the foreground objects from video frames. To deal with multi-instance cases, we extend our method with a spatial-temporal instance embedding model that further segments the foreground regions into instances and propagates instance labels. We further introduce a single-channel residual attention module in CRN to incorporate the coarse segmentation map as attention, which makes the network effective and efficient in both training and testing. We perform experiments on popular benchmarks and the results show that our method achieves state-of-the-art performance with high time efficiency. Ping Hu 0001, Gang Wang 0012, Xiangfei Kong, Jason Kuen, Yap-Peng Tan |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2018 | Motion-Guided Cascaded Refinement Network for Video Object SegmentationabstractDeep CNNs have achieved superior performance in many tasks of computer vision and image understanding. However, it is still difficult to effectively apply deep CNNs to video object segmentation(VOS) since treating video frames as separate and static will lose the information hidden in motion. To tackle this problem, we propose a Motion-guided Cascaded Refinement Network for VOS. By assuming the object motion is normally different from the background motion, for a video frame we first apply an active contour model on optical flow to coarsely segment objects of interest. Then, the proposed Cascaded Refinement Network(CRN) takes the coarse segmentation as guidance to generate an accurate segmentation of full resolution. In this way, the motion information and the deep CNNs can well complement each other to accurately segment objects from video frames. Furthermore, in CRN we introduce a Single-channel Residual Attention Module to incorporate the coarse segmentation map as attention, making our network effective and efficient in both training and testing. We perform experiments on the popular benchmarks and the results show that our method achieves state-of-the-art performance at a much faster speed. Ping Hu 0001, Gang Wang 0012, Xiangfei Kong, Jason Kuen, Yap-Peng Tan |
CVPR | 1 |
| 2018 | Recurrent Spatial Pyramid CNN for Optical Flow EstimationabstractOptical flow estimation plays an important role in many multimedia and computer vision tasks. Although great progress has been made in applying convolutional neural networks (CNNs) to estimate optical flow in recent works, it is still difficult for CNNs to generate optical flow with the desired effectiveness and efficiency. Compared to CNN-based methods, conventional variational methods normally perform to optimize an energy function and produce optical flow with more precise details. Inspired by the effectiveness of variational methods and deep CNNs, we propose a recurrent spatial pyramid (RecSPy) network for optical flow estimation. To deal with large displacements and to decrease the number of parameters, we formulate the spatial pyramid as a recurrent process, and adopt a CNN to refine optical flow at each spatial scale. Furthermore, to improve the results with more precise details, we propose an energy function that encodes structure and constancy constraints to help refine the optical flow at each spatial scale. The combination of the proposed RecSPy network and the proposed energy-based refinement enables our system to estimate optical flow effectively and efficiently. Experimental results on the benchmarks validate the effectiveness and efficiency of the proposed method. Ping Hu 0001, Gang Wang 0012, Yap-Peng Tan |
IEEE Trans. Multim. | 1 |
| 2017 | Deep Level Sets for Salient Object DetectionabstractDeep learning has been applied to saliency detection in recent years. The superior performance has proved that deep networks can model the semantic properties of salient objects. Yet it is difficult for a deep network to discriminate pixels belonging to similar receptive fields around the object boundaries, thus deep networks may output maps with blurred saliency and inaccurate boundaries. To tackle such an issue, in this work, we propose a deep Level Set network to produce compact and uniform saliency maps. Our method drives the network to learn a Level Set function for salient objects so it can output more accurate boundaries and compact saliency. Besides, to propagate saliency information among pixels and recover full resolution saliency map, we extend a superpixel-based guided filter to be a layer in the network. The proposed network has a simple structure and is trained end-to-end. During testing, the network can produce saliency maps by efficiently feedforwarding testing images at a speed over 12FPS on GPUs. Evaluations on benchmark datasets show that the proposed method achieves state-of-the-art performance. Ping Hu 0001, Bing Shuai, Jun Liu 0036, Gang Wang 0012 |
CVPR | 1 |
| 2017 | Global Context-Aware Attention LSTM Networks for 3D Action RecognitionabstractLong Short-Term Memory (LSTM) networks have shown superior performance in 3D human action recognition due to their power in modeling the dynamics and dependencies in sequential data. Since not all joints are informative for action analysis and the irrelevant joints often bring a lot of noise, we need to pay more attention to the informative ones. However, original LSTM does not have strong attention capability. Hence we propose a new class of LSTM network, Global Context-Aware Attention LSTM (GCA-LSTM), for 3D action recognition, which is able to selectively focus on the informative joints in the action sequence with the assistance of global contextual information. In order to achieve a reliable attention representation for the action sequence, we further propose a recurrent attention mechanism for our GCA-LSTM network, in which the attention performance is improved iteratively. Experiments show that our end-to-end network can reliably focus on the most informative joints in each frame of the skeleton sequence. Moreover, our network yields state-of-the-art performance on three challenging datasets for 3D action recognition. Jun Liu 0036, Gang Wang 0012, Ping Hu 0001, Ling-Yu Duan, Alex Chichung Kot |
CVPR | 3 |
| 2016 | Detecting Salient Objects via Color and Texture Compactness HypothesesabstractIn recent years, the object-level saliency detection has attracted much research attention, due to its usefulness in many high-level tasks. Existing methods are mostly based on the contrast hypothesis, which regards the regions with high contrast in a certain context as salient objects. Although the contrast hypothesis is effective in many scenarios, it cannot handle some difficult cases. As a remedy to address the weakness of contrast hypothesis, we propose a novel compactness hypothesis, which assumes salient regions are more compact than background from the perspectives of both color layout and texture layout. Based on the compactness hypotheses, we implement an effective object-level saliency detection method. In the proposed method, we first construct a weak saliency map based on the compact hypotheses, then collect samples from the weak saliency map to train a dedicated classifier. This classifier is applied on each individual pixel of the input image to produce a confidence score. Finally, the confidence scores are used to form a saliency map. This process is carried out at different scales, and the corresponding results are integrated into the formation of the final saliency map. The proposed approach is evaluated on eight benchmark data sets, where it delivers the competitive performance compared with the state-of-the-art methods. Ping Hu 0001, Weiqiang Wang 0001, Ke Lu 0002 |
IEEE Trans. Image Process. | 1 |
| 2015 | A novel binarization approach for text in imagesabstractAccurate recognition of scene text and overlaid text is still a challenging issue due to degradation and complex background, and text binarization is crucial for recognition accuracy. This paper presents an effective method to extract characters in images and video frames. Our method assumes that background pixels possess good spatial connectivity and high appearance similarity to boundary pixels in a cropped text string image. It first computes the confidence of pixels as text. Then the confidence map is exploited to partition text regions into characters. Further, each character region is clustered into different layers and background components are removed to generate candidate binarization results. The final result is obtained based on the scores of each layer. Our method is validated by better recognition rates and segmentation accuracy on the ICADR03 dataset and a big dataset of overlaid text. Ping Hu 0001, Weiqiang Wang 0001, Ke Lu 0002 |
ICIP | 1 |
| 2015 | Detecting Salient Objects via Spatial and Appearance Compactness HypothesesabstractObject-level saliency detection has been attracting a lot of attention, due to its potential enhancement in many high-level vision tasks. Many previous methods are based on the contrast hypothesis which regards the regions with high contrast in a certain context as salient. Although the contrast hypothesis is valid in many cases, it cannot handle some difficult cases. To make up for the weakness of contrast hypothesis, we propose a novel compactness hypothesis which assumes salient regions are more compact than background spatially and in appearance. Based on compactness hypotheses, we implement an effective object-level saliency detection method, which is demonstrated to be effective even in difficult cases. In addition, we present an adaptive multiple saliency maps fusion framework which can automatically select saliency maps of high quality according to three quality assessment rules. We evaluate the proposed method on four benchmark datasets and the comparable performance as the state-of-the-art methods has been achieved. Ping Hu 0001, Weiqiang Wang 0001, Ke Lu 0002 |
ACM Multimedia | 1 |