Hongyuan Zhu 0002

dblp:133/4527-2 · DBLP profile ↗
← Back
89ranked-venue papers
12as first author
62since 2021 · last 2026
0000-0001-5177-8320ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 55 · 10 first-author · 37 since 2021Artificial intelligence and machine learning · 53 · 6 first-author · 38 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Systems, architecture and hardware · 2 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1
YearPublicationVenuePosition
2026 Your AI-Generated Image Detector Can Secretly Achieve SOTA Accuracy, If Calibrated
abstract
Despite being trained on balanced datasets, existing AI-generated image detectors often exhibit systematic bias at test time, frequently misclassifying fake images as real. We hypothesize that this behavior stems from distributional shift in fake samples and implicit priors learned during training. Specifically, models tend to overfit to superficial artifacts that do not generalize well across different generation methods, leading to a misaligned decision threshold when faced with test-time distribution shift. To address this, we propose a theoretically grounded post-hoc calibration framework based on Bayesian decision theory. In particular, we introduce a learnable scalar correction to the model’s logits, optimized on a small validation set from the target distribution while keeping the backbone frozen. This parametric adjustment compensates for distributional shift in model output, realigning the decision boundary even without requiring ground-truth labels. Experiments on challenging benchmarks show that our approach significantly improves robustness without retraining, offering a lightweight and principled solution for reliable and adaptive AI-generated image detection in the open world.
Muli Yang, Gabriel James Goenawan, Henan Wang, Huaiyuan Qin, Yanhua Yang, Fen Fang, Ying Sun 0001, Joo-Hwee Lim, Hongyuan Zhu 0002
AAAI10
2026 Toward Accurate Procedure Planning in Instructional Videos: Visual State Generation Helps Task-Selective Diffusion
abstract
Procedure planning in instructional videos entails predicting an action sequence that transitions a given start state to a desired goal state. This task is particularly challenging due to two key sources of uncertainty: limited visual observations and an enormous decision space. The former results in multiple plausible plan variations due to missing intermediate visual states, while the latter complicates prediction by requiring selection from a large set of potential actions. Unlike prior work that addresses these issues implicitly, we propose an explicit solution. To mitigate the first challenge, we employ image generation models to synthesize diverse intermediate visual states using various text prompts, followed by a prompt selection module integrated within a diffusion model. To tackle the second challenge, we introduce a task-selective diffusion model that applies a task-specific mask to constrain the action space. As the effectiveness of this mask depends on accurate task classification, we further enhance visual representation by leveraging pre-trained vision-language models to generate action-aware, text-enriched multimodal embeddings. Extensive experiments on three benchmark datasets validate the superior performance of our proposed approach.
Fen Fang, Muli Yang, Min Wu 0008, Yanhua Yang, Qianli Xu, Joo-Hwee Lim, Xulei Yang, Hongyuan Zhu 0002
IEEE Trans. Pattern Anal. Mach. Intell.8
2026 Counterfactual Risk Minimization for Out-of-Distribution Generalization
abstract
The out-of-distribution (OOD) property in data is deemed as one main challenge hindering the generalization ability of machine learning algorithms. However, the underlying reasons for this property remain an intriguing and open question that has yet to be fully understood. In this paper, we seek to enhance our understanding of the OOD phenomenon by framing it as a problem of distribution shift and addressing it through two complementary causal perspectives. The first is a generative causal view that elucidates the data generation process. We introduce a novel three-dimensional coordinate system to represent three fundamental distribution shifts, illustrating their role in various OOD generalization problems. The second is an anti-causal view that focuses on the model learning process. We develop an effective approach dubbed Counterfactual Risk Minimization (CRM) to address arbitrary distribution shifts in a unified framework. Additionally, we introduce a new multi-domain visual recognition dataset called CONA to facilitate further exploration of OOD generalization. We conduct evaluations of CRM alongside several state-of-the-art competitors on four benchmark datasets under the three distribution shifts. The results not only affirm CRM's superiority but also shed light on potential future directions. Code and data: https://github.com/muliyangm/CRM.
Yanhua Yang, Muli Yang, Henan Wang, Cheng Deng 0002, Hongyuan Zhu 0002
IEEE Trans. Image Process.6
2026 Rendered 2D Semantic and Generative Priors Guided 3D Multi-Object Grounding
abstract
3D multi-object visual grounding aims to identify and localize all objects in a 3D scene that correspond to a given text description. Unlike traditional single-object grounding, this task presents additional challenges as point clouds inherently lack fine-grained details, making it difficult to capture subtle object features and contextual information. Moreover, textual descriptions are inherently limited in perceiving and understanding complex 3D environments, especially in scenarios with high object similarity or intricate spatial arrangements. To tackle the above challenges, we propose SGMG, a Rendered 2D Semantic and Generative priors guided 3D Multi-object Grounding Framework. The SGMG framework introduces two key innovations that work cohesively to enhance grounding accuracy. First, a Generative-Assistant(GA) Module leverages the capabilities of a generative model to provide enriched scene prior information and capture the fine-grained scene details. Second, the Semantic-Augment Fusion(SAF) Module is designed to improve the representation of text features from the vision features, thereby boosting the accuracy of multimodal information interactions. Furthermore, we introduce a multi-level fusion mechanism, ensuring that semantic and spatial relationships between objects are preserved and effectively leveraged during the grounding process. Experimental results demonstrate that SGMG achieves state-of-the-art performance in multi-object 3D grounding and competitive results in traditional single-object tasks, highlighting its effectiveness in diverse scenarios.
Peng Guo 0011, Hongyuan Zhu 0002, Hancheng Ye, Yihang Yang, Fukun Yin, Tao Chen 0003
IEEE Trans. Multim.2
2025 Balancing Privacy and Performance: A Many-in-One Approach for Image Anonymization
abstract
The effective utilization of data through Deep Neural Networks (DNNs) has profoundly influenced various aspects of society. The growing demand for high-quality, particularly personalized, data has spurred research efforts to prevent data leakage and protect privacy in recent years. Early privacy-preserving methods primarily relied on instance-wise modifications, such as erasing or obfuscating essential features for de-identification. However, this approach highlights an inherent trade-off: minimal modification offers insufficient privacy protection, while excessive modification significantly degrades task performance. In this paper, we propose a novel Recombining for Obfuscation (FRO) approach to address this trade-off. Unlike existing methods that generate one anonymized instance by perturbing the original data on a one-to-one basis, our FRO approach generates an anonymized instance by reassembling mixed ID-related features from multiple original data sources on a many-in-one basis. Instead of introducing additional noise for de-identification, our approach leverages the existing non-polluted features from other instances to anonymize data. Extensive experiments on identity identification tasks demonstrate that FRO outperforms previous state-of-the-art methods, not only in utility performance but also in visual anonymization.
Xuemei Jia, Jiawei Du 0002, Hui Wei 0004, Ruinian Xue, Zheng Wang 0007, Hongyuan Zhu 0002, Jun Chen 0001
AAAI6
2025 Detecting Open World Objects via Partial Attribute Assignment
abstract
Despite being trained on massive data, today’s vision foundation models still fall short in detecting open world objects. Apart from recognizing known objects from training, a successful Open World Object Detection (OWOD) system must also be able to detect unknown objects never seen before, without confusing them with the backgrounds. Unlike prevailing prior works that rely on probability models to learn "objectness", we focus on learning fine-grained, class-agnostic attributes, allowing the detection of both known and unknown objects in an explainable manner. In this paper, we propose Partial Attribute Assignment (PASS), aiming to automatically select and optimize a small, relevant subset of attributes from a large attribute pool. Specifically, we model attribute selection as a Partial Optimal Transport (POT) problem between known visual objects and the attribute pool, in which more relevant attributes signify more transported mass. PASS follows a curriculum schedule that progressively selects and optimizes a targeted subset of attributes during training, promoting stability and accuracy. Our method enjoys end-to-end optimization by minimizing the POT distance and the classification loss on known visual objects, demonstrating high training efficiency and superior OWOD performance among extensive experimental evaluations.‡
Muli Yang, Gabriel James Goenawan, Huaiyuan Qin, Xi Peng 0001, Yanhua Yang, Hongyuan Zhu 0002
CVPR7
2025 Object-Level Correlation for Few-Shot Segmentation
abstract
Few-shot semantic segmentation (FSS) aims to segment objects of novel categories in the query images given only a few annotated support samples. Existing methods primarily build the image-level correlation between the support target object and the entire query image. However, this correlation contains the hard pixel noise, \textit{i.e.}, irrelevant background objects, that is intractable to trace and suppress, leading to the overfitting of the background. To address the limitation of this correlation, we imitate the biological vision process to identify novel objects in the object-level information. Target identification in the general objects is more valid than in the entire image, especially in the low-data regime. Inspired by this, we design an Object-level Correlation Network (OCNet) by establishing the object-level correlation between the support target object and query general objects, which is mainly composed of the General Object Mining Module (GOMM) and Correlation Construction Module (CCM). Specifically, GOMM constructs the query general object feature by learning saliency and high-level similarity cues, where the general objects include the irrelevant background objects and the target foreground object. Then, CCM establishes the object-level correlation by allocating the target prototypes to match the general object feature. The generated object-level correlation can mine the query target feature and suppress the hard pixel noise for the final prediction. Extensive experiments on PASCAL-${5}^{i}$ and COCO-${20}^{i}$ show that our model achieves the state-of-the-art performance.
Chunlin Wen, Yu Zhang 0004, Hongyuan Zhu 0002, Xiu-Shen Wei, Zhiqiang Kou, Shuzhou Sun
ICCV4
2025 Deep Unsupervised Hashing via External Guidance
abstract
Recently, deep unsupervised hashing has gained considerable attention in image retrieval due to its advantages in cost-free data labeling, computational efficiency, and storage savings. Although existing methods achieve promising performance by leveraging inherent visual structures within the data, they primarily focus on learning discriminative features from unlabeled images through limited internal knowledge, resulting in an intrinsic upper bound on their performance. To break through this intrinsic limitation, we propose a novel method, called Deep Unsupervised Hashing with External Guidance (DUH-EG), which incorporates external textual knowledge as semantic guidance to enhance discrete representation learning. Specifically, our DUH-EG: i) selects representative semantic nouns from an external textual database by minimizing their redundancy, then matches images with them to extract more discriminative external features; and ii) presents a novel bidirectional contrastive learning mechanism to maximize agreement between hash codes in internal and external spaces, thereby capturing discrimination from both external and intrinsic structures in Hamming space. Extensive experiments on four benchmark datasets demonstrate that our DUH-EG remarkably outperforms existing state-of-the-art hashing methods.
Qihong Song, XitingLiu, Hongyuan Zhu 0002, Joey Tianyi Zhou, Xi Peng 0001, Peng Hu 0002
ICML3
2025 Multi-Modality Test-Time Adaptation for Semantic Segmentation in Robotic Perception
abstract
Test-Time Adaptation (TTA) adjusts pre-trained models in unlabeled unseen environments during the test phase, making it more practical for robotic applications. However, the constant changes of the physical world create significant domain gaps between the received data during robot deployment and the source data used for training. In addition, existing methods mainly focus on a single modality, e.g., RGB images, limiting the application of these methods in multi-modality input scenarios. In this work, we propose a Deep Multi-modality Aggregation Test-time Adaptation (DMATA) method to address the above mentioned issues. To prevent the domain shifts from disrupting the adaptation process, we first propose a Momentum-based Teacher-Student (MTS) framework. Since the teacher model and the student model contain complementary information, we design an Uncertainty-Guide (UG) feature fusion block to fuse the representation of the teacher model and student model of each modality. Finally, we introduce a 3D-Guide-2D (3G2) feature fusion block to leverage spatial information for enhancing 2D feature representation. Extensive experiments across three scenarios, including sensor-to-sensor, day-to-night, and city-to-city, demonstrate the effectiveness of our method in TTA multi-modality semantic segmentation tasks. Notably, under the scenario of sensor-to-sensor adaptation, our proposed DMATA obtains an$m$IoU of 54.2%, which is superior to the state-of-the-art test-time adaptation method.
Yan Liu 0043, Hongyuan Zhu 0002, Ye Zhang 0037, Yinjie Lei, Yulan Guo
ICRA2
2025 Robust Cross-modal Alignment Learning for Cross-Scene Spatial Reasoning and Grounding
abstract
Grounding target objects in 3D environments via natural language is a fundamental capability for autonomous agents to successfully fulfill user requests. Almost all existing works typically assume that the target object lies within a known scene and focus solely on in-scene localization. In practice, however, agents often encounter unknown or previously visited environments and need to search across a large archive of scenes to ground the described object, thereby invalidating this assumption. To address this, we reveal a novel task called Cross-Scene Spatial Reasoning and Grounding (CSSRG), which aims to locate a described object anywhere across an entire collection of 3D scenes rather than predetermined scenes. Due to the difference from existing 3D visual grounding, CSSRG poses two challenges: the prohibitive cost of exhaustively traversing all scenes and more complex cross-modal spatial alignment. To address the challenges, we propose a Cross-Scene 3D Object Reasoning Framework (CoRe), which adopts a matching-then-grounding pipeline to reduce computational overhead. Specifically, CoRe consists of i) a Robust Text-Scene Aligning (RTSA) module that learns global scene representations for robust alignment between object descriptions and the corresponding 3D scenes, enabling efficient retrieval of candidate scenes; and ii) a Tailored Word-Object Associating (TWOA) module that establishes fine-grained alignment between words and target objects to filter out redundant context, supporting precise object-level reasoning and alignment. Additionally, to benchmark CSSRG, we construct a new CrossScene-RETR dataset and evaluation protocol tailored for cross-scene grounding. Extensive experiments across four multimodal datasets demonstrate that CoRe dramatically reduces computational overhead while showing superiority in both scene retrieval and object grounding. Code is available at https://github.com/Yangl1nFeng/CoRe.
Yanglin Feng, Hongyuan Zhu 0002, Dezhong Peng, Xi Peng 0001, Xiaomin Song, Peng Hu 0002
NeurIPS2
2025 Dynamic Adapter Tuning for Long-Tailed Class-Incremental Learning
abstract
Long-tailed class-incremental learning (LT-CIL) aims to learn new classes continuously from a long-tailed data stream, while simultaneously dealing with challenges such as imbalanced learning of tail classes and catastrophic for-getting. To address these challenges, most existing methods employ a two-stage strategy by initializing model training from scratch with further balanced knowledge driven cali-bration. This strategy faces challenges in deriving discrim-inative features from cold-started backbones for the long-tailed distribution of data, consequently leading to relatively diminished performance. In this paper, with the pow-erful feature extraction capability of pre-trained foundation models, we have achieved a one-stage approach that de-livers superior performance. Specifically, we propose Dy-namic Adapter Tuning (DAT), which employs a dynamic adapter cache mechanism to adapt a pre-trained model to learn tasks sequentially. The adapter in the cache is either dynamically selected or created according to task similar-ity, and further compactified with the new task's adapter to mitigate cross-task and cross-class gaps in LT-CIL, sig-nificantly alleviating catastrophic forgetting and imbalance learning issues, respectively. With extensive experimental validation, our method consistently achieves state-of-the-art performance under the challenging LT-CIL setting.
Yanan Gu, Muli Yang, Xu Yang 0019, Hongyuan Zhu 0002, Gabriel James Goenawan, Cheng Deng 0002
WACV5
2025 Consistent Prompt Tuning for Generalized Category Discovery
Muli Yang, Yanan Gu, Cheng Deng 0002, Hanwang Zhang, Hongyuan Zhu 0002
Int. J. Comput. Vis.6
2025 Correction: Consistent Prompt Tuning for Generalized Category Discovery
Muli Yang, Yanan Gu, Cheng Deng 0002, Hanwang Zhang, Hongyuan Zhu 0002
Int. J. Comput. Vis.6
2025 Text-Driven Adaptive Semantic Alignment Network for Cross-Scene Hyperspectral Image Classification
abstract
Land cover in different scenes generally exhibits scene-invariant category semantic, typically represented and described consistently in a textual modality. Traditional cross-scene classification methods often treat categories as discrete class labels, neglecting their semantic information, or use category names merely as auxiliary textual modalities to enhance the discriminative representations of land cover. However, the cross-scene consistency of category semantic for land cover remains underexplored and underutilized. To address this issue, the text-driven adaptive semantic alignment network (TASA-Net) is proposed in this article for cross-scene hyperspectral image classification (HSIC). TASA-Net employs hand-crafted template prompts for stable category descriptions and vision-guided fine semantic prompts (VG-FSPs) for dynamic scene adaptation. Through a dual-gated adaptive mechanism, TASA-Net optimally weights coarse- and fine-grained semantics in a shared space, ensuring stable yet discriminative semantic representation. Additionally, cross-modal semantic alignment projects visual features into the shared semantic space, while a soft alignment strategy dynamically adjusts category correlations to enhance intraclass consistency and mitigate domain shifts. Ultimately, by leveraging text-driven semantic consistency representation, TASA-Net achieves zero-shot cross-scene transfer for unsupervised classification. Experiments demonstrate superior performance across multiple hyperspectral datasets, validating the critical role of textual modality in enhancing model robustness and cross-scene generalization ability.
Wenzhen Wang, Fang Liu 0034, Hongyuan Zhu 0002, Liang Xiao 0001
IEEE Trans. Geosci. Remote. Sens.3
2025 Object Adaptive Self-Supervised Dense Visual Pre-Training
abstract
Self-supervised visual pre-training models have achieved significant success without employing expensive annotations. Nevertheless, most of these models focus on iconic single-instance datasets (e.g. ImageNet), ignoring the insufficient discriminative representation for non-iconic multi-instance datasets (e.g. COCO). In this paper, we propose a novel Object Adaptive Dense Pre-training (OADP) method to learn the visual representation directly on the multi-instance datasets (e.g., PASCAL VOC and COCO) for dense prediction tasks (e.g., object detection and instance segmentation). We present a novel object-aware and learning-adaptive random view augmentation to focus the contrastive learning to enhance the discrimination of object presentations from large to small scale during different learning stages. Furthermore, the representations across different scale and resolutions are integrated so that the method can learn diverse representations. In the experiment, we evaluated OADP pre-trained on PASCAL VOC and COCO. Results show that our method has better performances than most existing state-of-the-art methods when transferring to various downstream tasks, including image classification, object detection, instance segmentation and semantic segmentation.
Yu Zhang 0004, Hongyuan Zhu 0002, Siya Mi, Xi Peng 0001, Xin Geng 0001
IEEE Trans. Image Process.3
2025 PointCloud-Text Matching: Benchmark Dataset and Baseline
abstract
In this paper, we present and study a new instance-level retrieval task: PointCloud-Text Matching (PTM), which aims to identify the exact cross-modal instance that matches a given point-cloud query or text query. PTM has potential applications in various scenarios, such as indoor/urban-canyon localization and scene retrieval. However, there is a lack of suitable and targeted datasets for PTM in practice. To address this issue, we present a new PTM benchmark dataset, namely SceneDepict-3D2T. We observe that the data poses significant challenges due to its inherent characteristics, such as the sparsity, noise, or disorder of point clouds and the ambiguity, vagueness, or incompleteness of texts, which render existing cross-modal matching methods ineffective for PTM. To overcome these challenges, we propose a PTM baseline, namedRobust PointCloud-TextMatching method (RoMa). RoMa consists of two key modules: a Dual Attention Perception module (DAP) and a Robust Negative Contrastive Learning module (RNCL). Specifically, DAP leverages token-level and feature-level attention mechanisms to adaptively focus on useful local and global features, and aggregate them into common representations, thereby reducing the adverse impact of noise and ambiguity. To handle noisy correspondence, RNCL enhances robustness against mismatching by dividing negative pairs into clean and noisy subsets and assigning them forward and reverse optimization directions, respectively. We conduct extensive experiments on our benchmarks and demonstrate the superiority of our RoMa.
Yanglin Feng, Dezhong Peng, Hongyuan Zhu 0002, Xi Peng 0001, Peng Hu 0002
IEEE Trans. Multim.4
2025 WI3D: Weakly Incremental 3D Detection via Vision Foundation Models
abstract
Class-incremental 3D object detection demands a 3D detector tolocateandrecognizenovel categories in a stream fashion while preserving its base detection ability. However, existing methods require delicate 3D annotations for learning novel categories, resulting in significant labeling costs. To this end, we explore a label-efficient approach calledWeaklyIncremental3DDetection (WI3D), which teaches a 3D detector to learn incrementally with off-the-shelf vision foundation models. We propose a novel dual-teaching framework incorporating both intra-modal and inter-modal knowledge from pseudo labels and feature space. Specifically, our framework features a class-agnostic pseudo-label refinement module, designed for the generation of high-quality 3D pseudo labels. This module is built on a lightweight transformer that models the spatial relationships between pseudo labels and their interactions with rich contextual information in point clouds. Additionally, we introduce a cross-modal knowledge transfer module to enhance the representation learning of novel classes, along with a reweighting knowledge distillation strategy that dynamically assesses and distills knowledge from previously learned categories. Extensive experiments show that our approach can efficiently learn novel concepts while preserving knowledge of base classes in WI3D scenarios, and surpass baseline approaches on both SUN-RGBD and ScanNet.
Mingsheng Li, Sijin Chen, Shengji Tang, Hongyuan Zhu 0002, Yanyan Fang, Xin Chen 0040, Zhuoyuan Li 0006, Fukun Yin, Tao Chen 0003
IEEE Trans. Multim.4
2025 SF-City: A Source-Free Domain Adaptation Method for City-Scale Point Cloud Semantic Segmentation
abstract
City-scale point cloud semantic segmentation is an important yet challenging task. Despite progress, existing methods rely heavily on point-wise annotations. An alternative solution is to apply the Unsupervised Domain Adaptation (UDA) approach. Recently, the 2D foundation model has achieved significant progress with training with internet-scale images. Therefore, adapting 2D foundation models to 3D City-scale point clouds is an attempting idea. Due to the data protection and storage issue, 2D source domain data is typically unavailable. Thus, we focus on Source-Free Domain Adaptation (SFDA) and propose a Source-Free City-scale point cloud semantic segmentation method, namely SF-City. Our method leverages knowledge from 2D pre-trained models to generate point-wise pseudo labels for training a 3D semantic segmentation network. We convert point clouds into remote-sensing-like images using Bird's-Eye-View (BEV) projection. However, directly using source models for pseudo label generation is hindered by domain gaps such as viewpoint variations, concept divergences, and geometry loss. To tackle these problems, we propose a Multi-scale Content Feature Extractor (MCFE) to extract holistic and contextual feature representations. Then, an Uncertainty-guided Inter-Model Feature Integrator (UIFI) is introduced to integrate inherent knowledge across source models. Furthermore, the Geometric-guided Pseudo Label Generator (GPLG) is leveraged to introduce geometric information to regulate pseudo labels. Through extensive experiments on two public benchmarks, SF-City demonstrates superior performance, achieving an mIoU of 28.8% on the SensatUrban dataset, outperforming recent state-of-the-art methods CLIPFO3D by about 6.3%.
Yan Liu 0043, Hongyuan Zhu 0002, Yinjie Lei, Hao Liu 0061, Yun Pei 0001, Yulan Guo
IEEE Trans. Multim.2
2025 Evaluating Self-Supervised Learning for WiFi CSI-Based Human Activity Recognition
abstract
With the advancement of the Internet of Things, WiFi Channel State Information (CSI)-based Human Activity Recognition (HAR) has garnered increasing attention from both academic and industrial communities. However, the scarcity of labeled data remains a prominent challenge in CSI-based HAR, primarily due to privacy concerns and the incomprehensibility of CSI data. Concurrently, Self-Supervised Learning (SSL) has emerged as a promising approach for addressing the dilemma of insufficient labeled data. In this article, we undertake a comprehensive inventory and analysis of different categories of SSL algorithms, encompassing both previously studied and unexplored approaches within the field. We provide an in-depth investigation and evaluation of SSL algorithms in the context of WiFi CSI-based HAR, utilizing publicly available datasets that encompass various tasks and environmental settings. To ensure relevance to real-world applications, we design experiment settings aligned with specific requirements. Furthermore, our experimental findings uncover several limitations and blind spots in existing work, shedding light on the barriers that need to be addressed before SSL can be effectively deployed in real-world WiFi-based HAR applications. Our results also serve as practical guidelines and provide valuable insights for future research endeavors in this field.
Jiangtao Wang 0001, Hongyuan Zhu 0002, Dingchang Zheng
ACM Trans. Sens. Networks3
2024 PrefAce: Face-Centric Pretraining with Self-Structure Aware Distillation
abstract
Video-based facial analysis is important for autonomous agents to understand human expressions and sentiments. However, limited labeled data is available to learn effective facial representations. This paper proposes a novel self-supervised face-centric pretraining framework, called PrefAce, which learns transferable video facial representation without labels. The self-supervised learning is performed with an effective landmark-guided global-local tube distillation. Meanwhile, a novel instance-wise update FaceFeat Cache is built to enforce more discriminative and diverse representations for downstream tasks. Extensive experiments demonstrate that the proposed framework learns universal instance-aware facial representations with fine-grained landmark details from videos. The point is that it can transfer across various facial analysis tasks, e.g., Facial Attribute Recognition (FAR), Facial Expression Recognition (FER), DeepFake Detection (DFD), and Lip Synchronization (LS). Our framework also outperforms the state-of-the-art on various downstream tasks, even in low data regimes. Code is available at https://github.com/siyuan-h/PrefAce.
Zheng Wang 0007, Peng Hu 0002, Xi Peng 0001, Hongyuan Zhu 0002, Yew-Soon Ong
AAAI6
2024 Contributing Dimension Structure of Deep Feature for Coreset Selection
abstract
Coreset selection seeks to choose a subset of crucial training samples for efficient learning. It has gained traction in deep learning, particularly with the surge in training dataset sizes. Sample selection hinges on two main aspects: a sample's representation in enhancing performance and the role of sample diversity in averting overfitting. Existing methods typically measure both the representation and diversity of data based on similarity metrics, such as L2-norm. They have capably tackled representation via distribution matching guided by the similarities of features, gradients, or other information between data. However, the results of effectively diverse sample selection are mired in sub-optimality. This is because the similarity metrics usually simply aggregate dimension similarities without acknowledging disparities among the dimensions that significantly contribute to the final similarity. As a result, they fall short of adequately capturing diversity. To address this, we propose a feature-based diversity constraint, compelling the chosen subset to exhibit maximum diversity. Our key lies in the introduction of a novel Contributing Dimension Structure (CDS) metric. Different from similarity metrics that measure the overall similarity of high-dimensional features, our CDS metric considers not only the reduction of redundancy in feature dimensions, but also the difference between dimensions that contribute significantly to the final similarity. We reveal that existing methods tend to favor samples with similar CDS, leading to a reduced variety of CDS types within the coreset and subsequently hindering model performance. In response, we enhance the performance of five classical selection methods by integrating the CDS constraint. Our experiments on three datasets demonstrate the general effectiveness of the proposed method in boosting existing methods.
Zhijing Wan, Zhixiang Wang 0001, Yuran Wang 0003, Zheng Wang 0007, Hongyuan Zhu 0002, Shin'ichi Satoh 0001
AAAI5
2024 LL3DA: Visual Interactive Instruction Tuning for Omni-3D Understanding, Reasoning, and Planning
abstract
Recent progress in Large Multimodal Models (LMM) has opened up great possibilities for various applications in the field of human-machine interactions. However, developing LMMs that can comprehend, reason, and plan in complex and diverse 3D environments remains a challenging topic, especially considering the demand for understanding permutation-invariant point cloud representations of the 3D scene. Existing works seek help from multi-view images by projecting 2D features to 3D space, which inevitably leads to huge computational overhead and performance degradation. In this paper, we present LL3DA, a Large Language 3D Assistant that takes point cloud as the direct input and responds to both text instructions and visual interactions. The additional visual interaction enables LMMs to better comprehend human interactions with the 3D environment and further remove the ambiguities within plain texts. Experiments show that LL3DA achieves remarkable results and surpasses various 3D vision-language models on both 3D Dense Captioning and 3D Question Answering.
Sijin Chen, Xin Chen 0040, Chi Zhang 0007, Mingsheng Li, Gang Yu 0002, Hao Fei 0001, Hongyuan Zhu 0002, Jiayuan Fan 0001, Tao Chen 0003
CVPR7
2024 M3DBench: Towards Omni 3D Assistant with Interleaved Multi-modal Instructions
Mingsheng Li, Xin Chen 0040, Chi Zhang 0007, Sijin Chen, Hongyuan Zhu 0002, Fukun Yin, Zhuoyuan Li 0006, Gang Yu 0002, Tao Chen 0003
ECCV (58)5
2024 Direct Distillation Between Different Domains
Jialiang Tang, Shuo Chen 0003, Gang Niu 0001, Hongyuan Zhu 0002, Joey Tianyi Zhou, Chen Gong 0002, Masashi Sugiyama
ECCV (80)4
2024 G-Former: A Grouping Transformer for Weakly Supervised Point Cloud Segmentation
abstract
Recent advancements in weakly supervised point cloud semantic segmentation have diminished the reliance on extensive annotations, thereby enhancing the efficacy of understanding the real-world environment. However, existing approaches, such as PSD [25], SQN [5] and OTOC [10], often overlook the valuable global class-related prior knowledge present in point clouds beyond the scope of labels. To fully leverage this prior knowledge, which suggests that points of the same class should be close in feature space and each class should have a representative feature, we propose G-Former, a grouping transformer model. G-Former incorporates the idea of clustering into the overall model by defining clusters aligned to classes and assigning learnable tensors as cluster centers. Points are then grouped into these clusters based on the similarity of their features to the cluster centers. The core components of G-Former include a Hierarchy Cluster Structure (HCS) and a Grouping Module (GM). The former consists of two sets of clusters, one for classes while the other serves as a middle layer to help class clusters handle large-scale point features. The latter facilitates grouping the point cloud into different clusters. With the help of the grouping transformer model, G-Former further proposes a series of cluster center constraints to augment inter-class distances and diminish intra-class distances to enhance the discriminability of points. Experimental results on ScanNet v2 and S3DIS datasets demonstrate that G-Former outperforms previous methods with limited labels (0.1% or 1%) by a significant margin and is even comparable to fully supervised methods.
Zehan Huang, Fukun Yin, Jiayuan Fan 0001, Xin Chen 0040, Hongyuan Zhu 0002, Bin Wang 0008, Tao Chen 0003
IJCNN5
2024 Robust Variational Contrastive Learning for Partially View-unaligned Clustering
abstract
Although multi-view learning has achieved remarkable progress over the past decades, most existing methods implicitly assume that all views (or modalities) are well-aligned. In practice, however, collecting fully aligned views is challenging due to complexities and discordances in time and space, resulting in the Partially View-unaligned Problem (PVP), such as audio-video asynchrony caused by network congestion. While some methods are proposed to align the unaligned views by learning view-invariant representations, almost all of them overlook specific information across different views for complementarity, limiting performance improvement. To address these problems, we propose a robust framework, dubbed VariatIonal ConTrAstive Learning (VITAL), designed to learn both common and specific information simultaneously. To be specific, each data sample is first modeled as a Gaussian distribution in the latent space, where the mean estimates the most probable common information, while the variance indicates view-specific information. Second, by using variational inference, VITAL conducts intra- and inter-view contrastive learning to preserve common and specific semantics in the distribution representations, thereby achieving comprehensive perception. As a result, the common representation (mean) could be used to guide category-level realignment, while the specific representation (variance) complements sample semantic information, thereby boosting overall performance. Finally, considering the abundance of False Negative Pairs (FNPs) generated by unsupervised contrastive learning, we propose a robust loss function that seamlessly incorporates FNP rectification into the contrastive learning paradigm. Empirical evaluations on eight benchmark datasets reveal that VITAL outperforms ten state-of-the-art deep clustering baselines, demonstrating its efficacy in both partially and fully aligned scenarios. The Code is available at https://github.com/He-Changhao/2024-MM-VITAL.
Changhao He, Hongyuan Zhu 0002, Peng Hu 0002, Xi Peng 0001
ACM Multimedia2
2024 Synergistic Dual Spatial-aware Generation of Image-to-text and Text-to-image
abstract
In the visual spatial understanding (VSU) field, spatial image-to-text (SI2T) and spatial text-to-image (ST2I) are two fundamental tasks that appear in dual form. Existing methods for standalone SI2T or ST2I perform imperfectly in spatial understanding, due to the difficulty of 3D-wise spatial feature modeling. In this work, we consider modeling the SI2T and ST2I together under a dual learning framework. During the dual framework, we then propose to represent the 3D spatial scene features with a novel 3D scene graph (3DSG) representation that can be shared and beneficial to both tasks. Further, inspired by the intuition that the easier 3D$\to$image and 3D$\to$text processes also exist symmetrically in the ST2I and SI2T, respectively, we propose the Spatial Dual Discrete Diffusion (SD$^3$) framework, which utilizes the intermediate features of the 3D$\to$X processes to guide the hard X$\to$3D processes, such that the overall ST2I and SI2T will benefit each other. On the visual spatial understanding dataset VSD, our system outperforms the mainstream T2I and I2T methods significantly. Further in-depth analysis reveals how our dual learning strategy advances.
Yu Zhao 0043, Hao Fei 0001, Xiangtai Li, Libo Qin 0004, Jiayi Ji, Hongyuan Zhu 0002, Meishan Zhang, Min Zhang 0005, Jianguo Wei
NeurIPS6
2024 Revisiting 3D visual grounding with Context-aware Feature Aggregation
Peng Guo 0011, Hongyuan Zhu 0002, Hancheng Ye, Taihao Li, Tao Chen 0003
Neurocomputing2
2024 Vote2Cap-DETR++: Decoupling Localization and Describing for End-to-End 3D Dense Captioning
abstract
3D dense captioning requires a model to translate its understanding of an input 3D scene into several captions associated with different object regions. Existing methods adopt a sophisticated "detect-then-describe" pipeline, which builds explicit relation modules upon a 3D detector with numerous hand-crafted components. While these methods have achieved initial success, the cascade pipeline tends to accumulate errors because of duplicated and inaccurate box estimations and messy 3D scenes. In this paper, we first propose Vote2Cap-DETR, a simple-yet-effective transformer framework that decouples the decoding process of caption generation and object localization through parallel decoding. Moreover, we argue that object localization and description generation require different levels of scene understanding, which could be challenging for a shared set of queries to capture. To this end, we propose an advanced version, Vote2Cap-DETR++, which decouples the queries into localization and caption queries to capture task-specific features. Additionally, we introduce the iterative spatial refinement strategy to vote queries for faster convergence and better localization performance. We also insert additional spatial information to the caption head for more accurate descriptions. Without bells and whistles, extensive experiments on two commonly used datasets, ScanRefer and Nr3D, demonstrate Vote2Cap-DETR and Vote2Cap-DETR++ surpass conventional "detect-then-describe" methods by a large margin.
Sijin Chen, Hongyuan Zhu 0002, Mingsheng Li, Xin Chen 0040, Peng Guo 0011, Yinjie Lei, Gang Yu 0002, Taihao Li, Tao Chen 0003
IEEE Trans. Pattern Anal. Mach. Intell.2
2024 Blessing few-shot segmentation via semi-supervised learning with noisy support images
Runtong Zhang, Hongyuan Zhu 0002, Hanwang Zhang, Chen Gong 0002, Joey Tianyi Zhou, Fanman Meng
Pattern Recognit.2
2024 Multi-View Vision Fusion Network: Can 2D Pre-Trained Model Boost 3D Point Cloud Data-Scarce Learning?
abstract
Point cloud based 3D deep model has wide applications in many applications such as autonomous driving, house robot, etc. Inspired by the recent prompt learning in natural language processing, this work proposes a novel Multi-view Vision Fusion Network (MvNet) for few-shot 3D point cloud classification. MvNet investigates the possibility of leveraging the off-the-shelf 2D pre-trained models to achieve the few-shot classification, which can alleviate the over-dependence issue of the existing baseline models towards the large-scale annotated 3D point cloud data. Specifically, MvNet first encodes a 3D point cloud into multi-view image features for a number of different views. Then, a novel multi-view prompt fusion module is developed to fuse information from different views effectively to bridge the gap between 3D point cloud data and 2D pre-trained models. A set of 2D image prompts can then be derived to better describe the suitable prior knowledge for a large-scale pre-trained image model for few-shot 3D point cloud classification. Extensive experiments on ModelNet, ScanObjectNN, and ShapeNet datasets demonstrate that MvNet achieves new state-of-the-art performance for 3D few-shot point cloud image classification. The source code of this work is available at https://github.com/invictus717/MetaTransformer.
Haoyang Peng, Baopu Li, Bo Zhang 0069, Xin Chen 0040, Tao Chen 0003, Hongyuan Zhu 0002
IEEE Trans. Circuits Syst. Video Technol.6
2024 Deep Supervised Multi-View Learning With Graph Priors
abstract
This paper presents a novel method for supervised multi-view representation learning, which projects multiple views into a latent common space while preserving the discrimination and intrinsic structure of each view. Specifically, an apriori discriminant similarity graph is first constructed based on labels and pairwise relationships of multi-view inputs. Then, view-specific networks progressively map inputs to common representations whose affinity approximates the constructed graph. To achieve graph consistency, discrimination, and cross-view invariance, the similarity graph is enforced to meet the following constraints: 1) pairwise relationship should be consistent between the input space and common space for each view; 2) within-class similarity is larger than any between-class similarity for each view; 3) the inter-view samples from the same (or different) classes are mutually similar (or dissimilar). Consequently, the intrinsic structure and discrimination are preserved in the latent common space using an apriori approximation schema. Moreover, we present a sampling strategy to approach a sub-graph sampled from the whole similarity structure instead of approximating the graph of the whole dataset explicitly, thus benefiting lower space complexity and the capability of handling large-scale multi-view datasets. Extensive experiments show the promising performance of our method on five datasets by comparing it with 18 state-of-the-art methods.
Peng Hu 0002, Liangli Zhen, Xi Peng 0001, Hongyuan Zhu 0002, Jie Lin 0001, Xu Wang 0028, Dezhong Peng
IEEE Trans. Image Process.4
2024 Learning Student Network Under Universal Label Noise
abstract
Data-free knowledge distillation aims to learn a small student network from a large pre-trained teacher network without the aid of original training data. Recent works propose to gather alternative data from the Internet for training student network. In a more realistic scenario, the data on the Internet contains two types of label noise, namely: 1) closed-set label noise, where some examples belong to the known categories but are mislabeled; and 2) open-set label noise, where the true labels of some mislabeled examples are outside the known categories. However, the latter is largely ignored by existing works, leading to limited student network performance. Therefore, this paper proposes a novel data-free knowledge distillation paradigm by utilizing a webly-collected dataset under universal label noise, which means both closed-set and open-set label noise should be tackled. Specifically, we first split the collected noisy dataset into clean set, closed noisy set, and open noisy set based on the prediction uncertainty of various data types. For the closed-set noisy examples, their labels are refined by teacher network. Meanwhile, a noise-robust hybrid contrastive learning is performed on the clean set and refined closed noisy set to encourage student network to learn the categorical and instance knowledge inherited by teacher network. For the open-set noisy examples unexplored by previous work, we regard them as unlabeled and conduct self-supervised learning on them to enrich the supervision signal for student network. Intensive experimental results on image classification tasks demonstrate that our approach can achieve superior performance to state-of-the-art data-free knowledge distillation methods.
Jialiang Tang, Ning Jiang 0002, Hongyuan Zhu 0002, Joey Tianyi Zhou, Chen Gong 0002
IEEE Trans. Image Process.3
2023 Towards Debiasing Frame Length Bias in Text-Video Retrieval via Causal Intervention
Burak Satar, Hongyuan Zhu 0002, Hanwang Zhang, Joo-Hwee Lim
BMVC2
2023 End-to-End 3D Dense Captioning with Vote2Cap-DETR
abstract
3D dense captioning aims to generate multiple captions localized with their associated object regions. Existing methods follow a sophisticated “detect-then-describe” pipeline equipped with numerous hand-crafted components. However, these hand-crafted components would yield sub-optimal performance given cluttered object spatial and class distributions among different scenes. In this paper, we propose a simple-yet-effective transformer framework Vote2Cap-DETR based on recent popular DEtection TRansformer (DETR). Compared with prior arts, our framework has several appealing advantages: 1) Without resorting to numerous hand-crafted components, our method is based on a full transformer encoder-decoder architecture with a learnable vote query driven object decoder, and a caption decoder that produces the dense captions in a set-prediction manner. 2) In contrast to the two-stage scheme, our method can perform detection and captioning in one-stage. 3) Without bells and whistles, extensive experiments on two commonly used datasets, ScanRefer and Nr3D, demonstrate that our Vote2Cap-DETR surpasses current state-of-the-arts by 11.13% and 7.11% in [email protected], respectively. Codes will be released soon.
Sijin Chen, Hongyuan Zhu 0002, Xin Chen 0040, Yinjie Lei, Gang Yu 0002, Tao Chen 0003
CVPR2
2023 RONO: Robust Discriminative Learning with Noisy Labels for 2D-3D Cross-Modal Retrieval
abstract
Recently, with the advent of Metaverse and AI Generated Content, cross-modal retrieval becomes popular with a burst of 2D and 3D data. However, this problem is challenging given the heterogeneous structure and semantic discrepancies. Moreover, imperfect annotations are ubiquitous given the ambiguous 2D and 3D content, thus inevitably producing noisy labels to degrade the learning performance. To tackle the problem, this paper proposes a robust 2D-3D retrieval framework (RONO) to robustly learn from noisy multimodal data. Specifically, one novel Robust Discriminative Center Learning mechanism (RDCL) is proposed in RONO to adaptively distinguish clean and noisy samples for respectively providing them with positive and negative optimization directions, thus mitigating the negative impact of noisy labels. Besides, we present a Shared Space Consistency Learning mechanism (SSCL) to capture the intrinsic information inside the noisy data by minimizing the cross-modal and semantic discrepancy between common space and label space simultaneously. Comprehensive mathematical analyses are given to theoretically prove the noise tolerance of the proposed method. Furthermore, we conduct extensive experiments on four 3D-model multimodal datasets to verify the effectiveness of our method by comparing it with 15 state-of-the-art methods. Code is available at https://github.com/penghu-cs/RONO.
Yanglin Feng, Hongyuan Zhu 0002, Dezhong Peng, Xi Peng 0001, Peng Hu 0002
CVPR2
2023 Rethinking Image Super Resolution from Long-Tailed Distribution Learning Perspective
abstract
Existing studies have empirically observed that the resolution of the low-frequency region is easier to enhance than that of the high-frequency one. Although plentiful works have been devoted to alleviating this problem, little understanding is given to explain it. In this paper, we try to give a feasible answer from a machine learning perspective, i.e., the twin fitting problem caused by the long-tailed pixel distribution in natural images. With this explanation, we reformulate image super resolution (SR) as a long-tailed distribution learning problem and solve it by bridging the gaps of the problem between in low- and high-level vision tasks. As a result, we design a long-tailed distribution learning solution, that rebalances the gradients from the pixels in the low- and high-frequency region, by introducing a static and a learnable structure prior. The learned SR model achieves better balance on the fitting of the low- and high-frequency region so that the overall performance is improved. In the experiments, we evaluate the solution on four CNN- and one Transformer-based SR models w.r.t. six datasets and three tasks, and experimental results demonstrate its superiority.
Yuanbiao Gou, Peng Hu 0002, Jiancheng Lv 0001, Hongyuan Zhu 0002, Xi Peng 0001
CVPR4
2023 Semi-Supervised Few-Shot Segmentation with Noisy Support Images
abstract
Motivated by the semi-supervised learning that uses the unlabeled data and pseudo annotations to improve the image classification, this paper proposes a new semi-supervised few-shot segmentation (FSS) framework of which the training process uses not only the annotated images, but also the unlabeled images, e.g. images from other available datasets, to enhance the training of the FSS model. Furthermore, in the test phase, more support images and pseudo-annotations can also be generated by the proposed framework to enrich the support set of novel classes and therefore benefit the inference. However, unlabeled images are not a free lunch. The noisy intra-class samples and inter-class samples existed in the unlabeled images as well as the interferences of the bad quality of pseudo annotations make it difficult to utilize the correct images and pseudo annotations for a certain class. To this end, we further propose a ranking algorithm consisting of an inter-class confidence term and an intra-class confidence term to efficiently utilize the pseudo annotations of the class with high quality. Extensive experiments on COCO-20idataset demonstrate that the proposed semi-supervised FSS framework is superior to many state-of-the-art methods.
Runtong Zhang, Hongyuan Zhu 0002, Hanwang Zhang, Chen Gong 0002, Joey Tianyi Zhou, Fanman Meng
ICIP2
2023 ROAD: Robust Unsupervised Domain Adaptation with Noisy Labels
abstract
In recent years, Unsupervised Domain Adaptation (UDA) has emerged as a popular technique for transferring knowledge from a labeled source domain to an unlabeled target domain. However, almost all of the existing approaches implicitly assume that the source domain is correctly labeled, which is expensive or even impossible to satisfy in open-world applications due to ubiquitous imperfect annotations (i.e., noisy labels). In this paper, we reveal that noisy labels interfere with learning from the source domain, thus leading to noisy knowledge being transferred from the source domain to the target domain, termed Dual Noisy Information (DNI). To address this issue, we propose a robust unsupervised domain adaptation framework (ROAD), which prevents the network model from overfitting noisy labels to capture accurate discrimination knowledge for domain adaptation. Specifically, a Robust Adaptive Weighted Learning mechanism (RSWL) is proposed to adaptively assign weights to each sample based on its reliability to enforce the model to focus more on reliable samples and less on unreliable samples, thereby mining robust discrimination knowledge against noisy labels in the source domain. In order to prevent noisy knowledge from misleading domain adaptation, we present a Robust Domain-adapted Prediction Learning mechanism (RDPL) to reduce the weighted decision uncertainty of predictions in the target domain, thus ensuring the accurate knowledge of source domain transfer into the target domain, rather than uncertain knowledge from noise impact. Comprehensive experiments are conducted on three widely-used UDA benchmarks to demonstrate the effectiveness and robustness of our ROAD against noisy labels by comparing it with 13 state-of-the-art methods. Code is available at https://github.com/penghu-cs/ROAD.
Yanglin Feng, Hongyuan Zhu 0002, Dezhong Peng, Xi Peng 0001, Peng Hu 0002
ACM Multimedia2
2023 HCMA '23: 4th International Workshop on Human-Centric Multimedia Analysis
abstract
Understanding human interactions within diverse media contexts has emerged as a fundamental challenge. The explosive growth of multimedia data not only provides opportunities for human-centirc analysis but also increases the complexity of processing multimodal data. To address this pivotal challenge and explore its multifaceted dimensions, the Fourth International Workshop on Human-Centric Multimedia Analysis is concentrated on the tasks of human-centric analysis with multimedia and multimodal information. By delving into the nuances of human behavior within multimedia, this workshop aims to uncover novel insights, showcase innovative methodologies, and discuss future directions. With a spotlight on cutting-edge research and a focus on real-world applications, the workshop seeks to equip researchers and practitioners with the tools and knowledge to navigate the intricacies of human-centric multimedia analysis.
Jingkuan Song, Wu Liu 0005, Xinchen Liu, Dingwen Zhang, Chaowei Fang, Hongyuan Zhu 0002, Wenbing Huang 0001, John R. Smith, Xin Wang 0019
ACM Multimedia6
2023 A Closer Look at Few-Shot 3D Point Cloud Classification
Chuangguan Ye, Hongyuan Zhu 0002, Bo Zhang 0069, Tao Chen 0003
Int. J. Comput. Vis.2
2023 Unsupervised Contrastive Cross-Modal Hashing
abstract
In this paper, we study how to make unsupervised cross-modal hashing (CMH) benefit from contrastive learning (CL) by overcoming two challenges. To be exact, i) to address the performance degradation issue caused by binary optimization for hashing, we propose a novel momentum optimizer that performs hashing operation learnable in CL, thus making on-the-shelf deep cross-modal hashing possible. In other words, our method does not involve binary-continuous relaxation like most existing methods, thus enjoying better retrieval performance; ii) to alleviate the influence brought by false-negative pairs (FNPs), we propose a Cross-modal Ranking Learning loss (CRL) which utilizes the discrimination from all instead of only the hard negative pairs, where FNP refers to the within-class pairs that were wrongly treated as negative pairs. Thanks to such a global strategy, CRL endows our method with better performance because CRL will not overuse the FNPs while ignoring the true-negative pairs. To the best of our knowledge, the proposed method could be one of the first successful contrastive hashing methods. To demonstrate the effectiveness of the proposed method, we carry out experiments on five widely-used datasets compared with 13 state-of-the-art methods. The code is available at https://github.com/penghu-cs/UCCH.
Peng Hu 0002, Hongyuan Zhu 0002, Jie Lin 0001, Dezhong Peng, Yin-Ping Zhao, Xi Peng 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 LPCL: Localized prominence contrastive learning for self-supervised dense visual pre-training
Hongyuan Zhu 0002, Siya Mi, Yu Zhang 0004, Xin Geng 0001
Pattern Recognit.2
2023 A Closer Look at Video Sampling for Sequential Action Recognition
abstract
In recent years, sequential action recognition has attracted increasingly attention as it requires long-term sequential and compositional reasoning of human actions and object interactions. Existing methods perform reasoning either by using snippets that cover very short consecutive frames or key frames sampled from segments, which take a bias process of local and global temporal information. We also find ad-hoc training and ensembling of two separate networks using existing sampling strategies can easily outperform complex state-of-the-art methods, which reveals the complementary nature of current sampling strategies. Motivated by this observation, we propose a simple yet efficient strategy named Dense Segmental Sampling (DSS) and a novel network architecture named Temporal Dense Segment Network (TDSN) to capture the complementary information from DSS. Our TDSN achieves excellent results on benchmark action recognition datasets, which not only validate the proposed strategy but also help highlight the importance along this direction for sequential video reasoning.
Yu Zhang 0004, Zhengjie Chen, Siya Mi, Hongyuan Zhu 0002, Xin Geng 0001
IEEE Trans. Circuits Syst. Video Technol.5
2023 Dual-Stream Contrastive Learning for Channel State Information Based Human Activity Recognition
abstract
WiFi-based human activity recognition (HAR) has been extensively studied due to its far-reaching applications in health domains, including elderly monitoring, exercise supervision and rehabilitation monitoring, etc. Although existing supervised deep learning techniques have achieved remarkable performances for these tasks, they are however data-hungry and hence are notoriously difficult due to the privacy and incomprehensibility of WiFi-based HAR data. Existing contrastive learning models, mainly designed for computer vision, cannot guarantee their performance on channel state information (CSI) data. To this end, we propose a new dual-stream contrastive learning model that can process and learn the raw WiFi CSI data in a self-supervised manner. More specifically, our proposed method, coined as DualConFi, takes raw WiFI CSI data as input and incorporates channel and temporal streams to learn highly-discriminative spatiotemporal features under a mutual information constraint using unlabeled data. We exhibit the effectiveness of our model on three publicly available CSI data sets in various experiment settings, including linear evaluation, semi-supervised, and transfer learning. We show that DualConFi is able to perform favourably against challenging baselines in each setting. Moreover, by studying the effects of different transform functions on CSI data, we finally verify the effectiveness of highly-discriminative features.
Jiangtao Wang 0001, Le Zhang 0001, Hongyuan Zhu 0002, Dingchang Zheng
IEEE J. Biomed. Health Informatics4
2022 CRAFT: Cross-Attentional Flow Transformer for Robust Optical Flow
abstract
Optical flow estimation aims to find the 2D motion field by identifying corresponding pixels between two images. Despite the tremendous progress of deep learning-based optical flow methods, it remains a challenge to accurately estimate large displacements with motion blur. This is mainly because the correlation volume, the basis of pixel matching, is computed as the dot product of the convolutional features of the two images. The locality of convolutional features makes the computed correlations susceptible to various noises. On large displacements with motion blur, noisy correlations could cause severe errors in the estimated flow. To overcome this challenge, we propose a new architecture “CRoss-Attentional Flow Trans-former” (CRAFT), aiming to revitalize the correlation volume computation. In CRAFT, a Semantic Smoothing Trans-former layer transforms the features of one frame, making them more global and semantically stable. In addition, the dot-product correlations are replaced with trans-former Cross-Frame Attention. This layer filters out feature noises through the Query and Key projections, and computes more accurate correlations. On Sintel (Final) and KITTI (foreground) benchmarks, CRAFT has achieved new state-of-the-art performance. Moreover, to test the robust-ness of different models on large motions, we designed an image shifting attack that shifts input images to generate large artificial motions. Under this attack, CRAFT per-forms much more robustly than two representative meth-ods, RAFT and GMA. The code of CRAFT is is available at https://github.com/askerlee/craft.
Xiuchao Sui, Shaohua Li 0003, Xue Geng, Yan Wu 0002, Xinxing Xu, Yong Liu 0026, Rick Siow Mong Goh, Hongyuan Zhu 0002
CVPR8
2022 HCMA'22: 3rd International Workshop on Human-Centric Multimedia Analysis
abstract
The Third International Workshop on Human-Centric Multimedia Analysis concentrates on the tasks of human-centric analysis with multimedia and multimodal information. It involves multiple tasks such as face detection and recognition, human body pattern analysis, person re-identification, human action detection, etc. Today, multiple multimedia sensing technologies and large-scale computing infrastructures are emerging at a rapid velocity a wide variety of big multi-modality data for human-centric analysis, which provides rich knowledge to help tackle these challenges. Researchers have strived to push the limits of human-centric multimedia analysis in a wide variety of applications, such as intelligent surveillance, retailing, fashion design, and services. Therefore, this workshop aims to provide a platform to bridge the gap between the communities of human analysis and multimedia.
Dingwen Zhang, Chaowei Fang, Wu Liu 0005, Xinchen Liu, Jingkuan Song, Hongyuan Zhu 0002, Wenbing Huang 0001, John R. Smith
ACM Multimedia6
2022 What Makes for Effective Few-shot Point Cloud Classification?
abstract
Due to the emergence of powerful computing resources and large-scale annotated datasets, deep learning has seen wide applications in our daily life. However, most current methods require extensive data collection and retraining when dealing with novel classes never seen before. On the other hand, we humans can quickly recognize new classes by looking at a few samples, which motivates the recent popularity of few-shot learning (FSL) in machine learning communities. Most current FSL approaches work on 2D image domain, however, its implication in 3D perception is relatively under-explored. Not only needs to recognize the unseen examples as in 2D domain, 3D few-shot learning is more challenging with unordered structures, high intra-class variances and subtle inter-class differences. Moreover, different architectures and learning algorithms make it difficult to study the effectiveness of existing 2D methods when migrating to the 3D domain.In this work, for the first time, we perform systematic and extensive studies of recent 2D FSL and 3D backbone networks for benchmarking few-shot point cloud classification, and we suggest a strong baseline and learning architectures for 3D FSL. Then, we propose a novel plug-and-play component called Cross-Instance Adaptation (CIA) module, to address the high intra-class variances and subtle inter-class differences issues, which can be easily inserted into current baselines with significant performance improvement. Extensive experiments on two newly introduced benchmark datasets, ModelNet40-FS and ShapeNet70-FS, demonstrate the superiority of our proposed network for 3D FSL.
Chuangguan Ye, Hongyuan Zhu 0002, Yongbin Liao, Yanggang Zhang, Tao Chen 0003, Jiayuan Fan 0001
WACV2
2022 XAI Beyond Classification: Interpretable Neural Clustering
abstract
In this paper, we study two challenging problems in explainable AI (XAI) and data clustering. The first is how to directly design a neural network with inherent interpretability, rather than giving post-hoc explanations of a black-box model. The second is implementing discrete $k$-means with a differentiable neural network that embraces the advantages of parallel computing, online clustering, and clustering-favorable representation learning. To address these two challenges, we design a novel neural network, which is a differentiable reformulation of the vanilla $k$-means, called inTerpretable nEuraL cLustering (TELL). Our contributions are threefold. First, to the best of our knowledge, most existing XAI works focus on supervised learning paradigms. This work is one of the few XAI studies on unsupervised learning, in particular, data clustering. Second, TELL is an interpretable, or the so-called intrinsically explainable and transparent model. In contrast, most existing XAI studies resort to various means for understanding a black-box model with post-hoc explanations. Third, from the view of data clustering, TELL possesses many properties highly desired by $k$-means, including but not limited to online clustering, plug-and-play module, parallel computing, and provable convergence. Extensive experiments show that our method achieves superior performance comparing with 14 clustering approaches on three challenging data sets. The source code could be accessed at www.pengxi.me.
Xi Peng 0001, Yunfan Li 0003, Ivor W. Tsang, Hongyuan Zhu 0002, Jiancheng Lv 0001, Joey Tianyi Zhou
J. Mach. Learn. Res.4
2022 Point Cloud Instance Segmentation With Semi-Supervised Bounding-Box Mining
abstract
Point cloud instance segmentation has achieved huge progress with the emergence of deep learning. However, these methods are usually data-hungry with expensive and time-consuming dense point cloud annotations. To alleviate the annotation cost, unlabeled or weakly labeled data is still less explored in the task. In this paper, we introduce the first semi-supervised point cloud instance segmentation framework (SPIB) using both labeled and unlabelled bounding boxes as supervision. To be specific, our SPIB architecture involves a two-stage learning procedure. For stage one, a bounding box proposal generation network is trained under a semi-supervised setting with perturbation consistency regularization (SPCR). The regularization works by enforcing an invariance of the bounding box predictions over different perturbations applied to the input point clouds, to provide self-supervision for network learning. For stage two, the bounding box proposals with SPCR are grouped into some subsets, and the instance masks are mined inside each subset with a novel semantic propagation module and a property consistency graph module. Moreover, we introduce a novel occupancy ratio guided refinement module to refine the instance masks. Extensive experiments on the challenging ScanNet v2 dataset demonstrate our method can achieve competitive performance compared with the recent fully-supervised methods.
Yongbin Liao, Hongyuan Zhu 0002, Yanggang Zhang, Chuangguan Ye, Tao Chen 0003, Jiayuan Fan 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2022 Locality-Aware Crowd Counting
abstract
Imbalanced data distribution in crowd counting datasets leads to severe under-estimation and over-estimation problems, which has been less investigated in existing works. In this paper, we tackle this challenging problem by proposing a simple but effective locality-based learning paradigm to produce generalizable features by alleviating sample bias. Our proposed method is locality-aware in two aspects. First, we introduce a locality-aware data partition (LADP) approach to group the training data into different bins via locality-sensitive hashing. As a result, a more balanced data batch is then constructed by LADP. To further reduce the training bias and enhance the collaboration with LADP, a new data augmentation method called locality-aware data augmentation (LADA) is proposed where the image patches are adaptively augmented based on the loss. The proposed method is independent of the backbone network architectures, and thus could be smoothly integrated with most existing deep crowd counting approaches in an end-to-end paradigm to boost their performance. We also demonstrate the versatility of the proposed method by applying it for adversarial defense. Extensive experiments verify the superiority of the proposed method over the state of the arts.
Joey Tianyi Zhou, Le Zhang 0001, Jiawei Du 0002, Xi Peng 0001, Zhiwen Fang, Hongyuan Zhu 0002
IEEE Trans. Pattern Anal. Mach. Intell.7
2022 Deep Semisupervised Multiview Learning With Increasing Views
abstract
In this article, we study two challenging problems in semisupervised cross-view learning. On the one hand, most existing methods assume that the samples in all views have a pairwise relationship, that is, it is necessary to capture or establish the correspondence of different views at the sample level. Such an assumption is easily isolated even in the semisupervised setting wherein only a few samples have labels that could be used to establish the correspondence. On the other hand, almost all existing multiview methods, including semisupervised ones, usually train a model using a fixed dataset, which cannot handle the data of increasing views. In practice, the view number will increase when new sensors are deployed. To address the above two challenges, we propose a novel method that employs multiple independent semisupervised view-specific networks (ISVNs) to learn representation for multiple views in a view-decoupling fashion. The advantages of our method are two-fold. Thanks to our specifically designed autoencoder and pseudolabel learning paradigm, our method shows an effective way to utilize both the labeled and unlabeled data while relaxing the data assumption of the pairwise relationship, that is, correspondence. Furthermore, with our view decoupling strategy, the proposed ISVNs could be separately trained, thus efficiently handling the data of increasing views without retraining the entire model. To the best of our knowledge, our ISVN could be one of the first attempts to make handling increasing views in the semisupervised setting possible, as well as an effective solution to the noncorresponding problem. To verify the effectiveness and efficiency of our method, we conduct comprehensive experiments by comparing 13 state-of-the-art approaches on four multiview datasets in terms of retrieval and classification.
Peng Hu 0002, Xi Peng 0001, Hongyuan Zhu 0002, Liangli Zhen, Jie Lin 0001, Huaibai Yan, Dezhong Peng
IEEE Trans. Cybern.3
2021 OPQ: Compressing Deep Neural Networks with One-shot Pruning-Quantization
abstract
As Deep Neural Networks (DNNs) usually are overparameterized and have millions of weight parameters, it is challenging to deploy these large DNN models on resource-constrained hardware platforms, e.g., smartphones. Numerous network compression methods such as pruning and quantization are proposed to reduce the model size significantly, of which the key is to find suitable compression allocation (e.g., pruning sparsity and quantization codebook) of each layer. Existing solutions obtain the compression allocation in an iterative/manual fashion while finetuning the compressed model, thus suffering from the efficiency issue. Different from the prior art, we propose a novel One-shot Pruning-Quantization (OPQ) in this paper, which analytically solves the compression allocation with pre-trained weight parameters only. During finetuning, the compression module is fixed and only weight parameters are updated. To our knowledge, OPQ is the first work that reveals pre-trained model is sufficient for solving pruning and quantization simultaneously, without any complex iterative/manual optimization at the finetuning stage. Furthermore, we propose a unified channel-wise quantization method that enforces all channels of each layer to share a common codebook, which leads to low bit-rate allocation without introducing extra overhead brought by traditional channel-wise quantization. Comprehensive experiments on ImageNet with AlexNet/MobileNet-V1/ResNet-50 show that our method improves accuracy and training efficiency while obtains significantly higher compression rates compared to the state-of-the-art.
Peng Hu 0002, Xi Peng 0001, Hongyuan Zhu 0002, Mohamed M. Sabry, Jie Lin 0001
AAAI3
2021 Learning Cross-Modal Retrieval With Noisy Labels
abstract
Recently, cross-modal retrieval is emerging with the help of deep multimodal learning. However, even for unimodal data, collecting large-scale well-annotated data is expensive and time-consuming, and not to mention the additional challenges from multiple modalities. Although crowd-sourcing annotation, e.g., Amazon’s Mechanical Turk, can be utilized to mitigate the labeling cost, but leading to the unavoidable noise in labels for the non-expert annotating. To tackle the challenge, this paper presents a general Multi-modal Robust Learning framework (MRL) for learning with multimodal noisy labels to mitigate noisy samples and correlate distinct modalities simultaneously. To be specific, we propose a Robust Clustering loss (RC) to make the deep networks focus on clean samples instead of noisy ones. Besides, a simple yet effective multimodal loss function, called Multimodal Contrastive loss (MC), is proposed to maxi-mize the mutual information between different modalities, thus alleviating the interference of noisy samples and cross-modal discrepancy. Extensive experiments are conducted on four widely-used multimodal datasets to demonstrate the effectiveness of the proposed approach by comparing to 14 state-of-the-art methods.
Peng Hu 0002, Xi Peng 0001, Hongyuan Zhu 0002, Liangli Zhen, Jie Lin 0001
CVPR3
2021 A Diagnostic Study Of Visual Question Answering With Analogical Reasoning
abstract
The deep learning community has made rapid progress in low-level visual perception tasks such as object localization, detection and segmentation. However, for tasks such as Visual Question Answering (VQA) and visual language grounding that require high-level reasoning abilities, huge gaps still exist between artificial systems and human intelligence. In this work, we perform a diagnostic study on recent popular VQA in terms of analogical reasoning. We term it as Analogical VQA, where a system needs to reason on a group of images to find analogical relations among them in order to correctly answer a natural language question. To study the task in depth, we propose an initial diagnostic synthetic dataset CLEVR-Analogy, which tests a range of analogical reasoning abilities (e.g. reasoning on object attributes, spatial relationships, existence, and arithmetic analogies). We benchmark various recent state-of-the-art methods on our dataset and compare the results against human performance, and discover that existing systems fall shorts when facing analogical reasoning involving spatial relationships. The dataset and code will be publicly available to facilitate future research.
Hongyuan Zhu 0002, Ying Sun 0001, Dongkyu Choi, Cheston Tan, Joo-Hwee Lim
ICIP2
2021 Spcr: semi-supervised point cloud instance segmentation with perturbation consistency regularization
abstract
Point cloud instance segmentation is steadily improving with the development of deep learning. However, current progress is hindered by the expensive cost of collecting dense point cloud labels. To this end, we propose the first semi-supervised point cloud instance segmentation architecture, which is called semi-supervised point cloud instance segmentation with perturbation consistency regularization (SPCR). It is capable to alleviate the data-hungry bottleneck of existing strongly supervised methods. Specifically, SPCR enforces an invariance of the predictions over different perturbations applied to the input point clouds. We firstly introduce various perturbation schemes on inputs to force the network to be robust and easily generalized to the unseen and unlabeled data. Further, perturbation consistency regularization is then conducted on predicted instance masks from various transformed inputs to provide self-supervision for network learning. Extensive experiments on the challenging ScanNet v2 dataset demonstrate our method can achieve competitive performance compared with the state-of-the-art of fully supervised methods.
Yongbin Liao, Hongyuan Zhu 0002, Tao Chen 0003, Jiayuan Fan 0001
ICIP2
2021 Semantic Role Aware Correlation Transformer For Text To Video Retrieval
abstract
With the emergence of social media, voluminous video clips are uploaded every day, and retrieving the most relevant visual content with a language query becomes critical. Most approaches aim to learn a joint embedding space for plain textual and visual contents without adequately exploiting their intra-modality structures and inter-modality correlations. This paper proposes a novel transformer that explicitly disentangles the text and video into semantic roles of objects, spatial contexts and temporal contexts with an attention scheme to learn the intra- and inter-role correlations among the three roles to discover discriminative features for matching at different levels. The preliminary results on popular YouCook2 indicate that our approach surpasses a current state-of-the-art method, with a high margin in all metrics. It also overpasses two SOTA methods in terms of two metrics.
Burak Satar, Hongyuan Zhu 0002, Xavier Bresson, Joo-Hwee Lim
ICIP2
2021 A comprehensive survey of procedural video datasets
Hui Li Tan, Hongyuan Zhu 0002, Joo-Hwee Lim, Cheston Tan
Comput. Vis. Image Underst.2
2021 Cross-modal discriminant adversarial network
Peng Hu 0002, Xi Peng 0001, Hongyuan Zhu 0002, Jie Lin 0001, Liangli Zhen, Wei Wang 0283, Dezhong Peng
Pattern Recognit.3
2021 Joint Versus Independent Multiview Hashing for Cross-View Retrieval
abstract
Thanks to the low storage cost and high query speed, cross-view hashing (CVH) has been successfully used for similarity search in multimedia retrieval. However, most existing CVH methods use all views to learn a common Hamming space, thus making it difficult to handle the data with increasing views or a large number of views. To overcome these difficulties, we propose a decoupled CVH network (DCHN) approach which consists of a semantic hashing autoencoder module (SHAM) and multiple multiview hashing networks (MHNs). To be specific, SHAM adopts a hashing encoder and decoder to learn a discriminative Hamming space using either a few labels or the number of classes, that is, the so-called flexible inputs. After that, MHN independently projects all samples into the discriminative Hamming space that is treated as an alternative ground truth. In brief, the Hamming space is learned from the semantic space induced from the flexible inputs, which is further used to guide view-specific hashing in an independent fashion. Thanks to such an independent/decoupled paradigm, our method could enjoy high computational efficiency and the capacity of handling the increasing number of views by only using a few labels or the number of classes. For a newly coming view, we only need to add a view-specific network into our model and avoid retraining the entire model using the new and previous views. Extensive experiments are carried out on five widely used multiview databases compared with 15 state-of-the-art approaches. The results show that the proposed independent hashing paradigm is superior to the common joint ones while enjoying high efficiency and the capacity of handling newly coming views.
Peng Hu 0002, Xi Peng 0001, Hongyuan Zhu 0002, Jie Lin 0001, Liangli Zhen, Dezhong Peng
IEEE Trans. Cybern.3
2021 Single-Image Dehazing via Compositional Adversarial Network
abstract
Single-image dehazing has been an important topic given the commonly occurred image degradation caused by adverse atmosphere aerosols. The key to haze removal relies on an accurate estimation of global air-light and the transmission map. Most existing methods estimate these two parameters using separate pipelines which reduces the efficiency and accumulates errors, thus leading to a suboptimal approximation, hurting the model interpretability, and degrading the performance. To address these issues, this article introduces a novel generative adversarial network (GAN) for single-image dehazing. The network consists of a novel compositional generator and a novel deeply supervised discriminator. The compositional generator is a densely connected network, which combines fine-scale and coarse-scale information. Benefiting from the new generator, our method can directly learn the physical parameters from data and recover clean images from hazy ones in an end-to-end manner. The proposed discriminator is deeply supervised, which enforces that the output of the generator to look similar to the clean images from low-level details to high-level structures. To the best of our knowledge, this is the first end-to-end generative adversarial model for image dehazing, which simultaneously outputs clean images, transmission maps, and air-lights. Extensive experiments show that our method remarkably outperforms the state-of-the-art methods. Furthermore, to facilitate future research, we create the HazeCOCO dataset which is currently the largest dataset for single-image dehazing.
Hongyuan Zhu 0002, Xi Peng 0001, Joey Tianyi Zhou, Zhao Kang 0001, Shijian Lu, Zhiwen Fang, Liyuan Li, Joo-Hwee Lim
IEEE Trans. Cybern.1
2021 Deep Spectral Representation Learning From Multi-View Data
abstract
Multi-view representation learning (MvRL) aims to learn a consensus representation from diverse sources or domains to facilitate downstream tasks such as clustering, retrieval, and classification. Due to the limited representative capacity of the adopted shallow models, most existing MvRL methods may yield unsatisfactory results, especially when the labels of data are unavailable. To enjoy the representative capacity of deep learning, this paper proposes a novel multi-view unsupervised representation learning method, termed as Multi-view Laplacian Network (MvLNet), which could be the first deep version of the multi-view spectral representation learning method. Note that, such an attempt is nontrivial because simply combining Laplacian embedding (i.e., spectral representation) with neural networks will lead to trivial solutions. To solve this problem, MvLNet enforces an orthogonal constraint and reformulates it as a layer with the help of Cholesky decomposition. The orthogonal layer is stacked on the embedding network so that a common space could be learned for consensus representation. Compared with numerous recent-proposed approaches, extensive experiments on seven challenging datasets demonstrate the effectiveness of our method in three multi-view tasks including clustering, recognition, and retrieval. The source code could be found at www.pengxi.me.
Zhenyu Huang 0005, Joey Tianyi Zhou, Hongyuan Zhu 0002, Changqing Zhang 0002, Jiancheng Lv 0001, Xi Peng 0001
IEEE Trans. Image Process.3
2020 Semi-Supervised Multi-Modal Learning with Balanced Spectral Decomposition
abstract
Cross-modal retrieval aims to retrieve the relevant samples across different modalities, of which the key problem is how to model the correlations among different modalities while narrowing the large heterogeneous gap. In this paper, we propose a Semi-supervised Multimodal Learning Network method (SMLN) which correlates different modalities by capturing the intrinsic structure and discriminative correlation of the multimedia data. To be specific, the labeled and unlabeled data are used to construct a similarity matrix which integrates the cross-modal correlation, discrimination, and intra-modal graph information existing in the multimedia data. What is more important is that we propose a novel optimization approach to optimize our loss within a neural network which involves a spectral decomposition problem derived from a ratio trace criterion. Our optimization enjoys two advantages given below. On the one hand, the proposed approach is not limited to our loss, which could be applied to any case that is a neural network with the ratio trace criterion. On the other hand, the proposed optimization is different from existing ones which alternatively maximize the minor eigenvalues, thus overemphasizing the minor eigenvalues and ignore the dominant ones. In contrast, our method will exactly balance all eigenvalues, thus being more competitive to existing methods. Thanks to our loss and optimization strategy, our method could well preserve the discriminative and instinct information into the common space and embrace the scalability in handling large-scale multimedia data. To verify the effectiveness of the proposed method, extensive experiments are carried out on three widely-used multimodal datasets comparing with 13 state-of-the-art approaches.
Peng Hu 0002, Hongyuan Zhu 0002, Xi Peng 0001, Jie Lin 0001
AAAI2
2020 6D Pose Estimation with Correlation Fusion
abstract
6D object pose estimation is widely applied in robotic tasks such as grasping and manipulation. Prior methods using RGB-only images are vulnerable to heavy occlusion and poor illumination, so it is important to complement them with depth information. However, existing methods using RGB-D data cannot adequately exploit consistent and complementary information between RGB and depth modalities. In this paper, we present a novel method to effectively consider the correlation within and across both modalities with attention mechanism to learn discriminative and compact multi-modal features. Then, effective fusion strategies for intra- and inter-correlation modules are explored to ensure efficient information flow between RGB and depth. To our best knowledge, this is the first work to explore effective intra- and inter-modality fusion in 6D pose estimation. The experimental results show that our method can achieve the state-of-the-art performance on LineMOD and YCB-Video dataset. We also demonstrate that the proposed method can benefit a real-world robot grasping task by providing accurate object pose estimation.
Hongyuan Zhu 0002, Ying Sun 0001, Cihan Acar, Yan Wu 0002, Liyuan Li, Cheston Tan, Joo-Hwee Lim
ICPR2
2020 Partition level multiview subspace clustering
Zhao Kang 0001, Xinjia Zhao, Chong Peng 0001, Hongyuan Zhu 0002, Joey Tianyi Zhou, Xi Peng 0001, Wenyu Chen 0001, Zenglin Xu
Neural Networks4
2020 A novel hybrid approach for crack detection
Fen Fang, Liyuan Li, Hongyuan Zhu 0002, Joo-Hwee Lim
Pattern Recognit.4
2020 Improving Night-Time Pedestrian Retrieval With Distribution Alignment and Contextual Distance
abstract
Night-time pedestrian retrieval is a cross-modality retrieval task of retrieving person images between day-time visible images and night-time thermal images. It is a very challenging problem due to modality difference, camera variations, and person variations, but it plays an important role in night-time video surveillance. The existing cross-modality retrieval usually focuses on learning modality sharable feature representations to bridge the modality gap. In this article, we propose to utilize auxiliary information to improve the retrieval performance, which consistently improves the performance with different baseline loss functions. Our auxiliary information contains two major parts: cross-modality feature distribution and contextual information. The former aligns the cross-modality feature distributions between two modalities to improve the performance, and the latter optimizes the cross-modality distance measurement with the contextual information. We also demonstrate that abundant annotated visible pedestrian images, which are easily accessible, help to improve the cross-modality pedestrian retrieval as well. The proposed method is featured in two aspects: the auxiliary information does not need additional human intervention or annotation; it learns discriminative feature representations in an end-to-end deep learning manner. Extensive experiments on two cross-modality pedestrian retrieval datasets demonstrate the superiority of the proposed method, achieving much better performance than the state-of-the-arts.
Mang Ye, Xiangyuan Lan, Hongyuan Zhu 0002
IEEE Trans. Ind. Informatics4
2020 Holistic Multi-Modal Memory Network for Movie Question Answering
abstract
Answering questions using multi-modal context is a challenging problem as it requires a deep integration of diverse data sources. Existing approaches only consider a subset of all possible interactions among data sources during one attention hop. In this paper, we present a Holistic Multi-modal Memory Network (HMMN) framework that fully considers interactions between different input sources (multi-modal context, question) at each hop. In addition, to hone in on relevant information, our framework takes answer choices into consideration during the context retrieval stage. Our HMMN framework effectively integrates information from the multi-modal context, question, and answer choices, enabling more informative context to be retrieved for question answering. Experimental results on the MovieQA and TVQA datasets validate the effectiveness of our HMMN framework. Extensive ablation studies show the importance of holistic reasoning and reveal the contributions of different attention strategies to model performance.
Anran Wang 0001, Anh Tuan Luu, Chuan-Sheng Foo, Hongyuan Zhu 0002, Yi Tay, Vijay Chandrasekhar 0001
IEEE Trans. Image Process.4
2020 Combining Faster R-CNN and Model-Driven Clustering for Elongated Object Detection
abstract
While analyzing the performance of state-of-the-art R-CNN based generic object detectors, we find that the detection performance for objects with low object-region-percentages (ORPs) of the bounding boxes are much lower than the overall average. Elongated objects are examples. To address the problem of low ORPs for elongated object detection, we propose a hybrid approach which employs a Faster R-CNN to achieve robust detections of object parts, and a novel model-driven clustering algorithm to group the related partial detections and suppress false detections. First, we train a Faster R-CNN with partial region proposals of suitable and stable ORPs. Next, we introduce a deep CNN (DCNN) for orientation classification on the partial detections. Then, on the outputs of the Faster R-CNN and DCNN, the algorithm of adaptive model-driven clustering first initializes a model of an elongated object with a data-driven process on local partial detections, and refines the model iteratively by model-driven clustering and data-driven model updating. By exploiting Faster R-CNN to produce robust partial detections and model-driven clustering to form a global representation, our method is able to generate a tight oriented bounding box for elongated object detection. We evaluate the effectiveness of our approach on two typical elongated objects in the COCO dataset, and other typical elongated objects, including rigid objects (pens, screwdrivers and wrenches) and non-rigid objects (cracks). Experimental results show that, compared with the state-of-the-art approaches, our method achieves a large margin of improvements for both detection and localization of elongated objects in images.
Fen Fang, Liyuan Li, Hongyuan Zhu 0002, Joo-Hwee Lim
IEEE Trans. Image Process.3
2020 Zero-Shot Image Dehazing
abstract
In this paper, we study two less-touched challenging problems in single image dehazing neural networks, namely, how to remove haze from a given image in an unsupervised and zeroshot manner. To the ends, we propose a novel method based on the idea of layer disentanglement by viewing a hazy image as the entanglement of several "simpler" layers, i.e., a hazy-free image layer, transmission map layer, and atmospheric light layer. The major advantages of the proposed ZID are two-fold. First, it is an unsupervised method that does not use any clean images including hazy-clean pairs as the ground-truth. Second, ZID is a "zero-shot" method, which just uses the observed single hazy image to perform learning and inference. In other words, it does not follow the conventional paradigm of training deep model on a large scale dataset. These two advantages enable our method to avoid the labor-intensive data collection and the domain shift issue of using the synthetic hazy images to address the real-world images. Extensive comparisons show the promising performance of our method compared with 15 approaches in the qualitative and quantitive evaluations. The source code could be found at www.pengxi.me.
Boyun Li, Yuanbiao Gou, Zitao Liu 0001, Hongyuan Zhu 0002, Joey Tianyi Zhou, Xi Peng 0001
IEEE Trans. Image Process.4
2020 Deep Clustering With Sample-Assignment Invariance Prior
abstract
Most popular clustering methods map raw image data into a projection space in which the clustering assignment is obtained with the vanilla k-means approach. In this article, we discovered a novel prior, namely, there exists a common invariance when assigning an image sample to clusters using different metrics. In short, different distance metrics will lead to similar soft clustering assignments on the manifold. Based on such a novel prior, we propose a novel clustering method by minimizing the discrepancy between pairwise sample assignments for each data point. To the best of our knowledge, this could be the first work to reveal the sample-assignment invariance prior based on the idea of treating labels as ideal representations. Furthermore, the proposed method is one of the first end-to-end clustering approaches, which jointly learns clustering assignment and representation. Extensive experimental results show that the proposed method is remarkably superior to 16 state-of-the-art clustering methods on five image data sets in terms of four evaluation metrics.
Xi Peng 0001, Hongyuan Zhu 0002, Jiashi Feng, Chunhua Shen, Haixian Zhang, Joey Tianyi Zhou
IEEE Trans. Neural Networks Learn. Syst.2
2019 Singe Image Rain Removal with Unpaired Information: A Differentiable Programming Perspective
abstract
Single image rain-streak removal is an extremely challenging problem due to the presence of non-uniform rain densities in images. Previous works solve this problem using various hand-designed priors or by explicitly mapping synthetic rain to paired clean image in a supervised way. In practice, however, the pre-defined priors are easily violated and the paired training data are hard to collect. To overcome these limitations, in this work, we propose RainRemoval-GAN (RRGAN), the first end-to-end adversarial model that generates realistic rain-free images using only unpaired supervision. Our approach alleviates the paired training constraints by introducing a physical-model which explicitly learns a recovered images and corresponding rain-streaks from the differentiable programming perspective. The proposed network consists of a novel multiscale attention memory generator and a novel multiscale deeply supervised discriminator. The multiscale attention memory generator uses a memory with attention mechanism to capture the latent rain streaks context at different stages to recover the clean images. The deeply supervised multiscale discriminator imposes constraints at the recovered output in terms of local details and global appearance to the clean image set. Together with the learned rainstreaks, a reconstruction constraint is employed to ensure the appearance consistent with the input image. Experimental results on public benchmark demonstrates our promising performance compared with nine state-of-the-art methods in terms of PSNR, SSIM, visual qualities and running time.
Hongyuan Zhu 0002, Xi Peng 0001, Joey Tianyi Zhou, Songfan Yang, Vijay Chanderasekh, Liyuan Li, Joo-Hwee Lim
AAAI1
2019 Dual Adversarial Neural Transfer for Low-Resource Named Entity Recognition
abstract
We propose a new neural transfer method termed Dual Adversarial Transfer Network (DATNet) for addressing low-resource Named Entity Recognition (NER).Specifically, two variants of DATNet, i.e., DATNet-F and DATNet-P, are investigated to explore effective feature fusion between high and low resource.To address the noisy and imbalanced training data, we propose a novel Generalized Resource-Adversarial Discriminator (GRAD).Additionally, adversarial training is adopted to boost model generalization.In experiments, we examine the effects of different components in DATNet across domains and languages, and show that significant improvement can be obtained especially for lowresource data, without augmenting any additional hand-crafted features and pre-trained language model.
Joey Tianyi Zhou, Hao Zhang 0048, Di Jin 0005, Hongyuan Zhu 0002, Rick Siow Mong Goh, Kenneth Kwok
ACL (1)4
2019 Spatial Fusion GAN for Image Synthesis
abstract
Recent advances in generative adversarial networks (GANs) have shown great potentials in realistic image synthesis whereas most existing works address synthesis realism in either appearance space or geometry space but few in both. This paper presents an innovative Spatial Fusion GAN (SF-GAN) that combines a geometry synthesizer and an appearance synthesizer to achieve synthesis realism in both geometry and appearance spaces. The geometry synthesizer learns contextual geometries of background images and transforms and places foreground objects into the background images unanimously. The appearance synthesizer adjust the color, brightness and styles of the foreground objects and embeds them into background images harmoniously, where a guided filter is incorporated for detail preserving. The two synthesizers are inter-connected as mutual references which can be trained end-to-end with little supervision. The SF-GAN has been evaluated in two tasks: (1) realistic scene text image synthesis for training better recognition models; (2) glass and hat wearing for realistic matching glasses and hats with real portraits. Qualitative and quantitative comparisons with the state-of-the-art demonstrate the superiority of the proposed SF-GAN.
Fangneng Zhan, Hongyuan Zhu 0002, Shijian Lu
CVPR2
2019 COMIC: Multi-view Clustering Without Parameter Selection
abstract
In this paper, we study two challenges in clustering analysis, namely, how to cluster multi-view data and how to perform clustering without parameter selection on cluster size. To this end, we propose a novel objective function to project raw data into one space in which the projection embraces the geometric consistency (GC) and the cluster assignment consistency (CAC). To be specific, the GC aims to learn a connection graph from a projection space wherein the data points are connected if and only if they belong to the same cluster. The CAC aims to minimize the discrepancy of pairwise connection graphs induced from different views based on the view-consensus assumption, i.e., different views could produce the same cluster assignment structure as they are different portraits of the same object. Thanks to the view-consensus derived from the connection graph, our method could achieve promising performance in learning view-specific representation and eliminating the heterogeneous gaps across different views. Furthermore, with the proposed objective, it could learn almost all parameters including the cluster number from data without labor-intensive parameter selection. Extensive experimental results show the promising performance achieved by our method on five datasets comparing with nine state-of-the-art multi-view clustering approaches.
Xi Peng 0001, Zhenyu Huang 0005, Jiancheng Lv 0001, Hongyuan Zhu 0002, Joey Tianyi Zhou
ICML4
2019 Multi-view Spectral Clustering Network
abstract
Multi-view clustering aims to cluster data from diverse sources or domains, which has drawn considerable attention in recent years. In this paper, we propose a novel multi-view clustering method named multi-view spectral clustering network (MvSCN) which could be the first deep version of multi-view spectral clustering to the best of our knowledge. To deeply cluster multi-view data, MvSCN incorporates the local invariance within every single view and the consistency across different views into a novel objective function, where the local invariance is defined by a deep metric learning network rather than the Euclidean distance adopted by traditional approaches. In addition, we enforce and reformulate an orthogonal constraint as a novel layer stacked on an embedding network for two advantages, i.e. jointly optimizing the neural network and performing matrix decomposition and avoiding trivial solutions. Extensive experiments on four challenging datasets demonstrate the effectiveness of our method compared with 10 state-of-the-art approaches in terms of three evaluation metrics.
Zhenyu Huang 0005, Joey Tianyi Zhou, Xi Peng 0001, Changqing Zhang 0002, Hongyuan Zhu 0002, Jiancheng Lv 0001
IJCAI5
2019 Clustering with similarity preserving
Zhao Kang 0001, Boyu Wang 0004, Hongyuan Zhu 0002, Zenglin Xu
Neurocomputing4
2019 AnomalyNet: An Anomaly Detection Network for Video Surveillance
abstract
Sparse coding-based anomaly detection has shown promising performance, of which the keys are feature learning, sparse representation, and dictionary learning. In this paper, we propose a new neural network for anomaly detection (termed AnomalyNet) by deeply achieving feature learning, sparse representation, and dictionary learning in three joint neural processing blocks. Specifically, to learn better features, we design a motion fusion block accompanied by a feature transfer block to enjoy the advantages of eliminating noisy background, capturing motion, and alleviating data deficiency. Furthermore, to address some disadvantages (e.g., nonadaptive updating) of the existing sparse coding optimizers and embrace the merits of neural network (e.g., parallel computing), we design a novel recurrent neural network to learn sparse representation and dictionary by proposing an adaptive iterative hard-thresholding algorithm (adaptive ISTA) and reformulating the adaptive ISTA as a new long short-term memory (LSTM). To the best of our knowledge, this could be one of the first works to bridge the$\ell _{1}$ -solver and LSTM and may provide novel insight into understanding LSTM and model-based optimization (or named differentiable programming), as well as sparse coding-based anomaly detection. Extensive experiments show the state-of-the-art performance of our method in the abnormal events detection task.
Joey Tianyi Zhou, Jiawei Du 0002, Hongyuan Zhu 0002, Xi Peng 0001, Yong Liu 0026, Rick Siow Mong Goh
IEEE Trans. Inf. Forensics Secur.3
2018 DehazeGAN: When Image Dehazing Meets Differential Programming
abstract
Single image dehazing has been a classic topic in computer vision for years. Motivated by the atmospheric scattering model, the key to satisfactory single image dehazing relies on an estimation of two physical parameters, i.e., the global atmospheric light and the transmission coefficient. Most existing methods employ a two-step pipeline to estimate these two parameters with heuristics which accumulate errors and compromise dehazing quality. Inspired by differentiable programming, we re-formulate the atmospheric scattering model into a novel generative adversarial network (DehazeGAN). Such a reformulation and adversarial learning allow the two parameters to be learned simultaneously and automatically from data by optimizing the final dehazing performance so that clean images with faithful color and structures are directly produced. Moreover, our reformulation also greatly improves the GAN’s interpretability and quality for single image dehazing. To the best of our knowledge, our method is one of the first works to explore the connection among generative adversarial models, image dehazing, and differentiable programming, which advance the theories and application of these areas. Extensive experiments on synthetic and realistic data show that our method outperforms state-of-the-art methods in terms of PSNR, SSIM, and subjective visual quality.
Hongyuan Zhu 0002, Xi Peng 0001, Vijay Chandrasekhar 0001, Liyuan Li, Joo-Hwee Lim
IJCAI1
2017 TORNADO: A Spatio-Temporal Convolutional Regression Network for Video Action Proposal
abstract
Given a video clip, action proposal aims to quickly generate a number of spatio-temporal tubes that enclose candidate human activities. Recently, the regression-based networks and long-term recurrent convolutional network (L-RCN) have demonstrated superior performance in object detection and action recognition. However, the regression-based detectors perform inference without considering the temporal context among neighboring frames, and the LRC-N using global visual percepts lacks the capability to capture local temporal dynamics. In this paper, we present a novel framework called TORNADO for human action proposal detection in un-trimmed video clips. Specifically, we propose a spatio-temporal convolutional network that combines the advantages of regression-based detector and L-RCN by empowering Convolutional LSTM with regression capability. Our approach consists of a temporal convolutional regression network (T-CRN) and a spatial regression network (S-CRN) which are trained end-to-end on both RGB and optical flow streams. They fuse appearance, motion and temporal contexts to regress the bounding boxes of candidate human actions simultaneously in 28 FPS. The action proposals are constructed by solving dynamic programming with peak trimming of the generated action boxes. Extensive experiments on the challenging UCF-101 and UCF-Sports datasets show that our method achieves superior performance as compared with the state-of-the-arts.
Hongyuan Zhu 0002, Romain Vial, Shijian Lu
ICCV1
2017 Search video action proposal with recurrent and static YOLO
abstract
In this paper, we propose a new approach for searching action proposals in unconstrained videos. Our method first produces snippet action proposals by combining state-of-the-art YOLO detector (Static YOLO) and our regression based RNN detector (Recurrent YOLO). Then, these short action proposals are integrated to form final action proposals by solving two-pass dynamic programming which maximizes actioness score and temporal smoothness concurrently. Our experimental comparison with other state-of-the-arts on challenging UCF101 dataset shows that our method advances state-of-the-art proposal generation performance while maintaining low computational cost.
Romain Vial, Hongyuan Zhu 0002, Yonghong Tian 0001, Shijian Lu
ICIP2
2016 Discriminative Multi-modal Feature Fusion for RGBD Indoor Scene Recognition
abstract
RGBD scene recognition has attracted increasingly attention due to the rapid development of depth sensors and their wide application scenarios. While many research has been conducted, most work used hand-crafted features which are difficult to capture high-level semantic structures. Recently, the feature extracted from deep convolutional neural network has produced state-of-the-art results for various computer vision tasks, which inspire researchers to explore incorporating CNN learned features for RGBD scene understanding. On the other hand, most existing work combines rgb and depth features without adequately exploiting the consistency and complementary information between them. Inspired by some recent work on RGBD object recognition using multi-modal feature fusion, we introduce a novel discriminative multi-modal fusion framework for rgbd scene recognition for the first time which simultaneously considers the inter-and intra-modality correlation for all samples and meanwhile regularizing the learned features to be discriminative and compact. The results from the multimodal layer can be back-propagated to the lower CNN layers, hence the parameters of the CNN layers and multimodal layers are updated iteratively until convergence. Experiments on the recently proposed large scale SUN RGB-D datasets show that our method achieved the state-of-the-art without any image segmentation.
Hongyuan Zhu 0002, Jean-Baptiste Weibel, Shijian Lu
CVPR1
2016 Beyond pixels: A comprehensive survey from bottom-up to semantic image segmentation and cosegmentation
Hongyuan Zhu 0002, Fanman Meng, Jianfei Cai 0001, Shijian Lu
J. Vis. Commun. Image Represent.1
2016 Multiple Human Identification and Cosegmentation: A Human-Oriented CRF Approach With Poselets
abstract
Localizing, identifying, and extracting humans with consistent appearance jointly from a personal photo stream is an important problem and has wide applications. The strong variations in foreground and background and irregularly occurring foreground humans make this realistic problem challenging. Inspired by advancements in object detection, scene understanding, and image cosegmentation, we explore explicit constraints to label and segment human objects rather than other nonhuman objects and “stuff.” We refer to such a problem as multiple human identification and cosegmentation (MHIC). To identify specific human subjects, we propose an efficient human instance detector by combining an extended color line model with a poselet-based human detector. Moreover, to capture high-level human shape information, a novel soft shape cue is proposed. It is initialized by the human detector, then further enhanced through a generalized geodesic distance transform, and finally refined with a joint bilateral filter. We also propose to capture the rich feature context around each pixel by using an adaptive cross-region data structure, which gives a higher discriminative power than a single pixel-based estimation. The high-level object cues from the detector and the shape are then integrated with the low-level pixel cues and midlevel contour cues into a principled conditional random field (CRF) framework, which can be efficiently solved by using fast graph cut algorithms. We evaluate our method over a newly created NTU-MHIC human dataset, which contains 351 images with manually annotated groundtruth segmentation. Both visual and quantitative results demonstrate that our method achieves state-of-the-art performance for the MHIC task.
Hongyuan Zhu 0002, Jiangbo Lu, Jianfei Cai 0001, Jianmin Zheng, Shijian Lu, Nadia Magnenat-Thalmann
IEEE Trans. Multim.1
2015 Diagnosing state-of-the-art object proposal methods
abstract
Recent top performing methods in PASCAL VOC [6] and ImageNet [13] make use of object proposal to replace exhaustive window search. Object proposal’s effectiveness is rooted in the assumption that there are general cues to differentiate objects from the background. Since the very first work by Alexe et al. [1], many object proposal methods have been proposed [2, 3, 4, 5, 7, 8, 10, 11, 12, 14, 15] and tested on various large scale datasets [6, 9, 13], and their overall detection rates versus different thresholds or window number have also been reported. Yet such partial performance summaries give us little idea of a method’s strengths and weaknesses for further improvement, and users are still facing difficulties in choosing methods for their applications. Therefore, more detailed analysis of existing state-of-the-arts is critical for future research and applications. Our contributions can be summarized in three aspects. First, we investigate the influence of object-level characteristics over state-of-the-art object proposal methods for the first time. Although there are some similar works in categorical object detection, few research has been conducted on object proposal side to the best of our knowledge. Second, we introduce the concept of localization latency to evaluate a method’s localization efficiency and accuracy. Third, we create a fully annotated PASCAL VOC dataset with various object-level characteristics to facilitate our analysis. The annotations take us nearly one month’s time which will be released to facilitate further related research. Our experiments are based on PASCAL VOC2007 test set, which has been widely used in evaluating object proposal methods. A proposed window B is treated as detected if its Intersection-over-Union (IoU) with a ground truth bounding box B: IoU(B,B) = area(B ∩ B) area(B ∪ B) is above a certain threshold T . We first study the localization accuracy of the existing methods. The region based methods have higher localization accuracy than window based methods. MCG and SelectiveSearch are the top performing region based methods, though window based EdgeBox shows comparable performance. The localization accuracy for region based methods are similar. One potential explanation is that all region based methods follow similar pipeline by grouping superpixels with either learned or handcrafted edge measures. A good object proposal method should not only produce candidates with high accuracy, but also use as less windows as possible. To summarize a method’s performance in terms of the accuracy and window number, we propose the localization latency metric:
Hongyuan Zhu 0002, Shijian Lu, Jianfei Cai 0001, Guangqing Lee
BMVC1
2014 Poselet-based multiple human identification and cosegmentation
abstract
Localizing, identifying and extracting human groups with consistent appearance jointly from a personal photo stream is an important problem and has wide applications. Inspired by recent advances in object detection, scene understanding and image cosegmentation, in this paper we explore explicit constraints to label and segment human objects rather than other non-human objects and “stuff”. We propose a novel soft human shape cue, which is initialized by color line poselet-based human part detection, further processed through a generalized geodesic distance transform, and refined finally with a joint bilateral filter. Such a high-level object cue is then integrated with other low-level unary and pairwise terms into a principled conditional random field framework, which can be efficiently solved by fast graph cut algorithms. We evaluate our algorithm over the FlickrMFC human dataset, and show that it achieves state-of-the-art performance for this challenging task.
Hongyuan Zhu 0002, Jiangbo Lu, Jianfei Cai 0001, Jianmin Zheng, Nadia Magnenat-Thalmann
ICIP1
2014 Multiple foreground recognition and cosegmentation: An object-oriented CRF model with robust higher-order potentials
abstract
Localizing, recognizing, and segmenting multiple foreground objects jointly from a general user's photo stream that records a specific event is an important task with many useful applications. As argued in recent Multiple Foreground Cosegmentation (MFC) work by Kim and Xing, this task is very challenging in that it contrasts substantially from the classical cosegmentation problem, and aims to parse a set of realistic event photos but each containing irregularly occurring multiple foregrounds with high appearance and scene configuration variations. Inspired by the impressive advance in scene understanding and object recognition, this paper casts the multiple foreground recognition and cosegmentation (MFRC) problem within a conditional random fields (CRFs) framework in a principled manner. We capitalize centrally on the key objective that MFRC is to segment out and annotate foreground objects or “things” rather than “stuff”. To this end, we exploit a few complementary objectness cues (e.g. contours, object detectors and layout) and propose novel and efficient methods to capture object-level information. Integrating object potentials as soft constraints (e.g. robust higher-order potentials defined over detected object regions) with low-level unary and pairwise terms holistically, we solve the MFRC task with a probabilistic CRF model. The inference for such a CRF model is performed efficiently with graph cut based move making algorithms. With a minimal amount of user annotations on just a few example photos, the proposed approach produces spatially coherent, boundary-aligned segmentation results with correct and consistent object labeling. Experiments on the FlickrMFC dataset justify that our method achieves state-of-the-art performance.
Hongyuan Zhu 0002, Jiangbo Lu, Jianfei Cai 0001, Jianmin Zheng, Nadia Magnenat-Thalmann
WACV1
2013 Salient object cutout using Google images
abstract
Given any image input by users, how to automatically cutout the object-of-interest is a challenging problem due to lack of information of the object-of-interest and the background. Saliency detection techniques are able to provide some rough information about object-of-interest since they highlight high-contrast or high attention regions or pixels. However, the generated saliency map is often noisy and directly applying it for segmentation often leads to erroneous results. Motivated by the recent progress on image co-segmentation and internet image retrieval techniques, in this paper, we propose to use the user input image for segmentation as a query image to Google Images and then employ the top returned Google images to build up the knowledge about the object-of-interest in the user input image. Particularly, we develop a lightweight algorithm to learn the knowledge of the object-of-interest in the retrieved images to enhance the saliency map of the input image. Then, the enhanced saliency map is used to initialize the graph-cut to extract the object-of-interest. Experiments with the Mcgill dataset and multiple challenge cases demonstrate the effectiveness of our method in terms of producing a clean cutout.
Hongyuan Zhu 0002, Jianfei Cai 0001, Jianmin Zheng, Jianxin Wu 0001, Nadia Magnenat-Thalmann
ISCAS1
2013 Object-Level Image Segmentation Using Low Level Cues
abstract
This paper considers the problem of automatically segmenting an image into a small number of regions that correspond to objects conveying semantics or high-level structure. Although such object-level segmentation usually requires additional high-level knowledge or learning process, we explore what low level cues can produce for this purpose. Our idea is to construct a feature vector for each pixel, which elaborately integrates spectral attributes, color Gaussian mixture models, and geodesic distance, such that it encodes global color and spatial cues as well as global structure information. Then, we formulate the Potts variational model in terms of the feature vectors to provide a variational image segmentation algorithm that is performed in the feature space. We also propose a heuristic approach to automatically select the number of segments. The use of feature attributes enables the Potts model to produce regions that are coherent in color and position, comply with global structures corresponding to objects or parts of objects and meanwhile maintain a smooth and accurate boundary. We demonstrate the effectiveness of our algorithm against the state-of-the-art with the data set from the famous Berkeley benchmark.
Hongyuan Zhu 0002, Jianmin Zheng, Jianfei Cai 0001, Nadia Magnenat-Thalmann
IEEE Trans. Image Process.1