Yuanfang Guo

dblp:78/8545 · also Andy Yuanfang Guo · DBLP profile ↗
← Back
100ranked-venue papers
6as first author
52since 2021 · last 2026
0000-0003-4592-8083ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 64 · 3 first-author · 27 since 2021Artificial intelligence and machine learning · 37 · 26 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 7 since 2021Systems, architecture and hardware · 6 · 1 first-author · 3 since 2021Security and privacy · 6 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 5 · 3 since 2021Computer networks · 4 · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Activation Manipulation Attack: Penetrating and Harmful Jailbreak Attack Against Large Vision-Language Models
abstract
Recently, Large Vision-Language Models (LVLMs) have been demonstrated to be vulnerable to jailbreak attacks, highlighting the urgent need for further research to comprehensively identify and mitigate these threats. Unfortunately, existing jailbreak studies primarily focus on coarse-grained input manipulation to elicit specific responses, overlooking the exploitation of internal representations, i.e., intermediate activations, which constrains their ability to penetrate alignment safeguards and generate harmful responses. To tackle this issue, we propose the Activation Manipulation (ActMan) Attack framework, which performs fine-grained activation manipulations inspired by the perception and cognition stages of human decision-making, enhancing both the penetration capability and harmfulness of attacks. To improve penetration capability, we introduce a Deceptive Visual Camouflage module inspired by the masking effect in human perception. This module uses a benign activation-guided attention redirection strategy to conceal abnormal activation patterns, thereby suppressing LVLM's defense detection during early-stage decoding. To enhance harmfulness, we design a Malicious Semantic Induction module drawing from the framing effect in human cognition, which reconstructs jailbreak instructions using malicious activation guidance to change LVLM’s risk assessment during late-stage decoding, thereby amplifying the harmfulness of model responses. Extensive experiments on six mainstream LVLMs demonstrate that our method remarkably outperforms state-of-the-art baselines, achieving an average relative ASR improvement of 12.06%.
Haojie Hao, Jiakai Wang, Aishan Liu, Yuqing Ma, Haotong Qin, Yuanfang Guo, Xianglong Liu 0001
AAAI6
2026 Towards multi-language repository-level code generation: From-scratch to guided tasks
Silin Li, Zeming Liu, Yuhang Guo 0001, Yuanfang Guo, Yunhong Wang 0001, Haifeng Wang 0001
Neurocomputing6
2026 Characterization of the heterogeneity in SARS-CoV-2 fitness dynamics via graph representation learning
abstract
Understanding the heterogeneity of population-level viral fitness dynamics, which reflect the interplay between intrinsic viral properties and population immunity, is critical for pandemic preparedness. However, how these dynamics vary across diverse immune backgrounds and mutational landscapes remain poorly characterized. We present Geno-GNN, a graph representation learning approach for retrospectively characterizing the viral fitness dynamics of severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2). Geno-GNN accurately predicts angiotensin-converting enzyme 2 (ACE2) binding affinity and immune escape potential across multiple external datasets. Using Geno-GNN, we identified temporal patterns in SARS-CoV-2 fitness and detected varying rates of fitness change associated with distinct immune backgrounds. Virtual mutation scanning revealed two fitness trajectories: broad immune evasion at the cost of ACE2 affinity and ACE2 affinity maintenance at or above the Wuhan-Hu-1 level along with moderate immune escape. Notably, real-world SARS-CoV-2 variants predominantly followed the latter trajectory, sustaining ACE2 affinity via fixed mutations. These findings underscore the heterogeneous, immune-contextualized nature of viral fitness dynamics and the complex evolutionary pathways of SARS-CoV-2.
Zengmiao Wang, Ziqin Zhou, Junfu Wang, Lingyue Yang, Zhirui Zhang, Weina Xu, Zeming Liu, Yuxi Ge, Liang Yang 0002, Quanyi Wang, Yunlong Cao, Yuanfang Guo, Huaiyu Tian
PLoS Comput. Biol.14
2026 HFA2RE: Enhancing adversarial robustness via Hyperspherical Feature Aggregation
Heqi Peng, Mingxuan Chen, Yunhong Wang 0001, Yuanfang Guo
Pattern Recognit.4
2026 Continual Adversarial Example Detection via Incremental Attack Configuration Within a Knowledge Distillation Framework
abstract
Adversarial example detection has emerged as a prominent defense strategy owing to its efficiency in training and deployment. Nevertheless, existing detectors are typically developed under a single-step paradigm, where models become static after training on adversarial examples generated by a single attack. This paradigm is infeasible in dynamic real-world scenarios, since retraining from scratch for each newly encountered attack is impractical and computationally prohibitive. To address this limitation, we propose Continual Adversarial example Detection via Incremental Attack Configuration (IAC-CAD), which pioneers to exploit continual learning for adversarial detection within a knowledge distillation framework. IAC-CAD constructs a sequence of continuous detection tasks which require only a limited number of samples per task. Moreover, the proposed Incremental Attack Configuration (IAC) mechanism selects the representative attacks which maximally cover the entire adversarial feature space and optimizes their training sequence through the Memory-aware Attack Ordering. This design simultaneously mitigates catastrophic forgetting of known attacks and enhances generalization ability against unseen attacks. Extensive experiments verify the superiority and practicality of IAC-CAD.
Heqi Peng, Yunhong Wang 0001, Jiantao Zhou 0001, Yuanfang Guo
IEEE Signal Process. Lett.5
2026 A Perceptual Distortion Reduction Framework: Toward Generating Adversarial Examples With High Perceptual Quality and Attack Success Rate
Ruijie Yang, Yuanfang Guo, Ruikui Wang, Jiantao Zhou 0001, Yunhong Wang 0001
IEEE Trans. Dependable Secur. Comput.2
2026 A Multi-Grained Parallel Spatio-Temporal Learning Architecture for Deepfake Video Detection
abstract
With advances in generation techniques, malicious users can easily generate deepfake videos, which can cause severe social problems and trust issues. Therefore, deepfake video detection has received increasing attention in recent years. Given that forgery clues are often subtle and imperceptible, effective detection relies heavily on multi-grained learning. However, existing approaches fail to systematically incorporate multi-grained learning across the key components of network training—namely, the training data, network structure, and supervision strategy—thus limiting their performance. In this article, we propose a multi-grained parallel spatio-temporal deepfake video detection architecture, which introduces a novel framework to mine more discriminative deepfake cues throughout the training pipeline. Firstly, we design a parallel spatio-temporal network combined with a cross-guided mechanism to concurrently extract frame-level spatial features and patch-level temporal features, while leveraging the relationship between spatial artifacts and temporal inconsistencies to enable multi-grained spatio-temporal synchronous learning. Secondly, we propose segment-level data augmentation strategies, including frame-random consistent self-blending and spatio-temporal data augmentation, which improve training data diversity at both frame and patch levels, thereby improving the model’s ability to learn comprehensive deepfake representations. Finally, we construct a multi-grained supervision, comprising a patch-level temporal loss, a distance-based frame-level spatial loss, and a standard segment-level loss, for subtle deepfake feature learning. Extensive experiments demonstrate that our method possesses strong robustness and the generalization ability outperforms the current state-of-the-art methods across a series of deepfake datasets, including FaceForensics++, CelebDF, DFDC, DeeperForensics, and Faceshifter, on average.
Yuanfang Guo, Leo Yu Zhang, Jiantao Zhou 0001, Yunhong Wang 0001
ACM Trans. Multim. Comput. Commun. Appl.2
2025 Multi-modal Deepfake Detection via Multi-task Audio-Visual Prompt Learning
abstract
With the malicious use and dissemination of multi-modal deepfake videos, researchers start to investigate multi-modal deepfake detection. Unfortunately, most of the existing methods tune all the parameters of the deep network with limited speech video datasets and are trained under coarse-grained consistency supervision, which hinders their generalization ability in practical scenarios. To solve these problems, in this paper, we propose the first multi-task audio-visual prompt learning method for multi-modal deepfake video detection, by exploiting multiple foundation models. Specifically, we construct a two-stream multi-task learning architecture and propose sequential visual prompts and short-time audio prompts to extract multi-modal features, which are aligned at the frame level and utilized in subsequent fine-grained feature matching and fusion. Due to the natural alignment of visual content and audio signal in real data, we propose a frame-level cross-modal feature matching loss function to learn the fine-grained audio-visual consistency. Comprehensive experiments demonstrate the effectiveness and superior generalization ability of our method against the state-of-the-art methods.
Yuanfang Guo, Zeming Liu, Yunhong Wang 0001
AAAI2
2025 EntropyMark: Towards More Harmless Backdoor Watermark via Entropy-based Constraint for Open-source Dataset Copyright Protection
abstract
High-quality open-source datasets are essential for advancing deep neural networks. However, the unauthorized commercial use of these datasets has raised significant concerns about copyright protection. One promising approach is backdoor watermark-based dataset ownership verification (BW-DOV), in which dataset protectors implant specific backdoors into illicit models through dataset watermarking, enabling the tracing of these models through abnormal prediction behaviors. Unfortunately, the targeted nature of these BW-DOV methods can be maliciously exploited, potentially leading to harmful side effects. While existing harmless methods attempt to mitigate these risks, watermarked datasets can still negatively affect prediction results, partially compromising dataset functionality. In this paper, we propose a more harmless backdoor watermark, called EntropyMark, which improves prediction confidence without altering the final prediction results. For this purpose, an entropy-based constraint is introduced to regulate the probability distribution. Specifically, we design an iterative clean-label dataset watermarking framework. Our framework employs gradient matching and adaptive data selection to optimize backdoor injection. In parallel, we introduce a hypothesis test method grounded in entropy inconsistency to verify dataset ownership. Extensive experiments on benchmark datasets demonstrate the effectiveness, transferability, and defense resistance of our approach. Our code is available at https://github.com/AaronSun2000/EntropyMark.
Ming Sun 0010, Rui Wang 0032, Zixuan Zhu 0002, Lihua Jing, Yuanfang Guo
CVPR5
2025 RETAIL: Towards Real-world Travel Planning for Large Language Models
abstract
Although large language models have enhanced automated travel planning abilities, current systems remain misaligned with real-world scenarios.First, they assume users provide explicit queries, while in reality requirements are often implicit.Second, existing solutions ignore diverse environmental factors and user preferences, limiting the feasibility of plans.Third, systems can only generate plans with basic POI arrangements, failing to provide all-in-one plans with rich details.To mitigate these challenges, we construct a novel dataset RETAIL, which supports decision-making for implicit queries while covering explicit queries, both with and without revision needs.It also enables environmental awareness to ensure plan feasibility under real-world scenarios, while incorporating detailed POI information for allin-one travel plans.Furthermore, we propose a topic-guided multi-agent framework, termed TGMA.Our experiments reveal that even the strongest existing model achieves merely a 1.0% pass rate, indicating real-world travel planning remains extremely challenging.In contrast, TGMA demonstrates substantially improved performance 2.72%, offering promising directions for real-world travel planning.1 sistant for personalized travel planning.
Yizhe Feng, Zeming Liu, Xiangrong Zhu 0002, Yuanfang Guo, Yunhong Wang 0001
EMNLP7
2025 Robust Deepfake Detection via Perturbation Domain Alignment
abstract
Deepfake detection becomes vital in distinguishing the real image/videos from the fake ones, which are produced via advanced deep learning based face manipulation techniques. Although existing approaches exhibit decent generalization, they struggle to maintain good robustness against diverse perturbations in practical scenarios. Perturbations, which can induce distortions on the original image/video, such as compression, Gaussian noise, blur, etc., tend to introduce negative impacts on the performance of deepfake detection models. Therefore, in this paper, we propose a novel deepfake detection method, named Robust Deepfake Detection via Perturbation Domain Alignment (PDA-RDD), by exploiting the mechanism of domain alignment. Our approach consider different perturbations as distinct domains, and proposes a paired instance momentum whitening (PIMW) module to align these domains, to effectively remove the sensitive information associated with these perturbations. To further enhance PIMW, we construct an MLP projector (MLPP) to project the encoded feature into a more optimal latent vector. Extensive experiments demonstrate the effectiveness of our method on multiple widely used datasets.
Yunhong Wang 0001, Yuanfang Guo
ICASSP4
2025 DUALFormer: Dual Graph Transformer
abstract
Graph Transformers (GTs), adept at capturing the locality and globality of graphs, have shown promising potential in node classification tasks. Most state-of-the-art GTs succeed through integrating local Graph Neural Networks (GNNs) with their global Self-Attention (SA) modules to enhance structural awareness. Nonetheless, this architecture faces limitations arising from scalability challenges and the trade-off between capturing local and global information. On the one hand, the quadratic complexity associated with the SA modules poses a significant challenge for many GTs, particularly when scaling them to large-scale graphs. Numerous GTs necessitated a compromise, relinquishing certain aspects of their expressivity to garner computational efficiency. On the other hand, GTs face challenges in maintaining detailed local structural information while capturing long-range dependencies. As a result, they typically require significant computational costs to balance the local and global expressivity. To address these limitations, this paper introduces a novel GT architecture, dubbed DUALFormer, featuring a dual-dimensional design of its GNN and SA modules. Leveraging approximation theory from Linearized Transformers and treating the query as the surrogate representation of node features, DUALFormer \emph{efficiently} performs the computationally intensive global SA module on feature dimensions. Furthermore, by such a separation of local and global modules into dual dimensions, DUALFormer achieves a natural balance between local and global expressivity. In theory, DUALFormer can reduce intra-class variance, thereby enhancing the discriminability of node representations. Extensive experiments on eleven real-world datasets demonstrate its effectiveness and efficiency over existing state-of-the-art GTs.
Jiaming Zhuo, Yintong Lu, Ziyi Ma, Chuan Wang 0002, Yuanfang Guo, Zhen Wang 0004, Xiaochun Cao, Liang Yang 0002
ICLR7
2025 Disentangled Graph Spectral Domain Adaptation
abstract
The distribution shifts and the scarcity of labels prevent graph learning methods, especially graph neural networks (GNNs), from generalizing across domains. Compared to Unsupervised Domain Adaptation (UDA) with embedding alignment, Unsupervised Graph Domain Adaptation (UGDA) becomes more challenging in light of the attribute and topology entanglement in the representation. Beyond embedding alignment, UGDA turns to topology alignment but is limited by the ability of the employed topology model and the estimation of pseudo labels. To alleviate this issue, this paper proposed a Disentangled Graph Spectral Domain adaptation (DGSDA) by disentangling attribute and topology alignments and directly aligning flexible graph spectral filters beyond topology. Specifically, Bernstein polynomial approximation, which mimics the behavior of the function to be approximated to a remarkable degree, is employed to capture complicated topology characteristics and avoid the expensive eigenvalue decomposition. Theoretical analysis reveals the tight GDA bound of DGSDA and the rationality of polynomial coefficient regularization. Quantitative and qualitative experiments justify the superiority of the proposed DGSDA.
Liang Yang 0002, Jiaming Zhuo, Di Jin 0001, Chuan Wang 0002, Xiaochun Cao, Zhen Wang 0004, Yuanfang Guo
ICML8
2025 Universal Graph Self-Contrastive Learning
abstract
As a pivotal architecture in Self-Supervised Learning (SSL), Graph Contrastive Learning (GCL) has demonstrated substantial application value in scenarios with limited labeled nodes (samples). However, existing GCLs encounter critical issues in the graph augmentation and positive and negative sampling stemming from the lack of explicit supervision, which collectively restrict their efficiency and universality. On the one hand, the reliance on graph augmentations in existing GCLs can lead to increased training times and memory usage, while potentially compromising the semantic integrity. On the other hand, the difficulty in selecting TRUE positive and negative samples for GCLs limits their universality to both homophilic and heterophilic graphs. To address these drawbacks, this paper introduces a novel GCL framework called GRAph learning via Self-contraSt (GRASS). The core mechanism is node-attribute self-contrast, which specifically involves increasing the feature similarities between nodes and their included attributes while decreasing the similarities between nodes and their non-included attributes. Theoretically, the self-contrast mechanism implicitly ensures accurate node-node contrast by capturing high-hop co-inclusion relationships, thereby enabling GRASS to be universally applicable to graphs with varying degrees of homophily. Evaluations on diverse benchmark datasets demonstrate the universality and efficiency of GRASS. The dataset and code are available at URL: https://github.com/YukunCai/GRASS.
Liang Yang 0002, Yukun Cai, Hui Ning, Jiaming Zhuo, Di Jin 0001, Ziyi Ma, Yuanfang Guo, Chuan Wang 0002, Zhen Wang 0004
IJCAI7
2025 Feature Perturbation Agent based Adversarial Attack Method for Weakly Supervised Video Anomaly Detection
abstract
Weakly supervised video anomaly detection (WS-VAD) techniques, based on video backbone models, are widely used in surveillance but are vulnerable to adversarial attacks. However, directly applying existing methods causes high memory consumption and low efficiency, and adversarial attacks on WS-VAD models have yet to be specifically studied. In this paper, we pioneer to propose a two-staged Feature Perturbation Agent based Adversarial Attack (FPAgent) method for WS-VAD. To better deceive detection models, we explore the deceivable feature spaces. To describe the locations of the deceivable feature spaces, we propose a feature perturbation agent, which also transforms the complex video-level attack into a simple segment-level attack. Besides, we propose a perturbation guider strategy to guide the feature vectors into the deceivable feature spaces, by computing the perturbation from the first segment of each video. The experiments have verified the effectiveness, as well as the attack efficiency and low memory consumption of our method.
Zhen Yang 0037, Yuanfang Guo, Ruijie Yang, Di Huang 0001, Jiantao Zhou 0001
ISCAS2
2025 Anomaly-aware self-supervised feature learning for weakly supervised video anomaly detection
Zhen Yang 0037, Guodong Wang 0006, Yuanfang Guo, Xiuguo Bao, Di Huang 0001
Comput. Vis. Image Underst.3
2025 Common knowledge learning for generating transferable adversarial examples
Ruijie Yang, Yuanfang Guo, Junfu Wang, Jiantao Zhou 0001, Yunhong Wang 0001
Frontiers Comput. Sci.2
2025 Vector Quantization Based Query-Efficient Attack via Direct Preference Optimization
abstract
This work studies black-box adversarial attacks against deep neural networks, where the attacker only has access to the query feedback from the target model. The current state-of-the-art (SOTA) query-efficient attacks usually combine transfer-based and query-based methods by utilizing the gradient or initializations of surrogate models. However, these strategies typically incur significant computational costs and require a large number of queries during the attack process. In this paper, we propose a novel query-efficient method for generating black-box adversarial perturbations, named Vector Quantization based Query-efficient Adversarial Perturbation generation (VQQAP). Specifically, we propose a Nucleus Sampling based Discretization Module (NSDM) to create diverse adversarial examples in the discrete latent space. To directly optimize the latent vector, we formulate the optimization problem as a direct preference optimization (DPO) problem, and iteratively solve this problem based on the target model feedback. Experimental evaluations demonstrate the effectiveness and efficiency of our method.
Ruijie Yang, Yuanfang Guo, Guohao Li 0010, Yunhong Wang 0001
IEEE Signal Process. Lett.2
2025 ALD-GCN: Graph Convolutional Networks With Attribute-Level Defense
abstract
Graph Neural Networks(GNNs), such as Graph Convolutional Network, have exhibited impressive performance on various real-world datasets. However, many researches have confirmed that deliberately designed adversarial attacks can easily confuse GNNs on the classification of target nodes (targeted attacks) or all the nodes (global attacks). According to our observations, different attributes tend to be differently treated when the graph is attacked. Unfortunately, most of the existing defense methods can only defend at the graph or node level, which ignores the diversity of different attributes within each node. To address this limitation, we propose to leverage a new property, named Attribute-level Smoothness (ALS), which is defined based on the local differences of graph. We then propose a novel defense method, named GCN with Attribute-level Defense (ALD-GCN), which utilizes the ALS property to provide attribute-level protection to each attributes. Extensive experiments on real-world graphs have demonstrated the superiority of the proposed work and the potentials of our ALS property in the attacks.
Yuanfang Guo, Junfu Wang, Shihao Nie, Liang Yang 0002, Di Huang 0001, Yunhong Wang 0001
IEEE Trans. Big Data2
2025 AED-PADA: Improving Generalizability of Adversarial Example Detection via Principal Adversarial Domain Adaptation
abstract
Adversarial example detection, which can be conveniently applied in many scenarios, is important in the area of adversarial defense. Unfortunately, existing detection methods suffer from poor generalization performance because their training process usually relies on the examples generated from a single known adversarial attack and there exists a large discrepancy between the training and unseen testing adversarial examples. To address this issue, we propose a novel method, named Adversarial Example Detection via Principal Adversarial Domain Adaptation (AED-PADA). Specifically, our approach identifies the Principal Adversarial Domains (PADs), i.e., a combination of features of the adversarial examples generated by different attacks, which possesses a large portion of the entire adversarial feature space. Subsequently, we pioneer to exploit Multi-source Unsupervised Domain Adaptation in adversarial example detection, with PADs as the source domains. Experimental results demonstrate the superior generalization ability of our proposed AED-PADA. Note that this superiority is particularly achieved in challenging scenarios characterized by employing the minimal magnitude constraint for the perturbations.
Heqi Peng, Yunhong Wang 0001, Ruijie Yang, Beichen Li 0001, Rui Wang 0032, Yuanfang Guo
ACM Trans. Multim. Comput. Commun. Appl.6
2024 AGS: Affordable and Generalizable Substitute Training for Transferable Adversarial Attack
abstract
In practical black-box attack scenarios, most of the existing transfer-based attacks employ pretrained models (e.g. ResNet50) as the substitute models. Unfortunately, these substitute models are not always appropriate for transfer-based attacks. Firstly, these models are usually trained on a largescale annotated dataset, which is extremely expensive and time-consuming to construct. Secondly, the primary goal of these models is to perform a specific task, such as image classification, which is not developed for adversarial attacks. To tackle the above issues, i.e., high cost and over-fitting on taskspecific models, we propose an Affordable and Generalizable Substitute (AGS) training framework tailored for transferbased adversarial attack. Specifically, we train the substitute model from scratch by our proposed adversary-centric constrastive learning. This proposed learning mechanism introduces another sample with slight adversarial perturbations as an additional positive view of the input image, and then encourages the adversarial view and two benign views to interact comprehensively with each other. To further boost the generalizability of the substitute model, we propose adversarial invariant learning to maintain the representations of the adversarial example invariants under augmentations with various strengths. Our AGS model can be trained solely with unlabeled and out-of domain data and avoid overfitting to any task-specific models, because of its inherently self-supervised nature. Extensive experiments demonstrate that our AGS achieves comparable or superior performance compared to substitute models pretrained on the complete ImageNet training set, when executing attacks across a diverse range of target models, including ViTs, robustly trained models, object detection and segmentation models. Our source codes are available at https://github.com/lwmming/AGS.
Ruikui Wang, Yuanfang Guo, Yunhong Wang 0001
AAAI2
2024 Improving Graph Contrastive Learning via Adaptive Positive Sampling
abstract
Graph Contrastive Learning (GCL), a Self-Supervised Learning (SSL) architecture tailored for graphs, has shown notable potential for mitigating label scarcity. Its core idea is to amplify feature similarities between the positive sample pairs and reduce them between the negative sample pairs. Unfortunately, most existing GCLs consistently present sub-optimal performances on both homophilic and heterophilic graphs. This is primarily attributed to two limitations of positive sampling, that is, incomplete local sampling and blind sampling. To address these limitations, this paper introduces a novel GCL framework with an adaptive positive sampling module, named grapH contrastivE Adaptive Positive Samples (HEATS). Motivated by the observation that the affinity matrix corresponding to optimal positive sample sets has a block-diagonal structure with equal weights within each block, a self-expressive learning objective incorporating the block and idempotent constraint is presented. This learning objective and the contrastive learning objective are iteratively optimized to improve the adaptability and robustness of HEATS. Extensive experiments on graphs and images validate the effectiveness and generality of HEATS.
Jiaming Zhuo, Feiyang Qin, Can Cui 0005, Bingxin Niu, Mengzhu Wang, Yuanfang Guo, Chuan Wang 0002, Zhen Wang 0004, Xiaochun Cao, Liang Yang 0002
CVPR7
2024 Deepfake Detection Via Separable Self-Consistency Learning
abstract
Deepfake detection technologies have been developed rapidly in recent years, due to the potential severe security threats induced by the realistic deep facial forgeries. Among the existing deepfake detection methods, self-supervised methods have drawn significant attentions from researchers, because of their better generalization ability against the deep forgeries produced via unseen deepfake techniques. Unfortunately, existing state-of-the-art self-supervised approaches have not properly considered that different pairs of patches from different regions actually give different contributions. Thus, their learned representations are coarse and the generalization performances are less decent. In this paper, we propose a new self-supervised deepfake detection method, named deepfake detection via separable self-consistency learning (SSCLDFD), to improve the generalization ability of deepfake detection. Specifically, to effectively extract detection features, we construct a multi-scale Texture Enhanced Feature Extraction Network (TEFEN), by forming a Central-Difference based Convolution Module (CDCM) to enhance the texture information, which contain rich forgery cues. Since different pairs of patches from different regions (i.e. background and facial regions) tend to give various consistencies, we propose a separable self-consistency loss to explicitly constrain the representation learning. Extensive experiments demonstrate that our SSCL-DFD can give superior generalization performances compared to the state-of-the-art methods.
Yunhong Wang 0001, Wenqi Zhuo, Guangshuai Gao, Yuanfang Guo
ICIP6
2024 Understanding Heterophily for Graph Neural Networks
abstract
Graphs with heterophily have been regarded as challenging scenarios for Graph Neural Networks (GNNs), where nodes are connected with dissimilar neighbors through various patterns. In this paper, we present theoretical understandings of heterophily for GNNs by incorporating the graph convolution (GC) operations into fully connected networks via the proposed Heterophilous Stochastic Block Models (HSBM), a general random graph model that can accommodate diverse heterophily patterns. Our theoretical investigation comprehensively analyze the impact of heterophily from three critical aspects. Firstly, for the impact of different heterophily patterns, we show that the separability gains are determined by two factors, i.e., the Euclidean distance of the neighborhood distributions and $\sqrt{\mathbb{E}\left[\operatorname{deg}\right]}$, where $\mathbb{E}\left[\operatorname{deg}\right]$ is the averaged node degree. Secondly, we show that the neighborhood inconsistency has a detrimental impact on separability, which is similar to degrading $\mathbb{E}\left[\operatorname{deg}\right]$ by a specific factor. Finally, for the impact of stacking multiple layers, we show that the separability gains are determined by the normalized distance of the $l$-powered neighborhood distributions, indicating that nodes still possess separability in various regimes, even when over-smoothing occurs. Extensive experiments on both synthetic and real-world data verify the effectiveness of our theory.
Junfu Wang, Yuanfang Guo, Liang Yang 0002, Yunhong Wang 0001
ICML2
2024 Unified Graph Augmentations for Generalized Contrastive Learning on Graphs
abstract
In real-world scenarios, networks (graphs) and their tasks possess unique characteristics, requiring the development of a versatile graph augmentation (GA) to meet the varied demands of network analysis. Unfortunately, most Graph Contrastive Learning (GCL) frameworks are hampered by the specificity, complexity, and incompleteness of their GA techniques. Firstly, GAs designed for specific scenarios may compromise the universality of models if mishandled. Secondly, the process of identifying and generating optimal augmentations generally involves substantial computational overhead. Thirdly, the effectiveness of the GCL, even the learnable ones, is constrained by the finite selection of GAs available. To overcome the above limitations, this paper introduces a novel unified GA module dubbed UGA after reinterpreting the mechanism of GAs in GCLs from a message-passing perspective. Theoretically, this module is capable of unifying any explicit GAs, including node, edge, attribute, and subgraph augmentations. Based on the proposed UGA, a novel generalized GCL framework dubbed Graph cOntrastive UnifieD Augmentations (GOUDA) is proposed. It seamlessly integrates widely adopted contrastive losses and an introduced independence loss to fulfill the common requirements of consistency and diversity of augmentation across diverse scenarios. Evaluations across various datasets and tasks demonstrate the generality and efficiency of the proposed GOUDA over existing state-of-the-art GCLs.
Jiaming Zhuo, Yintong Lu, Hui Ning, Bingxin Niu, Dongxiao He, Chuan Wang 0002, Yuanfang Guo, Zhen Wang 0004, Xiaochun Cao, Liang Yang 0002
NeurIPS8
2024 GAUSS: GrAph-customized Universal Self-Supervised Learning
abstract
To make Graph Neural Networks (GNNs) meet the requirements of the Web, the universality and the generalization become two important research directions. On one hand, many universal GNNs are presented for semi-supervised tasks on both homophilic and non-homophilic graphs by distinguishing homophilic and heterophilic edges with the help of labels. On the other hand, self-supervised learning (SSL) algorithms on graphs are presented by leveraging the self-supervised learning schemes from computer vision and natural language processing. Unfortunately, graph universal self-supervised learning remains resolved. Most existing SSL methods on graphs, which often employ two-layer GCN as the encoder and train the mapping functions, can't alter the low-passing filtering characteristic of GCN. Therefore, to be universal, SSL must becustomized for the graph, i.e., learning the graph. However, learning the graph via universal GNNs is disabled in SSL, since their distinguishability on homophilic and heterophilic edges disappears without the labels. To overcome this difficulty, this paper proposes novel GrAph-customized Universal Self-Supervised Learning (GAUSS) by exploiting local attribute distribution. The main idea is to replace the global parameters with locally learnable propagation. To make the propagation matrix demonstrate the affinity between the nodes, the self-representative learning framework is employed with k-block diagonal regularization. Extensive experiments on synthetic and real-world datasets demonstrate its effectiveness, universality and robustness to noises.
Liang Yang 0002, Weixiao Hu, Jizhong Xu, Runjie Shi, Dongxiao He, Chuan Wang 0002, Xiaochun Cao, Zhen Wang 0004, Bingxin Niu, Yuanfang Guo
WWW10
2024 Graph Contrastive Learning Reimagined: Exploring Universality
abstract
Real-world graphs exhibit diverse structures, including homophilic and heterophilic patterns, necessitating the development of a universal Graph Contrastive Learning (GCL) framework. Nonetheless, the existing GCLs, especially those with a local focus, lack universality due to the mismatch between the input graph structure and the homophily assumption for two primary components of GCLs. Firstly, the encoder, commonly Graph Convolution Network (GCN), operates as a low-pass filter, which assumes the input graph to be homophilic. This makes it challenging to aggregate features from neighbor nodes of the same class on heterophilic graphs. Secondly, the local positive sampling regards neighbor nodes as positive samples, which is inspired by the homophily assumption. This results in feature similarity amplification for the samples from the different classes (i.e., FALSE positive samples). Therefore, it is crucial to feed the encoder and positive sampling of GCLs with homophilic graph structures. This paper presents a novel GCL framework, named gRaph cOntraStive Exploring uNiversality (ROSEN), designed to achieve this objective. Specifically, ROSEN equips a local graph structure inference module, utilizing the Block Diagonal Property (BDP) of the affinity matrix extracted from node ego networks. This module can generate the homophilic graph structure by selectively removing disassortative edges. Extensive evaluations validate the effectiveness and universality of ROSEN across node classification and node clustering tasks.
Jiaming Zhuo, Can Cui 0005, Bingxin Niu, Dongxiao He, Chuan Wang 0002, Yuanfang Guo, Zhen Wang 0004, Xiaochun Cao, Liang Yang 0002
WWW7
2024 Exploring transferable and robust adversarial perturbation generation across network hierarchy
Ruikui Wang, Yuanfang Guo, Ruijie Yang, Yunhong Wang 0001
Neurocomputing2
2024 Binary Graph Convolutional Network With Capacity Exploration
abstract
The current success of Graph Neural Networks (GNNs) usually relies on loading the entire attributed graph for processing, which may not be satisfied with limited memory resources, especially when the attributed graph is large. This paper pioneers to propose a Binary Graph Convolutional Network (Bi-GCN), which binarizes both the network parameters and input node attributes and exploits binary operations instead of floating-point matrix multiplications for network compression and acceleration. Meanwhile, we also propose a new gradient approximation based back-propagation method to properly train our Bi-GCN. According to the theoretical analysis, our Bi-GCN can reduce the memory consumption by an average of ∼ 31x for both the network parameters and input data, and accelerate the inference speed by an average of ∼ 51x, on three citation networks, i.e., Cora, PubMed, and CiteSeer. Besides, we introduce a general approach to generalize our binarization method to other variants of GNNs, and achieve similar efficiencies. Although the proposed Bi-GCN and Bi-GNNs are simple yet efficient, these compressed networks may also possess a potential capacity problem, i.e., they may not have enough storage capacity to learn adequate representations for specific tasks. To tackle this capacity problem, an Entropy Cover Hypothesis is proposed to predict the lower bound of the width of Bi-GNN hidden layers. Extensive experiments have demonstrated that our Bi-GCN and Bi-GNNs can give comparable performances to the corresponding full-precision baselines on seven node classification datasets and verified the effectiveness of our Entropy Cover Hypothesis for solving the capacity problem.
Junfu Wang, Yuanfang Guo, Liang Yang 0002, Yunhong Wang 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2024 Heterophily-aware graph attention network
Junfu Wang, Yuanfang Guo, Liang Yang 0002, Yunhong Wang 0001
Pattern Recognit.2
2024 Towards Video Anomaly Detection in the Real World: A Binarization Embedded Weakly-Supervised Network
abstract
In this letter, we pioneer to propose a binarization embedded weakly-supervised video anomaly detection (BE-WSVAD) method by constructing a binarized GCN-based anomaly detection module. Compared to the existing weakly-supervised video anomaly detection (WS-VAD) methods, BE-WSVAD focuses on the detection efficiency, which is ignored by the existing literature yet vital in real applications. Specifically, to improve the detection performance of the binary anomaly detection module, we propose a binary network augmentation strategy in the training process. Due to the weakly supervision mechanism, the videos employed in the training process are usually lengthy, in which the lengthy-input dependencies tend to be exploited to improve the detection performance with extra memory consumption. Then, we propose the short-input inference modes, which can largely reduce the desired length of the input video. Experimental results demonstrate the superiority of our BE-WSVAD in terms of the memory and computational consumptions while giving comparable accuracies.
Zhen Yang 0037, Yuanfang Guo, Junfu Wang, Di Huang 0001, Xiuguo Bao, Yunhong Wang 0001
IEEE Trans. Circuits Syst. Video Technol.2
2023 Deepfake Video Detection via Facial Action Dependencies Estimation
abstract
Deepfake video detection has drawn significant attention from researchers due to the security issues induced by deepfake videos. Unfortunately, most of the existing deepfake detection approaches have not competently modeled the natural structures and movements of human faces. In this paper, we formulate the deepfake video detection problem into a graph classification task, and propose a novel paradigm named Facial Action Dependencies Estimation (FADE) for deepfake video detection. We propose a Multi-Dependency Graph Module (MDGM) to capture abundant dependencies among facial action units, and extracts subtle clues in these dependencies. MDGM can be easily integrated into the existing frame-level detection schemes to provide significant performance gains. Extensive experiments demonstrate the superiority of our method against the state-of-the-art methods.
Lingfeng Tan, Yunhong Wang 0001, Junfu Wang, Liang Yang 0002, Xunxun Chen, Yuanfang Guo
AAAI6
2023 Global-Local Characteristic Excited Cross-Modal Attacks from Images to Videos
abstract
The transferability of adversarial examples is the key property in practical black-box scenarios. Currently, numerous methods improve the transferability across different models trained on the same modality of data. The investigation of generating video adversarial examples with imagebased substitute models to attack the target video models, i.e., cross-modal transferability of adversarial examples, is rarely explored. A few works on cross-modal transferability directly apply image attack methods for each frame and no factors especial for video data are considered, which limits the cross-modal transferability of adversarial examples. In this paper, we propose an effective cross-modal attack method which considers both the global and local characteristics of video data. Firstly, from the global perspective, we introduce inter-frame interaction into attack process to induce more diverse and stronger gradients rather than perturb each frame separately. Secondly, from the local perspective, we disrupt the inherently local correlation of frames within a video, which prevents black-box video model from capturing valuable temporal clues. Extensive experiments on the UCF-101 and Kinetics-400 validate the proposed method significantly improves cross-modal transferability and even surpasses strong baseline using video models as substitute model. Our source codes are available at https://github.com/lwmming/Cross-Modal-Attack.
Ruikui Wang, Yuanfang Guo, Yunhong Wang 0001
AAAI2
2023 A Dual Domain Attention Mechanism for Face Forgery Detection
abstract
Recently, deep face forgery detection has been attracting considerable attentions, due to the potential security consequences induced by this type of forgeries. Unfortunately, the existing techniques have not specifically considered the intrinsic differences between the frequency and spatial domain information. To explicitly accommodate different feature representations from different domains, we propose a novel Dual Domain Attention Mechanism (DDAM) for deep face forgery detection. Inspired by digital image processing, we construct a “soft” filter to adaptively filter the frequency information, which is irrelevant to our forgery detection. Besides, we construct a FC-based Attention Module to maintain a receptive field of the entire feature map, to better leverage contextual information from different domains. Extensive experiments demonstrate the effectiveness of the proposed method on widely used datasets.
Yucong Suo, Xiaohan Zhao, Yuanfang Guo, Yangxi Li, Yunhong Wang 0001
IJCB3
2023 Long Short-Term Graph Memory Against Class-imbalanced Over-smoothing
abstract
Most Graph Neural Networks (GNNs) follow the message-passing scheme. Residual connection is an effective strategy to tackle GNNs' over-smoothing issue and performance reduction issue on non-homophilic networks. Unfortunately, the coarse-grained residual connection still suffers from class-imbalanced over-smoothing issue, due to the fixed and linear combination of topology and attribute in node representation learning. To make the combination flexible to capture complicated relationship, this paper reveals that the residual connection needs to be node-dependent, layer-dependent, and related to both topology and attribute. To alleviate the difficulty in specifying complicated relationship, this paper presents a novel perspective on GNNs, i.e., the representations of one node in different layers can be seen as a sequence of states. From this perspective, existing residual connections are not flexible enough for sequence modeling. Therefore, a novel node-dependent residual connection, i.e., Long Short-Term Graph Memory Network (LSTGM) is proposed to employ Long Short-Term Memory (LSTM), to model the sequence of node representation. To make the graph topology fully employed, LSTGM innovatively enhances the updated memory and three gates with graph topology. A speedup version is also proposed for effective training. Experimental evaluations on real-world datasets demonstrate their effectiveness in preventing over-smoothing issue and handling networks with heterophily.
Liang Yang 0002, Dongxiao He, Chuan Wang 0002, Yuanfang Guo, Xiaochun Cao, Bingxin Niu, Zhen Wang 0004
ACM Multimedia6
2023 Propagation is All You Need: A New Framework for Representation Learning and Classifier Training on Graphs
abstract
Graph Neural Networks (GNNs) have been the standard toolkit for processing non-euclidean spatial data since their powerful capability in graph representation learning. Unfortunately, their training strategy for network parameters is inefficient since it is directly inherited from classic Neural Networks (NNs), ignoring the characteristic of GNNs. To alleviate this issue, experimental analyses are performed to investigate the knowledge captured in classifier parameters during network training. We conclude that the parameter features, i.e., the column vectors of the classifier parameter matrix, are cluster representations with high discriminability. And after a theoretical analysis, we conclude that the discriminability of these features is obtained from the feature propagation from nodes to parameters. Furthermore, an experiment verifies that compared with cluster centroids, the parameter features are more potential for augmenting the feature propagation between nodes. Accordingly, a novel GNN-specific training framework is proposed by simultaneously updating node representations and classifier parameters via a unified feature propagation scheme. Moreover, two augmentation schemes are implemented for the framework, named Full Propagation Augmentation (FPA) and Simplified Full Propagation Augmentation (SFPA). Specifically, FPA augmentates the feature propagation of each node with the updated classifier parameters. SFPA only augments nodes with the classifier parameters corresponding to their clusters. Theoretically, FPA is equivalent to optimizing a novel graph learning objective, which demonstrates the universality of the proposed framework to existing GNNs. Extensive experiments demonstrate the superior performance and the universality of the proposed framework.
Jiaming Zhuo, Can Cui 0005, Bingxin Niu, Dongxiao He, Yuanfang Guo, Zhen Wang 0004, Chuan Wang 0002, Xiaochun Cao, Liang Yang 0002
ACM Multimedia6
2023 Deepfake Detection via Fine-Grained Classification and Global-Local Information Fusion
Tonghui Li, Yuanfang Guo, Yunhong Wang 0001
PRCV (6)2
2023 Graph Neural Networks without Propagation
abstract
Due to the simplicity, intuition and explanation, most Graph Neural Networks (GNNs) are proposed by following the pipeline of message passing. Although they achieve superior performances in many tasks, propagation-based GNNs possess three essential drawbacks. Firstly, the propagation tends to produce smooth effect, which meets the inductive bias of homophily, and causes two serious issues: over-smoothing issue and performance drop on networks with heterophily. Secondly, the propagations to each node are irrelevant, which prevents GNNs from modeling high-order relation, and cause the GNNs fragile to the attributes noises. Thirdly, propagation-based GNNs may be fragile to topology noise, since they heavily relay on propagation over the topology. Therefore, the propagation, as the key component of most GNNs, may be the essence of some serious issues in GNNs. To get to the root of these issue, this paper attempts to replace the propagation with a novel local operation. Quantitative experimental analysis reveals: 1) the existence of low-rank characteristic in the node attributes from ego-networks and 2) the performance improvement by reducing its rank. Motivated by this finding, this paper propose the Low-Rank GNNs, whose key component is the low-rank attribute matrix approximation in ego-network. The graph topology is employed to construct the ego-networks instead of message propagation, which is sensitive to topology noises. The proposed Low-Rank GNNs posses some attractive characteristics, including robust to topology and attribute noises, parameter-free and parallelizable. Experimental evaluations demonstrate the superior performance, robustness to noises and universality of the proposed Low-Rank GNNs.
Liang Yang 0002, Qiuliang Zhang, Runjie Shi, Wenmiao Zhou, Bingxin Niu, Chuan Wang 0002, Xiaochun Cao, Dongxiao He, Zhen Wang 0004, Yuanfang Guo
WWW10
2023 Enabling Homogeneous GNNs to Handle Heterogeneous Graphs via Relation Embedding
abstract
Graph Neural Networks (GNNs) have been generalized to process the heterogeneous graphs by various approaches. Unfortunately, these approaches usually model the heterogeneity via various complicated modules. This article aims to propose a simple yet effective framework to assign adequate ability to the homogeneous GNNs to handle the heterogeneous graphs. Specifically, we propose Relation Embedding based Graph Neural Network (RE-GNN), which employs only one parameter per relation to embed the importance of distinct types of relations and node-type-specific self-loop connections. To optimize these relation embeddings and the model parameters simultaneously, a gradient scaling factor is proposed to constrain the embeddings to converge to suitable values. Besides, we interpret the proposed RE-GNN from two perspectives, and theoretically demonstrate that our RE-GCN possesses more expressive power than GTN (which is a typical heterogeneous GNN, and it can generate meta-paths adaptively). Extensive experiments demonstrate that our RE-GNN can effectively and efficiently handle the heterogeneous graphs and can be applied to various homogeneous GNNs.
Junfu Wang, Yuanfang Guo, Liang Yang 0002, Yunhong Wang 0001
IEEE Trans. Big Data2
2022 Self-Supervised Graph Neural Networks via Diverse and Interactive Message Passing
abstract
By interpreting Graph Neural Networks (GNNs) as the message passing from the spatial perspective, their success is attributed to Laplacian smoothing. However, it also leads to serious over-smoothing issue by stacking many layers. Recently, many efforts have been paid to overcome this issue in semi-supervised learning. Unfortunately, it is more serious in unsupervised node representation learning task due to the lack of supervision information. Thus, most of the unsupervised or self-supervised GNNs often employ \textit{one-layer GCN} as the encoder. Essentially, the over-smoothing issue is caused by the over-simplification of the existing message passing, which possesses two intrinsic limits: blind message and uniform passing. In this paper, a novel Diverse and Interactive Message Passing (DIMP) is proposed for self-supervised learning by overcoming these limits. Firstly, to prevent the message from blindness and make it interactive between two connected nodes, the message is determined by both the two connected nodes instead of the attributes of one node. Secondly, to prevent the passing from uniformness and make it diverse over different attribute channels, different propagation weights are assigned to different elements in the message. To this end, a natural implementation of the message in DIMP is the element-wise product of the representations of two connected nodes. From the perspective of numerical optimization, the proposed DIMP is equivalent to performing an overlapping community detection via expectation-maximization (EM). Both the objective function of the community detection and the convergence of EM algorithm guarantee that DMIP can prevent from over-smoothing issue. Extensive evaluations on node-level and graph-level tasks demonstrate the superiority of DIMP on improving performance and overcoming over-smoothing issue.
Liang Yang 0002, Weixun Li, Bingxin Niu, Junhua Gu, Chuan Wang 0002, Dongxiao He, Yuanfang Guo, Xiaochun Cao
AAAI8
2022 Exploring the Impact of Adding Adversarial Perturbation onto Different Image Regions
abstract
Adversarial attack has been a hot topic for a long time in machine learning and deep learning. Studying adversarial attack has vital significance to artificial intelligence security. Existing methods mainly pursue a higher attack success rate. Few researches pay attention to the region where adversarial perturbations are added. Actually, different pixels in an image usually have different contributions in results, which motivates us to apply region constraint in the image for adversarial perturbations generation. In this paper, we present an easy-to-implement way to decrease the unnecessary adversarial perturbations while preserving a relatively high attack success rate. Specifically, we do not use the same constraint of perturbations in the input image but set specific constraint for specific region. Furthermore, we point that adversarial examples work in a different way to normal images. Directly using the activated region in normal images is not optimal. Then, to get the crucial area in adversarial attacks, we propose six transformation schemes to revise the activated region which is generated by the normal image. We launch extensive experiments on ImageNet dataset and the results show that our methods can get better attack strength under the same perturbation level when compared to the baseline methods.
Ruijie Yang, Yuanfang Guo, Ruikui Wang, Xiaohan Zhao, Yunhong Wang 0001
ISCAS2
2022 Difference Residual Graph Neural Networks
abstract
Graph Neural Networks have been widely employed for multimodal fusion and embedding. To overcome over-smoothing issue, residual connections, which are designed for alleviating vanishing gradient problem in NNs, are adopted in Graph Neural Networks (GNNs) to incorporate local node information. However, these simple residual connections are ineffective on networks with heterophily, since the roles of both convolutional operations and residual connections in GNNs are significantly different from those in classic NNs. By considering the specific smoothing characteristic of graph convolutional operation, deep layers in GNNs are expected to focus on the data which can't be properly handled in shallow layers. To this end, a novel and universal Difference Residual Connections (DRC), which feed the difference of the output and input of previous layer as the input of the next layer, is proposed. Essentially, Difference Residual Connections is equivalent to inserting layers with opposite effect (e.g., sharpening) into the network to prevent the excessive effect (e.g., over-smoothing issue) induced by too many layers with the similar role (e.g., smoothing) in GNNs. From the perspective of optimization, DRC is the gradient descent method to minimize an objective function with both smoothing and sharpening terms. The analytic solution to this objective function is determined by both graph topology and node attributes, which theoretically proves that DRC can prevent over-smoothing issue. Extensive experiments demonstrate the superiority of DRC on real networks with both homophily and heterophily, and show that DRC can automatically determine the model depth and be adaptive to both shallow and deep models with two complementary components.
Liang Yang 0002, Wenmiao Zhou, Bingxin Niu, Junhua Gu, Chuan Wang 0002, Yuanfang Guo, Dongxiao He, Xiaochun Cao
ACM Multimedia7
2022 OPEN: Orthogonal Propagation with Ego-Network Modeling
abstract
To alleviate the unfavorable effect of noisy topology in Graph Neural networks (GNNs), some efforts perform the local topology refinement through the pairwise propagation weight learning and the multi-channel extension. Unfortunately, most of them suffer a common and fatal drawback: irrelevant propagation to one node and in multi-channels. These two kinds of irrelevances make propagation weights in multi-channels free to be determined by the labeled data, and thus the GNNs are exposed to overfitting. To tackle this issue, a novel Orthogonal Propagation with Ego-Network modeling (OPEN) is proposed by modeling relevances between propagations. Specifically, the relevance between propagations to one node is modeled by whole ego-network modeling, while the relevance between propagations in multi-channels is modeled via diversity requirement. By interpreting the propagations to one node from the perspective of dimension reduction, propagation weights are inferred from principal components of the ego-network, which are orthogonal to each other. Theoretical analysis and experimental evaluations reveal four attractive characteristics of OPEN as modeling high-order relationships beyond pairwise one, preventing overfitting, robustness, and high efficiency.
Liang Yang 0002, Lina Kang, Qiuliang Zhang, Mengzhe Li, Bingxin Niu, Dongxiao He, Zhen Wang 0004, Chuan Wang 0002, Xiaochun Cao, Yuanfang Guo
NeurIPS10
2022 JoinTW: A Joint Image-to-Image Translation and Watermarking Method
Xiaohan Zhao, Yunhong Wang 0001, Ruijie Yang, Yuanfang Guo
PRCV (3)4
2022 Probabilistic Graph Convolutional Network via Topology-Constrained Latent Space Model
abstract
Although many graph convolutional neural networks (GCNNs) have achieved superior performances in semisupervised node classification, they are designed from either the spatial or spectral perspective, yet without a general theoretical basis. Besides, most of the existing GCNNs methods tend to ignore the ubiquitous noises in the network topology and node content and are thus unable to model these uncertainties. These drawbacks certainly reduce their effectiveness in integrating network topology and node content. To provide a probabilistic perspective to the GCNNs, we model the semisupervised node classification problem as a topology-constrained probabilistic latent space model, probabilistic graph convolutional network (PGCN). By representing the nodes in a more efficient distribution form, the proposed framework can seamlessly integrate the node content and network topology. When specifying the distribution in PGCN to be a Gaussian distribution, the transductive node classification problems can be solved by the general framework and a specific method, called PGCN with the Gaussian distribution representation (PGCN-G), is proposed. To overcome the overfitting problem in covariance estimation and reduce the computational complexity, PGCN-G is further improved to PGCN-G+ by imposing the covariance matrices of all vertices to possess the identical singular vectors. The optimization algorithm based on expectation-maximization indicates that the proposed method can iteratively denoise the network topology and node content with respect to each other. Besides the effectiveness of this top-down framework demonstrated via extensive experiments, it can also be deduced to cover the existing methods, graph convolutional network, graph attention network, and Gaussian mixture model and elaborate their characteristics and relationships by specific derivations.
Liang Yang 0002, Yuanfang Guo, Junhua Gu, Di Jin 0001, Bo Yang 0002, Xiaochun Cao
IEEE Trans. Cybern.2
2021 Bi-GCN: Binary Graph Convolutional Network
abstract
Graph Neural Networks (GNNs) have achieved tremendous success in graph representation learning. Unfortunately, current GNNs usually rely on loading the entire attributed graph into network for processing. This implicit assumption may not be satisfied with limited memory resources, especially when the attributed graph is large. In this paper, we pioneer to propose a Binary Graph Convolutional Network (Bi-GCN), which binarizes both the network parameters and input node features. Besides, the original matrix multiplications are revised to binary operations for accelerations. According to the theoretical analysis, our Bi-GCN can reduce the memory consumption by an average of ~30x for both the network parameters and input data, and accelerate the inference speed by an average of ~47x, on the citation networks. Meanwhile, we also design a new gradient approximation based back-propagation method to train our Bi-GCN well. Extensive experiments have demonstrated that our Bi-GCN can give a comparable performance compared to the full-precision baselines. Besides, our binarization approach can be easily applied to other GNNs, which has been verified in the experiments.
Junfu Wang, Yunhong Wang 0001, Zhen Yang 0037, Liang Yang 0002, Yuanfang Guo
CVPR5
2021 Multi-Scale Background Suppression Anomaly Detection In Surveillance Videos
abstract
Video anomaly detection has been widely applied in various surveillance systems for public security. However, the existing weakly supervised video anomaly detection methods tend to ignore the interference of the background frames and possess limited ability to extract effective temporal information among the video snippets. In this paper, a multi-scale background suppression based anomaly detection (MSBSAD) method is proposed to suppress the interference of the background frames. We propose a multi-scale temporal convolution module to effectively extract more temporal information among the video snippets for the anomaly events with different durations. A modified hinge loss is constructed in the suppression branch to help our model to better differentiate the abnormal samples from the confusing samples. Experiments on UCF Crime demonstrate the superiority of our MS-BSAD method in the video anomaly detection task.
Yuanfang Guo, Jinjie Wei, Xiuguo Bao, Di Huang 0001
ICIP2
2021 Heterogeneous Graph Information Bottleneck
abstract
Most attempts on extending Graph Neural Networks (GNNs) to Heterogeneous Information Networks (HINs) implicitly take the direct assumption that the multiple homogeneous attributed networks induced by different meta-paths are complementary. The doubts about the hypothesis of complementary motivate an alternative assumption of consensus. That is, the aggregated node attributes shared by multiple homogeneous attributed networks are essential for node representations, while the specific ones in each homogeneous attributed network should be discarded. In this paper, a novel Heterogeneous Graph Information Bottleneck (HGIB) is proposed to implement the consensus hypothesis in an unsupervised manner. To this end, information bottleneck (IB) is extended to unsupervised representation learning by leveraging self-supervision strategy. Specifically, HGIB simultaneously maximizes the mutual information between one homogeneous network and the representation learned from another homogeneous network, while minimizes the mutual information between the specific information contained in one homogeneous network and the representation learned from this homogeneous network. Model analysis reveals that the two extreme cases of HGIB correspond to the supervised heterogeneous GNN and the infomax on homogeneous graph, respectively. Extensive experiments on real datasets demonstrate that the consensus-based unsupervised HGIB significantly outperforms most semi-supervised SOTA methods based on complementary assumption.
Liang Yang 0002, Zichen Zheng, Bingxin Niu, Junhua Gu, Chuan Wang 0002, Xiaochun Cao, Yuanfang Guo
IJCAI8
2021 Diverse Message Passing for Attribute with Heterophily
abstract
Most of the existing GNNs can be modeled via the Uniform Message Passing framework. This framework considers all the attributes of each node in its entirety, shares the uniform propagation weights along each edge, and focuses on the uniform weight learning. The design of this framework possesses two prerequisites, the simplification of homophily and heterophily to the node-level property and the ignorance of attribute differences. Unfortunately, different attributes possess diverse characteristics. In this paper, the network homophily rate defined with respect to the node labels is extended to attribute homophily rate by taking the attributes as weak labels. Based on this attribute homophily rate, we propose a Diverse Message Passing (DMP) framework, which specifies every attribute propagation weight on each edge. Besides, we propose two specific strategies to significantly reduce the computational complexity of DMP to prevent the overfitting issue. By investigating the spectral characteristics, existing spectral GNNs are actually equivalent to a degenerated version of DMP. From the perspective of numerical optimization, we provide a theoretical analysis to demonstrate DMP's powerful representation ability and the ability of alleviating the over-smoothing issue. Evaluations on various real networks demonstrate the superiority of our DMP on handling the networks with heterophily and alleviating the over-smoothing issue, compared to the existing state-of-the-arts.
Liang Yang 0002, Mengzhe Li, Liyang Liu, Bingxin Niu, Chuan Wang 0002, Xiaochun Cao, Yuanfang Guo
NeurIPS7
2021 Graph-CAT: Graph Co-Attention Networks via local and global attribute augmentations
Liang Yang 0002, Weixun Li, Yuanfang Guo, Junhua Gu
Future Gener. Comput. Syst.3
2021 Hierarchical Reasoning Network for Pedestrian Attribute Recognition
abstract
Pedestrian attribute recognition, which can benefit other tasks such as person re-identification and pedestrian retrieval, is very important in video surveillance related tasks. In this paper, we observe that the existing methods tackle this problem from the perspective of multi-label classification without considering the hierarchical relationships among the attributes. In human cognition, the attributes can be categorized according to their semantic/abstraction levels. The high-level attributes can be predicted by reasoning from the low-level and medium-level attributes, while the recognition of the low-level and medium-level attributes can be guided by the high-level attributes. Based on this attribute categorization, we propose a novel Hierarchical Reasoning Network (HR-Net), which can hierarchically predict the attributes at different abstraction levels in different stages of the network. We also propose an attribute reasoning structure to exploit the relationships among the attributes at different semantic levels. Experimental results demonstrate that the proposed network gives superior performances compared to the state-of-the-art techniques.
Haoran An, Hai-Miao Hu, Yuanfang Guo, Qianli Zhou, Bo Li 0006
IEEE Trans. Multim.3
2021 Semantic Correspondence with Geometric Structure Analysis
abstract
This article studies the correspondence problem for semantically similar images, which is challenging due to the joint visual and geometric deformations. We introduce the Flip-aware Distance Ratio method (FDR) to solve this problem from the perspective of geometric structure analysis. First, a distance ratio constraint is introduced to enforce the geometric consistencies between images with large visual variations, whereas local geometric jitters are tolerated via a smoothness term. For challenging cases with symmetric structures, our proposed method exploits Curl to suppress the mismatches. Subsequently, image correspondence is formulated as a permutation problem, for which we propose a Gradient Guided Simulated Annealing (GGSA) algorithm to perform a robust discrete optimization. Experiments on simulated and real-world datasets, where both visual and geometric deformations are present, indicate that our method significantly improves the baselines for both visually and semantically similar images.
Rui Wang 0032, Xiaochun Cao, Yuanfang Guo
ACM Trans. Multim. Comput. Commun. Appl.4
2020 Distraction-Aware Feature Learning for Human Attribute Recognition via Coarse-to-Fine Attention Mechanism
abstract
Recently, Human Attribute Recognition (HAR) has become a hot topic due to its scientific challenges and application potentials, where localizing attributes is a crucial stage but not well handled. In this paper, we propose a novel deep learning approach to HAR, namely Distraction-aware HAR (Da-HAR). It enhances deep CNN feature learning by improving attribute localization through a coarse-to-fine attention mechanism. At the coarse step, a self-mask block is built to roughly discriminate and reduce distractions, while at the fine step, a masked attention branch is applied to further eliminate irrelevant regions. Thanks to this mechanism, feature learning is more accurate, especially when heavy occlusions and complex backgrounds exist. Extensive experiments are conducted on the WIDER-Attribute and RAP databases, and state-of-the-art results are achieved, demonstrating the effectiveness of the proposed approach.
Mingda Wu, Di Huang 0001, Yuanfang Guo, Yunhong Wang 0001
AAAI3
2020 Toward Unsupervised Graph Neural Network: Interactive Clustering and Embedding via Optimal Transport
abstract
Most of the existing Graph Neural Networks (GNNs) are deliberately designed for semi-supervised learning tasks, where supervision information (labelled node) is utilized to mitigate the oversmoothing problem of message passing. Unfortunately, the oversmoothing problem tends to be more severe in unsupervised tasks, since supervision information is not available. Since community structure/cluster is an essential characteristic of network, a natural approach to reduce the oversmoothing problem is to also constrain the node embeddings to maintain their own characteristics to prevent all the node embeddings from becoming too similar to be distinguished. In this paper, a novel Optimal Transport based Graph Neural Network (OT-GNN) is proposed to overcome the oversmoothing problem in unsupervised GNNs by imposing the equal-sized clustering constraints to the obtained node embeddings. To solve the combinatorial optimization problem, the constrained objective function of unsupervised GNN is relaxed to an Optimal Transport problem, and a fast version of the Sinkhorm-Knopp algorithm is adopted to handle large networks. Extensive experiments on node clustering and classification demonstrate the superior performance of our proposed OT-GNN.
Liang Yang 0002, Junhua Gu, Chuan Wang 0002, Xiaochun Cao, Lu Zhai, Di Jin 0001, Yuanfang Guo
ICDM7
2020 Fake Generated Painting Detection Via Frequency Analysis
abstract
With the development of deep neural networks, digital fake paintings can be generated by various style transfer algorithms. To detect the fake generated paintings, we analyze the fake generated and real paintings in Fourier frequency domain and observe statistical differences and artifacts. Based on our observations, we propose Fake Generated Painting Detection via Frequency Analysis (FGPD-FA) by extracting three types of features in frequency domain. Besides, we also propose a digital fake painting detection database for assessing the proposed method. Experimental results demonstrate the excellence of the proposed method in different testing conditions.
Yuanfang Guo, Jinjie Wei, Rui Wang 0032, Yunhong Wang 0001
ICIP2
2020 JANE: Jointly Adversarial Network Embedding
abstract
Motivated by the capability of Generative Adversarial Network on exploring the latent semantic space and capturing semantic variations in the data distribution, adversarial learning has been adopted in network embedding to improve the robustness. However, this important ability is lost in existing adversarially regularized network embedding methods, because their embedding results are directly compared to the samples drawn from perturbation (Gaussian) distribution without any rectification from real data. To overcome this vital issue, a novel Joint Adversarial Network Embedding (JANE) framework is proposed to jointly distinguish the real and fake combinations of the embeddings, topology information and node features. JANE contains three pluggable components, Embedding module, Generator module and Discriminator module. The overall objective function of JANE is defined in a min-max form, which can be optimized via alternating stochastic gradient. Extensive experiments demonstrate the remarkable superiority of the proposed JANE on link prediction (3% gains in both AUC and AP) and node clustering (5% gain in F1 score).
Liang Yang 0002, Yuexue Wang, Junhua Gu, Chuan Wang 0002, Xiaochun Cao, Yuanfang Guo
IJCAI6
2020 Graph Attention Topic Modeling Network
abstract
Existing topic modeling approaches possess several issues, including the overfitting issue of Probablistic Latent Semantic Indexing (pLSI), the failure of capturing the rich topical correlations among topics in Latent Dirichlet Allocation (LDA), and high inference complexity. In this paper, we provide a new method to overcome the overfitting issue of pLSI by using the amortized inference with word embedding as input, instead of the Dirichlet prior in LDA. For generative topic model, the large number of free latent variables is the root of overfitting. To reduce the number of parameters, the amortized inference replaces the inference of latent variable with a function which possesses the shared (amortized) learnable parameters. The number of the shared parameters is fixed and independent of the scale of the corpus. To overcome the limited application of amortized inference to independent and identically distributed (i.i.d) data, a novel graph neural network, Graph Attention TOpic Network (GATON), is proposed to model the topic structure of non-i.i.d documents according to the following two observations. First, pLSI can be interpreted as stochastic block model (SBM) on a specific bi-partite graph. Second, graph attention network (GAT) can be explained as the semi-amortized inference of SBM, which relaxes the i.i.d data assumption of vanilla amortized inference. GATON provides a novel scheme, i.e. graph convolution operation based scheme, to integrate word similarity and word co-occurrence structure. Specifically, the bag-of-words document representation is modeled as a bi-partite graph topology. Meanwhile, word embedding, which captures the word similarity, is modeled as attribute of the word node and the term frequency vector is adopted as the attribute of the document node. Based on the weighted (attention) graph convolution operation, the word co-occurrence structure and word similarity patterns are seamlessly integrated for topic identification. Extensive experiments demonstrate that the effectiveness of GATON on topic identification not only benefits the document classification, but also significantly refines the input word embedding.
Liang Yang 0002, Junhua Gu, Chuan Wang 0002, Xiaochun Cao, Di Jin 0001, Yuanfang Guo
WWW7
2020 A New Polyphase Down-Sampling-Based Multiple Description Image Coding
abstract
Multiple description coding (MDC) is an efficient source coding technique for error-prone transmission over multiple channels. In this paper, we focus on the design of a new polyphase down-sampling based MDC (NPDS-MDC) for image signals. The encoding of our proposed NPDS-MDC consists of three steps. First, we perform down-sampling on each N×N image block according to the quincunx down-sampling pattern. Second, we propose a new transform and apply it to the down-sampled pixels to produce the side descriptions. Third, we develop an error compensation algorithm to reduce the compression distortion occurring on the down-sampled pixels. In our scheme, the side decoding is performed posterior to image interpolation with reference to the down-sampled compressed pixels. Moreover, the central decoding is achieved by interlacing the side descriptions. We also propose a compression-constrained central deblocking algorithm to further improve the efficiency of the central decoding. The experimental results indicate that our proposed MDC scheme offers clearly superior performance, especially at high bit rates, as compared to the state-of-the-art methods for various types of images.
Shuyuan Zhu, Zhiying He, Xiandong Meng, Jiantao Zhou 0001, Yuanfang Guo, Bing Zeng 0001
IEEE Trans. Image Process.5
2019 Scene Graph Generation via Convolutional Message Passing and Class-Aware Memory Embeddings
Yunhong Wang 0001, Yuanfang Guo
ICANN (3)3
2019 Dual Self-Paced Graph Convolutional Network: Towards Reducing Attribute Distortions Induced by Topology
abstract
The success of graph convolutional neural networks (GCNNs) based semi-supervised node classification is credited to the attribute smoothing (propagating) over the topology. However, the attributes may be interfered by the utilization of the topology information. This distortion will induce a certain amount of misclassifications of the nodes, which can be correctly predicted with only the attributes. By analyzing the impact of the edges in attribute propagations, the simple edges, which connect two nodes with similar attributes, should be given priority during the training process compared to the complex ones according to curriculum learning. To reduce the distortions induced by the topology while exploit more potentials of the attribute information, Dual Self-Paced Graph Convolutional Network (DSP-GCN) is proposed in this paper. Specifically, the unlabelled nodes with confidently predicted labels are gradually added into the training set in the node-level self-paced learning, while edges are gradually, from the simple edges to the complex ones, added into the graph during the training process in the edge-level self-paced learning. These two learning strategies are designed to mutually reinforce each other by coupling the selections of the edges and unlabelled nodes. Experimental results of transductive semi-supervised node classification on many real networks indicate that the proposed DSP-GCN has successfully reduced the attribute distortions induced by the topology while it gives superior performances with only one graph convolutional layer.
Liang Yang 0002, Junhua Gu, Yuanfang Guo
IJCAI4
2019 Topology Optimization based Graph Convolutional Network
abstract
In the past few years, semi-supervised node classification in attributed network has been developed rapidly. Inspired by the success of deep learning, researchers adopt the convolutional neural network to develop the Graph Convolutional Networks (GCN), and they have achieved surprising classification accuracy by considering the topological information and employing the fully connected network (FCN). However, the given network topology may also induce a performance degradation if it is directly employed in classification, because it may possess high sparsity and certain noises. Besides, the lack of learnable filters in GCN also limits the performance. In this paper, we propose a novel Topology Optimization based Graph Convolutional Networks (TO-GCN) to fully utilize the potential information by jointly refining the network topology and learning the parameters of the FCN. According to our derivations, TO-GCN is more flexible than GCN, in which the filters are fixed and only the classifier can be updated during the learning process. Extensive experiments on real attributed networks demonstrate the superiority of the proposed TO-GCN against the state-of-the-art approaches.
Liang Yang 0002, Zesheng Kang, Xiaochun Cao, Di Jin 0001, Bo Yang 0002, Yuanfang Guo
IJCAI6
2019 Masked Graph Convolutional Network
abstract
Semi-supervised classification is a fundamental technology to process the structured and unstructured data in machine learning field. The traditional attribute-graph based semi-supervised classification methods propagate labels over the graph which is usually constructed from the data features, while the graph convolutional neural networks smooth the node attributes, i.e., propagate the attributes, over the real graph topology. In this paper, they are interpreted from the perspective of propagation, and accordingly categorized into symmetric and asymmetric propagation based methods. From the perspective of propagation, both the traditional and network based methods are propagating certain objects over the graph. However, different from the label propagation, the intuition ``the connected data samples tend to be similar in terms of the attributes", in attribute propagation is only partially valid. Therefore, a masked graph convolution network (Masked GCN) is proposed by only propagating a certain portion of the attributes to the neighbours according to a masking indicator, which is learned for each node by jointly considering the attribute distributions in local neighbourhoods and the impact on the classification results. Extensive experiments on transductive and inductive node classification tasks have demonstrated the superiority of the proposed method.
Liang Yang 0002, Yingkui Wang, Junhua Gu, Yuanfang Guo
IJCAI5
2019 Pedestrian Attribute Recognition via Hierarchical Multi-task Learning and Relationship Attention
abstract
Pedestrian Attribute Recognition (PAR) is an important task in surveillance video analysis. In this paper, we propose a novel end-to-end hierarchical deep learning approach to PAR. The proposed network introduces semantic segmentation into PAR and formulates it as a multi-task learning problem, which brings in pixel-level supervision in feature learning for attribute localization. According to the spatial properties of local and global attributes, we present a two stage learning mechanism to decouple coarse attribute localization and fine attribute recognition into successive phases within a single model, which strengthens feature learning. Besides, we design an attribute relationship attention module to efficiently capture and emphasize the latent relations among different attributes, further enhancing the discriminative power of the feature. Extensive experiments are conducted and very competitive results are reached on the RAP and PETA databases, indicating the effectiveness and superiority of the proposed approach.
Lian Gao, Di Huang 0001, Yuanfang Guo, Yunhong Wang 0001
ACM Multimedia3
2019 Color Image Compression with Transform Domain Down-Sampling and Deep Convolutional Reconstruction
abstract
In this paper, we build up a new block-based color image compression scheme based on our proposed transform domain down-sampling method and deep convolutional reconstruction algorithm. Specifically, our proposed down-sampling scheme aims to down-sample each N × N transform block into the N/2 × N/2 block for the saving of bit-cost. On the other hand, the proposed deep convolutional reconstruction algorithm is employed to reconstruct the down-sampled block for a full- resolution reconstruction. We apply our proposed methods to both the chrominance components to compress color images. Experimental results show that our proposed method achieves excellent results when used in practice.
Shuyuan Zhu, Xiandong Meng, Bing Zeng 0001, Yuanfang Guo, Ruiqin Xiong
VCIP5
2019 Efficient Chroma Sub-Sampling and Luma Modification for Color Image Compression
abstract
In color image compression, the chroma components are often sub-sampled before compression and up-sampled after compression. Although sub-sampling the chroma components saves the bit-cost for compression, it often induces extra color distortions in the compressed images. In this paper, we propose two approaches to tackle this problem. First, we propose a sub-sampling method in the transform domain and apply it to both chroma components. Then, based on this sub-sampling, we propose a novel method to modify the luma component. In our proposed luma modification algorithm, the distortions that occurred in the two chroma components can be coupled together and utilized to modify the luma component. With our proposed chroma sub-sampling and luma modification algorithms, we can achieve a low RGB distortion in practical image coding. The experimental results demonstrate that our proposed methods offer more significant coding gains compared with the state-of-the-art methods for the compression of color images.
Shuyuan Zhu, Chang Cui, Ruiqin Xiong, Yuanfang Guo, Bing Zeng 0001
IEEE Trans. Circuits Syst. Video Technol.4
2019 High-Quality Color Image Compression by Quantization Crossing Color Spaces
abstract
Coding of a color image usually happens in the YCbCr space so that the rate-distortion optimization is conducted in this space. Due to the use of a non-unitary matrix in the RGB-to-YCbCr conversion, an optimal coding performance achieved in the YCbCr space does not guarantee an optimal quality in the RGB space, which would impact most display devices that need RGB signals as the inputs. In this paper, we first study the relationship between the coding distortions of the compressed RGB signals and the quantization errors occurred in the coded YCbCr signals. Then, we design a new quantization scheme crossing the RGB and YCbCr spaces to achieve a high-quality color image compression with the YCbCr 4:4:4 format. Although our proposed quantization takes place in the YCbCr space, it aims at reducing the coding distortion in the RGB space as much as possible. Experimental results demonstrate that our proposed method offers a significant quality gain over the existing block-based coding methods for various images.
Shuyuan Zhu, Zhiying He, Chen Chen 0015, Shuaicheng Liu, Jiantao Zhou 0001, Yuanfang Guo, Bing Zeng 0001
IEEE Trans. Circuits Syst. Video Technol.6
2018 Multi-Facet Network Embedding: Beyond the General Solution of Detection and Representation
abstract
In network analysis, community detection and network embedding are two important topics. Community detection tends to obtain the most noticeable partition, while network embedding aims at seeking node representations which contains as many diverse properties as possible. We observe that the current community detection and network embedding problems are being resolved by a general solution, i.e., "maximizing the consistency between similar nodes while maximizing the distance between the dissimilar nodes." This general solution only exploits the most noticeable structure (facet) of the network, which effectively satisfies the demands of the community detection. Unfortunately, most of the specific embedding algorithms, which are developed from the general solution, cannot achieve the goal of network embedding by exploring only one facet of the network. To improve the general solution for better modeling the real network, we propose a novel network embedding method, Multi-facet Network Embedding (MNE), to capture the multiple facets of the network. MNE learns multiple embeddings simultaneously, with the Hilbert Schmidt Independence Criterion (HSIC) being the a diversity constraint. To efficiently solve the optimization problem, we propose a Binary HSIC with linear complexity and solve the MNE objective function by adopting the Augmented Lagrange Multiplier (ALM) method. The overall complexity is linear with the scale of the network. Extensive results demonstrate that MNE gives efficient performances and outperforms the state-of-the-art network embedding methods.
Liang Yang 0002, Yuanfang Guo, Xiaochun Cao
AAAI2
2018 3-in-1 Correlated Embedding via Adaptive Exploration of the Structure and Semantic Subspaces
abstract
Combinational network embedding, which learns the node representation by exploring both topological and non-topological information, becomes popular due to the fact that the two types of information are complementing each other. Most of the existing methods either consider the topological and non-topological information being aligned or possess predetermined preferences during the embedding process.Unfortunately, previous methods fail to either explicitly describe the correlations between topological and non-topological information or adaptively weight their impacts. To address the existing issues, three new assumptions are proposed to better describe the embedding space and its properties. With the proposed assumptions, nodes, communities and topics are mapped into one embedding space. A novel generative model is proposed to formulate the generation process of the network and content from the embeddings, with respect to the Bayesian framework. The proposed model automatically leans to the information which is more discriminative.The embedding result can be obtained by maximizing the posterior distribution by adopting the variational inference and reparameterization trick. Experimental results indicate that the proposed method gives superior performances compared to the state-of-the-art methods when a variety of real-world networks is analyzed.
Liang Yang 0002, Yuanfang Guo, Di Jin 0001, Huazhu Fu, Xiaochun Cao
IJCAI2
2018 iHuman3D: Intelligent Human Body 3D Reconstruction using a Single Flying Camera
abstract
Aiming at autonomous, adaptive and real-time human body reconstruction technique, this paper presents iHuman3D: an intelligent human body 3D reconstruction system using a single aerial robot integrated with an RGB-D camera. Specifically, we propose a real-time and active view planning strategy based on a highly efficient ray casting algorithm in GPU and a novel information gain formulation directly in TSDF. We also propose the human body reconstruction module by revising the traditional volumetric fusion pipeline with a compactly-designed non-rigid deformation for slight motion of the human target. We unify both the active view planning and human body reconstruction in the same TSDF volume-based representation. Quantitative and qualitative experiments are conducted to validate that the proposed iHuman3D system effectively removes the constraint of extra manual labor, enabling real-time and autonomous reconstruction of human body.
Lan Xu 0003, Yuanfang Guo, Lu Fang 0001
ACM Multimedia4
2018 Fake Colorized Image Detection
abstract
Image forensics aims to detect the manipulation of digital images. Currently, splicing detection, copy-move detection, and image retouching detection are attracting significant attention from researchers. However, image editing techniques develop over time. An emerging image editing technique is colorization, in which grayscale images are colorized with realistic colors. Unfortunately, this technique may also be intentionally applied to certain images to confound object recognition algorithms. To the best of our knowledge, no forensic technique has yet been invented to identify whether an image is colorized. We observed that, compared with natural images, colorized images, which are generated by three state-of-the-art methods, possess statistical differences for the hue and saturation channels. Besides, we also observe statistical inconsistencies in the dark and bright channels, because the colorization process will inevitably affect the dark and bright channel values. Based on our observations, i.e., potential traces in the hue, saturation, dark, and bright channels, we propose two simple yet effective detection methods for fake colorized images: Histogram-based fake colorized image detection and feature encoding-based fake colorized image detection. Experimental results demonstrate that both proposed methods exhibit a decent performance against multiple state-of-the-art colorization approaches.
Yuanfang Guo, Xiaochun Cao, Wei Zhang 0031, Rui Wang 0032
IEEE Trans. Inf. Forensics Secur.1
2018 Halftone Image Watermarking by Content Aware Double-Sided Embedding Error Diffusion
abstract
In this paper, we carry out a performance analysis from a probabilistic perspective to introduce the error diffusion-based halftone visual watermarking (EDHVW) methods' expected performances and limitations. Then, we propose a new general EDHVW method, content aware double-sided embedding error diffusion (CaDEED), via considering the expected watermark decoding performance with specific content of the cover images and watermark, different noise tolerance abilities of various cover image content, and the different importance levels of every pixel (when being perceived) in the secret pattern (watermark). To demonstrate the effectiveness of CaDEED, we propose CaDEED with expectation constraint (CaDEED-EC) and CaDEED-noise visibility function (NVF) and importance factor (IF) (CaDEED-N&I). Specifically, we build CaDEED-EC by only considering the expected performances of specific cover images and watermark. By adopting the NVF and proposing the IF to assign weights to every embedding location and watermark pixel, respectively, we build the specific method CaDEED-N&I. In the experiments, we select the optimal parameters for NVF and IF via extensive experiments. In both the numerical and visual comparisons, the experimental results demonstrate the superiority of our proposed work.
Yuanfang Guo, Oscar C. Au, Rui Wang 0032, Lu Fang 0001, Xiaochun Cao
IEEE Trans. Image Process.1
2017 Image Deblurring via Extreme Channels Prior
abstract
Camera motion introduces motion blur, affecting many computer vision tasks. Dark Channel Prior (DCP) helps the blind deblurring on scenes including natural, face, text, and low-illumination images. However, it has limitations and is less likely to support the kernel estimation while bright pixels dominate the input image. We observe that the bright pixels in the clear images are not likely to be bright after the blur process. Based on this observation, we first illustrate this phenomenon mathematically and define it as the Bright Channel Prior (BCP). Then, we propose a technique for deblurring such images which elevates the performance of existing motion deblurring algorithms. The proposed method takes advantage of both Bright and Dark Channel Prior. This joint prior is named as extreme channels prior and is crucial for achieving efficient restorations by leveraging both the bright and dark information. Extensive experimental results demonstrate that the proposed method is more robust and performs favorably against the state-of-the-art image deblurring methods on both synthesized and natural images.
Yanyang Yan, Wenqi Ren, Yuanfang Guo, Rui Wang 0032, Xiaochun Cao
CVPR3
2017 Binarized Mode Seeking for Scalable Visual Pattern Discovery
abstract
This paper studies visual pattern discovery in large-scale image collections via binarized mode seeking, where images can only be represented as binary codes for efficient storage and computation. We address this problem from the perspective of binary space mode seeking. First, a binary mean shift (bMS) is proposed to discover frequent patterns via mode seeking directly in binary space. The binomial-based kernel and binary constraint are introduced for binarized analysis. Second, we further extend bMS to a more general form, namely contrastive binary mean shift (cbMS), which maximizes the contrastive density in binary space, for finding informative patterns that are both frequent and discriminative for the dataset. With the binarized algorithm and optimization, our methods demonstrate significant computation (50×) and storage (32×) improvement compared to standard techniques operating in Euclidean space, while the performance does not largely degenerate. Furthermore, cbMS discovers more informative patterns by suppressing low discriminative modes. We evaluate our methods on both annotated ILSVRC (1M images) and un-annotated blind Flickr (10M images) datasets with million scale images, which demonstrates both the scalability and effectiveness of our algorithms for discovering frequent and informative patterns in large scale collection.
Wei Zhang 0031, Xiaochun Cao, Rui Wang 0032, Yuanfang Guo, Zhineng Chen
CVPR4
2017 Contextual approach for identifying malicious Inter-Component privacy leaks in Android apps
abstract
Inter-Component Communication (ICC) enables developers to create rich and innovative applications in Android platform. However, some privacy problems occur because of the interactions among multiple components. Since the flow of sensitive data across components may be legal or malicious, it is necessary to perform a precise ICC analysis to identify the malicious flow of sensitive data. In this paper, we propose a static taint analysis method, named IccChecker, to identify the malicious ICC-based privacy leaks in Android applications. IccChecker first tracks the potential flow of sensitive data across components and extracts the contextual factors which trigger the sensitive behavior. By leveraging the context information, our approach differentiates the malicious privacy leaks from the legal privacy information exchanges according to the proposed contextual policy. Moreover, we present a comprehensive assessment with benchmarks and real-world applications. Our evaluation results with benchmarks demonstrate that IccChecker improves the precision of ICC-based privacy leak detection. In the evaluation with real-world applications, our approach identifies 4 apps with ICC-based privacy leaks among 168 Google Play apps (2.3%) while 31 apps are identified from 49 malwares (63.3%).
Daojuan Zhang, Yuanfang Guo, Dianjie Guo, Rui Wang 0032, Guangming Yu
ISCC2
2017 egoPortray: Visual Exploration of Mobile Communication Signature from Egocentric Network Perspective
Qing Wang 0038, Jiansu Pu, Yuanfang Guo, Zheng Hu 0001, Hui Tian 0003
MMM (1)3
2016 An Adaptive Reversible Data Hiding Scheme for JPEG Images
Jiaxin Yin, Rui Wang 0032, Yuanfang Guo, Feng Liu 0001
IWDW3
2016 Halftone image watermarking via optimization
Yuanfang Guo, Oscar C. Au, Jiantao Zhou 0001, Ketan Tang, Xiaopeng Fan 0001
Signal Process. Image Commun.1
2016 Adaptive Block Coding Order for Intra Prediction in HEVC
abstract
In this paper, an adaptive block coding order for intra prediction is proposed. Modern video coding standards, including the most recent High Efficiency Video Coding (HEVC), utilize fixed scan orders in processing blocks during intra coding. However, the fixed scan orders typically result in residual blocks with noticeable edge patterns. That means the fixed scan orders cannot fully exploit the content-adaptive spatial correlations between adjacent blocks, thus the bitrate after compression tends to be large. To reduce the bitrate induced by inaccurate intra prediction, the proposed approach adaptively chooses both the block and subblock coding orders by minimizing the coding cost. Specifically, determining the block coding order is formulated as a traveling salesman problem that is solved using dynamic programming. Besides the block coding order, we also design the subblock coding order in each block with an adaptive manner. The experimental results demonstrate a Bjøntegaard-Delta-rate reduction of up to 4.4% compared with HEVC anchor.
Amin Zheng, Yuan Yuan 0002, Jiantao Zhou 0001, Yuanfang Guo, Haitao Yang 0001, Oscar C. Au
IEEE Trans. Circuits Syst. Video Technol.4
2016 Subpixel-Based Image Scaling for Grid-like Subpixel Arrangements: A Generalized Continuous-Domain Analysis Model
abstract
Subpixel-based image scaling can improve the apparent resolution of displayed images by controlling individual subpixels rather than whole pixels. However, improved luminance resolution brings chrominance distortion, making it crucial to suppress color error while maintaining sharpness. Moreover, it is challenging to develop a scheme that is applicable for various subpixel arrangements and for arbitrary scaling factors. In this paper, we address the aforementioned issues by proposing a generalized continuous-domain analysis model, which considers the low-pass nature of the human visual system (HVS). Specifically, given a discrete image and a grid-like subpixel arrangement, the signal perceived by the HVS is modeled as a 2D continuous image. Minimizing the difference between the perceived image and the continuous target image leads to the proposed scheme, which we call continuous-domain analysis for subpixel-based scaling (CASS). To eliminate the ringing artifacts caused by the ideal low-pass filtering in CASS, we propose an improved scheme, which we call CASS with Laplacian-of-Gaussian filtering. Experiments show that the proposed methods provide sharp images with negligible color fringing artifacts. Our methods are comparable with the state-of-the-art methods when applied on the RGB stripe arrangement, and outperform existing methods when applied on other subpixel arrangements.
Jiahao Pang, Lu Fang 0001, Jin Zeng 0004, Yuanfang Guo, Ketan Tang
IEEE Trans. Image Process.4
2014 Analysis of sampling pattern and Luma-Chroma filter design for subpixel-based image downsampling
abstract
Subpixel-based image downsampling is attractive in that it produces higher apparent resolution of down-sampled images on LCD displays. However increased luminance resolution is achieved at the price of color fringing artifacts. In this paper, we propose an algorithm to find a pleasing balance between increased resolution and color fidelity. We separate the subpixel-based downsampling into two stages, shifting followed by downsampling with anti-aliasing filtering. In stage one, we find special characteristics of the luminance and chrominance spectra of the shifted image, based on which the optimal sampling pattern is found. In stage two, anti-aliasing filters for luminance and chrominance are designed respectively. Experimental results verify that the proposed method manages to suppress color artifacts while maintaining high luminance sharpness.
Jin Zeng 0004, Oscar C. Au, Yuanfang Guo, Jiahao Pang, Ketan Tang, Yonggen Ling
ICASSP3
2014 Self-similarity-based image colorization
abstract
In this work, we tackle the problem of coloring black-and-white images, which is image colorization. Existing image colorization algorithms can be categorized into two types: scribble-based colorization algorithms and example-based colorization algorithms. Differently, we propose a hybrid scheme that combines the advantages of both categories. Given the grayscale image to be colorized and a few color scribbles (or scattered color labels) as input, the proposed method manages to colorize the grayscale image with high quality. Similar to the mechanisms in example-based colorization methods, our algorithm firstly propagates chrominance information based on the assumption that similar image patches should have similar colors. Therefore colors of some pixels can be transferred from similar patches with known colors. After that, we apply scribble-based colorization algorithm to fully colorize the grayscale image, with different confidences assigned onto the transferred color labels. Experimental results show that, the proposed method effectively utilizes the known chrominance, and provides pleasant colorizations with very few user interventions.
Jiahao Pang, Oscar C. Au, Yukihiko Yamashita, Yonggen Ling, Yuanfang Guo, Jin Zeng 0004
ICIP5
2014 Fast algorithm of arbitrary factor subpixel downsampling based on frequency analysis
abstract
Subpixel-based downsampling has shown its advantages over pixel-based downsampling in terms of preserving more spatial details along edges and generating sharper images, at the cost of certain amount of color-fringing artifacts in the downsampled image. To balance the sharpness and color-fringing artifacts, some algorithms are proposed to design optimal anti-aliasing (AA) filters, which are either image independent, or computationally too expensive. And all of the existing AA filters are designed for fixed downsampling factor, which makes them impractical for real applications. In this paper we propose two fast algorithms to design AA filter for arbitrary factor subpixel downsampling based on frequency analysis of the input image. The proposed algorithms generate image dependent AA filter which is as good as the state-of-the-art algorithm, but much faster.
Ketan Tang, Oscar C. Au, Lu Fang 0001, Jiahao Pang, Yuanfang Guo
ICME5
2014 Adaptive Predictor Structure Based Interpolation for Reversible Data Hiding
Sunil Prasad Jaiswal, Oscar C. Au, Vinit Jakhetiya, Yuanfang Guo, Anil Kumar Tiwari
IWDW4
2013 Image colorization using sparse representation
abstract
Image colorization is the task to color a grayscale image with limited color cues. In this work, we present a novel method to perform image colorization using sparse representation. Our method first trains an over-complete dictionary in YUV color space. Then taking a grayscale image and a small subset of color pixels as inputs, our method colorizes overlapping image patches via sparse representation; it is achieved by seeking sparse representations of patches that are consistent with both the grayscale image and the color pixels. After that, we aggregate the colorized patches with weights to get an intermediate result. This process iterates until the image is properly colorized. Experimental results show that our method leads to high-quality colorizations with small number of given color pixels. To demonstrate one of the applications of the proposed method, we apply it to transfer the color of one image onto another to obtain a visually pleasing image.
Jiahao Pang, Oscar C. Au, Ketan Tang, Yuanfang Guo
ICASSP4
2013 Arbitrary factor image interpolation using geodesic distance weighted 2D autoregressive modeling
abstract
Least square regression has been widely used in image interpolation. Some existing regression-based interpolation methods used ordinary least squares (OLS) to formulate cost functions. These methods usually have difficulties at object boundaries because OLS is sensitive to outliers. Weighted least squares (WLS) is then adopted to solve the outlier problem. Some weighting schemes have been proposed in the literature. In this paper we propose to use geodesic distance weighting in that geodesic distance can simultaneously measure both the spatial distance and color difference. Another contribution of this paper is that we propose an optimization scheme that can handle arbitrary factor interpolation. The idea is to separate the problem into two parts, an adaptive pixel correlation model and a convolution based image degradation model. Geodesic distance weighted 2D autoregressive model is used to model the pixel correlation which preserves local geometry. The convolution based image degradation model provides the flexibility to handle arbitrary interpolation factor. The entire problem is formulated as a WLS problem constrained by a linear equality.
Ketan Tang, Oscar C. Au, Yuanfang Guo, Jiahao Pang
ICASSP3
2013 Efficient adaptive prediction based reversible image watermarking
abstract
In this paper, we propose a new reversible watermarking algorithm based on additive prediction-error expansion which can recover original image after extracting the hidden data. Embedding capacity of such algorithms depend on the prediction accuracy of the predictor. We observed that the performance of a predictor based on full context prediction is preciser as compared to that of partial context prediction. In view of this observation, we propose an efficient adaptive prediction (EAP) method based on full context, that exploits local characteristics of neighboring pixels much effectively than other prediction methods reported in literature. Experimental results demonstrate that the proposed algorithm has a better embedding capacity and also gives better Peak Signal to Noise Ratio (PSNR) as compared to state-of-the-art reversible watermarking schemes.
Sunil Prasad Jaiswal, Oscar C. Au, Vinit Jakhetiya, Yuanfang Guo, Anil Kumar Tiwari, Yue Kong
ICIP4
2013 Arbitrary factor image interpolation by convolution kernel constrained 2-D autoregressive modeling
abstract
Among existing interpolation methods, convolution-based methods are able to perform arbitrary factor interpolation but the results are usually blurry or jaggy, adaptive interpolation methods usually can reduce the blurry and jaggy artifacts but cannot handle arbitrary factor interpolation. In this paper we propose an arbitrary factor adaptive interpolation algorithm by combining 2-D piecewise autoregressive (PAR) modeling and convolution kernel constraint. PAR model ensures local geometries are well preserved thus the resultant image is not blurry or jaggy. Convolution kernel constraint ensures the recovered high resolution image consistent with the low resolution image, and also provides the flexibility to handle arbitrary interpolation factor. Experiment results show that our algorithm achieves state-of-the-art performance for any interpolation factor.
Ketan Tang, Oscar C. Au, Yuanfang Guo, Jiahao Pang, Lu Fang 0001
ICIP3
2013 Data hiding in error diffused color halftone images
abstract
Halftone image watermarking has been explored and developed rapidly over the past decade. However, there are still issues to be studied. This paper presents a data hiding method called Data Hiding by Dual Color Conjugate Error Diffusion (DHDCCED) to hide a binary secret pattern into two error diffused color halftone images, such that when the two color halftone images are overlaid, the secret pattern will be revealed. The experimental results show that DHDCCED can significantly improve the performances when comparing both the correct decoding rate and the visual quality of the revealed secret pattern to the existing method Color Conjugate Error Diffusion (CCED).
Yuanfang Guo, Oscar C. Au, Ketan Tang, Jiahao Pang, Wenxiu Sun, Lingfeng Xu
ISCAS1
2013 Stereo matching by adaptive weighting selection based cost aggregation
abstract
Cost aggregation is the most essential step for dense stereo correspondence searching, which measures the similarity between pixels in the stereo images. In this paper, based on the analysis of the optimal adaptive weight, we propose a novel support aggregation strategy by adaptive weighting selection. The proposed method calculates the aggregation cost by the joint optimization of both left and right matching cost. By assigning more reasonable weighting coefficients, we exclude the occlusion pixels while preserving sufficient support region for accurate matching. The proposed optimal strategy can be integrated by any other adaptive weighting based cost aggregation method to generate more reasonable similarity measurement. Experimental results show that, compare with traditional methods, our algorithm can reduce the foreground fatten phenomenon while increasing the accuracy in the high texture regions.
Lingfeng Xu, Oscar C. Au, Wenxiu Sun, Lu Fang 0001, Ketan Tang, Yuanfang Guo
ISCAS7
2013 HEVC-based adaptive quantization for screen content by detecting low contrast edge regions
abstract
High-Efficiency Video Coding (HEVC) is the newest video coding standard which can significantly reduce the bit rate by 50% compared with existing standards. The key features and new tools in HEVC are designed for natural video sequences captured by a real camera. Different from natural videos, screen content contain much more edges in text and icon regions. The current video coding standards may blur or even remove low contrast edges, which are very important in screen content for human eyes to recognize the character and the icon. Therefore, this paper proposes an effective modification on HEVC to preserve the low contrast edges in screen content. First, discrete laplacian filter is adopted for edge detection, and then we adaptively adjust QPs for low contrast edge regions, which can be detected based on our designed measurement for edge contrast. Experimental results show that nearly all the regions containing low contrast edges can be detected, and the adjustment of QPs for these regions can greatly protect the edges with no RD performance reduction.
Hong Zhang 0024, Oscar C. Au, Yongfang Shi, Ketan Tang, Yuanfang Guo
ISCAS6
2013 Hiding a Secret Pattern into Color Halftone Images
Yuanfang Guo, Oscar C. Au, Ketan Tang, Jiahao Pang
IWDW1
2013 Chroma Replacing and adaptive Chroma Blending for subpixel-based downsampling
abstract
Subpixel-based downsampling generates images with higher apparent resolution with the expense of annoying color-fringing artifacts near strong edges. In this paper we propose two methods that find a balance in the tradeoff of apparent resolution and color-fringing artifacts. The first method is called Chroma Replacing in which the color-fringing artifacts are completely removed but the subpixel rendering effect is also removed. The second one is called Chroma Blending in which only the color-fringing artifacts that are strong enough to be noticed are removed, and also the subpixel rendering effect is retained. We also propose two objective measures for measuring the similarity of downsampled image to the original image. Experiment results show that the proposed methods are effective in removing color-fringing artifacts, without harming the high apparent resolution.
Ketan Tang, Oscar C. Au, Lu Fang 0001, Yuanfang Guo, Jiahao Pang
MMSP4
2012 Image de-quantization via spatially varying sparsity prior
abstract
We address the problem of image de-quantization, which is also known as bit-depth expansion if the reconstructed 2D signal is re-quantized into higher bit-precision. In this paper, a novel image de-quantization method based on convex optimization theory is proposed, which exploits the spatially varying characteristics of image surface. We test our method on image bit-depth expansion problems, and the experimental results show that proposed method can achieve superior PSNR and SSIM performance.
Pengfei Wan 0001, Oscar C. Au, Ketan Tang, Yuanfang Guo
ICIP4
2012 From 2D Extrapolation to 1D Interpolation: Content Adaptive Image Bit-Depth Expansion
abstract
In this paper, we address the problem of image bit-depth expansion and present a novel method to generate high bit-depth (HBD) images from a single low bit-depth (LBD) image. We expand image bit-depth by reconstructing the least significant bits (LSBs) for the LBD image after it is rescaled to high bit-depth. For image regions whose intensities are neither locally maximum nor minimum, neighborhood flooding is applied to convert 2D interpolation problem into 1D interpolation, for local maxima/minima (LMM) regions where interpolation is not applicable, a virtual skeleton marking algorithm is proposed to convert problematic 2D extrapolation problem into 1D interpolation. At last, a content-adaptive reconstruction model is proposed to obtain the output HBD image. The experimental results show that proposed method significantly outperforms existing methods in PSNR and SSIM without contouring artifacts.
Pengfei Wan 0001, Oscar C. Au, Ketan Tang, Yuanfang Guo, Lu Fang 0001
ICME4
2011 Image Interpolation Using Autoregressive Model and Gauss-Seidel Optimization
abstract
In this paper we propose a simple yet effective image interpolation algorithm based on autoregressive model. Unlike existing algorithms which rely on low resolution pixels to estimate interpolation coefficients, we optimize the interpolation coefficients and high resolution pixel values jointly from one optimization problem. Although the two sets of variables are coupled in the cost function, the problem can be effectively solved using Gauss-Seidel method. We prove the iterations are guaranteed to converge. Experiments show that on average we have over 3dB gain compared to bicubic interpolation and over 0.1dB gain compared to SAI.
Ketan Tang, Oscar C. Au, Lu Fang 0001, Zhiding Yu, Yuanfang Guo
ICIG5
2011 Multi-scale analysis of color and texture for salient object detection
abstract
In this paper we propose a multi-scale segment-based framework for salient object detection. In this framework texture and color features are used together to provide diverse information of salient object. Segmentation is performed on three different scales so that the object boundary can be accurately captured with high probability. Besides, we propose a novel adaptive feature combination mechanism to combine the saliency maps produced with different features, in which the combining weight of each saliency map is learned using online learning. Experiment results demonstrate that the proposed method significantly outperforms the state-of-the-art methods.
Ketan Tang, Oscar C. Au, Lu Fang 0001, Zhiding Yu, Yuanfang Guo
ICIP5
2011 Data hiding in dot diffused halftone images
abstract
In this paper, we propose two halftone image watermarking methods. Data Hiding by Conjugate Dot Diffusion (DHCDD) and Data Hiding by Dual Conjugate Dot Diffusion (DHD-CDD). DHDCDD is an improved method of DHCDD. Both of these two methods can embed a secret pattern into two halftone images. When the two halftone images are overlaid, the secret pattern will be revealed. Compared to the recent method Noise Balanced Dot Diffusion, the experimental results show that the proposed methods are better in both Correct Decoding Rate and the visual quality of the revealed hidden pattern.
Yuanfang Guo, Oscar C. Au, Ketan Tang, Lu Fang 0001, Zhiding Yu
ICME1
2011 How anti-aliasing filter affects image contrast: An analysis from majorization theory perspective
abstract
When we design an anti-aliasing low pass filter, it is usually an IIR filter. We need to truncate the filter to an FIR filter. One may think that the more taps there are, the better the image quality is. However, we find that there exists an optimal value of tap number that will give the best visual quality. Filters with larger or smaller number of taps will degrade the image quality, due to the fact that the image contrast is reduced. In this paper we analyze this phenomenon using majorization theory and find that the image contrast can be formulated as a Schur convex function on filter coefficients. We also propose an effective method to choose the best filter so that the image contrast is maximized, so as to give best visual quality.
Ketan Tang, Oscar C. Au, Lu Fang 0001, Zhiding Yu, Yuanfang Guo
ICME5
2011 Towards robust and efficient segmentation: An approach based on inter-region contour and intra-region content analysis
abstract
We address the problem of boundary estimation by formulating it as inter-region contour and intra-region information analysis in the framework of graph-based segmentation. Given an image without any prior information about object model and class, we seek to approximate one's instant perception of visual similarity. The method can serve as a preprocessing step for many higher level operations that require regional support, such as scene understanding and object recognition. We show in this paper that the defined region comparison predicate makes a better boundary estimator than efficient graph-based image segmentation (EGS) - a well known and widely used segmentation method. We further illustrate, by making a small relaxation, further improvement of segmentation performance can be achieved. Experimental results have demonstrated the effectiveness of our proposed method.
Zhiding Yu, Oscar C. Au, Ketan Tang, Lingfeng Xu, Wenxiu Sun, Yuanfang Guo
ICME6
2010 Inter-channel demosaicking traces for digital image forensics
abstract
Digital image forensics seeks to detect statistical traces left by image acquisition or post-processing in order to establish an images source and authenticity. Digital cameras acquire an image with one sensor overlayed with a color filter array (CFA), capturing at each spatial location one sample from the three necessary color channels. The missing pixels must be interpolated in a process known as demosaicking. This process is highly nonlinear and can vary greatly between different camera brands and models. Most practical algorithms, however, introduce correlations between the color channels, which are often different between algorithms. In this paper, we show how these correlations can be used to construct a characteristic map that is useful in matching an image to its source. Results show that our method employing inter-channel traces can distinguish between sophisticated demosaicking algorithms. It can complement existing classifiers based on inter-pixel correlations by providing a new feature dimension.
John S. Ho, Oscar C. Au, Jiantao Zhou 0001, Yuanfang Guo
ICME4