Yanfu Zhang

dblp:154/7698 · DBLP profile ↗
← Back
32ranked-venue papers
10as first author
26since 2021 · last 2025
0000-0002-0183-925XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 24 · 6 first-author · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 6 first-author · 13 since 2021Databases, data management, data science and information retrieval · 5 · 4 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 first-author · 3 since 2021Systems, architecture and hardware · 2 · 1 since 2021
YearPublicationVenuePosition
2025 Dynamic Retriever for In-Context Knowledge Editing via Policy Optimization
abstract
Large language models (LLMs) excel at factual recall yet still propagate stale or incorrect knowledge.In-context knowledge editing offers a gradient-free remedy suitable for black-box APIs, but current editors rely on static demonstration sets chosen by surface-level similarity, leading to two persistent obstacles: (i) a quantity-quality trade-off, and (ii) lack of adaptivity to task difficulty.We address these issues by dynamically selecting supporting demonstrations according to their utility for the edit.We propose Dynamic Retriever for In-Context Knowledge Editing (DR-IKE), a lightweight framework that (1) trains a BERT retriever with REINFORCE to rank demonstrations by editing the reward, and (2) employs a learnable threshold to prune low-value examples, shortening the prompt when the edit is easy and expanding it when the task is hard.DR-IKE performs editing without modifying model weights, relying solely on forward passes for compatibility with black-box LLMs.On the COUNTERFACT benchmark, it improves edit success by up to 17.1%, reduces latency by 41.6%, and preserves accuracy on unrelated queries, demonstrating scalable and adaptive knowledge editing.
Mahmud Wasif Nafee, Maiqi Jiang, Haipeng Chen 0001, Yanfu Zhang
EMNLP4
2025 Controllable Memorization in LLMs via Weight Pruning
abstract
The evolution of pre-trained large language models (LLMs) has significantly transformed natural language processing.However, these advancements pose challenges, particularly the unintended memorization of training data, which raises ethical and privacy concerns.While prior research has largely focused on mitigating memorization or extracting memorized information, the deliberate control of memorization has been underexplored.This study addresses this gap by introducing a novel and unified gradient-based weight pruning framework to freely control memorization rates in LLMs.Our method enables fine-grained control over pruning parameters, allowing models to suppress or enhance memorization based on application-specific requirements.Experimental results demonstrate that our approach effectively balances the trade-offs between memorization and generalization, with an increase of up to 89.3% in Fractional ER suppression and 40.9% in Exact ER amplification compared to the original models.
Chenjie Ni, Zhepeng Wang 0001, Runxue Bao, Shangqian Gao, Yanfu Zhang
EMNLP5
2025 HCNQA: Enhancing 3D VQA with Hierarchical Concentration Narrowing Supervision
Jianuo Zhu, Fangjing Wang, Yanfu Zhang, Feng Zheng 0001
ICANN (2)5
2025 Safe Screening Rules for Group OWL Models
abstract
Group Ordered Weighted L1-Norm (Group OWL) regularized models have emerged as a useful procedure for high-dimensional sparse multi-task learning with correlated features. Proximal gradient methods are used as standard approaches to solving Group OWL models. However, Group OWL models usually suffer huge computational costs and memory usage when the feature size is large in the high-dimensional scenario. To address this challenge, in this paper, we are the first to propose the safe screening rule for Group OWL models by effectively tackling the structured non-separable penalty, which can quickly identify the inactive features that have zero coefficients across all the tasks. Thus, by removing the inactive features during the training process, we may achieve substantial computational gain and memory savings. More importantly, the proposed screening rule can be directly integrated with the existing solvers both in the batch and stochastic settings. Theoretically, we prove our screening rule is safe and also can be safely applied to the existing iterative optimization algorithms. Our experimental results demonstrate that our screening rule can effectively identify the inactive features and leads to a significant computational speedup without any loss of accuracy.
Runxue Bao, Quanchao Lu 0002, Yanfu Zhang
IJCNN3
2025 CLEP-DG: Contrastive Learning for Speech Emotion Domain Generalization via Soft Prompt Tuning
Jiacheng Shi 0005, Yanfu Zhang, Ye Gao 0001
INTERSPEECH2
2025 Safe Screening Rules for Group SLOPE
Runxue Bao, Quanchao Lu 0002, Yanfu Zhang
ECML/PKDD (2)3
2025 UniAV: Unified Audio-Visual Perception for Multi-Task Video Event Localization
abstract
Video event localization tasks include temporal action localization (TAL), sound event detection (SED) and audio-visual event localization (AVEL). Existing methods tend to over-specialize on individual tasks, neglecting the equal importance of these different events for a complete understanding of video content. In this work, we aim to develop a unified framework to solve TAL, SED and AVEL tasks together to facilitate holistic video understanding. However, it is challenging since different tasks emphasize distinct event characteristics and there are substantial disparities in existing task-specific datasets (size/domain/duration). It leads to unsatisfactory results when applying a naive multi-task strategy. To tackle the problem, we introduce UniAV, a Unified Audio-Visual perception network to effectively learn and share mutually beneficial knowledge across tasks and modalities. Concretely, we propose a unified audio-visual encoder to derive generic representations from multiple temporal scales for videos from all tasks. Meanwhile, task-specific experts are designed to capture the unique knowledge specific to each task. Besides, instead of using separate prediction heads, we develop a novel unified language-aware classifier by utilizing semantic-aligned task prompts, enabling our model to flexibly localize various instances across tasks with an impressive open-set ability to localize novel categories. Extensive experiments demonstrate that UniAV, with its unified architecture, significantly outperforms both single-task models and the naive multi-task baseline across all three tasks. It achieves superior or on-par performances compared to the state-of-the-art task-specific methods on ActivityNet 1.3, DESED and UnAV-100 benchmarks.
Tiantian Geng, Teng Wang 0007, Jinming Duan 0001, Yanfu Zhang, Weili Guan, Feng Zheng 0001, Ling Shao 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2024 Device-Wise Federated Network Pruning
abstract
Neural network pruning, particularly channel pruning, is a widely used technique for compressing deep learning models to enable their deployment on edge devices with limited resources. Typically, redundant weights or structures are removed to achieve the target resource budget. Although data-driven pruning approaches have proven to be more effective, they cannot be directly applied to federated learning (FL), which has emerged as a popular technique in edge computing applications, because of distributed and confidential datasets. In response to this challenge, we design a new network pruning method for FL. We propose device-wise sub-networks for each device, assuming that the data distribution is similar within each device. These sub-networks are generated through sub-network embeddings and a hypernetwork. To further minimize memory usage and communication costs, we permanently prune the full model to remove weights that are not useful for all devices. During the FL process, we simultaneously train the device-wise sub-networks and the base sub-network to facilitate the pruning process. We then finetune the pruned model with device-wise sub-networks to regain performance. Moreover, we provided the theoretical guarantee of convergence for our method. Our method achieves better performance and resource trade-off than other well-established network pruning baselines, as demonstrated through extensive experiments on CIFAR-10, CIFAR-100, and TinyImageNet.
Shangqian Gao, Junyi Li 0002, Yanfu Zhang, Tom Weidong Cai, Heng Huang 0001
CVPR4
2024 BilevelPruning: Unified Dynamic and Static Channel Pruning for Convolutional Neural Networks
abstract
Most existing dynamic or runtime channel pruning meth-ods have to store all weights to achieve efficient inference, which brings extra storage costs. Static pruning methods can reduce storage costs directly, but their performance is limited by using a fixed sub-network to approximate the orig-inal model. Most existing pruning works suffer from these drawbacks because they were designed to only conduct ei-ther static or dynamic pruning. In this paper, we propose a novel method to solve both efficiency and storage challenges via simultaneously conducting dynamic and static channel pruning for convolutional neural networks. We propose a new bi-level optimization based model to naturally integrate the static and dynamic channel pruning. By doing so, our method enjoys benefits from both sides, and the disadvan-tages of dynamic and static pruning are reduced. After pruning, we permanently remove redundant parameters and then finetune the model with dynamic flexibility. Experimental results on CIFAR-10 and ImageNet datasets suggest that our method can achieve state-of-the-art performance compared to existing dynamic and static channel pruning methods.
Shangqian Gao, Yanfu Zhang, Feihu Huang 0001, Heng Huang 0001
CVPR2
2024 Depth-Aware Concealed Crop Detection in Dense Agricultural Scenes
abstract
Concealed Object Detection (COD) aims to identify objects visually embedded in their background. Existing COD datasets and methods predominantly focus on animals or humans, ignoring the agricultural domain, which often contains numerous, small, and concealed crops with severe occlusions. In this paper, we introduce Concealed Crop Detection (CCD), which extends classic COD to agricultural domains. Experimental study shows that unimodal data provides insufficient information for CCD. To address this gap, we first collect a large-scale RGB-D dataset, ACOD-12K, containing high-resolution crop images and depth maps. Then, we propose a foundational framework named Recurrent Iterative Segmentation Network (RISNet). To tackle the challenge of dense objects, we employ multi-scale receptive fields to capture objects of varying sizes, thus enhancing the detection performance for dense objects. By fusing depth features, our method can acquire spatial information about concealed objects to mitigate disturbances caused by intricate backgrounds and occlusions. Furthermore, our model adopts a multi-stage iterative approach, using predictions from each stage as gate attention to reinforce position information, thereby improving the detection accuracy for small objects. Extensive experimental results demonstrate that our RISNet achieves new state-of-the-art performance on both newly proposed CCD and classic COD tasks. All resources will be available at https://github.com/Kki2Eve/RISNet.
Liqiong Wang, Yanfu Zhang, Fangyi Wang, Feng Zheng 0001
CVPR3
2024 Auto- Train-Once: Controller Network Guided Automatic Network Pruning from Scratch
abstract
Current techniques for deep neural network (DNN) pruning often involve intricate multi-step processes that re-quire domain-specific expertise, making their widespread adoption challenging. To address the limitation, the Only-Train-Once (OTO) and OTOv2 are proposed to eliminate the need for additional fine-tuning steps by directly training and compressing a general DNN from scratch. Never-theless, the static design of optimizers (in OTO) can lead to convergence issues of local optima. In this paper, we proposed the Auto-Train-Once (A TO), an innovative net-work pruning algorithm designed to automatically reduce the computational and storage costs of DNNs. During the model training phase, our approach not only trains the tar-get model but also leverages a controller network as an ar-chitecture generator to guide the learning of target model weights. Furthermore, we developed a novel stochastic gradient algorithm that enhances the coordination between model training and controller network training, thereby im-proving pruning performance. We provide a comprehen-sive convergence analysis as well as extensive experiments, and the results show that our approach achieves state-of-the-art performance across various model architectures (including ResNet18, ResNet34, ResNet50, ResNet56, and MobileNetv2) on standard benchmark datasets (CIFAR-10, CIFAR-100, and ImageNet). The code is available at https: 11 g i thub. comlxidon gwul Auto Train Once.
Xidong Wu, Shangqian Gao, Runxue Bao, Yanfu Zhang, Xiaoqian Wang 0001, Heng Huang 0001
CVPR6
2024 Unlocking Memorization in Large Language Models with Dynamic Soft Prompting
abstract
Zhepeng Wang, Runxue Bao, Yawen Wu, Jackson Taylor, Cao Xiao, Feng Zheng, Weiwen Jiang, Shangqian Gao, Yanfu Zhang. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Zhepeng Wang 0001, Runxue Bao, Yawen Wu, Jackson Taylor, Cao Xiao, Feng Zheng 0001, Weiwen Jiang, Shangqian Gao, Yanfu Zhang
EMNLP9
2024 Edge-Aware Dual Branch Network for Nucleus Instance Segmentation
abstract
In mobile healthcare and remote diagnosis, nucleus segmentation is a critical step for pathological analysis, diagnosis, and classification, requiring real-time processing and high accuracy. However, variations in nucleus size, blurred contours, uneven staining, cell clustering, and overlapping cells hinder precise segmentation. Additionally, existing deep learning models often prioritize accuracy at the cost of increased complexity, making them unsuitable for resource-limited edge devices and real-world deployment. To address the aforementioned issues, we propose an edge-aware dual branch network for nucleus instance segmentation. The network simultaneously predicts target information and target contours. Within the network, we propose a context fusion block (CF-block) that effectively extracts and merges contextual information from the network. Additionally, we introduce a post-processing method that combines the target information and target contours to distinguish overlapping nuclei and generate an instance segmentation image. Extensive quantitative evaluations are conducted to assess the performance of our method. Experimental results demonstrate the superior performance of the proposed method compared to state-of-the-art approaches on the BNS, MoNuSeg, and CPM-17 datasets.
Junzhou Chen 0002, Yanfu Zhang, Sidi Lu
SEC2
2024 Self-guided Knowledge-Injected Graph Neural Network for Alzheimer's Diseases
Zhepeng Wang 0001, Runxue Bao, Yawen Wu, Lei Yang 0018, Liang Zhan, Feng Zheng 0001, Weiwen Jiang, Yanfu Zhang
MICCAI (2)9
2023 Demystify the Gravity Well in the Optimization Landscape (Student Abstract)
abstract
We provide both empirical and theoretical insights to demystify the gravity well phenomenon in the optimization landscape. We start from describe the problem setup and theoretical results (an escape time lower bound) of the Softmax Gravity Well (SGW) in the literature. Then we move toward the understanding of a recent observation called ASR gravity well. We provide an explanation of why normal distribution with high variance can lead to suboptimal plateaus from an energy function point of view. We also contribute to the empirical insights of curriculum learning by comparison of policy initialization by different normal distributions. Furthermore, we provide the ASR escape time lower bound to understand the ASR gravity well theoretically. Future work includes more specific modeling of the reward as a function of time and quantitative evaluation of normal distribution’s influence on policy initialization.
Jason Xiaotian Dou, Runxue Bao, Susan Song, Shuran Yang, Yanfu Zhang, Paul Pu Liang, Haiyi Harry Mao
AAAI5
2023 Structural Alignment for Network Pruning through Partial Regularization
abstract
In this paper, we propose a novel channel pruning method to reduce the computational and storage costs of Convolutional Neural Networks (CNNs). Many existing one-shot pruning methods directly remove redundant structures, which brings a huge gap between the model before and after network pruning. This gap will no doubt result in performance loss for network pruning. To mitigate this gap, we first learn a target sub-network during the model training process, and then we use this sub-network to guide the learning of model weights through partial regularization. The target sub-network is learned and produced by using an architecture generator, and it can be optimized efficiently. In addition, we also derive the proximal gradient for our proposed partial regularization to facilitate the structural alignment process. With these designs, the gap between the pruned model and the sub-network is reduced, thus improving the pruning performance. Empirical results also suggest that the sub-network found by our method has a much higher performance than the one-shot pruning setting. Extensive experiments show that our method can achieve state-of-the-art performances on CIFAR-10 and ImageNet with ResNets and MobileNet-V2.
Shangqian Gao, Yanfu Zhang, Feihu Huang 0001, Heng Huang 0001
ICCV3
2022 Disentangled Differentiable Network Pruning
Shangqian Gao, Feihu Huang 0001, Yanfu Zhang, Heng Huang 0001
ECCV (11)3
2022 Recover Fair Deep Classification Models via Altering Pre-trained Structure
Yanfu Zhang, Shangqian Gao, Heng Huang 0001
ECCV (13)1
2022 Toward Unified Data and Algorithm Fairness via Adversarial Data Augmentation and Adaptive Model Fine-tuning
abstract
There is some recent research interest in algorithmic fairness for biased data. There are a variety of pre-, in-, and post-processing methods designed for this problem. However, these methods are exclusively targeting data unfairness and algorithmic unfairness. In this paper, we propose a novel intra-processing method to broaden the application scenario of fairness methods, which can simultaneously address the two bias sources. Since training modern deep models from scratch is expensive due to the enormous training data and the complicated structures, we propose an augmentation and fine-tuning framework. First, we design an adversarial attack to generate weighted samples disentangled with the protected attribute. Next, we identify the fair sub-structure in the biased model and fine-tune the model via weight reactivation. At last, we provide an optional joint training scheme for the augmentation and the fine-tuning. Our method can be combined with a variety of fairness measures. We benchmark our method and some related baselines to show the advantage and the scalability. Experimental results on several standard datasets demonstrate that our approach can effectively learn fair augmentation and achieve superior results to the state-of-the-art baselines. Our method also generalizes well to different types of data.
Yanfu Zhang, Runxue Bao, Jian Pei 0001, Heng Huang 0001
ICDM1
2022 Improving Social Network Embedding via New Second-Order Continuous Graph Neural Networks
abstract
Graph neural networks (GNN) are powerful tools in many web research problems. However, existing GNNs are not fully suitable for many real-world web applications. For example, over-smoothing may affect personalized recommendations and the lack of an explanation for the GNN prediction hind the understanding of many business scenarios. To address these problems, in this paper, we propose a new second-order continuous GNN which naturally avoids over-smoothing and enjoys better interpretability. There is some research interest in continuous graph neural networks inspired by the recent success of neural ordinary differential equations (ODEs). However, there are some remaining problems w.r.t. the prevailing first-order continuous GNN frameworks. Firstly, augmenting node features is an essential, however heuristic step for the numerical stability of current frameworks; secondly, first-order methods characterize a diffusion process, in which the over-smoothing effect w.r.t. node representations are intrinsic; and thirdly, there are some difficulties to integrate the topology of graphs into the ODEs. Therefore, we propose a framework employing second-order graph neural networks, which usually learn a less stiff transformation than the first-order counterpart. Our method can also be viewed as a coupled first-order model, which is easy to implement. We propose a semi-model-agnostic method based on our model to enhance the prediction explanation using high-order information. We construct an analog between continuous GNNs and some famous partial differential equations and discuss some properties of the first and second-order models. Extensive experiments demonstrate the effectiveness of our proposed method, and the results outperform related baselines.
Yanfu Zhang, Shangqian Gao, Jian Pei 0001, Heng Huang 0001
KDD1
2022 Robust Self-Supervised Structural Graph Neural Network for Social Network Prediction
abstract
The self-supervised graph representation learning has achieved much success in recent web based research and applications, such as recommendation system, social networks, and anomaly detection. However, existing works suffer from two problems. Firstly, in social networks, the influential neighbors are important, but the overwhelming routine in graph representation-learning utilizes the node-wise similarity metric defined on embedding vectors that cannot exactly capture the subtle local structure and the network proximity. Secondly, existing works implicitly assume a universal distribution across datasets, which presumably leads to sub-optimal models considering the potential distribution shift. To address these problems, in this paper, we learn structural embeddings in which the proximity is characterized by 1-Wasserstein distance. We propose a distributionally robust self-supervised graph neural network framework to learn the representations. More specifically, in our method, the embeddings are computed based on subgraphs centering at the node of interest and represent both the node of interest and its neighbors, which better preserves the local structure of nodes. To make our model end-to-end trainable, we adopt a deep implicit layer to compute the Wasserstein distance, which can be formulated as a differentiable convex optimization problem. Meanwhile, our distributionally robust formulation explicitly constrains the maximal diversity for matched queries and keys. As such, our model is insensitive to the data distributions and has better generalization abilities. Extensive experiments demonstrate that the graph encoder learned by our approach can be utilized for various downstream analyses, including node classification, graph classification, and top-k similarity search. The results show our algorithm outperforms state-of-the-art baselines, and the ablation study validates the effectiveness of our design.
Yanfu Zhang, Hongchang Gao, Jian Pei 0001, Heng Huang 0001
WWW1
2021 Learning Better Visual Data Similarities via New Grouplet Non-Euclidean Embedding
abstract
In many computer vision problems, it is desired to learn the effective visual data similarity such that the prediction accuracy can be enhanced. Deep Metric Learning (DML) methods have been actively studied to measure the data similarity. Pair-based and proxy-based losses are the two major paradigms in DML. However, pair-wise methods involve expensive training costs, while proxy-based methods are less accurate in characterizing the relationships between data points. In this paper, we provide a hybrid grouplet paradigm, which inherits the accurate pair-wise relationship in pair-based methods and the efficient training in proxy-based methods. Our method also equips a non-Euclidean space to DML, which employs a hierarchical representation manifold. More specifically, we propose a unified graph perspective — different DML methods learn different local connecting patterns between data points. Based on the graph interpretation, we construct a flexible subset of data points, dubbed grouplet. Our grouplet doesn’t require explicit pair-wise relationships, instead, we encode the data relationships in an optimal transport problem regarding the proxies, and solve this problem via a differentiable implicit layer to automatically determine the relationships. Extensive experimental results show that our method significantly outperforms state-of-the-art baselines on several benchmarks. The ablation studies also verify the effectiveness of our method.
Yanfu Zhang, Lei Luo 0001, Wenhan Xian, Heng Huang 0001
ICCV1
2021 Exploration and Estimation for Model Compression
abstract
Deep neural networks achieve great success in many visual recognition tasks. However, the model deployment is usually subject to some computational resources. Model pruning under computational budget has attracted growing attention. In this paper, we focus on the discrimination-aware compression of Convolutional Neural Networks (CNNs). In prior arts, directly searching the optimal sub-network is an integer programming problem, which is non-smooth, non-convex, and NP-hard. Meanwhile, the heuristic pruning criterion lacks clear interpretability and doesn’t generalize well in applications. To address this problem, we formulate sub-networks as samples from a multivariate Bernoulli distribution and resort to the approximation of continuous problem. We propose a new flexible search scheme via alternating exploration and estimation. In the exploration step, we employ stochastic gradient Hamiltonian Monte Carlo with budget-awareness to generate sub-networks, which allows large search space with efficient computation. In the estimation step, we deduce the sub-network sampler to a near-optimal point, to promote the generation of high-quality sub-networks. Unifying the exploration and estimation, our approach avoids early falling into local minimum via a fast gradient-based search in a larger space. Extensive experiments on CIFAR-10 and ImageNet show that our method achieves state-of-the-art performances on pruning several popular CNNs.
Yanfu Zhang, Shangqian Gao, Heng Huang 0001
ICCV1
2021 Unified Fairness from Data to Learning Algorithm
abstract
In classification problems, individual fairness prevents discrimination against individuals based on protected attributes. Fairness-aware methods usually consist of two stages, first determining a fair metric concerning the similarity between different instances and then learning the fairness-aware model. However, existing works usually consider these two stages separately and only focus on improving the individual stage. Moreover, the choice of fair metric is heavily dependent on the task or dataset of interest, which requires ad-hoc domain knowledge and introduces extra difficulty into algorithm designing. As such, this discrepancy presumably leads to sub-optimal fairness-aware pipelines for different applications. In this paper, we propose to fill in the fairness learning gap between these two stages by automatically learning an effective metric integrated into the fairness of both data and classifiers. Specifically, we formulate the fairness-aware classification as a distributional robustness optimization problem based on deep metric learning and propose an effective optimization algorithm to solve it. Meanwhile, we establish the asymptotically unbiased generalization bounds for the proposed algorithm using the techniques of U-statistics. The experimental results on popular benchmark datasets demonstrate that the proposed approach achieves consistent improvement concerning several fairness assessments.
Yanfu Zhang, Lei Luo 0001, Heng Huang 0001
ICDM1
2021 Disentangled and Proportional Representation Learning for Multi-view Brain Connectomes
Yanfu Zhang, Liang Zhan, Shandong Wu, Paul M. Thompson, Heng Huang 0001
MICCAI (7)1
2021 A Faster Decentralized Algorithm for Nonconvex Minimax Problems
abstract
In this paper, we study the nonconvex-strongly-concave minimax optimization problem on decentralized setting. The minimax problems are attracting increasing attentions because of their popular practical applications such as policy evaluation and adversarial training. As training data become larger, distributed training has been broadly adopted in machine learning tasks. Recent research works show that the decentralized distributed data-parallel training techniques are specially promising, because they can achieve the efficient communications and avoid the bottleneck problem on the central node or the latency of low bandwidth network. However, the decentralized minimax problems were seldom studied in literature and the existing methods suffer from very high gradient complexity. To address this challenge, we propose a new faster decentralized algorithm, named as DM-HSGD, for nonconvex minimax problems by using the variance reduced technique of hybrid stochastic gradient descent. We prove that our DM-HSGD algorithm achieves stochastic first-order oracle (SFO) complexity of $O(\kappa^3 \epsilon^{-3})$ for decentralized stochastic nonconvex-strongly-concave problem to search an $\epsilon$-stationary point, which improves the exiting best theoretical results. Moreover, we also prove that our algorithm achieves linear speedup with respect to the number of workers. Our experiments on decentralized settings show the superior performance of our new algorithm.
Wenhan Xian, Feihu Huang 0001, Yanfu Zhang, Heng Huang 0001
NeurIPS3
2020 Adversarial Nonnegative Matrix Factorization
abstract
Nonnegative Matrix Factorization (NMF) has become an increasingly important research topic in machine learning. Despite all the practical success, most of existing NMF models are still vulnerable to adversarial attacks. To overcome this limitation, we propose a novel Adversarial NMF (ANMF) approach in which an adversary can exercise some control over the perturbed data generation process. Different from the traditional NMF models which focus on either the regular input or certain types of noise, our model considers potential test adversaries that are beyond the pre-defined constraints, which can cope with various noises (or perturbations). We formulate the proposed model as a bilevel optimization problem and use Alternating Direction Method of Multipliers (ADMM) to solve it with convergence analysis. Theoretically, the robustness analysis of ANMF is established under mild conditions dedicating asymptotically unbiased prediction. Extensive experiments verify that ANMF is robust to a broad categories of perturbations, and achieves state-of-the-art performances on distinct real-world benchmark datasets.
Lei Luo 0001, Yanfu Zhang, Heng Huang 0001
ICML2
2019 Improved Generalization of Heading Direction Estimation for Aerial Filming Using Semi-Supervised Regression
abstract
In the task of Autonomous aerial filming of a moving actor (e.g. a person or a vehicle), it is crucial to have a good heading direction estimation for the actor from the visual input. However, the models obtained in other similar tasks, such as pedestrian collision risk analysis and human-robot interaction, are very difficult to generalize to the aerial filming task, because of the difference in data distributions. Towards improving generalization with less amount of labeled data, this paper presents a semi-supervised algorithm for heading direction estimation problem. We utilize temporal continuity as the unsupervised signal to regularize the model and achieve better generalization ability. This semi-supervised algorithm is applied to both training and testing phases, which increases the testing performance by a large margin. We show that by leveraging unlabeled sequences, the amount of labeled data required can be significantly reduced. We also discuss several important details on improving the performance by balancing labeled and unlabeled loss, and making good combinations. Experimental results show that our approach robustly outputs the heading direction for different types of actor. The aesthetic value of the video is also improved in the aerial filming task.
Aayush Ahuja, Yanfu Zhang, Rogerio Bonatti, Sebastian A. Scherer
ICRA3
2019 Brain Dynamics Through the Lens of Statistical Mechanics by Unifying Structure and Function
Igor Fortel, Mitchell Butler, Laura E. Korthauer, Liang Zhan, Olusola Ajilore, Ira Driscoll, Anastasios Sidiropoulos, Yanfu Zhang, Lei Guo 0028, Heng Huang 0001, Dan Schonfeld, Alex D. Leow
MICCAI (5)8
2019 Integrating Heterogeneous Brain Networks for Predicting Brain Disease Conditions
Yanfu Zhang, Liang Zhan, Tom Weidong Cai, Paul M. Thompson, Heng Huang 0001
MICCAI (4)1
2017 HazeRD: An outdoor scene dataset and benchmark for single image dehazing
abstract
In this paper, a new dataset, HazeRD, is proposed for benchmarking dehazing algorithms under more realistic haze conditions. HazeRD contains fifteen real outdoor scenes, for each of which five different weather conditions are simulated. As opposed to prior datasets that made use of synthetically generated images or indoor images with unrealistic parameters for haze simulation, our outdoor dataset allows for more realistic simulation of haze with parameters that are physically realistic and justified by scattering theory. All images are of high resolution, typically six to eight megapixels. We test the performance of several state-of-the-art dehazing techniques on HazeRD. The results exhibit a significant difference among algorithms across the different datasets, reiterating the need for more realistic datasets such as ours and for more careful benchmarking of the methods.
Yanfu Zhang, Li Ding 0009, Gaurav Sharma 0001
ICIP1
2015 Multi-Focus Image Fusion Based on Spatial Frequency in Discrete Cosine Transform Domain
abstract
Multi-focus image fusion in wireless visual sensor networks (WVSN) is a process of fusing two or more images to obtain a new one which contains a more accurate description of the scene than any of the individual source images. In this letter, we propose an efficient algorithm to fuse multi-focus images or videos using discrete cosine transform (DCT) based standards in WVSN. The spatial frequencies of the corresponding blocks from source images are calculated as the contrast criteria, and the blocks with the larger spatial frequencies compose the DCT presentation of the output image. Experiments on plenty of pairs of multi-focus images coded in Joint Photographic Experts Group (JPEG) standard are conducted to evaluate the fusion performance. The results show that our fusion method improves the quality of the output image visually and outperforms the previous DCT based techniques and the state-of-art methods in terms of the objective evaluation.
Liu Cao, Longxu Jin, Hongjiang Tao, Guoning Li, Zhuang Zhuang, Yanfu Zhang
IEEE Signal Process. Lett.6