EDBT 2026 Demo / reviewers in the wild / expert
Hao Zhang 0063
dblp:55/2270-63
· DBLP profile ↗
31ranked-venue papers
2as first author
29since 2021 · last 2026
0009-0007-1175-5918ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 26 · 2 first-author · 24 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 1 first-author · 17 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Bidirectional Noise Injection: Enhancing Diffusion Models via Coordinated Input-Output PerturbationabstractDiffusion models have demonstrated remarkable success in image generation, yet a persistent challenge remains: the bias between model predictions and the target distribution. In this paper, we propose a Bidirectional Noise Injection framework for enhancing diffusion models, implemented via Coordinated Input-Output Perturbation (CIOP). Our approach mitigates this bias by randomly applying synchronized noise injection to both the model inputs and the prediction targets during the training stage. This stochastic, synchronized noise injected acts as a smoothing mechanism that effectively reduces the 2-Wasserstein distance between the predicted and target distributions, as substantiated by our theoretical analysis based on optimal transport theory. Extensive experiments on multiple benchmark datasets and various generative tasks demonstrate that our method improves generation quality and training efficiency without incurring additional computational cost. Furthermore, the design of CIOP enables seamless integration with existing diffusion model improvements and advanced frameworks, thereby broadening its applicability. These results highlight the potential of Bidirectional Noise Injection via CIOP to alleviate bias in diffusion-based generative models across a wide range of settings. Tianyi Zheng 0001, Jiayang Gao, Peng-Tao Jiang, Fengxiang Yang, Ben Wan, Hao Zhang 0063, Jinwei Chen 0003, Jia Wang 0004, Bo Li 0130 |
AAAI | 6 |
| 2026 | Diffusion for automatic matting
Longfei Huang, Hao Zhang 0063, Xinning Zhu, Lunde Chen |
Neurocomputing | 3 |
| 2026 | Bidirectional Beta-Tuned Diffusion ModelabstractDiffusion models have gained significant attention in the field of generative modeling due to their capability to produce high-quality samples. However, recent studies show that applying a uniform treatment to all distributions during the training of diffusion models is sub-optimal. In this paper, we present a comprehensive theoretical analysis of the forward process in diffusion models. Our findings indicate that distribution variations are not uniform throughout the diffusion process, with the sharpest changes occurring during the initial stages. Moreover, we observe that the initial distribution converges to a Gaussian distribution at an exponential rate, indicating that different initial distributions rapidly become quite similar during the forward diffusion process. Consequently, employing a uniform timestep sampling strategy does not effectively capture these dynamics, potentially leading to sub-optimal training outcomes for diffusion models. To remedy this, we introduce the Bidirectional Beta-Tuned Diffusion Model (BB-TDM). The BB-TDM leverages the Beta distribution to design the timestep sampling distribution and enhance the separation between different initial distributions during the diffusion process. By selecting appropriate parameters, the BB-TDM ensures that the timestep sampling distribution is aligned with the properties of the forward diffusion process and moderates the convergence speed of different initial distributions. Extensive experiments across various benchmark datasets on different diffusion models confirm the efficacy of the proposed BB-TDM. Tianyi Zheng 0001, Jiayang Zou, Peng-Tao Jiang, Hao Zhang 0063, Jinwei Chen 0003, Jia Wang 0004, Bo Li 0115 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2026 | DepthMaster: Taming Diffusion Models for Monocular Depth EstimationabstractMonocular depth estimation within the diffusion-denoising paradigm demonstrates impressive generalization ability but suffers from low inference speed. Recent methods adopt a single-step deterministic paradigm to improve inference efficiency while maintaining comparable performance. However, they overlook the gap between generative and discriminative features, leading to suboptimal results. In this work, we propose DepthMaster, a single-step diffusion model designed to adapt generative features for the discriminative depth estimation task. First, to mitigate overfitting to texture details introduced by generative features, we propose a Feature Alignment module, which incorporates high-quality semantic features to enhance the denoising network's representation capability. Second, to address the lack of fine-grained details in the single-step deterministic framework, we propose a Fourier Enhancement module to adaptively balance low-frequency structure and high-frequency details. We adopt a two-stage training strategy to fully leverage the potential of the two modules. In the first stage, we focus on learning the global scene structure with the Feature Alignment module, while in the second stage, we exploit the Fourier Enhancement module to improve the visual quality. Through these efforts, our model achieves state-of-the-art performance in terms of generalization and detail preservation, outperforming other diffusion-based methods across various datasets. Our project page can be found at https://indu1ge.github.io/DepthMaster_page. Ziyang Song 0001, Zerong Wang, Bo Li 0115, Hao Zhang 0063, Ruijie Zhu 0002, Li Liu 0067, Peng-Tao Jiang, Tianzhu Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | SEMat: Semantic Enhanced Natural Image Interactive MattingabstractRecent approaches attempt to adapt powerful interactive segmentation models, such as SAM, to interactive matting and fine-tune the models based on synthetic matting datasets. However, models trained on synthetic data fail to generalize to complex and occlusion scenes. We address this challenge by proposing a new matting dataset based on the COCO dataset, namely COCO-Matting. It selects real-world complex images from COCO and converts semantic segmentation masks to matting labels. The built COCO-Matting comprises an extensive collection of 36,980 human instance-level alpha mattes in complex natural scenarios. Furthermore, existing SAM-based matting methods extract intermediate features and masks from a frozen SAM and only train a lightweight matting decoder by end-to-end matting losses, which do not fully exploit the potential of the pre-trained SAM. Thus, we propose SEMat which revamps the network architecture and training objectives. For network architecture, the proposed feature-aligned transformer learns to extract fine-grained edge and transparency features. The proposed matte-aligned decoder aims to segment matting-specific objects and convert coarse masks into high-precision mattes. For training objectives, the proposed regularization and trimap loss aim to retain the prior from the pre-trained model and push the matting logits extracted from the mask decoder to contain trimap-based semantic information. Extensive experiments across seven diverse datasets demonstrate the superior performance of our method, proving its efficacy in interactive natural image matting. Code is available at https://github.com/XiaRho/SEMat. Ruihao Xia, Peng-Tao Jiang, Hao Zhang 0063, Qianru Sun, Yang Tang 0001, Bo Li 0115, Pan Zhou 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Advancing Comprehensive Aesthetic Insight with Multi-Scale Text-Guided Self-Supervised LearningabstractImage Aesthetic Assessment (IAA) is a vital and intricate task that entails analyzing and assessing an image's aesthetic values, and identifying its highlights and areas for improvement. Traditional methods of IAA often concentrate on a single aesthetic task and suffer from inadequate labeled datasets, thus impairing in-depth aesthetic comprehension. Despite efforts to overcome this challenge through the application of Multi-modal Large Language Models (MLLMs), such models remain underdeveloped for IAA purposes. To address this, we propose a comprehensive aesthetic MLLM capable of nuanced aesthetic insight. Central to our approach is an innovative multi-scale text-guided self-supervised learning technique. This technique features a multi-scale feature alignment module and capitalizes on a wealth of unlabeled data in a self-supervised manner to structurally and functionally enhance aesthetic ability. The empirical evidence indicates that accompanied with extensive instruct-tuning, our model sets new state-of-the-art benchmarks across multiple tasks, including aesthetic scoring, aesthetic commenting, and personalized image aesthetic assessment. Remarkably, it also demonstrates zero-shot learning capabilities in the emerging task of aesthetic suggesting. Furthermore, for personalized image aesthetic assessment, we harness the potential of in-context learning and showcase its inherent advantages. Yuti Liu, Shice Liu, Junyuan Gao, Peng-Tao Jiang, Hao Zhang 0063, Jinwei Chen 0003, Bo Li 0130 |
AAAI | 5 |
| 2025 | Boosting Vision State Space Model with Fractal ScanningabstractRecently, foundational models have significantly advanced in different tasks, accompanied by Transformer as the general backbone. However, Transformer's quadratic complexity poses challenges for handling longer sequences and higher resolution images, which may limit foundational models further development. To alleviate this issue, various efficient State Space Models (SSMs) like Mamba have emerged, initially matching Transformer performance and gradually surpassing it. To improve the performance of SSMs in computer vision tasks, one crucial viewpoint is effective serialization of images. Existing vision Mambas, which rely on a linear scanning mechanism, often struggle to capture complex spatial relationships in 2D images. This results in feature loss during serialization and negatively impacts model performance. To overcome this limitation, we propose the use of fractal scanning curves for image serialization to enhance the Mambas’ ability to accurately model complex spatial dependencies. Additionally, unlike existing vision Mambas, which are designed with various curve scanning directions that increase the complexity, contradicting the original intent of Mamba to enhance model performance. We novelty introduce the Fractal Fusion Pathway (FFP) for our FractalMamba, which can enhance its performance efficiently. Extensive experiments underscore the superiority of our proposed FractalMamba. Haoke Xiao, Lv Tang, Peng-Tao Jiang, Hao Zhang 0063, Jinwei Chen 0003, Bo Li 0130 |
AAAI | 4 |
| 2025 | SDMATTE: Grafting Diffusion Models for Interactive MattingabstractRecent interactive matting methods have shown satisfactory performance in capturing the primary regions of objects, but they fall short in extracting fine-grained details in edge regions. Diffusion models trained on billions of image-text pairs, demonstrate exceptional capability in modeling highly complex data distributions and synthesizing realistic texture details, while exhibiting robust text-driven interaction capabilities, making them an attractive solution for interactive matting. To this end, we propose SDMatte, a diffusion-driven interactive matting model, with three key contributions. First, we exploit the powerful priors of diffusion models and transform the text-driven interaction capability into visual prompt-driven interaction capability to enable interactive matting. Second, we integrate coordinate embeddings of visual prompts and opacity embeddings of target objects into U-Net, enhancing SDMatte's sensitivity to spatial position information and opacity information. Third, we propose a masked self-attention mechanism that enables the model to focus on areas specified by visual prompts, leading to better performance. Extensive experiments on multiple datasets demonstrate the superior performance of our method, validating its effectiveness in interactive matting. Our code and model are available at https://github.com/vivoCameraResearch/SDMatte. Longfei Huang, Hao Zhang 0063, Jinwei Chen 0003, Lunde Chen, Peng-Tao Jiang |
ICCV | 3 |
| 2025 | Multi-Task Dense Predictions via Unleashing the Power of DiffusionabstractDiffusion models have exhibited extraordinary performance in dense prediction tasks. However, there are few works exploring the diffusion pipeline for multi-task dense predictions. In this paper, we unlock the potential of diffusion models in solving multi-task dense predictions and propose a novel diffusion-based method, called TaskDiffusion, which leverages the conditional diffusion process in the decoder. Instead of denoising the noisy labels for different tasks separately, we propose a novel joint denoising diffusion process to capture the task relations during denoising. To be specific, our method first encodes the task-specific labels into a task-integration feature space to unify the encoding strategy. This allows us to get rid of the cumbersome task-specific encoding process. In addition, we also propose a cross-task diffusion decoder conditioned on task-specific multi-level features, which can model the interactions among different tasks and levels explicitly while preserving efficiency. Experiments show that our TaskDiffusion outperforms previous state-of-the-art methods for all dense prediction tasks on the widely-used PASCAL-Context and NYUD-v2 datasets. Our code is available at https://github.com/YuqiYang213/TaskDiffusion. Peng-Tao Jiang, Qibin Hou, Hao Zhang 0063, Jinwei Chen 0003 |
ICLR | 4 |
| 2025 | High-Precision Dichotomous Image Segmentation via Probing Diffusion CapacityabstractIn the realm of high-resolution (HR), fine-grained image segmentation, the primary challenge is balancing broad contextual awareness with the precision required for detailed object delineation, capturing intricate details and the finest edges of objects. Diffusion models, trained on vast datasets comprising billions of image-text pairs, such as SD V2.1, have revolutionized text-to-image synthesis by delivering exceptional quality, fine detail resolution, and strong contextual awareness, making them an attractive solution for high-resolution image segmentation. To this end, we propose DiffDIS, a diffusion-driven segmentation model that taps into the potential of the pre-trained U-Net within diffusion models, specifically designed for high-resolution, fine-grained object segmentation. By leveraging the robust generalization capabilities and rich, versatile image representation prior of the SD models, coupled with a task-specific stable one-step denoising approach, we significantly reduce the inference time while preserving high-fidelity, detailed generation. Additionally, we introduce an auxiliary edge generation task to not only enhance the preservation of fine details of the object boundaries, but reconcile the probabilistic nature of diffusion with the deterministic demands of segmentation. With these refined strategies in place, DiffDIS serves as a rapid object mask generation model, specifically optimized for generating detailed binary maps at high resolutions, while demonstrating impressive accuracy and swift processing. Experiments on the DIS5K dataset demonstrate the superiority of DiffDIS, achieving state-of-the-art results through a streamlined inference process. The source code will be publicly available at \href{https://github.com/qianyu-dlut/DiffDIS}{DiffDIS}. Qian Yu 0015, Peng-Tao Jiang, Hao Zhang 0063, Jinwei Chen 0003, Bo Li 0115, Lihe Zhang, Huchuan Lu |
ICLR | 3 |
| 2025 | Learning Adaptive Lighting via Channel-Aware GuidanceabstractLearning lighting adaptation is a crucial step in achieving good visual perception and supporting downstream vision tasks. Current research often addresses individual light-related challenges, such as high dynamic range imaging and exposure correction, in isolation. However, we identify shared fundamental properties across these tasks: i) different color channels have different light properties, and ii) the channel differences reflected in the spatial and frequency domains are different. Leveraging these insights, we introduce the channel-aware Learning Adaptive Lighting Network (LALNet), a multi-task framework designed to handle multiple light-related tasks efficiently. Specifically, LALNet incorporates color-separated features that highlight the unique light properties of each color channel, integrated with traditional color-mixed features by Light Guided Attention (LGA). The LGA utilizes color-separated features to guide color-mixed features focusing on channel differences and ensuring visual consistency across all channels. Additionally, LALNet employs dual domain channel modulation for generating color-separated features and a mixed channel modulation and light state space module for producing color-mixed features. Extensive experiments on four representative light-related tasks demonstrate that LALNet significantly outperforms state-of-the-art methods on benchmark tests and requires fewer computational resources. We provide an anonymous online demo at LALNet. Peng-Tao Jiang, Hao Zhang 0063, Jinwei Chen 0003, Bo Li 0026, Huanjing Yue, Jing-Yu Yang 0002 |
ICML | 3 |
| 2025 | Interpretable Rotation-Equivariant Multiary-Valued Network for Attribute ObfuscationabstractThis paper focuses on the problem of preventing information leakage in neural networks, i.e., assuming that attackers have obtained intermediate-layer features of a neural network, and preventing attackers from inverting these features to the input with private information. We propose a generic method to slightly revise each arbitrary traditional neural network into a multiary-valued rotation-equivariant neural network (RENN) for preventing information leakage. Specifically, we convert real-valued features in the network into multi-ary features, and each element in the feature vector is a multi-ary number. We hide the input information into a certain phase of the multi-ary feature, and rotate the multi-ary feature for attribute obfuscation in the encryption process. The rotation axis and angle can be considered as the private key. In this way, even when attackers have obtained network parameters and intermediate-layer features, they still cannot extract input information without knowing the rotation information. More crucially, the encryption operation does not damage the spatial correlations between features, so that the encrypted features can be easily processed by convolution operations in the neural network without difficulties. In order to implement successful encryption and decryption, the RENN is designed to satisfy the rotation equivariance property. To this end, we propose a set of rules to revise classic operations in the neural network to ensure the rotation equivariance property. Besides, we prove that the $d$d-ary RENN is downward compatible with the $d^{\prime }$d'-ary RENN when $d^{\prime }< d$d' Quanshi Zhang, Hao Zhang 0063, Yiting Chen 0003, Qihan Ren, Jie Ren 0018, Xu Cheng 0005, Liyao Xiang |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | Clarifying the Behavior and the Difficulty of Adversarial TrainingabstractAdversarial training is usually difficult to optimize. This paper provides conceptual and analytic insights into the difficulty of adversarial training via a simple theoretical study, where we derive an approximate dynamics of a recursive multi-step attack in a simple setting. Despite the simplicity of our theory, it still reveals verifiable predictions about various phenomena in adversarial training under real-world settings. First, compared to vanilla training, adversarial training is more likely to boost the influence of input samples with large gradient norms in an exponential manner. Besides, adversarial training also strengthens the influence of the Hessian matrix of the loss w.r.t. network parameters, which is more likely to make network parameters oscillate and boosts the difficulty of adversarial training. Xu Cheng 0005, Hao Zhang 0063, Wen Shen 0002, Quanshi Zhang |
AAAI | 2 |
| 2024 | Explaining Generalization Power of a DNN Using Interactive ConceptsabstractThis paper explains the generalization power of a deep neural network (DNN) from the perspective of interactions. Although there is no universally accepted definition of the concepts encoded by a DNN, the sparsity of interactions in a DNN has been proved, i.e., the output score of a DNN can be well explained by a small number of interactions between input variables. In this way, to some extent, we can consider such interactions as interactive concepts encoded by the DNN. Therefore, in this paper, we derive an analytic explanation of inconsistency of concepts of different complexities. This may shed new lights on using the generalization power of concepts to explain the generalization power of the entire DNN. Besides, we discover that the DNN with stronger generalization power usually learns simple concepts more quickly and encodes fewer complex concepts. We also discover the detouring dynamics of learning complex concepts, which explains both the high learning difficulty and the low generalization power of complex concepts. The code will be released when the paper is accepted. Huilin Zhou, Hao Zhang 0063, Huiqi Deng, Dongrui Liu, Wen Shen 0002, Shih-Han Chan, Quanshi Zhang |
AAAI | 2 |
| 2024 | Multi-Task Dense Prediction via Mixture of Low-Rank ExpertsabstractPrevious multitask dense prediction methods based on the Mixture of Experts (MoE) have received great performance but they neglect the importance of explicitly modeling the global relations among all tasks. In this paper, we present a novel decoder-focused method for multitask dense prediction, called Mixture-of-Low-Rank-Experts (MLoRE). To model the global task relationships, MLoRE adds a generic convolution path to the original MoE structure, where each task feature can go through this path for explicit parameter sharing. Furthermore, to control the parameters and computational cost brought by the increase in the number of experts, we take inspiration from LoRA and propose to leverage the low-rank format of a vanilla con-volution in the expert network. Since the low-rank experts have fewer parameters and can be dynamically parameter-ized into the generic convolution, the parameters and computational cost do not change much with the increase of experts. Benefiting from this design, we increase the number of experts and its reception field to enlarge the representation capacity, facilitating multiple dense tasks learning in a unified network. Extensive experiments on the PASCAL-Context and NYUD-v2 benchmarks show that our MLoRE achieves superior performance compared to previous state-of-the-art methods on all metrics. Our code is available at https://github.com/YuqiYang213/MLoRE. Peng-Tao Jiang, Qibin Hou, Hao Zhang 0063, Jinwei Chen 0003, Bo Li 0130 |
CVPR | 4 |
| 2024 | Revisiting Single Image Reflection Removal in the WildabstractThis research focuses on the issue of single-image reflection removal (SIRR) in real-world conditions, examining it from two angles: the collection pipeline of real reflection pairs and the perception of real reflection locations. We devise an advanced reflection collection pipeline that is highly adaptable to a wide range of real-world reflection scenarios and incurs reduced costs in collecting large-scale aligned reflection pairs. In the process, we develop a large-scale, high-quality reflection dataset named Reflection Removal in the Wild (RRW). RRW contains over 14,950 high-resolution real-world reflection pairs, a dataset forty-five times larger than its predecessors. Regarding perception of reflection locations, we identify that numerous virtual reflection objects visible in reflection images are not present in the corresponding ground-truth images. This observation, drawn from the aligned pairs, leads us to conceive the Maximum Reflection Filter (MaxRF). The MaxRF could accurately and explicitly characterize reflection locations from pairs of images. Building upon this, we design a reflection location-aware cascaded framework, specifically tailored for SIRR. Powered by these innovative techniques, our solution achieves superior performance than current leading methods across multiple real-world benchmarks. Codes and datasets are available at here. Yurui Zhu, Xueyang Fu, Peng-Tao Jiang, Hao Zhang 0063, Qibin Sun, Jinwei Chen 0003, Zhengjun Zha, Bo Li 0130 |
CVPR | 4 |
| 2024 | SAFNet: Selective Alignment Fusion Network for Efficient HDR Imaging
Lingtong Kong, Bo Li 0115, Yike Xiong, Hao Zhang 0063, Jinwei Chen 0003 |
ECCV (26) | 4 |
| 2024 | Beta-Tuned Timestep Diffusion Model
Tianyi Zheng 0001, Peng-Tao Jiang, Ben Wan, Hao Zhang 0063, Jinwei Chen 0003, Jia Wang 0004, Bo Li 0115 |
ECCV (3) | 4 |
| 2024 | Improving Adversarial Energy-Based Model via Diffusion ProcessabstractGenerative models have shown strong generation ability while efficient likelihood estimation is less explored. Energy-based models (EBMs) define a flexible energy function to parameterize unnormalized densities efficiently but are notorious for being difficult to train. Adversarial EBMs introduce a generator to form a minimax training game to avoid expensive MCMC sampling used in traditional EBMs, but a noticeable gap between adversarial EBMs and other strong generative models still exists. Inspired by diffusion-based models, we embedded EBMs into each denoising step to split a long-generated process into several smaller steps. Besides, we employ a symmetric Jeffrey divergence and introduce a variational posterior distribution for the generator's training to address the main challenges that exist in adversarial EBMs. Our experiments show significant improvement in generation compared to existing adversarial EBMs, while also providing a useful energy function for efficient density estimation. Cong Geng, Tian Han 0001, Peng-Tao Jiang, Hao Zhang 0063, Jinwei Chen 0003, Søren Hauberg, Bo Li 0115 |
ICML | 4 |
| 2024 | Chain of Visual Perception: Harnessing Multimodal Large Language Models for Zero-shot Camouflaged Object DetectionabstractIn this paper, we introduce a novel multimodal camo-perceptive framework (MMCPF) aimed at handling zero-shot Camouflaged Object Detection (COD) by leveraging the powerful capabilities of Multimodal Large Language Models (MLLMs). Recognizing the inherent limitations of current COD methodologies, which predominantly rely on supervised learning models demanding extensive and accurately annotated datasets, resulting in weak generalization, our research proposes a zero-shot MMCPF that circumvents these challenges. Although MLLMs hold significant potential for broad applications, their effectiveness in COD is hindered and they would make misinterpretations of camouflaged objects. To address this challenge, we further propose a strategic enhancement called the Chain of Visual Perception (CoVP), which significantly improves the perceptual capabilities of MLLMs in camouflaged scenes by leveraging both linguistic and visual cues more effectively. We validate the effectiveness of MMCPF on five widely used COD datasets, containing CAMO, COD10K, NC4K, MoCA-Mask and OVCamo. Experiments show that MMCPF can outperform all existing state-of-the-art zero-shot COD methods, and achieve competitive performance compared to weakly-supervised and fully-supervised methods, which demonstrates the potential of MMCPF. The Github link of this paper is https://github.com/luckybird1994/MMCPF. Lv Tang, Peng-Tao Jiang, Hao Zhang 0063, Jinwei Chen 0003, Bo Li 0115 |
ACM Multimedia | 4 |
| 2024 | Non-uniform Timestep Sampling: Towards Faster Diffusion Model TrainingabstractDiffusion models have garnered significant success in generative tasks, emerging as the predominant model in this domain. Despite their success, the substantial computational resources required for training diffusion models restrict their practical applications. In this paper, we resort to the optimal transport theory to accelerate the training of diffusion models, providing an in-depth analysis of the forward diffusion process. It shows that the upper bound on the Wasserstein distance of the distribution between any two timesteps in the diffusion process is an exponential decrease of the initial distance by a factor of times. This finding suggests that the state distribution of the diffusion model has a non-uniform rate of change at different points in time, thus highlighting the different importance of the diffusion timestep. To this end, we propose a novel non-uniform timestep sampling method based on the Bernoulli distribution, which favors more frequent sampling in significant timestep intervals. The key idea is to make the model focus on timesteps with larger differences, thus accelerating the training of the diffusion model. Experiments on benchmark datasets reveal that the proposed method significantly reduces the computational overhead while improving the quality of the generated images. Tianyi Zheng 0001, Cong Geng, Peng-Tao Jiang, Ben Wan, Hao Zhang 0063, Jinwei Chen 0003, Jia Wang 0004, Bo Li 0115 |
ACM Multimedia | 5 |
| 2024 | Crafter: Facial Feature Crafting against Inversion-based Identity Theft on Deep Models
Liyao Xiang, Hao Zhang 0063, Xinbing Wang, Chenghu Zhou, Bo Li 0115 |
NDSS | 4 |
| 2024 | Unsupervised Modality Adaptation with Text-to-Image Diffusion Models for Semantic SegmentationabstractDespite their success, unsupervised domain adaptation methods for semantic segmentation primarily focus on adaptation between image domains and do not utilize other abundant visual modalities like depth, infrared and event. This limitation hinders their performance and restricts their application in real-world multimodal scenarios. To address this issue, we propose Modality Adaptation with text-to-image Diffusion Models (MADM) for semantic segmentation task which utilizes text-to-image diffusion models pre-trained on extensive image-text pairs to enhance the model's cross-modality capabilities. Specifically, MADM comprises two key complementary components to tackle major challenges. First, due to the large modality gap, using one modal data to generate pseudo labels for another modality suffers from a significant drop in accuracy. To address this, MADM designs diffusion-based pseudo-label generation which adds latent noise to stabilize pseudo-labels and enhance label accuracy. Second, to overcome the limitations of latent low-resolution features in diffusion models, MADM introduces the label palette and latent regression which converts one-hot encoded labels into the RGB form by palette and regresses them in the latent space, thus ensuring the pre-trained decoder for up-sampling to obtain fine-grained features. Extensive experimental results demonstrate that MADM achieves state-of-the-art adaptation performance across various modality tasks, including images to depth, infrared, and event modalities. We open-source our code and models at https://github.com/XiaRho/MADM. Ruihao Xia, Peng-Tao Jiang, Hao Zhang 0063, Bo Li 0130, Yang Tang 0001, Pan Zhou 0002 |
NeurIPS | 4 |
| 2022 | Discovering and Explaining the Representation Bottleneck of DNNS
Huiqi Deng, Qihan Ren, Hao Zhang 0063, Quanshi Zhang |
ICLR | 3 |
| 2022 | Quantification and Analysis of Layer-wise and Pixel-wise Information DiscardingabstractThis paper presents a method to explain how the information of each input variable is gradually discarded during the forward propagation in a deep neural network (DNN), which provides new perspectives to explain DNNs. We define two types of entropy-based metrics, i.e. (1) the discarding of pixel-wise information used in the forward propagation, and (2) the uncertainty of the input reconstruction, to measure input information contained by a specific layer from two perspectives. Unlike previous attribution metrics, the proposed metrics ensure the fairness of comparisons between different layers of different DNNs. We can use these metrics to analyze the efficiency of information processing in DNNs, which exhibits strong connections to the performance of DNNs. We analyze information discarding in a pixel-wise manner, which is different from the information bottleneck theory measuring feature information w.r.t. the sample distribution. Experiments have shown the effectiveness of our metrics in analyzing classic DNNs and explaining existing deep-learning techniques. The code is available at https://github.com/haotianSustc/deepinfo. Hao Zhang 0063, Yinqing Zhang, Quanshi Zhang |
ICML | 2 |
| 2021 | Interpreting Multivariate Shapley Interactions in DNNsabstractThis paper aims to explain deep neural networks (DNNs) from the perspective of multivariate interactions. In this paper, we define and quantify the significance of interactions among multiple input variables of the DNN. Input variables with strong interactions usually form a coalition and reflect prototype features, which are memorized and used by the DNN for inference. We define the significance of interactions based on the Shapley value, which is designed to assign the attribution value of each input variable to the inference. We have conducted experiments with various DNNs. Experimental results have demonstrated the effectiveness of the proposed method. Hao Zhang 0063, Yichen Xie 0002, Longjie Zheng, Die Zhang, Quanshi Zhang |
AAAI | 1 |
| 2021 | Building Interpretable Interaction Trees for Deep NLP ModelsabstractThis paper proposes a method to disentangle and quantify interactions among words that are encoded inside a DNN for natural language processing. We construct a tree to encode salient interactions extracted by the DNN. Six metrics are proposed to analyze properties of interactions between constituents in a sentence. The interaction is defined based on Shapley values of words, which are considered as an unbiased estimation of word contributions to the network prediction. Our method is used to quantify word interactions encoded inside the BERT, ELMo, LSTM, CNN, and Transformer networks. Experimental results have provided a new perspective to understand these DNNs, and have demonstrated the effectiveness of our method. Die Zhang, Hao Zhang 0063, Huilin Zhou, Xiaoyi Bao, Da Huo 0002, Ruizhao Chen, Xu Cheng 0005, Mengyue Wu, Quanshi Zhang |
AAAI | 2 |
| 2021 | Interpreting Attributions and Interactions of Adversarial AttacksabstractThis paper aims to explain adversarial attacks in terms of how adversarial perturbations contribute to the attacking task. We estimate attributions of different image regions to the decrease of the attacking cost based on the Shapley value. We define and quantify interactions among adversarial perturbation pixels, and decompose the entire perturbation map into relatively independent perturbation components. The decomposition of the perturbation map shows that adversarially-trained DNNs have more perturbation components in the foreground than normally-trained DNNs. Moreover, compared to the normally-trained DNN, the adversarially-trained DNN have more components which mainly decrease the score of the true category. Above analyses provide new insights into the understanding of adversarial attacks. Xin Wang 0108, Shuyun Lin, Hao Zhang 0063, Quanshi Zhang |
ICCV | 3 |
| 2021 | Interpreting and Boosting Dropout from a Game-Theoretic View
Hao Zhang 0063, Yinchao Ma, Yichen Xie 0002, Quanshi Zhang |
ICLR | 1 |
| 2020 | Interpretable Complex-Valued Neural Networks for Privacy Protection
Liyao Xiang, Hao Zhang 0063, Jie Ren 0018, Quanshi Zhang |
ICLR | 2 |
| 2018 | Mining deep And-Or object structures via cost-sensitive question-answer-based active annotations
Quanshi Zhang, Ying Nian Wu, Hao Zhang 0063, Song-Chun Zhu |
Comput. Vis. Image Underst. | 3 |