EDBT 2026 Demo / reviewers in the wild / expert
Chaoning Zhang
dblp:263/2846
· DBLP profile ↗
51ranked-venue papers
16as first author
45since 2021 · last 2027
0000-0001-6007-6099ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 33 · 14 first-author · 28 since 2021Artificial intelligence and machine learning · 27 · 12 first-author · 22 since 2021Computer networks · 4 · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Pseudo-image spatiotemporal augmentation for cellular network traffic prediction
Maryam Qamar, Tahir Khalil, Md. Atikuzzaman, Chaoning Zhang, Sung-Ho Bae |
Expert Syst. Appl. | 4 |
| 2026 | From Dialogue to Destination: Geography-Aware Large Language Models with Multimodal Fusion for Conversational RecommendationabstractConversational Recommender Systems (CRS) aim to provide personalized recommendations by interacting with users through natural language dialogue. However, in scenarios requiring deep geospatial awareness, existing methods, including those based on Large Language Models (LLMs), still face significant challenges in effectively fusing heterogeneous, multimodal geographic information with dynamic dialogue context. Simple fusion strategies struggle to resolve the asymmetric dependencies between dynamic user intent and static geographic context and fail to bridge the semantic gap between LLMs and structured geospatial data. To address these issues, we propose a framework for geography-aware CRS, named GeoCRS. Our core idea is to empower a frozen LLM with powerful geospatial reasoning capabilities by conditioning it on a dynamic, multimodal guidance signal generated by an external fusion architecture, all without altering the LLM's internal parameters. Specifically, we first design a hierarchical geographical encoder to uniformly represent heterogeneous geographic data. Subsequently, we introduce a contextual feature modulation module that asymmetrically injects the geographic context into the user's dialogue intent via a novel modulation mechanism to improve conversational recommendation via both geographic and dialogue context. Extensive experiments on public benchmark datasets demonstrate that our proposed GeoCRS significantly outperforms state-of-the-art baselines on the geography-aware conversational recommendation task. Chenxi Liu 0003, Jie Zou 0001, Cheng Long 0001, Chaoning Zhang, Peng Wang 0023, Yang Yang 0002 |
AAAI | 5 |
| 2026 | Lightweight LLM Agent Memory with Small Language ModelsabstractJiaquan Zhang, Chaoning Zhang, Shuxu Chen, Zhenzhen Huang, Pengcheng Zheng, Zhicheng Wang, Ping Guo, Fan Mo, Sung-Ho Bae, Jie Zou, Jiwei Wei, Yang Yang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Jiaquan Zhang, Chaoning Zhang, Shuxu Chen, Zhenzhen Huang, Sung-Ho Bae, Jie Zou 0001, Jiwei Wei, Yang Yang 0002 |
ACL (1) | 2 |
| 2026 | Efficient dataset condensation with learnable color representation and subset matching
Linh-Tam Tran, Quang Hieu Vo, Maryam Qamar, Chaoning Zhang, Hui Yong Kim, Sung-Ho Bae |
Knowl. Based Syst. | 4 |
| 2026 | Internal-external boundary attention fusion for glass surface segmentation
Dongshen Han, Heechan Yoon, Hyukmin Kwon, Hyun-Cheol Kim, Hyon-Gon Choo, Seungkyu Lee 0001, Chaoning Zhang |
Neural Networks | 7 |
| 2026 | Compression in 3D Gaussian Splatting: A Survey of Methods, Trends, and Future Directionsabstract3D Gaussian Splatting (3DGS) has recently emerged as a pioneering approach in explicit scene rendering and computer graphics. Unlike traditional neural radiance field (NeRF) methods, which typically rely on implicit, coordinate-based models to map spatial coordinates to pixel values, 3DGS utilizes millions of learnable 3D Gaussians. Its differentiable rendering technique and inherent capability for explicit scene representation and manipulation positions 3DGS as a potential game-changer for the next generation of 3D reconstruction and representation technologies. This enables 3DGS to deliver real-time rendering speeds while offering unparalleled editability levels. However, despite its advantages, 3DGS suffers from substantial memory and storage requirements, posing challenges for deployment on resource-constrained devices. In this survey, we provide a comprehensive overview focusing on the scalability and compression of 3DGS. We begin with a detailed background overview of 3DGS, followed by a structured taxonomy of existing compression methods. Additionally, we analyze and compare current methods from the topological perspective, evaluating their strengths and limitations in terms of fidelity, compression ratios, and computational efficiency. Furthermore, we explore how advancements in efficient NeRF representations can inspire future developments in 3DGS optimization. Finally, we conclude with current research challenges and highlight key directions for future exploration. Muhammad Salman Ali, Chaoning Zhang, Marco Cagnazzo, Giuseppe Valenzise, Enzo Tartaglione, Sung-Ho Bae |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | Practical No-Box Adversarial Attacks With Training-Free Hybrid Image TransformationabstractRecently, the adversarial vulnerability of deep neu ral networks (DNNs) has raised increasing attention. Among all the threat models, no-box attacks are the most practical but extremely challenging since they neither rely on any knowledge of the target model or similar substitute model, nor access the dataset for training a new substitute model. Although a recent method has attempted such an attack in a loose sense, its performance is not good enough and computational overhead of training is expensive. In this paper, we move a step forward and show the existence of a training-free adversarial perturbation under the no-box threat model, which can be successfully used to attack different DNNs in real-time. Motivated by our observation that high-frequency component (HFC) is dominant in low-level features and plays a crucial role in classification, we attack an image mainly by suppression of the original HFC and adding of noisy HFC. We empirically and experimentally analyze the requirements of effective noisy HFC and show that it should be regionally homogeneous, repeating and dense. Remarkably, on ImageNet dataset, our method attacks ten well-known models with a success rate of 98.13% on average, which outperforms state-of-the-art no-box attacks by 6.41%. Furthermore, our method is even competitive to mainstream transfer-based black box attacks. Our code is publicly available1 Youheng Sun, Chaoning Zhang, Chaoqun Li 0007, Xuanhan Wang, Jingkuan Song, Lianli Gao |
IEEE Trans. Multim. | 3 |
| 2025 | KDA: Knowledge Diffusion Alignment with Enhanced Context for Video Temporal Grounding
Ran Ran 0001, Jiwei Wei, Shiyuan He, Zeyu Ma 0002, Chaoning Zhang, Ning Xie 0003, Yang Yang 0002 |
ICCV | 5 |
| 2025 | SyncGaussian: Stable 3D Gaussian-Based Talking Head Generation with Enhanced Lip Sync via Discriminative Speech FeaturesabstractGenerating high-fidelity talking heads that maintain stable head poses and achieve robust lip sync remains a significant challenge. Although methods based on 3D Gaussian Splatting (3DGS) offer a promising solution via point-based deformation, they suffer from inconsistent head dynamics and mismatched mouth movements due to unstable Gaussian initialization and incomplete speech features. To overcome these limitations, we introduce SyncGaussian, a 3DGS-based framework that ensures stable head poses, enhanced lip sync, and realistic appearances with real-time rendering. SyncGaussian employs a stable head Gaussian initialization strategy to mitigate head jitter by optimizing commonly used rough head pose parameters. To enhance lip sync, we propose a sync-enhanced encoder that leverages audio-to-text and audio-to-visual speech features. Guided by a tailored cosine similarity loss function, the encoder integrates discriminative speech features through a multi-level sync adaptation mechanism, enabling the learning of an adaptive speech feature space. Extensive experiments demonstrate that SyncGaussian outperforms state-of-the-art methods in image quality, dynamic motion, and lip sync, with the potential for real-time applications. Jiwei Wei, Shiyuan He, Zeyu Ma 0002, Chaoning Zhang, Ning Xie 0003, Yang Yang 0002 |
IJCAI | 5 |
| 2025 | InteractGuide: LLM-Enhanced Multimodal Reasoning for User-Centric Interaction Recommendations in AR-HRI AuthoringabstractAugmented Reality (AR) enhances Human-Robot Interaction (HRI) by offering diverse interaction methods. However, existing systems often fail to resolve the conflict between a user's implicit preferences and physical ergonomics, leading to suboptimal experiences. We introduce InteractGuide, a novel framework that, for the first time, uses a Large Language Model (LLM) as a central reasoning engine to dynamically balance these competing factors. Our system translates physiological signals into a symbolic ''Preference Memory'' that the LLM reasons over, alongside real-time ergonomic and contextual data, to provide personalized interaction recommendations. A 29-participant study confirms our architecture improves efficiency and experience compared to single-factor approaches, showing the potential of LLMs as reasoning engines for complex AR-HRI. This work presents a validated end-to-end architecture for user-centric interaction adaptation, demonstrating the potential of LLMs as reasoning engines in complex AR-HRI systems. Yunqiang Pei, Hongrong Yang, Guoqing Wang 0001, Peng Wang 0023, Chaoning Zhang, Yang Yang 0002, Heng Tao Shen |
ACM Multimedia | 6 |
| 2025 | Beyond Whole Dialogue Modeling: Contextual Disentanglement for Conversational RecommendationabstractConversational recommender systems aim to provide personalized recommendations by analyzing and utilizing contextual information related to dialogue. However, existing methods typically model the dialogue context as a whole, neglecting the inherent complexity and entanglement within the dialogue. Specifically, a dialogue comprises both focus information and background information, which mutually influence each other. Current methods tend to model these two types of information mixedly, leading to misinterpretation of users' actual needs, thereby lowering the accuracy of recommendations. To address this issue, this paper proposes a novel model to introduce contextual disentanglement for improving conversational recommender systems, named DisenCRS. The proposed model DisenCRS employs a dual disentanglement framework, including self-supervised contrastive disentanglement and counterfactual inference disentanglement, to effectively distinguish focus information and background information from the dialogue context under unsupervised conditions. Moreover, we design an adaptive prompt learning module to automatically select the most suitable prompt based on the specific dialogue context, fully leveraging the power of large language models. Experimental results on two widely used public datasets demonstrate that DisenCRS significantly outperforms existing conversational recommendation models, achieving superior performance on both item recommendation and response generation tasks. Guojia An, Jie Zou 0001, Jiwei Wei, Chaoning Zhang, Fuming Sun, Yang Yang 0002 |
SIGIR | 4 |
| 2025 | Adapting lightweight SAM with gradient map for mirror object segmentation
Dongshen Han, Chaoning Zhang, Fachrina Dewi Puspitasari, Shuxu Chen, Feng Qiao 0001, Sungyoung Lee 0001, Choong Seon Hong, Yang Yang 0002 |
Inf. Sci. | 2 |
| 2025 | FedMEKT: Distillation-based embedding knowledge transfer for multimodal federated learning
Huy Q. Le, Minh N. H. Nguyen, Chu Myaet Thwal, Yu Qiao 0004, Chaoning Zhang, Choong Seon Hong |
Neural Networks | 5 |
| 2025 | Mix-Based Training Strategies for Learning Implicit Neural RepresentationsabstractWith coordinates as the input and RGB pixel values as the output, a neural network can be used to represent an image, which is widely known as Implicit neural representations (INRs). Previous works on INR have mainly focused on learning an invariant image target without exploring the impact of learning strategies on learning INR. It is observed that there is a substantial variation in PSNR among different images, and our preliminary investigation shows that, in the early training stage, learning complex image content yields significantly better performance than simple image content. Inspired by this finding, we conjecture that increasing INR task complexity in the early stage of training might boost INR performance and thus propose to intentionally contaminate the target image with another complex image. Our proposed method is called Mix-INR, which adopts a two-stage training to first learn a pseudo-target image (contaminated target) and then learn the real-target image (uncontaminated target). To generate the pseudo-target image, we experiment with two contamination methods (blending and replacement), both of which show superior performance and verify our conjecture. INRs have gained popularity as a promising approach for representing a variety of data types, including images of the task complexity of the pseudo-target image, we set the contamination image from a complex natural image to a random-noise image. Moreover, we propose a dynamic contamination method to smoothly transition from the pseudo-target image to the real-target image. Experimental results demonstrate that our proposed method achieves competitive performance, which suggests that INR can be improved by manipulating the task complexity in the early stage of training. Dongshen Han, Chaoning Zhang, Fachrina Dewi Puspitasari, Yang Yang 0002, Heng Tao Shen |
IEEE Trans. Multim. | 2 |
| 2025 | Black-Box Targeted Adversarial Attack on Segment Anything (SAM)abstractDeep recognition models are widely vulnerable to adversarial examples, which change the model output by adding quasi-imperceptible perturbation to the image input. Recently, Segment Anything Model (SAM) has emerged to become a popular foundation model in computer vision due to its impressive generalization to unseen data and tasks. Realizing flexible attacks on SAM is beneficial for understanding the robustness of SAM in the adversarial context. To this end, this work aims to achieve a targeted adversarial attack (TAA) on SAM. Specifically, under a specific prompt, the goal is to make the predicted mask of an adversarial example resemble that of a given target image. The task of TAA on SAM has been realized in the white-box setup by assuming access topromptandmodel, which is thus less practical. To address the issue of prompt dependence, we propose a simple yet effective approach by only attacking the image encoder. Moreover, we propose a novel regularization loss to enhance the cross-model transferability by increasing the feature dominance of adversarial images over random natural images. Extensive experiments verify the effectiveness of our proposed method to conduct a successful black-box TAA on SAM. Chaoning Zhang, Xinhong Hao |
IEEE Trans. Multim. | 2 |
| 2025 | Exploring Kernel Transformations for Implicit Neural RepresentationsabstractImplicit neural representations (INRs), which leverage neural networks to represent signals by mapping coordinates to their corresponding attributes, have garnered significant attention. They are extensively utilized for image representation, with pixel coordinates as input and pixel values as output. In contrast to prior works focusing on investigating the effect of the model's inside components (activation function, for instance), this work pioneers the exploration of the effect of kernel transformation of input/output while keeping the model itself unchanged. A byproduct of our findings is a simple yet effective method that combines scale and shift to significantly boost INR with negligible computation overhead. Moreover, we present two perspectives, depth and normalization, to interpret the performance benefits caused by scale and shift transformation. Overall, our work provides a new avenue for future works to understand and improve INR through the lens of kernel transformation. Chaoning Zhang, Dongshen Han, Fachrina Dewi Puspitasari, Xinhong Hao, Yang Yang 0002, Heng Tao Shen |
IEEE Trans. Multim. | 2 |
| 2025 | Cyber Attacks Prevention Toward Prosumer-Based EV Charging Stations: An Edge-Assisted Federated Prototype Knowledge Distillation ApproachabstractIn this paper, cyber-attack prevention for the prosumer-based electric vehicle (EV) charging stations (EVCSs) is investigated, which covers two aspects: 1) cyber-attack detection on prosumers’ network traffic (NT) data, and 2) cyber-attack intervention. To establish an effective prevention mechanism, several challenges need to be tackled, for instance, the NT data per prosumer may be non-independent and identically distributed (non-IID), and the boundary between benign and malicious traffic becomes blurred. To this end, we propose an edge-assisted federated prototype knowledge distillation (E-FPKD) approach, where each client is deployed on a dedicated local edge server (DLES) and can report its availability for joining the federated learning (FL) process. Prior to the E-FPKD approach, to enhance accuracy, the Pearson Correlation Coefficient is adopted for feature selection. Regarding the proposed E-FPKD approach, we integrate the knowledge distillation and prototype aggregation technique into FL to deal with the non-IID challenge. To address the boundary issue, instead of directly calculating the distance between benign and malicious traffic, we consider maximizing the overall detection correctness of all prosumers (ODC), which can mitigate the computational cost compared with the former way. After detection, a rule-based method will be triggered at each DLES for cyber-attack intervention. Experimental analysis demonstrates that the proposed E-FPKD can achieve the largest ODC on NSL-KDD, UNSW-NB15, and IoTID20 datasets in both binary and multi-class classification, compared with baselines. For instance, the ODC for IoTID20 obtained via the proposed method is separately 0.3782% and 4.4471% greater than FedProto and FedAU in multi-class classification. Luyao Zou, Quang Hieu Vo, Kitae Kim 0001, Huy Q. Le, Chu Myaet Thwal, Chaoning Zhang, Choong Seon Hong |
IEEE Trans. Netw. Serv. Manag. | 6 |
| 2024 | Knowledge Distillation Assisted Robust Federated Learning: Towards Edge IntelligenceabstractFederated learning (FL) makes it possible to advance towards edge intelligence by enabling collaborative and privacy-preserving model training across distributed edge devices. One of the main challenges in FL is non-IID (not Independent and Identically Distributed) nature of data distribution across edge devices, which results in inconsistent update directions of local and global models, thus hindering model convergence. Moreover, recent studies have shown that FL models can significantly degrade performance under adversarial attacks, which further poses challenges for deployment at edge sides. In this work, we attempt to improve the robustness of FL model under adversarial attacks in non-IID settings by sharing knowledge between a central server and edge devices via knowledge distillation. Specifically, we propose a new knowledge distillation-based federated adversarial training (FAT) framework, termed FedAdv (Federated Adversarial), which involves an edge server collecting global prototypes by aggregating local prototypes obtained from participating devices after adversarial training (AT). These global prototypes are subsequently distributed to the edge devices for regularization. This regularization mechanism aims to encourage each device to align its local representation with the corresponding global prototype. By doing so, it helps prevent significant deviations of local model updates from the global model. Experimental results on MNIST and Fashion-MNIST show that our strategy yields comparable or superior performance gains in both natural and robust accuracy compared to several baselines. Yu Qiao 0004, Apurba Adhikary, Kitae Kim 0001, Chaoning Zhang, Choong Seon Hong |
ICC | 4 |
| 2024 | Towards Robust Federated Learning via Logits Calibration on Non-IID DataabstractFederated learning (FL) is a privacy-preserving distributed management framework based on collaborative model training of distributed devices in edge networks. However, recent studies have shown that FL is vulnerable to adversarial examples (AEs), leading to a significant drop in its performance. Meanwhile, the non-independent and identically distributed (non-IID) challenge of data distribution between edge devices can further degrade the performance of models. Consequently, both AEs and non-IID pose challenges to deploying robust learning models at the edge. In this work, we adopt the adversarial training (AT) framework to improve the robustness of FL models against adversarial example (AE) attacks, which can be termed as federated adversarial training (FAT). Moreover, we address the non-IID challenge by implementing a simple yet effective logits calibration strategy under the FAT framework, which can enhance the robustness of models when subjected to adversarial attacks. Specifically, we employ a direct strategy to adjust the logits output by assigning higher weights to classes with small samples during training. This approach effectively tackles the class imbalance in the training data, with the goal of mitigating biases between local and global models. Experimental results on three dataset benchmarks, MNIST, Fashion-MNIST, and CIFAR-10 show that our strategy achieves competitive results in natural and robust accuracy compared to several baselines. Yu Qiao 0004, Apurba Adhikary, Chaoning Zhang, Choong Seon Hong |
NOMS | 3 |
| 2024 | CDKT-FL: Cross-device knowledge transfer using proxy dataset in federated learning
Huy Q. Le, Minh N. H. Nguyen, Shashi Raj Pandey, Chaoning Zhang, Choong Seon Hong |
Eng. Appl. Artif. Intell. | 4 |
| 2024 | MP-FedCL: Multiprototype Federated Contrastive Learning for Edge IntelligenceabstractFederated learning-assisted edge intelligence enables privacy protection in modern intelligent services. However, not independent and identically distributed (non-IID) distribution among edge clients can impair the local model performance. The existing single prototype-based strategy represents a class by using the mean of the feature space. However, feature spaces are usually not clustered, and a single prototype may not represent a class well. Motivated by this, this article proposes a multiprototype federated contrastive learning approach (MP-FedCL) which demonstrates the effectiveness of using a multiprototype strategy over a single-prototype under non-IID settings, including both label and feature skewness. Specifically, a multiprototype computation strategy based on k-means is first proposed to capture different embedding representations for each class space, using multiple prototypes$(k$centroids) to represent a class in the embedding space. In each global round, the computed multiple prototypes and their respective model parameters are sent to the edge server for aggregation into a global prototype pool, which is then sent back to all clients to guide their local training. Finally, local training for each client minimizes their own supervised learning tasks and learns from shared prototypes in the global prototype pool through supervised contrastive learning, which encourages them to learn knowledge related to their own class from others and reduces the absorption of unrelated knowledge in each global iteration. Experimental results on MNIST, Digit-5, Office-10, and DomainNet show that our method outperforms multiple baselines, with an average test accuracy improvement of about 4.6% and 10.4% under feature and label non-IID distributions, respectively. Yu Qiao 0004, Md. Shirajum Munir, Apurba Adhikary, Huy Q. Le, Avi Deb Raha, Chaoning Zhang, Choong Seon Hong |
IEEE Internet Things J. | 6 |
| 2024 | Toward a deeper understanding: RetNet viewed through Convolution
Chaoning Zhang |
Pattern Recognit. | 2 |
| 2024 | Boosting Adversarial Training with Hardness-Guided Attack StrategyabstractThe susceptibility of deep neural networks (DNNs) to adversarial examples has raised significant concerns regarding the security and reliability of artificial intelligence systems. These examples contain maliciously crafted perturbations not perceptible to the human eye but can cause the model to make wrong predictions. Adversarial training (AT) is the de facto standard method for enhancing adversarial robustness. However, the improved robustness is often at the cost of a significant drop in standard accuracy for clean samples. Numerous works have attempted to alleviate this trade-off by identifying its causes. A key factor lies in the variability of clean samples, which leads to different adversarial examples being generated using the same attack strategy. The other factor is the disruption of the underlying data structure caused by adversarial perturbations. To overcome these challenges, we propose a novel adversarial training framework named Hardness-Guided Sample-Dependent Adversarial Training (HGSD-AT), which dynamically adjusts the attack strategy based on the hardness of the current adversarial sample to further improve the robustness of the model. By utilizing the two types of constraints which construct from a temporal perspective and spatial distribution perspective, our method directly learns the impact of attack methods on the model, rather than the indirect effects associated with sample distribution. This approach aims to improve the generation of adversarial examples while simultaneously enhancing the robustness and accuracy of DNNs. Our approach exhibits superior performance in terms of both robustness and natural accuracy compared to state-of-the-art defense methods, as validated through comprehensive experiments conducted on three benchmark datasets. Shiyuan He, Jiwei Wei, Chaoning Zhang, Xing Xu 0001, Jingkuan Song, Yang Yang 0002, Heng Tao Shen |
IEEE Trans. Multim. | 3 |
| 2024 | Adaptive Multi-scale Degradation-Based Attack for Boosting the Adversarial TransferabilityabstractThe vulnerability of deep neural networks to adversarial examples has raised huge concerns about the security of these algorithms. Black-box adversarial attacks have received a lot of attention as an influential method for evaluating model robustness. While various sophisticated adversarial attack methods have been proposed, the success rate in the black-box scenario still needs to be improved. To address these issues, we develop an Adaptive Multi-scale Degradation-based Attack method calledAMDA. The intuitive motivation behind our approach is that different models tend to have similar attention regions for low-scale images. Specifically, AMDA uses degraded images to generate perturbations at different scales and fuses these perturbations to generate adversarial examples that are insensitive to model changes. Furthermore, we design an adaptive multi-scale perturbation fusion that evaluates the transferability of perturbations at different scales based on noise and adaptively allocates fusion weights to prioritize strong transferability attacks and avoid being compromised by local optima. Extensive experimental results on the ImageNet, CIFAR-100, and CIFAR-10 datasets demonstrate that the proposed AMDA algorithm exhibits competitive performance for both normally trained models and defense models. Ran Ran 0001, Jiwei Wei, Chaoning Zhang, Guoqing Wang 0001, Yang Yang 0002, Heng Tao Shen |
IEEE Trans. Multim. | 3 |
| 2023 | Knowledge Distillation in Federated Learning: Where and How to Distill?
Yu Qiao 0004, Chaoning Zhang, Huy Q. Le, Avi Deb Raha, Apurba Adhikary, Choong Seon Hong |
APNOMS | 2 |
| 2023 | A Survey on Masked Autoencoder for Visual Self-supervised LearningabstractWith the increasing popularity of masked autoencoders, self-supervised learning (SSL) in vision undertakes a similar trajectory as in NLP. Specifically, generative pretext tasks with the masked prediction have become a de facto standard SSL practice in NLP (e.g., BERT). By contrast, early attempts at generative methods in vision have been outperformed by their discriminative counterparts (like contrastive learning). However, the success of masked image modeling has revived the autoencoder-based visual pretraining method. As a milestone to bridge the gap with BERT in NLP, masked autoencoder in vision has attracted unprecedented attention. This work conducts a survey on masked autoencoders for visual SSL. Chaoning Zhang, Chenshuang Zhang, Junha Song, John Seon Keun Yi, In-So Kweon |
IJCAI | 1 |
| 2023 | Simple Techniques are Sufficient for Boosting Adversarial TransferabilityabstractTransferable targeted adversarial attack against deep image classifiers has remained an open issue. Depending on the space to optimize the loss, the existing methods can be divided into two categories: (a) feature space attack and (b) output space attack. The feature space attack outperforms output space one by a large margin but at the cost of requiring the training of layer-wise auxiliary classifiers for each corresponding target class together with the greedy search for the optimal layers. In this work, we revisit the method of output space attack and improve it from two perspectives. First, we identify over-fitting as one major factor that hinders transferability, for which we propose to augment the network input and/or feature layers with noise. Second, we propose a new cross-entropy loss with two ends: one for pushing the sample far from the source class, i.e. ground-truth class, and the other for pulling it close to the target class. We demonstrate that simple techniques are sufficient enough for achieving very competitive performance. Chaoning Zhang, Philipp Benz, Adil Karjauv, In-So Kweon, Choong Seon Hong |
ACM Multimedia | 1 |
| 2023 | Towards Efficient Image Compression Without Autoregressive ModelsabstractRecently, learned image compression (LIC) has garnered increasing interest with its rapidly improving performance surpassing conventional codecs. A key ingredient of LIC is a hyperprior-based entropy model, where the underlying joint probability of the latent image features is modeled as a product of Gaussian distributions from each latent element. Since latents from the actual images are not spatially independent, autoregressive (AR) context based entropy models were proposed to handle the discrepancy between the assumed distribution and the actual distribution. Though the AR-based models have proven effective, the computational complexity is significantly increased due to the inherent sequential nature of the algorithm.
In this paper, we present a novel alternative to the AR-based approach that can provide a significantly better trade-off between performance and complexity. To minimize the discrepancy, we introduce a correlation loss that forces the latents to be spatially decorrelated and better fitted to the independent probability model. Our correlation loss is proved to act as a general plug-in for the hyperprior (HP) based learned image compression methods. The performance gain from our correlation loss is ‘free’ in terms of computation complexity for both inference time and decoding time. To our knowledge, our method gives the best trade-off between the complexity and performance: combined with the Checkerboard-CM, it attains **90%** and when combined with ChARM-CM, it attains **98%** of the AR-based BD-Rate gains yet is around **50 times** and **30 times** faster than AR-based methods respectively Muhammad Salman Ali, Yeongwoong Kim, Maryam Qamar, Sung-Chang Lim, Donghyun Kim 0017, Chaoning Zhang, Sung-Ho Bae, Hui Yong Kim |
NeurIPS | 6 |
| 2023 | Efficient deep-narrow residual networks using dilated pooling for scene recognition
Zhinan Qiao, Xiaohui Yuan 0001, Runmei Zhang, Chaoning Zhang |
Expert Syst. Appl. | 5 |
| 2022 | Investigating Top-k White-Box and Transferable Black-box AttackabstractExisting works have identified the limitation of top-1 attack success rate (ASR) as a metric to evaluate the attack strength but exclusively investigated it in the white-box setting, while our work extends it to a more practical black-box setting: transferable attack. It is widely reported that stronger I-FGSM transfers worse than simple FGSM, leading to a popular belief that transferability is at odds with the white-box attack strength. Our work challenges this belief with empirical finding that stronger attack actually transfers better for the general top-k ASR indicated by the interest class rank (ICR) after attack. For increasing the attack strength, with an intuitive analysis on the logit gradient from the geometric perspective, we identify that the weakness of the commonly used losses lie in prioritizing the speed to fool the network instead of maximizing its strength. To this end, we propose a new normalized CE loss that guides the logit to be updated in the direction of implicitly maximizing its rank distance from the ground-truth class. Extensive results in various settings have verified that our proposed new loss is simple yet effective for top-k attack. Code is available at: https://bit.ly/3uCiomP Chaoning Zhang, Philipp Benz, Adil Karjauv, Jae-Won Cho, Kang Zhang 0008, In-So Kweon |
CVPR | 1 |
| 2022 | Dual Temperature Helps Contrastive Learning Without Many Negative Samples: Towards Understanding and Simplifying MoCoabstractContrastive learning (CL) is widely known to require many negative samples, 65536 in MoCo for instance, for which the performance of a dictionary-free framework is often inferior because the negative sample size (NSS) is limited by its mini-batch size (MBS). To decouple the NSS from the MBS, a dynamic dictionary has been adopted in a large volume of CL frameworks, among which arguably the most popular one is MoCo family. In essence, MoCo adopts a momentum-based queue dictionary, for which we perform a fine-grained analysis of its size and consistency. We point out that InfoNCE loss used in MoCo implicitly attract anchors to their corresponding positive sample with various strength of penalties and identify such inter-anchor hardness-awareness property as a major reason for the necessity of a large dictionary. Our findings motivate us to simplify MoCo v2 via the removal of its dictionary as well as momentum. Based on an InfoNCE with the proposed dual temperature, our simplified frameworks, Sim-MoCo and SimCo, outperform MoCo v2 by a visible margin. Moreover, our work bridges the gap between CL and non-CL frameworks, contributing to a more unified under-standing of these two mainstream frameworks in SSL. Code is available at: https://bit.ly/3LkQbaT. Chaoning Zhang, Kang Zhang 0008, Trung X. Pham, Axi Niu, Zhinan Qiao, Chang Dong Yoo, In-So Kweon |
CVPR | 1 |
| 2022 | Decoupled Adversarial Contrastive Learning for Self-supervised Adversarial Robustness
Chaoning Zhang, Kang Zhang 0008, Chenshuang Zhang, Axi Niu, Jiu Feng, Chang Dong Yoo, In-So Kweon |
ECCV (30) | 1 |
| 2022 | How Does SimSiam Avoid Collapse Without Negative Samples? A Unified Understanding with Self-supervised Contrastive Learning
Chaoning Zhang, Kang Zhang 0008, Chenshuang Zhang, Trung X. Pham, Chang Dong Yoo, In-So Kweon |
ICLR | 1 |
| 2022 | Rotation-aware correlation filters for robust visual tracking
Jiawen Liao, Chun Qi, Jianzhong Cao, Long Ren, Chaoning Zhang |
J. Vis. Commun. Image Represent. | 6 |
| 2022 | MS2Net: Multi-Scale and Multi-Stage Feature Fusion for Blurred Image Super-ResolutionabstractAt present, most mainstream algorithms for single image super-resolution (SISR) assume the image degradation process as an ideal degradation process (e.g. bicubic downscaling), which violates the actual degeneration conditions. In real-world image capturing, objects often move in a dynamic environment, and camera shake also often occurs, which results in serious blurs. Our work focuses on the task of image super-resolution with heavy motion blur, for which we adopt a network with two branches: one branch for image deblurring and the other one for super-resolution. Since the features obtained by the deblurring are rich in details, we apply their features as supplementary information to the super-resolution branch. Based on the adopted dual-branch framework, our major technical novelties lie in two novel modules: Multi-Scale Feature Fusion (MSFF1) module which fuses features of different scale from the deblurring branch to get local and global information, and Multi-Stage Feature Fusion (MSFF2) module which further filters useful information with attention. We evaluate the proposed method under various blur scenarios on the benchmark datasets, demonstrating competitive performance against existing methods. Axi Niu, Yu Zhu 0004, Chaoning Zhang, Jinqiu Sun, In-So Kweon, Yanning Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2021 | Universal Adversarial Perturbations Through the Lens of Deep Steganography: Towards a Fourier PerspectiveabstractThe booming interest in adversarial attacks stems from a misalignment between human vision and a deep neural network (DNN), \ie~a human imperceptible perturbation fools the DNN. Moreover, a single perturbation, often called universal adversarial perturbation (UAP), can be generated to fool the DNN for most images. A similar misalignment phenomenon has also been observed in the deep steganography task, where a decoder network can retrieve a secret image back from a slightly perturbed cover image. We attempt explaining the success of both in a unified manner from the Fourier perspective. We perform task-specific and joint analysis and reveal that (a) frequency is a key factor that influences their performance based on the proposed entropy metric for quantifying the frequency distribution; (b) their success can be attributed to a DNN being highly sensitive to high-frequency content. We also perform feature layer analysis for providing deep insight on model generalization and robustness. Additionally, we propose two new variants of universal perturbations: (1) high-pass UAP (HP-UAP) being less visible to the human eye; (2) Universal Secret Adversarial Perturbation (USAP) that simultaneously achieves attack and hiding. Chaoning Zhang, Philipp Benz, Adil Karjauv, In-So Kweon |
AAAI | 1 |
| 2021 | Adversarial Robustness Comparison of Vision Transformer and MLP-Mixer to CNNs
Philipp Benz, Soomin Ham, Chaoning Zhang, Adil Karjauv, In-So Kweon |
BMVC | 3 |
| 2021 | Batch Normalization Increases Adversarial Vulnerability and Decreases Adversarial Transferability: A Non-Robust Feature PerspectiveabstractBatch normalization (BN) has been widely used in modern deep neural networks (DNNs) due to improved convergence. BN is observed to increase the model accuracy while at the cost of adversarial robustness. There is an increasing interest in the ML community to understand the impact of BN on DNNs, especially related to the model robustness. This work attempts to understand the impact of BN on DNNs from a non-robust feature perspective. Straightforwardly, the improved accuracy can be attributed to the better utilization of useful features. It remains unclear whether BN mainly favors learning robust features (RFs) or non-robust features (NRFs). Our work presents empirical evidence that supports that BN shifts a model towards being more dependent on NRFs. To facilitate the analysis of such a feature robustness shift, we propose a framework for disentangling robust usefulness into robustness and usefulness. Extensive analysis under the proposed framework yields valuable insight on the DNN behavior regarding robustness, e.g. DNNs first mainly learn RFs and then NRFs. The insight that RFs transfer better than NRFs, further inspires simple techniques to strengthen transfer-based black-box attacks. Philipp Benz, Chaoning Zhang, In-So Kweon |
ICCV | 2 |
| 2021 | Data-free Universal Adversarial Perturbation and Black-box AttackabstractUniversal adversarial perturbation (UAP), i.e. a single perturbation to fool the network for most images, is widely recognized as a more practical attack because the UAP can be generated beforehand and applied directly during the at-tack stage. One intriguing phenomenon regarding untargeted UAP is that most images are misclassified to a dominant label. This phenomenon has been reported in previous works while lacking a justified explanation, for which our work attempts to provide an alternative explanation. For a more practical universal attack, our investigation of untargeted UAP focuses on alleviating the dependence on the original training samples, from removing the need for sample labels to limiting the sample size. Towards strictly data-free untargeted UAP, our work proposes to exploit artificial Jigsaw images as the training samples, demonstrating competitive performance. We further investigate the possibility of exploiting the UAP for a data-free black-box attack which is arguably the most practical yet challenging threat model. We demonstrate that there exists optimization-free repetitive patterns which can successfully attack deep models. Code is available at https://bit.ly/3y0ZTIC. Chaoning Zhang, Philipp Benz, Adil Karjauv, In-So Kweon |
ICCV | 1 |
| 2021 | Universal Adversarial Training with Class-Wise PerturbationsabstractDespite their overwhelming success on a wide range of applications, convolutional neural networks (CNNs) are widely recognized to be vulnerable to adversarial examples. This intriguing phenomenon led to a competition between adversarial attacks and defense techniques. So far, adversarial training is the most widely used method for defending against adversarial attacks. It has also been extended to defend against universal adversarial perturbations (UAPs). The SOTA universal adversarial training (UAT) method optimizes a single perturbation for all training samples in the mini-batch. In this work, we find that a UAP does not attack all classes equally. Inspired by this observation, we identify it as the source of the model having unbalanced robustness. To this end, we improve the SOTA UAT by proposing to utilize class-wise UAPs during adversarial training. On multiple benchmark datasets, our class-wise UAT leads superior performance for both clean accuracy and adversarial robustness against universal attack. Philipp Benz, Chaoning Zhang, Adil Karjauv, In-So Kweon |
ICME | 2 |
| 2021 | Motionsnap: A Motion Sensor-Based Approach for Automatic Capture and Editing of Photos and Videos on SmartphonesabstractTaking photos and videos with smartphones has become part of our daily life. However, it is still challenging to capture a brief action at the right time (e.g. jump photos), and video editing (e.g. local slow motion) remains a manual, time-consuming task. To address this problem, we present a motion sensor-based approach that leverages the advanced technical features of modern smartphones to facilitate the capture and editing tasks. Concretely, we simultaneously record the motion sensor data from a smartphone carried by the object of interest, as well as a video from a remote smartphone cam-era. Taking advantage of the motion sensor data, our approach can automatically "snap" the video editing effect to the input video, i.e. apply the effect (e.g. local slow motion) at the right time. The proposed approach is shown to be effective in various applications. Moreover, we implemented it as an app for Android smartphones, running in a fully automatic manner. Adil Karjauv, Sanzhar Bakhtiyarov, Chaoning Zhang, Jean-Charles Bazin, In-So Kweon |
ICME | 3 |
| 2021 | A Survey on Universal Adversarial AttackabstractThe intriguing phenomenon of adversarial examples has attracted significant attention in machine learning and what might be more surprising to the community is the existence of universal adversarial perturbations (UAPs), i.e. a single perturbation to fool the target DNN for most images. With the focus on UAP against deep classifiers, this survey summarizes the recent progress on universal adversarial attacks, discussing the challenges from both the attack and defense sides, as well as the reason for the existence of UAP. We aim to extend this work as a dynamic survey that will regularly update its content to follow new works regarding UAP or universal attack in a wide range of domains, such as image, audio, video, text, etc. Relevant updates will be discussed at: https://bit.ly/2SbQlLG. We welcome authors of future works in this field to contact us for including your new findings. Chaoning Zhang, Philipp Benz, Chenguo Lin, Adil Karjauv, In-So Kweon |
IJCAI | 1 |
| 2021 | Towards Robust Deep Hiding Under Non-Differentiable Distortions for Practical Blind WatermarkingabstractData hiding is one widely used approach for proving ownership through blind watermarking. Deep learning has been widely used in data hiding, for which inserting an attack simulation layer (ASL) after the watermarked image has been widely recognized as the most effective approach for improving the pipeline robustness against distortions. Despite its wide usage, the gain of enhanced robustness from ASL is usually interpreted through the lens of augmentation, while our work explores this gain from a new perspective by disentangling the forward and backward propagation of such ASL. We find that the main influential component is forward propagation instead of backward propagation. This observation motivates us to use forward ASL to make the pipeline compatible with non-differentiable and/or black-box distortion, such as lossy (JPEG) compression and photoshop effects. Extensive experiments demonstrate the efficacy of our simple approach. Chaoning Zhang, Adil Karjauv, Philipp Benz, In-So Kweon |
ACM Multimedia | 1 |
| 2021 | Revisiting Batch Normalization for Improving Corruption RobustnessabstractThe performance of DNNs trained on clean images has been shown to decrease when the test images have common corruptions. In this work, we interpret corruption robustness as a domain shift and propose to rectify batch normalization (BN) statistics for improving model robustness. This is motivated by perceiving the shift from the clean domain to the corruption domain as a style shift that is represented by the BN statistics. We find that simply estimating and adapting the BN statistics on a few (32 for instance) representation samples, without retraining the model, improves the corruption robustness by a large margin on several benchmark datasets with a wide range of model architectures. For example, on ImageNet-C, statistics adaptation improves the top1 accuracy of ResNet50 from 39.2% to 48.7%. Moreover, we find that this technique can further improve state-of-the-art robust models from 58.1% to 63.3%. Philipp Benz, Chaoning Zhang, Adil Karjauv, In-So Kweon |
WACV | 2 |
| 2021 | ResNet or DenseNet? Introducing Dense Shortcuts to ResNetabstractResNet or DenseNet? Nowadays, most deep learning based approaches are implemented with seminal backbone networks, among them the two arguably most famous ones are ResNet and DenseNet. Despite their competitive performance and overwhelming popularity, inherent drawbacks exist for both of them. For ResNet, the identity shortcut that stabilizes training might limit its representation capacity, and DenseNet mitigates it with multi-layer feature concatenation. However, the dense concatenation causes a new problem of requiring high GPU memory and more training time. Partially due to this, it is not a trivial choice between ResNet and DenseNet. This paper provides a unified perspective of dense summation to analyze them, which facilitates a better understanding of their core difference. We further propose dense weighted normalized shortcuts as a solution to the dilemma between them. Our proposed dense shortcut inherits the design philosophy of simple design in ResNet and DenseNet. On several benchmark datasets, the experimental results show that the proposed DSNet achieves significantly better results than ResNet, and achieves comparable performance as DenseNet but requiring fewer computation resources. Chaoning Zhang, Philipp Benz, Dawit Mureja Argaw, Seokju Lee, Junsik Kim 0001, François Rameau, Jean-Charles Bazin, In-So Kweon |
WACV | 1 |
| 2020 | CD-UAP: Class Discriminative Universal Adversarial PerturbationabstractA single universal adversarial perturbation (UAP) can be added to all natural images to change most of their predicted class labels. It is of high practical relevance for an attacker to have flexible control over the targeted classes to be attacked, however, the existing UAP method attacks samples from all classes. In this work, we propose a new universal attack method to generate a single perturbation that fools a target network to misclassify only a chosen group of classes, while having limited influence on the remaining classes. Since the proposed attack generates a universal adversarial perturbation that is discriminative to targeted and non-targeted classes, we term it class discriminative universal adversarial perturbation (CD-UAP). We propose one simple yet effective algorithm framework, under which we design and compare various loss function configurations tailored for the class discriminative universal attack. The proposed approach has been evaluated with extensive experiments on various benchmark datasets. Additionally, our proposed approach achieves state-of-the-art performance for the original task of UAP attacking all classes, which demonstrates the effectiveness of our approach. Chaoning Zhang, Philipp Benz, Tooba Imtiaz, In-So Kweon |
AAAI | 1 |
| 2020 | Double Targeted Universal Adversarial Perturbations
Philipp Benz, Chaoning Zhang, Tooba Imtiaz, In-So Kweon |
ACCV (4) | 2 |
| 2020 | Understanding Adversarial Examples From the Mutual Influence of Images and PerturbationsabstractA wide variety of works have explored the reason for the existence of adversarial examples, but there is no consensus on the explanation. We propose to treat the DNN logits as a vector for feature representation, and exploit them to analyze the mutual influence of two independent inputs based on the Pearson correlation coefficient (PCC). We utilize this vector representation to understand adversarial examples by disentangling the clean images and adversarial perturbations, and analyze their influence on each other. Our results suggest a new perspective towards the relationship between images and universal perturbations: Universal perturbations contain dominant features, and images behave like noise to them. This feature perspective leads to a new method for generating targeted universal adversarial perturbations using random source images. We are the first to achieve the challenging task of a targeted universal attack without utilizing original training data. Our approach using a proxy dataset achieves comparable performance to the state-of-the-art baselines which utilize the original training dataset. Chaoning Zhang, Philipp Benz, Tooba Imtiaz, In-So Kweon |
CVPR | 1 |
| 2020 | UDH: Universal Deep Hiding for Steganography, Watermarking, and Light Field MessagingabstractNeural networks have been shown effective in deep steganography for hiding a full image in another. However, the reason for its success remains not fully clear. Under the existing cover ($C$) dependent deep hiding (DDH) pipeline, it is challenging to analyze how the secret ($S$) image is encoded since the encoded message cannot be analyzed independently. We propose a novel universal deep hiding (UDH) meta-architecture to disentangle the encoding of $S$ from $C$. We perform extensive analysis and demonstrate that the success of deep steganography can be attributed to a frequency discrepancy between $C$ and the encoded secret image. Despite $S$ being hidden in a cover-agnostic manner, strikingly, UDH achieves a performance comparable to the existing DDH. Beyond hiding one image, we push the limits of deep steganography. Exploiting its property of being \emph{universal}, we propose universal watermarking as a timely solution to address the concern of the exponentially increasing amount of images/videos. UDH is robust to a pixel intensity shift on the container image, which makes it suitable for challenging application of light field messaging (LFM). This is the first work demonstrating the success of (DNN-based) hiding a full image for watermarking and LFM. Code: \url{https://github.com/ChaoningZhang/Universal-Deep-Hiding} Chaoning Zhang, Philipp Benz, Adil Karjauv, In-So Kweon |
NeurIPS | 1 |
| 2020 | DeepPTZ: Deep Self-Calibration for PTZ CamerasabstractRotating and zooming cameras, also called PTZ (Pan-Tilt-Zoom) cameras, are widely used in modern surveillance systems. While their zooming ability allows acquiring detailed images of the scene, it also makes their calibration more challenging since any zooming action results in a modification of their intrinsic parameters. Therefore, such camera calibration has to be computed online; this process is called self-calibration. In this paper, given an image pair captured by a PTZ camera, we propose a deep learning based approach to automatically estimate the focal length and distortion parameters of both images as well as the rotation angles between them. The proposed approach relies on a dual-Siamese structure, imposing bidirectional constraints. The proposed network is trained on a large-scale dataset automatically generated from a set of panoramas. Empirically, we demonstrate that our proposed approach achieves competitive performance with respect to both deep learning based and traditional state-of-the art methods. Our code and model will be publicly available at https://github.com/ChaoningZhang/DeepPTZ. Chaoning Zhang, François Rameau, Junsik Kim 0001, Dawit Mureja Argaw, Jean-Charles Bazin, In-So Kweon |
WACV | 1 |
| 2019 | Revisiting Residual Networks with Nonlinear Shortcuts
Chaoning Zhang, François Rameau, Seokju Lee, Junsik Kim 0001, Philipp Benz, Dawit Mureja Argaw, Jean-Charles Bazin, In-So Kweon |
BMVC | 1 |