Hongjing Niu

dblp:267/9397 · DBLP profile ↗
← Back
10ranked-venue papers
4as first author
9since 2021 · last 2024
0000-0002-9480-6464ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 4 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 6 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2024 Rectifying Shortcut Learning through Cellular Differentiation in Deep Learning Neurons
Hongjing Niu, Hanting Li, Guoping Wu, Bin Li 0025, Feng Zhao 0004
BMVC1
2024 Frequency Decomposition to Tap the Potential of Single Domain for Generalization
Hongjing Niu, Qingyue Yang, Wei Zhang 0251, Bin Li 0025, Feng Zhao 0004
BMVC1
2024 Stable Preference: Redefining Training Paradigm of Human Preference Model for Text-to-Image Synthesis
Hanting Li, Hongjing Niu, Feng Zhao 0004
ECCV (28)2
2024 Idling Neurons, Appropriately Lenient Workload During Fine-Tuning Leads to Better Generalization
Hongjing Niu, Hanting Li, Bin Li 0025, Feng Zhao 0004
ECCV (53)1
2024 CLIPER: A Unified Vision-Language Framework for In-the-Wild Facial Expression Recognition
abstract
As one of the most informative behaviors of humans, facial expressions are often compound and variable, which is manifested by the fact that different people may express the same expression in very different ways. However, most facial expression recognition (FER) methods still use one-hot or soft labels as the supervision, which lack sufficient semantic descriptions of facial expressions and are less interpretable. Recently, contrastive vision-language pre-training models (e.g., CLIP) use text as the supervision and have injected new vitality into various computer vision tasks, benefiting from the rich semantics in text. Therefore, we propose CLIPER, a unified framework for both static and dynamic facial Expression Recognition based on CLIP. Besides, we introduce multiple expression text descriptors (METD) to learn fine-grained expression representations and a two-stage training paradigm to reserve the interpretability of CLIP. We conduct extensive experiments on several popular FER benchmarks to demonstrates the effectiveness of CLIPER. The source code will be available at https://github.com/muse1998/CLIPER.
Hanting Li, Hongjing Niu, Zhaoqing Zhu, Feng Zhao 0004
ICME2
2023 Intensity-Aware Loss for Dynamic Facial Expression Recognition in the Wild
abstract
Compared with the image-based static facial expression recognition (SFER) task, the dynamic facial expression recognition (DFER) task based on video sequences is closer to the natural expression recognition scene. However, DFER is often more challenging. One of the main reasons is that video sequences often contain frames with different expression intensities, especially for the facial expressions in the real-world scenarios, while the images in SFER frequently present uniform and high expression intensities. Nevertheless, if the expressions with different intensities are treated equally, the features learned by the networks will have large intra-class and small inter-class differences, which are harmful to DFER. To tackle this problem, we propose the global convolution-attention block (GCA) to rescale the channels of the feature maps. In addition, we introduce the intensity-aware loss (IAL) in the training process to help the network distinguish the samples with relatively low expression intensities. Experiments on two in-the-wild dynamic facial expression datasets (i.e., DFEW and FERV39k) indicate that our method outperforms the state-of-the-art DFER approaches. The source code will be available at https://github.com/muse1998/IAL-for-Facial-Expression-Recognition.
Hanting Li, Hongjing Niu, Zhaoqing Zhu, Feng Zhao 0004
AAAI2
2023 Enhancing Backdoor Attacks With Multi-Level MMD Regularization
abstract
While Deep Neural Networks (DNNs) excel in many tasks, the huge training resources they require become an obstacle for practitioners to develop their own models. It has become common to collect data from the Internet or hire a third party to train models. Unfortunately, recent studies have shown that these operations provide a viable pathway for maliciously injecting hidden backdoors into DNNs. Several defense methods have been developed to detect malicious samples, with the common assumption that the latent representations of benign and malicious samples extracted by the infected model exhibit different distributions. However, it is still an open question whether this assumption holds up. In this article, we investigate such differences thoroughly via answering three questions: 1) What are the characteristics of the distributional differences? 2) How can they be effectively reduced? 3) What impact does this reduction have on difference-based defense methods? First, the distributional differences of multi-level representations on the regularly trained backdoored models are verified to be significant by adopting Maximum Mean Discrepancy (MMD), Energy Distance (ED), and Sliced Wasserstein Distance (SWD) as the metrics. Then, ML-MMDR, a difference reduction method that adds multi-level MMD regularization into the loss, is proposed, and its effectiveness is testified on three typical difference-based defense methods. Across all the experimental settings, the F1 scores of these methods drop from 90%-100% on the regularly trained backdoored models to 60%-70% on the models trained with ML-MMDR. These results indicate that the proposed MMD regularization can enhance the stealthiness of existing backdoor attack methods. The prototype code of our method is now available athttps://github.com/xpf/Multi-Level-MMD-Regularization.
Hongjing Niu, Ziqiang Li 0001, Bin Li 0025
IEEE Trans. Dependable Secur. Comput.2
2022 Roadblocks for Temporarily Disabling Shortcuts and Learning New Knowledge
abstract
Deep learning models have been found with a tendency of relying on shortcuts, i.e., decision rules that perform well on standard benchmarks but fail when transferred to more challenging testing conditions. Such reliance may hinder deep learning models from learning other task-related features and seriously affect their performance and robustness. Although recent studies have shown some characteristics of shortcuts, there are few investigations on how to help the deep learning models to solve shortcut problems. This paper proposes a framework to address this issue by setting up roadblocks on shortcuts. Specifically, roadblocks are placed when the model is urged to learn to complete a gently modified task to ensure that the learned knowledge, including shortcuts, is insufficient the complete the task. Therefore, the model trained on the modified task will no longer over-rely on shortcuts. Extensive experiments demonstrate that the proposed framework significantly improves the training of networks on both synthetic and real-world datasets in terms of both classification accuracy and feature diversity. Moreover, the visualization results show that the mechanism behind the proposed our method is consistent with our expectations. In summary, our approach can effectively disable the shortcuts and thus learn more robust features.
Hongjing Niu, Hanting Li, Feng Zhao 0004, Bin Li 0025
NeurIPS1
2021 On the receptive field misalignment in CAM-based visual explanations
Hongjing Niu, Ziqiang Li 0001, Bin Li 0025
Pattern Recognit. Lett.2
2020 Interpreting the Latent Space of GANs via Correlation Analysis for Controllable Concept Manipulation
abstract
Generative adversarial nets (GANs) have been successfully applied in many fields like image generation, inpainting, super-resolution, and drug discovery, etc. By now, the inner process of GANs is far from being understood. To get a deeper insight into the intrinsic mechanism of GANs, in this paper, a method for interpreting the latent space of GANs by analyzing the correlation between latent variables and the corresponding semantic contents in generated images is proposed. Unlike previous methods that focus on dissecting models via feature visualization, the emphasis of this work is put on the variables in latent space, i.e. how the latent variables affect the quantitative analysis of generated results. Given a pre-trained GAN model with weights fixed, the latent variables are intervened to analyze their effect on the semantic content in generated images. A set of controlling latent variables can be derived for specific content generation, and the controllable semantic content manipulation is achieved. The proposed method is testified on the datasets Fashion-MNIST and UT Zappos50K, experiment results show its effectiveness.
Ziqiang Li 0001, Rentuo Tao, Hongjing Niu, Mingdao Yue, Bin Li 0025
ICPR3