EDBT 2026 Demo / reviewers in the wild / expert
Hanting Li
dblp:276/2638
· DBLP profile ↗
10ranked-venue papers
5as first author
10since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 8 since 2021Artificial intelligence and machine learning · 7 · 4 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Rectifying Shortcut Learning through Cellular Differentiation in Deep Learning Neurons
Hongjing Niu, Hanting Li, Guoping Wu, Bin Li 0025, Feng Zhao 0004 |
BMVC | 2 |
| 2024 | Stable Preference: Redefining Training Paradigm of Human Preference Model for Text-to-Image Synthesis
Hanting Li, Hongjing Niu, Feng Zhao 0004 |
ECCV (28) | 1 |
| 2024 | Idling Neurons, Appropriately Lenient Workload During Fine-Tuning Leads to Better Generalization
Hongjing Niu, Hanting Li, Bin Li 0025, Feng Zhao 0004 |
ECCV (53) | 2 |
| 2024 | CLIPER: A Unified Vision-Language Framework for In-the-Wild Facial Expression RecognitionabstractAs one of the most informative behaviors of humans, facial expressions are often compound and variable, which is manifested by the fact that different people may express the same expression in very different ways. However, most facial expression recognition (FER) methods still use one-hot or soft labels as the supervision, which lack sufficient semantic descriptions of facial expressions and are less interpretable. Recently, contrastive vision-language pre-training models (e.g., CLIP) use text as the supervision and have injected new vitality into various computer vision tasks, benefiting from the rich semantics in text. Therefore, we propose CLIPER, a unified framework for both static and dynamic facial Expression Recognition based on CLIP. Besides, we introduce multiple expression text descriptors (METD) to learn fine-grained expression representations and a two-stage training paradigm to reserve the interpretability of CLIP. We conduct extensive experiments on several popular FER benchmarks to demonstrates the effectiveness of CLIPER. The source code will be available at https://github.com/muse1998/CLIPER. Hanting Li, Hongjing Niu, Zhaoqing Zhu, Feng Zhao 0004 |
ICME | 1 |
| 2024 | Achieving accurate and balanced regional electric vehicle charging load forecasting with a dynamic road network: a case study of Lanzhou City
Hanting Li, Minan Tang, Yunfei Mu |
Appl. Intell. | 1 |
| 2023 | Intensity-Aware Loss for Dynamic Facial Expression Recognition in the WildabstractCompared with the image-based static facial expression recognition (SFER) task, the dynamic facial expression recognition (DFER) task based on video sequences is closer to the natural expression recognition scene. However, DFER is often more challenging. One of the main reasons is that video sequences often contain frames with different expression intensities, especially for the facial expressions in the real-world scenarios, while the images in SFER frequently present uniform and high expression intensities. Nevertheless, if the expressions with different intensities are treated equally, the features learned by the networks will have large intra-class and small inter-class differences, which are harmful to DFER. To tackle this problem, we propose the global convolution-attention block (GCA) to rescale the channels of the feature maps. In addition, we introduce the intensity-aware loss (IAL) in the training process to help the network distinguish the samples with relatively low expression intensities. Experiments on two in-the-wild dynamic facial expression datasets (i.e., DFEW and FERV39k) indicate that our method outperforms the state-of-the-art DFER approaches. The source code will be available at https://github.com/muse1998/IAL-for-Facial-Expression-Recognition. Hanting Li, Hongjing Niu, Zhaoqing Zhu, Feng Zhao 0004 |
AAAI | 1 |
| 2023 | AFNet-M: Adaptive Fusion Network with Masks for 2D+3D Facial Expression Recognitionabstract2D+3D facial expression recognition (FER) can effectively cope with illumination and pose changes by merging texture and robust depth information. Most deep learning-based approaches employ the simple fusion strategy that concatenates the multimodal features directly after fully-connected layers, without considering the different degrees of significance for each modality. Meanwhile, how to focus more on both 2D and 3D local features is still a great challenge. In this paper, we propose the adaptive fusion network with masks (AFNet-M) for 2D+3D FER. To enhance 2D and 3D local features, we take the masks annotating salient regions of the face as prior knowledge and design the mask attention module (MA) which can automatically learn two modulation vectors to scale the feature maps. We also introduce an adaptive fusion module (AF) at convolutional layers through the computed importance weights. Experimental results demonstrate that our AFNet-M achieves the state-of-the-art performance on BU-3DFE and Bosphorus datasets and requires fewer parameters in comparison with other models. Mingzhe Sui, Hanting Li, Zhaoqing Zhu, Feng Zhao 0004 |
ICIP | 2 |
| 2022 | CMANET: Curvature-Aware Soft Mask Guided Attention Fusion Network for 2D+3D Facial Expression RecognitionabstractAs 2D texture and 3D structural information can describe facial features complementarily, 2D+3D facial expression recognition (FER) has received widespread attention. Though recent methods for 2D+3D FER have reached excellent performance, they still face two challenges: the way for attending to critical face areas and the strategy for fusing multi-modal information. To address these issues, we propose a curvature-aware soft mask guided attention fusion network (CMANet), which mainly consists of two components: curvature-aware attention module and multi-modal attention fusion module. The former utilizes the curvature-aware soft mask guiding the homo-modal attention mechanism to focus on potentially important areas with soft weights, while the latter applies pixel-level fusion on multi-modal features to retain the significant information from different modalities and also allows multi-modal features to interact in a larger field of view. Extensive experimental results show that our CMANet achieves outstanding accuracies (90.24% on BU-3DFE and 89.36% on Bosphorus) and outperforms the state-of-the-art methods. Zhaoqing Zhu, Mingzhe Sui, Hanting Li, Feng Zhao 0004 |
ICME | 3 |
| 2022 | MMNet: Muscle Motion-Guided Network for Micro-Expression RecognitionabstractFacial micro-expressions (MEs) are involuntary facial motions revealing people’s real feelings and play an important role in the early intervention of mental illness, the national security, and many human-computer interaction systems. However, existing micro-expression datasets are limited and usually pose some challenges for training good classifiers. To model the subtle facial muscle motions, we propose a robust micro-expression recognition (MER) framework, namely muscle motion-guided network (MMNet). Specifically, a continuous attention (CA) block is introduced to focus on modeling local subtle muscle motion patterns with little identity information, which is different from most previous methods that directly extract features from complete video frames with much identity information. Besides, we design a position calibration (PC) module based on the vision transformer. By adding the position embeddings of the face generated by the PC module at the end of the two branches, the PC module can help to add position information to facial muscle motion-pattern features for the MER. Extensive experiments on three public micro-expression datasets demonstrate that our approach outperforms state-of-the-art methods by a large margin. Code is available at https://github.com/muse1998/MMNet. Hanting Li, Mingzhe Sui, Zhaoqing Zhu, Feng Zhao 0004 |
IJCAI | 1 |
| 2022 | Roadblocks for Temporarily Disabling Shortcuts and Learning New KnowledgeabstractDeep learning models have been found with a tendency of relying on shortcuts, i.e., decision rules that perform well on standard benchmarks but fail when transferred to more challenging testing conditions. Such reliance may hinder deep learning models from learning other task-related features and seriously affect their performance and robustness. Although recent studies have shown some characteristics of shortcuts, there are few investigations on how to help the deep learning models to solve shortcut problems. This paper proposes a framework to address this issue by setting up roadblocks on shortcuts. Specifically, roadblocks are placed when the model is urged to learn to complete a gently modified task to ensure that the learned knowledge, including shortcuts, is insufficient the complete the task. Therefore, the model trained on the modified task will no longer over-rely on shortcuts. Extensive experiments demonstrate that the proposed framework significantly improves the training of networks on both synthetic and real-world datasets in terms of both classification accuracy and feature diversity. Moreover, the visualization results show that the mechanism behind the proposed our method is consistent with our expectations. In summary, our approach can effectively disable the shortcuts and thus learn more robust features. Hongjing Niu, Hanting Li, Feng Zhao 0004, Bin Li 0025 |
NeurIPS | 2 |