Mang Ning

dblp:302/2427 · DBLP profile ↗
← Back
8ranked-venue papers
6as first author
8since 2021 · last 2025
0000-0001-6037-1661ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 4 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Generative modeling · 44% Probabilistic and Bayesian machine learning · 18% Deep learning architectures and training · 14%
Computer graphics and multimedia
1 paper
Image and video coding · 100%

Topics — the 9 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
diffusion model
3.242025
Consistent Story Generation: Unlocking the Potential of Zigzag Sampling · NeurIPS 2025
DCTdiff: Intriguing Properties of Image Generative Modeling in the DCT Space · ICML 2025
Elucidating the Exposure Bias in Diffusion Models · ICLR 2024
Machine learning › Probabilistic and Bayesian machine learning
sampling
1.622025
Consistent Story Generation: Unlocking the Potential of Zigzag Sampling · NeurIPS 2025
Elucidating the Exposure Bias in Diffusion Models · ICLR 2024
Natural language and speech › Language models and text generation › text generation
story generation
0.912025
Consistent Story Generation: Unlocking the Potential of Zigzag Sampling · NeurIPS 2025
Machine learning › Generative modeling › diffusion model
text-to-image generation
0.912025
Consistent Story Generation: Unlocking the Potential of Zigzag Sampling · NeurIPS 2025
Image and video coding › transform coding
discrete cosine transform
0.912025
DCTdiff: Intriguing Properties of Image Generative Modeling in the DCT Space · ICML 2025
Natural language and speech › Machine translation › neural machine translation
exposure bias
0.812024
Elucidating the Exposure Bias in Diffusion Models · ICLR 2024
Machine learning › Deep learning architectures and training › sequence modeling › sequence generation
exposure bias mitigation
0.712023
Input Perturbation Reduces Exposure Bias in Diffusion Models · ICML 2023
Machine learning › Trustworthy machine learning
input perturbation
0.712023
Input Perturbation Reduces Exposure Bias in Diffusion Models · ICML 2023
Machine learning › Deep learning architectures and training › regularization
training regularization
0.712023
Input Perturbation Reduces Exposure Bias in Diffusion Models · ICML 2023

Methods — techniques the papers use, named apart from their topics

dit · 1.7diffusion sampling · 1.7UViT · 1.7zigzag sampling · 0.9visual sharing · 0.9asymmetric prompting · 0.9epsilon scaling · 0.8input perturbation · 0.7denoising diffusion probabilistic model · 0.7
YearPublicationVenuePosition
2025 Revisiting Representation Learning and Identity Adversarial Training for Facial Behavior Understanding
abstract
Facial Action Unit (AU) detection has gained significant attention as it enables the breakdown of complex facial expressions into individual muscle movements. In this paper, we revisit two fundamental factors in AU detection: diverse and large-scale data and subject identity regularization. Motivated by recent advances in foundation models, we highlight the importance of data and introduce Face 9 M, a diverse dataset comprising 9 million facial images from multiple public sources. Pretraining a masked autoencoder on Face9M yields strong performance in AU detection and facial expression tasks. More importantly, we emphasize that the Identity Adversarial Training (IAT) has not been well explored in AU tasks. To fill this gap, we first show that subject identity in AU datasets creates shortcut learning for the model and leads to suboptimal solutions to AU predictions. Secondly, we demonstrate that strong IAT regularization is necessary to learn identityinvariant features. Finally, we elucidate the design space of IAT and empirically show that IAT circumvents the identity-based shortcut learning and results in a better solution. Our proposed methods, Facial Masked Autoencoder (FMAE) and IAT, are simple, generic and effective. Remarkably, the proposed FMAEIAT approach achieves new state-of-the-art F1 scores on BP4D ($67.1 \%$), BP4D+ ($66.8 \%$), and DISFA ($70.1 \%$) databases, significantly outperforming previous work. We release the code and model at https://github.com/forever208/FMAE-IAT.
Mang Ning, Albert Ali Salah, Itir Önal
FG1
2025 DCTdiff: Intriguing Properties of Image Generative Modeling in the DCT Space
abstract
This paper explores image modeling from the frequency space and introduces DCTdiff, an end-to-end diffusion generative paradigm that efficiently models images in the discrete cosine transform (DCT) space. We investigate the design space of DCTdiff and reveal the key design factors. Experiments on different frameworks (UViT, DiT), generation tasks, and various diffusion samplers demonstrate that DCTdiff outperforms pixel-based diffusion models regarding generative quality and training efficiency. Remarkably, DCTdiff can seamlessly scale up to 512$\times$512 resolution without using the latent diffusion paradigm and beats latent diffusion (using SD-VAE) with only 1/4 training cost. Finally, we illustrate several intriguing properties of DCT image modeling. For example, we provide a theoretical proof of why `image diffusion can be seen as spectral autoregression', bridging the gap between diffusion and autoregressive models. The effectiveness of DCTdiff and the introduced properties suggest a promising direction for image modeling in the frequency space. The code is at https://github.com/forever208/DCTdiff.
Mang Ning, Mingxiao Li 0002, Jianlin Su, Haozhe Jia, Lanmiao Liu, Martin Benes 0001, Wenshuo Chen, Albert Ali Salah, Itir Önal
ICML1
2025 Consistent Story Generation: Unlocking the Potential of Zigzag Sampling
abstract
Text-to-image generation models have made significant progress in producing high-quality images from textual descriptions, yet they continue to struggle with maintaining subject consistency across multiple images, a fundamental requirement for visual storytelling. Existing methods attempt to address this by either fine-tuning models on large-scale story visualization datasets, which is resource-intensive, or by using training-free techniques that share information across generations, which still yield limited success. In this paper, we introduce a novel training-free sampling strategy called Zigzag Sampling with Asymmetric Prompts and Visual Sharing to enhance subject consistency in visual story generation. Our approach proposes a zigzag sampling mechanism that alternates between asymmetric prompting to retain subject characteristics, while a visual sharing module transfers visual cues across generated images to enforce consistency. Experimental results, based on both quantitative metrics and qualitative evaluations, demonstrate that our method significantly outperforms previous approaches in generating coherent and consistent visual stories. The code is available at https://github.com/Mingxiao-Li/Asymmetry-Zigzag-StoryDiffusion.
Mingxiao Li 0002, Mang Ning, Marie-Francine Moens
NeurIPS2
2024 Elucidating the Exposure Bias in Diffusion Models
abstract
Diffusion models have demonstrated impressive generative capabilities, but their exposure bias problem, described as the input mismatch between training and sampling, lacks in-depth exploration. In this paper, we investigate the exposure bias problem in diffusion models by first analytically modelling the sampling distribution, based on which we then attribute the prediction error at each sampling step as the root cause of the exposure bias issue. Furthermore, we discuss potential solutions to this issue and propose an intuitive metric for it. Along with the elucidation of exposure bias, we propose a simple, yet effective, training-free method called Epsilon Scaling to alleviate the exposure bias. We show that Epsilon Scaling explicitly moves the sampling trajectory closer to the vector field learned in the training phase by scaling down the network output, mitigating the input mismatch between training and sampling. Experiments on various diffusion frameworks (ADM, DDIM, EDM, LDM, DiT, PFGM++) verify the effectiveness of our method. Remarkably, our ADM-ES, as a state-of-the-art stochastic sampler, obtains 2.17 FID on CIFAR-10 under 100-step unconditional generation. The code is at https://github.com/forever208/ADM-ES
Mang Ning, Mingxiao Li 0002, Jianlin Su, Albert Ali Salah, Itir Önal
ICLR1
2023 Automated Emotional Valence Estimation in Infants with Stochastic and Strided Temporal Sampling
abstract
We propose the first automated approach to estimate the emotional valence of infants from their facial behavior. We use the state-of-the-art transformer-based video masked autoencoder (VideoMAE) that is pre-trained on a large video dataset as a backbone, and finetune it on two large, well-annotated infant video datasets (SIBSMILE and MODELING). To augment the limited data, we propose a novel video temporal augmentation method called Stochastic and Strided Temporal Sampling (SSTS). We demonstrate the effectiveness of our approach for infant valence estimation by achieving 0.671 Concordance Correlation Coefficient (CCC) on SIBSMILE and MODELING. The experiments show that SSTS remarkably accelerates the training speed by 8 times while gaining the best valence estimation performance. Lastly, we suggest that face detection and cropping (coarse registration) is a promising alternative to landmark-based registration (i.e. fine registration) in data pre-processing when accurate infant facial landmark detectors are inaccessible.
Mang Ning, Itir Önal, Daniel S. Messinger, Jeffrey F. Cohn, Albert Ali Salah
ACII1
2023 Input Perturbation Reduces Exposure Bias in Diffusion Models
abstract
Denoising Diffusion Probabilistic Models have shown an impressive generation quality although their long sampling chain leads to high computational costs. In this paper, we observe that a long sampling chain also leads to an error accumulation phenomenon, which is similar to the exposure bias problem in autoregressive text generation. Specifically, we note that there is a discrepancy between training and testing, since the former is conditioned on the ground truth samples, while the latter is conditioned on the previously generated results. To alleviate this problem, we propose a very simple but effective training regularization, consisting in perturbing the ground truth samples to simulate the inference time prediction errors. We empirically show that, without affecting the recall and precision, the proposed input perturbation leads to a significant improvement in the sample quality while reducing both the training and the inference times. For instance, on CelebA 64x64, we achieve a new state-of-the-art FID score of 1.27, while saving 37.5% of the training time. The code is available at https://github.com/forever208/DDPM-IP
Mang Ning, Enver Sangineto, Angelo Porrello, Simone Calderara, Rita Cucchiara
ICML1
2022 METRIC: Toward a Drone-based Cyber-Physical Traffic Management System
abstract
Drone-based system has a big potential to be applied for traffic monitoring and other advanced applications in Intelligent Transport Systems (ITS). This paper introduces our latest efforts of digitalising road traffic by various types of sensing systems, among which visual detection by drones provides a promising technical solution. A platform, called METRIC, is under recent development to carry out real-time traffic measurement and prediction using drone-based data collection. The current system is designed as a cyber-physical system (CPS) with essential functions aiming for visual traffic detection and analysis, real-time traffic estimation and prediction as well as decision supports based on simulation. In addition to the computer vision functions developed in the earlier stage, this paper also presents the CPS system architecture and the current implementation of the drone front-end system and a simulation-based system being used for further drone operations.
Mang Ning, Andrei Radu
SMC3
2021 YOLOv4-object: an Efficient Model and Method for Object Discovery
abstract
Object discovery refers to recognising all unknown objects in images, which is crucial for robotic systems to explore the unseen environment. Recently, object detection models based on deep learning have shown remarkable achievements in object classification and localisation. However, these models have difficulties handling the unseen environment because it is infeasible to exhaustively predefine all types of objects. In this paper, we propose the model YOLOv4-object to recognise all objects in images by modifying the output space of YOLOv4 and related image labels. Experiments on COCO dataset demonstrate the effectiveness of our method by achieving 67.97% recall (6.49% higher than vanilla YOLOv4). We point out that the incomplete labels (COCO only labels for 80 categories) hurt the learning process of object discovery and a higher recall can be achieved by our method if the dataset is fully labelled. Moreover, our approach is transferable, extensible, and compressible, showing broad application scenarios. Finally, we conduct extensive experiments to illustrate the factors that affect the object discovery performance of our model and some suggestions on practical implementations are elaborated.
Mang Ning, Wenyuan Hou, Mihhail Matskin
COMPSAC1