Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Dongting Hu

dblp:159/6788 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
8since 2021 · last 2026
0009-0007-2119-1829ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 3 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Generative modeling · 33% Efficient and distributed learning · 28% Time series and sequential data · 14%
Network and information security
1 paper
Security and privacy of machine learning · 100%
Computer graphics and multimedia
1 paper
Rendering · 100%

Topics — the 18 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
diffusion model
2.532025
Stochastic Diffusion: A Diffusion Based Model for Stochastic Time Series Forecasting · KDD (2) 2025
SnapGen: Taming High-Resolution Text-to-Image Models for Mobile Devices with Efficient Architectures and Training · CVPR 2025
In-N-Out: Lifting 2D Diffusion Prior for 3D Object Removal via Tuning-Free Latents Alignment · NeurIPS 2024
Machine learning › Efficient and distributed learning
inference efficiency
0.912025
SnapGen: Taming High-Resolution Text-to-Image Models for Mobile Devices with Efficient Architectures and Training · CVPR 2025
Machine learning › Efficient and distributed learning › model compression
knowledge distillation
0.912025
SnapGen: Taming High-Resolution Text-to-Image Models for Mobile Devices with Efficient Architectures and Training · CVPR 2025
Machine learning › Efficient and distributed learning › model deployment
mobile deployment
0.912025
SnapGen: Taming High-Resolution Text-to-Image Models for Mobile Devices with Efficient Architectures and Training · CVPR 2025
Machine learning › Efficient and distributed learning
model compression
0.912025
SnapGen: Taming High-Resolution Text-to-Image Models for Mobile Devices with Efficient Architectures and Training · CVPR 2025
Machine learning › Time series and sequential data › time series modeling
probabilistic forecasting
0.912025
Stochastic Diffusion: A Diffusion Based Model for Stochastic Time Series Forecasting · KDD (2) 2025
Machine learning › Generative modeling › diffusion model
text-to-image generation
0.912025
SnapGen: Taming High-Resolution Text-to-Image Models for Mobile Devices with Efficient Architectures and Training · CVPR 2025
Machine learning › Time series and sequential data › time series analysis
time series forecasting
0.912025
Stochastic Diffusion: A Diffusion Based Model for Stochastic Time Series Forecasting · KDD (2) 2025
Security and privacy of machine learning
adversarial attack
0.912025
STBA: Towards Evaluating the Robustness of DNNs for Query-Limited Black-Box Scenario · IEEE Trans. Multim. 2025
Security and privacy of machine learning
adversarial example
0.912025
STBA: Towards Evaluating the Robustness of DNNs for Query-Limited Black-Box Scenario · IEEE Trans. Multim. 2025
Machine learning › Generative modeling › generative adversarial network
3d-aware image synthesis
0.812024
In-N-Out: Lifting 2D Diffusion Prior for 3D Object Removal via Tuning-Free Latents Alignment · NeurIPS 2024
Computer vision › 3D vision
neural radiance field
0.812024
In-N-Out: Lifting 2D Diffusion Prior for 3D Object Removal via Tuning-Free Latents Alignment · NeurIPS 2024
Machine learning › Deep learning architectures and training
multi-scale representation
0.712023
Multiscale Representation for Real-Time Anti-Aliasing Neural Rendering · ICCV 2023
Rendering
neural radiance fields
0.712023
Multiscale Representation for Real-Time Anti-Aliasing Neural Rendering · ICCV 2023
Computer vision › 3D vision
depth estimation
0.612022
Uncertainty Quantification in Depth Estimation via Constrained Ordinal Regression · ECCV (2) 2022
Machine learning › Trustworthy machine learning
uncertainty estimation
0.612022
Uncertainty Quantification in Depth Estimation via Constrained Ordinal Regression · ECCV (2) 2022
Machine learning › Trustworthy machine learning
robustness evaluation
0.312025
STBA: Towards Evaluating the Robustness of DNNs for Query-Limited Black-Box Scenario · IEEE Trans. Multim. 2025
Computer vision › 3D vision › multi-view geometry
multi-view consistency
0.212024
In-N-Out: Lifting 2D Diffusion Prior for 3D Object Removal via Tuning-Free Latents Alignment · NeurIPS 2024

Methods — techniques the papers use, named apart from their topics

gradient estimation · 1.7flow field estimation · 1.7sequential model · 0.9few-step generation · 0.9diffusion probabilistic model · 0.9cross-architecture knowledge distillation · 0.9adversarial guidance · 0.9patch-based hybrid loss · 0.8latent alignment · 0.8cross-view attention · 0.8quadrilinear interpolation · 0.7deferred architecture · 0.7
YearPublicationVenuePosition
2026 Probabilistic modeling of disparity uncertainty for robust and efficient stereo matching
Wenxiao Cai, Dongting Hu, Ruoyan Yin, Jiankang Deng, Huan Fu, Wankou Yang, Mingming Gong
Pattern Recognit.2
2026 AniFaceDiff: Animating stylized avatars via parametric conditioned diffusion models
abstract
Animating stylized head avatars with dynamic poses and expressions has become an important focus in recent research due to its broad range of applications (e.g. VR/AR, film and animation, privacy protection). Previous research has made significant progress by training controllable generative models to animate the reference avatar using the target pose and expression. However, existing portrait animation methods are mostly trained using human faces, making them struggle to generalize to stylized avatar references such as cartoon and painting. Moreover, the mechanisms used to animate avatars—namely, to control the pose and expression of the reference—often inadvertently introduce unintended features—such as facial shape—from the target, while also causing a loss of intended features, like expression-related details. This paper proposes AniFaceDiff, a Stable Diffusion (Rombach et al., 2022)-based method with a new conditioning module for animating stylized avatars. First, we propose a refined spatial conditioning approach by Facial Alignment to minimize identity mismatches, particularly between stylized avatars and human faces. Then, we introduce an Expression Adapter that incorporates additional cross-attention layers to address the potential loss of expression-related information. Extensive experiments demonstrate that our method achieves state-of-the-art performance, particularly in the most challenging out-of-domain stylized avatar animation, i.e., domains unseen during training. It delivers superior image quality, identity preservation, and expression accuracy. This work enhances the quality of virtual stylized avatar animation for constructive and responsible applications. To promote ethical use in virtual environments, we contribute to the advancement of detection for generative content by evaluating state-of-the-art detectors, highlighting potential areas for improvement, and suggesting solutions.
Sachith Seneviratne, Wei Wang 0133, Dongting Hu, Sanjay Saha, Md. Tarek Hasan, Sanka Rasnayaka, Tamasha Malepathirana, Mingming Gong, Saman K. Halgamuge
Pattern Recognit.4
2025 SnapGen: Taming High-Resolution Text-to-Image Models for Mobile Devices with Efficient Architectures and Training
abstract
Existing text-to-image (T2I) diffusion models face several limitations, including large model sizes, slow runtime, and low-quality generation on mobile devices. This paper aims to address all of these challenges by developing an extremely small and fast T2I model that generates high-resolution and high-quality images on mobile platforms. We propose several techniques to achieve this goal. First, we systematically examine the design choices of the network architecture to reduce model parameters and latency, while ensuring high-quality generation. Second, to further improve generation quality, we employ cross-architecture knowledge distillation from a much larger model, using a multi-level approach to guide the training of our model from scratch. Third, we enable a few-step generation by integrating adversarial guidance with knowledge distillation. For the first time, our model SnapGen, demonstrates the generation of 10242px images on a mobile device around 1.4 seconds. On ImageNet-1K, our model, with only 372M parameters, achieves an FID of 2.06 for 2562px generation. On T2I benchmarks (i.e., GenEval and DPG-Bench), our model with merely 379M parameters, surpasses large-scale models with billions of parameters at a significantly smaller size (e.g., 7× smaller than SDXL, 14× smaller than IF-XL).
Jierun Chen, Dongting Hu, Xijie Huang, Huseyin Coskun, Arpit Sahni, Aarush Gupta, Anujraaj Goyal, Dishani Lahiri, Yerlan Idelbayev, Junli Cao, Yanyu Li, Kwang-Ting Cheng, Shueng-Han Gary Chan, Mingming Gong, Sergey Tulyakov, Anil Kag, Yanwu Xu 0003, Jian Ren 0005
CVPR2
2025 Stochastic Diffusion: A Diffusion Based Model for Stochastic Time Series Forecasting
abstract
Recent successes in diffusion probabilistic models have demonstrated their strength in modeling and generating different types of data, paving the way for their application in generative time series forecasting.However, most existing diffusion based approaches rely on sequential models and unimodal latent variables to capture global dependencies and model entire observable data, resulting in difficulties when it comes to highly stochastic time series data.In this paper, we propose a novel Stochastic Diffusion (StochDiff) model that integrates the diffusion process into time series modeling stage and utilizes the representational power of the stochastic latent spaces to capture the variability of the stochastic time series data.Specifically, the model applies diffusion module at each time step within the sequential framework and learns a step-wise, datadriven prior for generative diffusion process.These features enable the model to effectively capture complex temporal dynamics and the multi-modal nature of the highly stochastic time series data.Through extensive experiments on real-world datasets, we demonstrate the effectiveness of our proposed model for probabilistic time series forecasting, particularly in scenarios with high stochasticity.Additionally, with a real-world surgical use case, we highlight the model's potential in a medical application.
Yuansan Liu, Sudanthi N. R. Wijewickrema, Dongting Hu, Christofer Bester, Stephen J. O'Leary, James Bailey 0001
KDD (2)3
2025 STBA: Towards Evaluating the Robustness of DNNs for Query-Limited Black-Box Scenario
abstract
Extensive studies have revealed that deep neural networks (DNNs) are vulnerable to adversarial attacks, especially black-box ones, which can heavily threaten the DNNs deployed in the real world. Many attack techniques have been proposed to explore the vulnerability of DNNs and further help to improve their robustness. Despite the significant progress made recently, existing black-box attack methods still suffer from unsatisfactory performance due to the vast number of queries needed to optimize desired perturbations. Besides, the other critical challenge is that adversarial examples built in a noise-adding manner are abnormal and struggle to successfully attack robust models, whose robustness is enhanced by adversarial training against small perturbations. There is no doubt that these two issues mentioned above will significantly increase the risk of exposure and result in a failure to dig deeply into the vulnerability of DNNs. Hence, it is necessary to evaluate DNNs' fragility sufficiently under query-limited settings in a non-additional way. In this paper, we propose the Spatial Transform Black-box Attack (STBA), a novel framework to craft formidable adversarial examples in the query-limited scenario. Specifically, STBA introduces a flow field to the high-frequency part of clean images to generate adversarial examples and adopts the following two processes to enhance their naturalness and significantly improve the query efficiency: a) we apply an estimated flow field to the high-frequency part of clean images to generate adversarial examples instead of introducing external noise to the benign image, and b) we leverage an efficient gradient estimation method based on a batch of samples to optimize such an ideal flow field under query-limited settings. Compared to existing score-based black-box baselines, extensive experiments indicated that STBA could effectively improve the imperceptibility of the adversarial examples and remarkably boost the attack success rate under query-limited settings.
Renyang Liu 0001, Kwok-Yan Lam, Wei Zhou 0011, Sixing Wu, Jun Zhao 0007, Dongting Hu, Mingming Gong
IEEE Trans. Multim.6
2024 In-N-Out: Lifting 2D Diffusion Prior for 3D Object Removal via Tuning-Free Latents Alignment
abstract
Neural representations for 3D scenes have made substantial advancements recently, yet object removal remains a challenging yet practical issue, due to the absence of multi-view supervision over occluded areas. Diffusion Models (DMs), trained on extensive 2D images, show diverse and high-fidelity generative capabilities in the 2D domain. However, due to not being specifically trained on 3D data, their application to multi-view data often exacerbates inconsistency, hence impacting the overall quality of the 3D output. To address these issues, we introduce "In-N-Out", a novel approach that begins by inpainting a prior, i.e., the occluded area from a single view using DMs, followed by outstretching it to create multi-view inpaintings via latents alignments. Our analysis identifies that the variability in DMs' outputs mainly arises from initially sampled latents and intermediate latents predicted in the denoising process. We explicitly align of initial latents using a Neural Radiance Field (NeRF) to establish a consistent foundational structure in the inpainted area, complemented by an implicit alignment of intermediate latents through cross-view attention during the denoising phases, enhancing appearance consistency across views. To further enhance rendering results, we apply a patch-based hybrid loss to optimize NeRF. We demonstrate that our techniques effectively mitigate the challenges posed by inconsistencies in DMs and substantially improve the fidelity and coherence of inpainted 3D representations.
Dongting Hu, Huan Fu, Jiaxian Guo, Liuhua Peng, Tingjin Chu, Feng Liu 0003, Tongliang Liu, Mingming Gong
NeurIPS1
2023 Multiscale Representation for Real-Time Anti-Aliasing Neural Rendering
abstract
The rendering scheme in neural radiance field (NeRF) is effective in rendering a pixel by casting a ray into the scene. However, NeRF yields blurred rendering results when the training images are captured at non-uniform scales, and produces aliasing artifacts if the test images are taken in distant views. To address this issue, Mip-NeRF proposes a multiscale representation as a conical frustum to encode scale information. Nevertheless, this approach is only suitable for offline rendering since it relies on integrated positional encoding (IPE) to query a multilayer perceptron (MLP). To overcome this limitation, we propose mip voxel grids (Mip-VoG), an explicit multiscale representation with a deferred architecture for real-time anti-aliasing rendering. Our approach includes a density Mip-VoG for scene geometry and a feature Mip-VoG with a small MLP for view-dependent color. Mip-VoG represents scene scale using the level of detail (LOD) derived from ray differentials and uses quadrilinear interpolation to map a queried 3D location to its features and density from two neighboring down-sampled voxel grids. To our knowledge, our approach is the first to offer multiscale training and real-time anti-aliasing rendering simultaneously. We conducted experiments on multiscale dataset, results show that our approach outperforms state-of-the-art real-time rendering baselines.
Dongting Hu, Zhenkai Zhang 0001, Tingbo Hou, Tongliang Liu, Huan Fu, Mingming Gong
ICCV1
2022 Uncertainty Quantification in Depth Estimation via Constrained Ordinal Regression
Dongting Hu, Liuhua Peng, Tingjin Chu, Yinian Mao, Howard D. Bondell, Mingming Gong
ECCV (2)1