EDBT 2026 Demo / reviewers in the wild / expert
Dongting Hu
dblp:159/6788
· DBLP profile ↗
8ranked-venue papers
3as first author
8since 2021 · last 2026
0009-0007-2119-1829ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 3 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Generative modeling · 33% Efficient and distributed learning · 28% Time series and sequential data · 14% | |
| Network and information security
1 paper |
Security and privacy of machine learning · 100% | |
| Computer graphics and multimedia
1 paper |
Rendering · 100% |
Topics — the 18 heaviest of 18, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling
diffusion model |
2.5 | 3 | 2025 | Stochastic Diffusion: A Diffusion Based Model for Stochastic Time Series Forecasting · KDD (2) 2025 SnapGen: Taming High-Resolution Text-to-Image Models for Mobile Devices with Efficient Architectures and Training · CVPR 2025 In-N-Out: Lifting 2D Diffusion Prior for 3D Object Removal via Tuning-Free Latents Alignment · NeurIPS 2024 |
Machine learning › Efficient and distributed learning
inference efficiency |
0.9 | 1 | 2025 | SnapGen: Taming High-Resolution Text-to-Image Models for Mobile Devices with Efficient Architectures and Training · CVPR 2025 |
Machine learning › Efficient and distributed learning › model compression
knowledge distillation |
0.9 | 1 | 2025 | SnapGen: Taming High-Resolution Text-to-Image Models for Mobile Devices with Efficient Architectures and Training · CVPR 2025 |
Machine learning › Efficient and distributed learning › model deployment
mobile deployment |
0.9 | 1 | 2025 | SnapGen: Taming High-Resolution Text-to-Image Models for Mobile Devices with Efficient Architectures and Training · CVPR 2025 |
Machine learning › Efficient and distributed learning
model compression |
0.9 | 1 | 2025 | SnapGen: Taming High-Resolution Text-to-Image Models for Mobile Devices with Efficient Architectures and Training · CVPR 2025 |
Machine learning › Time series and sequential data › time series modeling
probabilistic forecasting |
0.9 | 1 | 2025 | Stochastic Diffusion: A Diffusion Based Model for Stochastic Time Series Forecasting · KDD (2) 2025 |
Machine learning › Generative modeling › diffusion model
text-to-image generation |
0.9 | 1 | 2025 | SnapGen: Taming High-Resolution Text-to-Image Models for Mobile Devices with Efficient Architectures and Training · CVPR 2025 |
Machine learning › Time series and sequential data › time series analysis
time series forecasting |
0.9 | 1 | 2025 | Stochastic Diffusion: A Diffusion Based Model for Stochastic Time Series Forecasting · KDD (2) 2025 |
Security and privacy of machine learning
adversarial attack |
0.9 | 1 | 2025 | STBA: Towards Evaluating the Robustness of DNNs for Query-Limited Black-Box Scenario · IEEE Trans. Multim. 2025 |
Security and privacy of machine learning
adversarial example |
0.9 | 1 | 2025 | STBA: Towards Evaluating the Robustness of DNNs for Query-Limited Black-Box Scenario · IEEE Trans. Multim. 2025 |
Machine learning › Generative modeling › generative adversarial network
3d-aware image synthesis |
0.8 | 1 | 2024 | In-N-Out: Lifting 2D Diffusion Prior for 3D Object Removal via Tuning-Free Latents Alignment · NeurIPS 2024 |
Computer vision › 3D vision
neural radiance field |
0.8 | 1 | 2024 | In-N-Out: Lifting 2D Diffusion Prior for 3D Object Removal via Tuning-Free Latents Alignment · NeurIPS 2024 |
Machine learning › Deep learning architectures and training
multi-scale representation |
0.7 | 1 | 2023 | Multiscale Representation for Real-Time Anti-Aliasing Neural Rendering · ICCV 2023 |
Rendering
neural radiance fields |
0.7 | 1 | 2023 | Multiscale Representation for Real-Time Anti-Aliasing Neural Rendering · ICCV 2023 |
Computer vision › 3D vision
depth estimation |
0.6 | 1 | 2022 | Uncertainty Quantification in Depth Estimation via Constrained Ordinal Regression · ECCV (2) 2022 |
Machine learning › Trustworthy machine learning
uncertainty estimation |
0.6 | 1 | 2022 | Uncertainty Quantification in Depth Estimation via Constrained Ordinal Regression · ECCV (2) 2022 |
Machine learning › Trustworthy machine learning
robustness evaluation |
0.3 | 1 | 2025 | STBA: Towards Evaluating the Robustness of DNNs for Query-Limited Black-Box Scenario · IEEE Trans. Multim. 2025 |
Computer vision › 3D vision › multi-view geometry
multi-view consistency |
0.2 | 1 | 2024 | In-N-Out: Lifting 2D Diffusion Prior for 3D Object Removal via Tuning-Free Latents Alignment · NeurIPS 2024 |
Methods — techniques the papers use, named apart from their topics
gradient estimation · 1.7flow field estimation · 1.7sequential model · 0.9few-step generation · 0.9diffusion probabilistic model · 0.9cross-architecture knowledge distillation · 0.9adversarial guidance · 0.9patch-based hybrid loss · 0.8latent alignment · 0.8cross-view attention · 0.8quadrilinear interpolation · 0.7deferred architecture · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Probabilistic modeling of disparity uncertainty for robust and efficient stereo matching
Wenxiao Cai, Dongting Hu, Ruoyan Yin, Jiankang Deng, Huan Fu, Wankou Yang, Mingming Gong |
Pattern Recognit. | 2 |
| 2026 | AniFaceDiff: Animating stylized avatars via parametric conditioned diffusion modelsabstractAnimating stylized head avatars with dynamic poses and expressions has become an important focus in recent research due to its broad range of applications (e.g. VR/AR, film and animation, privacy protection). Previous research has made significant progress by training controllable generative models to animate the reference avatar using the target pose and expression. However, existing portrait animation methods are mostly trained using human faces, making them struggle to generalize to stylized avatar references such as cartoon and painting. Moreover, the mechanisms used to animate avatars—namely, to control the pose and expression of the reference—often inadvertently introduce unintended features—such as facial shape—from the target, while also causing a loss of intended features, like expression-related details. This paper proposes AniFaceDiff, a Stable Diffusion (Rombach et al., 2022)-based method with a new conditioning module for animating stylized avatars. First, we propose a refined spatial conditioning approach by Facial Alignment to minimize identity mismatches, particularly between stylized avatars and human faces. Then, we introduce an Expression Adapter that incorporates additional cross-attention layers to address the potential loss of expression-related information. Extensive experiments demonstrate that our method achieves state-of-the-art performance, particularly in the most challenging out-of-domain stylized avatar animation, i.e., domains unseen during training. It delivers superior image quality, identity preservation, and expression accuracy. This work enhances the quality of virtual stylized avatar animation for constructive and responsible applications. To promote ethical use in virtual environments, we contribute to the advancement of detection for generative content by evaluating state-of-the-art detectors, highlighting potential areas for improvement, and suggesting solutions. Sachith Seneviratne, Wei Wang 0133, Dongting Hu, Sanjay Saha, Md. Tarek Hasan, Sanka Rasnayaka, Tamasha Malepathirana, Mingming Gong, Saman K. Halgamuge |
Pattern Recognit. | 4 |
| 2025 | SnapGen: Taming High-Resolution Text-to-Image Models for Mobile Devices with Efficient Architectures and TrainingabstractExisting text-to-image (T2I) diffusion models face several limitations, including large model sizes, slow runtime, and low-quality generation on mobile devices. This paper aims to address all of these challenges by developing an extremely small and fast T2I model that generates high-resolution and high-quality images on mobile platforms. We propose several techniques to achieve this goal. First, we systematically examine the design choices of the network architecture to reduce model parameters and latency, while ensuring high-quality generation. Second, to further improve generation quality, we employ cross-architecture knowledge distillation from a much larger model, using a multi-level approach to guide the training of our model from scratch. Third, we enable a few-step generation by integrating adversarial guidance with knowledge distillation. For the first time, our model SnapGen, demonstrates the generation of 10242px images on a mobile device around 1.4 seconds. On ImageNet-1K, our model, with only 372M parameters, achieves an FID of 2.06 for 2562px generation. On T2I benchmarks (i.e., GenEval and DPG-Bench), our model with merely 379M parameters, surpasses large-scale models with billions of parameters at a significantly smaller size (e.g., 7× smaller than SDXL, 14× smaller than IF-XL). Jierun Chen, Dongting Hu, Xijie Huang, Huseyin Coskun, Arpit Sahni, Aarush Gupta, Anujraaj Goyal, Dishani Lahiri, Yerlan Idelbayev, Junli Cao, Yanyu Li, Kwang-Ting Cheng, Shueng-Han Gary Chan, Mingming Gong, Sergey Tulyakov, Anil Kag, Yanwu Xu 0003, Jian Ren 0005 |
CVPR | 2 |
| 2025 | Stochastic Diffusion: A Diffusion Based Model for Stochastic Time Series ForecastingabstractRecent successes in diffusion probabilistic models have demonstrated their strength in modeling and generating different types of data, paving the way for their application in generative time series forecasting.However, most existing diffusion based approaches rely on sequential models and unimodal latent variables to capture global dependencies and model entire observable data, resulting in difficulties when it comes to highly stochastic time series data.In this paper, we propose a novel Stochastic Diffusion (StochDiff) model that integrates the diffusion process into time series modeling stage and utilizes the representational power of the stochastic latent spaces to capture the variability of the stochastic time series data.Specifically, the model applies diffusion module at each time step within the sequential framework and learns a step-wise, datadriven prior for generative diffusion process.These features enable the model to effectively capture complex temporal dynamics and the multi-modal nature of the highly stochastic time series data.Through extensive experiments on real-world datasets, we demonstrate the effectiveness of our proposed model for probabilistic time series forecasting, particularly in scenarios with high stochasticity.Additionally, with a real-world surgical use case, we highlight the model's potential in a medical application. Yuansan Liu, Sudanthi N. R. Wijewickrema, Dongting Hu, Christofer Bester, Stephen J. O'Leary, James Bailey 0001 |
KDD (2) | 3 |
| 2025 | STBA: Towards Evaluating the Robustness of DNNs for Query-Limited Black-Box ScenarioabstractExtensive studies have revealed that deep neural networks (DNNs) are vulnerable to adversarial attacks, especially black-box ones, which can heavily threaten the DNNs deployed in the real world. Many attack techniques have been proposed to explore the vulnerability of DNNs and further help to improve their robustness. Despite the significant progress made recently, existing black-box attack methods still suffer from unsatisfactory performance due to the vast number of queries needed to optimize desired perturbations. Besides, the other critical challenge is that adversarial examples built in a noise-adding manner are abnormal and struggle to successfully attack robust models, whose robustness is enhanced by adversarial training against small perturbations. There is no doubt that these two issues mentioned above will significantly increase the risk of exposure and result in a failure to dig deeply into the vulnerability of DNNs. Hence, it is necessary to evaluate DNNs' fragility sufficiently under query-limited settings in a non-additional way. In this paper, we propose the Spatial Transform Black-box Attack (STBA), a novel framework to craft formidable adversarial examples in the query-limited scenario. Specifically, STBA introduces a flow field to the high-frequency part of clean images to generate adversarial examples and adopts the following two processes to enhance their naturalness and significantly improve the query efficiency: a) we apply an estimated flow field to the high-frequency part of clean images to generate adversarial examples instead of introducing external noise to the benign image, and b) we leverage an efficient gradient estimation method based on a batch of samples to optimize such an ideal flow field under query-limited settings. Compared to existing score-based black-box baselines, extensive experiments indicated that STBA could effectively improve the imperceptibility of the adversarial examples and remarkably boost the attack success rate under query-limited settings. Renyang Liu 0001, Kwok-Yan Lam, Wei Zhou 0011, Sixing Wu, Jun Zhao 0007, Dongting Hu, Mingming Gong |
IEEE Trans. Multim. | 6 |
| 2024 | In-N-Out: Lifting 2D Diffusion Prior for 3D Object Removal via Tuning-Free Latents AlignmentabstractNeural representations for 3D scenes have made substantial advancements recently, yet object removal remains a challenging yet practical issue, due to the absence of multi-view supervision over occluded areas. Diffusion Models (DMs), trained on extensive 2D images, show diverse and high-fidelity generative capabilities in the 2D domain. However, due to not being specifically trained on 3D data, their application to multi-view data often exacerbates inconsistency, hence impacting the overall quality of the 3D output. To address these issues, we introduce "In-N-Out", a novel approach that begins by inpainting a prior, i.e., the occluded area from a single view using DMs, followed by outstretching it to create multi-view inpaintings via latents alignments. Our analysis identifies that the variability in DMs' outputs mainly arises from initially sampled latents and intermediate latents predicted in the denoising process. We explicitly align of initial latents using a Neural Radiance Field (NeRF) to establish a consistent foundational structure in the inpainted area, complemented by an implicit alignment of intermediate latents through cross-view attention during the denoising phases, enhancing appearance consistency across views. To further enhance rendering results, we apply a patch-based hybrid loss to optimize NeRF. We demonstrate that our techniques effectively mitigate the challenges posed by inconsistencies in DMs and substantially improve the fidelity and coherence of inpainted 3D representations. Dongting Hu, Huan Fu, Jiaxian Guo, Liuhua Peng, Tingjin Chu, Feng Liu 0003, Tongliang Liu, Mingming Gong |
NeurIPS | 1 |
| 2023 | Multiscale Representation for Real-Time Anti-Aliasing Neural RenderingabstractThe rendering scheme in neural radiance field (NeRF) is effective in rendering a pixel by casting a ray into the scene. However, NeRF yields blurred rendering results when the training images are captured at non-uniform scales, and produces aliasing artifacts if the test images are taken in distant views. To address this issue, Mip-NeRF proposes a multiscale representation as a conical frustum to encode scale information. Nevertheless, this approach is only suitable for offline rendering since it relies on integrated positional encoding (IPE) to query a multilayer perceptron (MLP). To overcome this limitation, we propose mip voxel grids (Mip-VoG), an explicit multiscale representation with a deferred architecture for real-time anti-aliasing rendering. Our approach includes a density Mip-VoG for scene geometry and a feature Mip-VoG with a small MLP for view-dependent color. Mip-VoG represents scene scale using the level of detail (LOD) derived from ray differentials and uses quadrilinear interpolation to map a queried 3D location to its features and density from two neighboring down-sampled voxel grids. To our knowledge, our approach is the first to offer multiscale training and real-time anti-aliasing rendering simultaneously. We conducted experiments on multiscale dataset, results show that our approach outperforms state-of-the-art real-time rendering baselines. Dongting Hu, Zhenkai Zhang 0001, Tingbo Hou, Tongliang Liu, Huan Fu, Mingming Gong |
ICCV | 1 |
| 2022 | Uncertainty Quantification in Depth Estimation via Constrained Ordinal Regression
Dongting Hu, Liuhua Peng, Tingjin Chu, Yinian Mao, Howard D. Bondell, Mingming Gong |
ECCV (2) | 1 |