Tongyan Hua

dblp:340/7382 · DBLP profile ↗
← Back
3ranked-venue papers
2as first author
3since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
3D vision · 49% Generative modeling · 27% Robot navigation and mapping · 24%

Topics — the 7 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision › 3d generation
3d scene generation
0.912025
Sat2City: 3D City Generation from a Single Satellite Image with Cascaded Latent Diffusion · ICCV 2025
Machine learning › Generative modeling
diffusion model
0.912025
Sat2City: 3D City Generation from a Single Satellite Image with Cascaded Latent Diffusion · ICCV 2025
Machine learning › Generative modeling › diffusion model
latent diffusion model
0.912025
Sat2City: 3D City Generation from a Single Satellite Image with Cascaded Latent Diffusion · ICCV 2025
Computer vision › 3D vision
implicit neural representation
0.812024
Benchmarking Implicit Neural Representation and Geometric Rendering in Real-Time RGB-D SLAM · CVPR 2024
Computer vision › 3D vision
neural radiance field
0.812024
Benchmarking Implicit Neural Representation and Geometric Rendering in Real-Time RGB-D SLAM · CVPR 2024
Robotics › Robot navigation and mapping › SLAM › visual SLAM
RGB-D SLAM
0.812024
Benchmarking Implicit Neural Representation and Geometric Rendering in Real-Time RGB-D SLAM · CVPR 2024
Robotics › Robot navigation and mapping
SLAM
0.812024
Benchmarking Implicit Neural Representation and Geometric Rendering in Real-Time RGB-D SLAM · CVPR 2024

Methods — techniques the papers use, named apart from their topics

sparse voxel grid · 0.9re-hash · 0.9inverse sampling · 0.9cascaded latent diffusion · 0.9implicit neural representation · 0.8geometric rendering · 0.8
YearPublicationVenuePosition
2025 Sat2City: 3D City Generation from a Single Satellite Image with Cascaded Latent Diffusion
abstract
Recent advancements in generative models have enabled 3D urban scene generation from satellite imagery, unlocking promising applications in gaming, digital twins, and beyond. However, most existing methods rely heavily on neural rendering techniques, which hinder their ability to produce detailed 3D structures on a broader scale, largely due to the inherent structural ambiguity derived from relatively limited 2D observations. To address this challenge, we propose Sat2City, a novel framework that synergizes the representational capacity of sparse voxel grids with latent diffusion models, tailored specifically for our novel 3D city dataset. Our approach is enabled by three key components: (1) A cascaded latent diffusion framework that progressively recovers 3D city structures from satellite imagery, (2) a Re-Hash operation at its Variational Autoencoder (VAE) bottleneck to compute multi-scale feature grids for stable appearance optimization and (3) an inverse sampling strategy enabling implicit supervision for smooth appearance transitioning.To overcome the challenge of collecting real-world city-scale 3D models with high-quality geometry and appearance, we introduce a dataset of synthesized large-scale 3D cities paired with satellite-view height maps. Validated on this dataset, our framework generates detailed 3D structures from a single satellite image, achieving superior fidelity compared to existing city generation models.
Tongyan Hua, Lutao Jiang, Ying-Cong Chen, Wufan Zhao
ICCV1
2024 Benchmarking Implicit Neural Representation and Geometric Rendering in Real-Time RGB-D SLAM
abstract
Implicit neural representation (INR), in combination with geometric rendering, has recently been employed in real-time dense RGB-D SLAM. Despite active research endeavors being made, there lacks a unified protocol for fair evaluation, impeding the evolution of this area. In this work, we establish, to our knowledge, the first open-source benchmark framework to evaluate the performance of a wide spectrum of commonly used INRs and rendering functions for mapping and localization. The goal of our benchmark is to 1) gain an intuition of how different INRs and rendering functions impact mapping and localization and 2) establish a unified evaluation protocol w.r. t. the design choices that may impact the mapping and localization. With the framework, we conduct a large suite of experiments, offering various insights in choosing the INRs and geometric rendering functions: for example, the dense feature grid outperforms other INRs (e.g. tri-plane and hash grid), even when geometric and color features are jointly encoded for memory efficiency. To extend the findings into the practical scenario, a hybrid encoding strategy is proposed to bring the best of the accuracy and completion from the grid-based and decomposition-based INRs. We further propose explicit hybrid encoding for high-fidelity dense grid mapping to comply with the RGB-D SLAM system that puts the premise on robustness and computation efficiency.
Tongyan Hua, Lin Wang 0025
CVPR1
2023 CDHD: Contrastive Dreamer for Hint Distillation
abstract
Replaying previous training data is the most effective approach for Class-Incremental Learning (CIL), with its performance bounded by data availability. Therefore, many recent studies consider the Data-Free Class-Incremental Learning (DFCIL) problem that requires no previous data. However, the existing methods do not consider synthesising data of heterogeneity, thus limiting models’ generalizability. Such homogenous images further hinder the knowledge distillation process when regularising only the deeper layers close to the output, resulting in catastrophic forgetting. To address these issues, we present CDHD: a contrastive dreamer for hint distillation. Our approach starts with training a generator for data synthesis. A model inversion technique is introduced to obtain a generator capable of producing heterogeneous images from the classifier by imposing the ContRastive Loss. Moreover, to better transfer the previous knowledge to the current model, we force the teacher network to provide more general knowledge to its students by enforcing the Hint Loss in shallower layers rather than only in deeper ones. We validate the performance of CDHD on CIFAR-100 for various tasks and compare it against the SOTA baseline for DFCIL, demonstrating our superiorities and thus constituting a new benchmark.
Tongyan Hua, Wenming Yang, Qingmin Liao
ICASSP2