Dongjin Huang

dblp:14/10437 · DBLP profile ↗
← Back
34ranked-venue papers
22as first author
24since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 28 · 20 first-author · 19 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 ImageNetion: Retrieval-as-Policy for Creative Gollin Figure Completion via Generative Feedback
abstract
Human vision excels at deriving polysemous meanings from sparse structures, such as those in the Gollin Test. However, existing models struggle to balance semantic retrieval with structural consistency under extreme sparsity. We propose ImageNetion, a framework that leverages retrieval-driven generation for creative completion. The core algorithm consists of two stages: (1) Static Training, which establishes basic sketch-text semantic associations through contrastive learning; and (2) Reinforcement Learning, where retrieval is modeled as a decision process. In this stage, a diffusion model integrated with ControlNet and LoRA serves as a feedback source. A self-feedback mechanism enables the recognition model to learn generator priors unsupervised, enhancing both retrieval rationality and completion quality. Evaluated on a new 400-class synthetic dataset and zero-shot hand-drawn tests, ImageNetion outperforms traditional sketch inpainting in structural fidelity and Top-K semantic consistency. Furthermore, user studies validate the framework’s superiority in creative diversity and human preference.
Dongjin Huang, Jiantao Qu, Qinghang Wu
ICMR1
2026 Guided by structure: boundary-aware modeling for moment retrieval and highlight detection
Youxian Di, Zhenzhen Jin, Youdong Ding, Dongjin Huang
Mach. Vis. Appl.6
2026 Joint pluralistic generation and realistic inpainting of occluded facial images
Yongsheng Shi, Dongjin Huang, Jinhua Liu 0002, Jiantao Qu, Wen Tang 0004
Pattern Recognit.2
2025 DreamDancer: Music-Driven Dance Video Intelligent Generation
Dongjin Huang, Jiyu Qian, Wenyun Tu
PRCV (6)1
2025 ClotheDreamer: Text-guided garment generation with 3D gaussians
Junshu Tang, Chu Zheng, Chengjie Wang 0001, Dongjin Huang
Appl. Intell.6
2025 SketchTailor: Lightweight sketch-driven modeling for high-fidelity garment pattern reconstruction
Dongjin Huang, Jiantao Qu, Ansheng Wang, Yixiang Tang
Comput. Graph.1
2025 MonoNeRF-DDP: Neural radiance fields from monocular endoscopic images with dense depth priors
Jinhua Liu 0002, Dongjin Huang, Yongsheng Shi, Jiantao Qu
Comput. Graph.2
2025 Hi3DFace: High-Realistic 3D Face Reconstruction From a Single Occluded Image
abstract
Abstract We propose Hi3DFace, a novel framework for simultaneous de‐occlusion and high‐fidelity 3D face reconstruction. To address real‐world occlusions, we construct a diverse facial dataset by simulating common obstructions and present TMANet, a transformer‐based multi‐scale attention network that effectively removes occlusions and restores clean face images. For the 3D face reconstruction stage, we propose a coarse‐medium‐fine self‐supervised scheme. In the coarse reconstruction pipeline, we adopt a face regression network to predict 3DMM coefficients for generating a smooth 3D face. In the medium‐scale reconstruction pipeline, we propose a novel depth displacement network, DDFTNet, to remove noise and restore rich details to the smooth 3D geometry. In the fine‐scale reconstruction pipeline, we design a GCN (graph convolutional network) refiner to enhance the fidelity of 3D textures. Additionally, a light‐aware network (LightNet) is proposed to distil lighting parameters, ensuring illumination consistency between reconstructed 3D faces and input images. Extensive experimental results demonstrate that the proposed Hi3DFace significantly outperforms state‐of‐the‐art reconstruction methods on four public datasets, and five constructed occlusion‐type datasets. Hi3DFace achieves robustness and effectiveness in removing occlusions and reconstructing 3D faces from real‐world occluded facial images.
Dongjin Huang, Yongsheng Shi, Jiantao Qu, Jinhua Liu 0002, Wen Tang 0004
Comput. Graph. Forum1
2024 YOLOv8-MGH: Dense Crowd Object Detection
Dongjin Huang, Jiantao Qu
CGI (1)1
2024 TexDreamer: Towards Zero-Shot High-Fidelity 3D Human Texture Generation
Junshu Tang, Jiangning Zhang, Weijian Cao, Chengjie Wang 0001, Yunsheng Wu, Dongjin Huang
ECCV (47)9
2024 Vrefine: A Self-Refinement Approach for Enhanced Clarity and Quality in Text-to-Speech Models
abstract
Driven by advancements in the Large Language Model (LLM), there has been significant global attention on integrating the Generative Pre-trained Transformer (GPT) concept into Text-to-Speech (TTS) technologies. However, existing TTS models face issues such as inconsistent training data quality and reliance on autoregressive models, which makes controlling the quality of audio generation challenging. To address these challenges, this study introduces a novel TTS framework known as TTS-Vrefine, which is a new type of TTS architecture based on a self-feedback mechanism, aimed at enhancing the inference capabilities of the model and the quality of generated audio. The Vrefine framework iteratively refines the output, allowing the TTS system to self-train using its own generated data, significantly improving the clarity and quality of the audio. It expands the training dataset and enhances the self-optimization potential of the audio generation model, reducing the Word Error Rate (WER) of the base model by 2.4%, increasing the Perceptual Evaluation of Speech Quality (PESQ) by 2.22%, and improving the Short-Time Objective Intelligibility (STOI) by 4.55%. Additionally, the architecture optimizes the use of low-quality resources through a self-refinement mechanism, effectively expanding the training dataset.
Dongjin Huang, Yuhua Liu, Jixu Qian
SMC1
2024 Mesh-controllable multi-level-of-detail text-to-3D generation
Dongjin Huang, Xinghan Huang, Jiantao Qu
Comput. Graph.1
2024 Learning to Play Guitar with Robotic Hands
abstract
Abstract Playing the guitar is a dexterous human skill that poses significant challenges in computer graphics and robotics due to the precision required in finger positioning and coordination between hands. Current methods often rely on motion capture data to replicate specific guitar playing segments, which restricts the range of performances and demands intricate post‐processing. In this paper, we introduce a novel reinforcement learning model that can play the guitar using robotic hands, without the need for motion capture datasets, from input tablatures. To achieve this, we divide the simulation task for playing guitar into three stages. (a): for an input tablature, we first generate corresponding fingerings that align with human habits. (b): based on the generated fingerings as the guidance, we train a neural network for controlling the fingers of the left hand using deep reinforcement learning, and (c): we generate plucking movements for the right hand based on inverse kinematics according to the tablature. We evaluate our method by employing precision, recall, and F1 scores as quantitative metrics to thoroughly assess its performance in playing musical notes. In addition, we conduct qualitative analysis through user studies to evaluate the visual and auditory effects of guitar performance. The results demonstrate that our model excels in playing most moderately difficult and easier musical pieces, accurately playing nearly all notes.
Chaoyi Luo, Pengbin Tang, Dongjin Huang
Comput. Graph. Forum4
2024 GPSwap: High-resolution face swapping based on StyleGAN prior
abstract
Abstract Existing high‐resolution face‐swapping works are still challenges in preserving identity consistency while maintaining high visual quality. We present a novel high‐resolution face‐swapping method GPSwap, which is based on StyleGAN prior. To better preserves identity consistency, the proposed facial feature recombination network fully leverages the properties of both w space and encoders to decouple identities. Furthermore, we presents the image reconstruction module aligns and blends images in FS space, which further supplements facial details and achieves natural blending. It not only improves image resolution but also optimizes visual quality. Extensive experiments and user studies demonstrate that GPSwap is superior to state‐of‐the‐art high‐resolution face‐swapping methods in terms of image quality and identity consistency. In addition, GPSwap saves nearly 80% of training costs compared to other high‐resolution face‐swapping works.
Dongjin Huang, Chuanman Liu, Jinhua Liu 0002
Comput. Animat. Virtual Worlds1
2024 Self-supervised learning for fine-grained monocular 3D face reconstruction in the wild
Dongjin Huang, Yongsheng Shi, Jinhua Liu 0002, Wen Tang 0004
Multim. Syst.1
2023 Staged Transformer Network with Color Harmonization for Image Outpainting
Wangyidai Lv, Dongjin Huang, Youdong Ding
CGI3
2023 TG-Dance: TransGAN-Based Intelligent Dance Generation with Music
Dongjin Huang, Zhenyan Li, Jinhua Liu 0002
MMM (1)1
2022 DDCNet: A Lightweight Network with Variable Receptive Field for Real-Time Portrait Segmentation in Complex Environment
Dongjin Huang, Jinhua Liu 0002, Yushan Lv
CGI1
2022 Physically-guided Disentangled Implicit Rendering for 3D Face Modeling
abstract
This paper presents a novel Physically-guided Disentangled Implicit Rendering (PhyDIR) framework for highfidelity 3D face modeling. The motivation comes from two observations: Widely-used graphics renderers yield excessive approximations against photo-realistic imaging, while neural rendering methods produce superior appearances but are highly entangled to perceive 3D-aware operations. Hence, we learn to disentangle the implicit rendering via explicit physical guidance, while guaranteeing the properties of: (1) 3D-aware comprehension and (2) high-reality image formation. For the former one, PhyDIR explicitly adopts 3D shading and rasterizing modules to control the renderer, which disentangles the light, facial shape, and viewpoint from neural reasoning. Specifically, PhyDIR proposes a novel multi-image shading strategy to compensate for the monocular limitation, so that the lighting variations are accessible to the neural renderer. For the latter, PhyDIR learns the face-collection implicit texture to avoid ill-posed intrinsic factorization, then leverages a series of consistency losses to constrain the rendering robustness. With the disentangled method, we make 3D face modeling benefit from both kinds of rendering strategies. Extensive experiments on benchmarks show that PhyDIR obtains superior performance than state-of-the-art explicit/implicit methods on geometry/texture modeling.
Zhenyu Zhang 0005, Yanhao Ge, Ying Tai, Weijian Cao, Renwang Chen, Kunlin Liu, Hao Tang 0005, Chengjie Wang 0001, Dongjin Huang
CVPR11
2022 Learning to Restore 3D Face from In-the-Wild Degraded Images
abstract
In-the-wild 3D face modelling is a challenging problem as the predicted facial geometry and texture suffer from a lack of reliable clues or priors, when the input images are degraded. To address such a problem, in this paper we propose a novel Learning to Restore (L2R) 3D face framework for unsupervised high-quality face reconstruction from low-resolution images. Rather than directly refining 2D image appearance, L2R learns to recover fine-grained 3D details on the proxy against degradation via extracting generative facial priors. Concretely, L2R proposes a novel albedo restoration network to model high-quality 3D facial texture, in which the diverse guidance from the pre-trained Generative Adversarial Networks (GANs) is leveraged to complement the lack of input facial clues. With the finer details of the restored 3D texture, L2R then learns displacement maps from scratch to enhance the significant facial structure and geometry. Both of the procedures are mutually optimized with a novel 3D-aware adversarial loss, which further improves the modelling performance and suppresses the potential uncertainty. Extensive experiments on benchmarks show that L2R outperforms state-of-the-art methods under the condition of low-quality inputs, and obtains superior performances than 2D pre-processed modelling approaches with limited 3D proxy.
Zhenyu Zhang 0005, Yanhao Ge, Ying Tai, Chengjie Wang 0001, Hao Tang 0005, Dongjin Huang
CVPR7
2022 Efficient angiography simulation for complex vessels
abstract
Abstract Angiography simulation is a critical step in virtual interventional surgery. However, the blood vessels are always complex and irregular, which poses a significant challenge to the accuracy and efficiency of simulation. In this paper, we present a novel method to simulate contrast media propagation, which can efficiently and accurately deal with various complex vascular structures and the coupling between blood and contrast media. Our method represents the vascular structures by signed distance functions and to impose boundary conditions, and then we compute the boundary volume by Gauss–Kronrod quadrature, which yields more accurate results even in complex vessels. Furthermore, we improve the simulation's initialization efficiency and real‐time force computational efficiency by precomputing the boundary volume values throughout the vascular structure and saving them to a volume map, which can be queried very efficiently during runtime. Moreover, we add term to the multiple‐fluid model to handle the vascular viscosity, which achieves a more realistic effect. Experiments show that our method obtains more smooth and accurate particle states, and significantly improves initialization efficiency. Finally, we invited 12 clinicians to evaluate the algorithm and verify the clinical value of our algorithm.
Dongjin Huang, Shuhua Zhou, Ziyang Zeng, Yongsheng Shi
Comput. Animat. Virtual Worlds1
2022 Virtual reality safety training using deep EEG-net and physiology data
Dongjin Huang, Jinhua Liu 0002, Jinyao Li, Wen Tang 0004
Vis. Comput.1
2021 ADD-Net: Attention U-Net with Dilated Skip Connection and Dense Connected Decoder for Retinal Vessel Segmentation
Dongjin Huang
CGI1
2021 Multi-scale Fusion Attention Network for Polyp Segmentation
Dongjin Huang, Kaili Han, Yongjie Xi, Wenqi Che
ICONIP (6)1
2020 Virtual Reality for Training and Fitness Assessments for Construction Safety
abstract
Reducing accident rate is a primary goal of construction safety. In this paper, we present a large scale study of using virtual reality technology for safety training. Beyond the training, a technology framework is proposed to assess the fitness of construction workers (e.g. suitability of people with underlining health conditions to work under particular construction environments). The new virtual construction system consists of a Brain-Computer Interface (BCI) of electroencephalography (EEG) neural network to capture EEG signals of users during the virtual simulation training continuously to achieve user profiling. For real-time assessment of the accident susceptibility of a worker under various construction environments, a deep learning neural network is trained to process the EEG crops and a clipping training algorithm that classifies small segments of the EEG dataset is used to improve the computational performance of the system. Physiology data of the person during the training, i.e. blood pressure and heart rate, is also recorded. Based on the EEG data and the physiology data, a statistic model is used in the safety assessment framework to set up the risk standard. The study has tested 117 workers who were employed by the construction sites in Shanghai. People who were tested in the risk group were further underwent medical examinations for risk related medical conditions that deemed unsuitable for working in construction sites. Results show six of the nine workers identified by the VR system have been medically confirmed unsuitable, thus, over 80% accuracy of our virtual reality training and assessment system. Our proposed system can be used as a tool for understanding risk conditions of workers and safety training.
Dongjin Huang, Jinyao Li, Wen Tang 0004
CW1
2019 New haptic syringe device for virtual angiography training
Dongjin Huang, Pengbin Tang, Tao Ruan Wan, Wen Tang 0004
Comput. Graph.1
2017 A 3D Tube-Object Centerline Extraction Algorithm Based on Steady Fluid Dynamics
Dongjin Huang, Ruobin Gong, Hejuan Li, Wen Tang 0004, Youdong Ding
ICIG (3)1
2017 Photographic Appearance Enhancement via Detail-Based Dictionary Learning
Shi Tang, Dongjin Huang, Youdong Ding, Lizhuang Ma
J. Comput. Sci. Technol.3
2015 Modeling and Simulation of Multi-frictional Interaction Between Guidewire and Vasculature
Dongjin Huang, Pengbin Tang, Wen Tang 0004, Youdong Ding
ICIG (2)1
2015 A Unified Fidelity Optimization Model for Global Color Transfer
Sheng Du, Dongjin Huang, Youdong Ding, Lizhuang Ma
ICIG (1)3
2012 Real-time simulation of long thin flexible objects in interactivevirtual environments
abstract
Many virtual reality-based applications involve simulations of micro-structures such as hair, fibers and textile yarns, as well as ropes, flexible wires and tubes. In virtual surgery, for example, flexible wires and tubes are common medical instruments and devices. Core to the simulations is the robust physics-based computation of elastic rods. In this paper, we present a volumetric finite element based approach to simulating rod-like objects with real-time performance suitable for interactive virtual environments. A sequence of Cosserat joints (tiny volumetric elastic joints) linked by rigid bar segments are used to compute the elastic rod objects. By construction, each of the joints is equipped with its own mass, degrees of freedom (DOFs) with a small volumetric deformation field to measure deformation energies due to stretching, shearing, bending, and twisting about the centerline curve of the long flexible object. Therefore, a generalized continuum formulation is derived to compute both bending and twisting deformations of elastic rods, resulting a simple and general simulation model to facilitate efficient physics computations, whereas conversional simulation methods for elastic rods require explicitly decoupling between bending and twisting deformations. In this paper, we show simulations of a wide range of object behaviors for interactive virtual reality applications.
Tao Ruan Wan, Wen Tang 0004, Dongjin Huang
VRST3
2011 An Interactive 3D Preoperative Planning and Training System for Minimally Invasive Vascular Surgery
abstract
Virtual reality based preoperative planning for Minimally Invasive Vascular Intervention is useful, not only for increasing the success rate of operation, but also used as a training tool for improving doctors' skills. In this paper we present an interactive 3D preoperative planning and training system with haptic device. In this system, we present an intelligent trajectory planning algorithm for searching an optimal path automatically along the centerline in two ways: from the suitable inserting point to the specific target location or to the most of objectives. Also, for the purpose of interactive training, we connect the haptic device to the end of guide wire and propose an algorithm that enables the simulator to model guide wire and catheter insertions realistically through essential operations i.e. pushing, pulling and twisting actions. We demonstrate experiment results to show that the 3D preoperative planning and training system is usable for simulating guide wire insertion procedures with complex blood vessel structures.
Dongjin Huang, Wen Tang 0004, Youdong Ding, Tao Ruan Wan
CAD/Graphics1
2011 Motion Capture of Hand Movements Using Stereo Vision for Minimally Invasive Vascular Interventions
abstract
A virtual reality (VR) based training system for Minimally Invasive Vascular Surgery can be a very useful training tool for improving skills and reducing errors in operation. Computer vision techniques have the potential to be incorporated into a VR based training system for developing low cost, high accuracy and flexible systems in this area. In this paper, we present an interactive 3D training system that uses stereo vision to capture hand movements as the input operations for the system. The standard operations i.e. pushing, pulling and twisting are captured with stereo vision based on the improved Camshift tracking algorithm and parallel alignment model theory to acquire hand gestures information. We present a new approach to calculate virtual pushing/pulling force and turning angle as extra inputs for understanding these essential operations. In addition, an algorithm that enables the simulator to model guide wire and catheter insertions realistically is presented through these basic actions. The experiment results demonstrate that stereo vision based training system is useful and effective for simulating guide wire insertion procedures with low system cost and flexible operations.
Dongjin Huang, Wen Tang 0004, Youdong Ding, Tao Ruan Wan, Xuechun Wu
ICIG1
2011 A new approach to haptic rendering of guidewires for use in minimally invasive surgical simulation
abstract
Abstract Guidewire insertion is an imperative task of minimally invasive medical procedures. During the procedure, surgeons need to steer long flexible thin wires through patient's blood vessels to reach a clinical target. In this paper, we present a novel approach to model haptics of guidewire insertion process for training simulation. The algorithm also allows for the analysis of the insertion process through subtle physical behaviours of guidewires via force feedbacks. The method includes a 6‐DoF dynamic coupling between a rigid body, i.e. the virtual tool and the deformation of the wire simulated as an elastic rod. Instead of using the frictional contact force or the acceleration of the guidewire tip for haptic feedbacks, we compute constrained forces by directly connecting the virtual tool to the end of the guidewire. Therefore, the coupling scheme transmits haptic interactions through constrained dynamics between the virtual tool and the guidewire. Both positional and rotational control modes are implemented and evaluated with respect to the dynamics of the guidewire, user inputs and feedback forces. Experiments highlight the usability of our algorithm for an insertion procedure simulation with complex blood vessel structures. Copyright © 2011 John Wiley & Sons, Ltd.
Dongjin Huang, Wen Tang 0004, Tao Ruan Wan, Nigel W. John, Derek Gould, Youdong Ding
Comput. Animat. Virtual Worlds1