VLDB 2026 Research / reviewers in the wild / expert
Tai-Ming Huang
dblp:325/4511
· DBLP profile ↗
7ranked-venue papers
1as first author
7since 2021 · last 2026
0000-0003-0525-0466ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Dragonite: Single-Step Drag-based Image Editing with Geometric-Semantic GuidanceabstractRecent interactive image editing methods have made notable progress, yet achieving both precise control and real-time performance remains a challenge. Drag-based methods offer detailed geometric manipulations but suffer from low image fidelity and slow runtime performance, while text-based approaches enhance realism but limit precise and pixel-level control. To overcome these limitations, we introduce Dragonite, an intuitive and efficient framework that seamlessly unifies geometric and semantic manipulation for image editing. Dragonite leverages a Dual Guidance Module that fuses geometric deformation vectors with semantic guidance cues into a joint representation space, ensuring precise manipulation of both content and semantics. By combining a single-step latent optimization mechanism with a enhanced interpolation method, Dragonite achieves efficient interactive image editing while maintaining high precision through integrated geometric and semantic guidance. Extensive evaluations on the DragBench benchmark demonstrate that Dragonite effectively resolves the trade-off between speed and accuracy, enabling real-time, high-fidelity image editing. Meng-Ting Jhong, Tai-Ming Huang, Shung-Fu Chen, Wen-Huang Cheng, Kai-Lung Hua |
WACV | 2 |
| 2026 | DirectDrag: High-Fidelity, Mask-Free, Prompt-Free Drag-based Image Editing via Readout-Guided Feature AlignmentabstractDrag-based image editing using generative models provides intuitive control over image structures. However, existing methods rely heavily on manually provided masks and textual prompts to preserve semantic fidelity and motion precision. Removing these constraints creates a fundamental trade-off: visual artifacts without masks and poor spatial control without prompts. To address these limitations, we propose DirectDrag, a novel mask-and prompt-free editing framework. DirectDrag enables precise and efficient manipulation with minimal user input while maintaining high image fidelity and accurate point alignment. DirectDrag introduces two key innovations. First, we design an Auto Soft Mask Generation module that intelligently infers editable regions from point displacement, automatically localizing deformation along movement paths while preserving contextual integrity through the generative model’s inherent capacity. Second, we develop a Readout-Guided Feature Alignment mechanism that leverages intermediate diffusion activations to maintain structural consistency during point-based edits, substantially improving visual fidelity. Despite operating without manual mask or prompt, DirectDrag achieves superior image quality compared to existing methods while maintaining competitive drag accuracy. Extensive experiments on DragBench and real-world scenarios demonstrate the effectiveness and practicality of DirectDrag for high-quality, interactive image manipulation. Code is available at: https://github.com/frakw/DirectDrag. Sheng-Hao Liao, Shang-Fu Chen, Tai-Ming Huang, Wen-Huang Cheng, Kai-Lung Hua |
WACV | 3 |
| 2025 | Towards More General Video-based Deepfake Detection through Facial Component Guided Adaptation for Foundation ModelabstractThe current deep generative models have enabled the creation of synthetic facial images with remarkable photorealism, raising significant societal concerns over their potential misuse. Despite rapid advancements in the field of deepfake detection, developing an efficient and effective approach for the generalized deepfake detection of unseen forgery samples remains challenging. To address this challenge, we leverage the rich semantic priors of foundation models and propose a novel side-network-based decoder that extracts spatial and temporal cues using the CLIP image encoder for generalized video-based Deepfake detection. Additionally, we introduce Facial Component Guidance (FCG) to enhance spatial learning generalizability by encouraging the model to focus on key facial regions. By leveraging the generic features of a vision-language foundation model, our approach demonstrates promising generalizability on challenging Deepfake datasets while also exhibiting superiority in training data efficiency, parameter efficiency, and model robustness. The source code is available at: https://github.com/aiiu-lab/DFD-FCG Yue-Hua Han, Tai-Ming Huang, Kai-Lung Hua, Jun-Cheng Chen |
CVPR | 2 |
| 2025 | EXDF: Explainable Deepfake Detection with Vision-Language ModelabstractAlthough many deepfake detection methods have been proposed to fight against severe misuse of generative AI, none provide detailed human-interpretable explanations beyond simple real/fake responses. This limitation makes it challenging for humans to assess the accuracy of detection results, especially when the models encounter unseen deepfakes. To address this issue, we propose a novel deepfake detector based on a large Vision-Language Model (VLM), capable of explaining manipulated facial regions. We frame the deepfake detection task as Visual Question Answering (VQA) and perform visual instruction tuning to train the model on our collected Explainable Deepfake Face (ExDF) dataset. The dataset consists of fake images from diverse generative adversarial networks (GANs) and diffusion models (DMs), with explanations produced by GPT-4o guided by the corresponding ground-truth masks of the manipulated regions. Moreover, a facial mask encoder is introduced to guide the model to focus on key facial features, thereby improving the detection and explanation performances. Extensive experiments demonstrate that training the proposed model on the full ExDF dataset not only enhances detection accuracy compared to baseline methods but also provides detailed, human-interpretable explanations. To our knowledge, ExDF is the first explainable deepfake face dataset covering both GANs and DMs with comprehensive descriptions of altered facial regions. Our code and dataset are available at https://github.com/aiiu-lab/ExDF. Shu-Tzu Lo, Tai-Ming Huang, Yue-Hua Han, Kai-Lung Hua, Jun-Cheng Chen |
ICIP | 2 |
| 2024 | Generalized Image-Based Deepfake Detection Through Foundation Model Adaptation
Tai-Ming Huang, Yue-Hua Han, Ernie Chu, Shu-Tzu Lo, Kai-Lung Hua, Jun-Cheng Chen |
ICPR (21) | 1 |
| 2023 | Adjustable Model Compression Using Multiple Genetic AlgorithmabstractGenerative Adversarial Networks (GAN) is a popular machine learning method that possesses powerful image generation ability, which is useful for different multimedia applications (e.g., photographic filters, image editing). However, typical GAN models have a large memory footprint that limits their practical applications for resource-constrained devices (e.g., smartphones). To deploy GAN models on devices with various hardware constraints, we propose our method, AdjustableGAN, which can compress a pretrained GAN model to different compression ratios. Our method compresses GAN by performing filter-wise pruning that follows these objectives: (1) deactivate convolutional filters for minimal performance decrease, (2) reactivate convolutional filters for maximal performance increase. We implement multiple Genetic Algorithms (GA) to perform each of these objectives— Downsize GA for best filter deactivations, while Upsize GA searches for best filter reactivations. By selective utilization of Upsize/Downsize GA, we could explicitly control the compression ratio of the model. For finalization, we fine-tune the compressed output model using the training dataset of the original input model. Our experimental results show that our method can reliably compress generative networks with minimal accuracy drop compared to other state-of-the-art compression algorithms. Jose Jaena Mari Ople, Tai-Ming Huang, Ming-Chih Chiu, Yi-Ling Chen 0002, Kai-Lung Hua |
IEEE Trans. Multim. | 2 |
| 2022 | VDNet: video deinterlacing network based on coarse adaptive module and deformable recurrent residual network
Yin-Chen Yeh, Jilyan Bianca Dy, Tai-Ming Huang, Yung-Yao Chen, Kai-Lung Hua |
Neural Comput. Appl. | 3 |