VLDB 2026 Research / reviewers in the wild / expert
Shang-Fu Chen
dblp:203/9102
· DBLP profile ↗
13ranked-venue papers
6as first author
11since 2021 · last 2026
0000-0002-6319-9294ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DirectDrag: High-Fidelity, Mask-Free, Prompt-Free Drag-based Image Editing via Readout-Guided Feature AlignmentabstractDrag-based image editing using generative models provides intuitive control over image structures. However, existing methods rely heavily on manually provided masks and textual prompts to preserve semantic fidelity and motion precision. Removing these constraints creates a fundamental trade-off: visual artifacts without masks and poor spatial control without prompts. To address these limitations, we propose DirectDrag, a novel mask-and prompt-free editing framework. DirectDrag enables precise and efficient manipulation with minimal user input while maintaining high image fidelity and accurate point alignment. DirectDrag introduces two key innovations. First, we design an Auto Soft Mask Generation module that intelligently infers editable regions from point displacement, automatically localizing deformation along movement paths while preserving contextual integrity through the generative model’s inherent capacity. Second, we develop a Readout-Guided Feature Alignment mechanism that leverages intermediate diffusion activations to maintain structural consistency during point-based edits, substantially improving visual fidelity. Despite operating without manual mask or prompt, DirectDrag achieves superior image quality compared to existing methods while maintaining competitive drag accuracy. Extensive experiments on DragBench and real-world scenarios demonstrate the effectiveness and practicality of DirectDrag for high-quality, interactive image manipulation. Code is available at: https://github.com/frakw/DirectDrag. Sheng-Hao Liao, Shang-Fu Chen, Tai-Ming Huang, Wen-Huang Cheng, Kai-Lung Hua |
WACV | 2 |
| 2026 | A novel simulation tool for low-coverage whole-genome sequencing using multivariate Gaussian mixture modelsabstractLow-coverage whole-genome sequencing (lcWGS) has emerged as a cost-effective and robust approach for population genomic studies. Despite its advantages, publicly available resources for large-scale lcWGS datasets remain limited. To our knowledge, there is yet a bioinformatics tool capable of directly simulating lcWGS datasets from variant call format (VCF) files. To address this gap, we developed a tool called lcSimVCF, which leverages multivariate Gaussian mixture models (MGMMs) to simulate lcWGS genotype likelihood distributions directly from high-coverage whole-genome sequencing (hcWGS) VCF files. In this study, we introduced a tool called lcSimVCF that aim to use 30× hcWGS VCF files to simulate 1× genotype likelihood distributions as a demonstration. We trained an MGMM framework for three genotypes: homozygous reference, heterozygous, and homozygous non-reference, each simulating its respective lcWGS genotype likelihood distribution. Our results demonstrate the robustness of the MGMM approach in capturing genotype likelihood distributions compared to single Gaussian mixture models (SGM). Simulated lcWGS data exhibited representative patterns in population stratification analyses and showed potential for applications in polygenic risk score (PRS) modeling. Analysis of 81,271,745 imputed single nucleotide polymorphisms (SNPs) revealed strong correlations among the PRS derived from different phenotypes: PRS CAD , PRS T2D , and PRS AF . The correlations demonstrated high coefficients of determination ( \({R}^{2}\) ) of 0.94, 0.87, and 0.85, respectively. Furthermore, we assessed simulation time across various configurations (1-16 CPUs, 1-800 individuals, 100-10,000 variants), finding its capability in simulating 10,000 variants across 800 individuals in under 30 seconds with 16 CPUs. The developed simulation tool has demonstrated its capability in generating large-scale lcWGS datasets. This tool holds significant potential to facilitate the development and evaluation of bioinformatics tools and analytical pipelines that rely on access to extensive lcWGS data resources. Kai-Yu Chen, Shang-Fu Chen, Raquel Dias, Ali Torkamani |
BMC Bioinform. | 2 |
| 2026 | Restoring Noisy Demonstration for Imitation Learning With Diffusion ModelsabstractImitation learning (IL) aims to learn a policy from expert demonstrations and has been applied to various applications. By learning from the expert policy, IL methods do not require environmental interactions or reward signals. However, most existing IL algorithms assume perfect expert demonstrations, but expert demonstrations often contain imperfections caused by errors from human experts or sensor/control system inaccuracies. To address the above problems, this work proposes a filter-and-restore framework to best leverage expert demonstrations with inherent noise. Our proposed method first filters clean samples from the demonstrations and then learns conditional diffusion models to recover the noisy ones. We evaluate our proposed framework and existing methods in various domains, including robot arm manipulation, dexterous manipulation, and locomotion. The experiment results show that our proposed framework consistently outperforms existing methods across all the tasks. Ablation studies further validate the effectiveness of each component and demonstrate the framework's robustness to different noise types and levels. These results confirm the practical applicability of our framework to noisy offline demonstration data. Shang-Fu Chen, Co Yong, Shao-Hua Sun |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2025 | HERO: Human-Feedback Efficient Reinforcement Learning for Online Diffusion Model FinetuningabstractControllable generation through Stable Diffusion (SD) fine-tuning aims to improve fidelity, safety, and alignment with human guidance. Existing reinforcement learning from human feedback methods usually rely on predefined heuristic reward functions or pretrained reward models built on large-scale datasets, limiting their applicability to scenarios where collecting such data is costly or difficult. To effectively and efficiently utilize human feedback, we develop a framework, HERO, which leverages online human feedback collected on the fly during model learning. Specifically, HERO features two key mechanisms: (1) Feedback-Aligned Representation Learning, an online training method that captures human feedback and provides informative learning signals for fine-tuning, and (2) Feedback-Guided Image Generation, which involves generating images from SD's refined initialization samples, enabling faster convergence towards the evaluator's intent. We demonstrate that HERO is 4x more efficient in online feedback for body part anomaly correction compared to the best existing method. Additionally, experiments show that HERO can effectively handle tasks like reasoning, counting, personalization, and reducing NSFW content with only 0.5K online feedback. The code and project page are available at [https://hero-dm.github.io/](https://hero-dm.github.io/). Ayano Hiranaka, Shang-Fu Chen, Chieh-Hsin Lai, Naoki Murata, Takashi Shibuya 0001, Wei-Hsiang Liao 0001, Shao-Hua Sun, Yuki Mitsufuji |
ICLR | 2 |
| 2024 | Representation and Boundary Enhancement for Action Segmentation Using TransformerabstractIn the task of action segmentation, the goal is to partition a lengthy, untrimmed video into a series of action segments. Recently, Transformer-based methods have outperformed the previous temporal convolutional networks (TCNs) in terms of overall performance. However, both TCNs and Transformers encounter the challenge of over-segmentation. Prior approaches often relied on post-processing techniques to address this issue, but these methods are not universally applicable to every model and may sometimes result in performance degradation. Therefore, in this paper, we propose a set of loss functions to enhance representation learning and employ a multi-task learning approach to strengthen the model’s ability to identify action boundaries. Through extensive experiments, we validate that our method demonstrates significant improvements, particularly in addressing the challenge of over-segmentation. Shang-Fu Chen, Cheng-Xun Wen, Wen-Huang Cheng, Kai-Lung Hua |
ICASSP | 1 |
| 2024 | Diffusion Model-Augmented Behavioral CloningabstractImitation learning addresses the challenge of learning by observing an expert’s demonstrations without access to reward signals from environments. Most existing imitation learning methods that do not require interacting with environments either model the expert distribution as the conditional probability p(a|s) (e.g., behavioral cloning, BC) or the joint probability p(s, a). Despite the simplicity of modeling the conditional probability with BC, it usually struggles with generalization. While modeling the joint probability can improve generalization performance, the inference procedure is often time-consuming, and the model can suffer from manifold overfitting. This work proposes an imitation learning framework that benefits from modeling both the conditional and joint probability of the expert distribution. Our proposed Diffusion Model-Augmented Behavioral Cloning (DBC) employs a diffusion model trained to model expert behaviors and learns a policy to optimize both the BC loss (conditional) and our proposed diffusion model loss (joint). DBC outperforms baselines in various continuous control tasks in navigation, robot arm manipulation, dexterous manipulation, and locomotion. We design additional experiments to verify the limitations of modeling either the conditional probability or the joint probability of the expert distribution, as well as compare different generative models. Ablation studies justify the effectiveness of our design choices. Shang-Fu Chen, Hsiang-Chun Wang, Ming-Hao Hsu, Chun-Mao Lai, Shao-Hua Sun |
ICML | 1 |
| 2023 | Controllable Model Compression for Roadside Camera Depth EstimationabstractIn the Cooperative Intelligent Transportation System (C-ITS) paradigm, vehicles could communicate with roadside units to augment their traffic knowledge. Smart roadside units could provide second-order information (e.g., vehicle count) from raw first-order data (e.g., visual feed, point clouds), and this “smart” feature is usually provided using deep neural network models. However, implementing these useful models implies a cost for computational complexity that could hinder the future deployment of smart roadside units needed for sustainability in transportation systems. In this paper, we propose to use model compression on deep image processing models to promote its feasibility for usage in smart sensors. We formulated a controllable convolutional model compression (CCMC) algorithm that can perform filter-wise evolutionary pruning on image processing networks, along with a predefined compression ratio. CCMC is applicable for image processing networks, which have multiple possible traffic data sources (e.g., road camera surveillance). Furthermore, CCMC has a definable target compression ratio that is useful for controlling the trade-off between resource consumption and output performance. We tested our proposed method on depth estimation, which is useful for scene understanding and mapping the locations of objects in the 3D space. Our experiments show that the pruned model has minimal performance discrepancy from the original one, supporting the sustainability features needed for intelligent transportation systems. Jose Jaena Mari Ople, Shang-Fu Chen, Yung-Yao Chen, Kai-Lung Hua, Mohammad Hijji, Po Yang 0001, Khan Muhammad 0001 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2022 | Domain-Generalized Textured Surface Anomaly DetectionabstractAnomaly detection aims to identify abnormal data that deviates from the normal ones, while typically requiring a sufficient amount of normal data to train the model for performing this task. Despite the success of recent anomaly detection methods, performing anomaly detection in an unseen domain remain a challenging task. In this paper, we address the task of domain-generalized textured surface anomaly detection. By observing normal and abnormal surface data across multiple source domains, our model is expected to be generalized to an unseen textured surface of interest, in which only a small number of normal data can be observed during testing. Although with only image-level labels observed in the training data, our patch-based meta-learning model exhibits promising generalization ability: not only can it generalize to unseen image domains, but it can also localize abnormal regions in the query image. Our experiments verify that our model performs favorably against state-of-the-art anomaly detection and domain generalization approaches in various settings. Shang-Fu Chen, Yu-Min Liu, Chia-Ching Lin, Trista Pei-Chun Chen, Yu-Chiang Frank Wang |
ICME | 1 |
| 2022 | Learning Facial Liveness Representation for Domain Generalized Face Anti-SpoofingabstractFace anti-spoofing (FAS) aims at distinguishing face spoof attacks from the authentic ones, which is typically approached by learning proper models for performing the associated classification task. In practice, one would expect such models to be generalized to FAS in different image domains. Moreover, it is not practical to assume that the type of spoof attacks would be known in advance. In this paper, we propose a deep learning model for addressing the aforementioned domain-generalized face anti-spoofing task. In particular, our proposed network is able to disentangle facial liveness representation from the irrelevant ones (i.e., facial content and image domain features). The resulting liveness representation exhibits sufficient domain invariant properties, and thus it can be applied for performing domain-generalized FAS. In our experiments, we conduct experiments on five benchmark datasets with various settings, and we verify that our model performs favorably against state-of-the-art approaches in identifying novel types of spoof attacks in unseen image domains. Zih-Ching Chen, Lin-Hsi Tsao, Chin-Lun Fu, Shang-Fu Chen, Yu-Chiang Frank Wang |
ICME | 4 |
| 2021 | Representation Decomposition For Image Manipulation And BeyondabstractRepresentation disentanglement aims at learning interpretable features, so that the output can be recovered or manipulated accordingly. While existing works like infoGAN [1] and ACGAN [2] exist, they choose to derive disjoint attribute code for feature disentanglement, which is not applicable for existing/trained generative models. In this paper, we propose a decomposition-GAN (dec-GAN), which is able to achieve the decomposition of an existing latent representation into content and attribute features. Guided by the classifier pre-trained on the attributes of interest, our dec-GAN decomposes the attributes of interest from the latent representation, while data recovery and feature consistency objectives enforce the learning of our proposed method. Our experiments on multiple image datasets confirm the effectiveness and robustness of our dec-GAN over recent representation disentanglement models. Shang-Fu Chen, Jia-Wei Yan, Ya-Fan Su, Yu-Chiang Frank Wang |
ICIP | 1 |
| 2021 | Robust Image Outpainting With Learnable Image MarginsabstractGiven a partial image input, image outpainting is to produce the desirable output by recovering or extending the surrounding image regions. While existing image outpainting methods achieve impressive results based on the recent advances of deep learning, they either lack the ability to extend image regions in arbitrary directions or require the filling image margins to be given in advance. To address this challenging task, we propose a unique deep learning framework for robust image outpainting, which consists of a margin prediction network and a teacher-student-based network for producing outpainted images. Our proposed model does not require image filling margins to be known beforehand, while both image appearance and perceptual feature consistencies can be jointly enforced. Our experiments quantitatively and qualitatively verify the effectiveness of our method, which is shown to perform favorably against baseline and state-of-the-art image outpainting works. Cheng-Yo Tan, Chiao-An Yang, Shang-Fu Chen, Meng-Lin Wu, Yu-Chiang Frank Wang |
ICIP | 3 |
| 2019 | Learning Hierarchical Self-Attention for Video SummarizationabstractVideo summarization still remains a challenging task. Due to sufficient video data on the Internet, such task draws significant attention in the vision community and benefits a wide range of applications, e.g., video retrieval, search, etc. To effectively perform video summarization by deriving the keyframes which represent the given input video, we propose a novel framework named Hierarchical Multi-Attention Network (H-MAN) which comprises the shot-level reconstruction model and multi-head attention model. While our designed attention model is two-stage hierarchical structure for producing various attention maps, we are among the first to utilize the multi-attention mechanism in the video summarization task, which brings improved performance. The quantitative and qualitative results demonstrate the effectiveness of our model, which performs favorably against state-of-the-art approaches. Yen-Ting Liu, Yu-Jhe Li, Fu-En Yang, Shang-Fu Chen, Yu-Chiang Frank Wang |
ICIP | 4 |
| 2018 | Order-Free RNN With Visual Attention for Multi-Label ClassificationabstractWe propose a recurrent neural network (RNN) based model for image multi-label classification. Our model uniquely integrates and learning of visual attention and Long Short Term Memory (LSTM) layers, which jointly learns the labels of interest and their co-occurrences, while the associated image regions are visually attended. Different from existing approaches utilize either model in their network architectures, training of our model does not require pre-defined label orders. Moreover, a robust inference process is introduced so that prediction errors would not propagate and thus affect the performance. Our experiments on NUS-WISE and MS-COCO datasets confirm the design of our network and its effectiveness in solving multi-label classification problems. Shang-Fu Chen, Chih-Kuan Yeh, Yu-Chiang Frank Wang |
AAAI | 1 |