VLDB 2026 Research / reviewers in the wild / expert
Minghao Liu 0009
dblp:283/6232
· DBLP profile ↗
8ranked-venue papers
4as first author
8since 2021 · last 2025
0000-0002-1629-2429ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Theory of computation · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Human and AI Perceptual Differences in Image Classification ErrorsabstractArtificial intelligence (AI) models for computer vision trained with supervised machine learning are assumed to solve classification tasks by imitating human behavior learned from training labels. Most efforts in recent vision research focus on measuring the model task performance using standardized benchmarks such as accuracy. However limited work has sought to understand the perceptual difference between humans and machines. To fill this gap, this study first analyzes the statistical distributions of mistakes from the two sources, and then explores how task difficulty level affects these distributions. We find that even when AI learns an excellent model from the training data, one that outperforms humans in overall accuracy, these AI models have significant and consistent differences from human perception. We demonstrate the importance of studying these differences with a simple human-AI teaming algorithm that outperforms humans alone, AI alone, or AI-AI teaming. Minghao Liu 0009, Jiaheng Wei, Yang Liu 0018, James Davis 0001 |
AAAI | 1 |
| 2025 | Neurosymbolic Tag-Based Annotation for Interpretable Avatar CreationabstractAvatar creation from human images presents challenges for direct neural approaches, which suffer from inconsistent predictions and poor interpretability due to the large parameter space with hundreds of ambiguous options. We propose a neurosymbolic tag-based annotation method that combines neural perceptual learning with symbolic semantic reasoning. Instead of directly predicting avatar parameters, our approach uses a neural network to predict semantic tags (hair length, curliness, direction) as an intermediate symbolic representation, then applies symbolic search algorithms to match optimal avatar assets. This neurosymbolic design produces higher annotator agreements (96.7% vs 31.0% for direct annotation), enables more consistent model predictions, and provides interpretable avatar selection with ranked alternatives. The tag-based system generalizes easily across rendering systems, requiring only new asset annotation while reusing human image tags. Experimental results demonstrate superior convergence, consistency, and visual quality compared to direct prediction methods, showing how neurosymbolic approaches can improve trustworthiness and interpretability in creative AI applications. Minghao Liu 0009, Zeyu Cheng, Shen Sang, Jing Liu 0053, James Davis 0001 |
NeSy | 1 |
| 2025 | GenIR: Generative Visual Feedback for Mental Image RetrievalabstractVision-language models (VLMs) have shown strong performance on text-to-image retrieval benchmarks. However, bridging this success to real-world applications remains a challenge. In practice, human search behavior is rarely a one-shot action. Instead, it is often a multi-round process guided by clues in mind. That is, a mental image ranging from vague recollections to vivid mental representations of the target image. Motivated by this gap, we study the task of Mental Image Retrieval (MIR), which targets the realistic yet underexplored setting where users refine their search for a mentally envisioned image through multi-round interactions with an image search engine. Central to successful interactive retrieval is the capability of machines to provide users with clear, actionable feedback; however, existing methods rely on indirect or abstract verbal feedback, which can be ambiguous, misleading, or ineffective for users to refine the query. To overcome this, we propose GenIR, a generative multi-round retrieval paradigm leveraging diffusion-based image generation to explicitly reify the AI system's understanding at each round. These synthetic visual representations provide clear, interpretable feedback, enabling users to refine their queries intuitively and effectively. We further introduce a fully automated pipeline to generate a high-quality multi-round MIR dataset. Experimental results demonstrate that GenIR significantly outperforms existing interactive methods in the MIR scenario. This work establishes a new task with a dataset and an effective generative retrieval method, providing a foundation for future research in this direction Diji Yang, Minghao Liu 0009, Chung-Hsiang Lo, Yi Zhang 0001, James Davis 0001 |
NeurIPS | 2 |
| 2025 | Controllable Biophysical Human FacesabstractWe present a novel generative model that synthesizes photorealistic, biophysically plausible faces by capturing the intricate relationships between facial geometry and biophysical attributes. Our approach models facial appearance in a biophysically grounded manner, allowing for the editing of both high‐level attributes such as age and gender, as well as low‐level biophysical properties such as melanin level and blood content. This enables continuous modeling of physical skin properties that correlate changes in skin properties with shape changes. We showcase the capabilities of our framework beyond its role as a generative model through two practical applications: editing the texture maps of 3D faces that have already been captured, and serving as a strong prior for face reconstruction when combined with differentiable rendering. Our model allows for the creation of physically‐based relightable, editable faces with consistent topology and uv layout that can be integrated into traditional computer graphics pipelines. Minghao Liu 0009, Stephane Grabli, Sébastien Speierer, Nikolaos Sarafianos, Lukas Bode, Matt Jen-Yuan Chiang, Christophe Hery, James Davis 0001, Carlos Aliaga |
Comput. Graph. Forum | 1 |
| 2022 | How much does input data type impact final face model accuracy?abstractFace models are widely used in image processing and other domains. The input data to create a 3D face model ranges from accurate laser scans to simple 2D RGB photographs. These input data types are typically deficient either due to missing regions, or because they are underconstrained. As a result, reconstruction methods include embedded priors encoding the valid domain of faces. System designers must choose a source of input data and then choose a reconstruction method to obtain a usable 3D face. If a particular application domain requires accuracy X, which kinds of input data are suitable? Does the input data need to be 3D, or will 2D data suffice? This paper takes a step toward answering these questions using synthetic data. A ground truth dataset is used to analyze accuracy obtainable from 2D landmarks, 3D landmarks, low quality 3D, high quality 3D, texture color, normals, dense 2D image data, and when regions of the face are missing. Since the data is synthetic it can be analyzed both with and without measurement error. This idealized synthetic analysis is then compared to real results from several methods for constructing 3D faces from 2D photographs. The experimental results suggest that accuracy is severely limited when only 2D raw input data exists. Jiahao Luo, Fahim Hasan Khan, Issei Mori, Akila de Silva, Eric Ruezga, Minghao Liu 0009, Alex T. Pang, James Davis 0001 |
CVPR | 6 |
| 2022 | DuelGAN: A Duel Between Two Discriminators Stabilizes the GAN Training
Jiaheng Wei, Minghao Liu 0009, Jiahao Luo, Andrew Zhu, James Davis 0001, Yang Liu 0018 |
ECCV (23) | 2 |
| 2022 | Low-light Image Enhancement Using Chain-consistent Adversarial NetworksabstractThe capability to generate clear and bright images in low light situations is crucial for photographers, engineers, and researchers. When it is not possible to modify the imaging conditions, an algorithm to enhance images is needed. Traditional methods require manually adjusting parameters to tune the image. Supervised learning methods need to collect a large amount of paired data for training. In this paper, we demonstrate an semi-supervised method for low light image enhancement, using a chain of cycle consistent generators. We show the effectiveness of our method by comparing it to existing image enhancement methods, both using standard image quality metrics and by using human perceptual judgements. We include an ablation study for features in our model. Our proposed method is computationally efficient and does not require paired training data. Minghao Liu 0009, Jiahao Luo, Xiaohan Zhang 0003, Yang Liu 0018, James Davis 0001 |
ICPR | 1 |
| 2022 | AgileAvatar: Stylized 3D Avatar Creation via Cascaded Domain BridgingabstractStylized 3D avatars have become increasingly prominent in our modern life. Creating these avatars manually usually involves laborious selection and adjustment of continuous and discrete parameters and is time-consuming for average users. Self-supervised approaches to automatically create 3D avatars from user selfies promise high quality with little annotation cost but fall short in application to stylized avatars due to a large style domain gap. We propose a novel self-supervised learning framework to create high-quality stylized 3D avatars with a mix of continuous and discrete parameters. Our cascaded domain bridging framework first leverages a modified portrait stylization approach to translate input selfies into stylized avatar renderings as the targets for desired 3D avatars. Next, we find the best parameters of the avatars to match the stylized avatar renderings through a differentiable imitator we train to mimic the avatar graphics engine. To ensure we can effectively optimize the discrete parameters, we adopt a cascaded relaxation-and-search pipeline. We use a human preference study to evaluate how well our method preserves user identity compared to previous work as well as manual creation. Our results achieve much higher preference scores than previous work and close to those of manual creation. We also provide an ablation study to justify the design choices in our pipeline. Shen Sang, Tiancheng Zhi, Guoxian Song, Minghao Liu 0009, Chun-Pong Lai, Jing Liu 0053, James Davis 0001, Linjie Luo |
SIGGRAPH Asia | 4 |