VLDB 2026 Research / reviewers in the wild / expert
Duc Minh Vo
dblp:138/9093
· DBLP profile ↗
20ranked-venue papers
9as first author
14since 2021 · last 2025
0000-0003-4839-032XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 17 · 7 first-author · 12 since 2021Artificial intelligence and machine learning · 13 · 7 first-author · 8 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | LAVA Grand Challenge 2025: Benchmarking Japanese-English Document Understanding with Large Vision-Language ModelsabstractThe advent of Large Vision-Language Models (LVLMs) has demonstrated significant capabilities in multimodal understanding. However, their application to complex, multi-page documents, particularly in non-English languages like Japanese, remains a significant challenge due to the scarcity of suitable benchmarks. To address this gap, we organized the ''Large Vision---Language Model Learning and Applications (LAVA) Grand Challenge'' at ACM MultiMedia 2025. We present an overview of the competition. We designed a novel, challenging task: a 10-way multiple-choice Visual Question Answering (VQA) task on multi-page Japanese PDF documents. The task demands that models integrate information across multiple pages, text, and figures. We detail the dataset construction, including an annotation and filtering process designed to ensure questions are visually grounded and non-trivial. We also present the competition results, including an analysis of the leaderboard, and discuss the baseline performance of representative models. The LAVA Grand Challenge highlighted both the current capabilities and limitations of LVLMs in practical document understanding scenarios, thereby stimulating future research and providing a robust benchmark in this important domain. Daichi Sato, Duc Minh Vo, Khan Md. Anwarus Salam, Hidenori Shoji, Yuma Matsuoka, Takara Taniguchi, Kaito Baba, Hideki Nakayama |
ACM Multimedia | 2 |
| 2024 | Evcap: Retrieval-Augmented Image Captioning with External Visual-Name Memory for Open-World ComprehensionabstractLarge language models (LLMs)-based image captioning has the capability of describing objects not explicitly observed in training data; yet novel objects occur frequently, necessitating the requirement of sustaining up-to-date object knowledge for open-world comprehension. Instead of relying on large amounts of data and/or scaling up network parameters, we introduce a highly effective retrieval-augmented image captioning method that prompts LLMs with object names retrieved from External Visual-name memory (EVCAP). We build ever-changing object knowledge memory using objects' visuals and names, enabling us to (i) update the memory at a minimal cost and (ii) effort-lessly augment LLMs with retrieved object names by uti-lizing a lightweight and fast-to-train model. Our model, which was trained only on the COCO dataset, can adapt to out-of-domain without requiring additional fine-tuning or retraining. Our experiments conducted on benchmarks and synthetic commonsense-violating data show that EV-CAP, with only 3.97M trainable parameters, exhibits superior performance compared to other methods based on frozen pretrained LLMs. Its performance is also competitive to specialist SOTAs that require extensive training. Jiaxuan Li 0004, Duc Minh Vo, Akihiro Sugimoto, Hideki Nakayama |
CVPR | 2 |
| 2024 | A Compact Dynamic 3D Gaussian Representation for Real-Time Dynamic View Synthesis
Kai Katsumata, Duc Minh Vo, Hideki Nakayama |
ECCV (86) | 2 |
| 2024 | Soft Curriculum for Learning Conditional GANs with Noisy-Labeled and Uncurated Unlabeled DataabstractLabel-noise or curated unlabeled data are used to compensate for the assumption of clean labeled data in training the conditional generative adversarial network; however, satisfying such an extended assumption is occasionally laborious or impractical. As a step towards generative modeling accessible to everyone, we introduce a novel conditional image generation framework that accepts noisy-labeled and uncurated unlabeled data during training: (i) closed-set and open-set label noise in labeled data and (ii) closed-set and open-set unlabeled data. To combat it, we propose soft curriculum learning, which assigns instance-wise weights for adversarial training while assigning new labels for unlabeled data and correcting wrong labels for labeled data. Unlike popular curriculum learning, which uses a threshold to pick the training samples, our soft curriculum controls the effect of each training instance by using the weights predicted by the auxiliary classifier, resulting in the preservation of useful samples while ignoring harmful ones. Our experiments show that our approach outperforms existing semi-supervised and label-noise robust methods in terms of both quantitative and qualitative performance. In particular, the proposed approach matches the performance of (semi-)supervised GANs even with less than half the labeled data.1 Kai Katsumata, Duc Minh Vo, Tatsuya Harada, Hideki Nakayama |
WACV | 2 |
| 2024 | Revisiting Latent Space of GAN Inversion for Robust Real Image EditingabstractWe present a generative adversarial network (GAN) inversion with high reconstruction and editing quality. GAN inversion algorithms with expressive latent spaces produce near-perfect inversion but are not robust to editing operations in a latent space, leading to undesirable edited images, a phenomenon known as the trade-off between reconstruction and editing quality. To cope with the trade-off, we revisit the hyperspherical prior of StyleGANs $\mathcal{Z}$ and propose to combine an extended space of $\mathcal{Z}$ with highly capable inversion algorithms. Our approach maintains the reconstruction quality of seminal GAN inversion methods while improving their editing quality owing to the constrained nature of $\mathcal{Z}$. Through comprehensive experiments with several GAN inversion algorithms, we demonstrate that our approach enhances the image editing quality in 2D/3D GANs.1 Kai Katsumata, Duc Minh Vo, Bei Liu 0001, Hideki Nakayama |
WACV | 2 |
| 2024 | Label Augmentation as Inter-class Data Augmentation for Conditional Image Synthesis with Imbalanced DataabstractConditional image synthesis performs admirably when trained on well-constructed and balanced datasets. However, in practice, training datasets frequently contain minorities (i.e., a class with a few samples), known as imbalanced data, which causes difficulties in learning generative models. To address conditional image synthesis with imbalanced data, we analyze a diversity issue of label-preserving data augmentation and an affinity issue of non-label-preserving data augmentation. From this observation, we present label augmentation, which works as inter-class data augmentation that effectively augments data by predicting a new label for a given image using the prediction of a pretrained image classification model (i.e., probabilities for each class). We incorporate our label augmentation into the discriminator of a seminal conditional generative adversarial network (GAN) model, proposing Softlabel-GAN. Using class probabilities extracts class-invariant and shared features between similar classes, achieving data augmentation with high affinity and diversity. Our experiments on imbalanced datasets show that Softlabel-GAN produces images with high quality and diversity while being hardly affected by the number of samples in each class. Code: https://github.com/raven38/softlabel-gan. Kai Katsumata, Duc Minh Vo, Hideki Nakayama |
WACV | 2 |
| 2023 | A-CAP: Anticipation Captioning with Commonsense KnowledgeabstractHumans possess the capacity to reason about the future based on a sparse collection of visual cues acquired over time. In order to emulate this ability, we introduce a novel task called Anticipation Captioning, which generates a caption for an unseen oracle image using a sparsely temporally-ordered set of images. To tackle this new task, we propose a model called A-CAP, which incorporates commonsense knowledge into a pre-trained vision-language model, allowing it to anticipate the caption. Through both qualitative and quantitative evaluations on a customized visual storytelling dataset, A-CAP out-performs other image captioning methods and establishes a strong baseline for anticipation captioning. We also address the challenges inherent in this task. Duc Minh Vo, Quoc-An Luong, Akihiro Sugimoto, Hideki Nakayama |
CVPR | 1 |
| 2023 | Partition-and-Debias: Agnostic Biases Mitigation via A Mixture of Biases-Specific ExpertsabstractBias mitigation in image classification has been widely researched, and existing methods have yielded notable results. However, most of these methods implicitly assume that a given image contains only one type of known or unknown bias, failing to consider the complexities of real-world biases. We introduce a more challenging scenario, agnostic biases mitigation, aiming at bias removal regardless of whether the type of bias or the number of types is unknown in the datasets. To address this difficult task, we present the Partition-and-Debias (PnD) method that uses a mixture of biases-specific experts to implicitly divide the bias space into multiple subspaces and a gating module to find a consensus among experts to achieve debiased classification. Experiments on both public and constructed benchmarks demonstrated the efficacy of the PnD. Code is available at: https://github.com/Jiaxuan-Li/PnD. Jiaxuan Li 0004, Duc Minh Vo, Hideki Nakayama |
ICCV | 2 |
| 2023 | Indirect Adversarial Losses via an Intermediate Distribution for Training GANsabstractIn this study, we consider the weak convergence characteristics of the Integral Probability Metrics (IPM) methods in training Generative Adversarial Networks (GANs). We first concentrate on a successful IPM-based GAN method that employs a repulsive version of the Maximum Mean Discrepancy (MMD) as the discriminator loss (called repulsive MMD-GAN). We reinterpret its repulsive metrics as an indirect discriminator loss function toward an intermediate distribution. This allows us to propose a novel generator loss via such an intermediate distribution based on our reinterpretation. Our indirect adversarial losses use a simple known distribution (i.e., the Normal or Uniform distribution in our experiments) to simulate indirect adversarial learning between three parts – real, fake, and intermediate distributions. Furthermore, we found the Kernelized Stein Discrepancy (KSD) from the IPM family as the adversarial loss function to avoid randomness from intermediate distribution samples because the target side (intermediate one) is sample-free in KSD. Experiments on several real-world datasets show that our methods can successfully train GANs with the intermediate-distribution-based KSD and MMD and can outperform previous loss metrics. Duc Minh Vo, Hideki Nakayama |
WACV | 2 |
| 2022 | OSSGAN: Open-Set Semi-Supervised Image GenerationabstractWe introduce a challenging training scheme of conditional GANs, called open-set semi-supervised image generation, where the training dataset consists of two parts: (i) labeled data and (ii) unlabeled data with samples belonging to one of the labeled data classes, namely, a closed-set, and samples not belonging to any of the labeled data classes, namely, an open-set. Unlike the existing semi-supervised image generation task, where unlabeled data only contain closed-set samples, our task is more general and lowers the data collection cost in practice by allowing open-set samples to appear. Thanks to entropy regularization, the classifier that is trained on labeled data is able to quantify sample-wise importance to the training of cGAN as confidence, allowing us to use all samples in un-labeled data. We design OSSGAN, which provides decision clues to the discriminator on the basis of whether an unlabeled image belongs to one or none of the classes of interest, smoothly integrating labeled and unlabeled data during training. The results of experiments on Tiny ImageNet and ImageNet show notable improvements over supervised Big-GAN and semi-supervised methods. Our code is available at https://github.com/raven38/OSSGAN. Kai Katsumata, Duc Minh Vo, Hideki Nakayama |
CVPR | 2 |
| 2022 | NOC-REK: Novel Object Captioning with Retrieved Vocabulary from External KnowledgeabstractNovel object captioning aims at describing objects absent from training data, with the key ingredient being the provision of object vocabulary to the model. Although existing methods heavily rely on an object detection model, we view the detection step as vocabulary retrieval from an external knowledge in the form of embeddings for any object's definition from Wiktionary, where we use in the retrieval image region features learned from a transformers model. We propose an end-to-end Novel Object Captioning with Retrieved vocabulary from External Knowledge method (NOC-REK), which simultaneously learns vocabulary retrieval and caption generation, successfully describing novel objects outside of the training dataset. Furthermore, our model eliminates the requirement for model retraining by simply updating the external knowledge whenever a novel object appears. Our comprehensive experiments on held-out COCO and Nocaps datasets show that our NOCREK is considerably effective against SOTAs. Duc Minh Vo, Hong Chen 0017, Akihiro Sugimoto, Hideki Nakayama |
CVPR | 1 |
| 2022 | StoryER: Automatic Story Evaluation via Ranking, Rating and ReasoningabstractExisting automatic story evaluation methods place a premium on story lexical level coherence, deviating from human preference.We go beyond this limitation by considering a novel Story Evaluation method that mimics human preference when judging a story, namely StoryER, which consists of three sub-tasks: Ranking, Rating and Reasoning.Given either a machine-generated or a human-written story, StoryER requires the machine to output 1) a preference score that corresponds to human preference, 2) specific ratings and their corresponding confidences and 3) comments for various aspects (e.g., opening, character-shaping).To support these tasks, we introduce a wellannotated dataset comprising (i) 100k ranked story pairs; and (ii) a set of 46k ratings and comments on various aspects of the story.We finetune Longformer-Encoder-Decoder (LED) on the collected dataset, with the encoder responsible for preference score and aspect prediction and the decoder for comment generation.Our comprehensive experiments result in a competitive benchmark for each task, showing the high correlation to human preference.In addition, we have witnessed the joint learning of the preference scores, the aspect ratings, and the comments brings gain in each single task.Our dataset and benchmarks are publicly available to advance the research of story evaluation tasks. 1 Hong Chen 0017, Duc Minh Vo, Hiroya Takamura, Yusuke Miyao, Hideki Nakayama |
EMNLP | 2 |
| 2022 | PPCD-GAN: Progressive Pruning and Class-Aware Distillation for Large-Scale Conditional GANs CompressionabstractWe push forward neural network compression research by exploiting a novel challenging task of large-scale conditional generative adversarial networks (GANs) compression. To this end, we propose a gradually shrinking GAN (PPCD-GAN) by introducing progressive pruning residual block (PP-Res) and class-aware distillation. The PP-Res is an extension of the conventional residual block where each convolutional layer is followed by a learnable mask layer to progressively prune network parameters as training proceeds. The class-aware distillation, on the other hand, enhances the stability of training by transferring immense knowledge from a well-trained teacher model through instructive attention maps. We train the pruning and distillation processes simultaneously on a well-known GAN architecture in an end-to-end manner. After training, all redundant parameters as well as the mask layers are discarded, yielding a lighter network while retaining the performance. We comprehensively illustrate, on ImageNet 128 × 128 dataset, PPCD-GAN reduces up to 5.2 ×(81%) parameters against state-of-the-arts while keeping better performance. Duc Minh Vo, Akihiro Sugimoto, Hideki Nakayama |
WACV | 1 |
| 2022 | Paired-D++ GAN for image manipulation with text
Duc Minh Vo, Akihiro Sugimoto |
Mach. Vis. Appl. | 1 |
| 2020 | Visual-Relation Conscious Image Generation from Structured-Text
Duc Minh Vo, Akihiro Sugimoto |
ECCV (28) | 1 |
| 2020 | Stylized-Colorization for Line ArtsabstractWe address a novel problem of stylized-colorization which colorizes a given line art using a given coloring style in text. This problem can be stated as multi-domain image translation and is more challenging than the current colorization problem because it requires not only capturing the illustration distribution but also satisfying the required coloring styles specific to anime such as lightness, shading, or saturation. We propose a GAN-based end-to-end model for stylized-colorization where the model has one generator and two discriminators. Our generator is based on the U-Net architecture and receives a pair of a line art and a coloring style in text as its input to produce a stylized-colorization image of the line art. Two discriminators, on the other hand, share weights at early layers to judge the stylized-colorization image in two different aspects: one for color and one for style. One generator and two discriminators are jointly trained in an adversarial and end-to-end manner. Extensive experiments demonstrate the effectiveness of our proposed model. Tzu-Ting Fang, Duc Minh Vo, Akihiro Sugimoto, Shang-Hong Lai |
ICPR | 2 |
| 2020 | Two-stream FCNs to balance content and style for style transfer
Duc Minh Vo, Akihiro Sugimoto |
Mach. Vis. Appl. | 1 |
| 2018 | Paired-D GAN for Semantic Image Synthesis
Duc Minh Vo, Akihiro Sugimoto |
ACCV (4) | 1 |
| 2018 | Balancing Content and Style with Two-Stream FCNs for Style TransferabstractStyle transfer is to render given image contents in given styles, and it has an important role in both computer vision fundamental research and industrial applications. Following the success ofdeep learning based approaches, this problem has been re-launched very recently, but still remains a difficult task because of trade-of between preserving contents and faithful rendering of styles. In this paper, we propose an end-to-end two-stream Fully Convolutional Networks (FCNs) aiming at balancing the contributions of the content and the style in rendered images. Our proposed network consists ofthe encoder and decoder parts. The encoder part utilizes a FCN for content and a FCN for style where the two FCNs are independently trained to preserve the semantic content and to learn the faithful style representation in each. The semantic content feature and the style representationfeature are then concatenated adaptively and fed into the decoder to generate style-transferred (stylized) images. In order to train our proposed network, we employ a loss network, the pre-trained VGG-I6, to compute content loss and style loss, both of which are efficiently used for the feature concatenation. Our intensive experiments show that our proposed model generates more balanced stylized images in content and style than state-of-theart methods. Moreover, our proposed network achieves efficiency in speed. Duc Minh Vo, Trung-Nghia Le, Akihiro Sugimoto |
WACV | 1 |
| 2016 | Facial expression recognition by re-ranking with global and local generic featuresabstractRecognizing the facial expression plays an important role in human computer interaction. Following the recent success of the Convolutional Neural Network (CNN) in image classification and object recognition, this paper proposes a facial expression recognition method that makes full use of CNNs to detect face features globally and locally and that combines global and local generic features for improving accuracy in recognition. Our method uses global generic features with the Support Vector Machine (SVM) classifier to generate most plausible candidates in expression class while local generic features with the SVM classifier to look into the candidates to re-rank them for recognition. Experimental results using data-sets available in public support the effectiveness of our proposed method by demonstrating improved accuracy against the state-of-the-arts. Duc Minh Vo, Akihiro Sugimoto |
ICPR | 1 |