VLDB 2026 Research / reviewers in the wild / expert
Yuan Gong 0002
dblp:98/8660-2
· DBLP profile ↗
5ranked-venue papers
2as first author
5since 2021 · last 2023
0009-0009-9097-4805ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Generative modeling · 27% Vision and language · 17% Image recognition and object detection · 14% | |
| Computer graphics and multimedia
1 paper |
Visual content generation and editing · 91% Computer animation and physical simulation · 9% |
Topics — the 17 heaviest of 18, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › 3D vision
3d face reconstruction |
0.7 | 1 | 2023 | 3D GAN Inversion with Facial Symmetry Prior · CVPR 2023 |
Machine learning › Generative modeling › generative adversarial network › GAN inversion
3D GAN inversion |
0.7 | 1 | 2023 | 3D GAN Inversion with Facial Symmetry Prior · CVPR 2023 |
Computer vision › Vision and language
cross-modal alignment |
0.7 | 1 | 2023 | MAP: Multimodal Uncertainty-Aware Vision-Language Pre-training Model · CVPR 2023 |
Computer vision › Face, body and person analysis › face manipulation
face reenactment |
0.7 | 1 | 2023 | ToonTalker: Cross-Domain Face Reenactment · ICCV 2023 |
Machine learning › Generative modeling › generative adversarial network
GAN inversion |
0.7 | 1 | 2023 | 3D GAN Inversion with Facial Symmetry Prior · CVPR 2023 |
Machine learning › Generative modeling › video generation
motion transfer |
0.7 | 1 | 2023 | ToonTalker: Cross-Domain Face Reenactment · ICCV 2023 |
Machine learning › Trustworthy machine learning › uncertainty estimation › uncertainty-aware learning
multimodal uncertainty modeling |
0.7 | 1 | 2023 | MAP: Multimodal Uncertainty-Aware Vision-Language Pre-training Model · CVPR 2023 |
Computer vision › Vision and language
vision-language pretraining |
0.7 | 1 | 2023 | MAP: Multimodal Uncertainty-Aware Vision-Language Pre-training Model · CVPR 2023 |
Visual content generation and editing
controllable generation |
0.7 | 1 | 2023 | Interactive Story Visualization with Multiple Characters · SIGGRAPH Asia 2023 |
Visual content generation and editing › multimodal content generation
story visualization |
0.7 | 1 | 2023 | Interactive Story Visualization with Multiple Characters · SIGGRAPH Asia 2023 |
Visual content generation and editing › image generation
text-to-image generation |
0.7 | 1 | 2023 | Interactive Story Visualization with Multiple Characters · SIGGRAPH Asia 2023 |
Machine learning › Efficient and distributed learning › model compression
knowledge distillation |
0.6 | 1 | 2022 | Focal and Global Knowledge Distillation for Detectors · CVPR 2022 |
Computer vision › Image recognition and object detection › object detection
knowledge distillation for detection |
0.6 | 1 | 2022 | Focal and Global Knowledge Distillation for Detectors · CVPR 2022 |
Machine learning › Efficient and distributed learning
model compression |
0.6 | 1 | 2022 | Focal and Global Knowledge Distillation for Detectors · CVPR 2022 |
Computer vision › Image recognition and object detection
object detection |
0.6 | 1 | 2022 | Focal and Global Knowledge Distillation for Detectors · CVPR 2022 |
Machine learning › Generative modeling › generative adversarial network
3d-aware image synthesis |
0.2 | 1 | 2023 | 3D GAN Inversion with Facial Symmetry Prior · CVPR 2023 |
Machine learning › Transfer learning and domain adaptation
domain adaptation |
0.2 | 1 | 2023 | ToonTalker: Cross-Domain Face Reenactment · ICCV 2023 |
Methods — techniques the papers use, named apart from their topics
transformer · 0.7text-to-image model · 0.7probability distribution encoder · 0.7motion encoder · 0.7masked language modeling · 0.7large language model · 0.7facial symmetry prior · 0.7depth-guided 3d warping · 0.7contrastive learning · 0.7analogy constraint · 0.7focal distillation · 0.6feature map distillation · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | MAP: Multimodal Uncertainty-Aware Vision-Language Pre-training ModelabstractMultimodal semantic understanding often has to deal with uncertainty, which means the obtained messages tend to refer to multiple targets. Such uncertainty is problematic for our interpretation, including inter- and intra-modal uncertainty. Little effort has studied the modeling of this uncertainty, particularly in pretraining on unlabeled datasets and fine-tuning in task-specific downstream datasets. In this paper, we project the representations of all modalities as probabilistic distributions via a Probability Distribution Encoder (PDE) by utilizing sequence-level interactions. Compared to the existing deterministic methods, such uncertainty modeling can convey richer multimodal semantic information and more complex relationships. Furthermore, we integrate uncertainty modeling with popular pretraining frameworks and propose suitable pretraining tasks: Distribution-based Vision-Language Contrastive learning (D-VLC), Distribution-based Masked Language Modeling (D-MLM), and Distribution-based Image-Text Matching (D-ITM). The fine-tuned models are applied to challenging downstream tasks, including image-text retrieval, visual question answering, visual reasoning, and visual entailment, and achieve state-of-the-art results. Yatai Ji, Junjie Wang 0011, Yuan Gong 0002, Yanru Zhu, Hongfa Wang, Jiaxing Zhang 0001, Tetsuya Sakai, Yujiu Yang 0001 |
CVPR | 3 |
| 2023 | 3D GAN Inversion with Facial Symmetry PriorabstractRecently, a surge of high-quality 3D-aware GANs have been proposed, which leverage the generative power of neural rendering. It is natural to associate 3D GANs with GAN inversion methods to project a real image into the generator's latent space, allowing free-view consistent synthesis and editing, referred as 3D GAN inversion. Although with the facial prior preserved in pre-trained 3D GANs, reconstructing a 3D portrait with only one monocular image is still an ill-pose problem. The straightforward application of 2D GAN inversion methods focuses on texture similarity only while ignoring the correctness of 3D geometry shapes. It may raise geometry collapse effects, especially when reconstructing a side face under an extreme pose. Besides, the synthetic results in novel views are prone to be blurry. In this work, we propose a novel method to promote 3D GAN inversion by introducing facial symmetry prior. We design a pipeline and constraints to make full use of the pseudo auxiliary view obtained via image flipping, which helps obtain a view-consistent and well-structured geometry shape during the inversion process. To enhance texture fidelity in unobserved viewpoints, pseudo labels from depth-guided 3D warping can provide extra supervision. We design constraints to filter out conflict areas for optimization in asymmetric situations. Comprehensive quantitative and qualitative evaluations on image reconstruction and editing demonstrate the superiority of our method. Yong Zhang 0034, Xuan Wang 0009, Tengfei Wang 0002, Xiaoyu Li 0002, Yuan Gong 0002, Yanbo Fan, Xiaodong Cun, Ying Shan, A. Cengiz Öztireli, Yujiu Yang 0001 |
CVPR | 6 |
| 2023 | ToonTalker: Cross-Domain Face ReenactmentabstractWe target cross-domain face reenactment in this paper, i.e., driving a cartoon image with the video of a real person and vice versa. Recently, many works have focused on one-shot talking face generation to drive a portrait with a real video, i.e., within-domain reenactment. Straightforwardly applying those methods to cross-domain animation will cause inaccurate expression transfer, blur effects, and even apparent artifacts due to the domain shift between cartoon and real faces. Only a few works attempt to settle cross-domain face reenactment. The most related work AnimeCeleb [13] requires constructing a dataset with pose vector and cartoon image pairs by animating 3D characters, which makes it inapplicable anymore if no paired data is available. In this paper, we propose a novel method for cross-domain reenactment without paired data. Specifically, we propose a transformer-based framework to align the motions from different domains into a common latent space where motion transfer is conducted via latent code addition. Two domain-specific motion encoders and two learnable motion base memories are used to capture domain properties. A source query transformer and a driving one are exploited to project domain-specific motion to the canonical space. The edited motion is projected back to the domain of the source with a transformer. Moreover, since no paired data is provided, we propose a novel cross-domain training scheme using data from two domains with the designed analogy constraint. Besides, we contribute a cartoon dataset in Disney style. Extensive evaluations demonstrate the superiority of our method over competing methods. Yuan Gong 0002, Yong Zhang 0034, Xiaodong Cun, Yanbo Fan, Xuan Wang 0009, Baoyuan Wu, Yujiu Yang 0001 |
ICCV | 1 |
| 2023 | Interactive Story Visualization with Multiple CharactersabstractAccurate Story visualization requires several necessary elements, such as identity consistency across frames, the alignment between plain text and visual content, and a reasonable layout of objects in images. Most previous works endeavor to meet these requirements by fitting a text-to-image (T2I) model on a set of videos in the same style and with the same characters, e.g., the FlintstonesSV dataset. However, the learned T2I models typically struggle to adapt to new characters, scenes, and styles, and often lack the flexibility to revise the layout of the synthesized images. This paper proposes a system for generic interactive story visualization, capable of handling multiple novel characters and supporting the editing of layout and local structure. It is developed by leveraging the prior knowledge of large language and T2I models, trained on massive corpora. The system comprises four interconnected components: story-to-prompt generation (S2P), text-to-layout generation (T2L), controllable text-to-image generation (C-T2I), and image-to-video animation (I2V). First, the S2P module converts concise story information into detailed prompts required for subsequent stages. Next, T2L generates diverse and reasonable layouts based on the prompts, offering users the ability to adjust and refine the layout to their preferences. The core component, C-T2I, enables the creation of images guided by layouts, sketches, and actor-specific identifiers to maintain consistency and detail across visualizations. Finally, I2V enriches the visualization process by animating the generated images. Extensive experiments and a user study are conducted to validate the effectiveness and flexibility of interactive editing of the proposed system. Yuan Gong 0002, Youxin Pang, Xiaodong Cun, Menghan Xia, Yingqing He, Haoxin Chen, Longyue Wang, Yong Zhang 0034, Xintao Wang 0002, Ying Shan, Yujiu Yang 0001 |
SIGGRAPH Asia | 1 |
| 2022 | Focal and Global Knowledge Distillation for DetectorsabstractKnowledge distillation has been applied to image classification successfully. However, object detection is much more sophisticated and most knowledge distillation methods have failed on it. In this paper, we point out that in object detection, the features of the teacher and student vary greatly in different areas, especially in the foreground and background. If we distill them equally, the uneven differences between feature maps will negatively affect the distillation. Thus, we propose Focal and Global Distillation (FGD). Focal distillation separates the foreground and background, forcing the student to focus on the teacher's critical pixels and channels. Global distillation rebuilds the relation between different pixels and transfers it from teachers to students, compensating for missing global information in focal distillation. As our method only needs to calculate the loss on the feature map, FGD can be applied to various detectors. We experiment on various detectors with different backbones and the results show that the student detector achieves excellent mAP improvement. For example, ResNet-50 based RetinaNet, Faster RCNN, RepPoints and Mask RCNN with our distillation method achieve 40.7%, 42.0%, 42.0% and 42.1% mAP on COCO2017, which are 3.3, 3.6, 3.4 and 2.9 higher than the baseline, respectively. Our codes are available at https://github.com/yzd-v/FGD. Zhendong Yang, Xiaohu Jiang, Yuan Gong 0002, Zehuan Yuan, Danpei Zhao, Chun Yuan 0003 |
CVPR | 4 |