VLDB 2026 Research / reviewers in the wild / expert
Jiali Yao
dblp:26/7253
· DBLP profile ↗
12ranked-venue papers
1as first author
9since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Software engineering, systems software and programming languages · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | High-Fidelity Image Inpainting with Multimodal Guided GAN Inversion
Libo Zhang 0001, Jiali Yao, Heng Fan 0001 |
Int. J. Comput. Vis. | 3 |
| 2024 | Beyond MOT: Semantic Multi-object Tracking
Hao Wang 0093, Jiali Yao, Shaohua Dong, Heng Fan 0001, Libo Zhang 0001 |
ECCV (35) | 5 |
| 2023 | An Image Extraction Method for Traditional Dress Pattern Line Drawings Based on Improved CycleGAN
Xingquan Cai, Sichen Jia, Jiali Yao, Haiyan Sun |
CGI | 3 |
| 2023 | An Ancient Murals Inpainting Method Based on Bidirectional Feature Adaptation and Adversarial Generative Networks
Xingquan Cai, Qingtao Lu, Jiali Yao |
CGI | 3 |
| 2023 | A Driver Abnormal Behavior Detection Method Based on Improved YOLOv7 and OpenPose
Xingquan Cai, Jiali Yao, Pengyan Cheng |
ICIC (5) | 3 |
| 2023 | An Extended Labanotation Generation Method Based on 3D Human Pose Estimation for Intangible Cultural Heritage Dance VideosabstractTo address the issues of low accuracy in existing 3D human pose estimation (HPE) methods and the limited level of details in Labanotation, we propose an extended Labanotation generation method for intangible cultural heritage dance videos based on 3D HPE. First, a 2D human pose sequence of the performer is inputted along with spatial location embeddings, where multiple spatial transformer modules are employed to extract spatial features of human joints and generate cross-joint multiple hypotheses. Afterward, temporal features are extracted by a self-attentive module and the correlation between different hypotheses is learned using bilinear pooling. Finally, the 3D joint coordinates of the performer are predicted, which are matched with the corresponding extended Labanotation symbols using the Laban template matching method to generate extended Labanotation. Experimental results show that, compared with VideoPose and CrossFormer algorithms, the Mean Per Joint Position Error (MPJPE) of the proposed method is reduced by 3.7[Formula: see text]mm and 0.6[Formula: see text]mm, respectively on Human3.6M dataset, and the generated extended Labanotation can better describe the movement details compared with the basic Labanotation. Xingquan Cai, Pengyan Cheng, Jiali Yao |
Int. J. Pattern Recognit. Artif. Intell. | 4 |
| 2023 | A Digital Simulation and Re-Editing Method for Clothing Patterns Based on Deep Learning and Somatosensory InteractionabstractTo address the issues in clothing pattern style migration, this paper proposes a digital simulation and re-editing method for clothing patterns based on deep learning and somatosensory interaction. First, the proposed method encodes the black-and-white line drawing image, generating random noise images through a diffusion process, introducing color information for synthesis, and using a decoder to reconstruct a colored image. Afterwards, an improved VGG19 model is used to reconstruct content features and perform linear color transformation on style images, enabling pattern style migration through the construction of a Gram matrix and resulting in colored clothing texture patterns. Finally, a KinectV2 is utilized for fabric simulation, overlaying colorful clothing texture patterns to achieve 3D virtual dressing. The experimental results show that the proposed method improves the structural similarity index measure (SSIM) by 9–11% and the peak signal-to-noise ratio (PSNR) by 3–8% when compared to existing algorithms. The experiments provide evidence that the proposed method effectively mitigates color overflow, delivers precise image coloring, and accomplishes realistic restoration of clothing texture. Furthermore, the method offers an improved garment fit to fulfill the user’s interaction requirements. Haiyan Sun, Jiali Yao, Xingquan Cai |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2021 | Phone Distribution Estimation for Low Resource Languages
Juncheng Li 0001, Jiali Yao, Alan W. Black, Florian Metze |
ICASSP | 3 |
| 2021 | Towards Realistic Visual Dubbing with Heterogeneous SourcesabstractThe task of few-shot visual dubbing focuses on synchronizing the lip movements with arbitrary speech input for any talking head video. Albeit moderate improvements in current approaches, they commonly require high-quality homologous data sources of videos and audios, thus causing the failure to leverage heterogeneous data sufficiently. In practice, it may be intractable to collect the perfect homologous data in some cases, for example, audio-corrupted or picture-blurry videos. To explore this kind of data and support high-fidelity few-shot visual dubbing, in this paper, we novelly propose a simple yet efficient two-stage framework with a higher flexibility of mining heterogeneous data. Specifically, our two-stage paradigm employs facial landmarks as intermediate prior of latent representations and disentangles the lip movements prediction from the core task of realistic talking head generation. By this means, our method makes it possible to independently utilize the training corpus for two-stage sub-networks using more available heterogeneous data easily acquired. Besides, thanks to the disentanglement, our framework allows a further fine-tuning for a given talking head, thereby leading to better speaker-identity preserving in the final synthesized results. Moreover, the proposed method can also transfer appearance features from others to the target speaker. Extensive experimental results demonstrate the superiority of our proposed method in generating highly realistic videos synchronized with the speech over the state-of-the-art. Tianyi Xie, Liucheng Liao, Benlai Tang, Xiang Yin 0006, Jianfei Yang 0001, Jiali Yao, Yang Zhang 0088, Zejun Ma 0001 |
ACM Multimedia | 8 |
| 2020 | Universal Phone Recognition with a Multilingual Allophone SystemabstractMultilingual models can improve language processing, particularly for low resource situations, by sharing parameters across languages. Multilingual acoustic models, however, generally ignore the difference between phonemes (sounds that can support lexical contrasts in a particular language) and their corresponding phones (the sounds that are actually spoken, which are language independent). This can lead to performance degradation when combining a variety of training languages, as identically annotated phonemes can actually correspond to several different underlying phonetic realizations. In this work, we propose a joint model of both language-independent phone and language-dependent phoneme distributions. In multilingual ASR experiments over 11 languages, we find that this model improves testing performance by 2% phoneme error rate absolute in low-resource conditions. Additionally, because we are explicitly modeling language-independent phones, we can build a (nearly-)universal phone recognizer that, when combined with the PHOIBLE [1] large, manually curated database of phone inventories, can be customized into 2,000 language dependent recognizers. Experiments on two low-resourced indigenous languages, Inuktitut and Tusom, show that our recognizer achieves phone accuracy improvements of more than 17%, moving a step closer to speech recognition for all languages in the world.1 Siddharth Dalmia, Juncheng Li 0001, Matthew Lee 0012, Patrick Littell, Jiali Yao, Antonios Anastasopoulos, David R. Mortensen, Graham Neubig, Alan W. Black, Florian Metze |
ICASSP | 6 |
| 2013 | Test Quality Measurement Using TBPP-RabstractSoftware test quality measurement is the key for software release decision. It is hard to evaluate software test quality because full coverage test is impossible in practice. In this paper, we compare the existing methods for the software test quality measurement and discuss their drawbacks. Then we propose TBPP - a weighted tree based partition and proration method to evaluate software test quality. Furthermore, we evolve this method to TBPP-R by introducing a risk distribution function, which can reveal software test quality more accurately. In our experiments, we compare TBPP-R and TBPP with existing methods in our test projects. The results show that TPBB-R can measure software test quality much more accurately. Ruilong Huo, Jiali Yao, Shaosen Wu |
ICST | 3 |
| 2009 | A Distributed Render Farm System for Animation Production
Jiali Yao, Hongxin Zhang 0001 |
ICEC | 1 |