Yue Lin 0002

dblp:98/7264-2 · DBLP profile ↗
← Back
11ranked-venue papers
0as first author
10since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 9 since 2021Artificial intelligence and machine learning · 7 · 6 since 2021
YearPublicationVenuePosition
2024 Lamarckian Platform: Pushing the Boundaries of Evolutionary Reinforcement Learning Toward Asynchronous Commercial Games
abstract
Despite the emerging progress of integrating evolutionary computation into reinforcement learning, the absence of a high-performance platform endowing composability and massive parallelism causes nontrivial difficulties for research and applications related to asynchronous commercial games. Here, we introduce Lamarckian1—an open-source platform featuring support for evolutionary reinforcement learning scalable to distributed computing resources. To improve the training speed and data efficiency, Lamarckian adopts optimized communication methods and an asynchronous evolutionary reinforcement learning workflow. To meet the demand for an asynchronous interface by commercial games and various methods, Lamarckian tailors an asynchronous Markov Decision Process interface and designs an object-oriented software architecture with decoupled modules. In comparison with the state-of-the-art RLlib, we empirically demonstrate the unique advantages of Lamarckian on benchmark tests with up to 6000 CPU cores: 1) both the sampling efficiency and training speed are doubled when running proximal policy optimization (PPO) on Google football game; 2) the training speed is 13 times faster when running PBT+PPO onPonggame. Moreover, we also present two use cases: 1) how Lamarckian is applied to generating behavior-diverse game AI; 2) how Lamarckian is applied to game balancing tests for an asynchronous commercial game.
Ruimin Shen, Yue Lin 0002, Botian Xu, Ran Cheng 0004
IEEE Trans. Games3
2023 MIGA: A Unified Multi-Task Generation Framework for Conversational Text-to-SQL
abstract
Conversational text-to-SQL is designed to translate multi-turn natural language questions into their corresponding SQL queries. Most advanced conversational text-to-SQL methods are incompatible with generative pre-trained language models (PLMs), such as T5. In this paper, we present a two-stage unified MultI-task Generation frAmework (MIGA) that leverages PLMs’ ability to tackle conversational text-to-SQL. In the pre-training stage, MIGA first decomposes the main task into several related sub-tasks and then unifies them into the same sequence-to-sequence (Seq2Seq) paradigm with task-specific natural language prompts to boost the main task from multi-task training. Later in the fine-tuning stage, we propose four SQL perturbations to alleviate the error propagation problem. MIGA tends to achieve state-of-the-art performance on two benchmarks (SparC and CoSQL). We also provide extensive analyses and discussions to shed light on some new perspectives for conversational text-to-SQL.
Yingwen Fu, Wenjie Ou, Yue Lin 0002
AAAI4
2023 General Image-to-Image Translation with One-Shot Image Guidance
abstract
Large-scale text-to-image models pre-trained on massive text-image pairs show excellent performance in image synthesis recently. However, image can provide more intuitive visual concepts than plain text. People may ask: how can we integrate the desired visual concept into an existing image, such as our portrait? Current methods are inadequate in meeting this demand as they lack the ability to preserve content or translate visual concepts effectively. Inspired by this, we propose a novel framework named visual concept translator (VCT) with the ability to preserve content in the source image and translate the visual concepts guided by a single reference image. The proposed VCT contains a content-concept inversion (CCI) process to extract contents and concepts, and a content-concept fusion (CCF) process to gather the extracted information to obtain the target image. Given only one reference image, the proposed VCT can complete a wide range of general image-to-image translation tasks with excellent results. Extensive experiments are conducted to prove the superiority and effectiveness of the proposed methods. Codes are available at https://github.com/CrystalNeuro/visual-concept-translator.
Zuhao Liu 0002, Yunbo Peng, Yue Lin 0002
ICCV4
2023 DR-Occluder: Generating Occluders Using Differentiable Rendering
abstract
The target of the occluder is to use very few faces to maintain similar occlusion properties of the original 3D model. In this paper, we present DR-Occluder, a novel coarse-to-fine framework for occluder generation that leverages differentiable rendering to optimize a triangle set to an occluder. Unlike prior work, which has not utilized differentiable rendering for this task, our approach provides the ability to optimize a 3D shape to defined targets. Given a 3D model as input, our method first projects it to silhouette images, which are then processed by a convolution network to output a group of vertex offsets. These offsets are used to transform a group of distributed triangles into a preliminary occluder, which is further optimized by differentiable rendering. Finally, triangles whose area is smaller than a threshold are removed to obtain the final occluder. Our extensive experiments demonstrate that DR-Occluder significantly outperforms state-of-the-art methods in terms of occlusion quality. Furthermore, we compare the performance of our method with other approaches in a commercial engine, providing compelling evidence of its effectiveness.
Jiaxian Wu, Yue Lin 0002, Dehui Lu
ACM Trans. Graph.2
2022 Spatial-Temporal Parallel Transformer for Arm-Hand Dynamic Estimation
abstract
We propose an approach to estimate arm and hand dynamics from monocular video by utilizing the relationship between arm and hand. Although monocular full human motion capture technologies have made great progress in recent years, recovering accurate and plausible arm twists and hand gestures from in-the-wild videos still remains a challenge. To solve this problem, our solution is proposed based on the fact that arm poses and hand gestures are highly correlated in most real situations. To fully exploit arm-hand correlation as well as inter-frame information, we carefully design a Spatial-Temporal Parallel Arm-Hand Motion Transformer (PAHMT) to predict the arm and hand dynamics simultaneously. We also introduce new losses to encourage the estimations to be smooth and accurate. Besides, we collect a motion capture dataset including 200K frames of hand gestures and use this data to train our model. By integrating a 2D hand pose estimation model and a 3D human pose estimation model, the proposed method can produce plausible arm and hand dynamics from monocular video. Extensive evaluations demonstrate that the proposed method has advantages over previous state-of-the-art approaches and shows robustness under various challenging scenarios.
Shuying Liu, Jiaxian Wu, Yue Lin 0002
CVPR4
2022 DGC-Vector: A New Speaker Embedding for Zero-Shot Voice Conversion
abstract
Recently, more and more zero-shot voice conversion algorithms have been proposed. As a fundamental part of zero-shot voice conversion, speaker embeddings are the key to improving the converted speech’s speaker similarity. In this paper, we study the impact of speaker embeddings on zero-shot voice conversion performance. To better represent the characteristics of the target speaker and improve the speaker similarity in zero-shot voice conversion, we propose a novel speaker representation method in this paper. Our method combines the advantages of D-vector, global style token (GST) based speaker representation and auxiliary supervision. Objective and subjective evaluations show that the proposed method achieves a decent performance on zero-shot voice conversion and significantly improves speaker similarity over D-vector and GST-based speaker embedding.
Ruitong Xiao, Haitong Zhang, Yue Lin 0002
ICASSP3
2022 Improve Few-Shot Voice Cloning Using Multi-Modal Learning
abstract
Recently, few-shot voice cloning has achieved a significant improvement. However, most models for few-shot voice cloning are single-modal, and multi-modal few-shot voice cloning has been understudied. In this paper, we propose to use multi-modal learning to improve the few-shot voice cloning performance. Inspired by the recent works on un-supervised speech representation, the proposed multi-modal system is built by extending Tacotron2 with an unsupervised speech representation module. We evaluate our proposed system in two few-shot voice cloning scenarios, namely few-shot text-to-speech (TTS) and voice conversion (VC). Experimental results demonstrate that the proposed multi-modal learning can significantly improve the few-shot voice cloning performance over their counterpart single-modal systems.
Haitong Zhang, Yue Lin 0002
ICASSP2
2022 Data Augmentation for Long-Tailed and Imbalanced Polyphone Disambiguation in Mandarin
abstract
Polyphone disambiguation is an important module in Mandarin Chinese text-to-speech (TTS). Recently, neural-network-based (NN-based) models have achieved a great improvement on poly-phone disambiguation. However, a long-tailed and imbalanced distribution is usually observed in the training data of polyphone disambiguation, resulting in an unsatisfying performance on the low-frequent polyphone in the imbalanced pinyin set, and the least-frequent polyphonic characters and polyphones. In this paper, we proposed a simple data-augmentation method based on the pre-trained mask language model BERT to mitigate the long-tailed and imbalanced distribution problem. We incorporate a weighted sampling technique in the data augmentation method to balance the data distribution, and a useful filtering strategy to remove some noisy augmented data. Experimental results show that the proposed data-augmentation method can improve the prediction accuracy, especially for those low-frequent polyphone in the imbalanced pinyin set, and the least-frequent polyphonic characters and polyphones.
Haitong Zhang, Yue Lin 0002
ICASSP3
2022 Exploring Timbre Disentanglement in Non-Autoregressive Cross-Lingual Text-to-Speech
abstract
In this paper, we study the disentanglement of speaker and language representations in non-autoregressive cross-lingual TTS models from various aspects.We propose a phoneme length regulator that solves the length mismatch problem between IPA input sequence and monolingual alignment results.Using the phoneme length regulator, we present a FastPitch-based crosslingual model with IPA symbols as input representations.Our experiments show that language-independent input representations (e.g.IPA symbols), an increasing number of training speakers, and explicit modeling of speech variance information all encourage non-autoregressive cross-lingual TTS model to disentangle speaker and language representations.The subjective evaluation shows that our proposed model can achieve decent naturalness and speaker similarity in cross-language voice cloning.
Haoyue Zhan, Xinyuan Yu, Haitong Zhang, Yue Lin 0002
INTERSPEECH5
2021 Improve Cross-Lingual Text-To-Speech Synthesis on Monolingual Corpora with Pitch Contour Information
Haoyue Zhan, Haitong Zhang, Wenjie Ou, Yue Lin 0002
Interspeech4
2020 Unsupervised Learning for Sequence-to-Sequence Text-to-Speech for Low-Resource Languages
abstract
Recently, sequence-to-sequence models with attention have been successfully applied in Text-to-speech (TTS). These models can generate near-human speech with a large accurately-transcribed speech corpus. However, preparing such a large data-set is both expensive and laborious. To alleviate the problem of heavy data demand, we propose a novel unsupervised pre-training mechanism in this paper. Specifically, we first use Vector-quantization Variational-Autoencoder (VQ-VAE) to ex-tract the unsupervised linguistic units from large-scale, publicly found, and untranscribed speech. We then pre-train the sequence-to-sequence TTS model by using the pairs. Finally, we fine-tune the model with a small amount of paired data from the target speaker. As a result, both objective and subjective evaluations show that our proposed method can synthesize more intelligible and natural speech with the same amount of paired training data. Besides, we extend our proposed method to the hypothesized low-resource languages and verify the effectiveness of the method using objective evaluation.
Haitong Zhang, Yue Lin 0002
INTERSPEECH2