Yingru Liu

dblp:160/6116 · DBLP profile ↗
← Back
11ranked-venue papers
4as first author
4since 2021 · last 2025
0000-0002-3784-4886ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Speech recognition and synthesis · 37% Vision and language · 17% Representation and self-supervised learning · 9%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Smart cities and intelligent transportation · 100%

Topics — the 14 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language
image captioning
1.022022
ReFormer: The Relational Transformer for Image Captioning · ACM Multimedia 2022
Fashion Captioning: Towards Generating Accurate Descriptions with Semantic Rewards · ECCV (13) 2020
Machine learning › Representation and self-supervised learning › representation learning
disentangled representation learning
0.912025
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement · ICLR 2025
Natural language and speech › Speech recognition and synthesis › speech synthesis
text-to-speech
0.912025
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement · ICLR 2025
Natural language and speech › Speech recognition and synthesis
voice conversion
0.912025
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement · ICLR 2025
Natural language and speech › Speech recognition and synthesis › text-to-speech synthesis
zero-shot text-to-speech
0.912025
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement · ICLR 2025
Natural language and speech › Speech recognition and synthesis › text-to-speech synthesis
zero-shot voice cloning
0.912025
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement · ICLR 2025
Computer vision › Vision and language › image captioning › structured image captioning
scene graph-based captioning
0.612022
ReFormer: The Relational Transformer for Image Captioning · ACM Multimedia 2022
Computer vision › Segmentation and scene understanding
scene graph generation
0.612022
ReFormer: The Relational Transformer for Image Captioning · ACM Multimedia 2022
Machine learning › Learning paradigms
multi-task learning
0.412020
Adaptive Activation Network and Functional Regularization for Efficient and Flexible Deep Multi-Task Learning · AAAI 2020
Machine learning › Efficient and distributed learning › automated machine learning
neural architecture search
0.412020
Adaptive Activation Network and Functional Regularization for Efficient and Flexible Deep Multi-Task Learning · AAAI 2020
Machine learning › Graph learning
graph neural network
0.412019
Dynamic Spatial-Temporal Graph Convolutional Neural Networks for Traffic Forecasting · AAAI 2019
Natural language and speech › Machine translation
neural machine translation
0.412019
Latent Part-of-Speech Sequences for Neural Machine Translation · EMNLP/IJCNLP (1) 2019
Machine learning › Graph learning › graph neural network › graph convolutional network
spatial-temporal graph convolution
0.412019
Dynamic Spatial-Temporal Graph Convolutional Neural Networks for Traffic Forecasting · AAAI 2019
Smart cities and intelligent transportation
traffic prediction
0.112019
Dynamic Spatial-Temporal Graph Convolutional Neural Networks for Traffic Forecasting · AAAI 2019

Methods — techniques the papers use, named apart from their topics

flow matching · 0.9autoregressive transformer · 0.9VQ-VAE · 0.9HuBERT · 0.9graph convolutional neural network · 0.8relational transformer · 0.6semantic reward · 0.4reinforcement learning · 0.4functional regularization · 0.4activation function learning · 0.4tensor decomposition · 0.4
YearPublicationVenuePosition
2025 Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement
abstract
The imitation of voice, targeted on specific speech attributes such as timbre and speaking style, is crucial in speech generation. However, existing methods rely heavily on annotated data, and struggle with effectively disentangling timbre and style, leading to challenges in achieving controllable generation, especially in zero-shot scenarios. To address these issues, we propose Vevo, a versatile zero-shot voice imitation framework with controllable timbre and style. Vevo operates in two core stages: (1) Content-Style Modeling: Given either text or speech's content tokens as input, we utilize an autoregressive transformer to generate the content-style tokens, which is prompted by a style reference; (2) Acoustic Modeling: Given the content-style tokens as input, we employ a flow-matching transformer to produce acoustic representations, which is prompted by a timbre reference. To obtain the content and content-style tokens of speech, we design a fully self-supervised approach that progressively decouples the timbre, style, and linguistic content of speech. Specifically, we adopt VQ-VAE as the tokenizer for the continuous hidden features of HuBERT. We treat the vocabulary size of the VQ-VAE codebook as the information bottleneck, and adjust it carefully to obtain the disentangled speech representations. Solely self-supervised trained on 60K hours of audiobook speech data, without any fine-tuning on style-specific corpora, Vevo matches or surpasses existing methods in accent and emotion conversion tasks. Additionally, Vevo’s effectiveness in zero-shot voice conversion and text-to-speech tasks further demonstrates its strong generalization and versatility. Audio samples are available at https://versavoice.github.io/.
Xueyao Zhang, Kainan Peng, Vimal Manohar, Yingru Liu, Jeff Hwang, Dangna Li, Julian Chan, Zhizheng Wu 0001, Mingbo Ma
ICLR6
2023 AGGDN: A Continuous Stochastic Predictive Model for Monitoring Sporadic Time Series on Graphs
Yucheng Xing, Jacqueline Wu, Yingru Liu, Xin Wang 0001
ICONIP (1)3
2022 ReFormer: The Relational Transformer for Image Captioning
abstract
Image captioning is shown to be able to achieve a better performance by using scene graphs to represent the relations of objects in the image. The current captioning encoders generally use a Graph Convolutional Net (GCN) to represent the relation information and merge it with the object region features via concatenation or convolution to get the final input for sentence decoding. However, the GCN-based encoders in the existing methods are less effective for captioning due to two reasons. First, using the image captioning as the objective (i.e., Maximum Likelihood Estimation) rather than a relation-centric loss cannot fully explore the potential of the encoder. Second, using a pre-trained model instead of the encoder itself to extract the relationships is not flexible and cannot contribute to the explainability of the model. To improve the quality of image captioning, we propose a novel architecture ReFormer- a RElational transFORMER to generate features with relation information embedded and to explicitly express the pair-wise relationships between objects in the image. ReFormer incorporates the objective of scene graph generation with that of image captioning using one modified Transformer model. This design allows ReFormer to generate not only better image captions with the benefit of extracting strong relational image features, but also scene graphs to explicitly describe the pair-wise relationships. Experiments on publicly available datasets show that our model significantly outperforms state-of-the-art methods on image captioning and scene graph generation.
Yingru Liu, Xin Wang 0001
ACM Multimedia2
2021 Continuous-Time Stochastic Differential Networks for Irregular Time Series Modeling
Yingru Liu, Yucheng Xing, Xin Wang 0001, Zhaoyue Chen, Jacqueline Wu
ICONIP (5)1
2020 Adaptive Activation Network and Functional Regularization for Efficient and Flexible Deep Multi-Task Learning
abstract
Multi-task learning (MTL) is a common paradigm that seeks to improve the generalization performance of task learning by training related tasks simultaneously. However, it is still a challenging problem to search the flexible and accurate architecture that can be shared among multiple tasks. In this paper, we propose a novel deep learning model called Task Adaptive Activation Network (TAAN) that can automatically learn the optimal network architecture for MTL. The main principle of TAAN is to derive flexible activation functions for different tasks from the data with other parameters of the network fully shared. We further propose two functional regularization methods that improve the MTL performance of TAAN. The improved performance of both TAAN and the regularization methods is demonstrated by comprehensive experiments.
Yingru Liu, Dongliang Xie, Xin Wang 0001, Li Shen 0008, Hao-Zhi Huang 0001, Niranjan Balasubramanian
AAAI1
2020 Fashion Captioning: Towards Generating Accurate Descriptions with Semantic Rewards
Heming Zhang 0003, Yingru Liu, Chihao Wu 0001, Jianchao Tan, Dongliang Xie, Jue Wang 0001, Xin Wang 0001
ECCV (13)4
2020 Robust Adaptive Control of an Offshore Ocean Thermal Energy Conversion System
abstract
Boundary control strategy is developed to analyze the vibration problem of the offshore ocean thermal energy conversion (OTEC) system as well as to constrain the bottom tension and top motion. To provide an accurate dynamic behavior for the OTEC system, this distributed parameter system is modeled and formulated with a governing equation and boundary conditions (PDE-ODEs model). Two robust adaptive boundary controllers are designed and disposed at the endpoints of the system, and the stability of the controlled system under unknown disturbances is achieved. After selecting the relevant parameters appropriately, the offset of the offshore OTEC system can be suppressed to equilibrium position. Finally, the effectiveness of the proposed control is illustrated by simulation.
Xiuyu He, Wei He 0001, Yingru Liu, Guang Li 0002, Yu Wang 0062
IEEE Trans. Syst. Man Cybern. Syst.3
2019 Dynamic Spatial-Temporal Graph Convolutional Neural Networks for Traffic Forecasting
abstract
Graph convolutional neural networks (GCNN) have become an increasingly active field of research. It models the spatial dependencies of nodes in a graph with a pre-defined Laplacian matrix based on node distances. However, in many application scenarios, spatial dependencies change over time, and the use of fixed Laplacian matrix cannot capture the change. To track the spatial dependencies among traffic data, we propose a dynamic spatio-temporal GCNN for accurate traffic forecasting. The core of our deep learning framework is the finding of the change of Laplacian matrix with a dynamic Laplacian matrix estimator. To enable timely learning with a low complexity, we creatively incorporate tensor decomposition into the deep learning framework, where real-time traffic data are decomposed into a global component that is stable and depends on long-term temporal-spatial traffic relationship and a local component that captures the traffic fluctuations. We propose a novel design to estimate the dynamic Laplacian matrix of the graph with above two components based on our theoretical derivation, and introduce our design basis. The forecasting performance is evaluated with two realtime traffic datasets. Experiment results demonstrate that our network can achieve up to 25% accuracy improvement.
Zulong Diao, Xin Wang 0001, Da-Fang Zhang 0001, Yingru Liu, Kun Xie 0001, Shaoyao He
AAAI4
2019 Generalized Boltzmann Machine with Deep Neural Structure
abstract
Restricted Boltzmann Machine (RBM) is an essential component in many machine learning applications. As a probabilistic graphical model, RBM posits a shallow structure, which makes it less capable of modeling real-world applications. In this paper, to bridge the gap between RBM and artificial neural network, we propose an energy-based probabilistic model that is more flexible on modeling continuous data. By introducing the pair-wise inverse autoregressive flow into RBM, we propose two generalized continuous RBMs which contain deep neural network structure to more flexibly track the practical data distribution while still keeping the inference tractable. In addition, we extend the generalized RBM structures into sequential setting to better model the stochastic process of time series. Performance improvements on probabilistic modeling and representation learning are demonstrated by the experiments on diverse datasets.
Yingru Liu, Dongliang Xie, Xin Wang 0001
AISTATS1
2019 Latent Part-of-Speech Sequences for Neural Machine Translation
abstract
Xuewen Yang, Yingru Liu, Dongliang Xie, Xin Wang, Niranjan Balasubramanian. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Yingru Liu, Dongliang Xie, Xin Wang 0001, Niranjan Balasubramanian
EMNLP/IJCNLP (1)2
2019 Energy-Based Recurrent Model for Stochastic Modeling of Music
abstract
The aim of this work is to more accurately model the stochastic process of music-related data, which is essential for many AI applications in musicology. When music is naturally represented as a sequence of vectorized frames, existing models generally cannot well capture the correlation of the elements inside each frame. We propose an energy-based model called Chain Graphical Recurrent Neural Network (CGRNN) to explore the correlation of elements for more accurate modeling of the dynamics of music. In CGRNN, a probabilistic substructure named Conditional spike and slab Restricted Boltzmann Machine (C-ssRBM) is defined to better model the conditional covariance and joint distribution of elements in a frame. Besides, CGRNN is capable of tracking the evolution of music and extracting sparse features with an efficient design of temporal transition. With the estimated stochastic process of music, we further implement CGRNN to generate melodious music automatically. Extensive empirical evaluations of multiple unsupervised learning tasks are conducted on symbolic MIDI and audio sounds to demonstrate the performance of our model.
Yingru Liu, Dongliang Xie, Xin Wang 0001
ICME1