EDBT 2026 Demo / reviewers in the wild / expert
Jinglin Liu
dblp:138/6898
· DBLP profile ↗
50ranked-venue papers
6as first author
41since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 21 · 4 first-author · 18 since 2021Systems, architecture and hardware · 18 · 1 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 4 first-author · 15 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Impact of PM Bar Position and Dimension on the Electromagnetic Characteristics of Enhanced Stator PM Hybrid Stepping MotorabstractHybrid stepping motors (HSMs) have been widely used in the robotics field due to their high torque and resolution advantages over other types of stepping motors. Unfortunately, in some applications where space is limited, the motor is often expected to have higher torque density to meet design requirements. Therefore, this paper presents an enhanced stator permanent magnet HSM (ESPMHSM), which possesses PM bars in the slots of stator poles. The implementation of PM bars changes the air gap magnetic field, which affects the other main electromagnetic characteristics of the motor. Consequently, it is important to qualitatively and quantitatively analyze the position and dimensions of PM bars on the electromagnetic characteristics. Firstly, the structure of ESPMHSM is introduced and the different positions of PM bars in slots are descried. Secondly, the motor is designed and the structural parameters are determined. Thirdly, the electromagnetic characteristics including the distribution of magnetic flux and magnetic density, air-gap flux density, back electromotive force, detent torque and so on are analyzed based on the finite element method. The results show that the closer the PM bars are the air gap, the better the electromagnetic characteristics of the motor. Xiaobao Chai, Jinglin Liu, Maixia Shang, Zehua Yang |
IECON | 2 |
| 2025 | GAF-CNN-LSTM Hybrid Model for Power Load Forecasting in Non-Stationary GridsabstractWith the increasing penetration of renewable energy in power systems, load forecasting faces dual challenges of modeling non-stationary fluctuations and spatiotemporally coupled features. This paper proposes a hybrid CNN-LSTM architecture based on Gramian Angular Field transformation, achieving high-precision load forecasting through spatiotemporal feature decoupling and multi-scale fusion strategies. First, the GAF algorithm encodes one-dimensional load sequences into two-dimensional polar coordinate image matrices, preserving dynamic evolution characteristics and phase information of raw data. Second, convolutional neural networks extract local mutation gradient patterns combined with an asymmetric pooling strategy to compress redundant features. Finally, long short-term memory networks model long-term periodic dependencies and optimize predictions through residual correction mechanisms. Experimental results on the GEFCom2014 dataset demonstrate that the proposed model reduces RMSE by 51.6%, improves R2to 0.981, and decreases maximum error by 48.9% compared to the baseline LSTM. The study proves that GAF’s polar coordinate encoding effectively enhances characterization capabilities for abrupt load fluctuations, while the collaborative optimization mechanism of CNN-LSTM significantly improves the model’s adaptability to non-stationary fluctuations caused by renewable energy integration, providing reliable technical support for real-time smart grid dispatch. Qianyuan Dong, Jinglin Liu |
IECON | 2 |
| 2025 | Temperature Rise and Loss Study for Wet Rotor Cooling Method of High-speed Permanent Magnet Synchronous Motors for Robot JointsabstractHigh-speed permanent magnet synchronous motors (PMSMs) are widely used in the robotics industry due to their advantages such as high power density, high efficiency, and compact size. However, PMSMs applied in robotic joints generate significant heat at high speeds, making effective cooling techniques essential to control temperature rise, ensure motor stability and service life, and maintain the robot's operational capability in specific applications. This study focuses on the wet rotor cooling technique for high-speed PMSMs, investigating the temperature rise calculation method based on the static rotor equivalent dynamic rotor model, as well as the effect of coolant injection on rotor wall shear friction and power loss. Simulation results confirm the rationality of the proposed calculation model and verify the effectiveness of the cooling strategy. Jinglin Liu, Zehua Yang |
IECON | 2 |
| 2025 | A control strategy for the synchronous lifting of uneven objects by joint motors of double mechanical arms based on SMESOabstractAiming at the phenomenon of joint motor position asynchrony that occurs when dual robotic arms cooperatively handle objects with uneven mass distribution, this paper proposes a position cross-coupling cooperative control strategy based on a sliding mode extended state observer (SMESO) to ensure smooth handling of objects by the dual arms. Using the cross-coupling control strategy, the real-time angle difference between the two motors is fed back to the controller to adjust the speed command. In order to speed up the speed adjustment, this paper designs SMESO for load torque difference estimation and feed-forward compensation of the estimated value to the system. The final realization of dual robotic arm joint motor speed and position are synchronized. Simulation results show that the proposed control strategy in the cooperative handling of uneven mass distribution of objects under the working conditions, position cross-coupling control to achieve rapid synchronization of the two motors position, SMESO significantly improves the position synchronization dynamic performance of the two motors and reduces the system dynamic adjustment time by 88%. In addition, the scheme does not need to rely on the robotic arm dynamics model, providing a model-free solution for dual robotic arm joint motor position synchronization control. Jinglin Liu, Ruizhi Guan, Xiaotao Li, Zhenqi Bai |
IECON | 2 |
| 2025 | Strategy for Parameter Tuning of PI Controller for PMSMs Based on Improved Particle AlgorithmabstractThis paper proposes a parameter tuning strategy for the PI controller of permanent magnet synchronous motors (PMSMs) based on an improved particle swarm optimization (IPSO) algorithm. To address the limitations of traditional PI controllers with fixed parameters, which struggle to adapt to external disturbances and dynamic environmental changes, the study enhances the global search capability and convergence speed of the PSO algorithm by incorporating adaptive mutation and elite learning strategies. Experiments conducted on the MATLAB/Simulink platform compared the performance of electrical parameter self-tuning, the standard PSO algorithm, and the proposed IPSO algorithm. The results demonstrate that the PI controller tuned by the IPSO algorithm significantly outperforms conventional methods in terms of dynamic speed response, overshoot suppression (only 4.73%), and resilience to sudden load changes (settling time of 0.18 seconds). Additionally, the algorithm exhibits faster convergence and effectively mitigates the issue of local optima in time-varying multi-parameter optimization. This research provides an adaptive and robust parameter tuning solution for high-performance control of PMSMs. Jinglin Liu, Xinran Shi, Maixia Shang, Xiaobao Chai |
IECON | 2 |
| 2025 | Parameter Identification Based on Deep Deterministic Policy Gradient for PMSMabstractIn this paper, a deep deterministic policy gradient (DDPG) based parameter identification method for permanent magnet synchronous motor (PMSM) is proposed for the simultaneous estimation of stator resistance, inductance, and magnetic chain. The method's efficacy is demonstrated by its ability to accurately identify motor parameters while exhibiting good robustness and convergence. This is achieved by designing a continuous action space and a multi-objective reward function in conjunction with an Actor-Critic network structure. Zehua Yang, Jinglin Liu |
IECON | 2 |
| 2024 | AudioGPT: Understanding and Generating Speech, Music, Sound, and Talking HeadabstractLarge language models (LLMs) have exhibited remarkable capabilities across a variety of domains and tasks, challenging our understanding of learning and cognition. Despite the recent success, current LLMs are not capable of processing complex audio information or conducting spoken conversations (like Siri or Alexa). In this work, we propose a multi-modal AI system named AudioGPT, which complements LLMs (i.e., ChatGPT) with 1) foundation models to process complex audio information and solve numerous understanding and generation tasks; and 2) the input/output interface (ASR, TTS) to support spoken dialogue. With an increasing demand to evaluate multi-modal LLMs of human intention understanding and cooperation with foundation models, we outline the principles and processes and test AudioGPT in terms of consistency, capability, and robustness. Experimental results demonstrate the capabilities of AudioGPT in solving 16 AI tasks with speech, music, sound, and talking head understanding and generation in multi-round dialogues, which empower humans to create rich and diverse audio content with unprecedented ease. Code can be found in https://github.com/AIGC-Audio/AudioGPT Rongjie Huang 0001, Dongchao Yang, Jiatong Shi, Xuankai Chang, Zhenhui Ye, Yuning Wu 0001, Zhiqing Hong, Jiawei Huang 0008, Jinglin Liu, Yi Ren 0006, Yuexian Zou, Zhou Zhao 0001, Shinji Watanabe 0001 |
AAAI | 10 |
| 2024 | Mega-TTS 2: Boosting Prompting Mechanisms for Zero-Shot Speech SynthesisabstractZero-shot text-to-speech (TTS) aims to synthesize voices with unseen speech prompts, which significantly reduces the data and computation requirements for voice cloning by skipping the fine-tuning process. However, the prompting mechanisms of zero-shot TTS still face challenges in the following aspects: 1) previous works of zero-shot TTS are typically trained with single-sentence prompts, which significantly restricts their performance when the data is relatively sufficient during the inference stage. 2) The prosodic information in prompts is highly coupled with timbre, making it untransferable to each other.
This paper introduces Mega-TTS 2, a generic prompting mechanism for zero-shot TTS, to tackle the aforementioned challenges. Specifically, we design a powerful acoustic autoencoder that separately encodes the prosody and timbre information into the compressed latent space while providing high-quality reconstructions. Then, we propose a multi-reference timbre encoder and a prosody latent language model (P-LLM) to extract useful information from multi-sentence prompts. We further leverage the probabilities derived from multiple P-LLM outputs to produce transferable and controllable prosody.
Experimental results demonstrate that Mega-TTS 2 could not only synthesize identity-preserving speech with a short prompt of an unseen speaker from arbitrary sources but consistently outperform the fine-tuning method when the volume of data ranges from 10 seconds to 5 minutes. Furthermore, our method enables to transfer various speaking styles to the target timbre in a fine-grained and controlled manner. Audio samples can be found in https://boostprompt.github.io/boostprompt/. Ziyue Jiang 0001, Jinglin Liu, Yi Ren 0006, Jinzheng He, Zhenhui Ye, Shengpeng Ji, Qian Yang 0006, Chen Zhang 0020, Pengfei Wei 0001, Xiang Yin 0006, Zejun Ma 0001, Zhou Zhao 0001 |
ICLR | 2 |
| 2024 | Real3D-Portrait: One-shot Realistic 3D Talking Portrait SynthesisabstractOne-shot 3D talking portrait generation aims to reconstruct a 3D avatar from an unseen image, and then animate it with a reference video or audio to generate a talking portrait video. The existing methods fail to simultaneously achieve the goals of accurate 3D avatar reconstruction and stable talking face animation. Besides, while the existing works mainly focus on synthesizing the head part, it is also vital to generate natural torso and background segments to obtain a realistic talking portrait video. To address these limitations, we present Real3D-Potrait, a framework that (1) improves the one-shot 3D reconstruction power with a large image-to-plane model that distills 3D prior knowledge from a 3D face generative model; (2) facilitates accurate motion-conditioned animation with an efficient motion adapter; (3) synthesizes realistic video with natural torso movement and switchable background using a head-torso-background super-resolution model; and (4) supports one-shot audio-driven talking face generation with a generalizable audio-to-motion model. Extensive experiments show that Real3D-Portrait generalizes well to unseen identities and generates more realistic talking portrait videos compared to previous methods. Video samples are available at https://real3dportrait.github.io. Zhenhui Ye, Tianyun Zhong, Yi Ren 0006, Jiaqi Yang 0008, Weichuang Li, Jiawei Huang 0008, Ziyue Jiang 0001, Jinzheng He, Rongjie Huang 0001, Jinglin Liu, Chen Zhang 0020, Xiang Yin 0006, Zejun Ma 0001, Zhou Zhao 0001 |
ICLR | 10 |
| 2024 | MimicTalk: Mimicking a personalized and expressive 3D talking face in minutesabstractTalking face generation (TFG) aims to animate a target identity's face to create realistic talking videos. Personalized TFG is a variant that emphasizes the perceptual identity similarity of the synthesized result (from the perspective of appearance and talking style). While previous works typically solve this problem by learning an individual neural radiance field (NeRF) for each identity to implicitly store its static and dynamic information, we find it inefficient and non-generalized due to the per-identity-per-training framework and the limited training data. To this end, we propose MimicTalk, the first attempt that exploits the rich knowledge from a NeRF-based person-agnostic generic model for improving the efficiency and robustness of personalized TFG. To be specific, (1) we first come up with a person-agnostic 3D TFG model as the base model and propose to adapt it into a specific identity; (2) we propose a static-dynamic-hybrid adaptation pipeline to help the model learn the personalized static appearance and facial dynamic features; (3) To generate the facial motion of the personalized talking style, we propose an in-context stylized audio-to-motion model that mimics the implicit talking style provided in the reference video without information loss by an explicit style representation. The adaptation process to an unseen identity can be performed in 15 minutes, which is 47 times faster than previous person-dependent methods. Experiments show that our MimicTalk surpasses previous baselines regarding video quality, efficiency, and expressiveness. Video samples are available at https://mimictalk.github.io . Zhenhui Ye, Tianyun Zhong, Yi Ren 0006, Ziyue Jiang 0001, Jiawei Huang 0008, Rongjie Huang 0001, Jinglin Liu, Jinzheng He, Chen Zhang 0020, Zehan Wang 0001, Xize Cheng, Xiang Yin 0006, Zhou Zhao 0001 |
NeurIPS | 7 |
| 2023 | AV-TranSpeech: Audio-Visual Robust Speech-to-Speech TranslationabstractRongjie Huang, Huadai Liu, Xize Cheng, Yi Ren, Linjun Li, Zhenhui Ye, Jinzheng He, Lichao Zhang, Jinglin Liu, Xiang Yin, Zhou Zhao. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Rongjie Huang 0001, Huadai Liu, Xize Cheng, Yi Ren 0006, Linjun Li, Zhenhui Ye, Jinzheng He, Jinglin Liu, Xiang Yin 0006, Zhou Zhao 0001 |
ACL (1) | 9 |
| 2023 | CLAPSpeech: Learning Prosody from Text Context with Contrastive Language-Audio Pre-TrainingabstractZhenhui Ye, Rongjie Huang, Yi Ren, Ziyue Jiang, Jinglin Liu, Jinzheng He, Xiang Yin, Zhou Zhao. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Zhenhui Ye, Rongjie Huang 0001, Yi Ren 0006, Ziyue Jiang 0001, Jinglin Liu, Jinzheng He, Xiang Yin 0006, Zhou Zhao 0001 |
ACL (1) | 5 |
| 2023 | VarietySound: Timbre-Controllable Video to Sound Generation Via Unsupervised Information DisentanglementabstractVideo-to-sound generation aims to generate realistic and natural sound given a video input. However, previous video-to-sound generation methods can only generate a random or average timbre without any controls of the generated sound timbre, leading to the problem that people cannot obtain the desired timbre under these methods sometimes. In this paper, we propose the task of generating sound with a specific timbre given a silent video input and a reference audio sample. To solve this task, we first use three encoders to disentangle each target sound audio into temporal, acoustic, and background information respectively, then we use a decoder to reconstruct the audio given these disentangled representations. To make the generated result achieve better quality and temporal alignment, we also adopt a mel discriminator and a temporal discriminator for the adversarial training. Our experimental results on the VAS dataset demonstrate that our method can generate high-quality audio samples with good synchronization with events in video and high timbre similarity with the reference audio. Our demos have been published on https://conferencedemos.github.io/icassp23/. Chenye Cui, Zhou Zhao 0001, Yi Ren 0006, Jinglin Liu, Rongjie Huang 0001, Feiyang Chen 0001, Zhefeng Wang 0001, Baoxing Huai, Fei Wu 0001 |
ICASSP | 4 |
| 2023 | Overview of the ICASSP 2023 General Meeting Understanding and Generation Challenge (MUG)abstractICASSP2023 General Meeting Understanding and Generation Challenge (MUG) focuses on prompting a wide range of spoken language processing (SLP) research on meeting transcripts, as SLP applications are critical to improve users’ efficiency in grasping important information in meetings. MUG includes five tracks, including topic segmentation, topic-level and session-level extractive summarization, topic title generation, keyphrase extraction, and action item detection. To facilitate MUG, we construct and release a large-scale meeting dataset, the AliMeeting4MUG Corpus. We review the dataset, track settings and baselines, and summarize the challenge results and major techniques used in the submissions. Chong Deng, Jiaqing Liu, Qian Chen 0003, Wen Wang 0001, Zhijie Yan, Jinglin Liu, Yi Ren 0006, Zhou Zhao 0001 |
ICASSP | 8 |
| 2023 | MUG: A General Meeting Understanding and Generation BenchmarkabstractListening to long video/audio recordings from video conferencing and online courses for acquiring information is extremely inefficient. Even after ASR systems transcribe recordings into long-form spoken language documents, reading ASR transcripts only partly speeds up seeking information. It has been observed that a range of NLP applications, such as keyphrase extraction, topic segmentation, and summarization, significantly improve users’ efficiency in grasping important information. The meeting scenario is among the most valuable scenarios for deploying these spoken language processing (SLP) capabilities. However, the lack of large-scale public meeting datasets annotated for these SLP tasks severely hinders their advancement. To prompt SLP advancement, we establish a large-scale general Meeting Understanding and Generation Benchmark (MUG) to benchmark the performance of a wide range of SLP tasks, including topic segmentation, topic-level and session-level extractive summarization and topic title generation, keyphrase extraction, and action item detection. To facilitate the MUG benchmark, we construct and release a large-scale meeting dataset for comprehensive long-form SLP development, the AliMeeting4MUG Corpus, which consists of 654 recorded Mandarin meeting sessions with diverse topic coverage, with manual annotations for SLP tasks on manual transcripts of meeting recordings. To the best of our knowledge, the AliMeeting4MUG Corpus is so far the largest meeting corpus in scale and facilitates most SLP tasks. In this paper, we provide a detailed introduction of this corpus, SLP tasks and evaluation methods, baseline systems and their performance1. Chong Deng, Jiaqing Liu, Qian Chen 0003, Wen Wang 0001, Zhijie Yan, Jinglin Liu, Yi Ren 0006, Zhou Zhao 0001 |
ICASSP | 8 |
| 2023 | TranSpeech: Speech-to-Speech Translation With Bilateral Perturbation
Rongjie Huang 0001, Jinglin Liu, Huadai Liu, Yi Ren 0006, Jinzheng He, Zhou Zhao 0001 |
ICLR | 2 |
| 2023 | GeneFace: Generalized and High-Fidelity Audio-Driven 3D Talking Face Synthesis
Zhenhui Ye, Ziyue Jiang 0001, Yi Ren 0006, Jinglin Liu, Jinzheng He, Zhou Zhao 0001 |
ICLR | 4 |
| 2023 | Make-An-Audio: Text-To-Audio Generation with Prompt-Enhanced Diffusion ModelsabstractLarge-scale multimodal generative modeling has created milestones in text-to-image and text-to-video generation. Its application to audio still lags behind for two main reasons: the lack of large-scale datasets with high-quality text-audio pairs, and the complexity of modeling long continuous audio data. In this work, we propose Make-An-Audio with a prompt-enhanced diffusion model that addresses these gaps by 1) introducing pseudo prompt enhancement with a distill-then-reprogram approach, it alleviates data scarcity with orders of magnitude concept compositions by using language-free audios; 2) leveraging spectrogram autoencoder to predict the self-supervised audio representation instead of waveforms. Together with robust contrastive language-audio pretraining (CLAP) representations, Make-An-Audio achieves state-of-the-art results in both objective and subjective benchmark evaluation. Moreover, we present its controllability and generalization for X-to-Audio with "No Modality Left Behind", for the first time unlocking the ability to generate high-definition, high-fidelity audios given a user-defined modality input. Audio samples are available at https://Make-An-Audio.github.io Rongjie Huang 0001, Jiawei Huang 0008, Dongchao Yang, Yi Ren 0006, Luping Liu, Zhenhui Ye, Jinglin Liu, Xiang Yin 0006, Zhou Zhao 0001 |
ICML | 8 |
| 2023 | Elimination of Digital Delay Effect on Rotor Position for Two-Step Finite Control Set MPCC Used in PMSMsabstractThis paper proposes a novel concept to eliminate the digital delay effect on rotor position for the traditional two-step finite control set model predictive current control (FCS-MPCC) method used in the PMSMs. Specifically, that the digital delay affects the position, and further the prediction accuracy is discussed firstly, posing the necessity of compensating the position used for control. Then, a speed-prediction-based linear compensation method which takes the speed shift trend into account is proposed to eliminate the digital delay impact on position, which is conducive to the control performance. Finally, simulation is conducted on a three-phase PMSM to validate the proposed strategy in comparison with the traditional two-step FCS-MPCC. Jinglin Liu, Chao Gong 0001, Lefei Ge |
IECON | 1 |
| 2023 | A MPTC Optimization Scheme for PMSM Based on Flux Error Tracking and PSOabstractIn this paper, an improved model prediction scheme of permanent magnet synchronous motor(PMSM) is proposed, which simplifies the vector selection process and improves the model predictive control(MPC) cost function. The weighting factor of the cost function is dynamically adjusted by the particle swarm optimization (PSO) algorithm to effectively suppress torque ripple. Zhiman Lu, Jinglin Liu, Minglang Xiao |
IECON | 2 |
| 2023 | Multi-Objective Optimization of High Power Density Motor Based on Metamodel of Optimal PrognosisabstractBecause of multi-parameter, nonlinearity and strong coupling, one of the main difficulties in motor design is multi-objective optimization. In this paper, the high power density high lift motor (HLM) of distributed all-electric aircraft is taken as the research object. The multi-objective optimization process based on Metamodel of optimal prognosis (MOP) is given; The parameter variables sensitivity coefficient, the coefficient of prognosis (COP) of the model and the optimization results are analyzed and summarized. The optimization method has shorter calculation time, higher efficiency, and is more suitable for multi-parameter nonlinear objective optimization. Maixia Shang, Jinglin Liu |
IECON | 2 |
| 2023 | Hybrid Position Sensorless Control Based on Estimation Position Error Switching for PMSM in Full Speed RangeabstractIn permanent magnet synchronous motor (PMSM) full speed range position sensorless control, the conventional hybrid control method usually determines the switching speed point as 10%-30% of the rated speed in the switching process empirically, which fails to effectively reduce the pulsations of position and speed. In this paper, a method is proposed to determine the switching speed point based on the scalarized estimated position error to achieve smooth switching and reduce the pulsations of position and speed during switching. The method is divided into three ranges: In zero-low and medium-high speed ranges, the estimated speed and position are obtained using the high-frequency injection method and the sliding mode observer approach, respectively. In transition range, the switching speed is determined when the position errors of the two methods are similar. After weighing the two errors, the estimated position and speed are acquired through a phase-locked loop. Finally, the effectiveness of the proposed method is verified in MATLAB/Simulink based on the built-in PMSM. Xinran Shi, Jinglin Liu, Jiamin Xu |
IECON | 2 |
| 2023 | Proportional Resonant Filtering For Improved SMO With Optimized Critical Saturation Switching FunctionabstractThis paper proposes a sensorless speed control strategy for a permanent-magnet synchronous motor (PMSM) based on proportional resonant filtering and an improved sliding-mode observer (SMO) with an optimized critical saturation switching function. This strategy suppresses pulsation and decreases the delay to improve the accuracy of position estimation. First, an optimized critical saturation switching function was designed to output a sine wave with a boundary layer fixed at 1. This strategy kept the convergence speed of SMO constant and weakened the pulsation. However, the phase delay in the estimated back-EMF caused by introducing a low-pass filter (LPF) persists in the system. Therefore, this paper proposes a proportional resonant filter (PRF) to address this problem. The PRF, without phase delay, can extract the fundamental component and filter out the harmonic components in the estimated back electromotive force (back-EMF). The simulation results verify the effectiveness and accuracy of the proposed strategy. Jiamin Xu, Jinglin Liu, Xinran Shi |
IECON | 2 |
| 2023 | Optimization Design of Torque Fluctuation Suppression for Surface-Mounted Permanent Magnet Synchronous MotorabstractPermanent magnet motors inevitably generate cogging torque, which in turn causes harm such as harmonics and the complexity of control system increased by current suppression methods, this paper proposes a method to effectively suppress the cogging torque by cutting Angle at both ends of the permanent magnet. Firstly, establish its generation mechanism and formula based on finite element analysis method. Secondly, the finite element simulation analysis of a 12-slot 8-pole SPMSM was carried out using ANSYS Maxwell. Finally, the simulation results show that the design can weaken the cogging torque of the prototype and reduce air gap magnetic induction intensity fluctuation and torque fluctuation. It is proved that this design optimizes the cogging torque based on cutting Angle of permanent magnet shape is correct and feasibility. Jinglin Liu, Yuyuan Yang, Xiaobao Chai, Lanlan Zheng |
IECON | 2 |
| 2023 | Direct Torque Control of Permanent Magnet Synchronous Motor for Reducing Torque RippleabstractTraditional direct torque control (DTC) system of permanent magnet synchronous motor has two main problems, one is large torque ripple, and the other is the switching frequency is not constant. To improve the problem of large torque ripple, this paper proposes a new simple torque controller to replace the traditional one, which could be used to reduce torque ripple. At the same time, the extended Kalman filter is used instead of the traditional voltage integral to calculate the flux, which reduces the error of the flux calculation process and makes the flux estimation more accurate. Based on the least variance estimation theory, the extended Kalman filter provides a solution to accurately estimate the state of nonlinear systems and has good dynamic performance, high anti-interference, and accurate estimation ability. Finally, the effectiveness of the proposed method is verified in MATLAB/Simulink based on the built-in PMSM. Lanlan Zheng, Jinglin Liu, Xinyue Jin |
IECON | 2 |
| 2023 | UniSinger: Unified End-to-End Singing Voice Synthesis With Cross-Modality Information MatchingabstractThough previous works have shown remarkable achievements in singing voice generation, most existing models focus on one specific application and there is a lack of unified singing voice synthesis models. In addition to low relevance among tasks, different input modalities are one of the most intractable hindrances. Current methods suffer from information confusion and they can not perform precise control. In this work, we propose UniSinger, a unified end-to-end singing voice synthesizer, which integrates three abilities related to singing voice generation: singing voice synthesis (SVS), singing voice conversion (SVC), and singing voice editing (SVE) into a single framework. Specifically, we perform representation disentanglement for controlling different attributes of the singing voice. We further propose a cross-modality information matching method to close the distribution gap between multi-modal inputs and achieve end-to-end training. The experiments conducted on the OpenSinger dataset demonstrate that UniSinger achieves state-of-the-art results in three applications. Further extensive experiments verify the capability of representation disentanglement and information matching, reflecting that UniSinger enjoys great superiority in sample quality, timbre similarity, and multi-task compatibility. Audio samples can be found in https://unisinger.github.io/Samples/. Zhiqing Hong, Chenye Cui, Rongjie Huang 0001, Jinglin Liu, Jinzheng He, Zhou Zhao 0001 |
ACM Multimedia | 5 |
| 2022 | Flow-Based Unconstrained Lip to Speech GenerationabstractUnconstrained lip-to-speech aims to generate corresponding speeches based on silent facial videos with no restriction to head pose or vocabulary. It is desirable to generate intelligible and natural speech with a fast speed in unconstrained settings. Currently, to handle the more complicated scenarios, most existing methods adopt the autoregressive architecture, which is optimized with the MSE loss. Although these methods have achieved promising performance, they are prone to bring issues including high inference latency and mel-spectrogram over-smoothness. To tackle these problems, we propose a novel flow-based non-autoregressive lip-to-speech model (GlowLTS) to break autoregressive constraints and achieve faster inference. Concretely, we adopt a flow-based decoder which is optimized by maximizing the likelihood of the training data and is capable of more natural and fast speech generation. Moreover, we devise a condition module to improve the intelligibility of generated speech. We demonstrate the superiority of our proposed method through objective and subjective evaluation on Lip2Wav-Chemistry-Lectures and Lip2Wav-Chess-Analysis datasets. Our demo video can be found at https://glowlts.github.io/. Jinzheng He, Zhou Zhao 0001, Yi Ren 0006, Jinglin Liu, Baoxing Huai, Nicholas Jing Yuan |
AAAI | 4 |
| 2022 | DiffSinger: Singing Voice Synthesis via Shallow Diffusion MechanismabstractSinging voice synthesis (SVS) systems are built to synthesize high-quality and expressive singing voice, in which the acoustic model generates the acoustic features (e.g., mel-spectrogram) given a music score. Previous singing acoustic models adopt a simple loss (e.g., L1 and L2) or generative adversarial network (GAN) to reconstruct the acoustic features, while they suffer from over-smoothing and unstable training issues respectively, which hinder the naturalness of synthesized singing. In this work, we propose DiffSinger, an acoustic model for SVS based on the diffusion probabilistic model. DiffSinger is a parameterized Markov chain that iteratively converts the noise into mel-spectrogram conditioned on the music score. By implicitly optimizing variational bound, DiffSinger can be stably trained and generate realistic outputs. To further improve the voice quality and speed up inference, we introduce a shallow diffusion mechanism to make better use of the prior knowledge learned by the simple loss. Specifically, DiffSinger starts generation at a shallow step smaller than the total number of diffusion steps, according to the intersection of the diffusion trajectories of the ground-truth mel-spectrogram and the one predicted by a simple mel-spectrogram decoder. Besides, we propose boundary prediction methods to locate the intersection and determine the shallow step adaptively. The evaluations conducted on a Chinese singing dataset demonstrate that DiffSinger outperforms state-of-the-art SVS work. Extensional experiments also prove the generalization of our methods on text-to-speech task (DiffSpeech). Audio samples: https://diffsinger.github.io. Codes: https://github.com/MoonInTheRiver/DiffSinger. Jinglin Liu, Chengxi Li 0002, Yi Ren 0006, Feiyang Chen 0001, Zhou Zhao 0001 |
AAAI | 1 |
| 2022 | Parallel and High-Fidelity Text-to-Lip GenerationabstractAs a key component of talking face generation, lip movements generation determines the naturalness and coherence of the generated talking face video. Prior literature mainly focuses on speech-to-lip generation while there is a paucity in text-to-lip (T2L) generation. T2L is a challenging task and existing end-to-end works depend on the attention mechanism and autoregressive (AR) decoding manner. However, the AR decoding manner generates current lip frame conditioned on frames generated previously, which inherently hinders the inference speed, and also has a detrimental effect on the quality of generated lip frames due to error propagation. This encourages the research of parallel T2L generation. In this work, we propose a parallel decoding model for fast and high-fidelity text-to-lip generation (ParaLip). Specifically, we predict the duration of the encoded linguistic features and model the target lip frames conditioned on the encoded linguistic features with their duration in a non-autoregressive manner. Furthermore, we incorporate the structural similarity index loss and adversarial learning to improve perceptual quality of generated lip frames and alleviate the blurry prediction problem. Extensive experiments conducted on GRID and TCD-TIMIT datasets demonstrate the superiority of proposed methods. Jinglin Liu, Yi Ren 0006, Wencan Huang, Baoxing Huai, Nicholas Jing Yuan, Zhou Zhao 0001 |
AAAI | 1 |
| 2022 | Learning the Beauty in Songs: Neural Singing Voice BeautifierabstractWe are interested in a novel task, singing voice beautification (SVB).Given the singing voice of an amateur singer, SVB aims to improve the intonation and vocal tone of the voice, while keeping the content and vocal timbre.Current automatic pitch correction techniques are immature, and most of them are restricted to intonation but ignore the overall aesthetic quality.Hence, we introduce Neural Singing Voice Beautifier (NSVB), the first generative model to solve the SVB task, which adopts a conditional variational autoencoder as the backbone and learns the latent representations of vocal tone.In NSVB, we propose a novel time-warping approach for pitch correction: Shape-Aware Dynamic Time Warping (SADTW), which ameliorates the robustness of existing time-warping approaches, to synchronize the amateur recording with the template pitch curve.Furthermore, we propose a latent-mapping algorithm in the latent space to convert the amateur vocal tone to the professional one.To achieve this, we also propose a new dataset containing parallel singing recordings of both amateur and professional versions.Extensive experiments on both Chinese and English songs demonstrate the effectiveness of our methods in terms of both objective and subjective metrics. Jinglin Liu, Chengxi Li 0002, Yi Ren 0006, Zhou Zhao 0001 |
ACL (1) | 1 |
| 2022 | SingGAN: Generative Adversarial Network For High-Fidelity Singing Voice GenerationabstractDeep generative models have achieved significant progress in speech synthesis to date, while high-fidelity singing voice synthesis is still an open problem for its long continuous pronunciation, rich high-frequency parts, and strong expressiveness. Existing neural vocoders designed for text-to-speech cannot directly be applied to singing voice synthesis because they result in glitches and poor high-frequency reconstruction. In this work, we propose SingGAN, a generative adversarial network designed for high-fidelity singing voice synthesis. Specifically, 1) to alleviate the glitch problem in the generated samples, we propose source excitation with the adaptive feature learning filters to expand the receptive field patterns and stabilize long continuous signal generation; and 2) SingGAN introduces global and local discriminators at different scales to enrich low-frequency details and promote high-frequency reconstruction; and 3) To improve the training efficiency, SingGAN includes auxiliary spectrogram losses and sub-band feature matching penalty loss. To the best of our knowledge, SingGAN is the first work designed toward high-fidelity singing voice vocoding. Our evaluation of SingGAN demonstrates the state-of-the-art results with higher-quality (MOS 4.05) samples. Also, SingGAN enables a sample speed of 50x faster than real-time on a single NVIDIA 2080Ti GPU. We further show that SingGAN generalizes well to the mel-spectrogram inversion of unseen singers, and the end-to-end singing voice synthesis system SingGAN-SVS enjoys a two-stage pipeline to transform the music scores into expressive singing voices. Rongjie Huang 0001, Chenye Cui, Feiyang Chen 0001, Yi Ren 0006, Jinglin Liu, Zhou Zhao 0001, Baoxing Huai, Zhefeng Wang 0001 |
ACM Multimedia | 5 |
| 2022 | ProDiff: Progressive Fast Diffusion Model for High-Quality Text-to-SpeechabstractDenoising diffusion probabilistic models (DDPMs) have recently achieved leading performances in many generative tasks. However, the inherited iterative sampling process costs hinder their applications to text-to-speech deployment. Through the preliminary study on diffusion model parameterization, we find that previous gradient-based TTS models require hundreds or thousands of iterations to guarantee high sample quality, which poses a challenge for accelerating sampling. In this work, we propose ProDiff, on progressive fast diffusion model for high-quality text-to-speech. Unlike previous work estimating the gradient for data density, ProDiff parameterizes the denoising model by directly predicting clean data to avoid distinct quality degradation in accelerating sampling. To tackle the model convergence challenge with decreased diffusion iterations, ProDiff reduces the data variance in the target site via knowledge distillation. Specifically, the denoising model uses the generated mel-spectrogram from an N-step DDIM teacher as the training target and distills the behavior into a new model with N/2 steps. As such, it allows the TTS model to make sharp predictions and further reduces the sampling time by orders of magnitude. Our evaluation demonstrates that ProDiff needs only 2 iterations to synthesize high-fidelity mel-spectrograms, while it maintains sample quality and diversity competitive with state-of-the-art models using hundreds of steps. ProDiff enables a sampling speed of 24x faster than real-time on a single NVIDIA 2080Ti GPU, making diffusion models practically applicable to text-to-speech synthesis deployment for the first time. Our extensive ablation studies demonstrate that each design in ProDiff is effective, and we further show that ProDiff can be easily extended to the multi-speaker setting. Rongjie Huang 0001, Zhou Zhao 0001, Huadai Liu, Jinglin Liu, Chenye Cui, Yi Ren 0006 |
ACM Multimedia | 4 |
| 2022 | GenerSpeech: Towards Style Transfer for Generalizable Out-Of-Domain Text-to-SpeechabstractStyle transfer for out-of-domain (OOD) speech synthesis aims to generate speech samples with unseen style (e.g., speaker identity, emotion, and prosody) derived from an acoustic reference, while facing the following challenges: 1) The highly dynamic style features in expressive voice are difficult to model and transfer; and 2) the TTS models should be robust enough to handle diverse OOD conditions that differ from the source data. This paper proposes GenerSpeech, a text-to-speech model towards high-fidelity zero-shot style transfer of OOD custom voice. GenerSpeech decomposes the speech variation into the style-agnostic and style-specific parts by introducing two components: 1) a multi-level style adaptor to efficiently model a large range of style conditions, including global speaker and emotion characteristics, and the local (utterance, phoneme, and word-level) fine-grained prosodic representations; and 2) a generalizable content adaptor with Mix-Style Layer Normalization to eliminate style information in the linguistic content representation and thus improve model generalization. Our evaluations on zero-shot style transfer demonstrate that GenerSpeech surpasses the state-of-the-art models in terms of audio quality and style similarity. The extension studies to adaptive style transfer further show that GenerSpeech performs robustly in the few-shot data setting. Audio samples are available at \url{https://GenerSpeech.github.io/}. Rongjie Huang 0001, Yi Ren 0006, Jinglin Liu, Chenye Cui, Zhou Zhao 0001 |
NeurIPS | 3 |
| 2022 | Dict-TTS: Learning to Pronounce with Prior Dictionary Knowledge for Text-to-SpeechabstractPolyphone disambiguation aims to capture accurate pronunciation knowledge from natural text sequences for reliable Text-to-speech (TTS) systems. However, previous approaches require substantial annotated training data and additional efforts from language experts, making it difficult to extend high-quality neural TTS systems to out-of-domain daily conversations and countless languages worldwide. This paper tackles the polyphone disambiguation problem from a concise and novel perspective: we propose Dict-TTS, a semantic-aware generative text-to-speech model with an online website dictionary (the existing prior information in the natural language). Specifically, we design a semantics-to-pronunciation attention (S2PA) module to match the semantic patterns between the input text sequence and the prior semantics in the dictionary and obtain the corresponding pronunciations; The S2PA module can be easily trained with the end-to-end TTS model without any annotated phoneme labels. Experimental results in three languages show that our model outperforms several strong baseline models in terms of pronunciation accuracy and improves the prosody modeling of TTS systems. Further extensive analyses demonstrate that each design in Dict-TTS is effective. The code is available at https://github.com/Zain-Jiang/Dict-TTS. Ziyue Jiang 0001, Su Zhe, Zhou Zhao 0001, Qian Yang 0006, Yi Ren 0006, Jinglin Liu, Zhenhui Ye |
NeurIPS | 6 |
| 2022 | M4Singer: A Multi-Style, Multi-Singer and Musical Score Provided Mandarin Singing CorpusabstractThe lack of publicly available high-quality and accurately labeled datasets has long been a major bottleneck for singing voice synthesis (SVS). To tackle this problem, we present M4Singer, a free-to-use Multi-style, Multi-singer Mandarin singing collection with elaborately annotated Musical scores as well as its benchmarks. Specifically, 1) we construct and release a large high-quality Chinese singing voice corpus, which is recorded by 20 professional singers, covering 700 Chinese pop songs as well as all the four SATB types (i.e., soprano, alto, tenor, and bass); 2) we take extensive efforts to manually compose the musical scores for each recorded song, which are necessary to the study of the prosody modeling for SVS. 3) To facilitate the use and demonstrate the quality of M4Singer, we conduct four different benchmark experiments: score-based SVS, controllable singing voice (CSV), singing voice conversion (SVC) and automatic music transcription (AMT). Ruiqi Li 0002, Shoutong Wang, Liqun Deng, Jinglin Liu, Yi Ren 0006, Jinzheng He, Rongjie Huang 0001, Jieming Zhu, Zhou Zhao 0001 |
NeurIPS | 5 |
| 2021 | Denoispeech: Denoising Text to Speech with Frame-Level Noise ModelingabstractWhile neural-based text to speech (TTS) models can synthesize natural and intelligible voice, they usually require high-quality speech data, which is costly to collect. In many scenarios, only noisy speech of a target speaker is available, which presents challenges for TTS model training for this speaker. Previous works usually address the challenge using two methods: 1) training the TTS model using the speech denoised with an enhancement model; 2) taking a single noise embedding as input when training with noisy speech. However, they usually cannot handle speech with real-world complicated noise such as those with high variations along time. In this paper, we develop DenoiSpeech, a TTS system that can synthesize clean speech for a speaker with noisy speech data. In DenoiSpeech, we handle real-world noisy speech by modeling the fine-grained frame-level noise with a noise condition module, which is jointly trained with the TTS model. Experimental results on real-world data show that DenoiSpeech outperforms the previous two methods by 0.31 and 0.66 MOS respectively.1 Chen Zhang 0020, Yi Ren 0006, Xu Tan 0003, Jinglin Liu, Tao Qin 0001, Sheng Zhao 0002, Tie-Yan Liu |
ICASSP | 4 |
| 2021 | EMOVIE: A Mandarin Emotion Speech Dataset with a Simple Emotional Text-to-Speech ModelabstractRecently, there has been an increasing interest in neural speech synthesis. While the deep neural network achieves the state-of-the-art result in text-to-speech (TTS) tasks, how to generate a more emotional and more expressive speech is becoming a new challenge to researchers due to the scarcity of high-quality emotion speech dataset and the lack of advanced emotional TTS model. In this paper, we first briefly introduce and publicly release a Mandarin emotion speech dataset including 9,724 samples with audio files and its emotion human-labeled annotation. After that, we propose a simple but efficient architecture for emotional speech synthesis called EMSpeech. Unlike those models which need additional reference audio as input, our model could predict emotion labels just from the input text and generate more expressive speech conditioned on the emotion embedding. In the experiment phase, we first validate the effectiveness of our dataset by an emotion classification task. Then we train our model on the proposed dataset and conduct a series of subjective evaluations. Finally, by showing a comparable performance in the emotional speech synthesis task, we successfully demonstrate the ability of the proposed model. Chenye Cui, Yi Ren 0006, Jinglin Liu, Feiyang Chen 0001, Rongjie Huang 0001, Zhou Zhao 0001 |
Interspeech | 3 |
| 2021 | Multi-Singer: Fast Multi-Singer Singing Voice Vocoder With A Large-Scale CorpusabstractHigh-fidelity multi-singer singing voice synthesis is challenging for neural vocoder due to the singing voice data shortage, limited singer generalization, and large computational cost. Existing open corpora could not meet requirements for high-fidelity singing voice synthesis because of the scale and quality weaknesses. Previous vocoders have difficulty in multi-singer modeling, and a distinct degradation emerges when conducting unseen singer singing voice generation. To accelerate singing voice researches in the community, we release a large-scale, multi-singer Chinese singing voice dataset OpenSinger. To tackle the difficulty in unseen singer modeling, we propose Multi-Singer, a fast multi-singer vocoder with generative adversarial networks. Specifically, 1) Multi-Singer uses a multi-band generator to speed up both training and inference procedure. 2) to capture and rebuild singer identity from the acoustic feature (i.e., mel-spectrogram), Multi-Singer adopts a singer conditional discriminator and conditional adversarial training objective. 3) to supervise the reconstruction of singer identity in the spectrum envelopes in frequency domain, we propose an auxiliary singer perceptual loss. The joint training approach effectively works in GANs for multi-singer voices modeling. Experimental results verify the effectiveness of OpenSinger and show that Multi-Singer improves unseen singer singing voices modeling in both speed and quality over previous methods. The further experiment proves that combined with FastSpeech 2 as the acoustic model, Multi-Singer achieves strong robustness in the multi-singer singing voice synthesis pipeline. Rongjie Huang 0001, Feiyang Chen 0001, Yi Ren 0006, Jinglin Liu, Chenye Cui, Zhou Zhao 0001 |
ACM Multimedia | 4 |
| 2021 | SimulLR: Simultaneous Lip Reading Transducer with Attention-Guided Adaptive MemoryabstractLip reading, aiming to recognize spoken sentences according to the given video of lip movements without relying on the audio stream, has attracted great interest due to its application in many scenarios. Although prior works that explore lip reading have obtained salient achievements, they are all trained in a non-simultaneous manner where the predictions are generated requiring access to the full video. To breakthrough this constraint, we study the task of simultaneous lip reading and devise SimulLR, a simultaneous lip Reading transducer with attention-guided adaptive memory from three aspects: (1) To address the challenge of monotonic alignments while considering the syntactic structure of the generated sentences under simultaneous setting, we build a transducer-based model and design several effective training strategies including CTC pre-training, model warm-up and curriculum learning to promote the training of the lip reading transducer. (2) To learn better spatio-temporal representations for simultaneous encoder, we construct a truncated 3D convolution and time-restricted self-attention layer to perform the frame-to-frame interaction within a video segment containing fixed number of frames. (3) The history information is always limited due to the storage in real-time scenarios, especially for massive video data. Therefore, we devise a novel attention-guided adaptive memory to organize semantic information of history segments and enhance the visual representations with acceptable computation-aware latency. The experiments show that the SimulLR achieves the translation speedup 9.10x compared with the state-of-the-art non-simultaneous methods, and also obtains competitive results, which indicates the effectiveness of our proposed methods. Zhijie Lin 0001, Zhou Zhao 0001, Haoyuan Li 0002, Jinglin Liu, Meng Zhang 0019, Xingshan Zeng, Xiaofei He 0001 |
ACM Multimedia | 4 |
| 2021 | SimulSLT: End-to-End Simultaneous Sign Language TranslationabstractSign language translation as a kind of technology with profound social significance has attracted growing researchers' interest in recent years. However, the existing sign language translation methods need to read all the videos before starting the translation, which leads to a high inference latency and also limits their application in real-life scenarios. To solve this problem, we propose SimulSLT, the first end-to-end simultaneous sign language translation model, which can translate sign language videos into target text concurrently. SimulSLT is composed of a text decoder, a boundary predictor, and a masked encoder. We 1) use the wait-k strategy for simultaneous translation. 2) design a novel boundary predictor based on the integrate-and-fire module to output the gloss boundary, which is used to model the correspondence between the sign language video and the gloss. 3) propose an innovative re-encode method to help the model obtain more abundant contextual information, which allows the existing video features to interact fully. The experimental results conducted on the RWTH-PHOENIX-Weather 2014T dataset show that SimulSLT achieves BLEU scores that exceed the latest end-to-end non-simultaneous sign language translation model while maintaining low latency, which proves the effectiveness of our method. Aoxiong Yin, Zhou Zhao 0001, Jinglin Liu, Weike Jin, Meng Zhang 0019, Xingshan Zeng, Xiaofei He 0001 |
ACM Multimedia | 3 |
| 2021 | PortaSpeech: Portable and High-Quality Generative Text-to-SpeechabstractNon-autoregressive text-to-speech (NAR-TTS) models such as FastSpeech 2 and Glow-TTS can synthesize high-quality speech from the given text in parallel. After analyzing two kinds of generative NAR-TTS models (VAE and normalizing flow), we find that: VAE is good at capturing the long-range semantics features (e.g., prosody) even with small model size but suffers from blurry and unnatural results; and normalizing flow is good at reconstructing the frequency bin-wise details but performs poorly when the number of model parameters is limited. Inspired by these observations, to generate diverse speech with natural details and rich prosody using a lightweight architecture, we propose PortaSpeech, a portable and high-quality generative text-to-speech model. Specifically, 1) to model both the prosody and mel-spectrogram details accurately, we adopt a lightweight VAE with an enhanced prior followed by a flow-based post-net with strong conditional inputs as the main architecture. 2) To further compress the model size and memory footprint, we introduce the grouped parameter sharing mechanism to the affine coupling layers in the post-net. 3) To improve the expressiveness of synthesized speech and reduce the dependency on accurate fine-grained alignment between text and speech, we propose a linguistic encoder with mixture alignment combining hard word-level alignment and soft phoneme-level alignment, which explicitly extracts word-level semantic information. Experimental results show that PortaSpeech outperforms other TTS models in both voice quality and prosody modeling in terms of subjective and objective evaluation metrics, and shows only a slight performance degradation when reducing the model parameters to 6.7M (about 4x model size and 3x runtime memory compression ratio compared with FastSpeech 2). Our extensive ablation studies demonstrate that each design in PortaSpeech is effective. Yi Ren 0006, Jinglin Liu, Zhou Zhao 0001 |
NeurIPS | 2 |
| 2020 | SimulSpeech: End-to-End Simultaneous Speech to Text TranslationabstractIn this work, we develop SimulSpeech, an endto-end simultaneous speech to text translation system which translates speech in source language to text in target language concurrently.SimulSpeech consists of a speech encoder, a speech segmenter and a text decoder, where 1) the segmenter builds upon the encoder and leverages a connectionist temporal classification (CTC) loss to split the input streaming speech in real time, 2) the encoder-decoder attention adopts a wait-k strategy for simultaneous translation.SimulSpeech is more challenging than previous cascaded systems (with simultaneous automatic speech recognition (ASR) and simultaneous neural machine translation (NMT)).We introduce two novel knowledge distillation methods to ensure the performance: 1) Attention-level knowledge distillation transfers the knowledge from the multiplication of the attention matrices of simultaneous NMT and ASR models to help the training of the attention mechanism in SimulSpeech; 2) Data-level knowledge distillation transfers the knowledge from the full-sentence NMT model and also reduces the complexity of data distribution to help on the optimization of Simul-Speech.Experiments on MuST-C English-Spanish and English-German spoken language translation datasets show that SimulSpeech achieves reasonable BLEU scores and lower delay compared to full-sentence end-to-end speech to text translation (without simultaneous translation), and better performance than the two-stage cascaded simultaneous translation model in terms of BLEU scores and translation delay. Yi Ren 0006, Jinglin Liu, Xu Tan 0003, Chen Zhang 0020, Tao Qin 0001, Zhou Zhao 0001, Tie-Yan Liu |
ACL | 2 |
| 2020 | A Study of Non-autoregressive Model for Sequence GenerationabstractNon-autoregressive (NAR) models generate all the tokens of a sequence in parallel, resulting in faster generation speed compared to their autoregressive (AR) counterparts but at the cost of lower accuracy.Different techniques including knowledge distillation and source-target alignment have been proposed to bridge the gap between AR and NAR models in various tasks such as neural machine translation (NMT), automatic speech recognition (ASR), and text to speech (TTS).With the help of those techniques, NAR models can catch up with the accuracy of AR models in some tasks but not in some others.In this work, we conduct a study to understand the difficulty of NAR sequence generation and try to answer: (1) Why NAR models can catch up with AR models in some tasks but not all?(2) Why techniques like knowledge distillation and source-target alignment can help NAR models.Since the main difference between AR and NAR models is that NAR models do not use dependency among target tokens while AR models do, intuitively the difficulty of NAR sequence generation heavily depends on the strongness of dependency among target tokens.To quantify such dependency, we propose an analysis model called CoMMA to characterize the difficulty of different NAR sequence generation tasks.We have several interesting findings: 1) Among the NMT, ASR and TTS tasks, ASR has the most target-token dependency while TTS has the least.2) Knowledge distillation reduces the target-token dependency in target sequence and thus improves the accuracy of NAR models.3) Source-target alignment constraint encourages dependency * Equal contribution. Yi Ren 0006, Jinglin Liu, Xu Tan 0003, Zhou Zhao 0001, Sheng Zhao 0002, Tie-Yan Liu |
ACL | 2 |
| 2020 | Cogging Torque Reduction for Outer Rotor Interior Permanent Magnet Synchronous MotorabstractWith the characteristic of high precision, good regulation-speed performance, high efficiency, high stability and reliability, outer rotor Interior Permanent Magnet Synchronous Motor (IPMSM) has been widely used in the machine tools, robotics, automotive electronic control field and other servo applications. In this paper, two methods to reduce the cogging torque are proposed. The first method is to minimize the fluctuation of air gap and torque ripple by optimizing the shape of flux-barrier. The second method to reduce the cogging torque is to slot the rotor. For the accurate calculation of torque distortion, Finite-Element Method (FEM) is used. In the final, the effect of these two methods on motor performance is analyzed. The results show that these two methods can greatly reduce the cogging torque, while ensuring that other performance parameters are in the best range. ZhenZhi Cao, Jinglin Liu |
IECON | 2 |
| 2020 | A Novel Maximum Current Sharing Method for Dual-Redundancy PMSM based on the Sliding Mode Error CompensationabstractDual-redundancy permanent magnet synchronous motor (PMSM) plays an important role in the aerospace field because of its redundancy and high reliability. The special motor structure can lead to the current imbalance between the two stator windings, which can result in torque disequilibrium, the increase of motor torque ripple, and the poor control performance. A new current balance strategy has come forth. This paper proposes a dual-redundancy PMSM mathematical model to explain the current generation mechanism first. In this model, the performance characteristics of the classical current balance methods are evaluated, but it is found that the dynamic response performance will not be qualified. In order to realize the current balance of two sets of three-phase symmetric coils quickly and effectively, a maximum current sharing method based on sliding mode error compensation is proposed. The simulation results prove that the proposed algorithms are effective. Xiaoning Pei, Jinglin Liu |
IECON | 2 |
| 2020 | Task-Level Curriculum Learning for Non-Autoregressive Neural Machine TranslationabstractNon-autoregressive translation (NAT) achieves faster inference speed but at the cost of worse accuracy compared with autoregressive translation (AT). Since AT and NAT can share model structure and AT is an easier task than NAT due to the explicit dependency on previous target-side tokens, a natural idea is to gradually shift the model training from the easier AT task to the harder NAT task. To smooth the shift from AT training to NAT training, in this paper, we introduce semi-autoregressive translation (SAT) as intermediate tasks. SAT contains a hyperparameter k, and each k value defines a SAT task with different degrees of parallelism. Specially, SAT covers AT and NAT as its special cases: it reduces to AT when k=1 and to NAT when k=N (N is the length of target sentence). We design curriculum schedules to gradually shift k from 1 to N, with different pacing functions and number of tasks trained at the same time. We called our method as task-level curriculum learning for NAT (TCL-NAT). Experiments on IWSLT14 De-En, IWSLT16 En-De, WMT14 En-De and De-En datasets show that TCL-NAT achieves significant accuracy improvements over previous NAT baselines and reduces the performance gap between NAT and AT models to 1-2 BLEU points, demonstrating the effectiveness of our proposed method. Jinglin Liu, Yi Ren 0006, Xu Tan 0003, Chen Zhang 0020, Tao Qin 0001, Zhou Zhao 0001, Tie-Yan Liu |
IJCAI | 1 |
| 2020 | FastLR: Non-Autoregressive Lipreading Model with Integrate-and-FireabstractLipreading is an impressive technique and there has been a definite improvement of accuracy in recent years. However, existing methods for lipreading mainly build on autoregressive (AR) model, which generate target tokens one by one and suffer from high inference latency. To breakthrough this constraint, we propose FastLR, a non-autoregressive (NAR) lipreading model which generates all target tokens simultaneously. NAR lipreading is a challenging task that has many difficulties: 1) the discrepancy of sequence lengths between source and target makes it difficult to estimate the length of the output sequence; 2) the conditionally independent behavior of NAR generation lacks the correlation across time which leads to a poor approximation of target distribution; 3) the feature representation ability of encoder can be weak due to lack of effective alignment mechanism; and 4) the removal of AR language model exacerbates the inherent ambiguity problem of lipreading. Thus, in this paper, we introduce three methods to reduce the gap between FastLR and AR model: 1) to address challenges 1 and 2, we leverage integrate-and-fire (I&F) module to model the correspondence between source video frames and output text sequence. 2) To tackle challenge 3, we add an auxiliary connectionist temporal classification (CTC) decoder to the top of the encoder and optimize it with extra CTC loss. We also add an auxiliary autoregressive decoder to help the feature extraction of encoder. 3) To overcome challenge 4, we propose a novel Noisy Parallel Decoding (NPD) for I&F and bring Byte-Pair Encoding (BPE) into lipreading. Our experiments exhibit that FastLR achieves the speedup up to 10.97× comparing with state-of-the-art lipreading model with slight WER absolute increase of 1.5% and 5.5% on GRID and LRS2 lipreading datasets respectively, which demonstrates the effectiveness of our proposed method. Jinglin Liu, Yi Ren 0006, Zhou Zhao 0001, Chen Zhang 0020, Baoxing Huai, Nicholas Jing Yuan |
ACM Multimedia | 1 |
| 2019 | A Novel FCS Model Predictive Speed Control Strategy for IPMSM Drives in Electric VehiclesabstractInterior permanent magnet synchronous motors (IPMSM) are widely used in electric vehicles (EV) because of its advantages of high-power density and compact structure, etc. However, EVs have placed intensive demands upon good performance of their traction drive systems. Therefore, high-performance control strategies are required to be developed. In order to obtain simple control structure and avoid tedious parameter tuning work for PI controllers, this paper proposes a novel single-loop based finite control set (FCS) model predictive speed control (MPSC) strategy. Firstly, two challenges of achieving an FCS-MPSC controller is discussed. On this ground, an FCS-MPSC controller by estimating d, q-axis currents firstly and then predicting the future speed information is developed to solve the problem that the optimal switching state cannot be determined by the current states. Then, a sliding mode load torque observer (SM-LTO) for load torque estimation is proposed to to calculate the future speed. Simulation results prove the effectiveness of the proposed FCS-MPSC controller and load torque observer. Jinqiu Gao, Jinglin Liu |
IECON | 2 |
| 2019 | Improved High-Frequency Square-Wave Voltage Signal Injection Sensorless Strategy for Interior Permanent Magnet Synchronous MotorsabstractIn this paper, a novel sensorless control strategy is proposed to solve the problem of sampling delay in the traditional square-wave signal injection method. A square-wave voltage vector is injected in the estimated rotating reference frame. Then, Fourier decomposition and discretization are used to analyze the values of the currents in estimated d- and q-axis, and the rotor position error is obtained through the calculation of the discrete currents. The influence of the sampling delay is reduced. Meanwhile, the effect of the amplitude and frequency of the injected signal is eliminated in the demodulation process. Compared with the conventional demodulation method, the proposed strategy shows a better stability and dynamic property. Finally, experiments are carried out to verify the correctness and effectiveness of the proposed method. Wenzhen Li, Jinglin Liu |
IECON | 2 |
| 2018 | A High-efficiency PMSM Sensorless Control Approach Based on MPC ControllerabstractAs the applications of permanent magnet synchronous motor (PMSM) become wide, the demand for highperformance control accuracy of PMSM drive systems is getting higher. It is true that the traditional dual-proportional-integral (PI) loop control strategy is widely used in PMSM drive systems. However, tuning parameters is difficult. In addition, when the system parameters (e.g., motor resistance) change, the output performance of PI controller degrades dramatically. In order to improve the capability of resisting disturbance, this paper presents a sensorless model predictive vector control for PMSM. First, it builds a discrete model of PMSM by using first-order Euler method. Second, a new single-loop model predictive control (MPC) based on back-electromotive force and I-F control is proposed, which simplifies the system structure and control algorithm. Finally, the simulation and experimental results verify that the MPC has high dynamic and steady-state performance. Jinqiu Gao, Jinglin Liu, Chao Gong 0001 |
IECON | 2 |