Lu Lu 0015

dblp:01/2086-15 · DBLP profile ↗
← Back
24ranked-venue papers
0as first author
24since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 14 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 QualiSpeech: A Speech Quality Assessment Dataset with Natural Language Reasoning and Descriptions
abstract
This paper explores a novel perspective to speech quality assessment by leveraging natural language descriptions, offering richer, more nuanced insights than traditional numerical scoring methods. Natural language feedback provides instructive recommendations and detailed evaluations, yet existing datasets lack the comprehensive annotations needed for this approach. To bridge this gap, we introduce QualiSpeech, a comprehensive low-level speech quality assessment dataset encompassing 11 key aspects and detailed natural language comments that include reasoning and contextual insights. Additionally, we propose the QualiSpeech Benchmark to evaluate the low-level speech understanding capabilities of auditory large language models (LLMs). Experimental results demonstrate that finetuned auditory LLMs can reliably generate detailed descriptions of noise and distortion, effectively identifying their types and temporal characteristics. The results further highlight the potential for incorporating reasoning to enhance the accuracy and reliability of quality assessments. The dataset can be found at https://huggingface.co/datasets/tsinghua-ee/QualiSpeech.
Siyin Wang, Wenyi Yu, Xianzhao Chen, Xiaohai Tian, Jun Zhang 0066, Lu Lu 0015, Yu Tsao 0001, Junichi Yamagishi, Yuxuan Wang 0002, Chao Zhang 0031
ACL (1)6
2025 Spy Inside: Scalable Verification of Dependable Transformers for Event Time Series Systems
abstract
Event time series appear in many software scenarios and are a necessary data type in data analytics systems. Transformers are the preferred type of sequential neural network for advanced analytics on event time series, particularly due to their significant contributions to the recent surge of large language models (LLMs). Event series analytics heavily depends on the quality of input data, which may contain natural measurement errors or adversarial noises. Since the input data deviates from the true state, the opaque nature of neural networks presents a challenge in ensuring the reliability of output, which might be deemed untrustworthy. In this paper, we introduce an innovative formal verification framework for Transformer-based event series systems, leveraging sampling, linear programming, and the extreme value theorem. This framework can support the verification of the dependability of Transformers in managing inputs characterized by unpredictability and uncertainty. To exemplify its utility, we apply our verification approach to verify natural requirements from a real-world event series environments: network traffic classification. It outperforms the current state-of-the-art verifier in terms of effectiveness, providing more stringent verified bounds. Our experimental findings provide valuable benchmarks for guaranteeing reliable deployment of systems in scenarios where the credibility of event data is compromised, and for exposing specific cases in which the expected requirements are not satisfied.
Haodong Deng, Qi Qi 0001, Lu Lu 0015, Zirui Zhuang, Xingyu Zeng, Jinguang Wang, Bo He 0003, Wei Li 0119, Jingyu Wang 0001
ICASSP3
2025 Enabling Auditory Large Language Models for Automatic Speech Quality Evaluation
abstract
Speech quality assessment typically requires evaluating audio from multiple aspects, such as mean opinion score (MOS) and speaker similarity (SIM) etc., which can be challenging to cover using one small model designed for a single task. In this paper, we propose leveraging recently introduced auditory large language models (LLMs) for automatic speech quality assessment. By employing task-specific prompts, auditory LLMs are finetuned to predict MOS, SIM and A/B testing results, which are commonly used for evaluating text-to-speech systems. Additionally, the finetuned auditory LLM is able to generate natural language descriptions assessing aspects like noisiness, distortion, discontinuity, and overall quality, providing more interpretable outputs. Extensive experiments have been performed on the NISQA, BVCC, SOMOS and VoxSim speech quality datasets, using open-source auditory LLMs such as SALMONN, Qwen-Audio, and Qwen2-Audio. For the natural language descriptions task, a commercial model Google Gemini 1.5 Pro is also evaluated. The results demonstrate that auditory LLMs achieve competitive performance compared to state-of-the-art task-specific small models in predicting MOS and SIM, while also delivering promising results in A/B testing and natural language descriptions. Our data processing scripts and finetuned model checkpoints can be found at https://github.com/bytedance/SALMONN.
Siyin Wang, Wenyi Yu, Yudong Yang, Changli Tang, Jimin Zhuang, Xianzhao Chen, Xiaohai Tian, Guangzhi Sun, Lu Lu 0015, Chao Zhang 0031
ICASSP11
2025 SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation
abstract
In order to enable fluid and natural human-machine speech interaction, existing full-duplex conversational systems often adopt modular architectures with auxiliary components such as voice activity detectors, interrupters, conversation state predictors, or multiple LLMs. These systems, however, suffer from error accumulation across modules and struggle with key challenges such as context-dependent barge-in and echo cancellation. Recent approaches, most notably Moshi, simplify the pipeline by injecting audio codecs into the token space of a single LLM. However, such methods still incur significant performance degradation when operating on the speech rather than text modality. In this paper, we introduce SALMONN-omni, the first single, standalone full-duplex speech LLM that operates without audio codecs in its token space. It features a novel dynamic thinking mechanism within the LLM backbone, enabling the model to learn when to transition between speaking and listening states. Experiments on widely used benchmarks for spoken question answering and open-domain dialogue show that SALMONN-omni achieves at least 30\% relative performance improvement over existing open-source full-duplex models and performs highly competitively to half-duplex and turn-based systems, despite using substantially less training data. Moreover, SALMONN-omni demonstrates strong performance in complex conversational scenarios, including turn-taking, backchanneling, echo cancellation and context-dependent barge-in, with further improvements achieved through reinforcement learning. Some demo conversations between user and SALMONN-omni are provided in the following repository https://github.com/bytedance/SALMONN.
Wenyi Yu, Siyin Wang, Xianzhao Chen, Xiaohai Tian, Jun Zhang 0003, Guangzhi Sun, Lu Lu 0015, Yuxuan Wang 0002, Chao Zhang 0031
NeurIPS8
2024 SA-SOT: Speaker-Aware Serialized Output Training for Multi-Talker ASR
abstract
Multi-talker automatic speech recognition plays a crucial role in scenarios involving multi-party interactions, such as meetings and conversations. Due to its inherent complexity, this task has been receiving increasing attention. Notably, the serialized output training (SOT) stands out among various approaches because of its simplistic architecture and exceptional performance. However, the frequent speaker changes in token-level SOT (t-SOT) present challenges for the autoregressive decoder in effectively utilizing context to predict output sequences. To address this issue, we introduce a masked t-SOT label, which serves as the cornerstone of an auxiliary training loss. Additionally, we utilize a speaker similarity matrix to refine the self-attention mechanism of the decoder. This strategic adjustment enhances contextual relationships within the same speaker’s tokens while minimizing interactions between different speakers’ tokens. We denote our method as speaker-aware SOT (SA-SOT). Experiments on the Librispeech datasets demonstrate that our SA-SOT obtains a relative cpWER reduction ranging from 12.75% to 22.03% on the multi-talker test sets. Furthermore, with more extensive training, our method achieves an impressive cpWER of 3.41%, establishing a new state-of-the-art result on the LibrispeechMix dataset.
Zhiyun Fan, Linhao Dong, Jun Zhang 0066, Lu Lu 0015, Zejun Ma 0001
ICASSP4
2024 Extending Multilingual ASR to New Languages Using Supplementary Encoder and Decoder Components
abstract
Extending multilingual automatic speech recognition (mASR) systems to new languages poses challenges, particularly when training data for existing languages is limited or unavailable. To tackle this issue, we suggest utilizing supplementary encoder and decoder components. Specifically, we propose appending and fine-tuning a distinct decoder designed for new languages, while preserving the parameters of existing languages to minimize disruption to their performance. Furthermore, we advocate attaching an additional encoder component to enhance acoustic representation learning for new languages, resulting in substantial improvements in word error rate performance. Our experimental findings demonstrate the effectiveness of the proposed methods for the task of extending language support within mASR systems.
Yerbolat Khassanov, Tianfeng Chen, Tze Yuang Chong, Wei Li 0119, Lu Lu 0015, Zejun Ma 0001
ICASSP6
2024 Extending Large Language Models for Speech and Audio Captioning
abstract
Multimodal large language models (LLMs) have shown promising visual perception abilities by connecting with image encoders, but their performance on auditory tasks has not yet been widely investigated. Meanwhile, automatic speech recognition (ASR) and automatic audio captioning (AAC) are often achieved with separate systems, resulting in incomplete auditory perception abilities. To fill in these gaps, in this paper, we present the first study that achieves both ASR and AAC by connecting an LLM with auditory encoders. A dual auditory encoder structure is proposed, integrating the Whisper encoder for speech and the BEATs encoder for audio events with a high temporal resolution by using a Q-Former at the window level. Experiments for ASR and AAC are performed correspondingly on the widely used LibriSpeech, GigaSpeech, WavCaps, AudioCaps, and Clotho datasets and yield promising results. In particular, state-of-the-art results are achieved on GigaSpeech, AudioCaps and Clotho. Our model is also able to caption speech and audio events simultaneously from clips with mixed speech and background audio events, which is a step towards more complete machine auditory perception.
Changli Tang, Wenyi Yu, Guangzhi Sun, Xianzhao Chen, Tian Tan 0019, Wei Li 0119, Lu Lu 0015, Zejun Ma 0001, Chao Zhang 0031
ICASSP7
2024 Connecting Speech Encoder and Large Language Model for ASR
abstract
The impressive capability and versatility of large language models (LLMs) have aroused increasing attention in automatic speech recognition (ASR), with several pioneering studies attempting to build integrated ASR models by connecting a speech encoder with an LLM. This paper presents a comparative study of three commonly used structures as connectors, including fully connected layers, multi-head cross-attention, and Q-Former. Speech encoders from the Whisper model series as well as LLMs from the Vicuna model series with different model sizes were studied. Experiments were performed on the commonly used LibriSpeech, Common Voice, and GigaSpeech datasets, where the LLMs with Q-Formers demonstrated consistent and considerable word error rate (WER) reductions over LLMs with other connector structures. Q-Former-based LLMs can generalise well to out-of-domain datasets, where 12% relative WER reductions over the Whisper baseline ASR model were achieved on the Eval2000 test set without using any in-domain training data from Switchboard. Moreover, a novel segment-level Q-Former is proposed to enable LLMs to recognise speech segments with a duration exceeding the limitation of the encoders, which results in 17% relative WER reductions over other connector structures on 90-second-long speech data.
Wenyi Yu, Changli Tang, Guangzhi Sun, Xianzhao Chen, Tian Tan 0019, Wei Li 0119, Lu Lu 0015, Zejun Ma 0001, Chao Zhang 0031
ICASSP7
2024 PolyVoice: Language Models for Speech to Speech Translation
abstract
With the huge success of GPT models in natural language processing, there is a growing interest in applying language modeling approaches to speech tasks. Currently, the dominant architecture in speech-to-speech translation (S2ST) remains the encoder-decoder paradigm, creating a need to investigate the impact of language modeling approaches in this area. In this study, we introduce PolyVoice, a language model-based framework designed for S2ST systems. Our framework comprises three decoder-only language models: a translation language model, a duration language model, and a speech synthesis language model. These language models employ different types of prompts to extract learned information effectively. By utilizing unsupervised semantic units, our framework can transfer semantic information across these models, making it applicable even to unwritten languages. We evaluate our system on Chinese $\rightarrow$ English and English $\rightarrow$ Spanish language pairs. Experimental results demonstrate that \method outperforms the state-of-the-art encoder-decoder model, producing voice-cloned speech with high translation and audio quality. Speech samples are available at https://polyvoice.github.io.
Qianqian Dong, Zhiying Huang, Qi Tian 0001, Chen Xu 0008, Tom Ko, Yunlong Zhao 0004, Tang Li 0001, Xuxin Cheng, Fengpeng Yue, Ye Bai 0001, Lu Lu 0015, Zejun Ma 0001, Yuping Wang 0005, Mingxuan Wang, Yuxuan Wang 0002
ICLR14
2024 SALMONN: Towards Generic Hearing Abilities for Large Language Models
abstract
Hearing is arguably an essential ability of artificial intelligence (AI) agents in the physical world, which refers to the perception and understanding of general auditory information consisting of at least three types of sounds: speech, audio events, and music. In this paper, we propose SALMONN, a speech audio language music open neural network, built by integrating a pre-trained text-based large language model (LLM) with speech and audio encoders into a single multimodal model. SALMONN enables the LLM to directly process and understand general audio inputs and achieve competitive performances on a number of speech and audio tasks used in training, such as automatic speech recognition and translation, auditory-information-based question answering, emotion recognition, speaker verification, and music and audio captioning etc. SALMONN also has a diverse set of emergent abilities unseen in the training, which includes but is not limited to speech translation to untrained languages, speech-based slot filling, spoken-query-based question answering, audio-based storytelling, and speech audio co-reasoning etc. The presence of cross-modal emergent abilities is studied, and a novel few-shot activation tuning approach is proposed to activate such abilities. To our knowledge, SALMONN is the first model of its type and can be regarded as a step towards AI with generic hearing abilities. The source code, model checkpoints and data are available at https://github.com/bytedance/SALMONN.
Changli Tang, Wenyi Yu, Guangzhi Sun, Xianzhao Chen, Tian Tan 0019, Wei Li 0119, Lu Lu 0015, Zejun Ma 0001, Chao Zhang 0031
ICLR7
2024 Challenges in Training PINNs: A Loss Landscape Perspective
abstract
This paper explores challenges in training Physics-Informed Neural Networks (PINNs), emphasizing the role of the loss landscape in the training process. We examine difficulties in minimizing the PINN loss function, particularly due to ill-conditioning caused by differential operators in the residual term. We compare gradient-based optimizers Adam, L-BFGS, and their combination Adam+L-BFGS, showing the superiority of Adam+L-BFGS, and introduce a novel second-order optimizer, NysNewton-CG (NNCG), which significantly improves PINN performance. Theoretically, our work elucidates the connection between ill-conditioned differential operators and ill-conditioning in the PINN loss and shows the benefits of combining first- and second-order optimization methods. Our work presents valuable insights and more powerful optimization strategies for training PINNs, which could improve the utility of PINNs for solving difficult partial differential equations.
Pratik Rathore, Weimu Lei, Zachary Frangella, Lu Lu 0015, Madeleine Udell
ICML4
2024 video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models
abstract
Speech understanding as an element of the more generic video understanding using audio-visual large language models (av-LLMs) is a crucial yet understudied aspect. This paper proposes video-SALMONN, a single end-to-end av-LLM for video processing, which can understand not only visual frame sequences, audio events and music, but speech as well. To obtain fine-grained temporal information required by speech understanding, while keeping efficient for other video elements, this paper proposes a novel multi-resolution causal Q-Former (MRC Q-Former) structure to connect pre-trained audio-visual encoders and the backbone large language model. Moreover, dedicated training approaches including the diversity loss and the unpaired audio-visual mixed training scheme are proposed to avoid frames or modality dominance. On the introduced audio-visual evaluation benchmark, video-SALMONN achieves more than 25% absolute accuracy improvements on the video-QA task and over 30% absolute accuracy improvements on audio-visual QA tasks with human speech. In addition, video-SALMONN demonstrates remarkable video comprehension and reasoning abilities on tasks that are unprecedented by other av-LLMs. Our training code and model checkpoints are available at https://github.com/bytedance/SALMONN/
Guangzhi Sun, Wenyi Yu, Changli Tang, Xianzhao Chen, Tian Tan 0019, Wei Li 0119, Lu Lu 0015, Zejun Ma 0001, Yuxuan Wang 0002, Chao Zhang 0031
ICML7
2024 Can Large Language Models Understand Spatial Audio?
Changli Tang, Wenyi Yu, Guangzhi Sun, Xianzhao Chen, Tian Tan 0019, Wei Li 0119, Jun Zhang 0066, Lu Lu 0015, Zejun Ma 0001, Yuxuan Wang 0002, Chao Zhang 0031
INTERSPEECH8
2024 MINT: Boosting Audio-Language Model via Multi-Target Pre-Training and Instruction Tuning
Yifei Xin, Zhesong Yu, Bilei Zhu, Lu Lu 0015, Zejun Ma 0001
INTERSPEECH5
2024 SD-Eval: A Benchmark Dataset for Spoken Dialogue Understanding Beyond Words
abstract
Speech encompasses a wealth of information, including but not limited to content, paralinguistic, and environmental information.This comprehensive nature of speech significantly impacts communication and is crucial for human-computer interaction.Chat-Oriented Large Language Models (LLMs), known for their general-purpose assistance capabilities, have evolved to handle multi-modal inputs, including speech.Although these models can be adept at recognizing and analyzing speech, they often fall short of generating appropriate responses.We argue that this is due to the lack of principles on task definition and model development, which requires open-source datasets and metrics suitable for model evaluation.To bridge the gap, we present SD-Eval, a benchmark dataset aimed at multidimensional evaluation of spoken dialogue understanding and generation.SD-Eval focuses on paralinguistic and environmental information and includes 7,303 utterances, amounting to 8.76 hours of speech data. The data is aggregated from eight public datasets, representing four perspectives: emotion, accent, age, and background sound.To assess the SD-Eval benchmark dataset, we implement three different models and construct a training set following a process similar to that of SD-Eval. The training set contains 1,052.72 hours of speech data and 724.4k utterances. We also conduct a comprehensive evaluation using objective evaluation methods (e.g. BLEU and ROUGE), subjective evaluations and LLM-based metrics for the generated responses.Models conditioned with paralinguistic and environmental information outperform their counterparts in both objective and subjective measures.Moreover, experiments demonstrate that LLM-based metrics show a higher correlation with human evaluation compared to traditional metrics.We open-source SD-Eval at https://github.com/amphionspace/SD-Eval.
Junyi Ao, Yuancheng Wang, Xiaohai Tian, Dekun Chen, Jun Zhang 0066, Lu Lu 0015, Yuxuan Wang 0002, Haizhou Li 0001, Zhizheng Wu 0001
NeurIPS6
2024 Slice Sandwich: Jagged Slicing Multi-Tier Dynamic Resources for Diversified V2X Services
abstract
With the advancement of intelligent transportation systems, a series of diversified V2X applications come into being, which have different key performance indicators (KPIs) and transmission features. Moreover, multi-tier computing as a new system-level architecture distributes computing and communication capabilities anywhere between the cloud and the end-user. Unfortunately, the existing network paradigm for V2X services adopts a one-shot allocation of resources ignoring the inherent differences of V2X service. To cope with these problems, three types of refined network slices for V2X services are first proposed to simultaneously support heterogeneous service characteristics without excessively splitting resources. Considering the spatiotemporal correlation between service traffic and physical resources, a jagged slicing in multi-tier dynamic resources, which forms a “slice sandwich” brightly, is realized by a dual timescale intelligent resource management scheme. The inter-slice resource configuration is based on neural bandits with upper confidence bounds at each large-time period, while the exclusive resources are managed elastically by deep Q-learning in terms of the real-time changing network state in the small slot. We developed a simulation environment by Simulation of Urban Mobility (SUMO) including real-world road conditions and traffic models. The experiment results demonstrate that the proposed scheme can effectively guarantee KPIs of V2X services and improve the system revenue compared with benchmark algorithms.
Yu Liu 0016, Zirui Zhuang, Qi Qi 0001, Jingyu Wang 0001, Dezhi Chen, Lu Lu 0015, Jianxin Liao, Zhu Han 0001
IEEE Trans. Mob. Comput.6
2023 Improving Large-Scale Deep Biasing With Phoneme Features and Text-Only Data in Streaming Transducer
abstract
Deep biasing for the Transducer can improve the recognition performance of rare words or contextual entities, which is essential in practical applications, especially for streaming Automatic Speech Recognition (ASR). However, deep biasing with large-scale rare words remains challenging, as the performance drops significantly when more distractors exist and there are words with similar grapheme sequences in the bias list. In this paper, we combine the phoneme and textual information of rare words in Transducers to distinguish words with similar pronunciation or spelling. Moreover, the introduction of training with text-only data containing more rare words benefits large-scale deep biasing. The experiments on the Librispeech corpus demonstrate that the proposed method achieves state-of-the-art performance on rare word error rate for different scales and levels of bias lists.
Jin Qiu, Jun Zhang 0066, Lu Lu 0015, Zejun Ma 0001
ASRU5
2023 AudioQR: Deep Neural Audio Watermarks For QR Code
abstract
Image-based quick response (QR) code is frequently used, but creates barriers for the visual impaired people. With the goal of ``AI for good", this paper proposes the AudioQR, a barrier-free QR coding mechanism for the visually impaired population via deep neural audio watermarks. Previous audio watermarking approaches are mainly based on handcrafted pipelines, which is less secure and difficult to apply in large-scale scenarios. In contrast, AudioQR is the first comprehensive end-to-end pipeline that hides watermarks in audio imperceptibly and robustly. To achieve this, we jointly train an encoder and decoder, where the encoder is structured as a concatenation of transposed convolutions and multi-receptive field fusion modules. Moreover, we customize the decoder training with a stochastic data augmentation chain to make the watermarked audio robust towards different audio distortions, such as environment background, room impulse response when playing through the air, music surrounding, and Gaussian noise. Experiment results indicate that AudioQR can efficiently hide arbitrary information into audio without introducing significant perceptible difference. Our code is available at https://github.com/xinghua-qu/AudioQR.
Xinghua Qu, Xiang Yin 0006, Pengfei Wei 0001, Lu Lu 0015, Zejun Ma 0001
IJCAI4
2023 Knowledge Distillation Approach for Efficient Internal Language Model Estimation
Haihua Xu 0001, Yerbolat Khassanov, Lu Lu 0015, Zejun Ma 0001, Ji Wu 0002
INTERSPEECH5
2023 Language-specific Boundary Learning for Improving Mandarin-English Code-switching Speech Recognition
Zhiyun Fan, Linhao Dong, Chen Shen 0011, Zhenlin Liang, Jun Zhang 0066, Lu Lu 0015, Zejun Ma 0001
INTERSPEECH6
2023 Text-only Domain Adaptation using Unified Speech-Text Representation in Transducer
Jun Zhang 0066, Lu Lu 0015, Zejun Ma 0001
INTERSPEECH4
2023 Random Utterance Concatenation Based Data Augmentation for Improving Short-video Speech Recognition
abstract
One of limitations in end-to-end automatic speech recognition (ASR) framework is its performance would be compromised if train-test utterance lengths are mismatched.In this paper, we propose an on-the-fly random utterance concatenation (RUC) based data augmentation method to alleviate train-test utterance length mismatch issue for short-video ASR task.Specifically, we are motivated by observations that our human-transcribed training utterances tend to be much shorter for short-video spontaneous speech (∼3 seconds on average), while our test utterance generated from voice activity detection front-end is much longer (∼10 seconds on average).Such a mismatch can lead to suboptimal performance.Empirically, it's observed the proposed RUC method significantly improves long utterance recognition without performance drop on short one.Overall, it achieves 5.72% word error rate reduction on average for 15 languages and improved robustness to various utterance length.
Yist Y. Lin, Haihua Xu 0001, Van Tung Pham, Yerbolat Khassanov, Tze Yuang Chong, Lu Lu 0015, Zejun Ma 0001
INTERSPEECH8
2023 Towards Building Voice-based Conversational Recommender Systems: Datasets, Potential Solutions and Prospects
abstract
Conversational recommender systems (CRSs) have become crucial emerging research topics in the field of RSs, thanks to their natural advantages of explicitly acquiring user preferences via interactive conversations and revealing the reasons behind recommendations. However, the majority of current CRSs are text-based, which is less user-friendly and may pose challenges for certain users, such as those with visual impairments or limited writing and reading abilities. Therefore,for the first time, this paper investigates the potential of voice-based CRS (VCRSs) to revolutionize the way users interact with RSs in a natural, intuitive, convenient, and accessible fashion. To support such studies, we create two VCRSs benchmark datasets in the e-commerce and movie domains, after realizing the lack of such datasets through an exhaustive literature review. Specifically, we first empirically verify the benefits and necessity of creating such datasets. Thereafter, we convert the user-item interactions to text-based conversations through the ChatGPT-driven prompts for generating diverse and natural templates, and then synthesize the corresponding audios via the text-to-speech model. Meanwhile, a number of strategies are delicately designed to ensure the naturalness and high quality of voice conversations. On this basis, we further explore the potential solutions and point out possible directions to build end-to-end VCRSs by seamlessly extracting and integrating voice-based inputs, thus delivering performance-enhanced, self-explainable, and user-friendly VCRSs. Our study aims to establish the foundation and motivate further pioneering research in the emerging field of VCRSs. This aligns with the principles of explainable AI and AI for social good, viz., utilizing technology's potential to create a fair, sustainable, and just world. Our codes and datasets are available on GitHub (https://github.com/hyllll/VCRS ).
Xinghua Qu, Zhu Sun 0001, Xiang Yin 0006, Yew-Soon Ong, Lu Lu 0015, Zejun Ma 0001
SIGIR6
2023 Multi-SP Network Slicing Parallel Relieving Edge Network Conflict
abstract
Network slicing is rapidly prevailing in the edge network, which provides computing, network, and storage resources for various services. When the multiple service providers (SPs) respond to their tenants in parallel, individual decisions on the dynamic and shared edge network may lead to resource conflicts, which affects the delivery of network slicing services. Existing works ignore resource interaction and coordination in the multi-SP scenario, which is not in line with the actual situation. Indeed, the complexity of resource interaction caused by the coexistence of multiple SP policies increases the difficulty to solve the formulated optimization model. In this article, we focus on the multi-SP network slicing deployment in parallel. The coordination of network resources between SPs is designed as an effective multi-agent communication mechanism that is merged into multi-agent deep reinforcement learning (MADRL). To deal with dynamic edge networks, we design the neurons hotplugging learning which realizes scalability without a high cost of model retraining. Experiments on real and random networks demonstrate that the proposed multi-SP network slicing mechanism can successfully learn coordination policies and easily adapt to various network scales. It improves the accepted requests by 7.4%, reduces resource conflicts by 14.5%, and shortens the model convergence time by 83.3%.
Rongxin Han, Dezhi Chen, Song Guo 0001, Jingyu Wang 0001, Qi Qi 0001, Lu Lu 0015, Jianxin Liao
IEEE Trans. Parallel Distributed Syst.6