EDBT 2026 Demo / reviewers in the wild / expert
Jing Wang 0037
dblp:02/736-37
· DBLP profile ↗
34ranked-venue papers
3as first author
23since 2021 · last 2026
0000-0002-3653-9951ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 17 · 2 first-author · 11 since 2021Artificial intelligence and machine learning · 10 · 1 first-author · 10 since 2021Computer networks · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Controllable timbre cloning and style replication with reference speech examples for multimodal human-computer interaction
Tianwei Lan, Yuhang Guo 0001, Mengyuan Deng, Jing Wang 0037, Wenwu Wang 0001, Chong Feng 0001 |
Neurocomputing | 4 |
| 2026 | Rateless Deep Joint Source-Channel Coding for Task-Oriented Image CommunicationsabstractThe advance of vehicle-to-everything (V2X) networks has led to many emerging data-intensive applications at the network edge. To meet the soaring data rate requirements of these applications, numerous coding schemes has been developed. However, the high heterogeneity of edge users bring challenges to these methods, including adaptation to performance requirements, coping with unknown or varying channels, as well as inefficient multicasting. In this paper, we address those problems by developing aratelessdeep joint source-channel coding scheme featuring fine-grained control over rate and informativeness at the user. Towards this end, we first design a novel class of variational information bottleneck (VIB) by employing the multinomial-Gaussian (MG) distribution, to achieve rateless transmission over an erasure channel. We derived important results on the statistical properties of this latent distribution to facilitate efficient training of MG-VIB. Then, we apply this framework to multicasting, proposing MG-VIB-M to enhance adaptability and scalability. Simulations show that our proposed method is more flexible regarding rate-relevance tradeoffs, has greater robustness against channel imperfections, and reduces bandwidth requirements for task-oriented multicasting. Zijun Qin, Zesong Fei, Jingxuan Huang, Jing Wang 0037, Xianhao Chen, Zhi Zhang 0003, Ming Xiao 0001 |
IEEE Trans. Commun. | 4 |
| 2025 | M2PAIR: A High-Quality Acoustic Impulse Response Computation ModelabstractAcoustic Impulse Response (AIR) provides crucial spatial information about the environment, significantly enhancing audio immersion. However, achieving high perceptual quality while computing AIR in real-time for interactive audio-video media (IAVM) presents a challenging problem. This study proposes the Mesh to Parametric AIR (M2PAIR), a method for computing AIR designed for IAVM. M2PAIR integrates neural networks with psychoacoustics. It takes the 3D scene mesh, the listener positions, and the sound source positions as inputs, utilizes perceptual parameters as intermediaries, and computes the desired high-quality AIR signal based on these parameters. Experimental results demonstrate that M2PAIR improves the perceptual quality of AIR output compared to existing methods while reducing the model complexity. Additionally, it meets the requirements of IAVM, including real-time computation, high sampling rates, and flexible duration for the output AIR. Xinpei Zhao, Jing Wang 0037, Xinyuan Qian 0001 |
ICASSP | 3 |
| 2025 | Classification-Oriented Semantic Communication for Internet of ThingsabstractWith the rapid development of the Internet of Things (IoT), the number of connected devices has increased exponentially, bringing significant convenience to various aspects of daily life and business operations. However, communication between IoT devices requires a significant amount of bandwidth, putting a strain on the communication system. To address this challenge, we introduce a classification-oriented semantic communication approach that transmits only essential information. We present a novel end-to-end task-oriented semantic communication model, which efficiently serves the classification task at the receiver. In particular, the proposed model first utilizes a neural network-based semantic encoder to extract classification-related semantic features. A transformer-based semantic decoder is used at the receiver to retrieve semantic features and generate classification results. We further introduce a channel encoder and decoder module to improve the ability of a single model to deal with various channel conditions. Simulation results show that, compared with the traditional method, the proposed scheme achieves higher classification accuracy on the ESC-50 dataset and UrbanSound8K dataset and has better performance for various channel conditions. Jing Wang 0037, Jingxuan Huang, Ming Zeng 0004, Zhong Zheng 0001, Ming Xiao 0001 |
VTC2025-Spring | 2 |
| 2025 | Three-way experience replay for the prediction under concept drift
Jing Wang 0037, Yanbing Ju, Peiwu Dong, Tian Ju |
Appl. Intell. | 1 |
| 2025 | Low-Bitrate High-Quality Digital Semantic Communication Based on RVQGANabstractDigital semantic communication has attracted considerable attention attributed to its potential for integration with modern digital communication systems, which has demonstrated significant performance gains. However, despite its ability to save transmission bandwidth, digital semantic communication can degrade the performance of tasks at the receiver, particularly in low-bitrate scenarios. In this article, we propose a novel low-bitrate digital semantic communication method based on a generative model for speech transmission to achieve high-quality reconstructed speech at low-bitrate transmission. In particular, we first investigate a multiscale semantic codec based on residual vector quantization with a generative adversary network (RVQGAN) model for extracting semantic information and obtaining high speech reconstruction quality while transmitting at a low bitrate. We then, design a channel noise suppression (CNS) module based on U-Net to alleviate the channel effect at low signal-to-noise ratio (SNR) by restoring high-quality semantic features, which is capable of improving the performance of the proposed method under challenging channel conditions. Moreover, a Transformer-based code predictor is utilized to further improve the robustness of the proposed method by accounting for both the channel impact and reconstruction quality. Finally, a three-stage training strategy is also presented in this article to ensure the effective operation of the proposed multiscale semantic codec, CNS module, and code predictor module. Experimental results demonstrate that the proposed method operating at 3 kb/s can save at least 50% of bandwidth while achieving higher speech restoration quality than the baseline method. Jing Wang 0037, Jingxuan Huang, Ming Zeng 0004, Zhong Zheng 0001, Zesong Fei |
IEEE Internet Things J. | 2 |
| 2025 | Non-Intrusive Speech Quality Assessment Based on Deep Neural Networks for Speech CommunicationabstractTraditionally, speech quality evaluation relies on subjective assessments or intrusive methods that require reference signals or additional equipment. However, over recent years, non-intrusive speech quality assessment has emerged as a promising alternative, capturing much attention from researchers and industry professionals. This article presents a deep learning-based method that exploits large-scale intrusive simulated data to improve the accuracy and generalization of non-intrusive methods. The major contributions of this article are as follows. First, it presents a data simulation method, which generates degraded speech signals and labels their speech quality with the perceptual objective listening quality assessment (POLQA). The generated data is proven to be useful for pretraining the deep learning models. Second, it proposes to apply an adversarial speaker classifier to reduce the impact of speaker-dependent information on speech quality evaluation. Third, an autoencoder-based deep learning scheme is proposed following the principle of representation learning and adversarial training (AT) methods, which is able to transfer the knowledge learned from a large amount of simulated speech data labeled by POLQA. With the help of discriminative representations extracted from the autoencoder, the prediction model can be trained well on a relatively small amount of speech data labeled through subjective listening tests. Fourth, an end-to-end speech quality evaluation neural network is developed, which takes magnitude and phase spectral features as its inputs. This phase-aware model is more accurate than the model using only the magnitude spectral features. A large number of experiments are carried out with three datasets: one simulated with labels obtained using POLQA and two recorded with labels obtained using subjective listening tests. The results show that the presented phase-aware method improves the performance of the baseline model and the proposed model with latent representations extracted from the adversarial autoencoder (AAE) outperforms the state-of-the-art objective quality assessment methods, reducing the root mean square error (RMSE) by 10.5% and 12.2% on the Beijing Institute of Technology (BIT) dataset and Tencent Corpus, respectively. The code and supplementary materials are available at https://github.com/liushenme/AAE-SQA. Miao Liu 0007, Jing Wang 0037, Fei Wang 0030, Jingdong Chen |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | Non-Intrusive Speech Quality Assessment with Multi-Task Learning Based on Tensor NetworkabstractWith the growing significance of non-intrusive speech quality assessment in speech systems, existing methods predominantly rely on neural networks to extract low-order features. Typically, these features undergo a low-dimensional linear transformation, yielding the network’s output. However, the intercorrelation between feature points is often overlooked. In this paper, we explore the concept of kernel method, which maps features into high dimensional space through dot product, in order to enhance the extraction of relationships among all feature points. Considering the unique advantages of tensors in complex data representation, we extend the utilization of tensor network and propose a novel framework that incorporates a matrix product state (MPS) layer to predict mean opinion score (MOS). By integrating the MPS layer, our model can transform low-order features into higher-order representations, facilitating linear transformation in a high dimensional space without increasing the number of parameters. Furthermore, we propose a loss function that concurrently assesses regression and classification biases, along with correlation with real MOS labels. Experimental results demonstrate that our proposed model consistently outperforms the baseline system across all evaluation metrics and surpasses state-of-the-art models on the test set. Miao Liu 0007, Jing Wang 0037, Lidong Yang |
ICASSP | 3 |
| 2024 | Visually Guided Binaural Audio Generation with Cross-Modal ConsistencyabstractBinaural audio delivers an immersive spatial auditory experience to human listeners, but most existing videos lack binaural audio due to the expertise required for recording environments. Recent studies have been dedicated to converting monaural audio into binaural ones conditioned on the visual inputs. In this paper, we propose a novel audio-visual spatialization network with two added audio decoders, which rely on carefully designed visual features to generate audio outputs for the left and right channels, respectively. In addition, we propose an audio-visual matching loss to further explore the correlation between binaural audio and the scene visual input. Experiment results show that the proposed method outperforms several state-of-the-art binaural audio generation methods on two benchmark datasets FAIR-Play and MUSIC-Stereo. Qualitative results are also presented to demonstrate the effectiveness of the proposed method. Miao Liu 0007, Jing Wang 0037, Xinyuan Qian 0001 |
ICASSP | 2 |
| 2024 | Semi-supervised Cross-Lingual Speech Recognition Exploiting Articulatory Features
Xinmei Su, Chenguang Hu, Jing Wang 0037 |
ICPR (33) | 5 |
| 2024 | ListenFormer: Responsive Listening Head Generation with Non-autoregressive TransformersabstractAs one of the crucial elements in human-robot interaction, responsive listening head generation has attracted considerable attention from researchers. It aims to generate a listening head video based on speaker's audio and video as well as a reference listener image. However, existing methods exhibit two limitations: 1) the generation capability of their models is limited, resulting in generated videos that are far from real ones, and 2) they mostly employ autoregressive generative models, unable to mitigate the risk of error accumulation. To tackle these issues, we propose Listenformer that leverages the powerful temporal modeling capability of transformers for generation. It can perform non-autoregressive prediction with the proposed two-stage training method, simultaneously achieving temporal continuity and overall consistency in the outputs. To fully utilize the information from the speaker inputs, we designed an audio-motion attention fusion module, which improves the correlation of audio and motion features for accurate response. Additionally, a novel decoding method called sliding window with a large shift is proposed for Listenformer, demonstrating both excellent computational efficiency and effectiveness. Extensive experiments show that Listenformer outperforms the existing state-of-the-art methods on ViCo and L2L datasets. And a perceptual user study demonstrates the comprehensive performance of our method in generating diversity, identity preserving, speaker-listener synchronization, and attitude matching. Our code is available at https://liushenme.github.io/ListenFormer.github.io/. Miao Liu 0007, Jing Wang 0037, Xinyuan Qian 0001, Haizhou Li 0001 |
ACM Multimedia | 2 |
| 2024 | A Perceptually Motivated Approach for Low-Complexity Speech Semantic CommunicationabstractDeep learning-based semantic communication is an emerging communication method that achieves cooperative transmission between source and channel. The primary objectives of semantic communication are to enhance the efficiency of information transmission and ensure the accurate restoration of semantic content. Recent studies have shown that semantic communication performs well in enhancing transmission rates, especially in low signal-to-noise ratio environments. However, existing speech semantic communication methods neglect to account for speech perception at the receiver and the complexity of the method, which limits the practical implementation of semantic communication methods. In this paper, we propose a perceptually-motivated, low-complexity speech semantic communication method. Specifically, we employ an end-to-end communication approach to transmit the source speech and obtain the reconstructed speech at the receiver. To ensure the accurate extraction of semantic information, we present a low-complexity fully convolutional semantic encoder, which increases the accuracy of semantic information extraction and improves transmission efficiency. Considering the sensitivity of human perception, a multi-resolution joint loss function has been implemented to enhance the model’s performance and guarantee that the reconstructed speech aligns with the human ear’s auditory perception. Experimental results show that the proposed method performs better on objective and subjective metrics than existing speech transmission methods. Compared with existing neural semantic transmission methods, we improve the transmission efficiency, and the number of symbols needed for transmission is decreased by 60% without compromising the quality of speech. Furthermore, the proposed semantic communication method has a lower complexity and consumes less time to transmit. Jing Wang 0037, Jingxuan Huang, Zesong Fei |
IEEE Internet Things J. | 2 |
| 2024 | Audio-Visual Temporal Forgery Detection Using Embedding-Level Fusion and Multi-Dimensional Contrastive LossabstractAudio-visual deepfake detection is the process of identifying and detecting deepfakes that have been generated using both audio and visual content with AI algorithms. Most existing methods primarily focus on the overall authenticity while neglecting the position of forgeries in time. This can be particularly problematic, as even a small alteration in a clip can significantly impact its meaning. Such brand new attacks are dangerous and how to tackle such attacks remains an open question. In this paper, we present a novel neural network-based model to tackle the temporal forgery detection (TFD) problem. It consists of new audio and visual encoders with cross-modal attention for embedding extraction, and an embedding-level fusion mechanism with self-attention for forgery localization. Besides, a multi-dimensional contrastive loss is proposed which helps the model not only to capture audio-visual inconsistency for deepfake detection but also to exploit temporal inconsistency by coherently constraining the extracted embeddings. Extensive experiments on the LAV-DF dataset show that the presented method outperforms several state-of-the-art temporal forgery localization methods by up to 23.4% on [email protected] and 13.8% on AR@100. In addition, we also show the effectiveness of the proposed model on deepfake detection. Miao Liu 0007, Jing Wang 0037, Xinyuan Qian 0001, Haizhou Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Semi-Supervised Sound Event Detection with Pre-Trained ModelabstractSound event detection (SED) is an interesting but challenging task due to the scarcity of data and diverse sound events in real life. In this paper, we focus on the semi-supervised SED task, and combine pre-trained model from other field to assist in improving the detection effect. Pre-trained models have been widely used in various tasks in the field of speech, such as automatic speech recognition, audio tagging, etc. If the training dataset is large and general enough, the embedding features extracted by the pre-trained model will cover the potential information in the original task. We use pre-trained model PANNs which is suitable for SED task and proposed two methods to fuse the features from PANNs and original model, respectively. In addition, we also propose a weight raised temporal contrastive loss to improve the model’s switching speed at event boundaries and the smoothness within events. Experimental results show that using pre-trained model features outperforms the baseline by 8.5% and 9.1% in DESED public evaluation dataset in terms of polyphonic sound detection score (PSDS). Lizhong Wang, Sijun Bi, Jing Wang 0037 |
ICASSP | 5 |
| 2023 | Deep-Reinforcement-Learning-Based NOMA-Aided Slotted ALOHA for LEO Satellite IoT NetworksabstractThe low earth orbit (LEO) satellites have received extensive attention as an essential supplement to the terrestrial network for supporting global Internet of Things (IoT) services. Considering the rapid growth of IoT devices and the significant satellite-to-ground latency, proposing low-latency, low-overhead access protocols for LEO satellite IoT systems is challenging. In this article, we propose a multibeam random access (RA) framework and deploy the deep reinforcement learning (DRL) algorithm to control the nonorthogonal multiple access (NOMA) aided RA strategy. First, we divide the satellite coverage region into multiple beams and assume that the adjacent beams share parts of regions. Hence, the devices in the sharing region are allowed to transmit packets in two periods allocated for the two beams. Then, packets in multiple beams can be decoded jointly by an interslot successive interference cancelation (SIC) decoder. In addition, we consider the heterogeneity among devices and assign different power levels for heterogeneous types of devices, which enables power-domain NOMA and the intraslot SIC decoder in this system to mitigate the collision resolution. To maximize the average throughput, the deep deterministic policy gradient (DDPG) algorithm is adopted to achieve an online decision to optimize the RA protocol where the packet repetition strategies of devices are adjusted dynamically. The simulation results show that the proposed scheme outperforms the traditional benchmark schemes with significant throughput gain. Hanxiao Yu, Zesong Fei, Jing Wang 0037, Zhiming Chen 0001, Yuping Gong |
IEEE Internet Things J. | 4 |
| 2023 | Multi-Source Localization Using Optimized Time-Frequency Representation and Sparsity Component AnalysisabstractThis paper aims to address the multi-source localization problem by exploiting the sparsity of the speech signal in the time-frequency domain, where the challenge mainly lies in extracting the sparse component. An optimized time-frequency representation and sparsity component analysis-based multi-source localization method is proposed to overcome this challenge. Firstly, extracting the sparse components relies on the accurate representation in the time-frequency domain. However, the energy leakage problem caused by linear time-frequency transformation limits the accuracy of sparse component extraction. To tackle this problem, inspired by empirical mode decomposition, the proposed method classifies all the points in the time-frequency domain into four categories based on their phase feature and mode characteristics. Each type of the point is modeled separately, and a point-by-point analysis is conducted to remove all the points affected by energy leakage. Then, based on the optimized time-frequency representation, the phase coherence criterion is used to detect the sparse component in the point level. Following that, guided by the mode consistency characteristic of sparse components, an extension scheme is proposed to recover the falsely removed sparse components. Finally, the detected sparse components are applied for the multiple source localization. The objective evaluation is performed in both simulation and actual recording environments, and the proposed method can achieve better localization accuracy compared to several existing methods. Mao-shen Jia, Dingding Yao, Jing Wang 0037 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2023 | Multiple-Speech-Source DOA Estimation Based on Single-Source Cluster DetectionabstractThis study proposes multiple-speech-source direction-of-arrival (DOA) estimation based on the distribution characteristic of the time-frequency (TF) point dominated by a single-source component (i.e., single-source point, SSP). By exploring the TF distribution characteristics of SSPs, we found that most are distributed in clusters in the TF domain. Hence, the concept of a single-source cluster (SSC) is given, each composed of adjacent TF points from one dominant sound source. Considering that SSCs have different shapes and sizes, an SSC detection method is designed based on point-to-cluster expansion, which is the research focus of this paper. A two-dimensional Gaussian function is introduced to model the theoretical distribution of the DOAs of SSPs, and a cluster expansion rule is proposed based on hypothesis testing of the DOA of a source. Two-dimensional kernel density estimation and peak search are adopted to estimate the DOAs and the number of sources using the detected SSCs. Experimental results in both simulated and real environments show that the proposed method can achieve better DOA estimation performance than some current techniques. Mao-shen Jia, Jing Wang 0037, Ruiyuan Cao |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2022 | MOS Predictor for Synthetic Speech with I-Vector InputsabstractBased on deep learning technology, non-intrusive methods have received increasing attention for synthetic speech quality assessment since it does not need reference signals. Meanwhile, i-vector has been widely used in paralinguistic speech attribute recognition such as speaker and emotion recognition, but few studies have used it to estimate speech quality. In this paper, we propose a neural-network-based model that splices the deep features extracted by convolutional neural network (CNN) and i-vector on the time axis and uses Transformer encoder as time sequence model. To evaluate the proposed method, we improve the previous prediction models and conduct experiments on Voice Conversion Challenge (VCC) 2018 and 2016 dataset. Results show that i-vector contains information very related to the quality of synthetic speech and the proposed models that utilize i-vector and Transformer encoder highly increase the accuracy of MOSNet and MBNet on both utterance-level and system-level results. Miao Liu 0007, Jing Wang 0037, Shicong Li, Lidong Yang |
ICASSP | 2 |
| 2022 | BIT-MI Deep Learning-based Model to Non-intrusive Speech Quality Assessment Challenge in Online Conferencing Applications
Miao Liu 0007, Jing Wang 0037, Jianqian Zhang, Shicong Li |
INTERSPEECH | 2 |
| 2022 | Human Sound Classification based on Feature Fusion Method with Air and Bone Conducted Signal
Jing Wang 0037, Lizhong Wang, Sijun Bi, Jianqian Zhang, Qiuyue Ma |
INTERSPEECH | 2 |
| 2022 | ASA-Net: Deep representation learning between object silhouette and attributes
Shu Yang 0007, Jing Wang 0037, Lidong Yang, Zesong Fei |
Neurocomputing | 2 |
| 2021 | Multi-Stream Gated and Pyramidal Temporal Convolutional Neural Networks for Audio-Visual Speech Separation in Multi-Talker Environments
Yiyu Luo, Jing Wang 0037, Lidong Yang |
Interspeech | 2 |
| 2021 | Fusion of Machine Learning and Privacy Preserving for Secure Facial Expression RecognitionabstractThe interest in Facial Expression Recognition (FER) is increasing day by day due to its practical and potential applications, such as human physiological interaction diagnosis and mental disease detection. This area has received much attention from the research community in recent years and achieved remarkable results; however, a significant improvement is required in spatial problems. This research work presents a novel framework and proposes an effective and robust solution for FER under an unconstrained environment; it also helps us to classify facial images in the client/server model along with preserving privacy. There are a lot of cryptography techniques available but they are computationally expensive; on the other side, we have implemented a lightweight method capable of ensuring secure communication with the help of randomization. Initially, we perform preprocessing techniques to encounter the unconstrained environment. Face detection is performed for the removal of excessive background and it detects the face in the real-world environment. Data augmentation is for the insufficient data regime. A dual-enhanced capsule network is used to handle the spatial problem. The traditional capsule networks are unable to sufficiently extract the features, as the distance varies greatly between facial features. Therefore, the proposed network is capable of spatial transformation due to the action unit aware mechanism and thus forwards the most desiring features for dynamic routing between capsules. The squashing function is used for classification purposes. Simple classification is performed through a single party, whereas we also implemented the client/server model with privacy measurements. Both parties do not trust each other, as they do not know the input of each other. We have elaborated that the effectiveness of our method remains unchanged by preserving privacy by validating the results on four popular and versatile databases that outperform all the homomorphic cryptographic techniques. Jing Wang 0037, Muhammad Shahid Anwar, Arshad Ahmad 0002, Shah Nazir, Habib Ullah Khan, Zesong Fei |
Secur. Commun. Networks | 2 |
| 2020 | An efficient method for generating assembly precedence constraints on 3D models based on a block sequence structure
Jing Wang 0037, Muhammad Shahid Anwar, Zhongpeng Zheng |
Comput. Aided Des. | 2 |
| 2020 | Measuring quality of experience for 360-degree videos in virtual reality
Muhammad Shahid Anwar, Jing Wang 0037, Wahab Khan, Sadique Ahmad, Zesong Fei |
Sci. China Inf. Sci. | 2 |
| 2019 | An Interactive Virtual Training System for Assembly and Disassembly Based on Precedence Constraints
Jing Wang 0037, Zhaoyu Yan, Muhammad Shahid Anwar |
CGI | 2 |
| 2019 | Output-based speech quality assessment using autoencoder and support vector regression
Jing Wang 0037, Yahui Shan, Jingming Kuang 0001 |
Speech Commun. | 1 |
| 2017 | An objective assessment method based on multi-level factors for panoramic videosabstractWith the development of Virtual Reality (VR) technology, single-viewpoint videos have been replaced by the multi-sviewpoint panoramic video owing to the fact that the latter brings people more immersive experiences. To improve the quality of experience (QoE) of panoramic videos, the video quality assessment (VQA) method need to be investigated. However the design of the quality metric for panoramic videos is a more complicated and harder problem, since users' feelings are affected by more psychological and physiological factors. Traditional VQA methods cannot evaluate the quality of panoramic videos accurately. In this paper, we propose a general objective full-reference quality assessment method for panoramic videos. The proposed method is based on multi-level quality factors, which are calculated with region of interest (ROI) maps. The framework is flexible and expandable, and its objective output has a higher correlation with subjective scores than that of traditional VQA methods and existing panoramic video evaluation methods. Shu Yang 0007, Junzhe Zhao, Tingting Jiang 0001, Jing Wang 0037, Tariq Rahim, Bo Zhang 0042, Zhaoji Xu, Zesong Fei |
VCIP | 4 |
| 2017 | HAS QoE prediction based on dynamic video features with data mining in LTE network
Fei Wang 0030, Zesong Fei, Jing Wang 0037, Zhikun Wu |
Sci. China Inf. Sci. | 3 |
| 2014 | A Dynamic Clustering Algorithm Design for C-RAN Based on Multi-Objective Optimization TheoryabstractCloud radio access network (C-RAN) is a new concept of network architecture, which brings a technical revolution into the wireless communication market and leads to some kind of all new mode of the future wireless communications. In this paper the clustering algorithm based on multi-objective optimization is investigated. The proposed algorithm aims at maximizing the throughput contribution of the Remote RF Head (RRH) to the whole system and minimizing its total power consumption with guaranteed energy efficiency of RRH. Using the novel greedy dynamic clustering algorithm, the joint capacity of RRHs is improved. The throughput of each RRH is first given using the pricing mechanism and the Pascoletti and Serafini Scalarization method is then implemented to solve the multiobjective optimization problem. Finally, the performance of the algorithm is assessed by the simulation results. It is shown that the novel dynamic clustering algorithm based on multiobjective optimization in the C-RAN architecture outperforms the traditional greedy clustering approach. Na Li 0001, Jing Wang 0037, Chengwen Xing, Ming Lei 0002 |
VTC Spring | 3 |
| 2014 | A real-time QoE methodology for AMR codec voice in mobile network
Wenzhi Li, Jing Wang 0037, Chengwen Xing, Zesong Fei, Jingming Kuang 0001 |
Sci. China Inf. Sci. | 2 |
| 2014 | Tensor-based blind signal recovery for multi-carrier amplify-and-forward relay networks
Jing Wang 0037, Chengwen Xing, Zesong Fei, Jingming Kuang 0001 |
Sci. China Inf. Sci. | 2 |
| 2013 | Multichannel audio signal compression based on tensor decompositionabstractThis paper proposes a novel multichannel audio signal compression method based on tensor decomposition. The multichannel audio tensor space is established with three factors (channel, time, and frequency) and is decomposed into the core tensor and three factor matrices based on tucker model. Only the truncated core tensor is transmitted to the decoder which is multiplied by the factor matrices trained before processing. The performance of the proposed method is evaluated with approximation errors, compression degree and listening tests. When the core tensor is smaller, the compression degree will be higher. A very noticeable compression capability will be achieved with an acceptable retrieved quality. The novelty of the proposed method is that it enables both high compression capability and backward compatibility with little signal distortion to the hearing. Jing Wang 0037, Chundong Xu, Jingming Kuang 0001 |
ICASSP | 1 |
| 2012 | Comparison and optimization of packet loss recovery methods based on AMR-WB for VoIP
Zhongbo Li, Stefan Bruhn, Jing Wang 0037, Jingming Kuang 0001 |
Speech Commun. | 4 |