VLDB 2026 Research / reviewers in the wild / expert
Mengyao Ma
dblp:42/6228
· DBLP profile ↗
41ranked-venue papers
11as first author
18since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 25 · 5 first-author · 8 since 2021Systems, architecture and hardware · 6 · 2 first-authorSecurity and privacy · 5 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ClieND: Client-Side Neuron-Level Detection against Poisoning Attacks on Cross-Silo Federated LearningabstractPoisoning attacks have been shown to pose significant threats to federated learning (FL), including both untargeted and targeted attacks. To mitigate such threats, the majority of existing defenses are implemented on the server side during the aggregation process. However, as data remains inherently local in FL, these defenses either rely on unreliable statistical or structural properties of local updates, or make strong assumptions that rely on dataset information. As a result, the unique advantage of clients having access to trusted local data and training dynamics has been largely overlooked. In this work, we propose ClieND1, a novel client-side detection framework that shifts the detection of poisoning attacks from the server to the client. Specifically, ClieND enables each client to maintain the neuron importance scores of the aggregated global model in each round by leveraging its local dataset. Through tracking the inter-round changes in these scores, each client detects abnormal behavior by identifying significant discrepancy spikes, which serve as indicators of potential poisoning attacks. To evaluate the effectiveness of ClieND, we compare it against three state-of-the-art baselines under both targeted and untargeted poisoning attacks, and also assess its ability to detect attacks at an early stage. Experimental results show that ClieND achieves a false positive rate (FPR) as low as 0.01 and a true positive rate (TPR) of up to 0.99, which significantly outperforms server-side defenses, demonstrating its high detection accuracy. Additionally, the existence of poisoning attacks, even those launched by adaptive settings, can be detected in an early stage within 5 communication rounds. Mengyao Ma, Shuofeng Liu, Viet Vo, Minghong Fang, Surya Nepal, Guangdong Bai |
AsiaCCS | 1 |
| 2026 | SecureSplit: Mitigating Backdoor Attacks in Split LearningabstractSplit Learning (SL) offers a framework for collaborative model training that respects data privacy by allowing participants to share the same dataset while maintaining distinct feature sets. However, SL is susceptible to backdoor attacks, in which malicious clients subtly alter their embeddings to insert hidden triggers that compromise the final trained model. To address this vulnerability, we introduce SecureSplit, a defense mechanism tailored to SL. SecureSplit applies a dimensionality transformation strategy to accentuate subtle differences between benign and poisoned embeddings, facilitating their separation. With this enhanced distinction, we develop an adaptive filtering approach that uses a majority-based voting scheme to remove contaminated embeddings while preserving clean ones. Rigorous experiments across four datasets (CIFAR-10, MNIST, CINIC-10, and ImageNette), five backdoor attack scenarios, and seven alternative defenses confirm the effectiveness of SecureSplit under various challenging conditions. Zhihao Dou, Dongfei Cui, Weida Wang, Anjun Gao, Yueyang Quan, Mengyao Ma, Viet Vo, Guangdong Bai, Zhuqing Liu, Minghong Fang |
WWW | 6 |
| 2026 | Geometry-Aware Cholesky Projection for Indoor Radio Map SamplingabstractMulti-frequency radio maps are vital for integrated sensing and communication, offering potential applications in indoor localization and smart homes. However, it is challenging to sample at sparse measurement locations and estimate the indoor radio map at unmeasured locations. Previous sampling algorithms do not consider the rough geometry of the furniture in the indoor environment. In order to utilize the room geometry to reduce the cost of precise on-site measurements, this work proposes a Geometry-Aware Cholesky Projection algorithm that effectively utilizes inaccurate indoor geometry information to suggest better measurement locations. Additionally, this study statistically analyzes the data distribution characteristics of 3D radio environments and radio map datasets, revealing a correlation between the room geometry and worst-case error variance. These insights justify the use of geometry information to enhance sampling efficiency in radio map reconstruction. With the proposed sampling algorithm and an autoencoder pretrained on uniform randomly masked radio maps, we find that the proposed algorithm outperforms state-of-the-art sampling algorithms. Shitong Chai, Mengyao Ma, Jiahui Li 0001, Shitong Wu, Wenxue Cui, Xiaopeng Fan 0001 |
IEEE Signal Process. Lett. | 2 |
| 2025 | Spatial Frequency Interleaving Residual Autoencoder for Indoor Radio Map ReconstructionabstractIndoor radio maps with frequency domain data are difficult to reconstruct when only limited measurements at a few locations are available. Naive convolutional neural networks suffer from flawed structures in the frequency domain when predicting these radio maps, resulting in overly smoothed predictions. We propose a Spatial Frequency Interleaving Residual Autoencoder (SFIRA) architecture to tackle this problem, along with a Procedural Radio Map Generation (PRMG) method to address the data deficiency of indoor radio maps. Experimental results indicate that the proposed architecture achieves lower Normalized Root Mean Square Error (NRMSE). Visualization of the reconstruction further suggests that the proposed methods are effective. Shitong Chai, Jiahui Li 0006, Mengyao Ma, Junwen Xie, Xiaopeng Fan 0001, Xianqi Zhang |
ICASSP | 3 |
| 2025 | Modifier Unlocked: Jailbreaking Text-to-Image Models Through PromptsabstractThe unprecedented image generation capability of text-to-image models makes them double-edged swords. While these models allow users to create exquisite images through simple prompts, they also provide adversaries with opportunities to generate Not-Safe-for-Work (NSFW) content, referred to as the jailbreak attack. Despite built-in safety filters serving as a mitigation, their vulnerabilities and associated safety issues remain a significant concern. In this work, we propose MODX, the first modifier-based attack framework for jailbreaking text-to-image models. Modx leverages a heuristic algorithm with two heuristic functions (constraints) to identify modifiers that adjust the artistic genre to subtly introduce unsafe elements that drive the generated images towards NSFW. This approach takes advantage of the fact that filters are unlikely to reject images in certain styles or artistic forms, effectively inducing the models to generate NSFW content. We demonstrate the feasibility of modifier-based jailbreaking with a theoretical analysis, and provide experimental evidence of the effectiveness of MODX. Our results show that MODX outperforms existing methods in successfully achieving jailbreaking across four state-of-the-art text-to-image models. Moreover, we evaluate MODX across additional NSFW categories and on more models or model versions, demonstrating its strong scalability and generalization. Disclaimer: This paper contains NSFW language and imagery that could be offensive, distressing, and/or upsetting. Reader discretion is advised. Shuofeng Liu, Mengyao Ma, Minhui Xue 0001, Guangdong Bai |
SP | 2 |
| 2025 | Practical Poisoning Attacks with Limited Byzantine Clients in Clustered Federated LearningabstractThe presence of non-independent and identically distributed (non-IID) data among clients poses a critical challenge to the deployment of Federated Learning (FL) in practice. In response, state-of-the-art solutions known as Clustered Federated Learning (CFL) schemes, such as FL+HC and PACFL, have emerged to tackle this issue. Their main innovation is to cluster non-IID clients into groups of IID clients, such that techniques designated for IID scenarios can be easily applicable. Nonetheless, the robustness of CFL schemes remains largely unexplored, and existing Byzantine-robust defence mechanisms prove inadequate in CFL schemes and non-IID data settings. In this work, we present novel powerful CFL-specific poisoning attacks, named Cluster-U-M and Cluster-U-D. These attacks are designed to significantly reduce the model utility, measured in terms of test accuracy, for benign clients participating in the CFL schemes. Notably, these attacks remain agnostic, requiring no adversarial knowledge regarding defense solutions and benign clients themselves. At a high level, the attacks involve two steps, including cluster poisoning attacks and client-drift exploitation within clusters. The former induces the grouping of clients with different training distributions, and the latter amplifies the difference between each client's optimum and their group's average aggregation. We extensively evaluate the impact of these attacks using FL+HC and PACFL schemes on both small and large scales. The evaluation results demonstrate that the attacks can compromise up to 54% of clients, with a maximum accuracy loss of 48%. Even with only 0.1% clients compromised, which represents a minimal practical adversarial effort, these attacks can still victimize around 4% clients. We evaluate the effectiveness of two state-of-the-art Byzantine-robust defence mechanisms, i.e., FLTrust and FLAME, in countering Cluster-U-M and Cluster-U-D, and find that the attacks can victimize up to 38% of clients with an accuracy loss of 18-38% under the FL+HC scheme. Viet Vo, Mengyao Ma, Guangdong Bai, Ryan Kok Leong Ko, Surya Nepal |
SP | 2 |
| 2025 | Effects of user personality traits and usage scenarios on middle-aged and older women using voice assistantsabstractAbstract This study investigated the effects of user personality traits, usage scenarios, and age on middle-aged and older women using voice assistants. Seventy-two middle-aged and older women (mean age = 57.7 years, SD = 7.6 years) participated in a 2 (user personality: introverted and extraverted, between-subjects) × 2 (user personality traits: emotionally stable and unstable, between-subjects) × 2 (age group: middle-aged and older women, between-subjects) × 2 (usage scenario: entertainment and medical, within-subjects) mixed design experiment. The Wizard of Oz method was used in this study. The study found that introverted and emotionally stable middle-aged women were least likely to self-disclose when interacting with voice assistants. Compared with medical scenarios, participants in entertainment scenarios had higher trust in the ability and integrity of voice assistants and lower mental workload. Middle-aged women had higher trust in ability and integrity, and lower mental workload than older women. Middle-aged and older women were affected differently by user personality traits in self-disclosure and by usage scenarios in acceptance. This study has implications for policymakers, designers, and manufacturers. Mengyao Ma, Runting Zhong, Xinhui Yang |
Interact. Comput. | 1 |
| 2025 | Multi-Scale Spiking Pyramid Wireless Communication Framework for Food RecognitionabstractFood recognition applications in human health have recently garnered significant attention in the field of computer vision. With the advancement of mobile devices, robust food recognition in wireless communication has become a practical and challenging application scenario. We propose a novel Multi-scale Spiking Pyramid Transmission Network (MSPTN) to tackle this challenge. The MSPTN learns diverse and complementary local and global feature maps simultaneously, generating a comprehensive description of food images that capture the correlations of feed-specific features. The feature sender uses a three-layer Spiking Neural Network (SNN). The proposed sender compresses features into sparse and discrete spike trains, significantly reducing the required transmission bandwidth and improving channel utilization and energy efficiency. Our model introduces the Compressed Factorized Bilinear block (CFB), which employs a low-rank feature approximation to reduce computational complexity and feature transmission volume while preserving the discriminate features. The enhancement reasoning module is proposed to enhance the received features by projecting them into a higher-dimensional space and utilizing the self-attention mechanism and sum pooling to compress them back to the original dimension. We conduct extensive experiments on the ETH Food-101 and Food2k datasets. Our results reveal that the MSPTN demonstrates state-of-the-art recognition performance, even with binary spike trains. Meanwhile, the MSPTN also exhibits remarkable robustness in wireless communication scenarios. With the combination of CFB, SNN, and EFB, our model achieves significant efficiency gains, including a nearly nine-fold decrease in feature transmission volume and a three-fold improvement in runtime & computational memory speed. Wenrui Li 0001, Jiahui Li 0001, Mengyao Ma, Xiaopeng Hong, Xiaopeng Fan 0001 |
IEEE Trans. Multim. | 3 |
| 2024 | Uncovering Gradient Inversion Risks in Practical Language Model TrainingabstractThe gradient inversion attack has been demonstrated as a significant privacy threat to federated learning (FL), particularly in continuous domains such as vision models. In contrast, it is often considered less effective or highly dependent on impractical training settings when applied to language models, due to the challenges posed by the discrete nature of tokens in text data. As a result, its potential privacy threats remain largely underestimated, despite FL being an emerging training method for language models. In this work, we propose a domain-specific gradient inversion attack named GRAB (gradient inversion with hybrid optimization). GRAB features two alternating optimization processes to address the challenges caused by practical training settings, including a simultaneous optimization on dropout masks between layers for improved token recovery and a discrete optimization for effective token sequencing. GRAB can recover a significant portion (up to 92.9% recovery rate) of the private training data, outperforming the attack strategy of utilizing discrete optimization with an auxiliary model by notable improvements of up to 28.9% recovery rate in benchmark settings and 48.5% recovery rate in practical settings. GRAB provides a valuable step forward in understanding this privacy threat in the emerging FL training mode of language models. Xinguo Feng, Zhongkui Ma, Eu Joe Chegne, Mengyao Ma, Alsharif Abuadbba, Guangdong Bai |
CCS | 5 |
| 2024 | Unveiling Intellectual Property Vulnerabilities of GAN-Based Distributed Machine Learning through Model Extraction AttacksabstractGenerative Adversarial Networks (GANs), as a cornerstone of artificial intelligence (AI), are widely recognized as the intellectual property (IP) of their owners, given the sensitivity of the training data and the commercial value tied to the models. Model extraction attacks, which aim to steal well-trained proprietary models, pose a significant threat to model IP. Nevertheless, current research predominately focuses on the context of machine learning as a service (MLaaS), where the emphasis lies in understanding the attack knowledge acquired through black-box API queries. This restricted perspective exposes a critical gap in investigating model extraction attacks within realistic distributed settings for generative tasks. In this work, we present the first investigation into model extraction attacks against GANs in distributed settings. We provide a comprehensive attack taxonomy, considering three different levels of knowledge the adversary can obtain in practice. Based on it, we introduce a novel model extraction attack named MoEx, which focuses on the GAN-based distributed learning scenario, i.e., Multi-Discriminator GANs, a typical asymmetric distributed setting. MoEx uses the objective function simulation, leveraging data exchanged during the learning process, to approximate the GAN generator owned by the server. We define two attack goals for MoEx, fidelity extraction and accuracy extraction . Then we comprehensively evaluate the effectiveness of MoEx's two goals with real-world datasets. Our results demonstrate its robust capabilities in extracting generators with high fidelity and accuracy compared with existing methods. Mengyao Ma, Shuofeng Liu, Mahawaga Arachchige Pathum Chamikara, Mohan Baruwal Chhetri, Guangdong Bai |
CIKM | 1 |
| 2023 | LoDen: Making Every Client in Federated Learning a Defender Against the Poisoning Membership Inference AttacksabstractFederated learning (FL) is a widely used distributed machine learning framework. However, recent studies have shown its susceptibility to poisoning membership inference attacks (MIA). In MIA, adversaries maliciously manipulate the local updates on selected samples and share the gradients with the server (i.e., poisoning). Since honest clients perform gradient descent on samples locally, an adversary can distinguish whether the attacked sample is a training sample based on observation of the change of the sample’s prediction. This type of attack exacerbates traditional passive MIA, yet the defense mechanisms remain largely unexplored. Mengyao Ma, Yanjun Zhang 0002, Mahawaga Arachchige Pathum Chamikara, Leo Yu Zhang, Mohan Baruwal Chhetri, Guangdong Bai |
AsiaCCS | 1 |
| 2023 | Formalizing Robustness Against Character-Level Perturbations for Neural Network Language Models
Zhongkui Ma, Xinguo Feng, Shuofeng Liu, Mengyao Ma, Hao Guan 0001, Mark Huasong Meng |
ICFEM | 5 |
| 2023 | AgrEvader: Poisoning Membership Inference against Byzantine-robust Federated LearningabstractThe Poisoning Membership Inference Attack (PMIA) is a newly emerging privacy attack that poses a significant threat to federated learning (FL). An adversary conducts data poisoning (i.e., performing adversarial manipulations on training examples) to extract membership information by exploiting the changes in loss resulting from data poisoning. The PMIA significantly exacerbates the traditional poisoning attack that is primarily focused on model corruption. However, there has been a lack of a comprehensive systematic study that thoroughly investigates this topic. In this work, we conduct a benchmark evaluation to assess the performance of PMIA against the Byzantine-robust FL setting that is specifically designed to mitigate poisoning attacks. We find that all existing coordinate-wise averaging mechanisms fail to defend against the PMIA, while the detect-then-drop strategy was proven to be effective in most cases, implying that the poison injection is memorized and the poisonous effect rarely dissipates. Inspired by this observation, we propose AgrEvader, a PMIA that maximizes the adversarial impact on the victim samples while circumventing the detection by Byzantine-robust mechanisms. AgrEvader significantly outperforms existing PMIAs. For instance, AgrEvader achieved a high attack accuracy of between 72.78% (on CIFAR-10) to 97.80% (on Texas100), which is an average accuracy increase of 13.89% compared to the strongest PMIA reported in the literature. We evaluated AgrEvader on five datasets across different domains, against a comprehensive list of threat models, which included black-box, gray-box and white-box models for targeted and non-targeted scenarios. AgrEvader demonstrated consistent high accuracy across all settings tested. The code is available at: https://github.com/PrivSecML/AgrEvader. Yanjun Zhang 0002, Guangdong Bai, Mahawaga Arachchige Pathum Chamikara, Mengyao Ma, Liyue Shen, Jingwei Wang 0003, Surya Nepal, Minhui Xue 0001, Joseph K. Liu |
WWW | 4 |
| 2022 | Distributed Audio-Visual Parsing Based On Multimodal Transformer and Deep Joint Source Channel CodingabstractAudio-visual parsing (AVP) is a newly emerged multimodal perception task, which detects and classifies audio-visual events in video. However, most existing AVP networks only use a simple attention mechanism to guide audio-visual multimodal events, and are implemented in a single end. This makes it unable to effectively capture the relationship between audio-visual events, and is not suitable for implementation in the network transmission scenario. In this paper, we focus on these problems and propose a distributed audio-visual parsing network (DAVPNet) based on multimodal transformer and deep joint source channel coding (DJSCC). Multimodal transformers are used to enhance the attention calculation between audio-visual events, and DJSCC is used to apply DAVP tasks to network transmission scenarios. Finally, the Look, Listen, and Parse (LLP) dataset is used to test the algorithm performance, and the experimental results show that the DAVPNet has superior parsing performance. Penghong Wang, Jiahui Li 0006, Mengyao Ma, Xiaopeng Fan 0001 |
ICASSP | 3 |
| 2022 | Constellation Design for Deep Joint Source-Channel CodingabstractDeep learning-based joint source-channel coding (JSCC) has shown excellent performance in image and feature transmission. However, the output values of the JSCC encoder are continuous, which makes the constellation of modulation complex and dense. It is hard and expensive to design radio frequency chains for transmitting such full-resolution constellation points. In this paper, two methods of mapping the full-resolution constellation to finite constellation are proposed for real system implementation. The constellation mapping results of the pro- posed methods correspond to regular constellation and irregular constellation, respectively. We apply the methods to existing deep JSCC models and evaluate them on AWGN channels with different signal-to-noise ratios (SNRs). Experimental results show that the proposed methods outperform the traditional uniform quadrature amplitude modulation (QAM) constellation mapping method by only adding a few additional parameters. Jiahui Li 0006, Mengyao Ma, Xiaopeng Fan 0001 |
IEEE Signal Process. Lett. | 3 |
| 2021 | SNR-Adaptive Deep Joint Source-Channel Coding for Wireless Image TransmissionabstractConsidering the problem of joint source-channel coding (JSCC) for multi-user transmission of images over noisy channels, an autoencoder-based novel deep joint source-channel coding scheme is proposed in this paper. In the proposed JSCC scheme, the decoder can estimate the signal-to-noise ratio (SNR) and use it to adaptively decode the transmitted image. Experiments demonstrate that the proposed scheme achieves impressive results in adaptability for different SNRs and is robust to the noise in the SNR estimation of the decoder. To the best of our knowledge, this is the first deep JSCC scheme that focuses on the adaptability for different SNRs and can be applied to multi-user scenarios. Mingze Ding, Jiahui Li 0006, Mengyao Ma, Xiaopeng Fan 0001 |
ICASSP | 3 |
| 2021 | MSFC: Deep Feature Compression in Multi-Task NetworkabstractWith the remarkable success of deep learning, a novel AI-deployment strategy on mobile devices called collaborative intelligence (CI) is proposed recently, which can greatly improve the efficiency of neural network by distributing work-loads between mobile devices and the cloud. In order to reduce transmission overhead, feature maps obtained from mobile devices need to be compressed before being transmitted to the cloud. In this paper, we propose a multi-scale feature compression (MSFC) framework for applying complex multi-task learning network in CI deployment scenarios, which consists of a multi-scale feature fusion (MSFF) module, a single-stream feature codec (SSFC) and a multi-scale feature reconstruction (MSFR) module. When applied to the popular multi-task network Mask R-CNN, experimental results show that with less than 2% accuracy degradation, the proposed MSFC can compress the 32-bit floating point feature to 0.012 bits on average. Mengyao Ma, Jiahui Li 0006, Xiaopeng Fan 0001 |
ICME | 3 |
| 2021 | Deep Joint Source-Channel Coding for Multi-Task NetworkabstractMulti-task learning (MTL) is an efficient way to improve the performance of related tasks by sharing knowledge. However, most existing MTL networks run on a single end and are not suitable for collaborative intelligence (CI) scenarios. In this work, we propose an MTL network with a deep joint source-channel coding (JSCC) framework, which allows operating under CI scenarios. We first propose a feature fusion based MTL network (FFMNet) for joint object detection and semantic segmentation. Compared with other MTL networks, FFMNet gets higher performance with fewer parameters. Then FFMNet is split into two parts, which run on a mobile device and an edge server respectively. The feature generated by the mobile device is transmitted through the wireless channel to the edge server. To reduce the transmission overhead of the intermediate feature, a deep JSCC network is designed. By combining two networks together, the whole model achieves 512 compression for the intermediate feature and a performance loss within 2% on both tasks. At last, by training with noise, the FFMNet with JSCC is robust to various channel conditions and outperforms the separate source and channel coding scheme. Jiahui Li 0006, Mengyao Ma, Xiaopeng Fan 0001 |
IEEE Signal Process. Lett. | 4 |
| 2010 | Integration of Recursive Temporal LMMSE Denoising Filter Into Video CodecabstractThe presence of noise can dramatically affect the efficiency of video compression systems. For performance improvement, most practical video compression systems adopt a denoising filter as a pre-processing module for the video encoder, or as a post-processing module for the video decoder, but the complexity introduced by denoising can be very high. This paper first presents a recursive temporal linear minimum mean squared error (LMMSE) filter for video denoising. Based on the analysis of the hybrid video compression process, two novel schemes are presented, one for video encoding and the other for video decoding, in which the proposed recursive temporal LMMSE filter is seamlessly integrated into the encoding and the decoding processes, respectively. For both of these two schemes, the denoising is implemented with nearly no extra computation introduced. Experimental results validate the effectiveness of the proposed schemes on encoding and decoding noisy video sequences. Oscar C. Au, Mengyao Ma, Peter Hon-Wah Wong |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2010 | Edge-Directed Error ConcealmentabstractIn this paper we propose an edge-directed error concealment (EDEC) algorithm, to recover lost slices in video sequences encoded by flexible macroblock ordering. First, the strong edges in a corrupted frame are estimated based on the edges in the neighboring frames and the received area of the current frame. Next, the lost regions along these estimated edges are recovered using both spatial and temporal neighboring pixels. Finally, the remaining parts of the lost regions are estimated. Simulation results show that compared to the existing boundary matching algorithm [1] and the exemplar-based inpainting approach [2] , the proposed EDEC algorithm can reconstruct the corrupted frame with both a better visual quality and a higher decoder peak signal-to-noise ratio. Mengyao Ma, Oscar C. Au, Shueng-Han Gary Chan, Ming-Ting Sun |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2009 | Transcoding based robust streaming of compressed videoabstractA variety of techniques have been proposed to enhance the error robustness of the video streaming system. However, most of them improves the error resilience during compression rather than after compression. In this paper, we propose a novel transcoding based scheme called lossless inter frame transcoding (LIFT) scheme to improve the error resilience of existing compressed video stream. In the LIFT scheme, inter coded blocks are selectively transcoded into new kind of blocks called ‘L-block’. At the decoder, the L-block can be transcoded back to the original P-block when the prediction is available and can also be robustly decoded as I-block when the prediction is unavailable. By offline transcoding and online adjusting the ratio of P-blocks and L-blocks, the proposed streaming server achieves error robustness scalability. Experimental results demonstrate the correctness and effectiveness of the proposed method. Xiaopeng Fan 0001, Oscar C. Au, Mengyao Ma, Ling Hou, Jiantao Zhou 0001, Ngai-Man Cheung |
ICASSP | 3 |
| 2009 | On Improving the Robustness of Compressed Video by Slepian-Wolf based Lossless TranscodingabstractA variety of techniques have been proposed to enhance the error robustness of the video streaming system. However, most of them improve the error resilience during compression rather than after compression. In this paper, we propose a novel transcoding based scheme called Slepian-Wolf based inter frame transcoding (SWIFT) to improve the error resilience of existing compressed video stream. In the SWIFT scheme, inter coded blocks are selectively transcoded into new kind of blocks called ‘X-block’. At the decoder, the X-block can be transcoded back to the original P-block when there is no error in the prediction, and can also be robustly decoded as I-block when there are errors in the prediction. In the experiments, the proposed SWIFT scheme does not introduce transcoding distortion as expected, and always improves the robustness of the compressed video at all packet loss rate. Compared with the H.264 based transcoder, SWIFT achieves better RD performance and error resilience performance. Xiaopeng Fan 0001, Oscar C. Au, Mengyao Ma, Ling Hou, Jiantao Zhou 0001, Ngai-Man Cheung |
ISCAS | 3 |
| 2009 | A Novel Ray-space based Color Correction Algorithm for Multi-view VideoabstractIn multi-view video, color inconsistency among different views always exists because of imperfect camera calibration, CCD noise, etc. Since color inconsistency greatly reduces the coding efficiency and rendering quality of multi-view video, a novel ray-space based color correction algorithm is proposed in this paper. Firstly, for each epipolar plane image (EPI) in ray-space domain, feature points are extracted to form a corresponding feature EPI (FEPI). Secondly, radon transform is applied to each FEPI to detect corresponding points from different views and the average color is calculated from the detected corresponding points. Finally, for each viewpoint image, the optimal color correction matrix is calculated by minimizing the error energy between the color of the current view and the average color based on the least square error criteria. Experimental results show that the proposed algorithm greatly improves the color consistency among different views. Moreover, the coding efficiency of the corrected multi-view images is greatly improved compared to that of the original ones and the ones corrected by histogram matching method [1]. Ling Hou, Oscar C. Au, Xiaopeng Fan 0001, Mengyao Ma |
ISCAS | 4 |
| 2009 | Wyner-Ziv-based bidirectionally decodable video coding
Xiaopeng Fan 0001, Oscar C. Au, Yan Chen 0007, Jiantao Zhou 0001, Mengyao Ma, Peter Hon-Wah Wong |
J. Vis. Commun. Image Represent. | 5 |
| 2009 | A Novel Analytic Quantization-Distortion Model for Hybrid Video CodingabstractA proper theoretical quantization-distortion model for hybrid video coding is always desirable, since this allows us to explain the behavior of existing codecs and to design better ones. However, due to the existence of motion-compensated prediction, hybrid video coding introduces interframe dependency into the encoded video, which makes its quantization-distortion characteristics difficult to analyze. In this paper, a joint analysis of quantization and motion-compensated prediction is presented. For a complete analysis, we investigate not only the distortion that quantization introduces into video signal, but also its effect on motion-compensated prediction. Based on the joint analysis, a quantization-distortion model of hybrid video coding is proposed. Our extensive experimental results show that the proposed model can estimate the quantization-distortion curve of hybrid video coding with high accuracy. Furthermore, the estimation accuracy remains high for various video sequences and encoder configurations. Oscar C. Au, Mengyao Ma, Zhiqin Liang, Peter Hon-Wah Wong |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2009 | Error Resilient Video Coding Using B Pictures in H.264abstractSince the quality of compressed video is vulnerable to errors, video transmission over unreliable Internet is very challenging today. Multi-hypothesis motion-compensated prediction (MHMCP) has been shown to have error resilience capability for video transmission, where each macroblock is predicted by a linear combination of multiple signals (hypotheses). B picture prediction is a special case of MHMCP. In H.264/AVC, the prediction of B pictures is generalized such that both of the two predictions can be selected from the past pictures or from the subsequent pictures. The multiple reference picture framework in H.264/AVC also allows previously decoded B pictures to be used as references for B picture coding. In this paper, we will discuss the error resilience characteristics of the generalized B pictures in H.264/AVC. Three prediction patterns of B pictures are analyzed in terms of their error-suppressing abilities. Both theoretical models (picture level error propagation) and simulation results are given for the comparison. Mengyao Ma, Oscar C. Au, Shueng-Han Gary Chan |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2008 | Temporal search range prediction based on a linear model of motion-compensated residueabstractAn efficient temporal search range prediction method is proposed to reduce the complexity of multiple reference frames motion estimation (MRFME) in video coding. Based on a linear model of motion-compensated residue, the behavior of residues under MRFME is investigated, and the gain of multiple reference frames is analyzed. The proposed method utilizes the current residue to estimate the gain of searching more reference frames, and predicts the temporal search range that maintains the coding performance with minimum complexity. Experimental results show that the proposed scheme can significantly reduce the complexity in motion estimation while the degradation of the coding performance is negligible. Oscar C. Au, Mengyao Ma, Zhiqin Liang, Peter Hon-Wah Wong |
ICASSP | 3 |
| 2008 | Image deblocking using convex optimizationabstractImages encoded at low-bit rate may suffer from blocking artifacts, which can dramatically degrade the visual quality. In this paper, a novel approach to image deblocking is presented. Based on the analysis of image coding process and the property of natural images, an objective function and a set of constraint functions are proposed, and image deblocking is formulated as a convex optimization problem which can be easily solved using numerical methods. The feasibility of the convex optimization problem is utilized to detect the true object edges and avoid blurring. Experimental results demonstrates the effectiveness of the proposed approach. Oscar C. Au, Mengyao Ma, Xiaopeng Fan 0001, Peter Hon-Wah Wong, Daniel Pérez Palomar |
ICIP | 3 |
| 2008 | Improved bidirectionally decodable Wyner-Ziv video codingabstractReverse playback is one of the most common video cassette recording (VCR) functions for video streaming systems. However, the predictive processing techniques employed in traditional hybrid video coding schemes severely complicate the reverse-play operation. In this paper, we enhance our previously proposed bidirectionally decodable Wyner-Ziv video coding scheme which supports both forward decoding and backward decoding. We derive that in our scheme the optimal Lagrangian multiplier for the backward motion estimation should be averagely two times larger than for the forward motion estimation. The new multiplier contributes 0.2dB gain in average. We propose an optimal P-frame/M-frame selection scheme to improve rate-distortion performance when the video is transmitted over error prone channels. The new scheme outperforms both H.264 and our previous scheme at all tested loss rate, and gain up to 0.55dB over our previous scheme in low loss rate case. Xiaopeng Fan 0001, Oscar C. Au, Jiantao Zhou 0001, Mengyao Ma |
ICME | 4 |
| 2008 | Bidirectionally decodable Wyner-Ziv video codingabstractInter frame prediction technique significantly improves the compression efficiency in the hybrid video coding schemes. However, this technique causes the decoding dependency of each inter frame on all of its reference frames. This dependency complicates the reverse play operation which is the most common video cassette recording (VCR) functions. This dependency also causes error propagation when the video is transmitted over error prone channel. In this paper, we propose a novel bidirectionally decodable Wyner-Ziv video coding scheme which relaxes this inter frame dependency. The proposed bidirectionally decodable Wyner-Ziv frame can be decoded by using whether forward prediction or backward prediction as side information at the decoder, i.e. the proposed stream supports forward decoding and backward decoding simultaneously. Compared with the other schemes which support reverse playback, our scheme requires much lower bandwidth and smaller storage space. In error resilient test, our scheme outperforms H.264 up to 4dB at same bitrate. Our proposed frames also support video splicing and stream switching at arbitrary time point like I-frames. Xiaopeng Fan 0001, Oscar C. Au, Yan Chen 0007, Jiantao Zhou 0001, Mengyao Ma |
ISCAS | 5 |
| 2008 | Video decoder embedded with temporal LMMSE denoising filterabstractUnder noisy circumstance, the encoded video can be corrupted by noise and the decoded video may be very noisy. In this paper, a novel scheme is proposed for decoding noisy video bitstreams. First a temporal linear minimum mean squared error (LMMSE) estimator for video denoising is presented. Based on the analysis of hybrid video decoder, this temporal LMMSE denoising filtering is seamlessly embedded into hybrid video decoding process. The operations performed for denoising are very few, and the total complexity of the proposed scheme is very similar to that of a regular decoder. Experimental results show that compared to the standard decoder, with the proposed scheme, the subjective quality of the decoded video can be significantly improved. Oscar C. Au, Mengyao Ma, Peter Hon-Wah Wong |
ISCAS | 3 |
| 2008 | A multi-hypothesis decoder for multiple description video codingabstractMultiple Description Coding (MDC) can be used as an Error Resilience (ER) technique for video coding. In case of transmission errors, Error Concealment (EC) can be combined with MDC to reconstruct the lost frame, such that the propagated error to the following frames is reduced. In this paper we propose a novel algorithm based on a Multi-hypothesis Decoder (MHD), to improve the reconstructed video quality of MDC over packet loss networks. Both subjective and objective results show that MHD can help to achieve a better video quality than a traditional EC algorithm. Mengyao Ma, Oscar C. Au, Xiaopeng Fan 0001, Ling Hou, Shueng-Han Gary Chan |
ISCAS | 1 |
| 2008 | A novel radon based Ray-Space interpolation algorithmabstractRay-Space interpolation is one of the key technologies to make Ray-Space based FTV (Free Viewpoint Television) system feasible. Since Ray-Space data is composed of straight lines with different slopes, the problem in Ray-Space interpolation is to find the slope of the straight lines which tells us the interpolation direction. In this paper, a radon based Ray-Space interpolation algorithm is proposed. First, feature points of epipolar plane image (EPI) are extracted to form a feature EPI (FEPI). Then, radon transform is applied to FEPI to find the possible interpolation direction. Finally, a new cost function is proposed to improve the smoothness of disparity map and find the optimal interpolation direction. Experimental results show that both the performance and the robustness of the proposed algorithm are much higher than that of the traditional pixel matching based interpolation (PMI) and block matching based interpolation (BMI). Ling Hou, Oscar C. Au, Xiaopeng Fan 0001, Mengyao Ma |
MMSP | 5 |
| 2008 | Alternate motion-compensated prediction for error resilient video coding
Mengyao Ma, Oscar C. Au, Shueng-Han Gary Chan, Xiaopeng Fan 0001, Ling Hou |
J. Vis. Commun. Image Represent. | 1 |
| 2008 | Error Concealment for Frame Losses in MDCabstractMultiple description coding(MDC) is an effectiveerror resilience(ER) technique for video coding. In case of frame loss,error concealment(EC) techniques can be used in MDC to reconstruct the lost frame, with error, from which subsequent frames can be decoded directly. With such direct decoding, the subsequent decoded frames will gradually recover from the frame loss, though slowly. In this paper we propose a novel algorithm usingmultihypothesis error concealment(MHC) to improve the error recovery rate of any EC in the temporal subsampling MDC. In MHC, the simultaneous temporal-interpolated frame is used as an additional hypothesis to improve the reconstructed video quality after the lost frame. Both subjective and objective results show that MHC can achieve significantly better video quality than direct decoding. Mengyao Ma, Oscar C. Au, Shueng-Han Gary Chan, Peter Hon-Wah Wong |
IEEE Trans. Multim. | 1 |
| 2007 | Joint Decoding of Multiple Video StreamsabstractDue to the storage limit, the original raw video may not exist after compressing it into multiple video bitstreams using different compression parameters. In this paper, we suggest a least square error (LSE) algorithm to jointly decode the multiple video bitstreams, aiming to achieve better quality of the reconstructed video. The experimental results by joint decoding of multiple H.263 streams show that the proposed algorithm can significantly enhance the decoded video quality. Zhiqin Liang, Jiantao Zhou 0001, Mengyao Ma, Oscar C. Au |
ICME | 4 |
| 2007 | Temporal Video Denoising Based on Multihypothesis Motion CompensationabstractDenoising module is required by any practical video processing systems. Most existing denoising schemes are spatio-temporal filters which operate on data over three dimensions. However, to limit the number of inputs, these filters only utilize one reference frame and cannot fully exploit temporal correlation. In this paper, a recursive temporal denoising filter named multihypothesis motion compensated filter (MHMCF) is proposed. To fully exploit temporal correlation, MHMCF performs motion estimation in a number of reference frames to construct multiple hypotheses (temporal predictions) of the current pixel. These hypotheses are combined by weighted averaging to suppress noise and estimate the actual current pixel value. Based on the multihypothesis motion compensated residue model presented in this paper, we investigate the efficiency of MHMCF, and some numerical evaluations are revealed. Experimental results show that MHMCF demonstrates quite good denoising performance while the inputs are much fewer than spatio-temporal filters. Moreover, as a purely temporal filter, it can well preserve spatial details and achieve satisfactory visual quality. Oscar C. Au, Mengyao Ma, Zhiqin Liang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2006 | A Multihypothesis Motion-Compensated Temporal Filter for Video DenoisingabstractMost existing filters for video denoising are spatio-temporal filters which operate on data over 3 dimensions. This paper presents a purely linear temporal filter with multihypothesis motion compensation (MC). Compared to the spatio-temporal filters, the proposed one needs much fewer inputs and much simpler operation. Experimental results show that just by involving 3 pixels, the proposed filter can achieve very good noise suppression performance. Moreover, as a purely temporal filter, it can well preserve spatial details and achieve satisfactory visual quality. Oscar C. Au, Mengyao Ma, Zhiqin Liang, Carman K. M. Yuk |
ICIP | 3 |
| 2006 | An Encoder-Embedded Video Denoising Filter Based on the Temporal LMMSE EstimatorabstractNoise not only degrades the visual quality of video contents, but also significantly affects the coding efficiency. Based on the temporal linear minimum mean square error (LMMSE) estimator, an innovative denoising filter is proposed in this paper. The proposed filter only requires simple operations manipulating on the individual residue coefficients and can be seamlessly integrated into video encoders. Compared to traditional filter-encoder cascaded scheme, embedding the proposed filter into the video encoder can save a large amount of computation. The experimental results show that with the proposed filter embedded, both the noise suppression capability and the coding efficiency of the video encoder can be dramatically improved. Furthermore, as a purely temporal filter, it can well preserve the fine details of video contents and satisfactory visual quality can be achieved Oscar C. Au, Mengyao Ma, Zhiqin Liang |
ICME | 3 |
| 2006 | Three-loop temporal interpolation for error concealment of MDCabstractMultiple description coding (MDC) can be used as an error resilience (ER) technique for video coding. In case of transmission errors, error concealment can be combined with MDC to reconstruct the lost frame, such that the propagated error to the following frames is reduced. In this paper, we propose a new temporal error concealment method named three-loop temporal interpolation (TLTI). TLTI can be well combined with temporal sub-sampling ER methods, such as MDC and alternative motion-compensated prediction. In the simulation, we compare the performance of TLTI with unidirectional motion compensated temporal interpolation (UMCTI). Both visual and quantitive results show that TLTI can achieve a better video quality than UMCTI. Mengyao Ma, Oscar C. Au, Shueng-Han Gary Chan, Zhiqin Liang |
ISCAS | 1 |
| 2005 | A new motion compensation approach for error resilient video codingabstractMultihypothesis motion-compensated prediction (MHMCP) can be used as an error resilience technique for video coding. Motivated by MHMCP, we propose a new error resilience approach named alternative motion-compensated prediction (AMCP), where two-hypothesis and one-hypothesis predictions are alternatively used with some mechanism. Both theory and simulation results show that in case of one frame loss, the expected converged error using AMCP is smaller than that using two-hypothesis MCP. Mengyao Ma, Oscar C. Au, Shueng-Han Gary Chan |
ICIP (1) | 1 |