VLDB 2026 Research / reviewers in the wild / expert
Rongshan Yu
dblp:69/6265
· DBLP profile ↗
64ranked-venue papers
19as first author
27since 2021 · last 2026
0000-0003-2179-173XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 33 · 14 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 22 · 2 first-author · 18 since 2021Artificial intelligence and machine learning · 16 · 3 first-author · 9 since 2021Computer networks · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Hybrid Routing for a Mixture of LoRA ExpertsabstractCombining Mixture of Experts (MoE) with Low-Rank Adaptation (LoRA) has shown promising efficiency in multi-task instruction tuning for Large Language Models (LLMs). While existing routing schemes for such MoE systems employ auxiliary functions to ensure both expert selection certainty and workload balance among experts, they are hindered by two critical challenges: (1) Existing methods overlook the evolving cross-expert relationships across layers, leading to inefficient expert utilization. (2) The auxiliary functions fail to incorporate cross-task semantic characteristics during expert assignment, leading to suboptimal task adaptation. To address these challenges, we propose Hybrid routing for a Mixture of LoRA Experts (HotMoE), a novel multi-task instruction tuning framework that adapts hierarchical routing to the distinct characteristics of different LLM layers. First, we design a hybrid routing module. In lower layers, expert-expert attention facilitates cross-task collaboration and generalization. In higher layers, token-expert attention enables precise alignment between task semantics and specialized experts. Second, we introduce a similarity-guided auxiliary loss module to regularize routing decisions by exploiting hidden state similarities. This loss synergistically reinforces expert specialization without sacrificing certainty of expert selection by promoting cohesive activation patterns among semantically related tasks while sharpening distinctions between conflicting ones. Experiments across two multi-task instruction tuning scenarios covering seven NLP benchmarks demonstrate that HotMoE consistently outperforms all baselines, improving Mean Relative Difference by up to 1.68% with only 3.1% of trainable parameters. Yitong Huang, Jianzhong Qi 0001, Rongshan Yu, Xiaoliang Fan, Cheng Wang 0003 |
AAAI | 5 |
| 2026 | BioMTAN: A Biological Knowledge-Guided Multi-Task Attention Network for Co-Enhanced Cancer Diagnosis and PrognosisabstractWith the advancement of precision medicine, gene expression data have become a crucial tool in both cancer diagnosis and prognosis for different cancer types. The incorporation of biological pathways as prior knowledge has gained increasing interest in tackling the difficulties of high dimensionality and noisy information within gene expression data. However, most existing approaches guided by biological pathways ignore the intrinsic link between diagnostic and prognostic tasks in cancer research. They fail to capitalize on the potential of leveraging shared biological information from both tasks to enhance gene pathway representations. To this end, we introduce the Biological Knowledge-guided Multi-task Attention Network (BioMTAN), a novel multi-task learning framework designed for simultaneous prediction of molecular subtypes and survival risk. Specifically, we compile tailored knowledge collections that comprise multiple pathways for the two tasks, model them as unique subgraphs and use a multi-level information fusion strategy to provide a wealth of biological insights. Moreover, we develop a Multi-task Attention Module, which extracts essential global information functioning as the key and value by interacting with biological pathways from different collections, and utilizes task-specific local information as the query, efficiently decoding task-awareness feature for each task and facilitating communication across tasks within cancer diagnosis and prognosis. Extensive validation on the public The Cancer Genome Atlas (TCGA) datasets confirms the enhanced performance of BioMTAN and highlights the significant pathways in each task, underscoring its potential as an instrumental asset in precision oncology. Jiajing Xie, Rongshan Yu |
IEEE J. Biomed. Health Informatics | 6 |
| 2025 | SurvMamba: State Space Model with Multi-Grained Multi-Modal Interaction for Survival PredictionabstractMulti-modal learning that combines pathological images with genomic data has significantly enhanced the accuracy of survival prediction. Nevertheless, existing methods have not fully utilized the inherent hierarchical structure within both whole slide images (WSIs) and transcriptomic data, from which better intra-modal representations and inter-modal integration could be derived. Moreover, many existing studies attempt to improve multi-modal representations through attention mechanisms, which inevitably lead to high complexity when processing high-dimensional WSIs and transcriptomic data. Recently, a structured state space model named Mamba emerged as a promising approach for its superior performance in modeling long sequences with low complexity. In this study, we propose Mamba with multi-grained multi-modal interaction (SurvMamba) for survival prediction. SurvMamba is implemented with a Hierarchical Interaction Mamba (HIM) module that facilitates efficient intra-modal interactions at different granularities, thereby capturing more detailed local features as well as rich global representations. In addition, an Interaction Fusion Mamba (IFM) module is used for cascaded inter-modal interactive fusion, yielding more comprehensive features for survival prediction. Comprehensive evaluations on five TCGA datasets demonstrate that SurvMamba outperforms other existing methods in terms of performance and computational cost. Our code is available at https://github.com/CYing18/SurvMamba. Jiajing Xie, Rongshan Yu |
BIBM | 7 |
| 2025 | SlideChat: A Large Vision-Language Assistant for Whole-Slide Pathology Image UnderstandingabstractDespite the progress made by multimodal large language models (MLLMs) in computational pathology, they remain limited by a predominant focus on patch-level analysis, missing essential contextual information at the whole-slide level. The lack of large-scale instruction datasets and the gigapixel scale of whole slide images (WSIs) pose significant developmental challenges. In this paper, we present SlideChat, the first vision-language assistant capable of understanding gigapixel whole-slide images, exhibiting excellent multimodal conversational capability and response complex instruction across diverse pathology scenarios. To support its development, we created SlideInstruction, the largest instruction-following dataset for WSIs consisting of 4.2K WSI captions and 176K VQA pairs with multiple categories. Furthermore, we propose SlideBench, a multimodal benchmark that incorporates captioning and VQA tasks to assess SlideChat’s capabilities in various settings such as microscopy, diagnosis and clinical. Compared to both general and specialized MLLMs, SlideChat exhibits exceptional capabilities, achieving state-of-the-art performance on 18 of 22 tasks. For example, it achieved an overall accuracy of 81.17% on SlideBench-VQA (TCGA), and 54.15% on SlideBench-VQA (BCNB). Our code, data, and model is publicly accessible at https://uni-medical.github.io/SlideChat.github.io. Guoan Wang, Yuanfeng Ji, Yanjun Li 0007, Jin Ye 0002, Tianbin Li, Rongshan Yu, Yu Qiao 0001, Junjun He |
CVPR | 8 |
| 2024 | Generalizable Whole Slide Image Classification with Fine-Grained Visual-Semantic InteractionabstractWhole Slide Image (WSI) classification is often formu-lated as a Multiple Instance Learning (MIL) problem. Re-cently, Vision-Language Models (VLMs) have demonstrated remarkable performance in WSI classification. However, existing methods leverage coarse-grained pathogenetic de-scriptions for visual representation supervision, which are insufficient to capture the complex visual appearance of pathogenetic images, hindering the generalizability of mod-els on diverse downstream tasks. Additionally, processing high-resolution WSIs can be computationally expensive. In this paper, we propose a novel “Fine-grained Visual-Semantic Interaction” (FiVE) framework for WSI classi-fication. It is designed to enhance the model's general-izability by leveraging the interaction between localized visual patterns and fine-grained pathological semantics. Specifically, with meticulously designed queries, we start by utilizing a large language model to extract fine-grained pathological descriptions from various non-standardized raw reports. The output descriptions are then reconstructed into fine-grained labels used for training. By introducing a Task-specific Fine-grained Semantics (TFS) module, we enable prompts to capture crucial visual information in WSIs, which enhances representation learning and aug-ments generalization capabilities significantly. Further-more, given that pathological visual patterns are redun-dantly distributed across tissue slices, we sample a subset of visual instances during training. Our method demon-strates robust generalizability and strong transferability, dominantly outperforming the counterparts on the TCGA Lung Cancer dataset with at least 9.19% higher accu-racy in few-shot experiments. The code is available at: https://github.com/lslrius/WSI_FiVE. Rongshan Yu, Liansheng Wang 0002, Yuchen Han 0003 |
CVPR | 4 |
| 2024 | Harmonizing Tradition with Technology: Using AI in Traditional Music PreservationabstractTraditional music plays a unique role in preserving our history, connecting us to our roots, and fostering a sense of identity and continuity in a rapidly changing world. However, the inheritance of traditional music is extremely challenging due to the limited availability of literature and the small population of practitioners and audiences. In this paper, we investigated the possibility of using generative models in traditional music inheritance. In particular, we studied whether singing voice conversion (SVC) models are capable of producing high-quality Nanyin, an ancient and endangered music genre that can be found in the southern Fujian province of China. Our results show that SVC models can produce Nanyin audio with relatively acceptable quality based on our subjective tests. Furthermore, our objective evaluation results show that SVC models can effectively capture F0, which means that it can faithfully capture the original audio melody. They effectively retain low-frequency information in audio, ensuring consistency in pitch and rhythm, while some high-frequency details may be slightly inconsistent. Finally, we proposed an XGBoost regression based objective test algorithm that can automatically extract audio features and generate estimated Mean Opinion Scores (MOS) of SVC produced Nanyin audio based on predesigned metrics. Our work suggests a potential new approach in traditional vocal music preservation. Tiexin Yu, Xinxia Wang, Rongshan Yu |
IJCNN | 4 |
| 2024 | FedSAC: Dynamic Submodel Allocation for Collaborative Fairness in Federated LearningabstractCollaborative fairness stands as an essential element in federated learning to encourage client participation by equitably distributing rewards based on individual contributions. Existing methods primarily focus on adjusting gradient allocations among clients to achieve collaborative fairness. However, they frequently overlook crucial factors such as maintaining consistency across local models and catering to the diverse requirements of high-contributing clients. This oversight inevitably decreases both fairness and model accuracy in practice. To address these issues, we propose FedSAC, a novel Federated learning framework with dynamic Submodel Allocation for Collaborative fairness, backed by a theoretical convergence guarantee. First, we present the concept of "bounded collaborative fairness (BCF)", which ensures fairness by tailoring rewards to individual clients based on their contributions. Second, to implement the BCF, we design a submodel allocation module with a theoretical guarantee of fairness. This module incentivizes high-contributing clients with high-performance submodels containing a diverse range of crucial neurons, thereby preserving consistency across local models. Third, we further develop a dynamic aggregation module to adaptively aggregate submodels, ensuring the equitable treatment of low-frequency neurons and consequently enhancing overall model accuracy. Extensive experiments conducted on three public benchmarks demonstrate that FedSAC outperforms all baseline methods in both fairness and model accuracy. We see this work as a significant step towards incentivizing broader client participation in federated learning. The source code is available at https://github.com/wangzihuixmu/FedSAC. Zheng Wang 0076, Lingjuan Lyu, Zhaopeng Peng, Chenglu Wen, Rongshan Yu, Cheng Wang 0003, Xiaoliang Fan |
KDD | 7 |
| 2024 | PathMethy: an interpretable AI framework for cancer origin tracing based on DNA methylationabstractDespite advanced diagnostics, 3%-5% of cases remain classified as cancer of unknown primary (CUP). DNA methylation, an important epigenetic feature, is essential for determining the origin of metastatic tumors. We presented PathMethy, a novel Transformer model integrated with functional categories and crosstalk of pathways, to accurately trace the origin of tumors in CUP samples based on DNA methylation. PathMethy outperformed seven competing methods in F1-score across nine cancer datasets and predicted accurately the molecular subtypes within nine primary tumor types. It not only excelled at tracing the origins of both primary and metastatic tumors but also demonstrated a high degree of agreement with previously diagnosed sites in cases of CUP. PathMethy provided biological insights by highlighting key pathways, functional categories, and their interactions. Using functional categories of pathways, we gained a global understanding of biological processes. For broader access, a user-friendly web server for researchers and clinicians is available at https://cup.pathmethy.com. Jiajing Xie, Hailong Zheng, Rongshan Yu, Mengsha Tong |
Briefings Bioinform. | 7 |
| 2024 | FedAVE: Adaptive data value evaluation framework for collaborative fairness in federated learning
Zhaopeng Peng, Xiaoliang Fan, Zheng Wang 0076, Shangbin Wu, Rongshan Yu, Peizhen Yang, Chuanpan Zheng, Cheng Wang 0003 |
Neurocomputing | 6 |
| 2024 | ConTIG: Continuous representation learning on temporal interaction graphs
Peizhen Yang, Xiaoliang Fan, Zonghan Wu, Shirui Pan, Longbiao Chen, Cheng Wang 0003, Rongshan Yu |
Neural Networks | 10 |
| 2023 | Detect the Unseen: An Expandable Detection Model for Stem Cell ImagesabstractStem cell culture in vitro is essential for research in cell biology, drug toxicity and translational studies. In recent years, there has been a surge in the development of deep learning-based object detection algorithms tailored for image analysis of stem cell culture. However, many of these algorithms fall short in terms of performance and interpretability. To address these challenges, we present StemCellDet, an innovative multimodal-based method for stem cell detection. By harnessing the power of the pretrained CLIP model, StemCellDet uniquely encodes descriptions of stem cell categories into text embeddings, which are then synchronized with image embeddings. This synchronization enhances the model’s ability to identify critical features for accurate stem cell model categorization. Furthermore, by integrating knowledge distillation and introducing our proposed Semantic Fusion Module (SFM), StemCellDet can adeptly identify stem cell culture categories that were not present during its training phase using only their textual descriptions. Our experiments highlight StemCellDet’s robust detection capabilities and its advantages in terms of accuracy and generalizability. Yating Lin, Sijie Lin, Zhibin Huang, Rongshan Yu |
BIBM | 8 |
| 2023 | Multi-scope Analysis Driven Hierarchical Graph Transformer for Whole Slide Image Based Cancer Survival Prediction
Wentai Hou, Bingjian Yao, Lequan Yu, Rongshan Yu, Feng Gao 0023, Liansheng Wang 0002 |
MICCAI (6) | 5 |
| 2023 | Prioritizing prognostic-associated subpopulations and individualized recurrence risk signatures from single-cell transcriptomes of colorectal cancerabstractColorectal cancer (CRC) is one of the most common gastrointestinal malignancies. There are few recurrence risk signatures for CRC patients. Single-cell RNA-sequencing (scRNA-seq) provides a high-resolution platform for prognostic signature detection. However, scRNA-seq is not practical in large cohorts due to its high cost and most single-cell experiments lack clinical phenotype information. Few studies have been reported to use external bulk transcriptome with survival time to guide the detection of key cell subtypes in scRNA-seq data. We proposed scRankXMBD, a computational framework to prioritize prognostic-associated cell subpopulations based on within-cell relative expression orderings of gene pairs from single-cell transcriptomes. scRankXMBD achieves higher precision and concordance compared with five existing methods. Moreover, we developed single-cell gene pair signatures to predict recurrence risk for patients individually. Our work facilitates the application of the rank-based method in scRNA-seq data for prognostic biomarker discovery and precision oncology. scRankXMBD is available at https://github.com/xmuyulab/scRank-XMBD. (XMBD:Xiamen Big Data, a biomedical open software initiative in the National Institute for Data Science in Health and Medicine, Xiamen University, China.). Mengsha Tong, Jinsheng Song, Zheyang Zhang, Jiajing Xie, Jingyi Tian, Chenyu Liang 0004, Rongshan Yu |
Briefings Bioinform. | 11 |
| 2023 | SC-AIR-BERT: a pre-trained single-cell model for predicting the antigen-binding specificity of the adaptive immune receptorabstractAccurately predicting the antigen-binding specificity of adaptive immune receptors (AIRs), such as T-cell receptors (TCRs) and B-cell receptors (BCRs), is essential for discovering new immune therapies. However, the diversity of AIR chain sequences limits the accuracy of current prediction methods. This study introduces SC-AIR-BERT, a pre-trained model that learns comprehensive sequence representations of paired AIR chains to improve binding specificity prediction. SC-AIR-BERT first learns the 'language' of AIR sequences through self-supervised pre-training on a large cohort of paired AIR chains from multiple single-cell resources. The model is then fine-tuned with a multilayer perceptron head for binding specificity prediction, employing the K-mer strategy to enhance sequence representation learning. Extensive experiments demonstrate the superior AUC performance of SC-AIR-BERT compared with current methods for TCR- and BCR-binding specificity prediction. Yu Zhao 0009, Xiaona Su, Sijie Mai, Chenchen Qin, Rongshan Yu, Jianhua Yao 0001 |
Briefings Bioinform. | 7 |
| 2023 | cgMSI: pathogen detection within species from nanopore metagenomic sequencing dataabstractBACKGROUND: Metagenomic sequencing is an unbiased approach that can potentially detect all the known and unidentified strains in pathogen detection. Recently, nanopore sequencing has been emerging as a highly potential tool for rapid pathogen detection due to its fast turnaround time. However, identifying pathogen within species is nontrivial for nanopore sequencing data due to the high sequencing error rate. RESULTS: We developed the core gene alleles metagenome strain identification (cgMSI) tool, which uses a two-stage maximum a posteriori probability estimation method to detect pathogens at strain level from nanopore metagenomic sequencing data at low computational cost. The cgMSI tool can accurately identify strains and estimate relative abundance at 1× coverage. CONCLUSIONS: We developed cgMSI for nanopore metagenomic pathogen detection within species. cgMSI is available at https://github.com/ZHU-XU-xmu/cgMSI . Liansheng Wang 0002, Rongshan Yu |
BMC Bioinform. | 6 |
| 2023 | Hybrid Graph Convolutional Network With Online Masked Autoencoder for Robust Multimodal Cancer Survival PredictionabstractCancer survival prediction requires exploiting related multimodal information (e.g., pathological, clinical and genomic features, etc.) and it is even more challenging in clinical practices due to the incompleteness of patient's multimodal data. Furthermore, existing methods lack sufficient intra- and inter-modal interactions, and suffer from significant performance degradation caused by missing modalities. This manuscript proposes a novel hybrid graph convolutional network, entitled HGCN, which is equipped with an online masked autoencoder paradigm for robust multimodal cancer survival prediction. Particularly, we pioneer modeling the patient's multimodal data into flexible and interpretable multimodal graphs with modality-specific preprocessing. HGCN integrates the advantages of graph convolutional networks (GCNs) and a hypergraph convolutional network (HCN) through node message passing and a hyperedge mixing mechanism to facilitate intra-modal and inter-modal interactions between multimodal graphs. With HGCN, the potential for multimodal data to create more reliable predictions of patient's survival risk is dramatically increased compared to prior methods. Most importantly, to compensate for missing patient modalities in clinical scenarios, we incorporated an online masked autoencoder paradigm into HGCN, which can effectively capture intrinsic dependence between modalities and seamlessly generate missing hyperedges for model inference. Extensive experiments and analysis on six cancer cohorts from TCGA show that our method significantly outperforms the state-of-the-arts in both complete and missing modal settings. Our codes are made available at https://github.com/lin-lcx/HGCN. Wentai Hou, Chengxuan Lin, Lequan Yu, Harry Qin, Rongshan Yu, Liansheng Wang 0002 |
IEEE Trans. Medical Imaging | 5 |
| 2023 | MuRCL: Multi-Instance Reinforcement Contrastive Learning for Whole Slide Image ClassificationabstractMulti-instance learning (MIL) is widely adop- ted for automatic whole slide image (WSI) analysis and it usually consists of two stages, i.e., instance feature extraction and feature aggregation. However, due to the "weak supervision" of slide-level labels, the feature aggregation stage would suffer from severe over-fitting in training an effective MIL model. In this case, mining more information from limited slide-level data is pivotal to WSI analysis. Different from previous works on improving instance feature extraction, this paper investigates how to exploit the latent relationship of different instances (patches) to combat overfitting in MIL for more generalizable WSI classification. In particular, we propose a novel Multi-instance Rein- forcement Contrastive Learning framework (MuRCL) to deeply mine the inherent semantic relationships of different patches to advance WSI classification. Specifically, the proposed framework is first trained in a self-supervised manner and then finetuned with WSI slide-level labels. We formulate the first stage as a contrastive learning (CL) process, where positive/negative discriminative feature sets are constructed from the same patch-level feature bags of WSIs. To facilitate the CL training, we design a novel reinforcement learning-based agent to progressively update the selection of discriminative feature sets according to an online reward for slide-level feature aggregation. Then, we further update the model with labeled WSI data to regularize the learned features for the final WSI classification. Experimental results on three public WSI classification datasets (Camelyon16, TCGA-Lung and TCGA-Kidney) demonstrate that the proposed MuRCL outperforms state-of-the-art MIL models. In addition, MuRCL can achieve comparable performance to other state-of-the-art MIL models on TCGA-Esca dataset. Zhonghang Zhu, Lequan Yu, Rongshan Yu, Liansheng Wang 0002 |
IEEE Trans. Medical Imaging | 4 |
| 2022 | H^2-MIL: Exploring Hierarchical Representation with Heterogeneous Multiple Instance Learning for Whole Slide Image AnalysisabstractCurrent representation learning methods for whole slide image (WSI) with pyramidal resolutions are inherently homogeneous and flat, which cannot fully exploit the multiscale and heterogeneous diagnostic information of different structures for comprehensive analysis. This paper presents a novel graph neural network-based multiple instance learning framework (i.e., H^2-MIL) to learn hierarchical representation from a heterogeneous graph with different resolutions for WSI analysis. A heterogeneous graph with the “resolution” attribute is constructed to explicitly model the feature and spatial-scaling relationship of multi-resolution patches. We then design a novel resolution-aware attention convolution (RAConv) block to learn compact yet discriminative representation from the graph, which tackles the heterogeneity of node neighbors with different resolutions and yields more reliable message passing. More importantly, to explore the task-related structured information of WSI pyramid, we elaborately design a novel iterative hierarchical pooling (IHPool) module to progressively aggregate the heterogeneous graph based on scaling relationships of different nodes. We evaluated our method on two public WSI datasets from the TCGA project, i.e., esophageal cancer and kidney cancer. Experimental results show that our method clearly outperforms the state-of-the-art methods on both tumor typing and staging tasks. Wentai Hou, Lequan Yu, Chengxuan Lin, Helong Huang, Rongshan Yu, Harry Qin, Liansheng Wang 0002 |
AAAI | 5 |
| 2022 | Spatial-Hierarchical Graph Neural Network with Dynamic Structure Learning for Histological Image Classification
Wentai Hou, Helong Huang, Qiong Peng, Rongshan Yu, Lequan Yu, Liansheng Wang 0002 |
MICCAI (2) | 4 |
| 2022 | Application of individualized differential expression analysis in human cancer proteomeabstractLiquid chromatography-mass spectrometry-based quantitative proteomics can measure the expression of thousands of proteins from biological samples and has been increasingly applied in cancer research. Identifying differentially expressed proteins (DEPs) between tumors and normal controls is commonly used to investigate carcinogenesis mechanisms. While differential expression analysis (DEA) at an individual level is desired to identify patient-specific molecular defects for better patient stratification, most statistical DEP analysis methods only identify deregulated proteins at the population level. To date, robust individualized DEA algorithms have been proposed for ribonucleic acid data, but their performance on proteomics data is underexplored. Herein, we performed a systematic evaluation on five individualized DEA algorithms for proteins on cancer proteomic datasets from seven cancer types. Results show that the within-sample relative expression orderings (REOs) of protein pairs in normal tissues were highly stable, providing the basis for individualized DEA for proteins using REOs. Moreover, individualized DEA algorithms achieve higher precision in detecting sample-specific deregulated proteins than population-level methods. To facilitate the utilization of individualized DEA algorithms in proteomics for prognostic biomarker discovery and personalized medicine, we provide Individualized DEP Analysis IDEPAXMBD (XMBD: Xiamen Big Data, a biomedical open software initiative in the National Institute for Data Science in Health and Medicine, Xiamen University, China.) (https://github.com/xmuyulab/IDEPA-XMBD), which is a user-friendly and open-source Python toolkit that integrates individualized DEA algorithms for DEP-associated deregulation pattern recognition. Yachen Liu, Yalan Lin, Yujuan Wu, Zheyang Zhang, Nuoqi Lin, Xianlong Wang 0002, Mengsha Tong, Rongshan Yu |
Briefings Bioinform. | 10 |
| 2022 | A downsampling method enables robust clustering and integration of single-cell transcriptome dataabstractThe random noises, sampling biases, and batch effects often confound true biological variations in single-cell RNA-sequencing (scRNA-seq) data. Adjusting such biases is key to the robust discoveries in downstream analyses, such as cell clustering, gene selection and data integration. Here we propose a model-based downsampling algorithm based on minimal unbiased representative points (MURPXMBD). MURPXMBD is designed to retrieve a set of representative points by reducing gene-wise random independent errors, while retaining the covariance structure of biological origin hence provide an unbiased representation of the cell population. Subsequent validation using benchmark datasets shows that MURPXMBD can improve the quality and accuracy of clustering algorithms, and thus facilitate the discovery of new cell types. Besides, MURPXMBD also improves the performance of dataset integration algorithms. In summary, MURPXMBD serves as a useful noise-reduction method for single-cell sequencing analysis in biomedical studies. Yudi Hu, Xuejing Lyu, Hongkun Fang, Rongshan Yu, Xiaodong Shi |
J. Biomed. Informatics | 8 |
| 2021 | scSparkXMBD: High-Performance scRNA-seq Data Processing with SparkabstractHigh-throughput single-cell RNA sequencing (scRNA-seq) data processing pipelines integrate multiple modules to transform raw scRNA-seq data to gene expression matrices, including barcode processing, sequence quality control, genome alignment and transcript quantification. With the rapid growth in data volume, the speed of scRNA-seq data processing pipeline has become a major bottleneck to large-scale scRNA-seq studies. We present scSparkXMBD1(denoted as scSpark), a cloud computing based scRNA-seq data processing pipeline. By leveraging the in-memory computing capability of Apache Spark, scSpark significantly improves the processing speed of scRNA-seq data, and achieves around 5-20 times faster than the state-of-the-art processing pipelines under the same CPU core consumption. In addition, thanks to the inherent scalability of Spark in a cloud computing environment, scSpark can further reduce the processing time for a typical scRNA-seq dataset (e.g., 640 million reads) from hours to minutes when multiple computer nodes (e.g., 16) are used. Biological evaluation also confirmed that the results generated by scSpark are highly consistent with existing scRNA-seq data processing pipelines.1XMBD refers to Xiamen Big Data, which is a biomedical open software initiative in the National Institute for Data Science in Health and Medicine, Xiamen University, China Mingxuan Gao, Lixuan Tan, Hongjin Liu, Yating Lin, Rongshan Yu |
BIBM | 7 |
| 2021 | Federated Learning with Fair AveragingabstractFairness has emerged as a critical problem in federated learning (FL). In this work, we identify a cause of unfairness in FL -- conflicting gradients with large differences in the magnitudes. To address this issue, we propose the federated fair averaging (FedFV) algorithm to mitigate potential conflicts among clients before averaging their gradients. We first use the cosine similarity to detect gradient conflicts, and then iteratively eliminate such conflicts by modifying both the direction and the magnitude of the gradients. We further show the theoretical foundation of FedFV to mitigate the issue conflicting gradients and converge to Pareto stationary solutions. Extensive experiments on a suite of federated datasets confirm that FedFV compares favorably against state-of-the-art methods in terms of fairness, accuracy and efficiency. The source code is available at https://github.com/WwZzz/easyFL. Zheng Wang 0076, Xiaoliang Fan, Jianzhong Qi 0001, Chenglu Wen, Cheng Wang 0003, Rongshan Yu |
IJCAI | 6 |
| 2021 | Comparison of high-throughput single-cell RNA sequencing data processing pipelinesabstractWith the development of single-cell RNA sequencing (scRNA-seq) technology, it has become possible to perform large-scale transcript profiling for tens of thousands of cells in a single experiment. Many analysis pipelines have been developed for data generated from different high-throughput scRNA-seq platforms, bringing a new challenge to users to choose a proper workflow that is efficient, robust and reliable for a specific sequencing platform. Moreover, as the amount of public scRNA-seq data has increased rapidly, integrated analysis of scRNA-seq data from different sources has become increasingly popular. However, it remains unclear whether such integrated analysis would be biassed if the data were processed by different upstream pipelines. In this study, we encapsulated seven existing high-throughput scRNA-seq data processing pipelines with Nextflow, a general integrative workflow management framework, and evaluated their performance in terms of running time, computational resource consumption and data analysis consistency using eight public datasets generated from five different high-throughput scRNA-seq platforms. Our work provides a useful guideline for the selection of scRNA-seq data processing pipelines based on their performance on different real datasets. In addition, these guidelines can serve as a performance evaluation framework for future developments in high-throughput scRNA-seq data processing. Mingxuan Gao, Mingyi Ling, Xinwei Tang, Rongshan Yu |
Briefings Bioinform. | 8 |
| 2021 | Snipe: highly sensitive pathogen detection from metagenomic sequencing dataabstractMetagenomics data provide rich information for the detection of foodborne pathogens from food and environmental samples that are mixed with complex background bacteria strains. While pathogen detection from metagenomic sequencing data has become an activity of increasing interest, shotgun sequencing of uncultured food samples typically produces data that contain reads from many different organisms, making accurate strain typing a challenging task. Particularly, as many pathogens may contain a common set of genes that are highly similar to those from normal bacteria in food samples, traditional strain-level abundance profiling approaches do not perform well at detecting pathogens of very low abundance levels. To overcome this limitation, we propose an abundance correction method based on species-specific genomic regions to achieve high sensitivity and high specificity in target pathogen detection at low abundance. Liansheng Wang 0002, Rongshan Yu |
Briefings Bioinform. | 5 |
| 2021 | Diamond: a multi-modal DIA mass spectrometry data processing pipelineabstractSUMMARY: Currently, various software tools are used to support two mainstream workflows for data-independent acquisition (DIA) mass spectrometry (MS) data processing, namely, spectrum-centric scoring (SCS) and peptide-centric scoring (PCS). However, a fully automatic, easily reproducible and freely accessible pipeline that simultaneously integrates SCS and PCS strategies and supports both library-free and library-based modes is absent. We developed Diamond, a Nextflow-based, containerized, multi-modal DIA-MS data processing pipeline for peptide identification and quantification. Diamond integrated two mainstream workflows for DIA data analysis, namely, SCS and PCS, for use cases both with and without assay libraries. This multi-modal pipeline serves as a versatile, easy-to-use and easily extendable toolbox for large-scale DIA data processing. AVAILABILITY: Diamond is hosted on GitHub (https://github.com/xmuyulab/Diamond) and is released under the highly permissive MIT license to encourage further customization and modification. The Docker image for Diamond is freely accessible at https://hub.docker.com/r/zeroli/diamond. Chenxin Li, Mingxuan Gao, Chuanqi Zhong, Rongshan Yu |
Bioinform. | 5 |
| 2021 | Early neoplasia identification in Barrett's esophagus via attentive hierarchical aggregation and self-distillation
Wentai Hou, Liansheng Wang 0002, Shuntian Cai, Zhenyu Lin, Rongshan Yu, Harry Qin |
Medical Image Anal. | 5 |
| 2020 | ScaleQC: a scalable lossy to lossless solution for NGS data compressionabstractMOTIVATION: Per-base quality values in Next Generation Sequencing data take a significant portion of storage even after compression. Lossy compression technologies could further reduce the space used by quality values. However, in many applications, lossless compression is still desired. Hence, sequencing data in multiple file formats have to be prepared for different applications. RESULTS: We developed a scalable lossy to lossless compression solution for quality values named ScaleQC (Scalable Quality value Compression). ScaleQC is able to provide the so-called bit-stream level scalability that the losslessly compressed bit-stream by ScaleQC can be further truncated to lower data rates without incurring an expensive transcoding operation. Despite its scalability, ScaleQC still achieves comparable compression performance at both lossless and lossy data rates compared to the existing lossless or lossy compressors. AVAILABILITY AND IMPLEMENTATION: ScaleQC has been integrated with SAMtools as a special quality value encoding mode for CRAM. Its source codes can be obtained from our integrated SAMtools (https://github.com/xmuyulab/samtools) with dependency on integrated HTSlib (https://github.com/xmuyulab/htslib). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Rongshan Yu |
Bioinform. | 1 |
| 2020 | Performance evaluation of lossy quality compression algorithms for RNA-seq dataabstractBACKGROUND: Recent advancements in high-throughput sequencing technologies have generated an unprecedented amount of genomic data that must be stored, processed, and transmitted over the network for sharing. Lossy genomic data compression, especially of the base quality values of sequencing data, is emerging as an efficient way to handle this challenge due to its superior compression performance compared to lossless compression methods. Many lossy compression algorithms have been developed for and evaluated using DNA sequencing data. However, whether these algorithms can be used on RNA sequencing (RNA-seq) data remains unclear. RESULTS: In this study, we evaluated the impacts of lossy quality value compression on common RNA-seq data analysis pipelines including expression quantification, transcriptome assembly, and short variants detection using RNA-seq data from different species and sequencing platforms. Our study shows that lossy quality value compression could effectively improve RNA-seq data compression. In some cases, lossy algorithms achieved up to 1.2-3 times further reduction on the overall RNA-seq data size compared to existing lossless algorithms. However, lossy quality value compression could affect the results of some RNA-seq data processing pipelines, and hence its impacts to RNA-seq studies cannot be ignored in some cases. Pipelines using HISAT2 for alignment were most significantly affected by lossy quality value compression, while the effects of lossy compression on pipelines that do not depend on quality values, e.g., STAR-based expression quantification and transcriptome assembly pipelines, were not observed. Moreover, regardless of using either STAR or HISAT2 as the aligner, variant detection results were affected by lossy quality value compression, albeit to a lesser extent when STAR-based pipeline was used. Our results also show that the impacts of lossy quality value compression depend on the compression algorithms being used and the compression levels if the algorithm supports setting of multiple compression levels. CONCLUSIONS: Lossy quality value compression can be incorporated into existing RNA-seq analysis pipelines to alleviate the data storage and transmission burdens. However, care should be taken on the selection of compression tools and levels based on the requirements of the downstream analysis pipelines to avoid introducing undesirable adverse effects on the analysis results. Rongshan Yu |
BMC Bioinform. | 1 |
| 2020 | A novel approach combined transfer learning and deep learning to predict TMB from histology image
Liansheng Wang 0002, Yudi Jiao, Nianyin Zeng, Rongshan Yu |
Pattern Recognit. Lett. | 5 |
| 2018 | Performance Evaluation of IMP: A Rapid Secondary Analysis Pipeline for NGS Data
Rongshan Yu |
BIBM | 4 |
| 2018 | Improving Coding Efficiency of MPEG-G Standard Using Context-Based Arithmetic Coding
Yating Lin, Shiyao Wu, Rongshan Yu |
BIBM | 4 |
| 2017 | A new noise annoyance measurement metric for urban noise sensing and evaluationabstractThis paper investigates the problem of noise-induced annoyance level evaluation, and proposes a novel annoyance measurement metric for more efficient and accurate evaluation of annoyance level of different types of noises. Results from a large-scale subjective listening test using 90 different noise clips and 96 subjects show that the proposed method can produce more consistent and reliable annoyance ratings than the widely adopted ISO method. Based on the subjective test results, we further develop an objective noise annoyance level measurement model based on the selected psycho-acoustic features extracted from the noise samples. Our evaluation results show that the objective model produces satisfactory prediction accuracy on noise annoyance level. Rongshan Yu, Haiyan Shu |
ICASSP | 2 |
| 2016 | Enhanced vote count circuit based on nor flash memory for fast similarity searchabstractA memory-based search circuit is introduced in this paper. In this circuit, the conventional memory structure is customized to provide equality comparison for each column of memory array, and a counting circuit is included at each column to record the degree of matches between query and reference data patterns. This customized memory circuit can be used for similarity search applications. Due to its massive parallel processing of equality comparisons and counting operations, the search time using this circuit has O(1) complexity. In addition, it uses NOR flash memory structure and Enhanced Vote Count (EVC) interlocked design to achieve low power and high speed. Energy consumption is significantly reduced, by approximately m-fold (m is the number of simultaneously compared pattern bits in EVC), while matching speed is m times faster, compared to original vote count circuit implemented on NOR flash memory structure. Haiyan Shu, Xiaoming Bao, Rongshan Yu |
ICASSP | 5 |
| 2015 | k-Nearest Neighbors algorithm based on weak bit implementation on Enhanced Vote Count circuitabstractk-Nearest Neighbors (kNN) algorithm is a method to find the closest points in a dataset to a query point. The result of kNN can be used for classification and regression, both of which are commonly used in data mining and machine learning. In this paper, Enhanced Vote Count (EVC) circuit, which uses hardware to compare the quantized projected values of query and training/reference vectors instead of the vectors themselves, is considered to approximate the kNN search to provide a low complexity search solution. To improve the performance of EVC with limited projection number because projection number is directly related to implementation cost of EVC circuit, the concept of weak bit is considered and only reliable binary pattern matching is evaluated. The implementation of weak bit based on EVC circuit is also described. Simulation results show that, the performance of EVC can be significantly improved with weak bit implementation under limited projection number. Haiyan Shu, Rongshan Yu |
MMSP | 3 |
| 2014 | Optimal normalisation of prediction residual for predictive coding with random accessabstractLinear prediction serves as a mathematical operation to estimate the future values of a discrete‐time signal based on a linear function of previous samples. When applied to predictive coding of waveform such as speech and audio, a common issue that plagues compression performance is the non‐stationary characteristics of prediction residuals around the starting point of the random access frames. This is because dependencies between prediction residuals and the historical waveform are interrupted to satisfy the random access requirement. In such cases, the dynamic range of the prediction residuals will fluctuate dramatically in such frames, leading to substantially poor coding performance in the subsequent entropy coder. In this study, the authors developed a solution to this long‐standing issue by establishing a theoretical relationship between the energy envelope of linear prediction residuals in the random access frames and the prediction coefficients. Using the established relationship, an adaptive normalisation method is formulated as a preprocessor to the entropy coder to mitigate the poor coding performance in the random access frames. Simulation results confirm the superiority of the proposed method over existing solutions in terms of coding efficiency performance. Haiyan Shu, Rongshan Yu |
IET Signal Process. | 2 |
| 2014 | Low-Complexity Packet Scheduling Algorithms for Streaming Scalable Media Based on Time Utility FunctionabstractWe propose a time-utility function (TUF)-based packet scheduling algorithm for streaming scalable media. In the proposed system, the scalable media is partitioned into data units of different quality layers, which are then prioritized and transmitted according to their TUFs that capture both their quality contributions to the decoded media and urgencies with respect to their playback schedule. For optimal streaming quality while maintaining a reasonable computational complexity, packet transmissions are scheduled using a low-complexity algorithm based on utility accrual maximization. The computational complexity can be further reduced by considering the look-ahead window and a modified utility function. Simulation results show that the proposed scheduling algorithms achieve near-optimal performance when compared with the operational rate distortion bound of the stream source at any given bandwidth budget. Rongshan Yu, Haiyan Shu |
IEEE Trans. Multim. | 1 |
| 2013 | A novel scalable audio coding schemeabstractA new scalable audio coding scheme is introduced in this paper. Its core idea is to create one additional scalability dimension during the encoding process for the purpose of generating a plural of scalable sub-bitstreams. Based on the multiple sub-streams, a smart truncator is designed that can truncate these sub-bitstreams with optimal rate-distortion (R-D) tradeoff. Benefited from the flexible R-D trade-off, the proposed new scheme could, within a wide bitrate range, outperform those traditional scalable coding schemes, which usually provides a fixed R-D relationship designed at a specified bitrate. To verify the performance, the proposed scheme is further implemented based on a prior art scalable audio codec. Significant quality improvement is observed from the new codec via a series of subjective listening tests. Haiyan Shu, Rongshan Yu, Susanto Rahardja |
ICASSP | 3 |
| 2013 | Time utility function based packet scheduling algorithm for streaming scalable mediaabstractIn this paper, a packet scheduling algorithm that is based on a Time-Utility Function (TUF) is proposed. In the proposed system, the scalable media is partitioned into data units of different quality layers, which are then prioritized and transmitted according to their TUF's that capture both their quality contributions to the decoded media and urgencies with respect to their playback schedule. For optimal streaming quality and meanwhile maintaining a reasonable computational complexity, the scheduling of packet transmission is obtained from a low-complexity packet scheduling algorithm based on utility accrual maximization. Experimental results show that the proposed scheduling algorithm achieves near optimal performance when compared to the operational rate distortion bound of the stream source at the capacity of the network. Rongshan Yu, Haiyan Shu, Susanto Rahardja |
ICME | 1 |
| 2013 | Speech enhancement based on soft audible noise masking and noise power estimation
Rongshan Yu |
Speech Commun. | 1 |
| 2012 | Detecting Intelligibility by Linear Dimensionality Reduction and Normalized Voice Quality Hierarchical Features
Dong-Yan Huang, Yongwei Zhu, Dajun Wu, Rongshan Yu |
INTERSPEECH | 4 |
| 2011 | Low-complexity priority based packet scheduling for streaming MPEG-4 SLSabstractIn this paper, we propose a low-complexity priority based packet scheduling algorithm for streaming MPEG-4 Scalable to Lossless (SLS) encoded audio. In the proposed system, the SLS encoded frames are partitioned into data units of different quality layers, which are transmitted according to their quality contribution to the final decoded audio and their urgency relative to the playback progress. Experimental results show that the proposed scheduling algorithm has an even lower compared to traditional greedy algorithm for packet scheduling, while outperforms them by a significant margin in for terms of quality of the streamed audio. Rongshan Yu, Dajun Wu, Susanto Rahardja |
MMSP | 1 |
| 2010 | Beamforming performance of circular microphone array on spherical platform near bottom boundaryabstractSpherical microphone arrays have been extensively studied for multimedia applications by both academic and industrial communities. General assumption of such study is based upon a planar wave traveling in free space model. However, in practical applications, acoustic wave reflections from boundary of a confined environment such as rooms or nearby furniture's carrying the array may degrade array performance significantly. In this paper, we present an approach for beamforming and direction of arrival (DOA) estimation using circular microphone array mounted on the equator of a sphere near a bottom boundary where the bottom reflection is dominant over other reflection, which corresponds to the applications such as spherical microphone array set on table or floor. We examine impact of reverberation to the beamforming performance and propose an algorithm to improve the microphone array performance. We first introduce an acoustic spherical scattering model to provide a theoretical background. The boundary reflection model is subsequently developed based on two-path ray theory. We further propose an approach using corrected steering vector to improve the beamforming performance to compensate mismatch of plane wave assumption in free space that the normal beam-forming is based upon. The effectiveness of the proposed approach is evaluated by numerical simulations. Susanto Rahardja, Rongshan Yu |
ICME | 3 |
| 2009 | A low-complexity noise estimation algorithm based on smoothing of noise power estimation and estimation bias correctionabstractThis paper presents a low-complexity algorithm for tracking the noise spectral variance of speech contaminated by non-stationary noise sources. The proposed algorithm is based upon a recursive refinement process in which each step of the algorithm expectation of the instantaneous noise power is calculated based on information from the incoming signal and the current estimated distribution parameters, and estimation of the distribution parameter is refined accordingly to incorporate the expectation results. A bias estimation correction method is also introduced in the algorithm to avoid estimation errors that may occur when there is a significant mismatch between the statistics of the input signal and the current estimated distribution parameters. The proposed algorithm is compared to the Minimum Statistics method and it is found that the proposed algorithm achieves similar or better performances for various noise conditions and SNR settings. Rongshan Yu |
ICASSP | 1 |
| 2007 | Low-Complexity Binaural Decoding Using Time/Frequency Domain HRTF Equalization
Rongshan Yu, Charles Q. Robinson, Corey Cheng |
MMM (1) | 1 |
| 2007 | On Integer MDCT for Perceptual Audio CodingabstractIn MPEG-4 scalable lossless coding (SLS) which was recently published as an ISO standard in June 2006, the integer modified discrete cosine transform (IntMDCT) was adopted to enable efficient lossless reconstruction. In addition, there is an MDCT filterbank which is inherent to the advanced audio coding (AAC) core that is present in the SLS codec. The presence of two filterbanks have undoubtedly increased the complexity of the implementation, and it is for this reason that the MDCT is disabled and the IntMDCT is then the only type of filterbank that is employed in SLS for both lossy and lossless operations. Because of the rounding operations in the IntMDCT, there is a concern if the use of IntMDCT for perceptual audio coding will eventually degrade the fidelity of the audio codec. This paper addresses this concern by analyzing the performance of the IntMDCT in a lossy coding scenario. It is found that noise introduced by the IntMDCT does not affect the perceptual quality of the coded audio under standard playback circumstances. As such, it concludes that the MDCT and IntMDCT filterbanks are interchangeable at lossy bitrate, and the way of using only the IntMDCT filterbank in scalable audio coding is also justified. Susanto Rahardja, Rongshan Yu, Soo Ngee Koh |
IEEE Trans. Speech Audio Process. | 3 |
| 2006 | Cascaded RLS-LMS Prediction in MPEG-4 Lossless Audio CodingabstractA new MPEG-4 standard for lossless audio coding is going to be published in 2006. This coming international standard consists of two parts: the transform-domain scalable to lossless coding (SLS), and the time-domain audio lossless coding (ALS). In ALS, linear prediction is used to compress the dynamic ranges of the input audio signal. The prediction residual is coded by an entropy coder with either Rice code or arithmetic code. There are two prediction modes in ALS: linear predictive coding (LPC) and cascaded RLS-LMS. As the developer of the RLS-LMS prediction, we present this technology in this paper. In RLS-LMS prediction, the input audio samples go through the cascaded DPCM, RLS, and LMS predictors, whose output predictions are linearly combined to generate a prediction for the current input sample. Through MPLG testings, it has been found that ALS with RLS-LMS prediction provides the best lossless compression ratio compared with SLS, ALS with LPC, and several non-MPEG codecs Susanto Rahardja, Xiao Lin 0001, Rongshan Yu, Pasi Fränti |
ICASSP (5) | 4 |
| 2006 | Perceptually Enhanced Bit-Plane Coding for Scalable AudioabstractThe MPEG-4 scalable to lossless (SLS) audio coding is recently being developed to provide a unified solution for high-compression perceptual audio coding and high-quality lossless audio coding. SLS provides efficient fine granular scalable (FGS) coding from AAC core layer to lossless, and achieves reasonable perceptual quality at its scalable coding range using a sequential bit-plane scanning method, which minimizes the audio distortion according to the spectral shape of the core layer quantization errors. In this paper, it is shown that the perceptual quality performance of SLS at intermediate rates can be further improved by incorporating psycho acoustic model into the bit-plane coding process. In addition, it is also found that such an improvement can be achieved by slightly tweaking the original bit-plane coding process of SLS and hence preserving its nice features such as compatibility to lossless coding and low complexity Rongshan Yu, Susanto Rahardja |
ICME | 1 |
| 2006 | A fine granular scalable to lossless audio coderabstractThis paper presents Advanced Audio Zip (AAZ), a fine grained scalable to lossless (SLS) audio coder that has recently been adopted as the reference model for MPEG-4 audio SLS work. AAZ integrates the functionalities of high-compression perceptual audio coding, fine granular scalable audio coding, and lossless audio coding in a single framework, and simultaneously provides backward compatibility to MPEG-4 Advanced Audio Coding (AAC). AAZ provides the fine granular bit-rate scalability from lossy to lossless coding, and such a scalability is achieved in a perceptually meaningful way, i.e., better perceptual quality at higher bit-rates. Despite its abundant functionalities, AAZ only introduces negligible overhead in terms of lossless compression performance compared with a nonscalable, lossless only audio coder. As a result, AAZ provides a universal yet efficient solution for digital audio applications such as audio archiving, network audio streaming, portable audio playing, and music downloading which were previously catered for by several different audio coding technologies, and eliminates the need for any transcoding system to facilitate sharing of digital audio contents across these application domains. Rongshan Yu, Susanto Rahardja, Xiao Lin 0001, Chi Chung Ko |
IEEE Trans. Speech Audio Process. | 1 |
| 2005 | Improving coding efficiency for MPEG-4 Audio Scalable Lossless codingabstractThe recently introduced MPEG standard for lossless audio coding, MPEG-4 Audio Scalable to Lossless (SLS) coding technology, provides a universal audio format that integrates the functionalities of lossy audio coding, lossless audio coding and fine granular scalable audio coding in a single framework. We propose two coding methods that improve the coding efficiency of SLS, namely, a context-based arithmetic code (CBAC) method and a low energy mode code method. These two coding methods work harmonically with the current SLS framework and preserve all its desirable features, such as fine granular scalability, while successfully improving its lossless compression ratio performance. Rongshan Yu, Xiao Lin 0001, Susanto Rahardja, Chi Chung Ko |
ICASSP (3) | 1 |
| 2005 | A scalable watermarking scheme for the scalable audio coderabstractIn this paper, we describe a scalable (i.e., lossy-to-lossless) watermarking scheme which overcomes the problem of non-invertible distortion introduced by the watermark signal. The scheme is based on a standardized scalable audio coder (R.S. Yu, et al, 2004) -as a result, the embedded watermark inherits the scalability of the audio coder. We elaborate how the scalability can be used to realize the recovery of the lossless audio signal after watermark embedding. The experimental results demonstrate the validity of the proposed watermarking scheme in terms of robustness, data expansion and perceptual quality. Zhi Li 0001, Qibin Sun, Yong Lian 0001, Rongshan Yu |
ICC | 4 |
| 2005 | A New Bit-Plane Entropy Coder for Scalable Image CodingabstractCompression ratio and computational complexity are two major factors for a successful image coder. By exploring the Laplacian distribution of the wavelet coefficients, a new bit plane entropy coder is proposed in this paper. Compared with the state-of-the-art JPEG2000 entropy coder (EBCOT), the proposed coder achieves a 0.75% better loss less performance for 5 level 5/3 wavelet decomposition at block size 64 £ 64 and 2.56% at block size 16 £ 16. Experimental results also show PSNR improvements of about 0.13dB at 1bpp and 0.25dB at 2bpp on average for lossy compression. However, the gain in coding performance is not based on increasing computational complexity but in stead a reduction by using a static arithmetic coder which avoids complicated adaptive procedure. Rongshan Yu, Qibin Sun, Lawrence Wai-Choong Wong |
ICME | 2 |
| 2005 | Study on Rounding Errors of INTMDCT in Perceptual Audio CodingabstractWith the proliferation of broadband access and continuous decline of storage prize per gigabyte, there has been an increasing demand of audio solution that provides high sampling rate and high resolution. Lossless audio is undoubtedly the ultimate solution. In response to this demand, MPEG issued a call for proposal soliciting technology contributions that provides a state-of-art solution. At the technology end, lossless compression requires the usage of integer transform. The integer modified discrete cosine transform (IntMDCT) has been adopted in MPEG-4 scalable to lossless (SLS) coding to enable this efficient lossless operation. Because of rounding operations, rounding errors introduced by IntMDCT exist during the whole coding process. With the SLS having capability of using operations that spreads over the bitrate spectrum which ranges from lossy to lossless, it is of interest to study the effect of rounding errors in IntMDCT for operation of SLS in lossy mode. This paper analyzes the contributions of noise due to these errors. It is found that the noise introduced by rounding operations of IntMDCT does not affect the perceptual quality of the coded audio under any circumstances. As such, it concludes that the MDCT and IntMDCT filterbanks are interchangeable at lossy bitrate. With the fact that SLS uses both MDCT and IntMDCT, the finding in this paper suggests the possibility of using only IntMDCT filterbank. Rongshan Yu, Soo Ngee Koh |
ISM | 2 |
| 2005 | MPEG-4 Scalable to Lossless Audio Coding - Emerging International Standard for Digital Audio CompressionabstractRecently, with the advance of network and storage technologies, it is becoming realistic that people will enjoy high sampling rate, high resolution audio contents with lossless quality. Envisioning of such a need, the international standardization body MPEG has recently introduced a scalable tool for lossless audio coding, namely, MPEG-4 Audio Scalable to Lossless (SLS) coding. MPEG-4 SLS integrates the functionalities of lossless audio coding, perceptual audio coding, and fine granular scalable audio coding in a single framework; meanwhile it provides backward compatibility to MPEG-4 Advanced Audio Coding (AAC) at the bit-stream level. This new tool, in combination with the existing MPEG audio toolset, provides a universal digital audio format that can be used in a variety of application domains such as professional audio, Internet music, consumer electronics, broadcasting Rongshan Yu, Xiao Lin 0001, Susanto Rahardja |
MMSP | 1 |
| 2004 | A fast algorithm of integer MDCT for lossless audio codingabstractA new fast algorithm to implement integer modified discrete cosine transform (IntMDCT) is proposed. It is shown that the total rounding operations required for this algorithm are only 2.5N, where N is the block size. As a result, its approximation error is far less than that of directly converted integer transforms. At the same time, the complexity is greatly reduced, which results in improved performance of a lossless audio coding system employing the IntMDCT in terms of compression ratio and complexity costs. Susanto Rahardja, Rongshan Yu, Xiao Lin 0001 |
ICASSP (4) | 3 |
| 2004 | A scalable lossy to lossless audio coder for MPEG-4 lossless audio codingabstractIn this paper, we present Advanced Audio Zip (AAZ), a scalable lossless audio coding technology that was recently selected as the reference model for MPEG audio scalable lossless coding (SLS) work. AAZ provides excellent compression performance while delivering fine grain bit-rate scalability from lossy to lossless coding. Moreover, AAZ provides backward compatibility to the MPEG advanced audio coding (AAC) system by embedding an AAC compliant bit-stream into the lossless bit-stream. As a result, AAZ serves as a universal coding solution with functionalities that were previously offered by several distinct audio coding technologies such as lossless audio coding, perceptual audio coding, or scalable audio coding; and maximizes the interchangeability for digital audio contents migrating among these application domains. Rongshan Yu, Xiao Lin 0001, Susanto Rahardja, Chi Chong Ko |
ICASSP (3) | 1 |
| 2004 | A statistics study of the MDCT coefficient distribution for audioabstractThe modified discrete cosine transform (MDCT) has been widely used in many transform audio coding algorithms such as MPEG-1/2 layer III (mp3), MPEG-2/4 AAC, Dolby AC2/AC3, and numerous experimental audio coding algorithms. In this paper, we study the probabilistic distribution properties of the MDCT coefficient for audio signals. It is shown that the generalized Gaussian function with distribution parameter r=0.5 or r=1 (Laplacian) provides a good approximation to the distributions of MDCT coefficients for a variety of audio signals. Results from our study also show that although the distribution of these coefficients is not strictly Laplacian, the divergence between them is in fact very small. Therefore, it leads to only marginal redundancy if these coefficients are simply coded with some low-complexity codes designed for Laplacian sources. Rongshan Yu, Xiao Lin 0001, Susanto Rahardja, Chi Chung Ko |
ICME | 1 |
| 2003 | Bit-plane Golomb coding for sources with Laplacian distributionsabstractThis paper presents a bit-plane coding algorithm for Laplacian distributed sources that are commonly encountered in signal compression applications. By exploiting the statistical characteristics of the sources, the proposed algorithm achieves a rate-distortion performance that is essentially comparable to an optimal nonscalable scalar quantizer, while at the same time operates at a complexity level suitable for most practical implementations. Rongshan Yu, Chi Chong Ko, Susanto Rahardja, Xiao Lin 0001 |
ICASSP (4) | 1 |
| 2003 | Video streaming on embedded devices through GPRS networkabstractWe introduce a PDA-based live video streaming system on GPRS network based on MPEG-4 video compression standard. Due to the limited computational resources of PDA, all the key modules of MPEG-4 codec are efficiently implemented and optimized such as multithreading, buffer design, wireless communication, encoder and decoder. Several novel techniques are developed in the coding, streaming as well as the post- processing stages of the system. Keng-Pang Lim, Dajun Wu, Si Wu 0004, Susanto Rahardja, Xiao Lin 0001, Lijun Jiang, Rongshan Yu, Feng Pan 0002, Zhengguo Li, Susu Yao, Genan Feng, Chi Chung Ko |
ICME | 7 |
| 2003 | An adaptive rate control algorithm for video coding over personal digital assistants (PDA)abstractWith the recent development of third-generation communication technologies, encoding live video using a PDA and sharing it among friends has become a reality. However, the embedded processor inside a PDA is still not powerful enough and there are two major hurdles to overcome: (1) video coding needs to meet the rigorous constraint of the available computation capacity of a PDA; (2) In a PDA the computing power allocated to video coding may vary drastically (in bursts). In this paper, a new adaptive rate control algorithm is proposed for video coding over a PDA. This adaptive rate control scheme takes into account the time constraint of a PDA, and its bit allocation depends not only on the available data bits, but more importantly, on the available coding time. Experimental results show that, compared to the existing rate control scheme, the new algorithm can always achieve the maximum frame rate, maximize the utilization of the available bandwidth and computing power, increase the average PSNR, and improve the subjective perceptual quality of the reconstructed video. Feng Pan 0002, Zhengguo Li, Keng-Pang Lim, Dajun Wu, Rongshan Yu, Genan Feng |
ICME | 5 |
| 2003 | A fine granular scalable perceptually lossy and lossless audio coderabstractThis paper presents advanced audio zip (AAZ), an audio codec that provides the fine granular bit-rate scalability from lossy to lossless coding. Perceptually embedded coding principle is employed in AAZ to provide lossy reconstruction with optimal perceptual quality at intermediate bit-rates. AAZ also provides the backward compatibility where the lossless bit-stream embeds a compliant MPEG-4 AAC bit-stream. Rongshan Yu, Xiao Lin 0001, Susanto Rahardja, Chi Chung Ko |
ICME | 1 |
| 2003 | Lossless compression of digital audio using cascaded RLS-LMS predictionabstractThis paper proposes a cascaded RLS-LMS predictor for lossless audio coding. In this proposed predictor, a high-order LMS predictor is employed to model the ample tonal and harmonic components of the audio signal for optimal prediction gain performance. To solve the slow convergence problem of the LMS algorithm with colored inputs, a low-order RLS predictor is cascaded prior to the LMS predictor to remove the spectral tilt of the audio signal. This cascaded RLS-LMS structure effectively mitigates the slow convergence problem of the LMS algorithm and provides superior prediction gain performance compared with the conventional LMS predictor, resulting in a better overall compression performance. Rongshan Yu, Chi Chung Ko |
IEEE Trans. Speech Audio Process. | 1 |
| 2002 | A warped linear-prediction-based subband audio coding algorithmabstractA novel audio coding algorithm is proposed where the warped-linear prediction (WLP) technique is employed to construct a perceptual pre- and post-filter for subband audio coding. A modified signal-to-mask ratio (SMR) calculation is given for subband coding of the WLP residuals of audio signals. The concept of perceptual entropy (PE) is extended to subband coding, resulting in the subband perceptual entropy (SPE), which gives a short time estimate of the lowest possible bit rate for transparent subband audio coding. Two WLP models with frequency responses approximating the spectral shape of the masking threshold are investigated and it is found that the residual signals of both models contain less SPE compared with that of the original audio signals. Subjective tests show that the proposed audio codec operating at 56 kbps has a perceptual quality comparable to MPEG-1 audio Layer II operating at 64 kbps. Rongshan Yu, Chi Chung Ko |
IEEE Trans. Speech Audio Process. | 1 |
| 2001 | Improving Quality Of Low Bit Rate Audio Coding By Using Short-Time Spectral AttenuationabstractIn this paper we present a novel postfiltering algorithm based on the Short-Time Spectral Attenuation (STSA) for enhancing the perceptual quality of audio signals coded at low bit rates. In the proposed algorithm, the quantization noise of coded audio signal is estimated and subtracted from the coded signal by using STSA. For more accurate estimation of the noise spectrum, dithering technique is applied to the quantizer designs. Results show that the proposed algorithm effectively reduces the quantization noise level in the coded audio, resulting in better perceptual quality. The proposed postfiltering algorithm can be implemented in cooperation with existing or future audio coding schemes to improve their performance at low bit rates. Rongshan Yu |
ICME | 1 |