EDBT 2026 Demo / reviewers in the wild / expert
Shihao Shen
dblp:99/7970
· DBLP profile ↗
22ranked-venue papers
10as first author
18since 2021 · last 2026
0000-0003-1012-1028ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 13 · 6 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 1 since 2021Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 2 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Co-Design Framework for Container Deployment in Mobile Edge Computing NetworksabstractWith the rapid advancement of mobile technologies, including self-driving cars and drones, the deployment of mobile software has become increasingly complex. In this context, virtualization plays a pivotal role by simplifying service deployment through containers and enabling container orchestration plat forms to efficiently manage an expanding number of container clusters. This is achieved by leveraging standardized interfaces and minimizing resource optimization overhead. However, the use of distributed servers in mobile edge clusters introduces several challenges, such as bandwidth limitations, network performance fluctuations, and resource constraints, which complicate deployment in these dynamic and resource-constrained environments. In this paper, we rethink the layer-based structure, a fundamental container design, and analyze the challenges and potential of real edge platform traces. Consequently, we propose BREAK, an acceleration middleware for efficient container deployment. With the primary insight of enhancing layer-reuse and deriving benefits from it, we develop a co-design approach centered on layer structure for efficient deployment, ensuring backward compatibility: (i) a container image refactoring solution that optimizes efficiency while preserving the stack-of-layers structure, (ii) distributed shared layer-stack caches, dynamically optimized for collaborative container deployment among mobile edge clusters, (iii) a customized Kubernetes (K8s) scheduler extending awareness of network performance, disk space, and container layer cache for container placement, and (iv) a tailored storage-driver of the standard container runtime for efficient layer extraction. Results indicate that BREAK accelerates the deployment process by up to 2.1× and reduces redundant image size by up to 3.11× compared to the state-of-the-art approach. Shihao Shen, Yicheng Feng, Xiaoxu Ren, Xiaofei Wang 0001, Qiao Xiang, Hong Xu 0001, Chenren Xu |
IEEE Trans. Mob. Comput. | 1 |
| 2025 | CORES: A Collaborative Orchestration and Extraction Strategy for Image Layers in AI ServicesabstractAs the rapid development of artificial intelligence (AI) and large language models (LLM), how to deploy related applications onto computing nodes has become a hot topic, and containerized service provides an excellent approach for this. The most time-consuming step of this approach is image extraction, the procedure of decompressing all layers of image package downloaded from remote image registry. Therefore, achieving fast image extracting is crucial for the efficient deployment of AI services. In this paper, we introduce a collaborative orchestration and extraction strategy, CORES. Firstly, we eliminate the dependencies among image layers, which impede unordered extraction of image layers. Based on this, we model the image extraction as a mixed integer linear programming (MILP) problem, aiming to minimize total extraction time. Then we use improved Benders decomposition to iteratively obtain a near-optimal solution with lower time complexity. Extensive experiments conducted on the real system validate the superior performance of our strategy. Compared with our closest baseline LOPO, CORES reduces the average image extracting time by 19.60%, significantly enhancing the efficiency of AI service deployment. Mingjun Cai, Shihao Shen, Xiaofei Wang 0001, Cheng Zhang 0007, Chao Qiu |
GLOBECOM | 2 |
| 2025 | Vodcm: Value-Optimized Distributed Caching Mechanism for Containerized Aigc Services in Edge-Cloud EnvironmentsabstractWith the rise of AI-Generated Content (AIGC) services, deployments within edge-cloud environments are becoming increasingly prevalent. Containerization offers resource isolation, lightweight deployment, and portability, making it a suitable technology for AIGC services. However, deploying AIGC services often requires large container images, leading to high deployment latency and bandwidth consumption. Based on real-world trace analysis showing the long-tail effect, where a few popular images account for the majority of requests, there is strong potential for optimizing caching mechanisms. This pattern can result in frequent cache misses and increased bandwidth consumption, especially under heavy load. In this paper, we propose a ValueOptimized Distributed Caching Mechanism (VODCM), which dynamically optimizes caching policies through a value-driven framework combined with deep reinforcement learning (DRL). VODCM prioritizes high-value images based on access frequency, layer size, and network latency, significantly improving cache hit rates and reducing network overhead. Preliminary evaluations show that VODCM enhances cache efficiency and reduces network and resource demands, offering an effective solution for AIGC image management in edge-cloud environments. Shihao Shen, Chao Qiu, Xiaofei Wang 0001, Tao Luo 0010, Cheng Zhang 0007 |
ICC | 2 |
| 2025 | RESCUER: QoS-Aware Service Rescheduling in Serverless Crowdsourced Edge Cloud ClustersabstractCrowdsourced edge-cloud clusters utilize idle third-party resources to provide cost-effective, scalable environments. This decentralized model reduces capital expenditures and carbon footprints but introduces hardware instability, as servers may unpredictably join or leave, affecting service availability. While integrating serverless computing enables automated management to lower operational costs, the dynamic nature of server availability still significantly impacts service quality. In this paper, we present RESCUER, a QoS-aware service rescheduling framework for serverless crowdsourced edge-cloud clusters. Partnering with a real-world provider, we collected data from over 10,000 edge servers over 400+ days, enabling a detailed analysis of server availability patterns. Based on these insights, we propose a predictive algorithm that forecasts the future online duration of each node, assigning them labels based on their predicted availability. Additionally, we introduce a rescheduling algorithm that combines these labels with node resources, latency constraints, and other factors to select the most suitable node for deploying containers. Evaluations on the real-world trace show that RESCUER significantly improves service availability and system efficiency compared to existing methods. Shihao Shen, Chao Qiu, Xiaofei Wang 0001 |
MASS | 1 |
| 2025 | Task Allocation With Geography-Context-Capacity Awareness in Distributed Burstable Billing Edge-Cloud SystemsabstractThe new real-time interactive services, such as virtual and augmented reality, demand significantly higher network bandwidth and quality, which the traditional centralized cloud struggles to meet. In addition, centralized optimization management becomes inefficient as the scale of the scene continues to expand. In response, edge cloud systems have emerged, but distributed geographic locations, burstable billing business models, and large numbers of servers in large-scale scenarios pose new challenges for resource management. In this article, we proposeGeoCC, a novel strategy to save bandwidth overhead in burstable billing edge cloud systems.GeoCCaddresses challenges through a dual approach. First, a geography-aware graph construction and partitioning algorithm is used to organize server resources, and a large number of servers are reasonably divided into multiple server pools for parallel processing. Second, it introduces an enhanced burstable billing optimization mechanism that considers contextual factors and adaptive bandwidth capacity. Experiments based on real data from an edge cloud operator demonstrate the effectiveness ofGeoCC. Compared with the baseline,GeoCCcan effectively reduce bandwidth peaks, decreasing bandwidth costs by an average of 28.30% and up to 81.83% at the 95th percentile billing. Shihao Shen, Chenfei Gu, Yuanze Li, Chao Qiu, Xiaofei Wang 0001, Rui Tan 0001, Cheng Zhang 0019 |
IEEE Trans. Serv. Comput. | 1 |
| 2024 | Kubernetes Scheduling Design Based on Imitation Learning in Edge Cloud ScenariosabstractWith the rapid increase in user scale and the explosive rise of emerging applications, the contradiction between heavy load pressure and excellent network performance is becoming increasingly prominent, and task processing is gradually shifting towards the edge of the network. However, the resources of edge networks are limited, making it difficult to meet the huge computing and storage needs, and managing and allocating edge nodes is also a huge challenge. The Kubernetes (K8S) framework for deploying and orchestrating containerized applications provides a solution for this. How to improve the adaptability of K8S in edge networks, meet the demand of services for heterogeneous resources, and train decision models with better performance using limited datasets has become an urgent problem to be solved. Based on the above issues, we propose a distributed service migration architecture for multi-user access, and design a service migration algorithm based on imitation learning to achieve resource combination optimization and reduce the impact of insufficient data on model training. Design agent models based on diffusion models to accelerate model convergence and avoid the increase in training costs caused by constantly updating agent models. Our results show that the efficiency of the expert model is 92.0%, and the learning process of the agent model can converge within 100 training cycles with an accuracy of 97.89%. The service processing delay, throughput rate, and model convergence are all significantly better than those of classical algorithms. Ziyi Sang, Mingjun Cai, Shihao Shen, Cheng Zhang 0007, Xiaofei Wang 0001, Chao Qiu |
GLOBECOM | 3 |
| 2024 | BREAK: A Holistic Approach for Efficient Container Deployment among Edge CloudsabstractContainer technology has revolutionized service deployment, offering streamlined processes and enabling container orchestration platforms to manage a growing number of container clusters. However, the deployment of containers in distributed edge clusters presents challenges due to their unique characteristics, such as bandwidth limitations and resource constraints. Existing approaches designed for cloud environments often fall short in addressing the specific requirements of edge computing. Additionally, very few edge-oriented solutions explore fundamental changes to the container design, resulting in difficulties achieving backward compatibility.In this paper, we reevaluate the fundamental layer-based structure of containers. We identify that the proliferation of redundant files and operations within image layers hinders efficient container deployment. Drawing upon the crucial insight of enhancing layer reuse and extracting benefits from it, we introduce BREAK, a holistic approach centered on layer structure throughout the entire container deployment pipeline, ensuring backward compatibility. BREAK refactors image layers and proposes an edge-oriented cache solution to enable ubiquitous and shared layers. Moreover, it addresses the complete deployment pipeline by introducing a customized scheduler and a tailored storage driver. Our results demonstrate that BREAK accelerates the deployment process by up to 2.1× and reduces redundant image size by up to 3.11× compared to state-of-the-art approaches. Yicheng Feng, Shihao Shen, Xiaofei Wang 0001, Qiao Xiang, Hong Xu 0001, Chenren Xu |
INFOCOM | 2 |
| 2024 | EdgeOptimizer: A programmable containerized scheduler of time-critical tasks in Kubernetes-based edge-cloud clusters
Yufei Qiao, Shihao Shen, Cheng Zhang 0007, Tie Qiu 0001, Xiaofei Wang 0001 |
Future Gener. Comput. Syst. | 2 |
| 2024 | Tango: Harmonious Optimization for Mixed Services in Kubernetes-Based Edge CloudsabstractDeploying Latency-Critical (LC) services and Best-Effort (BE) services together is expected to improve resource utilization in edge clouds. However, co-locating LC and BE services on edge clouds presents unique challenges. Unlike cloud datacenters, edge clouds are heterogeneous, resource-constrained, and geographically distributed, leading to fiercer competition for resources and greater difficulty in balancing fluctuating co-located workloads. Due to the lack of consideration for the characteristics of edge environments, previous solutions designed for cloud datacenters are no longer applicable. To address these challenges, we introduceTango, a harmonious scheduling framework forKubernetes-based edge cloud systems with mixed services.Tangoincorporates novel components and mechanisms for elastic resource allocation on the edge, as well as two traffic scheduling algorithms that efficiently manage distributed edge resources.Tangofosters harmony not only by supporting compatible mixed services but also by offering collaborative solutions that complement each other. Based on a non-intrusive design forKubernetes,Tangofurther enhances it with automatic scaling and traffic scheduling capabilities. Compared to state-of-the-art approaches, experiments on large-scale hybrid edge clouds, driven by real workload traces, show thatTangoimproves system resource utilization by 36.9%, QoS-guarantee satisfaction rate by 11.3%, and throughput by 47.6%. Shihao Shen, Yicheng Feng, Mengwei Xu 0001, Yuanming Ren, Xiaofei Wang 0001, Victor C. M. Leung |
IEEE Trans. Serv. Comput. | 1 |
| 2023 | Quicklayer: A Layer-Stack-Oriented Accelerating Middleware for Fast Deployment in Edge CloudsabstractContainers are gaining popularity in edge computing due to their standardization and low overhead. This trend has brought new technologies such as container engines and container orchestration platforms (COPs). However, fast and effective container deployment remains a challenge, especially at the edge. Prior work, which was designed for cloud datacenters, is no longer suitable for container deployment in edge clouds due to bandwidth limitations, fluctuating network performance, resource constraints, and geo-distributed organization. These edge features make rapid deployment on the edge difficult. Additionally, integrating with COPs is crucial for successful deployment. Yicheng Feng, Shihao Shen, Cheng Zhang 0019, Xiaofei Wang 0001 |
APNet | 2 |
| 2023 | Tango: Harmonious Management and Scheduling for Mixed Services Co-located among Distributed Edge-CloudsabstractCo-locating Latency-Critical (LC) and Best-Effort (BE) services in edge-clouds is expected to enhance resource utilization. However, this mixed deployment encounters unique challenges. Edge-clouds are heterogeneous, distributed, and resource-constrained, leading to intense competition for edge resources, making it challenging to balance fluctuating co-located workloads. Previous works in cloud datacenters are no longer applicable since they do not consider the unique nature of edges. Although very few works explicitly provide specific schemes for edge workload co-location, these solutions fail to address the major challenges simultaneously. Yicheng Feng, Shihao Shen, Mengwei Xu 0001, Yuanming Ren, Xiaofei Wang 0001, Victor C. M. Leung |
ICPP | 2 |
| 2023 | DytanVO: Joint Refinement of Visual Odometry and Motion Segmentation in Dynamic EnvironmentsabstractLearning-based visual odometry (VO) algorithms achieve remarkable performance on common static scenes, benefiting from high-capacity models and massive annotated data, but tend to fail in dynamic, populated environments. Semantic segmentation is largely used to discard dynamic associations before estimating camera motions but at the cost of discarding static features and is hard to scale up to unseen categories. In this paper, we leverage the mutual dependence between camera ego-motion and motion segmentation and show that both can be jointly refined in a single learning-based framework. In particular, we present DytanVO, the first supervised learning-based VO method that deals with dynamic environments. It takes two consecutive monocular frames in real-time and predicts camera ego-motion in an iterative fashion. Our method achieves an average improvement of 27.7% in ATE over state-of-the-art VO solutions in real-world dynamic environments, and even performs competitively among dynamic visual SLAM systems which optimize the trajectory on the backend. Experiments on plentiful unseen environments also demonstrate our method's generalizability. Shihao Shen, Sebastian A. Scherer |
ICRA | 1 |
| 2023 | A Holistic QoS View of Crowdsourced Edge Cloud PlatformabstractEdge clouds have become a de-facto paradigm to deliver low and stable networks to delay-critical applications such as web services and AR/VR. A unique form of edge clouds is those crowdsourced from third parties, e.g., idle PCs or workstations. Such crowdsourced edge platforms can better sink computations closer to users, reduce the purchase cost, and eliminates the carbon generated during manufacturing. Yet, they also face the challenge of out-of-control hardware, e.g., a server dropping in/out anytime. In this paper, we perform the first-of-its-kind measurement of Quality of Service (QoS) for a large-scale crowdsourced edge platform, which covers over 10,000 edge servers, 100,000 users and 10,000,000 user requests. The measurement takes a holistic QoS view: (1) First, we look at how much hardware resources are provided by edge servers, how much time they are available for service deployment, and what are the major abnormal behaviors. (2) Second, we analyze the factors affecting service stability and quantify the resource utilization pattern of containerized services hosted on those edge servers. (3) Third, we investigate the spatial and temporal features of user requests handled by the platform. Many useful and somehow surprising findings are obtained through the above measurements. We also derive insightful implications that could help edge platforms and edge applications to better deliver their services to users. Shihao Shen, Yicheng Feng, Mengwei Xu 0001, Cheng Zhang 0007, Xiaofei Wang 0001, Victor C. M. Leung |
IWQoS | 1 |
| 2023 | EdgeMatrix: A Resource-Redefined Scheduling Framework for SLA-Guaranteed Multi-Tier Edge-Cloud Computing SystemsabstractWith the development of networking technology, the computing system has evolved towards the multi-tier paradigm gradually. However, challenges, such as multi-resource heterogeneity of devices, resource competition of services, and networked system dynamics, make it difficult to guarantee service-level agreement (SLA) for the applications. In this paper, we propose a multi-tier edge-cloud computing framework, EdgeMatrix, to maximize the throughput of the system while guaranteeing different SLA priorities. First, in order to reduce the impact of physical resource heterogeneity, EdgeMatrix introduces the Networked Multi-agent Actor-Critic (NMAC) algorithm to re-define physical resources with the same quality of service as logically isolated resource units and combinations, i.e., cells and channels. In addition, a multi-task mechanism is designed in EdgeMatrix to solve the problem of Joint Service Orchestration and Request Dispatch (JSORD) for matching the requests and services, which can significantly reduce the optimization runtime. For integrating above two algorithms, EdgeMatrix is designed with two time-scales, i.e., coordinating services and resources at the larger time-scale, and dispatching requests at the smaller time-scale. Realistic trace-based experiments proves that the overall throughput of EdgeMatrix is 36.7% better than that of the closest baseline, while the SLA priorities are guaranteed still. Shihao Shen, Yuanming Ren, Yanli Ju, Xiaofei Wang 0001, Victor C. M. Leung |
IEEE J. Sel. Areas Commun. | 1 |
| 2023 | Collaborative Learning-Based Scheduling for Kubernetes-Oriented Edge-Cloud NetworkabstractKubernetes (k8s) has the potential to coordinate distributed edge resources and centralized cloud resources, but currently lacks a specialized scheduling framework for edge-cloud networks. Besides, the hierarchical distribution of heterogeneous resources makes the modeling and scheduling of k8s-oriented edge-cloud network particularly challenging. In this paper, we introduce KaiS, a learning-based scheduling framework for such edge-cloud network to improve the long-term throughput rate of request processing. First, we design a coordinated multiagent actor-critic algorithm to cater to decentralized request dispatch and dynamic dispatch spaces within the edge cluster. Second, for diverse system scales and structures, we use graph neural networks to embed system state information, and combine the embedding results with multiple policy networks to reduce the orchestration dimensionality by stepwise scheduling. Finally, we adopt a two-time-scale scheduling mechanism to harmonize request dispatch and service orchestration, and present the implementation design of deploying the above algorithms compatible with native k8s components. Experiments using real workload traces show that KaiS can successfully learn appropriate scheduling policies, irrespective of request arrival patterns and system scales. Moreover, KaiS can enhance the average system throughput rate by 15.9% while reducing scheduling cost by 38.4% compared to baselines. Shihao Shen, Yiwen Han, Xiaofei Wang 0001, Shiqiang Wang 0001, Victor C. M. Leung |
IEEE/ACM Trans. Netw. | 1 |
| 2023 | A large-scale holistic measurement of crowdsourced edge cloud platform
Yicheng Feng, Shihao Shen, Mengwei Xu 0001, Cheng Zhang 0007, Xin Wang 0030, Xiaofei Wang 0001, Victor C. M. Leung |
World Wide Web (WWW) | 2 |
| 2022 | EdgeMatrix: A Resources Redefined Edge-Cloud System for Prioritized ServicesabstractThe edge-cloud system has the potential to com-bine the advantages of heterogeneous devices and truly realize ubiquitous computing. However, for service providers to guar-antee the Service-Level-Agreement (SLA) priorities, the complex networked environment brings inherent challenges such as multi-resource heterogeneity, resource competition, and networked sys-tem dynamics. In this paper, we design a framework for the edge-cloud system, namely EdgeMatrix, to maximize the throughput while guaranteeing various SLA priorities. First, EdgeMatrix introduces Networked Multi-agent Actor-Critic (NMAC) algorithm to redefines physical resources as logically isolated resource combinations, i.e., resource cells. Then, we use a clustering algorithm to group the cells with similar characteristics into various sets, i.e., resource channels, for different channels can offer different SLA guarantees. Besides, we design a multi-task mechanism to solve the problem of joint service orchestration and request dispatch (JSORD) among edge-cloud clusters, significantly reducing the runtime than traditional methods. To ensure stability, EdgeMatrix adopts a two-time-scale framework, i.e., coordinating resources and services at the large time scale and dispatching requests at the small time scale. The real trace-based experimental results verify that EdgeMatrix can improve system throughput in complex networked environments, reduce SLA violations, and significantly reduce the runtime than traditional methods. Yuanming Ren, Shihao Shen, Yanli Ju, Xiaofei Wang 0001, Victor C. M. Leung |
INFOCOM | 2 |
| 2021 | Tailored Learning-Based Scheduling for Kubernetes-Oriented Edge-Cloud SystemabstractKubernetes (k8s) has the potential to merge the distributed edge and the cloud but lacks a scheduling framework specifically for edge-cloud systems. Besides, the hierarchical distribution of heterogeneous resources and the complex dependencies among requests and resources make the modeling and scheduling of k8s-oriented edge-cloud systems particularly sophisticated. In this paper, we introduce KaiS, a learning-based scheduling framework for such edge-cloud systems to improve the long-term throughput rate of request processing. First, we design a coordinated multi-agent actor-critic algorithm to cater to decentralized request dispatch and dynamic dispatch spaces within the edge cluster. Second, for diverse system scales and structures, we use graph neural networks to embed system state information, and combine the embedding results with multiple policy networks to reduce the orchestration dimensionality by stepwise scheduling. Finally, we adopt a two-time-scale scheduling mechanism to harmonize request dispatch and service orchestration, and present the implementation design of deploying the above algorithms compatible with native k8s components. Experiments using real workload traces show that KaiS can successfully learn appropriate scheduling policies, irrespective of request arrival patterns and system scales. Moreover, KaiS can enhance the average system throughput rate by 14.3% while reducing scheduling cost by 34.7% compared to baselines. Yiwen Han, Shihao Shen, Xiaofei Wang 0001, Shiqiang Wang 0001, Victor C. M. Leung |
INFOCOM | 2 |
| 2020 | Computation Offloading with Multiple Agents in Edge-Computing-Supported IoTabstractWith the development of the Internet of Things (IoT) and the birth of various new IoT devices, the capacity of massive IoT devices is facing challenges. Fortunately, edge computing can optimize problems such as delay and connectivity by offloading part of the computational tasks to edge nodes close to the data source. Using this feature, IoT devices can save more resources while still maintaining the quality of service. However, since computation offloading decisions concern joint and complex resource management, we use multiple Deep Reinforcement Learning (DRL) agents deployed on IoT devices to guide their own decisions. Besides, Federated Learning (FL) is utilized to train DRL agents in a distributed fashion, aiming to make the DRL-based decision making practical and further decrease the transmission cost between IoT devices and Edge Nodes. In this article, we first study the problem of computation offloading optimization and prove the problem is an NP-hard problem. Then, based on DRL and FL, we propose an offloading algorithm that is different from the traditional method. Finally, we studied the effects of various parameters on the performance of the algorithm and verified the effectiveness of both the DRL and FL in the IoT system. Shihao Shen, Yiwen Han, Xiaofei Wang 0001, Yan Wang 0108 |
ACM Trans. Sens. Networks | 1 |
| 2017 | rMATS-DVR: rMATS discovery of differential variants in RNAabstractMOTIVATION: RNA sequences of a gene can have single nucleotide variants (SNVs) due to single nucleotide polymorphisms (SNPs) in the genome, or RNA editing events within the RNA. By comparing RNA-seq data of a given cell type before and after a specific perturbation, we can detect and quantify SNVs in the RNA and discover SNVs with altered frequencies between distinct cellular states. Such differential variants in RNA (DVRs) may reflect allele-specific changes in gene expression or RNA processing, as well as changes in RNA editing in response to cellular perturbations or stimuli. RESULTS: We have developed rMATS-DVR, a convenient and user-friendly software program to streamline the discovery of DVRs between two RNA-seq sample groups with replicates. rMATS-DVR combines a stringent GATK-based pipeline for calling SNVs including SNPs and RNA editing events in RNA-seq reads, with our rigorous rMATS statistical model for identifying differential isoform ratios using RNA-seq sequence count data with replicates. We applied rMATS-DVR to RNA-seq data of the human chronic myeloid leukemia cell line K562 in response to shRNA knockdown of the RNA editing enzyme ADAR1. rMATS-DVR discovered 1372 significant DVRs between knockdown and control. These DVRs encompassed known SNPs and RNA editing sites as well as novel SNVs, with the majority of DVRs corresponding to known RNA editing sites repressed after ADAR1 knockdown. AVAILABILITY AND IMPLEMENTATION: rMATS-DVR is at https://github.com/Xinglab/rMATS-DVR . CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Shihao Shen, Yi Xing |
Bioinform. | 3 |
| 2014 | Similarity of markers identified from cancer gene expression studies: observations from GEOabstractGene expression profiling has been extensively conducted in cancer research. The analysis of multiple independent cancer gene expression datasets may provide additional information and complement single-dataset analysis. In this study, we conduct multi-dataset analysis and are interested in evaluating the similarity of cancer-associated genes identified from different datasets. The first objective of this study is to briefly review some statistical methods that can be used for such evaluation. Both marginal analysis and joint analysis methods are reviewed. The second objective is to apply those methods to 26 Gene Expression Omnibus (GEO) datasets on five types of cancers. Our analysis suggests that for the same cancer, the marker identification results may vary significantly across datasets, and different datasets share few common genes. In addition, datasets on different cancers share few common genes. The shared genetic basis of datasets on the same or different cancers, which has been suggested in the literature, is not observed in the analysis of GEO data. Xingjie Shi, Shihao Shen, Jin Liu 0011, Jian Huang 0003, Shuangge Ma |
Briefings Bioinform. | 2 |
| 2010 | MADS+: discovery of differential splicing events from Affymetrix exon junction array dataabstractMOTIVATION: The Affymetrix Human Exon Junction Array is a newly designed high-density exon-sensitive microarray for global analysis of alternative splicing. Contrary to the Affymetrix exon 1.0 array, which only contains four probes per exon and no probes for exon-exon junctions, this new junction array averages eight probes per probeset targeting all exons and exon-exon junctions observed in the human mRNA/EST transcripts, representing a significant increase in the probe density for alternative splicing events. Here, we present MADS+, a computational pipeline to detect differential splicing events from the Affymetrix exon junction array data. For each alternative splicing event, MADS+ evaluates the signals of probes targeting competing transcript isoforms to identify exons or splice sites with different levels of transcript inclusion between two sample groups. MADS+ is used routinely in our analysis of Affymetrix exon junction arrays and has a high accuracy in detecting differential splicing events. For example, in a study of the novel epithelial-specific splicing regulator ESRP1, MADS+ detects hundreds of exons whose inclusion levels are dependent on ESRP1, with a RT-PCR validation rate of 88.5% (153 validated out of 173 tested). AVAILABILITY: MADS+ scripts, documentations and annotation files are available at http://www.medicine.uiowa.edu/Labs/Xing/MADSplus/. Shihao Shen, Claude C. Warzecha, Russ P. Carstens, Yi Xing |
Bioinform. | 1 |