EDBT 2026 Demo / reviewers in the wild / expert
Hsin-Wei Wang
dblp:61/2130
· DBLP profile ↗
20ranked-venue papers
3as first author
13since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 1 first-author · 10 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-authorSystems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | PRIME: Novel Prompting Strategies for Effective Biasing Word Recognition in Contextualized ASRabstractAccurately recognizing domain-specific words remains an arduous challenge facing current automatic speech recognition (ASR) systems. While there are a number of prior arts managing to incorporate domain context through crossattention mechanisms, these methods often add additional components that increase model complexity and focus solely on wordlevel cues, overlooking broader domain-level topic information. To address these limitations, we put forward PRIME, a simple yet effective prompt-tuning method that enriches ASR with domain context using LLM-generated topic descriptions. To mitigate the limited context window of the ASR decoder, we introduce a biasing word retriever that selects the most relevant domain-specific words to construct informative prompts. Notably, PRIME requires no architectural modifications and offers a lightweight, scalable solution for contextualized ASR. A series of experiments counducted on the AISHELL and SlideSpeech benchmark datasets show that PRIME considerably promotes biasing word recognition, outperforming some strong baselines. Yu-Chun Liu, Li-Ting Pai, Yi-Cheng Wang, Bi-Cheng Yan, Hsin-Wei Wang, Chi-Han Lin, Juan-Wei Xu, Berlin Chen |
ASRU | 5 |
| 2025 | ConPCO: Preserving Phoneme Characteristics For Automatic Pronunciation Assessment Leveraging Contrastive Ordinal RegularizationabstractAutomatic pronunciation assessment (APA) manages to evaluate the pronunciation proficiency of a second language (L2) learner in a target language. Existing efforts typically draw on regression models for proficiency score prediction, wherein the models are trained to estimate target values without explicitly accounting for phoneme-awareness in the feature space. In this paper, we propose a contrastive phonemic ordinal regularizer (ConPCO) tailored for regression-based APA models to generate more phoneme-discriminative features while factoring in the ordinal relationships among the regression targets. The proposed ConPCO first aligns the phoneme representations of an APA model and textual embeddings of phonetic transcriptions via contrastive learning. Afterward, the phoneme characteristics are retained by regulating the distances between inter- and intra-phoneme categories in the feature space while allowing for the ordinal relationships among the output targets. We further design and develop a hierarchical APA model to evaluate the effectiveness of our regularizer. A series of experiments conducted on the speechocean762 benchmark dataset suggests the feasibility and effectiveness of our approach in relation to several competitive baselines. Bi-Cheng Yan, Yi-Cheng Wang, Jiun-Ting Li, Meng-Shin Lin, Hsin-Wei Wang, Wei-Cheng Chao, Berlin Chen |
ICASSP | 5 |
| 2024 | An Effective Pronunciation Assessment Approach Leveraging Hierarchical Transformers and Pre-training StrategiesabstractBi-Cheng Yan, Jiun-Ting Li, Yi-Cheng Wang, Hsin Wei Wang, Tien-Hong Lo, Yung-Chang Hsu, Wei-Cheng Chao, Berlin Chen. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Bi-Cheng Yan, Jiun-Ting Li, Yi-Cheng Wang, Hsin-Wei Wang, Tien-Hong Lo, Yung-Chang Hsu, Wei-Cheng Chao, Berlin Chen |
ACL (1) | 4 |
| 2024 | DANCER: Entity Description Augmented Named Entity Corrector for Automatic Speech RecognitionabstractEnd-to-end automatic speech recognition (E2E ASR) systems often suffer from mistranscription of domain-specific phrases, such as named entities, sometimes leading to catastrophic failures in downstream tasks. A family of fast and lightweight named entity correction (NEC) models for ASR have recently been proposed, which normally build on pho-netic-level edit distance algorithms and have shown impressive NEC performance. However, as the named entity (NE) list grows, the problems of phonetic confusion in the NE list are exacerbated; for example, homophone ambiguities increase substantially. In view of this, we proposed a novel Description Augmented Named entity CorrEctoR (dubbed DANCER), which leverages entity descriptions to provide additional information to facilitate mitigation of phonetic con-fusion for NEC on ASR transcription. To this end, an efficient entity description augmented masked language model (EDA-MLM) comprised of a dense retrieval model is introduced, enabling MLM to adapt swiftly to domain-specific entities for the NEC task. A series of experiments conducted on the AISHELL-1 and Homophone datasets confirm the effectiveness of our modeling approach. DANCER outperforms a strong baseline, the phonetic edit-distance-based NEC model (PED-NEC), by a character error rate (CER) reduction of about 7% relatively on AISHELL-1 for named entities. More notably, when tested on Homophone that contain named entities of high phonetic confusion, DANCER offers a more pronounced CER reduction of 46% relatively over PED-NEC for named entities. The code is available at https://github.com/Amiannn/Dancer. Yi-Cheng Wang, Hsin-Wei Wang, Bi-Cheng Yan, Chi-Han Lin, Berlin Chen |
LREC/COLING | 2 |
| 2024 | An Effective Mixture-Of-Experts Approach For Code-Switching Speech Recognition Leveraging Encoder DisentanglementabstractWith the massive developments of end-to-end (E2E) neural networks, recent years have witnessed unprecedented breakthroughs in automatic speech recognition (ASR). However, the code-switching phenomenon remains a major obstacle that hinders ASR from perfection, as the lack of labeled data and the variations between languages often lead to degradation of ASR performance. In this paper, we focus exclusively on improving the acoustic encoder of E2E ASR to tackle the challenge caused by the code-switching phenomenon. Our main contributions are threefold: First, we introduce a novel disentanglement loss to enable the lower-layer of the encoder to capture inter-lingual acoustic information while mitigating linguistic confusion at the higher-layer of the encoder. Second, through comprehensive experiments, we verify that our proposed method outperforms the prior-art methods using pre-trained dual-encoders, meanwhile having access only to the code-switching corpus and consuming half of the parameterization. Third, the apparent differentiation of the encoders’ output features also corroborates the complementarity between the disentanglement loss and the mixture-of-experts (MoE) architecture. Tzu-Ting Yang, Hsin-Wei Wang, Yi-Cheng Wang, Chi-Han Lin, Berlin Chen |
ICASSP | 2 |
| 2024 | An Effective Context-Balanced Adaptation Approach for Long-Tailed Speech RecognitionabstractEnd-to-end (E2E) automatic speech recognition (ASR) models have become standard practice for various commercial applications. However, in real-world scenarios, the long-tailed nature of word distribution often leads E2E ASR models to perform well on common words but fall short in recognizing uncommon ones. Recently, the notion of a contextual adapter (CA) was proposed to infuse external knowledge represented by a context word list into E2E ASR models. Although CA can improve recognition performance on rare words, two crucial data imbalance problems remain. First, when using low-frequency words as context words during training, since these words rarely occur in the utterance, CA becomes prone to overfit on attending to thetoken due to higher-frequency words not being present in the context list. Second, the long-tailed distribution within the context list itself still causes the model to perform poorly on low-frequency context words. In light of this, we explore in-depth the impact of altering the context list to have words with different frequency distributions on model performance, and meanwhile extend CA with a simple yet effective context-balanced learning objective1. A series of experiments conducted on the AISHELL-1 benchmark dataset suggests that using all vocabulary words from the training corpus as the context list and pairing them with our balanced objective yields the best performance, demonstrating a significant reduction in character error rate (CER) by up to 1.21% and a more pronounced 9.44% reduction in the error rate of zero-shot words.1The code is available at: https://github.com/Amiannn/espnet/tree/context_balanced_adapter Yi-Cheng Wang, Li-Ting Pai, Bi-Cheng Yan, Hsin-Wei Wang, Chi-Han Lin, Berlin Chen |
SLT | 4 |
| 2024 | Enhancing Code-Switching ASR Leveraging Non-Peaky CTC Loss and Deep Language Posterior InjectionabstractCode-switching—where multilingual speakers alternately switch between languages during conversations—still poses significant challenges to end-to-end (E2E) automatic speech recognition (ASR) systems due to phenomena of both acoustic and semantic confusion. This issue arises because ASR systems struggle to handle the rapid alternation of languages effectively, which often leads to significant performance degradation. Our main contributions are at least threefold: First, we incorporate language identification (LID) information into several intermediate layers of the encoder, aiming to enrich output embeddings with more detailed language information. Secondly, through the novel application of language boundary alignment loss, the subsequent ASR modules are enabled to more effectively utilize the knowledge of internal language posteriors. Third, we explore the feasibility of using language posteriors to facilitate deep interaction between shared encoder and language-specific encoders. Through comprehensive experiments on the SEAME corpus, we have verified that our proposed method outperforms the prior-art method, disentangle based mixture-of-experts (D-MoE), further enhancing the acuity of the encoder to languages. Tzu-Ting Yang, Hsin-Wei Wang, Yi-Cheng Wang, Berlin Chen |
SLT | 2 |
| 2023 | Preserving Phonemic Distinctions For Ordinal Regression: A Novel Loss Function For Automatic Pronunciation AssessmentabstractAutomatic pronunciation assessment (APA) manages to quantify the pronunciation proficiency of a second language (L2) learner in a language. Prevailing approaches to APA normally leverage neural models trained with a regression loss function, such as the mean-squared error (MSE) loss, for proficiency level prediction. Despite most regression models can effectively capture the ordinality of proficiency levels in the feature space, they are confronted with a primary obstacle that different phoneme categories with the same proficiency level are inevitably forced to be close to each other, retaining less phoneme-discriminative information. On account of this, we devise a phonemic contrast ordinal (PCO) loss for training regression-based APA models, which aims to preserve better phonemic distinctions between phoneme categories meanwhile considering ordinal relationships of the regression target output. Specifically, we introduce a phoneme-distinct regularizer into the MSE loss, which encourages feature representations of different phoneme categories to be far apart while simultaneously pulling closer the representations belonging to the same phoneme category by means of weighted distances. An extensive set of experiments carried out on the speechocean 762 benchmark dataset demonstrate the feasibility and effectiveness of our model in relation to some existing state-of-the-art models. Bi-Cheng Yan, Hsin-Wei Wang, Yi-Cheng Wang, Jiun-Ting Li, Chi-Han Lin, Berlin Chen |
ASRU | 2 |
| 2023 | Effective Graph-Based Modeling of Articulation Traits for Mispronunciation Detection and DiagnosisabstractMispronunciation detection and diagnosis (MDD) manages to pinpoint phone-level erroneous pronunciation segmentations and provide instant and informative diagnostic feedback to L2 (second-language) learners. Among the various modeling paradigms for MDD, dictation-based neural methods have recently become a de facto standard, which identifies pronunciation errors and returns diagnostic feedback at the same time by aligning the recognized phone sequence uttered by an L2 learner to the corresponding canonical phone sequence of a given text prompt. Despite their decent efficacy, dictation-based methods have at least two downsides. First, the dictation process and alignment process are made independent of each other, often resulting in a poor diagnostic feedback. Second, prior knowledge about the articulation traits of the canonical phones in the text prompt is not fully utilized in MDD. On account of this, we propose a novel end-to-end MDD method that can streamline the dictation process and the alignment process in a non-autoregressive manner. In addition, knowledge about phone-level articulation traits are extracted with a graph convolutional network (GCN) to obtain more discriminative phonetic embeddings so as to promote the MDD performance. An extensive set of experiments conducted on the L2-ARCTIC benchmark dataset suggest the feasibility and effectiveness of our approach in relation to competitive baselines. Bi-Cheng Yan, Hsin-Wei Wang, Yi-Cheng Wang, Berlin Chen |
ICASSP | 2 |
| 2022 | Exploring Non-Autoregressive End-to-End Neural Modeling for English Mispronunciation Detection and DiagnosisabstractEnd-to-end (E2E) neural modeling has emerged as one predominant school of thought to develop computer-assisted pronunciation training (CAPT) systems, showing competitive performance to conventional pronunciation-scoring based methods. However, current E2E neural methods for CAPT are faced with at least two pivotal challenges. On one hand, most of the E2E methods operate in an autoregressive manner with left-to-right beam search to dictate the pronunciations of an L2 learners. This however leads to very slow inference speed, which inevitably hinders their practical use. On the other hand, E2E neural methods are normally data-hungry and meanwhile an insufficient amount of nonnative training data would often reduce their efficacy on mispronunciation detection and diagnosis (MD&D). In response, we put forward a novel MD&D method that leverages non-autoregressive (NAR) E2E neural modeling to dramatically speed up the inference time while maintaining performance in line with the conventional E2E neural methods. In addition, we design and develop a pronunciation modeling network stacked on top of the NAR E2E models of our method to further boost the effectiveness of MD&D. Empirical experiments conducted on the L2-ARCTIC English dataset seems to validate the feasibility of our method, in comparison to some top-of-the-line E2E models and an iconic pronunciation-scoring based method built on a DNN-HMM acoustic model. Hsin-Wei Wang, Bi-Cheng Yan, Hsuan-Sheng Chiu, Yung-Chang Hsu, Berlin Chen |
ICASSP | 1 |
| 2022 | Maximum F1-Score Training for End-to-End Mispronunciation Detection and Diagnosis of L2 English SpeechabstractEnd-to-end (E2E) neural models are increasingly attracting attention as a promising modeling approach for mispronunciation detection and diagnosis (MDD). Typically, these models are trained by optimizing a cross-entropy criterion, which corresponds to improving the log-likelihood of the training data. However, there is a discrepancy between the objectives of model training and the MDD evaluation, since the performance of an MDD model is commonly evaluated in terms of F1-score instead of phone or word error rate (PER/WER). In view of this, we in this paper explore the use of a discriminative objective function for training E2E MDD models, which aims to maximize the expected F1-score directly. A series of experiments conducted on the L2-ARCTIC dataset show that our proposed method can yield considerable performance improvements in relation to some state-of-the-art E2E MDD approaches and the celebrated GOP method. Bi-Cheng Yan, Hsin-Wei Wang, Shao-Wei Fan-Jiang, Fu-An Chao, Berlin Chen |
ICME | 2 |
| 2022 | Effective Cross-Utterance Language Modeling for Conversational Speech RecognitionabstractConversational speech normally is embodied with loose syntactic structures at the utterance level but simultaneously exhibits topical coherence relations across consecutive utterances. Prior work has shown that capturing longer context information with a recurrent neural network or long short-term memory language model (LM) may suffer from the recent bias while excluding the long-range context. In order to capture the long-term semantic interactions among words and across utterances, we put forward disparate conversation history fusion methods for language modeling in automatic speech recognition (ASR) of conversational speech. Furthermore, a novel audio-fusion mechanism is introduced, which manages to fuse and utilize the acoustic embeddings of a current utterance and the semantic content of its corresponding conversation history in a cooperative way. To flesh out our ideas, we frame the ASR N-best hypothesis rescoring task as a prediction problem, leveraging BERT, an iconic pre-trained LM, as the ingredient vehicle to facilitate selection of the oracle hypothesis from a given N-best hypothesis list. Empirical experiments conducted on the AMI benchmark dataset seem to demonstrate the feasibility and efficacy of our methods in relation to some current top-of-line methods. The proposed methods not only achieve significant inference time reduction but also improve the ASR performance for conversational speech. Bi-Cheng Yan, Hsin-Wei Wang, Shih-Hsuan Chiu, Hsuan-Sheng Chiu, Berlin Chen |
IJCNN | 2 |
| 2022 | Peppanet: Effective Mispronunciation Detection and Diagnosis Leveraging Phonetic, Phonological, and Acoustic CuesabstractMispronunciation detection and diagnosis (MDD) aims to detect erroneous pronunciation segments in an L2 learner's articulation and subsequently provide informative diagnostic feedback. Most existing neural methods follow a dictation-based modeling paradigm that finds out pronunciation errors and returns diagnostic feedback at the same time by aligning the recognized phone sequence uttered by an L2 learner to the corresponding canonical phone sequence of a given text prompt. However, the main downside of these methods is that the dictation process and alignment process are mostly made independent of each other. In view of this, we present a novel end-to-end neural method, dubbed PeppaNet, building on a unified structure that can jointly model the dictation process and the alignment process. The model of our method learns to directly predict the pronunciation correctness of each canonical phone of the text prompt and in turn provides its corresponding diagnostic feedback. In contrast to the conventional dictation-based methods that rely mainly on a free-phone recognition process, PeppaNet makes good use of an effective selective gating mechanism to simultaneously incorporate phonetic, phonological and acoustic cues to generate corrections that are more proper and phonetically related to the canonical pronunciations. Extensive sets of experiments conducted on the L2-ARCTIC benchmark dataset seem to show the merits of our proposed method in comparison to some recent top-of-the-line methods. Bi-Cheng Yan, Hsin-Wei Wang, Berlin Chen |
SLT | 2 |
| 2014 | A local average distance descriptor for flexible protein structure comparisonabstractBACKGROUND: Protein structures are flexible and often show conformational changes upon binding to other molecules to exert biological functions. As protein structures correlate with characteristic functions, structure comparison allows classification and prediction of proteins of undefined functions. However, most comparison methods treat proteins as rigid bodies and cannot retrieve similarities of proteins with large conformational changes effectively. RESULTS: In this paper, we propose a novel descriptor, local average distance (LAD), based on either the geodesic distances (GDs) or Euclidean distances (EDs) for pairwise flexible protein structure comparison. The proposed method was compared with 7 structural alignment methods and 7 shape descriptors on two datasets comprising hinge bending motions from the MolMovDB, and the results have shown that our method outperformed all other methods regarding retrieving similar structures in terms of precision-recall curve, retrieval success rate, R-precision, mean average precision and F1-measure. CONCLUSIONS: Both ED- and GD-based LAD descriptors are effective to search deformed structures and overcome the problems of self-connection caused by a large bending motion. We have also demonstrated that the ED-based LAD is more robust than the GD-based descriptor. The proposed algorithm provides an alternative approach for blasting structure database, discovering previously unknown conformational relationships, and reorganizing protein structure classification. Hsin-Wei Wang, Chia-Han Chu, Wen-Ching Wang, Tun-Wen Pai |
BMC Bioinform. | 1 |
| 2013 | Protein-ligand binding region prediction (PLB-SAVE) based on geometric features and CUDA accelerationabstractBACKGROUND: Protein-ligand interactions are key processes in triggering and controlling biological functions within cells. Prediction of protein binding regions on the protein surface assists in understanding the mechanisms and principles of molecular recognition. In silico geometrical shape analysis plays a primary step in analyzing the spatial characteristics of protein binding regions and facilitates applications of bioinformatics in drug discovery and design. Here, we describe the novel software, PLB-SAVE, which uses parallel processing technology and is ideally suited to extract the geometrical construct of solid angles from surface atoms. Representative clusters and corresponding anchors were identified from all surface elements and were assigned according to the ranking of their solid angles. In addition, cavity depth indicators were obtained by proportional transformation of solid angles and cavity volumes were calculated by scanning multiple directional vectors within each selected cavity. Both depth and volume characteristics were combined with various weighting coefficients to rank predicted potential binding regions. RESULTS: Two test datasets from LigASite, each containing 388 bound and unbound structures, were used to predict binding regions using PLB-SAVE and two well-known prediction systems, SiteHound and MetaPocket2.0 (MPK2). PLB-SAVE outperformed the other programs with accuracy rates of 94.3% for unbound proteins and 95.5% for bound proteins via a tenfold cross-validation process. Additionally, because the parallel processing architecture was designed to enhance the computational efficiency, we obtained an average of 160-fold increase in computational time. CONCLUSIONS: In silico binding region prediction is considered the initial stage in structure-based drug design. To improve the efficacy of biological experiments for drug development, we developed PLB-SAVE, which uses only geometrical features of proteins and achieves a good overall performance for protein-ligand binding region prediction. Based on the same approach and rationale, this method can also be applied to predict carbohydrate-antibody interactions for further design and development of carbohydrate-based vaccines. PLB-SAVE is available at http://save.cs.ntou.edu.tw. Ying-Tsang Lo, Hsin-Wei Wang, Tun-Wen Pai, Wen-Shyong Tzou, Hui-Huang Hsu, Hao-Teng Chang |
BMC Bioinform. | 2 |
| 2011 | A Hybrid Method of Propensity Scales and Support Vector Machine in a Linear Epitope PredictionabstractAn epitope activates B cells to amplify and induce antibodies which can neutralize the foreign molecules, particles and pathogens. It also plays a crucial role in developing synthetic peptides for vaccination. Identification of epitopes using biological screening approaches is time consuming and high cost. Therefore, bioinformatics approaches are developed to enhance the speed of identifying the epitopes and conserve time. Herein, a combinatorial methodology based on physico-chemical properties and SVM (Support Vector Machine) techniques was proposed to address the aim of this study. Datasets of epitope and non epitope segments with 2, 3 and 4 residues in length were trained and applied as statistical features of SVM. After training, three datasets including one curated and two public ones were employed to evaluate the performance of the proposed system which was also compared with four existing LE predictors, BepiPred, ABCpred, BCPred and FBCPred. Our proposed system has presented better specificity, accuracy, and positive prediction value (PPV) in most testing cases. High specificity and PPV of a linear epitope prediction can lead to an efficient and effective design on biological experiments. Hsin-Wei Wang, Ya-Chi Lin, Tun-Wen Pai, Pei-Wen Tsai, Hao-Teng Chang |
CISIS | 1 |
| 2010 | Using Solid Angles to Detect Protein Docking Regions by CUDA Parallel AlgorithmsabstractA novel approach based on solid angles and CUDA parallel technologies for detection of protein docking regions is proposed in this paper. A key feature of a solid angle reveals the geometrical characteristics of protein surface structures and provides an efficient criterion for identifying the anchor amino acids on surface regions. These solid angles as well as several corresponding physical properties facilitate rapid evaluation on matching shape complementary from two interactive proteins. According to the performance of the proposed algorithms on enzyme-inhibitor protein complexes, the retrieved and clustered potential docking regions and the counterpart of centralized anchor amino acids can be effectively determined. Most of identified candidate regions on protein surfaces reveal either with concave or convex characteristics, and the results conform to the regular conditions of protein docking phenomena with respect to enzyme-inhibitor complexes. More importantly, the main goal of this paper not only identifies possible protein docking regions by evaluating solid angle characteristics, but also provides a feasible and efficient way to identify possible binding regions by employing the CUDA parallel computing architecture. The system evaluation has shown that the proposed algorithms achieved an accuracy rate of 62.85% for identifying possible docking regions on two interacted proteins and an average of one-half saving in running time requirements by incorporating CUDA parallel algorithms. Ying-Tsang Lo, Yueh-Lin Tsai, Hsin-Wei Wang, Yu-Ping Hsu, Tun-Wen Pai |
ISPA | 3 |
| 2008 | Meta Analysis of Microarray Data Using Gene Regulation PathwaysabstractUsing microarray technology for genetic analysis in biological experiments requires computationally intensive tools to interpret results. The main objective here is to develop a ldquometa-analysisrdquo tool that enables researchers to ldquosprayrdquo microarray data over a network of relevant gene regulation relationships, extracted from a database of published gene regulatory pathway models. The consistency of the data from a microarray experiment is evaluated to determine if it agrees or contradicts with previous findings. The database is limited to ldquoactivaterdquo and ldquoinhibitrdquo gene regulatory relationships at this point and a heuristic graph based approach is developed for consistency checking. Predictions are made for the regulation of genes that were not a part of the microarray experiment, but are related to the experiment through regulatory relationships. This meta-analysis will not only highlight consistent findings but also pinpoint genes that were missed in earlier experiments and should be considered in subsequent analysis. Saira Ali Kazmi, Yoo-Ah Kim, Baikang Pei, Ravi Nori, David W. Rowe, Hsin-Wei Wang, Dong-Guk Shin |
BIBM | 6 |
| 2006 | A Computational Inference Framework for analyzing Gene Regulation Pathway using Microarray DataabstractMicroarray experiments produce gene expression data at such a high speed and volume that it is imperative to use highly specialized computational tools for their analyses. One group of such computational tools deals with, namely, "meta-analysis" of microarray data. This step attempts to extract biological interpretations from the identified gene expression pattern. One particular aspect of meta-analysis is sorting out which gene regulation pathways are active and/or inhibited. The focus on this paper is to propose a computational framework with which scientists can compare microarray data with known gene regulation networks that are formed by two known binary gene regulation relationships, activate and inhibit. Using this framework scientists can conduct numerous analysis tasks including (i) identify active or inhibited sub-networks out of massively interconnected gene regulation pathways, (ii) find key genes, namely hubs, that are inferred to be widely involved in multiple aspects of gene regulation, (iii) identify regions of the network that contradict the known regulation, and (iv) estimate the direction of expression of genes that were not included in the microarray experiment. One known utility of this meta-analysis is to help scientists to identify a group of genes that they have missed in earlier experiments and should include in their subsequent experiments. Another utility is to enable them to isolate the group of genes that should be followed more closely using higher accuracy gene expression assays. We introduce three separate but inter-related meta-analysis methodologies, namely, FCFS, Hub majority, and SDP-based. We illustrate our proposed framework using microarray data derived from noggin treated osteoblast cells. The example clearly shows finding sub-networks related to B-catenin which should be expected and thus demonstrates the effectiveness of our proposed framework Dong-Guk Shin, John Bluis, Yoo-Ah Kim, Winfried Krueger, Jeffrey Maddox, Ravi Nori, Nathan Viniconis, Hsin-Wei Wang, David W. Rowe |
BIBE | 8 |
| 2001 | Comparing Trees in a Phylogenetic Relationship RepositoryabstractScientists are able to use many different phylogenetic analysis tools to assist them in their research. The data collection and data analysis processes can take days or weeks to complete. Having this data available in a repository would reduce this process time and allow researchers to spend more time analyzing data instead of collecting it. We collected data for 21 completely sequenced genomes and created an intuitive interface for browsing the genomes. The interface allows users to search this data including the pre-calculated phylogenetic trees that are stored in the database. We have also developed a new method for comparing a large set of phylogenetic trees so users can search the database based on a user given hypothesis tree. John Bluis, Ravi Nori, Hsin-Wei Wang, Pinglei Zhou, J. Peter Gogarten, Dong-Guk Shin |
BIBE | 3 |