Vinhthuy T. Phan

dblp:20/1 · also Vinhthuy Phan · DBLP profile ↗
← Back
39ranked-venue papers
11as first author
11since 2021 · last 2026
0000-0001-6108-1228ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 25 · 7 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 12 · 1 first-author · 11 since 2021Theory of computation · 2 · 2 first-authorArtificial intelligence and machine learning · 1 · 1 first-author
YearPublicationVenuePosition
2026 Supporting K-12 Teachers in the Presidential AI Challenge: A Case Study of a Faculty-Mentored Workshop for AI Tool Creation
Yukyeong Song, Rachel Min Wong, Jinhee Kim, Jewoong Moon, Edward Patton, Vinhthuy T. Phan, Jennifer McCullum, Jess Day
AIED (5)6
2026 Using In-Class Exercise Data for Early Support of Struggling Students
abstract
This classroom experience report investigates how in-class programming exercises can function as early-warning systems for struggling students. We studied two graduate Computer Science courses (n = 57) where students' code was automatically captured every 20 seconds during short programming activities. These fine-grained behavioral snapshots created a continuous record of real-time engagement. Clustering analysis revealed three consistent profiles: Early Birds (44.4%) began quickly and achieved strong outcomes, Active Strugglers (33.3%) engaged steadily but solved fewer problems, and Delayed/Disengaged students (22.2%) delayed by 10+ minutes and performed poorly. These profiles emerged within the first six weeks of the semester and strongly predicted exam performance (exercise access delay correlated with midterm scores, r = -0.68). To explore timely intervention, we developed personalized practice materials for struggling students based on their observed patterns. Large language models (LLMs) were used to analyze code snapshots and generate tailored exercises. These materials were then reviewed for technical accuracy and pedagogical alignment before delivery. Among the 11 students receiving this support, 81.8% achieved high-performer status (≥85% on both exams), substantially exceeding baseline expectations. This study contributes to research on early warning systems, procrastination in computing education, and fine-grained behavioral analytics, while also demonstrating how LLMs can be integrated into classroom practice. Our findings suggest that routine programming activities can serve a dual purpose: supporting active learning while simultaneously providing a scalable, low-cost framework for early identification and intervention
Kritish Pahi, Vinhthuy T. Phan
SIGCSE (1)2
2025 Graduate Computer Science TA Perspectives on In-Person Pedagogical Training: An Experience Report
abstract
Computer science (CS) departments rely heavily on graduate teaching assistants (GTAs), yet many departments struggle to provide effective pedagogical training, particularly for international GTAs who face additional cultural and communication challenges. While pedagogical training improves student outcomes, engaging GTAs with diverse career priorities in such training remains difficult, and resource constraints often lead departments to default to less engaging online formats. This experience report examines the implementation of a low-cost, in-person training program for CS GTAs using a flipped classroom approach that supplemented existing online modules. Our cohort of 34 international GTAs actively engaged with the training, which focused on grading practices, feedback techniques, cultural competencies, and common teaching scenarios. Survey data collected before and after training and midway through the semester revealed that participants found the training highly beneficial, with 94% reporting increased preparedness for their teaching roles. GTAs consistently applied learned skills throughout the semester, particularly in providing effective feedback (90%) and using rubrics (76%). The program's structure-requiring minimal faculty resources while yielding significant improvements in GTA confidence and teaching practices-offers a practical, replicable model for other CS departments seeking to enhance GTA preparation without substantial resource investment.
Alina Zaman, Amy Cook, Vinhthuy T. Phan, Alistair Windsor
ITiCSE (1)3
2025 In-class Coding Exercises As A Mechanism To Inform Early Intervention In Programming Courses
abstract
Early intervention is critical in increasing student success in Computer Science (CS) courses, which have attracted a diverse student population. In-class exercises, which are often low-stake and quick assignments, are a popular method for active learning and formative assessment. This study explores the potential of using in-class coding exercises for early intervention in programming courses, particularly before midterm exams. We analyzed historical data from a CS1 course to evaluate whether in-class coding exercises can predict midterm exam performance. Our findings reveal that in-class coding exercises are effective predictors of midterm performance and can serve as valuable tools for early intervention. Specifically, exercise scores and time on task are sufficient indicators of student performance. Although in-class exercises are less powerful predictors than traditional metrics, they offer quicker actionable insights. Additionally, predicting students who are struggling is more feasible than forecasting those who will fail or achieve specific letter grades. This research underscores the potential for designing targeted intervention schemes to support students in CS1 and other programming courses, highlighting the importance of timely and data-driven support mechanisms.
Eric Hicks, Vinhthuy T. Phan
SIGCSE (1)2
2025 Enhancing Student Performance Prediction In CS1 Via In-Class Coding
abstract
Computer science's increased recognition as a prominent field of study has attracted students with diverse academic backgrounds. This has significantly increased the already high failure rates in introductory courses. To address this challenge, it is essential to identify struggling students early on. Incorporating in-class coding exercises in these courses not only offers additional practice opportunities to students but may also reveal their abilities and help teachers identify those in need of assistance. In this work, we seek to determine the extent to which the practice of using in-class coding exercises enhances the ability to predict student performance, especially early in the semester. Based on data obtained in a CS1 course taught at a mid-size American university, we found that in-class exercises could improve the prediction of students' eventual performance. In particular, we found relatively accurately predictions as early as academic weeks 3 through 5, making it possible to devise early intervention strategies. This work can benefit future studies on the impact of in-class exercises as well as intervention strategies throughout the semester.
Eric Hicks, Vinhthuy T. Phan, Kriangsiri Malasri
SIGCSE (1)2
2023 A Practical Strategy for Training Graduate CS Teaching Assistants to Provide Effective Feedback
abstract
Computer science (CS) relies heavily on teaching assistants (TAs) who are often untrained in CS pedagogy. Existing research on CS TA training typically studies American undergraduate TAs at high-resource universities, ignoring the many universities that use graduate TAs, who are often international students, and that don't have the resources to implement the training strategies discussed in the literature. We describe our approach to implement graduate TA training in a high-diversity, low-resource context. We present a needs assessment, design, pilot test, and deployment of our training course, and discuss implications for other similar departments hoping to train their TAs.
Alina Zaman, Amy Cook, Vinhthuy T. Phan, Alistair Windsor
ITiCSE (1)3
2023 A Cloud-Based Technology for Conducting In-class Exercises in Data Science and Machine Learning Courses
abstract
Teaching data science can be challenging partly due to a diverse student population and the difficulty of providing a hands-on coding experience on complex topics. To address these challenges, we introduce a software package that facilitates active learning in the form of in-class coding exercises. This approach provides a much-needed hands-on experience in courses with a diverse student population and a highly technical content. Utilizing a popular cloud-based technology, JupyterHub, this approach enables in-class exercises with personalized feedback from the instructor. We report a classroom experience of using the technology for the first time in a graduate-level Machine Learning course, consisting of a mix of Data Science and Computer Science students. We found that, to a great extent, the course instructor could conduct complex in-class exercises within 10-15 minutes of class time. The instructor was able to understand students' abilities and challenges better and provide them with meaningful personalized feedback as well as group feedback. Students felt that the experience provided valuable hands-on practice, helped them figure out coding mistakes, and prepared them better for homework assignments.
Kritish Pahi, Vinhthuy T. Phan
SIGCSE (1)2
2022 Improving TA Feedback on In-Class Coding Assignments for Introductory Computer Science
abstract
Teaching assistants (TAs) for introductory computer science courses are most often responsible for providing feedback on student code. TAs, however, lack teaching experience and are rarely trained in how to give effective feedback that positively impacts student learning. The lack of training is particularly problematic when TAs are asked to give feedback in real time, e.g. during in-class coding exercises. We analyzed data from multiple semesters, where CS1 TAs and instructors provided written feedback on in-class coding exercises. Importantly, a very small percentage of feedback met our gold standard for high quality. This finding reveals a need for training TAs to provide more effective feedback in introductory programming courses.
Amy Cook, Vinhthuy T. Phan, Alistair Windsor
ITiCSE (1)2
2022 Try That Again! How a Second Attempt on In-Class Coding Problems Benefits Students in CS1
abstract
One way to introduce active learning in large introductory computer science courses is for students to solve coding exercises in class. Although it is commonly understood that re-solving a problem after receiving feedback can deepen understanding and improve performance, students often do not have opportunities to make multiple attempts on in-class exercises due to practical classroom constraints in time and logistics. In this experience report, we share the results from our experience with multiple attempts in our CS1 course of 114 undergraduate students. In each of 2 lectures on arrays, students were given two in-class coding problems. The first was a practice problem, where they had either one attempt or two attempts to solve the problem, and the second was a test problem where all students had only one attempt. We measured how having one attempt or two attempts on the practice problem impacted student performance on the test problem. We observed that students who used a second attempt to try re-solving missed practice problems were more likely to succeed on the test problem, even if they missed both tries on the practice problem. This work suggests that, given the right context and tool, multiple attempts on in-class exercises in CS1 might improve student performance.
Amy Cook, Alina Zaman, Eric Hicks, Kriangsiri Malasri, Vinhthuy T. Phan
SIGCSE (1)5
2022 Keep It Relevant! Using In-class Exercises to Predict Weekly Performance in CS1
abstract
In large programming courses, it can be difficult for instructors to identify students who need help. Often the earliest indication of trouble is when a student fails an exam, which unfortunately can be too late. Using data from 7 sections of CS1 over multiple semesters, we show that performance on lab and in-class coding exercises can be used to accurately predict which students will fail or struggle on upcoming weekly lab assignments. We found that recent relevant in-class coding exercises were the best features for building accurate models. This approach has potential in helping CS1 instructors identify students who need help, determine which topics need additional attention, and formulate intervention plans, all on a weekly basis before each lab meeting.
Eric Hicks, Amy Cook, Kriangsiri Malasri, Alina Zaman, Vinhthuy T. Phan
SIGCSE (1)5
2022 Enabling In-Class Peer Feedback on Introductory Computer Science Coding Exercises
abstract
Instructors often implement active learning in CS1 by giving students in-class coding problems. Students need feedback on their work to improve. While some systems provide automated feedback, human feedback is more effective for novice learners. However, instructors cannot provide feedback quickly at a large scale. Peer feedback systems help students get prompt feedback during class. Existing CS peer feedback systems usually support feedback on completed code rather than work in progress, which limits opportunities to reflect on the feedback and correct their work. We introduce a novel system for giving peer feedback on code in progress during CS1 classes, as well as a pilot test of the peer feedback process in CS1. Our initial experience has implications for the delivery of in-class instruction and for teaching growth mindset in order to take full advantage of peer feedback.
Alina Zaman, Vinhthuy T. Phan, Amy Cook
SIGCSE (2)2
2019 icHET: interactive visualization of cytoplasmic heteroplasmy
abstract
SUMMARY: Although heteroplasmy has been studied extensively in animal systems, there is a lack of tools for analyzing, exploring and visualizing heteroplasmy at the genome-wide level in other taxonomic systems. We introduce icHET, which is a computational workflow that produces an interactive visualization that facilitates the exploration, analysis and discovery of heteroplasmy across multiple genomic samples. icHET works on short reads from multiple samples from any organism with an organellar reference genome (mitochondrial or plastid) and a nuclear reference genome. AVAILABILITY AND IMPLEMENTATION: The software is available at https://github.com/vtphan/HeteroplasmyWorkflow. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Vinhthuy T. Phan, Diem-Trang Pham, Caroline Melton, Adam J. Ramsey, Bernie J. Daigle Jr., Jennifer R. Mandel
Bioinform.1
2018 Code4Brownies: an active learning solution for teaching programming and problem solving in the classroom
abstract
Code4Brownies is a software solution designed to foster active learning, coding, and problem solving in the classroom. Through this active learning style and platform, teachers can instantly provide guided instruction that gradually assists and leads students through various steps of solving a problem before reaching a correct solution. Teachers can even provide individualized instruction that addresses different needs of students with different levels of preparation. Two different delivery modes of guided instruction (teacher controlled and on-demand at student request) support various classroom scenarios and teaching pedagogies. Our experience of using Code4Brownies over a period of several semesters suggests that this tool helped students become more engaged, perform better, and ultimately be more successful.
Vinhthuy T. Phan, Eric Hicks
ITiCSE1
2018 Leveraging known genomic variants to improve detection of variants, especially close-by Indels
abstract
Motivation: The detection of genomic variants has great significance in genomics, bioinformatics, biomedical research and its applications. However, despite a lot of effort, Indels and structural variants are still under-characterized compared to SNPs. Current approaches based on next-generation sequencing data usually require large numbers of reads (high coverage) to be able to detect such types of variants accurately. However Indels, especially those close to each other, are still hard to detect accurately. Results: We introduce a novel approach that leverages known variant information, e.g. provided by dbSNP, dbVar, ExAC or the 1000 Genomes Project, to improve sensitivity of detecting variants, especially close-by Indels. In our approach, the standard reference genome and the known variants are combined to build a meta-reference, which is expected to be probabilistically closer to the subject genomes than the standard reference. An alignment algorithm, which can take into account known variant information, is developed to accurately align reads to the meta-reference. This strategy resulted in accurate alignment and variant calling even with low coverage data. We showed that compared to popular methods such as GATK and SAMtools, our method significantly improves the sensitivity of detecting variants, especially Indels that are close to each other. In particular, our method was able to call these close-by Indels at a 15-20% higher sensitivity than other methods at low coverage, and still get 1-5% higher sensitivity at high coverage, at competitive precision. These results were validated using simulated data with variant profiles extracted from the 1000 Genomes Project data, and real data from the Illumina Platinum Genomes Project and ExAC database. Our finding suggests that by incorporating known variant information in an appropriate manner, sensitive variant calling is possible at a low cost. Availability and implementation: Implementation can be found in our public code repository https://github.com/namsyvo/IVC. Supplementary information: Supplementary data are available at Bioinformatics online.
Nam Sy Vo, Vinhthuy T. Phan
Bioinform.2
2017 Using 16S rRNA gene as marker to detect unknown bacteria in microbial communities
abstract
BACKGROUND: Quantification and identification of microbial genomes based on next-generation sequencing data is a challenging problem in metagenomics. Although current methods have mostly focused on analyzing bacteria whose genomes have been sequenced, such analyses are, however, complicated by the presence of unknown bacteria or bacteria whose genomes have not been sequence. RESULTS: We propose a method for detecting unknown bacteria in environmental samples. Our approach is unique in its utilization of short reads only from 16S rRNA genes, not from entire genomes. We show that short reads from 16S rRNA genes retain sufficient information for detecting unknown bacteria in oral microbial communities. CONCLUSION: In our experimentation with bacterial genomes from the Human Oral Microbiome Database, we found that this method made accurate and robust predictions at different read coverages and percentages of unknown bacteria. Advantages of this approach include not only a reduction in experimental and computational costs but also a potentially high accuracy across environmental samples due to the strong conservation of the 16S rRNA gene.
Quang Tran 0002, Diem-Trang Pham, Vinhthuy T. Phan
BMC Bioinform.3
2016 Analysis of optimal alignments unfolds aligners' bias in existing variant profiles
abstract
Efforts such as International HapMap Project and 1000 Genomes Project resulted in a catalog of millions of single nucleotides and insertion/deletion (INDEL) variants of the human population. Viewed as a reference of existing variants, this resource commonly serves as a gold standard for studying and developing methods to detect genetic variants. Our analysis revealed that this reference contained thousands of INDELs that were constructed in a biased manner. This bias occurred at the level of aligning short reads to reference genomes to detect variants. The bias is caused by the existence of many theoretically optimal alignments between the reference genome and reads containing alternative alleles at those INDEL locations. We examined several popular aligners and showed that these aligners could be divided into groups whose alignments yielded INDELs that agreed strongly or disagreed strongly with reported INDELs. This finding suggests that the agreement or disagreement between the aligners' called INDEL and the reported INDEL is merely a result of the arbitrary selection of one of the optimal alignments. The existence of bias in INDEL calling might have a serious influence in downstream analyses. As such, our finding suggests that this phenomenon should be further addressed.
Quang Tran 0002, Vinhthuy T. Phan
BMC Bioinform.3
2015 Alignment-free methods for metagenomic profiling
abstract
Background The primary goal of metagenomic studies is to analyze and evaluate the rich microbial communities present in all natural environments. The construction and utilization of a large index required by alignment-based methods for thousands of microbial genomes can be computationally prohibitive. To avoid this computational cost, we investigated three different variations of an alignmentfree method for profiling abundances of microbial communities.
Diem-Trang Pham, Vinhthuy T. Phan
BMC Bioinform.3
2015 How genome complexity can explain the difficulty of aligning reads to genomes
abstract
BACKGROUND: Although it is frequently observed that aligning short reads to genomes becomes harder if they contain complex repeat patterns, there has not been much effort to quantify the relationship between complexity of genomes and difficulty of short-read alignment. Existing measures of sequence complexity seem unsuitable for the understanding and quantification of this relationship. RESULTS: We investigated several measures of complexity and found that length-sensitive measures of complexity had the highest correlation to accuracy of alignment. In particular, the rate of distinct substrings of length k, where k is similar to the read length, correlated very highly to alignment performance in terms of precision and recall. We showed how to compute this measure efficiently in linear time, making it useful in practice to estimate quickly the difficulty of alignment for new genomes without having to align reads to them first. We showed how the length-sensitive measures could provide additional information for choosing aligners that would align consistently accurately on new genomes. CONCLUSIONS: We formally established a connection between genome complexity and the accuracy of short-read aligners. The relationship between genome complexity and alignment accuracy provides additional useful information for selecting suitable aligners for new genomes. Further, this work suggests that the complexity of genomes sometimes should be thought of in terms of specific computational problems, such as the alignment of short reads to genomes.
Vinhthuy T. Phan, Quang Tran 0002, Nam Sy Vo
BMC Bioinform.1
2015 A linear model for predicting performance of short-read aligners using genome complexity
abstract
Background The effectiveness and accuracy of aligning short reads to genomes have an important impact on many applications that rely on next-generation sequencing data. The computational requirements and material cost for aligning largescale short reads to genomes is also expensive. To prevent wasted time and resources for aligning short reads, we investigated the different measures of genome complexity [1] that correlated best to the performance of alignment to propose a linear model for each aligning method [2].
Quang Tran 0002, Nam Sy Vo, Vinhthuy T. Phan
BMC Bioinform.4
2015 Improving variant calling by incorporating known genetic variants into read alignment
Nam Sy Vo, Vinhthuy T. Phan
BMC Bioinform.2
2014 Determining gene response patterns of time series gene expression data using R
abstract
Background The rapid advancement of sequencing technologies has not been without challenges. The extensive size of genomes as well as the need to comprehend the vast complexities of genomic information has fostered the need to utilize more robust computational methods and statistical analysis tools. In a gene expression study involving multiple treatments, a time series analysis of differentially expressed genes provides great insights into the replicates. Such information is useful in identifying the potential sources of variation that cannot be easily extrapolated from a generalized experimental approach. It is also essential to be able to group the identified gene expression patterns in order to attain meaningful interpretations. Materials and methods The experiment employs the use of R software to identify gene expression patterns in a dataset involving parathyroid tumor treatments [1]. The treatments were generated in a three day time interval. The R tool provides important packages that are ideal for gene expression studies and also delivers significant visualization and inference resources [2]. The proposed approach for this analysis consists of two stages. First, the Kruskal-Wallis test is used to identify differentially expressed genes. Second, patterns of differentially expressed genes are determined using the Wilcoxon rank-sum test.
Kevin L. O'kello, Vinhthuy T. Phan
BMC Bioinform.2
2014 Alignment of short reads to multiple genomes using hashing
abstract
Materials and methods Inspired by [3], we propose a new method that attempts to take advantage of multiple genomes and SNV information to align reads. This approach is promising in that it allows us to distinguish between sequencing errors and SNV. Our proposed alignment algorithm uses read fragments to identify seeds and extend these seeds to find occurrences of reads in the genome. In this study, we have developed and implemented an algorithm using multiple genomes that captures genomic variations, indexes the multiple genomes and operates short read alignment on a collection of genomes. The preliminary result was validated on Aspergillus fumigatus.
Quang Tran 0002, Vinhthuy T. Phan
BMC Bioinform.2
2014 Exploiting the bootstrap method to analyze patterns of gene expression
abstract
Background High-throughput technologies like microarrays or the recent RNA-Seq provide large amounts of data for gene expression studies. Although there have been diverse methods to design gene-expression experiments and analyze gene-expression data, the prediction of true patterns of gene expression in case of having few samples remains a challenging problem [1,2]. Materials and methods We propose a method to predict response patterns of gene expression studies in the case of small sample size using a bootstrap method [3]. Our approach adopts partially order sets (posets) to represent gene patterns, which are determined based on pairwise comparisons [4]. Results We show that patterns that are not linearly orderable cannot be true patterns of gene response to treatments. From this result, we propose a strategy using bootstrap resampling to infer true responses of non-linearly-orderable patterns. Our experiments showed that this method produced gene lists with more biological functional enrichment than those obtained without bootstrap resampling. Conclusions Our method is useful in designing cost-effective experiments with small sample sizes. Researchers can still use a small sample size to determine true patterns for most genes. For highly-variantly expressed genes, their true patterns can be identified using the proposed method.
Nam Sy Vo, Vinhthuy T. Phan
BMC Bioinform.2
2014 Exploiting dependencies of pairwise comparison outcomes to predict patterns of gene response
abstract
The analysis of gene expression has played an important role in medical and bioinformatics research. Although it is known that a large number of samples is needed to determine the patterns of gene expression accurately, practical designs of gene expression studies occasionally have insufficient numbers of samples, making it difficult to ascertain true response patterns of variantly expressed genes. We describe an approach to cope with the challenge of predicting true orders of gene response to treatments. We show that true patterns of gene response must be orderable sets. In experiments with few samples, we modify the conventional pairwise comparison tests and increase the significance level α intelligently to deduce orderable patterns, which are most likely true orders of gene response. Additionally, motivated by the fact that a gene can be involved in multiple biological functions, our method further resamples experimental replicates and predicts multiple response patterns for each gene. Using a gene expression data set of Sprague-Dawley rats treated with chemopreventive chemical compounds and DAVID to annotate and validate gene sets, we showed that compared to the conventional method of fixing α , this method increased enrichment significantly. A comparison with hierarchical clustering showed that gene clusters labelled by response patterns produced by our method were much more enriched. One of the clusters contained 3 transcription factors, which hierarchical clustering failed to place into one cluster, that have been found to participate in multiple biological networks. One of the transcription factors is known to play an important role in pathways affected by the studied chemical compounds. This method can be useful in designing cost-effective experiments with small sample sizes. Patterns of highly-variantly expressed genes can be predicted by varying α intelligently. Furthermore, clusters are labeled meaningfully with patterns that describe precisely how genes in such clusters respond to treatments.
Nam Sy Vo, Vinhthuy T. Phan
BMC Bioinform.2
2014 An integrated approach for SNP calling based on population of genomes
abstract
Background The identification of genetic variants such as single nucleotide polymorphisms (SNPs) is a critical step in many applications based on NGS technologies [1]. Although many SNP calling programs have been developed, it is still challenging to accurately call SNPs, especially when coverage level is low [2]. Moreover, the determination of SNPs, which is performed through many separate steps, requires a careful selection of a diverse set of tools [3,4]. This can lead to several disadvantages, for example, one cannot incorporate information from the read alignment step into the SNP calling step or vice versa to help improve accuracy of called SNPs. Materials and methods We propose a novel integrated approach to detect more true SNPs while calling fewer false positives. Different from current methods that perform read alignment and SNP calling steps separately, our method combines them methodologically to improve the accuracy of SNP identification. To effectively exploit information from a population of genomes, databases of confirmed SNPs, such as dbSNP, are employed in both aligning reads to references as well as calling SNPs. This strategy allows us to develop a novel algorithm to align reads to references that can differentiate sequencing errors from SNPs. Results Based on this result, the method can call SNPs accurately and effectively even with low-coverage sequencing data. Our results on simulated data show that the method is able to call SNPs with very high precision and recall rate with low-coverage datasets. Conclusions With the existence of databases of confirmed SNPs for large amounts of sequenced species, our approach provides a promising method to call accurate SNP information even with low-coverage sequencing data. This approach can also help researchers facilitate the determination of SNPs by using an integrated SNP calling tool.
Nam Sy Vo, Quang Tran 0002, Vinhthuy T. Phan
BMC Bioinform.3
2013 Exploiting Dependencies of Patterns in Gene Expression Analysis Using Pairwise Comparisons
Nam Sy Vo, Vinhthuy T. Phan
ISBRA2
2013 Using partially ordered sets to represent and predict true patterns of gene response to treatments
abstract
Advances in biotechnology have empowered high-throughput measurement of gene expression levels for tens of thousands of genes simultaneously. This means that one sample size must be used for all genes in most experimental designs [ 1 , 2 ], which implies that patterns of response of highly variantly expressed genes might not be measured accurately. Response patterns of gene expression data with multiple treatments have been characterized using post hoc pairwise comparisons by several researchers [ 3 , 4 ]. Nevertheless, these researchers did not address how to cope with highly variantly expressed genes with inaccurate patterns due to having too few experimental samples. We show that dependencies of pairwise comparison outcomes in post hoc calculations can be exploited to infer true response patterns of genes with inaccurate patterns due to having too few experimental samples. Characterizing such response patterns as partially ordered sets, we show that linearly orderable patterns are more likely true patterns and those that are not linearly orderable cannot be true patterns. We propose a strategy to predict most likely linearly orderable extensions of such patterns. Using microarray data of rats' liver cells, we showed that this approach yielded more and better functionally enriched gene lists than a conventional approach. This approach opens up opportunities to design cost-effective experiments, in which only a conservatively large sample size is needed to collect expression levels of almost all genes. For most genes, such a sample size is sufficient. For highly variantly expressed genes, our method can help infer true response patterns.
Nam Sy Vo, Vinhthuy T. Phan
BMC Bioinform.2
2011 mDAG: a web-based tool for analyzing microarray data with multiple treatments
abstract
In microarray experiments involving multiple treatments, pairwise comparisons between all pairs of treatments are desirable but expensive. To cope with this, we previously introduced a method that performed all pairwise comparisons in a post hoc manner. This method employs directed graphs to represent gene response to pairs of treatments. It has been applied and found useful in identifying and differentiating genes sharing similar functional pathways [ 1 , 2 ]. mDAG is a web-based software based on this method. mDAG allows users to upload microarray data in GCT format through a web interface. From this data, the application performs calculations to assign graphical patterns to genes and outputs images and textual data for further analyses. These graphical patterns carry specific meanings in terms of how genes respond to pairs of treatments. The application is implemented using Python and web2py. mDAG is available at http://cetus.cs.memphis.edu:8080/mDAG . For experiments involved multiple treatments and replicates, mDAG allows researchers to analyze and visualize in graphical representations relationships of gene interactions to all pairs of treatments. The software can be used online or off-line.
Vinhthuy T. Phan, Nam Sy Vo, Thomas R. Sutter
BMC Bioinform.1
2011 Latent Semantic Indexing of PubMed abstracts for identification of transcription factor candidates from microarray derived gene sets
abstract
BACKGROUND: Identification of transcription factors (TFs) responsible for modulation of differentially expressed genes is a key step in deducing gene regulatory pathways. Most current methods identify TFs by searching for presence of DNA binding motifs in the promoter regions of co-regulated genes. However, this strategy may not always be useful as presence of a motif does not necessarily imply a regulatory role. Conversely, motif presence may not be required for a TF to regulate a set of genes. Therefore, it is imperative to include functional (biochemical and molecular) associations, such as those found in the biomedical literature, into algorithms for identification of putative regulatory TFs that might be explicitly or implicitly linked to the genes under investigation. RESULTS: In this study, we present a Latent Semantic Indexing (LSI) based text mining approach for identification and ranking of putative regulatory TFs from microarray derived differentially expressed genes (DEGs). Two LSI models were built using different term weighting schemes to devise pair-wise similarities between 21,027 mouse genes annotated in the Entrez Gene repository. Amongst these genes, 433 were designated TFs in the TRANSFAC database. The LSI derived TF-to-gene similarities were used to calculate TF literature enrichment p-values and rank the TFs for a given set of genes. We evaluated our approach using five different publicly available microarray datasets focusing on TFs Rel, Stat6, Ddit3, Stat5 and Nfic. In addition, for each of the datasets, we constructed gold standard TFs known to be functionally relevant to the study in question. Receiver Operating Characteristics (ROC) curves showed that the log-entropy LSI model outperformed the tf-normal LSI model and a benchmark co-occurrence based method for four out of five datasets, as well as motif searching approaches, in identifying putative TFs. CONCLUSIONS: Our results suggest that our LSI based text mining approach can complement existing approaches used in systems biology research to decipher gene regulatory networks by providing putative lists of ranked TFs that might be explicitly or implicitly associated with sets of DEGs derived from microarray experiments. In addition, unlike motif searching approaches, LSI based approaches can reveal TFs that may indirectly regulate genes.
Sujoy Roy, Kevin Heinrich, Vinhthuy T. Phan, Michael W. Berry, Ramin Homayouni
BMC Bioinform.3
2009 DNA Chips for Species Identification and Biological Phylogenies
Max H. Garzon, Tit-Yee Wong, Vinhthuy T. Phan
DNA3
2009 Motif Tool Manager: a web-based platform for motif discovery
Vinhthuy T. Phan
BMC Bioinform.1
2009 On codeword design in metric DNA spaces
Vinhthuy T. Phan, Max H. Garzon
Nat. Comput.1
2008 Synthetic Gene Design with a Large Number of Hidden Stop Codons
abstract
Hidden stop codons are nucleotide triples TAA, TAG, and TGA that appear in the second and third reading frames of a protein coding gene. Recent studies reported biological evidence suggesting that hidden stop codons are important in preventing misread of mRNA, which is often detrimental to the cell. We study the problem of designing protein-encoding genes with large number of hidden stop codons under biological constraints including GC content and codon usage of individual organism. In simpler models, we obtained provably optimal results. In more complex models, the designed genes have many more hidden stop codons than wild-type genes do, as observed in an experiment with 8 genomes with a wide range of GC content and codon usage.
Vinhthuy T. Phan, Sudip Saha, Ashutosh Pandey 0003, Tit-Yee Wong
BIBM1
2008 Motif Tool Manager: a web-based framework for motif discovery
abstract
MOTIVATION: Motif Tool Manager is a web-based framework for comparing and combining different approaches to discover novel DNA motifs. It comes with a set of five well-known approaches to motif discovery. It provides an easy mechanism for adding new motif finding tools to the framework through a web-interface and a minimal setup of the tools on the server. Users can execute the tools through the web-based framework and compare results from such executions. The framework provides a basic mechanism for identifying the most similar motif candidates found by a majority of themotif finding tools. AVAILABILITY: http://cetus.cs.memphis.edu/motif
Vinhthuy T. Phan, Nicholas A. Furlotte
Bioinform.1
2006 "Reasoning" and "Talking" DNA: Can DNA Understand English?
Kiranchand V. Bobba, Andrew Neel, Vinhthuy T. Phan, Max H. Garzon
DNA3
2006 In Search of Optimal Codes for DNA Computing
Max H. Garzon, Vinhthuy T. Phan, Sujoy Roy, Andrew Neel
DNA2
2003 A Model for Analyzing Black-Box Optimization
Vinhthuy T. Phan, Steven Skiena, Pavel Sumazin
WADS1
2002 A Time-Sensitive System for Black-Box Combinatorial Optimization
Vinhthuy T. Phan, Pavel Sumazin, Steven Skiena
ALENEX1
2001 Dealing with errors in interactive sequencing by hybridization
abstract
MOTIVATION: A realistic approach to sequencing by hybridization must deal with realistic sequencing errors. The results of such a method can surely be applied to similar sequencing tasks. RESULTS: We provide the first algorithms for interactive sequencing by hybridization which are robust in the presence of hybridization errors. Under a strong error model allowing both positive and negative hybridization errors without repeated queries, we demonstrate accurate and efficient reconstruction with error rates up to 7%. Under the weaker traditional error model of Shamir and Tsur (Proceedings of the Fifth International Conference on Computational Molecular Biology (RECOMB-01), pp 269-277, 2000), we obtain accurate reconstructions with up to 20% false negative hybridization errors. Finally, we establish theoretical bounds on the performance of the sequential probing algorithm of Skiena and Sundaram (J. Comput. Biol., 2, 333-353, 1995) under the strong error model. AVAILABILTY: Freely available upon request. CONTACT: [email protected].
Vinhthuy T. Phan, Steven Skiena
Bioinform.1