VLDB 2026 Research / reviewers in the wild / expert
Nam Sy Vo
dblp:121/1152 · also Nam S. Vo, Nam Vo
· DBLP profile ↗
20ranked-venue papers
10as first author
5since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 15 · 7 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | AMRViz enables seamless genomics analysis and visualization of antimicrobial resistanceabstractWe have developed AMRViz, a toolkit for analyzing, visualizing, and managing bacterial genomics samples. The toolkit is bundled with the current best practice analysis pipeline allowing researchers to perform comprehensive analysis of a collection of samples directly from raw sequencing data with a single command line. The analysis results in a report showing the genome structure, genome annotations, antibiotic resistance and virulence profile for each sample. The pan-genome of all samples of the collection is analyzed to identify core- and accessory-genes. Phylogenies of the whole genome as well as all gene clusters are also generated. The toolkit provides a web-based visualization dashboard allowing researchers to interactively examine various aspects of the analysis results. Availability: AMRViz is implemented in Python and NodeJS, and is publicly available under open source MIT license at https://github.com/amromics/amrviz . Duc Quang Le, Son Hoang Nguyen, Tam Thi Nguyen, Canh Hao Nguyen, Tho Huu Ho, Nam Sy Vo, Hoang Anh Nguyen, Minh Duc Cao |
BMC Bioinform. | 6 |
| 2022 | LmTag: functional-enrichment and imputation-aware tag SNP selection for population-specific genotyping arraysabstractDespite the rapid development of sequencing technology, single-nucleotide polymorphism (SNP) arrays are still the most cost-effective genotyping solutions for large-scale genomic research and applications. Recent years have witnessed the rapid development of numerous genotyping platforms of different sizes and designs, but population-specific platforms are still lacking, especially for those in developing countries. SNP arrays designed for these countries should be cost-effective (small size), yet incorporate key information needed to associate genotypes with traits. A key design principle for most current platforms is to improve genome-wide imputation so that more SNPs not included in the array (imputed SNPs) can be predicted. However, current tag SNP selection methods mostly focus on imputation accuracy and coverage, but not the functional content of the array. It is those functional SNPs that are most likely associated with traits. Here, we propose LmTag, a novel method for tag SNP selection that not only improves imputation performance but also prioritizes highly functional SNP markers. We apply LmTag on a wide range of populations using both public and in-house whole-genome sequencing databases. Our results show that LmTag improved both functional marker prioritization and genome-wide imputation accuracy compared to existing methods. This novel approach could contribute to the next generation genotyping arrays that provide excellent imputation capability as well as facilitate array-based functional genetic studies. Such arrays are particularly suitable for under-represented populations in developing countries or non-model species, where little genomics data are available while investment in genome sequencing or high-density SNP arrays is limited. $\textrm{LmTag}$ is available at: https://github.com/datngu/LmTag. Dat Thanh Nguyen, Quan Hoang Nguyen 0002, Duong Thuy Nguyen 0001, Nam Sy Vo |
Briefings Bioinform. | 4 |
| 2022 | Assessing polygenic risk score models for applications in populations with under-represented genomics data: an example of VietnamabstractMost polygenic risk score (PRS)models have been based on data from populations of European origins (accounting for the majority of the large genomics datasets, e.g. >78% in the UK Biobank and >85% in the GTEx project). Although several large-scale Asian biobanks were initiated (e.g. Japanese, Korean, Han Chinese biobanks), most other Asian countries have little or near-zero genomics data. To implement PRS models for under-represented populations, we explored transfer learning approaches, assuming that information from existing large datasets can compensate for the small sample size that can be feasibly obtained in developing countries, like Vietnam. Here, we benchmark 13 common PRS methods in meta-population strategy (combining individual genotype data from multiple populations) and multi-population strategy (combining summary statistics from multiple populations). Our results highlight the complementarity of different populations and the choice of methods should depend on the target population. Based on these results, we discussed a set of guidelines to help users select the best method for their datasets. We developed a robust and comprehensive software to allow for benchmarking comparisons between methods and proposed a computational framework for improving PRS performance in a dataset with a small sample size. This work is expected to inform the development of genomics applications in under-represented populations. PRSUP framework is available at: https://github.com/BiomedicalMachineLearning/VGP. Duy Pham, Buu Truong, Khai Tran, Guiyan Ni, Trang T. H. Tran, Mai H. Tran, Duong Thuy Nguyen 0001, Nam Sy Vo |
Briefings Bioinform. | 9 |
| 2022 | Extract antibody and antigen names from biomedical literatureabstractBACKGROUND: The roles of antibody and antigen are indispensable in targeted diagnosis, therapy, and biomedical discovery. On top of that, massive numbers of new scientific articles about antibodies and/or antigens are published each year, which is a precious knowledge resource but has yet been exploited to its full potential. We, therefore, aim to develop a biomedical natural language processing tool that can automatically identify antibody and antigen entities from articles. RESULTS: We first annotated an antibody-antigen corpus including 3210 relevant PubMed abstracts using a semi-automatic approach. The Inter-Annotator Agreement score of 3 annotators ranges from 91.46 to 94.31%, indicating that the annotations are consistent and the corpus is reliable. We then used the corpus to develop and optimize BiLSTM-CRF-based and BioBERT-based models. The models achieved overall F1 scores of 62.49% and 81.44%, respectively, which showed potential for newly studied entities. The two models served as foundation for development of a named entity recognition (NER) tool that automatically recognizes antibody and antigen names from biomedical literature. CONCLUSIONS: Our antibody-antigen NER models enable users to automatically extract antibody and antigen names from scientific articles without manually scanning through vast amounts of data and information in the literature. The output of NER can be used to automatically populate antibody-antigen databases, support antibody validation, and facilitate researchers with the most appropriate antibodies of interest. The packaged NER model is available at https://github.com/TrangDinh44/ABAG_BioBERT.git . Thuy Trang Dinh, Trang Phuong Vo-Chanh, Viet Quoc Huynh, Nam Sy Vo, Hoang Duc Nguyen |
BMC Bioinform. | 5 |
| 2022 | Does your dermatology classifier know what it doesn't know? Detecting the long-tail of unseen conditions
Abhijit Guha Roy, Jie Ren 0006, Shekoofeh Azizi, Aaron Loh, Vivek Natarajan, Basil Mustafa, Nick Pawlowski, Jan Freyberg, Zachary Beaver, Nam Sy Vo, Peggy Bui, Samantha Winter, Patricia MacWilliams, Gregory S. Corrado, Umesh Telang, Yun Liu 0013, A. Taylan Cemgil, Alan Karthikesalingam, Balaji Lakshminarayanan, Jim Winkens |
Medical Image Anal. | 11 |
| 2019 | Composing Text and Image for Image Retrieval - an Empirical OdysseyabstractIn this paper, we study the task of image retrieval, where the input query is specified in the form of an image plus some text that describes desired modifications to the input image. For example, we may present an image of the Eiffel tower, and ask the system to find images which are visually similar, but are modified in small ways, such as being taken at nighttime instead of during the day. o tackle this task, we embed the query (reference image plus modification text) and the target (images). The encoding function of the image text query learns a representation, such that the similarity with the target image representation is high iff it is a ``positive match''. We propose a new way to combine image and text through residual connection, that is designed for this retrieval task. We show this outperforms existing approaches on 3 different datasets, namely Fashion-200k, MIT-States and a new synthetic dataset we create based on CLEVR. We also show that our approach can be used to perform image classification with compositionally novel labels, and we outperform previous methods on MIT-States on this task. Nam Sy Vo, Lu Jiang 0004, Chen Sun 0002, Kevin Murphy 0002, Li-Jia Li 0001, Li Fei-Fei 0001, James Hays |
CVPR | 1 |
| 2019 | Generalization in Metric Learning: Should the Embedding Layer Be Embedding Layer?abstractThis work studies deep metric learning under small to medium scale as we believe that better generalization could be a contributing factor to the improvement of previous fine-grained image retrieval methods; it should be considered when designing future techniques. In particular, we investigate using other layers in a deep metric learning system (besides the embedding layer) for feature extraction and analyze how well they perform on training data and generalize to testing data. From this study, we suggest a new regularization practice where one can add or choose a more optimal layer for feature extraction. State-of-the-art performance is demonstrated on 3 fine-grained image retrieval benchmarks: Cars-196, CUB-200-2011, and Stanford Online Product. Nam Sy Vo, James Hays |
WACV | 1 |
| 2018 | Leveraging known genomic variants to improve detection of variants, especially close-by IndelsabstractMotivation: The detection of genomic variants has great significance in genomics, bioinformatics, biomedical research and its applications. However, despite a lot of effort, Indels and structural variants are still under-characterized compared to SNPs. Current approaches based on next-generation sequencing data usually require large numbers of reads (high coverage) to be able to detect such types of variants accurately. However Indels, especially those close to each other, are still hard to detect accurately. Results: We introduce a novel approach that leverages known variant information, e.g. provided by dbSNP, dbVar, ExAC or the 1000 Genomes Project, to improve sensitivity of detecting variants, especially close-by Indels. In our approach, the standard reference genome and the known variants are combined to build a meta-reference, which is expected to be probabilistically closer to the subject genomes than the standard reference. An alignment algorithm, which can take into account known variant information, is developed to accurately align reads to the meta-reference. This strategy resulted in accurate alignment and variant calling even with low coverage data. We showed that compared to popular methods such as GATK and SAMtools, our method significantly improves the sensitivity of detecting variants, especially Indels that are close to each other. In particular, our method was able to call these close-by Indels at a 15-20% higher sensitivity than other methods at low coverage, and still get 1-5% higher sensitivity at high coverage, at competitive precision. These results were validated using simulated data with variant profiles extracted from the 1000 Genomes Project data, and real data from the Illumina Platinum Genomes Project and ExAC database. Our finding suggests that by incorporating known variant information in an appropriate manner, sensitive variant calling is possible at a low cost. Availability and implementation: Implementation can be found in our public code repository https://github.com/namsyvo/IVC. Supplementary information: Supplementary data are available at Bioinformatics online. Nam Sy Vo, Vinhthuy T. Phan |
Bioinform. | 1 |
| 2015 | How genome complexity can explain the difficulty of aligning reads to genomesabstractBACKGROUND: Although it is frequently observed that aligning short reads to genomes becomes harder if they contain complex repeat patterns, there has not been much effort to quantify the relationship between complexity of genomes and difficulty of short-read alignment. Existing measures of sequence complexity seem unsuitable for the understanding and quantification of this relationship. RESULTS: We investigated several measures of complexity and found that length-sensitive measures of complexity had the highest correlation to accuracy of alignment. In particular, the rate of distinct substrings of length k, where k is similar to the read length, correlated very highly to alignment performance in terms of precision and recall. We showed how to compute this measure efficiently in linear time, making it useful in practice to estimate quickly the difficulty of alignment for new genomes without having to align reads to them first. We showed how the length-sensitive measures could provide additional information for choosing aligners that would align consistently accurately on new genomes. CONCLUSIONS: We formally established a connection between genome complexity and the accuracy of short-read aligners. The relationship between genome complexity and alignment accuracy provides additional useful information for selecting suitable aligners for new genomes. Further, this work suggests that the complexity of genomes sometimes should be thought of in terms of specific computational problems, such as the alignment of short reads to genomes. Vinhthuy T. Phan, Quang Tran 0002, Nam Sy Vo |
BMC Bioinform. | 4 |
| 2015 | A linear model for predicting performance of short-read aligners using genome complexityabstractBackground The effectiveness and accuracy of aligning short reads to genomes have an important impact on many applications that rely on next-generation sequencing data. The computational requirements and material cost for aligning largescale short reads to genomes is also expensive. To prevent wasted time and resources for aligning short reads, we investigated the different measures of genome complexity [1] that correlated best to the performance of alignment to propose a linear model for each aligning method [2]. Quang Tran 0002, Nam Sy Vo, Vinhthuy T. Phan |
BMC Bioinform. | 3 |
| 2015 | Improving variant calling by incorporating known genetic variants into read alignment
Nam Sy Vo, Vinhthuy T. Phan |
BMC Bioinform. | 1 |
| 2014 | Exploiting the bootstrap method to analyze patterns of gene expressionabstractBackground High-throughput technologies like microarrays or the recent RNA-Seq provide large amounts of data for gene expression studies. Although there have been diverse methods to design gene-expression experiments and analyze gene-expression data, the prediction of true patterns of gene expression in case of having few samples remains a challenging problem [1,2]. Materials and methods We propose a method to predict response patterns of gene expression studies in the case of small sample size using a bootstrap method [3]. Our approach adopts partially order sets (posets) to represent gene patterns, which are determined based on pairwise comparisons [4]. Results We show that patterns that are not linearly orderable cannot be true patterns of gene response to treatments. From this result, we propose a strategy using bootstrap resampling to infer true responses of non-linearly-orderable patterns. Our experiments showed that this method produced gene lists with more biological functional enrichment than those obtained without bootstrap resampling. Conclusions Our method is useful in designing cost-effective experiments with small sample sizes. Researchers can still use a small sample size to determine true patterns for most genes. For highly-variantly expressed genes, their true patterns can be identified using the proposed method. Nam Sy Vo, Vinhthuy T. Phan |
BMC Bioinform. | 1 |
| 2014 | Exploiting dependencies of pairwise comparison outcomes to predict patterns of gene responseabstractThe analysis of gene expression has played an important role in medical and bioinformatics research. Although it is known that a large number of samples is needed to determine the patterns of gene expression accurately, practical designs of gene expression studies occasionally have insufficient numbers of samples, making it difficult to ascertain true response patterns of variantly expressed genes. We describe an approach to cope with the challenge of predicting true orders of gene response to treatments. We show that true patterns of gene response must be orderable sets. In experiments with few samples, we modify the conventional pairwise comparison tests and increase the significance level α intelligently to deduce orderable patterns, which are most likely true orders of gene response. Additionally, motivated by the fact that a gene can be involved in multiple biological functions, our method further resamples experimental replicates and predicts multiple response patterns for each gene. Using a gene expression data set of Sprague-Dawley rats treated with chemopreventive chemical compounds and DAVID to annotate and validate gene sets, we showed that compared to the conventional method of fixing α , this method increased enrichment significantly. A comparison with hierarchical clustering showed that gene clusters labelled by response patterns produced by our method were much more enriched. One of the clusters contained 3 transcription factors, which hierarchical clustering failed to place into one cluster, that have been found to participate in multiple biological networks. One of the transcription factors is known to play an important role in pathways affected by the studied chemical compounds. This method can be useful in designing cost-effective experiments with small sample sizes. Patterns of highly-variantly expressed genes can be predicted by varying α intelligently. Furthermore, clusters are labeled meaningfully with patterns that describe precisely how genes in such clusters respond to treatments. Nam Sy Vo, Vinhthuy T. Phan |
BMC Bioinform. | 1 |
| 2014 | An integrated approach for SNP calling based on population of genomesabstractBackground The identification of genetic variants such as single nucleotide polymorphisms (SNPs) is a critical step in many applications based on NGS technologies [1]. Although many SNP calling programs have been developed, it is still challenging to accurately call SNPs, especially when coverage level is low [2]. Moreover, the determination of SNPs, which is performed through many separate steps, requires a careful selection of a diverse set of tools [3,4]. This can lead to several disadvantages, for example, one cannot incorporate information from the read alignment step into the SNP calling step or vice versa to help improve accuracy of called SNPs. Materials and methods We propose a novel integrated approach to detect more true SNPs while calling fewer false positives. Different from current methods that perform read alignment and SNP calling steps separately, our method combines them methodologically to improve the accuracy of SNP identification. To effectively exploit information from a population of genomes, databases of confirmed SNPs, such as dbSNP, are employed in both aligning reads to references as well as calling SNPs. This strategy allows us to develop a novel algorithm to align reads to references that can differentiate sequencing errors from SNPs. Results Based on this result, the method can call SNPs accurately and effectively even with low-coverage sequencing data. Our results on simulated data show that the method is able to call SNPs with very high precision and recall rate with low-coverage datasets. Conclusions With the existence of databases of confirmed SNPs for large amounts of sequenced species, our approach provides a promising method to call accurate SNP information even with low-coverage sequencing data. This approach can also help researchers facilitate the determination of SNPs by using an integrated SNP calling tool. Nam Sy Vo, Quang Tran 0002, Vinhthuy T. Phan |
BMC Bioinform. | 1 |
| 2013 | Exploiting Dependencies of Patterns in Gene Expression Analysis Using Pairwise Comparisons
Nam Sy Vo, Vinhthuy T. Phan |
ISBRA | 1 |
| 2013 | Using partially ordered sets to represent and predict true patterns of gene response to treatmentsabstractAdvances in biotechnology have empowered high-throughput measurement of gene expression levels for tens of thousands of genes simultaneously. This means that one sample size must be used for all genes in most experimental designs [ 1 , 2 ], which implies that patterns of response of highly variantly expressed genes might not be measured accurately. Response patterns of gene expression data with multiple treatments have been characterized using post hoc pairwise comparisons by several researchers [ 3 , 4 ]. Nevertheless, these researchers did not address how to cope with highly variantly expressed genes with inaccurate patterns due to having too few experimental samples. We show that dependencies of pairwise comparison outcomes in post hoc calculations can be exploited to infer true response patterns of genes with inaccurate patterns due to having too few experimental samples. Characterizing such response patterns as partially ordered sets, we show that linearly orderable patterns are more likely true patterns and those that are not linearly orderable cannot be true patterns. We propose a strategy to predict most likely linearly orderable extensions of such patterns. Using microarray data of rats' liver cells, we showed that this approach yielded more and better functionally enriched gene lists than a conventional approach. This approach opens up opportunities to design cost-effective experiments, in which only a conservatively large sample size is needed to collect expression levels of almost all genes. For most genes, such a sample size is sufficient. For highly variantly expressed genes, our method can help infer true response patterns. Nam Sy Vo, Vinhthuy T. Phan |
BMC Bioinform. | 1 |
| 2011 | Robust Visual Tracking Using Randomized Forest and Online Appearance Model
Nam Sy Vo, Thang Ba Dinh, Tien Ba Dinh |
ACIIDS (2) | 1 |
| 2011 | Context tracker: Exploring supporters and distracters in unconstrained environmentsabstractVisual tracking in unconstrained environments is very challenging due to the existence of several sources of varieties such as changes in appearance, varying lighting conditions, cluttered background, and frame-cuts. A major factor causing tracking failure is the emergence of regions having similar appearance as the target. It is even more challenging when the target leaves the field of view (FoV) leading the tracker to follow another similar object, and not reacquire the right target when it reappears. This paper presents a method to address this problem by exploiting the context on-the-fly in two terms: Distracters and Supporters. Both of them are automatically explored using a sequential randomized forest, an online template-based appearance model, and local features. Distracters are regions which have similar appearance as the target and consistently co-occur with high confidence score. The tracker must keep tracking these distracters to avoid drifting. Supporters, on the other hand, are local key-points around the target with consistent co-occurrence and motion correlation in a short time span. They play an important role in verifying the genuine target. Extensive experiments on challenging real-world video sequences show the tracking improvement when using this context information. Comparisons with several state-of-the-art approaches are also provided. Thang Ba Dinh, Nam Sy Vo, Gérard G. Medioni |
CVPR | 2 |
| 2011 | High resolution face sequences from a PTZ network cameraabstractWe propose here to acquire high resolution sequences of a person's face using a pan-tilt-zoom (PTZ) network camera. This capability should prove helpful in forensic analysis of video sequences as frames containing faces are tagged, and within a frame, windows containing faces can be retrieved. The system starts in pedestrian detector mode, where the lens angle is set widest, and detects people using a pedestrian detector module. The camera then changes to the region of interest (ROI) focusing mode where the parameters are automatically tuned to put the upper body of the detected person, where the face should appear, in the field of view (FOV). Then, in the face detection mode, the face is detected using a face detector module, and the system switches to an active tracking mode consisting a control loop to actively follow the detected face with two different modules: a tracker to track the face in the image, and a camera control module to adjust the camera parameters. During this loop, our tracker learns online the face appearance in multiple views under all condition changes. It runs robustly at 15 fps and is able to reacquire the face of interest after total occlusion or leaving FOV. We compare our tracker with various state-of-the-art tracking methods in terms of precision and running time performance. Extensive experiments in challenging indoor and outdoor conditions are also demonstrated to validate the complete system. Thang Ba Dinh, Nam Sy Vo, Gérard G. Medioni |
FG | 2 |
| 2011 | mDAG: a web-based tool for analyzing microarray data with multiple treatmentsabstractIn microarray experiments involving multiple treatments, pairwise comparisons between all pairs of treatments are desirable but expensive. To cope with this, we previously introduced a method that performed all pairwise comparisons in a post hoc manner. This method employs directed graphs to represent gene response to pairs of treatments. It has been applied and found useful in identifying and differentiating genes sharing similar functional pathways [ 1 , 2 ]. mDAG is a web-based software based on this method. mDAG allows users to upload microarray data in GCT format through a web interface. From this data, the application performs calculations to assign graphical patterns to genes and outputs images and textual data for further analyses. These graphical patterns carry specific meanings in terms of how genes respond to pairs of treatments. The application is implemented using Python and web2py. mDAG is available at http://cetus.cs.memphis.edu:8080/mDAG . For experiments involved multiple treatments and replicates, mDAG allows researchers to analyze and visualize in graphical representations relationships of gene interactions to all pairs of treatments. The software can be used online or off-line. Vinhthuy T. Phan, Nam Sy Vo, Thomas R. Sutter |
BMC Bioinform. | 2 |