VLDB 2026 Research / reviewers in the wild / expert
Han-Jen Lin
dblp:241/2124
· DBLP profile ↗
3ranked-venue papers
0as first author
0since 2021 · last 2020
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 3
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
2 papers |
Bioinformatics and computational biology · 100% |
Topics — the 4 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology
statistical genetics |
0.4 | 1 | 2020 | SparkINFERNO: a scalable high-throughput pipeline for inferring molecular mechanisms of non-coding genetic variants · Bioinform. 2020 |
Bioinformatics and computational biology › genomics
variant calling |
0.4 | 1 | 2019 | VCPA: genomic variant calling pipeline and data management tool for Alzheimer's Disease Sequencing Project · Bioinform. 2019 |
Bioinformatics and computational biology
functional genomics |
0.1 | 1 | 2020 | SparkINFERNO: a scalable high-throughput pipeline for inferring molecular mechanisms of non-coding genetic variants · Bioinform. 2020 |
Bioinformatics and computational biology › gene regulation › regulatory element
regulatory element annotation |
0.1 | 1 | 2020 | SparkINFERNO: a scalable high-throughput pipeline for inferring molecular mechanisms of non-coding genetic variants · Bioinform. 2020 |
Methods — techniques the papers use, named apart from their topics
giggle genomic indexing · 0.4apache spark · 0.4workflow description language · 0.4GATK · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2020 | SparkINFERNO: a scalable high-throughput pipeline for inferring molecular mechanisms of non-coding genetic variantsabstractSUMMARY: We report Spark-based INFERence of the molecular mechanisms of NOn-coding genetic variants (SparkINFERNO), a scalable bioinformatics pipeline characterizing non-coding genome-wide association study (GWAS) association findings. SparkINFERNO prioritizes causal variants underlying GWAS association signals and reports relevant regulatory elements, tissue contexts and plausible target genes they affect. To achieve this, the SparkINFERNO algorithm integrates GWAS summary statistics with large-scale collection of functional genomics datasets spanning enhancer activity, transcription factor binding, expression quantitative trait loci and other functional datasets across more than 400 tissues and cell types. Scalability is achieved by an underlying API implemented using Apache Spark and Giggle-based genomic indexing. We evaluated SparkINFERNO on large GWASs and show that SparkINFERNO is more than 60 times efficient and scales with data size and amount of computational resources. AVAILABILITY AND IMPLEMENTATION: SparkINFERNO runs on clusters or a single server with Apache Spark environment, and is available at https://bitbucket.org/wanglab-upenn/SparkINFERNO or https://hub.docker.com/r/wanglab/spark-inferno. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Pavel P. Kuksa, Chien-Yueh Lee, Alexandre Amlie-Wolf, Prabhakaran Gangadharan, Elizabeth E. Mlynarski, Yi-Fan Chou, Han-Jen Lin, Heather Issen, Emily Greenfest-Allen, Otto Valladares, Yuk Yee Leung, Li-San Wang |
Bioinform. | 7 |
| 2019 | VCPA: genomic variant calling pipeline and data management tool for Alzheimer's Disease Sequencing ProjectabstractBioinformatics, https://doi.org/10.1093/bioinformatics/bty894 In the supplementary material, the phrase “The pairwise variant disconcordance rate” has been changed to “The pairwise variant concordance rate”. The updated supplementary file is now available online. Yuk Yee Leung, Otto Valladares, Yi-Fan Chou, Han-Jen Lin, Amanda Kuzma, Laura Cantwell, Liming Qu, Prabhakaran Gangadharan, William J. Salerno, Gerard D. Schellenberg, Li-San Wang |
Bioinform. | 4 |
| 2019 | VCPA: genomic variant calling pipeline and data management tool for Alzheimer's Disease Sequencing ProjectabstractSUMMARY: We report VCPA, our SNP/Indel Variant Calling Pipeline and data management tool used for the analysis of whole genome and exome sequencing (WGS/WES) for the Alzheimer's Disease Sequencing Project. VCPA consists of two independent but linkable components: pipeline and tracking database. The pipeline, implemented using the Workflow Description Language and fully optimized for the Amazon elastic compute cloud environment, includes steps from aligning raw sequence reads to variant calling using GATK. The tracking database allows users to view job running status in real time and visualize >100 quality metrics per genome. VCPA is functionally equivalent to the CCDG/TOPMed pipeline. Users can use the pipeline and the dockerized database to process large WGS/WES datasets on Amazon cloud with minimal configuration. AVAILABILITY AND IMPLEMENTATION: VCPA is released under the MIT license and is available for academic and nonprofit use for free. The pipeline source code and step-by-step instructions are available from the National Institute on Aging Genetics of Alzheimer's Disease Data Storage Site (http://www.niagads.org/VCPA). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yuk Yee Leung, Otto Valladares, Yi-Fan Chou, Han-Jen Lin, Amanda Kuzma, Laura Cantwell, Liming Qu, William J. Salerno, Gerard D. Schellenberg, Li-San Wang |
Bioinform. | 4 |