EDBT 2026 Demo / reviewers in the wild / expert
Jeremy S. Archuleta
dblp:54/4530
· DBLP profile ↗
7ranked-venue papers
2as first author
0since 2021 · last 2010
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Distributed systems · 59% High-performance computing · 17% Storage systems · 17% | |
| Interdisciplinary, comprehensive, and emerging computing
2 papers |
Bioinformatics and computational biology · 100% |
Topics — the 9 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Distributed systems
fault tolerance |
0.1 | 1 | 2010 | MOON: MapReduce On Opportunistic eNvironments · HPDC 2010 |
Distributed systems › grid computing
volunteer computing |
0.1 | 1 | 2010 | MOON: MapReduce On Opportunistic eNvironments · HPDC 2010 |
Storage systems › file systems
distributed file system |
0.1 | 1 | 2008 | Semantics-based distributed I/O for mpiBLAST · PPoPP 2008 |
High-performance computing
distributed i/o |
0.1 | 1 | 2008 | Semantics-based distributed I/O for mpiBLAST · PPoPP 2008 |
Bioinformatics and computational biology › sequence analysis › sequence similarity search
genomic sequence search |
0.1 | 1 | 2006 | Grid applications - Parallel genomic sequence-searching on an ad-hoc grid: experiences, lessons learned, and implications · SC 2006 |
Bioinformatics and computational biology › multiple sequence alignment
parallel sequence alignment |
0.1 | 1 | 2006 | Grid applications - Parallel genomic sequence-searching on an ad-hoc grid: experiences, lessons learned, and implications · SC 2006 |
Distributed systems
grid computing |
0.1 | 1 | 2006 | Grid applications - Parallel genomic sequence-searching on an ad-hoc grid: experiences, lessons learned, and implications · SC 2006 |
Parallel and multicore computing › data-parallel programming
mapreduce |
0.0 | 1 | 2010 | MOON: MapReduce On Opportunistic eNvironments · HPDC 2010 |
Bioinformatics and computational biology › sequence analysis › sequence similarity search
sequence database search |
0.0 | 1 | 2008 | Semantics-based distributed I/O for mpiBLAST · PPoPP 2008 |
Methods — techniques the papers use, named apart from their topics
semantic compression · 0.2metadata generation · 0.2i/o worker partitioning · 0.2parallel BLAST · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2010 | MOON: MapReduce On Opportunistic eNvironmentsabstractMapReduce offers an ease-of-use programming paradigm for processing large data sets, making it an attractive model for distributed volunteer computing systems. However, unlike on dedicated resources, where MapReduce has mostly been deployed, such volunteer computing systems have significantly higher rates of node unavailability. Furthermore, nodes are not fully controlled by the MapReduce framework. Consequently, we found the data and task replication scheme adopted by existing MapReduce implementations woefully inadequate for resources with high unavailability. Heshan Lin, Xiaosong Ma, Jeremy S. Archuleta, Wu-chun Feng, Mark K. Gardner, Zhe Zhang 0005 |
HPDC | 3 |
| 2010 | Missing genes in the annotation of prokaryotic genomesabstractBACKGROUND: Protein-coding gene detection in prokaryotic genomes is considered a much simpler problem than in intron-containing eukaryotic genomes. However there have been reports that prokaryotic gene finder programs have problems with small genes (either over-predicting or under-predicting). Therefore the question arises as to whether current genome annotations have systematically missing, small genes. RESULTS: We have developed a high-performance computing methodology to investigate this problem. In this methodology we compare all ORFs larger than or equal to 33 aa from all fully-sequenced prokaryotic replicons. Based on that comparison, and using conservative criteria requiring a minimum taxonomic diversity between conserved ORFs in different genomes, we have discovered 1,153 candidate genes that are missing from current genome annotations. These missing genes are similar only to each other and do not have any strong similarity to gene sequences in public databases, with the implication that these ORFs belong to missing gene families. We also uncovered 38,895 intergenic ORFs, readily identified as putative genes by similarity to currently annotated genes (we call these absent annotations). The vast majority of the missing genes found are small (less than 100 aa). A comparison of select examples with GeneMark, EasyGene and Glimmer predictions yields evidence that some of these genes are escaping detection by these programs. CONCLUSIONS: Prokaryotic gene finders and prokaryotic genome annotations require improvement for accurate prediction of small genes. The number of missing gene families found is likely a lower bound on the actual number, due to the conservative criteria used to determine whether an ORF corresponds to a real gene. Andrew S. Warren, Jeremy S. Archuleta, Wu-chun Feng, João Carlos Setubal |
BMC Bioinform. | 2 |
| 2010 | Global-scale distributed I/O with ParaMEDICabstractAbstract Achieving high performance for distributed I/O on a wide‐area network continues to be an elusive holy grail. Despite enhancements in network hardware as well as software stacks, achieving high‐performance remains a challenge. In this paper, our worldwide team took a completely new and non‐traditional approach to distributed I/O, calledParaMEDIC: Parallel Metadata Environment for Distributed I/O and Computing, by utilizing application‐specifictransformationof data to orders of magnitude smaller metadata before performing the actual I/O. Specifically, this paper details our experiences in deploying a large‐scale system to facilitate the discovery of missing genes and constructing a genome similarity tree by encapsulating the mpiBLAST sequence‐search algorithm into ParaMEDIC. The overall project involved nine computational sites spread across the U.S. and generated more than a petabyte of data that was ‘teleported’ to a large‐scale facility in Tokyo for storage. Copyright © 2010 John Wiley & Sons, Ltd. Pavan Balaji, Wu-chun Feng, Heshan Lin, Jeremy S. Archuleta, Satoshi Matsuoka, Andrew S. Warren, João Carlos Setubal, Ewing L. Lusk, Rajeev Thakur, Ian T. Foster, Daniel S. Katz, Shantenu Jha, K. Shinpaugh, Susan Coghlan, Daniel A. Reed |
Concurr. Comput. Pract. Exp. | 4 |
| 2009 | Multi-dimensional characterization of temporal data mining on graphics processorsabstractThrough the algorithmic design patterns of data parallelism and task parallelism, the graphics processing unit (GPU) offers the potential to vastly accelerate discovery and innovation across a multitude of disciplines. For example, the exponential growth in data volume now presents an obstacle for high-throughput data mining in fields such as neuroscience and bioinformatics. As such, we present a characterization of a MapReduced-based data-mining application on a general-purpose GPU (GPGPU). Using neuroscience as the application vehicle, the results of our multi-dimensional performance evaluation show that a ldquoone-size-fits-allrdquo approach maps poorly across different GPGPU cards. Rather, a high-performance implementation on the GPGPU should factor in the 1) problem size, 2) type of GPU, 3) type of algorithm, and 4) data-access method when determining the type and level of parallelism. To guide the GPGPU programmer towards optimal performance within such a broad design space, we provide eight general performance characterizations of our data-mining application. Jeremy S. Archuleta, Yong Cao 0003, Thomas Scogland, Wu-chun Feng |
IPDPS | 1 |
| 2008 | Semantics-based distributed I/O for mpiBLASTabstractBLAST is a widely used software toolkit for genomic sequence search. mpiBLAST is a freely available, open-source parallelization of BLAST that uses database segmentation to allow different worker processes to search (in parallel) unique segments of the database. After searching, the workers write their output to a filesystem. While mpiBLAST has been shown to achieve high performance in clusters with fast local filesystems, its I/O processing remains a concern for scalability, especially in systems having limited I/O capabilities such as distributed filesystems spread across a wide-area network. Thus, we present ParaMEDIC---a novel environment that uses application-specific semantic information to compress I/O data and improve performance in distributed environments. Specifically, for mpiBLAST, ParaMEDIC partitions worker processes into compute and I/O workers. Compute workers, instead of directly writing the output to the filesystem, the workers process the output using semantic knowledge about the application to generate metadata and write the metadata to the filesystem. I/O workers, which physically reside closer to the actual storage, then process this metadata to re-create the actual output and write it to the filesystem. This approach allows ParaMEDIC to reduce I/O time, thus accelerating mpiBLAST by as much as 25-fold. Pavan Balaji, Wu-chun Feng, Jeremy S. Archuleta, Heshan Lin, Rajkumar Kettimuthu, Rajeev Thakur, Xiaosong Ma |
PPoPP | 3 |
| 2007 | A Maintainable Software Architecture for Fast and Modular Bioinformatics Sequence SearchabstractBioinformaticists use the Basic Local Alignment Search Tool (BLAST) to characterize an unknown sequence by comparing it against a database of known sequences, thus detecting evolutionary relationships and biological properties. mpiBLAST is a widely-used, high-performance, open-source parallelization of BLAST that runs on a computer cluster delivering super-linear speedups. However, the Achilles heel of mpiBLAST is its lack of modularity, thus adversely affecting maintainability and extensibility. Alleviating this shortcoming requires an architectural refactoring to improve maintenance and extensibility while preserving high performance. Toward that end, this paper evaluates five different software architectures and details how each satisfies our design objectives. In addition, we introduce a novel approach to using mixin layers to enable mixing-and-matching of modules in constructing sequence-search applications for a variety of high-performance computing systems. Our design, which we call "mixin layers with refined roles", utilizes mixin layers to separate functionality into complementary modules and the refined roles in each layer improve the inherently modular design by precipitating flexible and structured parallel development, a necessity for an open-source application. We believe that this new software architecture for mpiBLAST-2.0 will benefit both the users and developers of the package and that our evaluation of different software architectures will be of value to other software engineers faced with the challenges of creating maintainable and extensible, high-performance, bioinformatics software. Jeremy S. Archuleta, Eli Tilevich, Wu-chun Feng |
ICSM | 1 |
| 2006 | Grid applications - Parallel genomic sequence-searching on an ad-hoc grid: experiences, lessons learned, and implicationsabstractThe Basic Local Alignment Search Tool (BLAST) allows bioinformaticists to characterize an unknown sequence by comparing it against a database of known sequences. The similarity between sequences enables biologists to detect evolutionary relationships and infer biological properties of the unknown sequence.mpiBLAST, our parallel BLAST, decreases the search time of a 300 KB query on the current NT database from over two full days to under 10 minutes on a 128-processor cluster and allows larger query files to be compared. Consequently, we propose to compare the largest query available, the entire NT database, against the largest database available, the entire NT database. The result of this comparison will provide critical information to the biology community, including insightful evolutionary, structural, and functional relationships between every sequence and family in the NT database.Preliminary projections indicated that to complete the above task in a reasonable length of time required more processors than were available to us at a single site. Hence, we assembled GreenGene, an ad-hoc grid that was constructed "on the fly" from donated computational, network, and storage resources during last year's SC|05. GreenGene consisted of 3048 processors from machines that were distributed across the United States. This paper presents a case study of mpiBLAST on GreenGene --- specifically, a pre-run characterization of the computation, the hardware and software architectural design, experimental results, and future directions. Mark K. Gardner, Wu-chun Feng, Jeremy S. Archuleta, Heshan Lin, Xiaosong Ma |
SC | 3 |