VLDB 2026 Research / reviewers in the wild / expert
Weijia Xu
dblp:68/4886
· DBLP profile ↗
52ranked-venue papers
18as first author
17since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 29 · 10 first-author · 2 since 2021Artificial intelligence and machine learning · 28 · 12 first-author · 9 since 2021Databases, data management, data science and information retrieval · 23 · 7 first-author · 2 since 2021Systems, architecture and hardware · 7 · 5 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Novel Goal Creation and Evaluation in Open-Ended Games
Graham Todd, Junyi Chu, Guy Davidson, Weijia Xu |
CogSci | 4 |
| 2025 | fLSA: Learning Semantic Structures in Document Collections Using Foundation ModelsabstractHumans can learn to solve new tasks by inducing high-level strategies from example solutions to similar problems and then adapting these strategies to solve unseen problems.Can we use large language models to induce such high-level structure from example documents or solutions?We introduce fLSA, a foundationmodel-based Latent Semantic Analysis method that iteratively clusters and tags document segments based on document-level contexts.These tags can be used to model the latent structure of given documents and for hierarchical sampling of new texts.Our experiments on story writing, math, and multi-step reasoning datasets demonstrate that fLSA tags are more informative in reconstructing the original texts than existing tagging methods.Moreover, when used for hierarchical sampling, fLSA tags help expand the output space in the right directions that lead to correct solutions more often than direct sampling and hierarchical sampling with existing tagging methods. Weijia Xu, Nebojsa Jojic, Nicolas Le Roux |
EMNLP | 1 |
| 2025 | Learning to Solve Complex Problems via Dataset DecompositionabstractCurriculum learning is a class of training strategies that organizes the data being exposed to a model by difficulty, gradually from simpler to more complex examples.
This research explores a reverse curriculum generation approach that recursively decomposes complex datasets into simpler, more learnable components.
We propose a teacher-student framework where the teacher is equipped with the ability to reason step-by-step, which is used to recursively generate easier versions of examples, enabling the student model to progressively master difficult tasks. We propose a novel scoring system to measure data difficulty based on its structural complexity and conceptual depth, allowing curriculum construction over decomposed data.
Experiments on math datasets (MATH and AIME) and code generation datasets demonstrate that models trained with curricula generated by our approach exhibit superior performance compared to standard training on original datasets. Wanru Zhao, Lucas Caccia, Zhengyan Shi, Minseon Kim, Weijia Xu, Alessandro Sordoni |
NeurIPS | 5 |
| 2024 | GENEVA: GENErating and Visualizing branching narratives using LLMsabstractDialogue-based Role Playing Games (RPGs) require powerful storytelling. The narratives of these may take years to write and typically involve a large creative team. In this work, we demonstrate the potential of large generative text models to assist this process. GENEVA, a prototype tool, generates a rich narrative graph with branching and reconverging storylines that match a high-level narrative description and constraints provided by the designer. A large language model (LLM), GPT-4, is used to generate the branching narrative and to render it in a graph format in a two-step process. We illustrate the use of GENEVA in generating new branching narratives for four well-known stories under different contextual constraints. This tool has the potential to assist in game development, simulations, and other applications with game-like properties. Jorge Leandro, Sudha Rao, Michael Xu, Weijia Xu, Nebojsa Jojic, Chris Brockett, William B. Dolan |
CoG | 4 |
| 2024 | Player-Driven Emergence in LLM-Driven Game NarrativeabstractWe explore how interaction with large language models (LLMs) can give rise to emergent behaviors, empowering players to participate in the evolution of game narratives. Our testbed is a text-adventure game in which players attempt to solve a mystery under a fixed narrative premise, but can freely interact with non-player characters generated by GPT-4, a large language model. We recruit 28 gamers to play the game and use GPT-4 to automatically convert the game logs into a node-graph representing the narrative in the player’s gameplay. We find that through their interactions with the non-deterministic behavior of the LLM, players are able to discover interesting new emergent nodes that were not a part of the original narrative but have potential for being fun and engaging. Players that created the most emergent nodes tended to be those that often enjoy games that facilitate discovery, exploration and experimentation. Jessica Quaye, Sudha Rao, Weijia Xu, Portia Botchway, Chris Brockett, Nebojsa Jojic, Gabriel DesGarennes, Ken Lobb, Michael Xu, Jorge Leandro, Claire Jin, William B. Dolan |
CoG | 4 |
| 2024 | Reprompting: Automated Chain-of-Thought Prompt Inference Through Gibbs SamplingabstractWe introduce Reprompting, an iterative sampling algorithm that automatically learns the Chain-of-Thought (CoT) recipes for a given task without human intervention. Through Gibbs sampling, Reprompting infers the CoT recipes that work consistently well for a set of training samples by iteratively sampling new recipes using previously sampled recipes as parent prompts to solve other training problems. We conduct extensive experiments on 20 challenging reasoning tasks. Results show that Reprompting outperforms human-written CoT prompts substantially by +9.4 points on average. It also achieves consistently better performance than the state-of-the-art prompt optimization and decoding algorithms. Weijia Xu, Andrzej Banburski-Fahey, Nebojsa Jojic |
ICML | 1 |
| 2023 | Understanding and Detecting Hallucinations in Neural Machine Translation via Model IntrospectionabstractAbstract Neural sequence generation models are known to “hallucinate”, by producing outputs that are unrelated to the source text. These hallucinations are potentially harmful, yet it remains unclear in what conditions they arise and how to mitigate their impact. In this work, we first identify internal model symptoms of hallucinations by analyzing the relative token contributions to the generation in contrastive hallucinated vs. non-hallucinated outputs generated via source perturbations. We then show that these symptoms are reliable indicators of natural hallucinations, by using them to design a lightweight hallucination detector which outperforms both model-free baselines and strong classifiers based on quality estimation or large pre-trained models on manually annotated English-Chinese and German-English translation test beds. Weijia Xu, Sweta Agrawal, Eleftheria Briakou, Marianna J. Martindale, Marine Carpuat |
Trans. Assoc. Comput. Linguistics | 1 |
| 2022 | Accelerating Deep Learning Training Through Transparent Storage TieringabstractWe present Monarch,a framework-agnostic storage middleware that transparently employs storage tiering to accelerate Deep Learning (DL) training. It leverages existing storage tiers of modern supercomputers (i.e., compute node's local storage and shared parallel file system (PFS)), while considering the I/O patterns of DL frameworks to improve data placement across tiers. Monarchaims at accelerating DL training and decreasing the I/O pressure imposed over the PFS. We apply Monarchto TensorFlow and PyTorch, while validating its performance and applicability under different models and dataset sizes. Results show that, even when the training dataset can only be partially stored at local storage, Monarchreduces TensorFlow's and PyTorch's training time by up to 28% and 37% for I/O-intensive models, respectively. Furthermore, Monarchdecreases the number of I/O operations submitted to the PFS by up to 56%. Marco Dantas, Diogo Leitão, Peter Cui, Ricardo Macedo, Xinlian Liu, Weijia Xu, João Paulo 0001 |
CCGRID | 6 |
| 2022 | Constrained Regeneration for Cross-Lingual Query-Focused Extractive SummarizationabstractQuery-focused summaries of foreign-language, retrieved documents can help a user understand whether a document is actually relevant to the query term. A standard approach to this problem is to first translate the source documents and then perform extractive summarization to find relevant snippets. However, in a cross-lingual setting, the query term does not necessarily appear in the translations of relevant documents. In this work, we show that constrained machine translation and constrained post-editing can improve human relevance judgments by including a query term in a summary when its translation appears in the source document. We also present several strategies for selecting only certain documents for regeneration which yield further improvements Elsbeth Turcan, David Wan, Faisal Ladhak, Petra Galuscáková, Sukanta Sen, Svetlana Tchistiakova, Weijia Xu, Marine Carpuat, Kenneth Heafield, Douglas W. Oard, Kathy McKeown |
COLING | 7 |
| 2022 | BatchLens: A Visualization Approach for Analyzing Batch Jobs in Cloud SystemsabstractCloud systems are becoming increasingly powerful and complex. It is highly challenging to identify anomalous execution behaviors and pinpoint problems by examining the overwhelming intermediate results/states in complex application workflows. Domain scientists urgently need a friendly and functional interface to understand the quality of the computing services and the performance of their applications in real time. To meet these needs, we explore data generated by job schedulers and investigate general performance metrics (e.g., utilization of CPU, memory and disk I/O). Specifically, we propose an interactive visual analytics approach, BatchLens, to provide both providers and users of cloud service with an intuitive and effective way to explore the status of system batch jobs and help them conduct root-cause analysis of anomalous behaviors in batch jobs. We demonstrate the effectiveness of BatchLens through a case study on the public Alibaba bench workload trace datasets. Shaolun Ruan, Yong Wang 0021, Hailong Jiang, Weijia Xu, Qiang Guan |
DATE | 4 |
| 2022 | Deep Neural Network Training With Distributed K-FACabstractScaling deep neural network training to more processors and larger batch sizes is key to reducing end-to-end training time; yet, maintaining comparable convergence and hardware utilization at larger scales is challenging. Increases in training scales have enabled natural gradient optimization methods as a reasonable alternative to stochastic gradient descent and variants thereof. Kronecker-factored Approximate Curvature (K-FAC), a natural gradient method, preconditions gradients with an efficient approximation of the Fisher Information Matrix to improve per-iteration progress when optimizing an objective function. Here we propose a scalable K-FAC algorithm and investigate K-FAC’s applicability in large-scale deep neural network training. Specifically, we explore layer-wise distribution strategies, inverse-free second-order gradient evaluation, and dynamic K-FAC update decoupling, with the goal of preserving convergence while minimizing training time. We evaluate the convergence and scaling properties of our K-FAC gradient preconditioner, for image classification, object detection, and language modeling applications. In all applications, our implementation converges to baseline performance targets in 9–25% less time than the standard first-order optimizers on GPU clusters across a variety of scales. J. Gregory Pauloski, Lei Huang 0019, Weijia Xu, Kyle Chard, Ian T. Foster, Zhao Zhang 0007 |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2021 | Tracking Property Ownership Variance and Forecasting Housing Price with Machine Learning and Deep LearningabstractBig data and its production, management, and utilization are essential components in smart city planning. This paper presents a research framework for applying machine learning and deep learning using multiple big data sets on real estate. We built ensemble machine learning models to track property ownership variance in Austin, TX, USA. Then, the study employed the long short term memory (LSTM) model as a deep learning approach to forecasting property value in the same area. For model validation, root mean squared errors were calculated in both models. To avoid underfitting or overfitting of LSTM, we experimented with specific parameters settings. Bagging-based random forest machine learning model outperformed other ensemble machine learning models. Regarding property ownership variance, the Random Forest model’s highest feature importance generally comprised the race, age of residents, land use and built environment factors, number of schools, and neighborhood location. Our LSTM model predicted Austin to retain a rising curve in housing prices and identified which part of Austin experiences an increase or decrease in property value. The predictive models may help city planners to quantify and gain insights on future impacts of developing neighborhoods. Junfeng Jiao, Seung-Jun Choi, Weijia Xu |
IEEE BigData | 3 |
| 2021 | MONARCH: Hierarchical Storage Management for Deep Learning FrameworksabstractDue to convenience and usability, many deep learning (DL) jobs resort to the available shared parallel file system (PFS) for storing and accessing training data when running in HPC environments. Under such a scenario, however, where multiple I/O-intensive applications operate concurrently, the PFS can quickly get saturated with simultaneous storage requests and become a critical performance bottleneck, leading to throughput variability and performance loss.We present MONARCH, a framework-agnostic middleware for hierarchical storage management. This solution leverages the existing storage tiers present at modern supercomputers (e.g., compute node’s local storage, PFS) to improve DL training performance and alleviate the current I/O pressure of the shared PFS.We validate the applicability of our approach by developing and integrating an early prototype with the TensorFlow DL framework. Results show that MONARCH can reduce I/O operations submitted to the shared PFS by up to 45%, decreasing training time by 24% and 12%, for I/O-intensive models, namely LeNet and AlexNet. Marco Dantas, Diogo Leitão, Cláudia Correia, Ricardo Macedo, Weijia Xu, João Paulo 0001 |
CLUSTER | 5 |
| 2021 | The Case for Storage Optimization Decoupling in Deep Learning FrameworksabstractDeep Learning (DL) training requires efficient access to large collections of data, leading DL frameworks to implement individual I/O optimizations to take full advantage of storage performance. However, these optimizations are intrinsic to each framework, limiting their applicability and portability across DL solutions, while making them inefficient for scenarios where multiple applications compete for shared storage resources.We argue that storage optimizations should be decoupled from DL frameworks and moved to a dedicated storage layer. To achieve this, we propose a new Software-Defined Storage architecture for accelerating DL training performance. The data plane implements self-contained, generally applicable I/O optimizations, while the control plane dynamically adapts them to cope with workload variations and multi-tenant environments.We validate the applicability and portability of our approach by developing and integrating an early prototype with the TensorFlow and PyTorch frameworks. Results show that our I/O optimizations significantly reduce DL training time by up to 54% and 63% for TensorFlow and PyTorch baseline configurations, while providing similar performance benefits to framework-intrinsic I/O mechanisms provided by TensorFlow. Ricardo Macedo, Cláudia Correia, Marco Dantas, Cláudia Brito, Weijia Xu, Yusuke Tanimura, Jason H. Haga, João Paulo 0001 |
CLUSTER | 5 |
| 2021 | Rule-based Morphological Inflection Improves Neural Terminology TranslationabstractCurrent approaches to incorporating terminology constraints in machine translation (MT) typically assume that the constraint terms are provided in their correct morphological forms.This limits their application to real-world scenarios where constraint terms are provided as lemmas.In this paper, we introduce a modular framework for incorporating lemma constraints in neural MT (NMT) in which linguistic knowledge and diverse types of NMT models can be flexibly applied.It is based on a novel cross-lingual inflection module that inflects the target lemma constraints based on the source context.We explore linguistically motivated rule-based and data-driven neuralbased inflection modules and design English-German health and English-Lithuanian news test suites to evaluate them in domain adaptation and low-resource MT settings.Results show that our rule-based inflection module helps NMT models incorporate lemma constraints more accurately than a neural module and outperforms the existing end-to-end approach with lower training costs. 1 Weijia Xu, Marine Carpuat |
EMNLP (1) | 1 |
| 2021 | Improved incremental local outlier detection for data streams based on the landmark window model
Aihua Li, Weijia Xu, Zhidong Liu, Yong Shi 0001 |
Knowl. Inf. Syst. | 2 |
| 2021 | EDITOR: an Edit-Based Transformer with Repositioning for Neural Machine Translation with Soft Lexical ConstraintsabstractAbstract We introduce an Edit-Based TransfOrmer with Repositioning (EDITOR), which makes sequence generation flexible by seamlessly allowing users to specify preferences in output lexical choice. Building on recent models for non-autoregressive sequence generation (Gu et al., 2019), EDITOR generates new sequences by iteratively editing hypotheses. It relies on a novel reposition operation designed to disentangle lexical choice from word positioning decisions, while enabling efficient oracles for imitation learning and parallel edits at decoding time. Empirically, EDITOR uses soft lexical constraints more effectively than the Levenshtein Transformer (Gu et al., 2019) while speeding up decoding dramatically compared to constrained beam search (Post and Vilar, 2018). EDITOR also achieves comparable or better translation quality with faster decoding speed than the Levenshtein Transformer on standard Romanian-English, English-German, and English-Japanese machine translation tasks. Weijia Xu, Marine Carpuat |
Trans. Assoc. Comput. Linguistics | 1 |
| 2020 | A Study of Spoken Audio Processing using Machine Learning for Libraries, Archives and Museums (LAM)abstractAs the need to provide access to spoken word audio collections in libraries, archives, and museums (LAM) increases, so does the need to process them efficiently and consistently. Traditionally, audio processing involves listening to the audio files, conducting manual transcription, and applying controlled subject terms to describe them. This workflow takes significant time with each recording. In this study, we investigate if and how machine learning (ML) can facilitate processing of audio collections in a manner that corresponds with LAM best practices. We use the StoryCorps collection of oral histories "Las Historias," and fixed subjects (metadata) that are manually assigned to describe each of them. Our methodology has two main phases. First, audio files are automatically transcribed using two automatic speech recognition (ASR) methods. Next, we build different supervised ML models for label prediction using the transcription data and the existing metadata. Throughout these phases the results are analyzed quantitatively and qualitatively. The workflow is implemented within the flexible web framework IDOLS to lower technical barriers for LAM professionals. By allowing users to submit ML jobs to supercomputers, reproduce workflows, change configurations, and view and provide feedback transparently, this workflow allows users to be in sync with LAM professional values. The study has several outcomes including a comparison of the quality between different transcription methods and the impact of that quality on label prediction accuracy. The study also unveiled the limitations of using manually assigned metadata to build models, to which we suggest alternate strategies for building successful training data. Weijia Xu, Maria Esteva, Peter Cui, Eugene Castillo, Kewen Wang 0004, Hanna Robbins Hopkins, Tanya E. Clement, Aaron Choate, Ruizhu Huang |
IEEE BigData | 1 |
| 2020 | End-to-End Slot Alignment and Recognition for Cross-Lingual NLUabstractNatural language understanding (NLU) in the context of goal-oriented dialog systems typically includes intent classification and slot labeling tasks.Existing methods to expand an NLU system to new languages use machine translation with slot label projection from source to the translated utterances, and thus are sensitive to projection errors.In this work, we propose a novel end-to-end model that learns to align and predict target slot labels jointly for cross-lingual transfer.We introduce MultiATIS++, a new multilingual NLU corpus that extends the Multilingual ATIS corpus to nine languages across four language families, and evaluate our method using the corpus.Results show that our method outperforms a simple label projection method using fast-align on most languages, and achieves competitive performance to the more complex, state-of-the-art projection method with only half of the training time.We release our MultiATIS++ corpus to the community to continue future research on cross-lingual NLU. Weijia Xu, Batool Haider, Saab Mansour |
EMNLP (1) | 1 |
| 2020 | Convolutional neural network training with distributed K-FACabstractTraining neural networks with many processors can reduce time-to-solution; however, it is challenging to maintain convergence and efficiency at large scales. The Kroneckerfactored Approximate Curvature (K-FAC) was recently proposed as an approximation of the Fisher Information Matrix that can be used in natural gradient optimizers. We investigate here a scalable K-FAC design and its applicability in convolutional neural network (CNN) training at scale. We study optimization techniques such as layer-wise distribution strategies, inverse-free second-order gradient evaluation, and dynamic K-FAC update decoupling to reduce training time while preserving convergence. We use residual neural networks (ResNet) applied to the CIFAR10 and ImageNet-1k datasets to evaluate the correctness and scalability of our K-FAC gradient preconditioner. With ResNet-50 on the ImageNet-1k dataset, our distributed K-FAC implementation converges to the 75.9% MLPerf baseline in 18-25% less time than does the classic stochastic gradient descent (SGD) optimizer across scales on a GPU cluster. J. Gregory Pauloski, Zhao Zhang 0007, Lei Huang 0019, Weijia Xu, Ian T. Foster |
SC | 4 |
| 2019 | Performance Comparison of Julia Distributed Implementations of Dirichlet Process Mixture ModelsabstractThe Dirichlet process mixture model (DPMM), one of the nonparametric Bayesian mixture models, is receiving more and more attentions from the statistical learning community. It has been demonstrated its great potentials in clustering analysis. When computational complexity increases as numbers of observations and features grow, the serial algorithms of DPMM need long processing time and cannot handle large volume of data on a single machine. To improve the computational efficiency, several parallel methods and implementations were proposed and implemented with C++ and Julia programming languages by different authors and publicly available on GitHub or published as a Julia package for users to download. However, the scalability of multi-cores and multi-node has not been thoroughly evaluated and compared among different implementations, even for multiple implementations of the same proposed distributed PPMM method. We selected two recent Julia implementations of parallel sampler via sub-cluster splits method proposed by Change and Fisher and performed a scalability comparison on supercomputer clusters. This paper presents some insights on the applicability of both implementations in terms of increasing number of dimensions of the feature space and provides some potential improvement strategies on multi-node scalability. Ruizhu Huang, Weijia Xu, Yinzhi Wang, Silvia Liverani, Ann E. Stapleton |
IEEE BigData | 2 |
| 2019 | Detecting Pedestrian Crossing Events in Large Video Data from Traffic Monitoring CamerasabstractPedestrian safety on the road is a priority for transportation system managers and operators. While there are a number of treatments and technologies to effectively improve pedestrian safety, identifying the location where these are most needed remains a challenge. Mid-block locations, where safety countermeasures are often needed the most, are typically harder to monitor. Current practice often requires manual observation of candidate locations for limited time periods, leading to an identification process that is often time consuming, lags behind traffic pattern changes over time, and lacks scalability. As a result, target locations are often selected reactively, after serious traffic incidents reveal an underlying safety issue. We propose an approach to use data collected by existing traffic monitoring cameras to automatically identify pedestrian activities on the road. We propose an algorithm to detect pedestrian crossing events based on the detection of individuals on individual video frames using a deep neural network model. Resulting pedestrian locations and movement trajectories can be visualized on a background image, which is automatically extracted at the analyzed location from the video. We demonstrate and evaluate our approach with a real-world use case. The case study considered in this work uses cameras owned by the City of Austin, Texas to study pedestrian road use before and after the deployment of a pedestrian-hybrid beacon. We explore qualitative and quantitative metrics to describe pedestrian activity and corresponding changes, which may be used to prioritize the deployment of pedestrian safety solutions, or evaluate their performance. We compared the number of crossing events detected per hour with manually reviewed results from a selected day. The result shows 67 percent overall accuracy, although we observe significant variability across times-of-day. Despite observed limitations, our work illustrates how the value of existing traffic camera networks can be augmented beyond everyday traffic monitoring, and used to collect valuable information on road usage by pedestrians. Weijia Xu, Natalia Ruiz-Juri, Kelly A. Pierce, Ruizhu Huang, Joel Meyer, Jennifer C. Duthie |
IEEE BigData | 1 |
| 2019 | Quantifying the Impact of Memory Errors in Deep LearningabstractThe use of deep learning (DL) on HPC resources has become common as scientists explore and exploit DL methods to solve domain problems. On the other hand, in the coming exascale computing era, a high error rate is expected to be problematic for most HPC applications. The impact of errors on DL applications, especially DL training, remains unclear given their stochastic nature. In this paper, we focus on understanding DL training applications on HPC in the presence of silent data corruption. Specifically, we design and perform a quantification study with three representative applications by manually injecting silent data corruption errors (SDCs) across the design space and compare training results with the error-free baseline. The results show only 0.61-1.76% of SDCs cause training failures, and taking the SDC rate in modern hardware into account, the actual chance of a failure is one in thousands to millions of executions. With this quantitatively measured impact, computing centers can make rational design decisions based on their application portfolio, the acceptable failure rate, and financial constraints; for example, they might determine their confidence in the correctness of training results performed on processors without error correction code (ECC) RAM. We also discover that over 75-90% of the SDCs that cause catastrophic errors can be easily detected by a training loss in the next iteration. Thus we propose this error-aware software solution to correct catastrophic errors, as it has significantly lower time and space overhead compared to algorithm-based fault-tolerance (ABFT) and ECC. Zhao Zhang 0007, Lei Huang 0019, Ruizhu Huang, Weijia Xu, Daniel S. Katz |
CLUSTER | 4 |
| 2018 | Integrated HPC Scheduler Data Processing Workflow using Apache ZeppelinabstractBig data analytics pipeline often naturally involves the components with different programming language, various programming models, etc. And it presents steep learning curve not only on developing the tools but also on using them. Building a user friendly interface can hide these usage complexities, and provide a easy way to get the insights out of data. It's a challenging task to link those components together to make a smooth end-to-end workflow. Apache Zeppelin provides native support on multiple language and data processing backends so that different workflow components can be linked together on Zeppelin's framework. We developed a web interface for analyzing High Performance Computing data center scheduler log data through Apache Zeppelin's support on AngularJS, Spark, Python and Batch. An interactive PACE-Fast Analysis of Computational Trends (PACE-FACT) environment is built on the extension of our previous work, this environment seamlessly puts multiple log data analysis and visualization components together, and it allows to visualize the result data interactively without dealing with cumbersome command line user interface. In this work, we demonstrate that software ranking and analysis can be done through web GUI with user specified date range. And this system will be used in Georgia Institute of Technology (Georgia Tech)'s high performance computing (HPC) PACE center. Fang (Cherry) Liu, Yuanjie Sun, Adele Yunlan Sun, Weijia Xu |
IEEE BigData | 4 |
| 2018 | Enabling User Driven Big Data Application on Remote Computing ResourcesabstractDriven by the computing resource requirement, there are increasing demands of migrating data driven analysis from local computing resource to powerful remote resources such as cloud and high performance computing cluster. In addition to various commercial cloud services, there are also rich selections of high performance computing centers in academia providing cyberinfrastructure (CI) offerings. However, access barriers exist in bring those resources to data driven research community at large. To help lower those access barriers and increase the adoption of utilization of remote resources for data driven analysis, we propose a new service model for utilizing remote computing resources, which empower users to deploy and run their big data application as a web application on remote computing resources. There are several key design goals of this model including enabling interactivity, reusability and reproducibility. Compare to the traditional batch-processing model commonly supported by CI resource providers, supporting a web application interface enables interactive analysis capabilities. Users design the application through a configuration file utilizing a set of predefined task templates that are also extensible by users. The application generated from the configuration file is self-contained and can be deployed without alleviated system privilege. Therefore, ad-hoc analysis routines can be described and preserved in a format that can be shared and re-used. Remote resources can also be described and implemented through configuration files to automatically bridge the application with remote resources and facilitate migration with different resources in the future. Consequently, analysis tasks can be preserved through the configuration file for reproducibility. Here we detail our proposed application framework and its preliminary implementations. We demonstrated usage of this framework with a practical use case of aggregating and analyzing live tweets. Weijia Xu, Ruizhu Huang |
IEEE BigData | 1 |
| 2018 | Enabling User Driven Web Applications on Remote Computing ResourceabstractWhile CI providers have continued success with the infrastructure-as-a-service model (IaaS), there are increasing demands to offer more user driven service models from domain scientists. We propose a user driven web application that empowers users to run their interactive analytic tasks using CI resources dynamically. Ad-hoc analysis routines can be described with multiple pre-defined task modules in a configuration file that can be shared and re-used. A user can run the web application on CI resource without alleviated privilege or additional service deployment by administrators. The functions and user interface of the web application are automatically initialized based on the configuration file. Therefore, the framework offers a new way for a user to access and utilize remote resources. This new model can effectively reduce the access barrier to remote computing resource offered by CI. In this paper, we describe the proposed architecture of this framework and give a use case example. Weijia Xu, Ruizhu Huang |
SERVICES | 1 |
| 2017 | Enabling versatile analysis of large scale traffic video data with deep learning and HiveQLabstractWhile monocular roadside cameras have been widely deployed and used to monitor traffic conditions across the United States, the analysis of those video data are commonly implemented either manually or through commercial applications tailor-made for specific tasks. The goal of this project is to develop an efficient system that can meet dynamic content based video analysis needs and scale to large scale traffic camera video data. The proposed system utilizes deep learning methods to recognize objects in the video data. That information can then be processed and analyzed through an analysis layer implemented using Spark and Hive. The analysis layer supports HiveQL, which enables end users to conduct sophisticated analysis with customized queries. In this paper, we present the implementation of this prototype application in details. The application can utilize both GPU and multiple CPUs to accelerate its computation. We evaluated its performance and scalability with different hardware and parameter settings, including Intel Knights Landing, Intel Skylake, Nvidia K40 GPU, and Nvidia P100 GPU, for object recognition. To demonstrate its versatile, we show two practical use case examples: counting moving vehicles and identifying scenes including pedestrians and vehicles. We show the accuracy of the system by comparing vehicular counts produced by the analysis with manually annotated results. The comparison shows our methods can achieve over eighty percent accuracy comparing to manual results. Lei Huang 0019, Weijia Xu, Si Liu 0008, Venktesh Pandey, Natalia Ruiz-Juri |
IEEE BigData | 2 |
| 2017 | Big data system for information aggregation and model comparison for precison medicineabstractPrecision and personalized medicine have gained a lot of attention in the recent years. The number of prescription drugs which are affected by genetic compositions of the patients has significantly increased over the years. We propose a big data system for aggregating and analyzing dispersed public information on prescription drugs whose useful are sensitive to patient genetic types. The proposed system includes two main features. An intelligent information aggregator that can identify and organize relevant information from several data sources into an integrated dashboard. The second feature enables users to collect relevant public data and apply and compare machine learning models to predict optimal dosage based on historical data. In this poster, we present a prototype implementation and illustrate its functions and benefits with a use case of Warfarin dosages. Weijia Xu |
IEEE BigData | 2 |
| 2016 | A workload aware model of computational resource selection for big data applicationsabstractWorkload characterization of Big Data applications has always been a challenging research problem. Big data applications often have high demands on multiple computing components in concert, such as storage, memory, network and processors and have evolving performance characteristics along with the scale of the workload. To further complicate the problem, the increasing diversity of hardware technologies available makes side-by-side comparisons hard. Choosing right resources among a wide array of available systems is a decision that is likely to plague both end users and resources providers. In this paper, we propose a workload aware model for the computational infrastructure selection problem for a given application. Our model considers both features of the workload and features of the computational infrastructure and predicts expected performance for a given workload, based on historical performance results using Support Vector Machines (SVM). We tested our model with a practical application from the domain of Transportation research on two distinct computing resources. The application has significant requirements on both memory availability and processing power. Therefore the optimal performance of the application is a dedicated trade-off between different types of resources and it is workload specific. The two testing systems represent two main trends in high performance computing resources. One infrastructure is a traditional high end computing cluster consisting of moderate number of CPUs and memories running at high frequency and high bandwidth. The other system, based on the latest Intel Knights Landing processor, is a good representation of the trending Many-Core technology in which high number of processing cores running at lower frequencies are available. The memory allocation models are also often different between the two systems. Our results show that our proposed model can achieve over 90% accuracy in performance prediction with small training data sets for our test application. The results also indicate that our model is a viable approach to be extended to other classes of applications and to be potentially adopted by high performance computing resource providers. Amit Gupta 0002, Weijia Xu, Natalia Ruiz-Juri, Kenneth Perrine |
IEEE BigData | 2 |
| 2016 | Content-based comparison for collections identificationabstractAssigning global unique persistent identifiers (GUPIs) to datasets has the goal of improving their accessibility and simplifying how they are referenced and reused. However, as repositories receive more and complex data, attesting for the identity of datasets attached to persistent identifiers over time is becoming more challenging. This is due to the nature of scientific research data, which is generated through distributed research practices and evolves across different computational environments. This work presents a robust, automated computational service for data content comparison as a valuable addition to assigning, managing, and tracking persistent identifiers. We operationalized the functions of the service within the archival space by linking data provenance and identity to authenticity. The need for such service is shown through three genomics data use cases in which the results aided curators establishing the identity of datasets and inferring issues of provenance. We describe the system's design, implementation and performance, and report on lessons learned. Weijia Xu, Ruizhu Huang, Maria Esteva, Jawon Song, Ramona L. Walls |
IEEE BigData | 1 |
| 2016 | Supporting large scale connected vehicle data analysis using HIVEabstractConnected vehicles (CVs) are vehicles that can exchange messages containing location and other safety-related information with other vehicles, and with devices affixed to roadside infrastructure. While the main purpose of vehicle connectivity is to enhance safety, the data generated by CVs has an enormous potential to support transportation planning and operations. However, handling the vast volume of data produced by CVs presents considerable challenges for researchers in the transportation domain. This paper presents a case study of using HIVE to facilitate CV data analysis based on the largest CV data set publicly released to date. We characterize the data analytic tasks that are expected to enable transportation planning research, and investigate several approaches to increase the corresponding query efficiency and throughput. This study compares the use of HIVE in conjunction with the MapReduce and Spark programming frameworks, analyzes its performance using different data storage formats, and exemplifies potential use cases. Weijia Xu, Natalia Ruiz-Juri, Amit Gupta 0002, Amanda Deering, Chandra Bhat, James Kuhr, Jackson Archer |
IEEE BigData | 1 |
| 2015 | Performance evaluation of enabling logistic regression for big data with RabstractThe software package R is a free, powerful, open source software package with extensive statistical computing and graphics capabilities. Due to its high-level expressiveness and multitude of domain-specific packages, R has become a popular tool for data analysis in many scientific fields. While there are a number of packages enabling running R in parallel using message passing interface across multiple nodes, only few packages extend R to the new system and computing paradigm for data intensive computing, such as Hadoop and Spark. In this paper, we focus on three approaches RHadoop, RHIPE and SparkR that can scale R to distributed computing systems for solving Big Data problems. We presented an algorithm for enabling logistic regression over large set of data under MapReduce programming model. We implemented the algorithm with three packages in R to exploit the benefit of Hadoop and Spark cluster. Our implementations significantly improved the scale of the data that can be analyzed with R. We conducted a study on performance and scalability up to 1TB data with those implementations and three other common solutions for logistic regression problem. The results showed SparkR consistently outperformed other approaches and also demonstrated the advantages and limitations of each package. Ruizhu Huang, Weijia Xu |
IEEE BigData | 2 |
| 2015 | Wrangler's user environment: A software framework for management of data-intensive computing systemabstractThe growth in the capacity and capability of NAND Flash based storage systems have changed the face of data oriented computational systems. These systems have become both more capable and flexible in how they are used. With these changes comes both increased potential and user complexity. While many systems attempt to hide this complexity through the addition of more layers of storage caches, the design of the Wrangler system went a different route, choosing instead to build a simple yet flexible web based interface to allow users to easily configure this complex data computing system based on their service and software needs. This allows users to work in the environments best suited to their workflows while optimally utilizing the systems high performance and high capacity storage systems. This interface also allows users to schedule long term periods of reserved capacity, "data campaigns", for projects. Finally, the system has been designed to support the data storage and sharing capacities of the system to enable these key aspects of data research. We discuss the capabilities with respect to three already existing workflows on the system to highlight the diversity and flexibility provided by this environment to data researchers. Christopher Jordan, David Walling, Weijia Xu, Stephen A. Mock, Niall Gaffney, Daniel C. Stanzione Jr. |
IEEE BigData | 3 |
| 2014 | On scaling time dependent shortest path computations for Dynamic Traffic AssignmentabstractDynamic Traffic Assignment (DTA) models provide a powerful tool to realistically represent the complex interactions between travelers and the transportation infrastructure in large regions, and they have been increasingly adopted by transportation network planners and operators in the last decade. Time dependent shortest path (TDSP) calculations at the core of most DTA methodologies usually require storing and comparing millions of discovered paths. This makes the problem I/O intensive in addition to it inherently being computationally demanding. In this paper we present a use case on scaling up the TDSP calculations within an established existing DTA software framework with distributed computing. Our approach alleviates I/O bottlenecks by using RAM disks and improves a label correcting shortest path algorithm by using priority queues which also leads to better workload balancing among parallel processes. Tests with real-world transportation networks show drastic run time performance improvements, in some cases by a factor of 12x. This suggests that our methodology enables the analysis of much larger networks. Furthermore, the improvements were achieved with relatively minor modifications to the base code, which makes this approach appealing for the enhancement of other existing DTA implementations. Amit Gupta 0002, Weijia Xu, Kenneth Perrine, Dennis Bell, Natalia Ruiz-Juri |
IEEE BigData | 2 |
| 2014 | The Adaptive Projection Forest: Using adjustable exclusion and parallelism in metric space indexesabstractThis paper introduces an indexing method for searching diverse data types that is easily parallelizable for use with large data sets. This method, the Adaptive Projection Forest (APF) is a partition-based metric-space indexing method, which provides generic retrieval solutions for data sets for which similarity is defined by a metric-distance function. The APF is uniquely suited to alleviate problems typically encountered in metric-space indexing because it adaptively incorporates exclusion, a method that removes data near a partition boundary and creates multiple trees for use in parallel computing. The use of exclusion allows the index to be more effective when data falls near partition boundaries, where traditional pruning is not always possible. The APF's use of exclusion also allows it to have greater success in parallel environments, meaning that the APF algorithm can be more effectively used on large data sets with diverse data types. In the APF index, the proportion of excluded data is adjusted dynamically at each index node by locally determining the dimension, k, of the projection of the metric space onto the real numbers. The algorithm, which provides asymptotic algorithmic guarantees for nearest neighbor search, is presented along with a parallel implementation of the APF. Across a suite of real-world and synthetic benchmarks the APF demonstrates favorable empirical results, measured in number of calculations, when compared with the emVP, MVP, and SA indexes. Experiments also reveal that number of calculations can be minimized when a critical parameter, the width of the exclusion region, is set much smaller than the value suggested by asymptotic algorithmic analysis. Lee Parnell Thompson, Weijia Xu, Daniel P. Miranker |
IEEE BigData | 2 |
| 2013 | Performance evaluation of R with Intel Xeon Phi coprocessorabstractOver the years, R has been adopted as a major data analysis and mining tool in many domain fields. As Big Data overwhelms those fields, the computational needs and workload of existing R solutions increases significantly. With recent hardware and software developments, it is possible to enable massive parallelism with existing R solutions with little to no modification. In this paper, we evaluated approaches to speed up R computations with the utilization of the Intel Math Kernel Library and automatic offloading to Intel Xeon Phi SE10P Co-processor. The testing workload includes a popular R benchmark and a practical application in health informatics. There are up to five times speedup gains from using MKL with a 16 cores without modification to the existing code for certain computing tasks. Offloading to Phi co-processor further improves the performance. The performance gains through parallelization increases as the data size increases, a promising result for adopting R for big data problem in the future. Yaakoub El Khamra, Niall Gaffney, David Walling, Eric A. Wernert, Weijia Xu, Hui Zhang 0006 |
IEEE BigData | 5 |
| 2013 | Fast scalable selection algorithms for large scale dataabstractSelection finding, and its most common form median finding, are used as a measure of central tendency for problems in biology, databases, and graphics. These problems often require selection finding as a subcomponent where it can be called many times, and as such speed is important. The Map/Reduce framework has been shown to be an important tool for creating scalable applications. There are a number of valid implementations of the selection algorithms inside of a Map/Reduce framework, certain of which are compared in this paper. However, as the volume of data increases, subtle theoretical algorithmic implementation differences can lead to significant differences in practical application. Therefore, an efficient and scalable selection finding method has the potential to provide general benefit to a number of applications. This paper compares algorithms that have been redesigned or created for the Map/Reduce framework for the purpose of selection finding, or, finding the k-th ranked element in an unordered set. This paper takes the concepts used from two existing selection algorithms and translates them into a novel method using the Map/Reduce framework with two variations. Each approach uses a different methodology to reduce the total amount of workload needed for a selection. All the algorithms are compared together for scalability and efficiency in a computing cluster environment with up to 256 processing cores. The results show that the methods proposed in this paper outperform several common alternatives in identifying medians with Hadoop, including using sorting, Pig, and BinMedian methods. Our implementations are also available upon request. Lee Parnell Thompson, Weijia Xu, Daniel P. Miranker |
IEEE BigData | 2 |
| 2013 | A case study on entity Resolution for Distant Processing of big Humanities dataabstractAt the forefront of big data in the Humanities, collections management can directly impact collections access and reuse. However, curators using traditional data management methods for tasks such as identifying redundant from relevant and related records, a small increase in data volume can significantly increase their workload. In this paper, we present preliminary work aimed at assisting curators in making important data management decisions for organizing and improving the overall quality of large unstructured Humanities data collections. Using Entity Resolution as a conceptual framework, we created a similarity model that compares directories and files based on their implicit metadata, and clusters pairs of closely related directories. Useful relationships between data are identified and presented through a graphical user interface that allows qualitative evaluation of the clusters and provides a guide to decide on data management actions. To evaluate the model's performance, we experimented with a test collection and asked the curator to classify the clusters according to four model cluster configurations that consider the presence of related and duplicate information. Evaluation results suggest that the model is useful for making data management action decisions. Weijia Xu, Maria Esteva, Jessica Trelogan, Todd Swinson |
IEEE BigData | 1 |
| 2012 | Designing an Interface for Exploring Online Autism Support Communities
Bretagne Abirached, Weijia Xu, Yan Zhang 0005 |
AMIA | 3 |
| 2012 | An accurate scalable template-based alignment algorithmabstractThe rapid determination of nucleic acid sequences is increasing the number of sequences that are available. Inherent in a template or seed alignment is the culmination of structural and functional constraints that are selecting those mutations that are viable during the evolution of the RNA. While we might not understand these structural and functional, template-based alignment programs utilize the patterns of sequence conservation to encapsulate the characteristics of viable RNA sequences that are aligned properly. We have developed a program that utilizes the different dimensions of information in rCAD, a large RNA informatics resource, to establish a profile for each position in an alignment. The most significant include sequence identity and column composition in different phylogenetic taxa. We have compared our methods with a maximum of eight alternative alignment methods on different sets of 16S and 23S rRNA sequences with sequence percent identities ranging from 50% to 100%. The results showed that CRWAlign outperformed the other alignment methods in both speed and accuracy. A web-based alignment server is available at http://www.rna.ccbb.utexas.edu/SAE/2F/CRWAlign. David P. Gardner, Weijia Xu, Daniel P. Miranker, Stuart Ozer, Jamie J. Cannone, Robin Ray Gutell |
BIBM | 2 |
| 2012 | On automatically tagging web documents from examplesabstractAn emerging need in information retrieval is to identify a set of documents conforming to an abstract description. This task presents two major challenges to existing methods of document retrieval and classification. First, similarity based on overall content is less effective because there may be great variance in both content and subject of documents produced for similar functions, e.g. a presidential speech or a government ministry white paper. Second, the function of the document can be defined based on user interests or the specific data set through a set of existing examples, which cannot be described with standard categories. Additionally, the increasing volume and complexity of document collections demands new scalable computational solutions. We conducted a case study using web-archived data from the Latin American Government Documents Archive (LAGDA) to illustrate these problems and challenges. We propose a new hybrid approach based on Naïve Bayes inference that uses mixed n-gram models obtained from a training set to classify documents in the corpus. The approach has been developed to exploit parallel processing for large scale data set. The preliminary work shows promising results with improved accuracy for this type of retrieval problem. Nicholas Joel Woodward, Weijia Xu, Kent Norsworthy |
SIGIR | 2 |
| 2011 | R-PASS: A Fast Structure-Based RNA Sequence Alignment AlgorithmabstractWe present a fast pairwise RNA sequence alignment method using structural information, named R-PASS (RNA Pairwise Alignment of Structure and Sequence), which shows good accuracy on sequences with low sequence identity and significantly faster than alternative methods. The method begins by representing RNA secondary structure as a set of structure motifs. The motifs from two RNAs are then used as input into a bipartite graph-matching algorithm, which determines the structure matches. The matches are then used as constraints in a constrained dynamic programming sequence alignment procedure. The R-PASS method has an O(nm) complexity. We compare our method with two other structure-based alignment methods, LARA and ExpaLoc, and with a sequence-based alignment method, MAFFT, across three benchmarks and obtain favorable results in accuracy and orders of magnitude faster in speed. Weijia Xu, Lee Parnell Thompson, Robin Ray Gutell, Daniel P. Miranker |
BIBM | 2 |
| 2011 | RNA2DMap: A Visual Exploration Tool of the Information in RNA's Higher-Order StructureabstractA new and emerging paradigm in molecular biology is revealing that RNA is implicated in nearly every aspect of the metabolism in the cell. To enhance our understanding of the function of these RNA molecules in the cell, it is essential that we have a complete understanding of their higher-order structures. While many computational tools have been developed to predict and analyse these higher-order RNA structures, few are able to visualize them for analytical purposes. In this paper, we present an interactive visualization tool of the secondary structure of RNA, named RNA2DMap. This program enables multiple-dimensions of information about RNA structure to be selected, customized and displayed to visually identify patterns and relationships. RNA2DMap facilitates the comparative analysis and understanding of RNAs that cannot be readily obtained with other graphical or text output from computer programs. Three use cases are presented to illustrate how RNA2DMap aids structural analysis. Weijia Xu, Ame Wongsa, Jung Lee, Jamie J. Cannone, Robin Ray Gutell |
BIBM | 1 |
| 2011 | rCAD: A Novel Database Schema for the Comparative Analysis of RNAabstractBeyond its direct involvement in protein synthesis with mRNA, tRNA, and rRNA, RNA is now being appreciated for its significance in the overall metabolism and regulation of the cell. Comparative analysis has been very effective in the identification and characterization of RNA molecules, including the accurate prediction of their secondary structure. We are developing an integrative scalable data management and analysis system, the RNA Comparative Analysis Database (rCAD), implemented with SQL Server to support RNA comparative analysis. The platformagnostic database schema of rCAD captures the essential relationships between the different dimensions of information for RNA comparative analysis datasets. The rCAD implementation enables a variety of comparative analysis manipulations with multiple integrated data dimensions for advanced RNA comparative analysis workflows. In this paper, we describe details of the rCAD schema design and illustrate its usefulness with two usage scenarios. Stuart Ozer, Kishore J. Doshi, Weijia Xu, Robin Ray Gutell |
eScience | 3 |
| 2011 | Facilitating Understanding of Large Document CollectionsabstractLarge document collections containing multiple topics can be overwhelming to understand, requiring librarians and archivists significant time and efforts to develop access points. Efficient computational methods can aid this process by uncovering groups of documents that can be described for access. We investigate the use of density based clustering with document segmentation to identify points of access as dense clusters of information. The method returns stories and classes of cohesive clusters that can be described as precise points of access. We found that our method performs more efficiently than K-means clustering and topic model using Latent Dirichlet Allocation (LDA). We use Hadoop to process a large document collection. Jae Hyeon Bae, Weijia Xu, Maria Esteva |
ICDAR | 2 |
| 2009 | Covariant Evolutionary Event Analysis for Base Interaction Prediction Using a Relational Database Management System for RNA
Weijia Xu, Stuart Ozer, Robin Ray Gutell |
SSDBM | 1 |
| 2006 | On Integrating Peptide Sequence Analysis and Relational Distance-Based IndexingabstractManaging data with distance-based indexing methods has the potential to provide scalability and integration with relational database management systems and the SQL programming model. We previously demonstrated the advantages of such an approach for nucleotide sequences using Hamming distance (mismatch). However, the larger alphabet size of peptide sequences increases the dimensionality of the problem, making algorithmic results more challenging. The development of a metric-PAM substitution matrix enables metric-distance based indexing for peptide sequences. The performance of distance-based indexing for homologous protein retrieval entails trade-off among accuracy, scalability and computational cost. We investigate the application of the multi-vantage point (MVP) tree algorithm to index peptide k-mers based on global mPAM alignment. We show that k-mer retrieval can still maintain accuracy when k is at least as large as 6 that creates a domain of over 60 million key values and enables scalability sufficient for effective performance on large disk-resident sequence databases Weijia Xu, Rui Mao 0001, Daniel P. Miranker |
BIBE | 1 |
| 2006 | A fast coarse filtering method for peptide identification by mass spectrometryabstractMOTIVATION: We reformulate the problem of comparing mass-spectra by mapping spectra to a vector space model. Our search method leverages a metric space indexing algorithm to produce an initial candidate set, which can be followed by any fine ranking scheme. RESULTS: We consider three distance measures integrated into a multi-vantage point index structure. Of these, a semi-metric fuzzy-cosine distance using peptide precursor mass constraints performs the best. The index acts as a coarse, lossless filter with respect to the SEQUEST and ProFound scoring schemes, reducing the number of distance computations and returned candidates for fine filtering to about 0.5% and 0.02% of the database respectively. The fuzzy cosine distance term improves specificity over a peptide precursor mass filter, reducing the number of returned candidates by an order of magnitude. Run time measurements suggest proportional speedups in overall search times. Using an implementation of ProFound's Bayesian score as an example of a fine filter on a test set of Escherichia coli protein fragmentation spectra, the top results of our sample system are consistent with that of SEQUEST. Smriti R. Ramakrishnan, Rui Mao 0001, Aleksey A. Nakorchevskiy, John T. Prince, Willard S. Willard, Weijia Xu, Edward M. Marcotte, Daniel P. Miranker |
Bioinform. | 6 |
| 2004 | A metric model of amino acid substitutionabstractMOTIVATION: We address the question of whether there exists an effective evolutionary model of amino-acid substitution that forms a metric-distance function. There is always a trade-off between speed and sensitivity among competing computational methods of determining sequence homology. A metric model of evolution is a prerequisite for the development of an entire class of fast sequence analysis algorithms that are both scalable, O(log n) and sensitive. RESULTS: We have reworked the mathematics of the point accepted mutation model (PAM) by calculating the expected time between accepted mutations in lieu of calculating log-odds probabilities. The resulting substitution matrix (mPAM) forms a metric. We validate the application of the mPAM evolutionary model for sequence homology by executing sequence queries from a controlled yeast protein homology search benchmark. We compare the accuracy of the results of mPAM and PAM similarity matrices as well as three prior metric models. The experiment shows that mPAM significantly outperforms the other three metrics and sufficiently approaches the sensitivity of PAM250 to make it applicable to the management of protein sequence databases. Weijia Xu, Daniel P. Miranker |
Bioinform. | 1 |
| 2004 | A metric model of amino acid substitutionabstractWeijia Xu, Daniel P. Miranke; A metric model of amino acid substitution, Bioinformatics, Volume 20, Issue 18, 12 December 2004, Pages 3716, https://doi.org Weijia Xu, Daniel P. Miranker |
Bioinform. | 1 |
| 2003 | An Assessment of a Metric Space Database Index to Support Sequence HomologyabstractHierarchical metric-space clustering methods have been commonly used to organize proteomes into taxonomies. Consequently, it is often anticipated that hierarchical clustering can be leveraged as a basis for scalable database index structures capable of managing the hyper-exponential growth of sequence data. M-tree is one such data structure specialized for the management of large data sets on disk. We explore the application of M-trees to the storage and retrieval of peptide sequence data. Exploiting a technique first suggested by Myers (1994), we organize the database as records of fixed length substrings. Empirical results are promising. However, metric-space indexes are subject to "the curse of dimensionality" and the ultimate performance of an index is sensitive to the quality of the initial construction of the index. We introduce new hierarchical bulk-load algorithm that alternates between top-down and bottom-up clustering to initialize the index. Using the Yeast Proteomes, the bi-directional bulk load produces a more effective index than the existing M-tree initialization algorithms. Rui Mao 0001, Weijia Xu, Neha Singh 0001, Daniel P. Miranker |
BIBE | 2 |
| 2003 | MoBIoS: A Metric-Space DBMS to Support Biological DiscoveryabstractMoBIoS is a specialized database management system whose storage manager is based on metricspace indexing, and whose query language entails biological data types. When relational database management systems are used to support biological data, important data types are relegated to blob and unstructured text fields. Thus, even simple, but critical queries are executed by sequentially dumping the data to utilities outside the database. MoBIoS provides O(log n) physical access to diverse biological data types as well as uniform logical and syntactic access. Consequently, MoBIoS provides a framework where complex bioinformatic algorithms may be effectively expressed and executed as concise declarative SQL-like (Structured Query Language) queries. Daniel P. Miranker, Weijia Xu, Rui Mao 0001 |
SSDBM | 2 |