EDBT 2026 Demo / reviewers in the wild / expert
Jason H. Haga
dblp:78/3586 · also Jason Haga
· DBLP profile ↗
18ranked-venue papers
0as first author
7since 2021 · last 2026
0000-0002-6407-0003ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 1 since 2021Software engineering, systems software and programming languages · 6 · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FLaTEC: An efficient federated learning scheme across the Thing-Edge-Cloud environment
Van An Le, Jason H. Haga, Yusuke Tanimura, Truong Thao Nguyen |
Future Gener. Comput. Syst. | 2 |
| 2024 | SFETEC: Split-FEderated Learning Scheme Optimized for Thing-Edge-Cloud EnvironmentabstractThis paper introduces SFETEC, an innovative federated learning framework addressing the limitations of traditional methods like FedAvg. SFETEC splits training into base models and core models, reducing communication overhead, mitigating non-IID data issues, and enhancing training speed. Base models are trained at client devices, while core models are trained at edge servers and aggregated at the cloud. Preliminary results show that SFETEC significantly reduces communication overhead and training duration compared to the state-of-the-art baselines while enhancing privacy. Van An Le, Jason H. Haga, Yusuke Tanimura, Truong Thao Nguyen |
e-Science | 2 |
| 2023 | Taming Metadata-intensive HPC Jobs Through Dynamic, Application-agnostic QoS ControlabstractModern I/O applications that run on HPC infrastructures are increasingly becoming read and metadata intensive. However, having multiple applications submitting large amounts of metadata operations can easily saturate the shared parallel file system's metadata resources, leading to overall performance degradation and I/O unfairness. We present PADLL, an application and file system agnostic storage middleware that enables QoS control of data and metadata workflows in HPC storage systems. It adopts ideas from Software-Defined Storage, building data plane stages that mediate and rate limit POSIX requests submitted to the shared file system, and a control plane that holistically coordinates how all I/O workflows are handled. We demonstrate its performance and feasibility under multiple QoS policies using synthetic benchmarks, real-world applications, and traces collected from a production file system. Results show that PADLL can enforce complex storage QoS policies over concurrent metadata-aggressive jobs, ensuring fairness and prioritization. Ricardo Macedo, Mariana Miranda, Yusuke Tanimura, Jason H. Haga, Amit Ruhela, Stephen Lien Harrell, R. Todd Evans, José Pereira 0001, João Paulo 0001 |
CCGrid | 4 |
| 2022 | Protecting Metadata Servers From Harm Through Application-level I/O ControlabstractModern large-scale I/O applications that run on HPC infrastructures are increasingly becoming metadata-intensive. Unfortunately, having multiple concurrent applications submitting massive amounts of metadata operations can easily saturate the shared parallel file system's metadata resources, leading to unresponsiveness of the storage backend and overall performance degradation. To address these challenges, we present Padll, a storage middleware that enables system administrators to proactively control and ensure QoS over metadata workflows in HPC storage systems. We demonstrate its performance and feasibility by controlling the rate of both synthetic and realistic I/O workloads. Results show that Padll can dynamically control metadata-aggressive workloads, prevent I/O burstiness, and ensure I/O fairness and prioritization. Ricardo Macedo, Mariana Miranda, Yusuke Tanimura, Jason H. Haga, Amit Ruhela, Stephen Lien Harrell, R. Todd Evans, João Paulo 0001 |
CLUSTER | 4 |
| 2022 | PAIO: General, Portable I/O Optimizations With Minor Application Modifications
Ricardo Macedo, Yusuke Tanimura, Jason H. Haga, Vijay Chidambaram, José Pereira 0001, João Paulo 0001 |
FAST | 3 |
| 2021 | Grand Challenges in Immersive AnalyticsabstractImmersive Analytics is a quickly evolving field that unites several areas such as visualisation, immersive environments, and human-computer interaction to support human data analysis with emerging technologies. This research has thrived over the past years with multiple workshops, seminars, and a growing body of publications, spanning several conferences. Given the rapid advancement of interaction technologies and novel application domains, this paper aims toward a broader research agenda to enable widespread adoption. We present 17 key research challenges developed over multiple sessions by a diverse group of 24 international experts, initiated from a virtual scientific workshop at ACM CHI 2020. These challenges aim to coordinate future work by providing a systematic roadmap of current directions and impending hurdles to facilitate productive and effective applications for Immersive Analytics. Barrett Ens, Benjamin Bach, Maxime Cordeil, Ulrich Engelke, Marcos Serrano, Wesley Willett, Arnaud Prouzeau, Christoph Anthes, Wolfgang Büschel, Cody Dunne, Tim Dwyer, Jens Grubert, Jason H. Haga, Nurit Kirshenbaum, Dylan Kobayashi, Tica Lin, Monsurat Olaosebikan, Fabian Pointecker, David Saffo, Dieter Schmalstieg, Danielle Albers Szafir, Matt Whitlock, Yalong Yang 0001 |
CHI | 13 |
| 2021 | The Case for Storage Optimization Decoupling in Deep Learning FrameworksabstractDeep Learning (DL) training requires efficient access to large collections of data, leading DL frameworks to implement individual I/O optimizations to take full advantage of storage performance. However, these optimizations are intrinsic to each framework, limiting their applicability and portability across DL solutions, while making them inefficient for scenarios where multiple applications compete for shared storage resources.We argue that storage optimizations should be decoupled from DL frameworks and moved to a dedicated storage layer. To achieve this, we propose a new Software-Defined Storage architecture for accelerating DL training performance. The data plane implements self-contained, generally applicable I/O optimizations, while the control plane dynamically adapts them to cope with workload variations and multi-tenant environments.We validate the applicability and portability of our approach by developing and integrating an early prototype with the TensorFlow and PyTorch frameworks. Results show that our I/O optimizations significantly reduce DL training time by up to 54% and 63% for TensorFlow and PyTorch baseline configurations, while providing similar performance benefits to framework-intrinsic I/O mechanisms provided by TensorFlow. Ricardo Macedo, Cláudia Correia, Marco Dantas, Cláudia Brito, Weijia Xu, Yusuke Tanimura, Jason H. Haga, João Paulo 0001 |
CLUSTER | 7 |
| 2020 | Massively Parallel Causal Inference of Whole Brain Dynamics at Single Neuron ResolutionabstractEmpirical Dynamic Modeling (EDM) is a nonlinear time series causal inference framework. The latest implementation of EDM, cppEDM, has only been used for small datasets due to computational cost. With the growth of data collection capabilities, there is a great need to identify causal relationships in large datasets. We present mpEDM, a parallel distributed implementation of EDM optimized for modern GPU-centric supercomputers. We improve the original algorithm to reduce redundant computation and optimize the implementation to fully utilize hardware resources such as GPUs and SIMD units. As a use case, we run mpEDM on AI Bridging Cloud Infrastructure (ABCI) using datasets of an entire animal brain sampled at single neuron resolution to identify dynamical causation patterns across the brain. mpEDM is 1,530× faster than cppEDM and a dataset containing 101,729 neuron was analyzed in 199 seconds on 512 nodes. This is the largest EDM causal inference achieved to date. Wassapon Watanakeesuntorn, Keichi Takahashi, Kohei Ichikawa, Joseph Park, George Sugihara, Ryousei Takano, Jason H. Haga, Gerald M. Pao |
ICPADS | 7 |
| 2020 | Water Level Detection from CCTV Cameras using a Deep Learning ApproachabstractNatural disasters are a global problem that causes widespread losses and damage. A system to provide timely information is required in order to help reduce losses. Flooding is one of the major natural disasters that requires a monitoring and detection system. The traditional flood detection systems use remote sensors such as river water levels and rainfall to provide information to both disaster management professionals and the general public. There is an attempt to use visual information such as CCTV cameras to detect extreme flooding events; however, it requires human experts and consistent attention to monitor any changes. In this paper, we introduce an approach to the automatic river water level detection using deep learning to determine the water level from surveillance cameras. The model achieves 93% accuracy using a single camera location and 83% accuracy using multiple camera locations. Punyanuch Borwarnginn, Jason H. Haga, Worapan Kusakunniran |
TENCON | 2 |
| 2019 | Usage Patterns of Wideband Display Environments In e-Science Research, Development and TrainingabstractSAGE (the Scalable Adaptive Graphics Environment) and its successor SAGE2 (the Scalable Amplified Group Environment) are operating systems for managing content across wideband display environments. This paper documents the prevalent usage patterns of SAGE-enabled display walls in support of the e-Science enterprise, based on nearly 15 years of observations of the SAGE community. These patterns will help guide e-Science users and cyberinfrastructure developers on how best to leverage large tiled display walls, and the types of software services that could be provided in the future. Jason Leigh, Krishna Bharadwaj, Arthur Nishimoto, Lance Long, Jason H. Haga, Francis Cristobal, Jared H. McLean, Roberto Pelayo, Mahdi Belcaid, Dylan Kobayashi, Nurit Kirshenbaum, Troy Wooton, Luc Renambot, Andrew E. Johnson 0001, Maxine D. Brown, Andrew Thomas Burks |
eScience | 5 |
| 2018 | Sage River Disaster Information (SageRDI): Demonstrating Application Data Sharing In SAGE2abstractThe Scalable Amplified Group Environment (SAGE2) is an open-source, web-based middleware for driving tiled display walls. SAGE2 allows running multiple applications at once within its workspace. In large display walls, users tend to collaborate using multiple applications in the same space for simultaneous interaction and review. Unfortunately, many of these applications were created independently by different developers and were never intended to interoperate which greatly limits their potential reusability. SAGE2 developers face system limitations where applications are data segregated and cannot easily communicate with others. To counter this problem, we developed the SAGE2 data sharing components. We describe the Sage River Disaster Information (SageRDI) application and the SAGE2 architectural implementations necessary for its operation. SageRDI enables river.go.jp, an existing website that provides water sensor data for Japan, to interact with other SAGE2 applications without modifying the website's server, hosted files, nor any of the default SAGE2 applications. Dylan Kobayashi, Matthew Ready, Alberto Gonzalez Martinez, Nurit Kirshenbaum, Tyson Seto-Mook, Jason Leigh, Jason H. Haga |
ISS | 7 |
| 2017 | PRAGMA-ENT: An International SDN testbed for cyberinfrastructure in the Pacific RimabstractSummary The Pacific Rim Application and Grid Middleware Assembly (PRAGMA) is an international community of researchers that actively collaborate to address problems and challenges of common interest in eScience. The PRAGMA Experimental Network Testbed (PRAGMA‐ENT) was established with the goal of constructing an international software‐defined network (SDN) testbed to offer the necessary networking support to the PRAGMA cyberinfrastructure. PRAGMA‐ENT is isolated, and PRAGMA researchers have complete freedom to access network resources to develop, experiment, and evaluate new ideas without the concerns of interfering with production networks. In the first phase, PRAGMA‐ENT focused on establishing an international L2 backbone. With support from the Florida Lambda Rail, Internet2, PacificWave, Japan Gigabit Network, and TaiWan Advanced Research and Education Network, PRAGMA‐ENT backbone connects openflow‐enabled switches at University of Florida, University of California, San Diego, Nara Institute of Science and Technology (Japan), Osaka University (Japan), National Institute of Advanced Industrial Science and Technology (Japan), and National Applied Research Laboratories (Taiwan). The second phase of PRAGMA‐ENT consisted of an evaluation of technologies for the control plane that enables multiple experiments (ie, OpenFlow controllers) to coexist. Preliminary experiments with FlowVisor revealed some limitations leading to the development of a new approach, called AutoVFlow. This paper describes our experience in the establishment of PRAGMA‐ENT backbone (with international L2 links), its current status, and plans for the control plane. Discussion of preliminary application ideas, including optimization of routing control; multipath routing control; extending the backbone using overlay network; and remote visualization are also discussed. Kohei Ichikawa, Pongsakorn U.-Chupala, Che Huang, Chawanat Nakasan, Te-Lung Liu, Jo-Yu Chang, Li-Chi Ku, Whey-Fone Tsai, Jason H. Haga, Hiroaki Yamanaka, Eiji Kawai, Yoshiyuki Kido, Susumu Date, Shinji Shimojo, Philip M. Papadopoulos, Maurício O. Tsugawa, Matthew Collins, Kyuho Jeong, Renato J. O. Figueiredo, José A. B. Fortes |
Concurr. Comput. Pract. Exp. | 9 |
| 2017 | Interactive museum exhibits with embedded systems: A use-case scenarioabstractSummary The feasibility of using embedded systems in real‐life applications is becoming more widespread. These applications have grown from do‐it‐yourself projects of computer enthusiasts or robotics projects to larger scale efforts and deployments. This paper describes a scenario that deployed a prototype application that allows the public to interact with features of a model and view videos from a first‐person perspective on the train. Through testing the embedded systems and their usage in a public setting, it was demonstrated that interactive features could be implemented in model train exhibits, which are featured in traditional museum environments that lack technical infrastructure. Specifically, the Arduino and Raspberry Pi provide the necessary linkages between the Internet and hardware, allowing for a greater interactive experience for museum visitors. These results provide an important use‐case scenario and lessons learned that cultural heritage institutions can use when implementing embedded systems on a larger scale, for the purpose of increasing visitors' experience through greater interaction and engagement. Lok Wong, Shinji Shimojo, Yuuichi Teranishi, Tomoki Yoshihisa, Jason H. Haga |
Concurr. Comput. Pract. Exp. | 5 |
| 2015 | Deployment of a Multi-site Cloud Environment for Molecular Virtual ScreeningsabstractWith the constant increase in the number and variety of small molecule chemical compounds, drug discovery is becoming a very resource intensive endeavor. Performing molecular simulations of ligand-protein binding by virtual screening has become an integral part of the discovery process. Cloud computing is an efficient choice to execute these large-scale screenings, given that large compute allocations are not accessible to many researchers. This research focused on developing a multi-site cloud environment that combines small allocations of virtual machines in multiple locations connected through a virtual networking system (ViNe), and compared two parallelization approaches: Message Passing Interface (MPI) and MapReduce using Hadoop. Virtual screenings were conducted using DOCK, a protein-ligand molecular interaction simulation program. Multiple DOCK test simulations through MPI and Hadoop were run to assess the performance and flexibility of the environment. These tests indicated that MPI and MapReduce offer comparable scalability performance, and that network latency has a significant influence on low accuracy simulations. Furthermore, differences in performance at individual cloud resource sites were reduced on average because of the larger combined pool of resources. This project prototyped and assessed a fully functional multi-site cloud environment for virtual screenings, which can be used to guide small laboratories in deploying their own cloud-based screenings. Andréa M. Matsunaga, Maurício O. Tsugawa, Susumu Date, Kohei Ichikawa, Jason H. Haga |
e-Science | 6 |
| 2013 | Protein Structure Modeling in a Grid Computing EnvironmentabstractAdvances in sequencing technology have resulted in an exponential increase in the availability of protein sequence information. In order to fully utilize information, it is important to translate the primary sequences into high-resolution tertiary protein structures. MODELLER is a leading homology modeling method that produces high quality protein structures. In this study, the function of MODELLER was expanded by configuring and deploying it on a parallel grid computing platform using a custom four-step workflow. The workflow consisted of template selection through a protein BLAST algorithm, target-template protein sequence alignment, distribution of model generation jobs among the compute clusters, and final protein model optimization. To test the validity of this workflow, we used the Dual Specificity Phosphatase (DSP) protein family, which shares high homology among each other. Comparison of the DSP member SSH-2 with its model counterpart revealed a minimal 1.3% difference in output energy scores. Furthermore, the Dali Pair wise Comparison Program demonstrated a 98% match among amino acid features and a Z-score of 26.6 indicating very significant similarities between the model and actual protein structure. After confirming the accuracy of our workflow, we generated 23 previously unknown DSP family protein structure models. Over 40,000 models were generated 30 times faster than conventional computing. Virtual receptor-ligand screening results of modeled protein DSP21 were compared with two known structures that had either higher or lower structural homology to DSP21. There was a significant difference (p!0.001) between the average ligand ranking discrepancy of a more homologous protein pair and a less homologous protein pair, suggesting that the protein models generated were sufficiently accurate for virtual screening. These results demonstrate the accuracy and usability of a grid-enabled MODELLER program and the increased efficiency of processing protein structure models. This workflow will help increase the speed of future drug development pipelines. Brian Tsui, Charles Xue, Jason H. Haga, Kohei Ichikawa, Susumu Date |
e-Science | 4 |
| 2010 | ViewDock TDW: high-throughput visualization of virtual screening resultsabstractSUMMARY: ViewDock TDW is a modification of the pre-existing ViewDock Chimera extension (http://www.cgl.ucsf.edu/chimera/) used to visualize results of virtual screening experiments. By combing TDW hardware and an enhanced ViewDock interface, dozens of ligand-protein complexes are rendered simultaneously to parallelize the analysis of candidate ligands. The ViewDock TDW GUI allows the user to easily and interactively manipulate the molecules on the TDW as an entire set, a selected subset or a single ligand-protein complex and preserves all Chimera functionality. AVAILABILITY AND IMPLEMENTATION: ViewDock TDW is an open source software; freely available on the web at http://www.tdw-prime.webs.com. Chimera UCSF is also available, free of charge, at http://www.cgl.ucsf.edu/chimera/ Christopher D. Lau, Marshall J. Levesque, Shu Chien, Susumu Date, Jason H. Haga |
Bioinform. | 5 |
| 2008 | Virtual Screening for SHP-2 Specific Inhibitors Using Grid ComputingabstractSHP-2 is a protein tyrosine phosphatase (PTP) that plays an important role in many cellular functions such as development, growth, and death; thus SHP-2 has been hypothesized to play an important role in various diseases such as diabetes, neurodegeneration, and cancer. The importance of the individual roles of different PTPs is not well understood and this is complicated by the lack of specific inhibitors. In this study, we have utilized the multi-institutional PRAGMA Grid computation resources to virtually screen the ZINC 7 database using virtual docking software DOCK 6.2. Preliminary results suggest several SHP-2 specific inhibitors that can be further tested and validated under laboratory conditions. Complications during these multiple, virtual screenings on the grid as well as potential improvements are also discussed. These findings have future clinical significance in the creation of new drug therapies for the treatment of different diseases. Simon X. Han, Marshall J. Levesque, Kohei Ichikawa, Susumu Date, Jason H. Haga |
eScience | 5 |
| 2008 | Identification of a Specific Inhibitor for the Dual-Specificity Enzyme SSH-2 via Docking Experiments on the GridabstractThe slingshot-2 (SSH-2) protein plays a significant role in different cell functions such as growth and movement. SSH-2 is a phosphatase protein that belongs to a unique class of enzymes called dual specificity phosphatases (DSP) that target the phosphothreonine and phosphotyrosine residues of mitogen-activated protein (MAP) kinases, which regulate cell growth. Because of this, it is of great interest to find specific inhibitors of DSPs such as SSH-2. Implementing an in silico platform to screen a sizable pool of chemical compounds against SSH-2 on the grid environment with the molecular docking software DOCK 6, several chemical compounds have been identified as potential inhibitors of SSH-2 activity. The issues of performing routine virtual screenings on the grid and possible improvements are also presented. The most promising inhibitor determined from standard and AMBER DOCK screenings was 2-amino-3-phosphonooxy-propanoic acid, which will be verified with wet bench testing. Phillip D. Pham, Marshall J. Levesque, Kohei Ichikawa, Susumu Date, Jason H. Haga |
eScience | 5 |