VLDB 2026 Research / reviewers in the wild / expert
Tsengdar J. Lee
dblp:139/2378
· DBLP profile ↗
13ranked-venue papers
0as first author
6since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 8 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Multi-Layer Agent-Based Spatiotemporal UoW Recommendation for Workflow CompositionabstractSoftware service discovery and recommendation help data scientists build scientific workflows - multi-step data analytics procedures - by automating the manual selection of services. Previous research shows that recommending chainable units of work (UoWs), rather than individual services, improves efficiency and reduces data shimming issues. However, UoW recommendation remains an NP-hard problem. To tackle this challenge, this study introduces a novel framework tailored to recommend UoWs in a goal-driven, context-aware manner, thereby facilitating workflow development. The framework is built around layered structure of software service social networks. At its foundation lies a service dependency network, where each edge represents a dependency between a two-service UoW within a specific context. The next layer abstracts each of these edges into a node, with new edges now representing three-service UoWs. This layering process continues iteratively, with subsequent layers capturing increasingly complex UoWs at higher levels of granularity. At high-order layers, UoW nodes are clustered based on their semantic embeddings, with each cluster represented by an intelligent agent. This approach transforms the workflow recommendation problem into a multi-agent collaboration task, where agents work together to identify high-level UoW groupings before refining selections by navigating down the layered structure for finer-grained recommendations. Experimental results over a real-world dataset confirm the effectiveness of the proposed framework in enhancing workflow composition efficiency. Xihao Xie, Jia Zhang 0001, Rahul Ramachandran, Tsengdar J. Lee, Seungwon Lee 0005 |
SSE | 5 |
| 2024 | High-Order-Modal Knowledge Graph Powered API Recommendation for Mashup DevelopmentabstractAs increasingly more APIs are published on the Internet, effective API recommendation remains a challenge yet highly demanded for mashup developers. This paper formalizes API recommendation as an incremental context-aware recom-mendation problem starting from a set of descriptive words and a set of APIs selected to date, supported by a fine-grained mashup-oriented knowledge graph (MKG). In contrast to traditional knowledge graphs where nodes are coarse-grained entities, entity-and relationship-encapsulated features are extracted as first-class citizens in an MKG, so that implicit feature relationships can be turned into explicit structural relationships. Two models are trained to learn fine-grained API selection strategies through path type patterns in the MKG, starting from intended descriptions and APIs selected, respectively. Extensive experiments over real-world datasets have demonstrated the effectiveness of the method. Beichen Hu, Xihao Xie, Jia Zhang 0001, Tsengdar J. Lee, Seungwon Lee 0005 |
SSE | 5 |
| 2024 | RANGER: Context-Aware Service Unit of Work Recommendation for Incremental Scientific Workflow Composition
Xihao Xie, Jia Zhang 0001, Rahul Ramachandran, Tsengdar J. Lee, Seungwon Lee 0005 |
WISE (3) | 5 |
| 2022 | Learning Context-Aware Service Representation for Service Recommendation in Workflow CompositionabstractAs increasingly more software services have been published onto the Internet, it becomes critical yet highly challenging to recommend suitable services to facilitate scientific workflow composition. This paper proposes a novel Natural Language Processing (NLP)-inspired approach to recommending services throughout a workflow development process, based on incrementally learning latent service representation from workflow provenance. A work-flow composition process is formalized as a step-wise, context-aware service selection procedure, which is mapped to next-word prediction in a natural language sentence generation. Historical service dependencies are extracted from workflow provenance to build and enrich a knowledge graph. Each path in the knowledge graph reflects a scenario in a data analytics experiment, which is analogous to a sentence in a conversation. All paths are thus formalized as composable service sequences and are mined, using various patterns, from the established knowledge graph to construct a corpus. Service embeddings are then learned by applying deep learning model from the NLP field. Extensive experiments on the real-world dataset demonstrate the effectiveness and efficiency of the approach. Xihao Xie, Jia Zhang 0001, Rahul Ramachandran, Tsengdar J. Lee, Seungwon Lee 0005 |
ICIS | 4 |
| 2022 | Goal-Driven Context-Aware Service Recommendation for Mashup DevelopmentabstractAs service-oriented architecture becoming one prevalent technique to rapidly compose functionalities to customers, increasingly more reusable software components have been published online in the form of web services. To create a mashup, however, it gets not only time-consuming but also error-prone for developers to find suitable services components from such a sea of services. Service discovery and recommendation has thus attracted significant momentum in both academia and industry. This paper proposes a novel incremental recommend-as-you-go approach to recommending next potential service based on the context of a mashup under construction, considering services that have been selected up to the current step as well as the mashup goal. The core technique is an algorithm of learning the embedding of services, which learns their past goal-driven context-aware decision making behaviors in addition to their semantic descriptions and co-occurrence history. A goal exclusionary negative sampling mechanism tailored for mashup development is also developed to improve training performance. Extensive experiments on a real-world dataset demonstrate the effectiveness of this approach. Xihao Xie, Jia Zhang 0001, Rahul Ramachandran, Tsengdar J. Lee, Seungwon Lee 0005 |
SNPD | 4 |
| 2021 | Augmenting Data Systems with Prediction based EmbeddingsabstractOne of the challenges of improving the search and use of complex Earth science data is designing and incorporating semantic components in existing Earth science data systems. Many projects have addressed this by using a knowledge engineering approach. However, using ontologies has inherent limitations as a practical and scalable approach. Data-driven strategies based on natural language processing, coupled with Machine Learning, provide an alternative approach. Data-driven approaches utilize existing corpus available as unstructured text. This paper describes a hybrid strategy that uses a data-driven approach to build an embedding from a large corpus of Earth science journal publications while leveraging existing ontologies to develop validation tests to evaluate the embedding's robustness and correctness. The paper also describes the use of this embedding in two different applications. The first application provides a semantic mapping service to bridge the gap between a science application need and the appropriate instruments or datasets required to address that need. The second application is keyword recommender to make the data set tagging process efficient for the data operators and ensure keyword consistency within a data catalog. Rahul Ramachandran, Muthukumaran Ramasubramanian, Iksha Gurung, Carson Davis, Derek Koehl, Manil Maskey, Tsengdar J. Lee |
IGARSS | 7 |
| 2020 | GeoNEX: A Geostationary Earth Observatory at NASA Earth Exchange: Earth Monitoring from Operational Geostationary Satellite SystemsabstractThe latest generation of geostationary satellites (Himawari 8/9, GOES-16/17, FY-4, GK-2A) carries sensors that closely mimic the spatial and spectral characteristics of widely used polar-orbiting, global monitoring sensors such as MODIS and VIIRS. When combined, data from various currently operating/planned geostationary platforms provide a geo-ring of hyper-temporal (5-10 minutes), multispectral observations at spatial resolutions as high as 500 m. These high frequency observations offer exciting new possibilities for monitoring our planet, including better retrievals of geophysical variables by overcoming cloud cover, enabling studies of diurnally varying phenomena in the atmosphere, land, and the oceans, and support operational decision-making in agriculture, hydrology and disaster management. The NASA Earth Exchange (NEX) team, in collaboration with scientists from JAXA, KARI, NOAA and other international institutions, created the GeoNEX (www.nasa.gov/geonex) pipeline to integrate data from all available geostationary platforms and produce and distribute spatially, temporally, and radiometrically consistent data for the earth science community. We envision various institutions adapting the Geo component (e.g., GeoNOAA, GeoKARI, GeoChiba, GeoJAXA, GeoCMA) and customizing the pipeline and downstream products to serve the local/regional research and applied science communities. Ramakrishna R. Nemani, Weile Wang, Hirofumi Hashimoto, Andrew R. Michaelis, Thomas Vandal, Alexei I. Lyapustin, Jia Zhang 0001, Tsengdar J. Lee, Satya Kalluri, Hideaki Takenaka, Atsushi Higuchi, Kazuhito Ichii, Jong-Min Yeom |
IGARSS | 8 |
| 2018 | Unit of Work Supporting Generative Scientific Workflow Recommendation
Jia Zhang 0001, Maryam Pourreza, Seungwon Lee 0005, Ramakrishna R. Nemani, Tsengdar J. Lee |
ICSOC | 5 |
| 2017 | A Fine-Grained API Link Prediction Approach Supporting Mashup RecommendationabstractService (API) discovery and recommendation is key to the wide spread of service oriented architecture and service oriented software engineering. Service recommendation typically relies on service linkage prediction calculated by the semantic distances (or similarities) among services based on their collection of inherent attributes. Given a specific context (mashup goal), however, different attributes may contribute differently to a service linkage. In this paper, instead of training a model for all attributes as a whole, a novel approach is presented to simultaneously train separate models for individual attributes. Meanwhile, a latent attribute modeling method is developed to reveal context-aware attribute distribution. Experiments over real-world datasets have demonstrated that this fine-grained method yields higher link prediction accuracy. Qihao Bao, Jia Zhang 0001, Xiaoyi Duan, Rahul Ramachandran, Tsengdar J. Lee, Yankai Zhang, Seungwon Lee 0005, Patrick Gatlin, Manil Maskey |
ICWS | 5 |
| 2017 | Linking Design-Time and Run-Time: A Graph-Based Uniform Workflow Provenance ModelabstractWorkflow is an important way to mashup reusable software services to create value-added data analytics services. Workflow provenance is core to understand how services and workflows behaved in the past, which knowledge can be used to provide a better recommendation. Existing workflow provenance management systems handle various types of provenance separately. A typical data science exploration scenario, however, calls for an integrated view of provenance and seamless transition among different types of provenance. In this paper, a graph-based, uniform provenance model is proposed to link together design-time and run-time provenance, by combining retrospective provenance, prospective provenance, and evolution provenance. Such a unified provenance model will not only facilitate workflow mining and exploration, but also facilitate workflow interoperability. The model is formalized into colored Petri nets for verification and monitoring management. A SQL-like query language is developed, which supports basic queries, recursive queries, and cross-provenance queries. To verify the effectiveness of our model, A web-based, collaborative workflow prototyping system is developed as a proof-of-concept. Experiments have been conducted to evaluate the effectiveness of the proposed SQL-like graph query against SQL query. Xiaoyi Duan, Jia Zhang 0001, Qihao Bao, Rahul Ramachandran, Tsengdar J. Lee, Seungwon Lee 0005 |
ICWS | 5 |
| 2016 | A Bloom Filter-Powered Technique Supporting Scalable Semantic Service Discovery in Service NetworksabstractAs more and more reusable web services are published on the Internet, how to help users quickly identify appropriate candidate services has become an increasingly critical challenge. Most of the current research efforts on service discovery rely on syntax and semantics-based service matchmaking. In contrast, this paper presents a novel way of applying network routing mechanism to facilitate service discovery, featuring scalability and performance. Services annotated by Web Ontology Language for Services (OWL-S) are organized into a network based on semantic clustering. Virtual routers are created representing clusters, and Bloom Filters are generated for service routing. A service search request is thus transformed into a network routing problem to quickly locate semantic service cluster and in turn to candidate services. In addition, the deterministic annealing technique is applied to facilitate service classification in the network construction. Dynamic network adjustment is operated to ensure the search performance in the network. Empirical study over common testbed annotated in OWL-S is reported. Jia Zhang 0001, Runyu Shi, Shenggu Lu, Yuanchen Bai, Qihao Bao, Tsengdar J. Lee, Kiran Nagaraja, Nimish Radia |
ICWS | 7 |
| 2015 | Climate Analytics Workflow Recommendation as a Service - Provenance-Driven Automatic Workflow MashupabstractExisting scientific workflow tools, created by computer scientists, require that domain scientists meticulously design their multi-step experiments before analyzing data. However, this is oftentimes contradictory to a domain scientist's routine of conducting research and exploration. This paper presents a novel way to resolve this dispute, in the context of service-oriented science. After scrutinizing how Earth scientists conduct data analytics research in their daily work, a provenance model is developed to record their activities. Reverse-engineering the provenance, a technology is developed to automatically generate workflows for scientists to review and revise, supported by a Petri nets-based workflow verification instrument. In addition, dataset is proposed to be treated as first-class citizen to drive the knowledge sharing and recommendation. A data-centric repository infrastructure is established to catch richer provenance to further facilitate collaboration in the science community. In this way, we aim to revolutionize computer-supported Earth science. Jia Zhang 0001, Wei Wang 0208, Chris Lee 0002, Seungwon Lee 0005, Tsengdar J. Lee |
ICWS | 7 |
| 2013 | Bridging VisTrails Scientific Workflow Management System to High Performance ComputingabstractNASA Earth Exchange (NEX) is a collaboration platform whose goal is to accelerate Earth science research, by leveraging NASA's vast collections of global satellite data together with access to NASA's High-End Computing (HEC) facilities. NEX also aims to facilitate the sharing of experimental results as well as scientific processes (workflows) with the Earth science community through integration with VisTrails workflow management system. While VisTrails is used internally, it is not easily accessible from remote computers without directly logging into the NASA HEC systems through twofactor authentication and a bastion host. This paper describes the initial design of an extensible architecture that facilitates easier workflow interaction on NEX, by enabling users to develop and execute workflows in a supercomputing environment directly from their local VisTrails installation. This architecture helps domain scientists seamlessly leverage distributed computing and storage resources and it is potentially applicable to other scientific workflow management software. We further describe the architecture of the VisTrails-HEC plugin (as well as the VisTrails-Amazon plugin) and the implementation of a working prototype to demonstrate the feasibility of our solution. Jia Zhang 0001, Petr Votava, Tsengdar J. Lee, Owen Chu, Clyde Li, Kate Liu, Norman Xin, Ramakrishna R. Nemani |
SERVICES | 3 |