VLDB 2026 Research / reviewers in the wild / expert
Basel Shbita
dblp:266/0414
· DBLP profile ↗
9ranked-venue papers
3as first author
7since 2021 · last 2026
0009-0006-3154-3501ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 8 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Context-Aware Visual Multi-turn Conversation Generation from Wikipedia and Wikidata
Basel Shbita, Pengyuan Li 0001, Anna Lisa Gentile |
ESWC (2) | 1 |
| 2025 | Exploiting Polygon Metadata to Colorize Draft MapsabstractBlack-and-white draft geological maps, produced during fieldwork, often contain dense handwritten annotations overlaid on monochromatic contour basemaps. Although interpretable in grayscale, the lack of color makes it difficult to visually distinguish overlapping or adjacent geological units, especially when boundaries are unclear and annotation styles vary. However, colorizing these draft maps is labor-intensive but essential, as they may be the only source of detailed geological information for certain regions. This hinders both human interpretation and downstream tasks such as map digitization and critical mineral resource assessment. We target the problem of automated colorization of draft geological maps. The challenge lies in interpreting noisy visual cues from uncolored sketches and assigning appropriate colors according to their geological categories. We propose a novel machine learning approach that exploits polygon metadata, including map keys that explicitly define geological units and implicitly suggest their intended colors, along with the semantic interpretation of the sketch content in the maps, to assign colors to the draft maps accordingly. We evaluate our method on USGS draft geological maps; it outperforms comparative methods by 15.7%. In addition, our approach improves downstream polygon-extraction performance by 9% in F1 score. Fandel Lin, Craig A. Knoblock, Basel Shbita, Yao-Yi Chiang |
SIGSPATIAL/GIS | 4 |
| 2025 | Exploiting LLMs and Semantic Technologies to Build a Knowledge Graph of Historical Mining Data
Craig A. Knoblock, Basel Shbita, Yao-Yi Chiang, Pothula Punith Krishna, Goran Muric, Jiyoon Pyo, Adriana Trejo-Sheu, Meng Ye 0002 |
ISWC (2) | 3 |
| 2024 | Exploiting Distant Supervision to Learn Semantic Descriptions of Tables with Overlapping Data
Craig A. Knoblock, Basel Shbita, Fandel Lin |
ISWC (2) | 3 |
| 2023 | Understanding Customer Requirements - An Enterprise Knowledge Graph Approach
Basel Shbita, Anna Lisa Gentile, Pengyuan Li 0001, Chad DeLuca |
ESWC | 1 |
| 2023 | Exploiting Polygon Metadata to Understand Raster Maps - Accurate Polygonal Feature ExtractionabstractLocating undiscovered deposits of critical minerals requires accurate geological data. However, most of the 100,000 historical geological maps of the United States Geological Survey (USGS) are in raster format. This hinders critical mineral assessment. We target the problem of extracting geological features represented as polygons from raster maps. We exploit the polygon metadata that provides information on the geological features, such as the map keys indicating how the polygon features are represented, to extract the features. We present a metadata-driven machine-learning approach that encodes the raster map and map key into a series of bitmaps and uses a convolutional model to learn to recognize the polygon features. We evaluated our approach on USGS geological maps; our approach achieves a median F1 score of 0.809 and outperforms state-of-the-art methods by 4.52%. Fandel Lin, Craig A. Knoblock, Basel Shbita, Zekun Li 0007, Yao-Yi Chiang |
SIGSPATIAL/GIS | 3 |
| 2021 | Artificial Intelligence for Modeling Complex Systems: Taming the Complexity of Expert Models to Improve Decision MakingabstractMajor societal and environmental challenges involve complex systems that have diverse multi-scale interacting processes. Consider, for example, how droughts and water reserves affect crop production and how agriculture and industrial needs affect water quality and availability. Preventive measures, such as delaying planting dates and adopting new agricultural practices in response to changing weather patterns, can reduce the damage caused by natural processes. Understanding how these natural and human processes affect one another allows forecasting the effects of undesirable situations and study interventions to take preventive measures. For many of these processes, there are expert models that incorporate state-of-the-art theories and knowledge to quantify a system's response to a diversity of conditions. A major challenge for efficient modeling is the diversity of modeling approaches across disciplines and the wide variety of data sources available only in formats that require complex conversions. Using expert models for particular problems requires integration of models with third-party data as well as integration of models across disciplines. Modelers face significant heterogeneity that requires resolving semantic, spatiotemporal, and execution mismatches, which are largely done by hand today and may take more than 2 years of effort. We are developing a modeling framework that uses artificial intelligence (AI) techniques to reduce modeling effort while ensuring utility for decision making. Our work to date makes several innovative contributions: (1) an intelligent user interface that guides analysts to frame their modeling problem and assists them by suggesting relevant choices and automating steps along the way; (2) semantic metadata for models, including their modeling variables and constraints, that ensures model relevance and proper use for a given decision-making problem; and (3) semantic representations of datasets in terms of modeling variables that enable automated data selection and data transformations. This framework is implemented in the MINT (Model INTegration) framework, and currently includes data and models to analyze the interactions between natural and human systems involving climate, water availability, agricultural production, and markets. Our work to date demonstrates the utility of AI techniques to accelerate modeling to support decision-making and uncovers several challenging directions for future work. Yolanda Gil, Daniel Garijo, Deborah Khider, Craig A. Knoblock, Varun Ratnakar, Maximiliano Osorio, Hernán Vargas, Minh Pham 0004, Jay Pujara, Basel Shbita, Yao-Yi Chiang, Dan Feldman, Yijun Lin 0001, Hayley Song, Vipin Kumar 0001, Ankush Khandelwal, Michael S. Steinbach, Kshitij Tayal, Shaoming Xu, Suzanne A. Pierce, Lissa Pearson, Daniel Hardesty-Lewis, Ewa Deelman, Rafael Ferreira da Silva, Rajiv Mayani, Armen R. Kemanian, Lorne Leonard, Scott D. Peckham, Maria Stoica 0001, Kelly M. Cobourn, Zeya Zhang, Christopher J. Duffy, Lele Shu |
ACM Trans. Interact. Intell. Syst. | 10 |
| 2020 | Building Linked Spatio-Temporal Data from Vectorized Historical Maps
Basel Shbita, Craig A. Knoblock, Yao-Yi Chiang, Johannes H. Uhl, Stefan Leyk |
ESWC | 1 |
| 2020 | An Automatic Approach for Generating Rich, Linked Geo-Metadata from Historical Map ImagesabstractHistorical maps contain detailed geographic information difficult to find elsewhere covering long-periods of time (e.g., 125 years for the historical topographic maps in the US). However, these maps typically exist as scanned images without searchable metadata. Existing approaches making historical maps searchable rely on tedious manual work (including crowd-sourcing) to generate the metadata (e.g., geolocations and keywords). Optical character recognition (OCR) software could alleviate the required manual work, but the recognition results are individual words instead of location phrases (e.g., "Black'' and "Mountain'' vs. "Black Mountain''). This paper presents an end-to-end approach to address the real-world problem of finding and indexing historical map images. This approach automatically processes historical map images to extract their text content and generates a set of metadata that is linked to large external geospatial knowledge bases. The linked metadata in the RDF (Resource Description Framework) format support complex queries for finding and indexing historical maps, such as retrieving all historical maps covering mountain peaks higher than 1,000 meters in California. We have implemented the approach in a system called mapKurator. We have evaluated mapKurator using historical maps from several sources with various map styles, scales, and coverage. Our results show significant improvement over the state-of-the-art methods. The code has been made publicly available as modules of the Kartta Labs project at https://github.com/kartta-labs/Project. Zekun Li 0007, Yao-Yi Chiang, Sasan Tavakkol, Basel Shbita, Johannes H. Uhl, Stefan Leyk, Craig A. Knoblock |
KDD | 4 |