VLDB 2026 Research / reviewers in the wild / expert
Amara Tariq
dblp:67/10440
· DBLP profile ↗
18ranked-venue papers
12as first author
10since 2021 · last 2026
0000-0001-5932-2491ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 10 · 5 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 1 since 2021Artificial intelligence and machine learning · 4 · 4 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Extraction of distant recurrence sites for breast cancer patients from free-text clinical notes using large language models
Madhu Babu Sikha, Amara Tariq, Allison W. Kurian, Kevin C. Ward, Theresa H. M. Keegan, Daniel L. Rubin, Imon Banerjee |
J. Biomed. Informatics | 2 |
| 2025 | Patient-centric Summarization of Radiology Findings Using Two-step Training of Large Language ModelsabstractEducation-level or socioeconomic background of patients may dictate their ability to understand medical jargon. Inability to understand primary findings from a radiology report may lead to unnecessary anxiety among patients or missed follow up. We aim to meet this challenge by developing a patient-sensitive summarization model for radiology reports. We selected computed tomography (CT) exams of chest as a use-case and collected 7,000 studies from Mayo Clinic. Summarization model was built on top of the T5 large language model (LLM) as our experiments indicated that its text-to-text transfer architecture was suited for abstractive text summarization, resulting in a model with 0.77B trainable parameters. Noisy ground truth for model training was collected by prompting LLaMA-13B model. We recruited experts (board-certified radiologists) and laymen to manually evaluate model-generated summaries generated by model. Our model rarely missed information as marked by majority opinion of radiologists. Laymen indicated 63% improvement in their understanding by reading model-generated layman summaries. Comparison with zero-shot performance of ChatGPT indicated that the proposed model reduced the rate of hallucination by half and rate of missing important information by fivefold. The proposed model can generate reliable summaries for radiology reports understandable by patients with vastly different levels of medical knowledge. Amara Tariq, Shubham Trivedi, Aisha Urooj Khan, Gokul Ramasamy, Sam Fathizadeh, Matthew Stib, Nelly Tan, Bhavik N. Patel, Imon Banerjee |
ACM Trans. Comput. Heal. | 1 |
| 2025 | Adaptable graph neural networks design to support generalizability for clinical event prediction
Amara Tariq, Gurkiran Kaur, Leon Su, Judy Gichoya, Bhavik N. Patel, Imon Banerjee |
J. Biomed. Informatics | 1 |
| 2024 | Knowledge-Grounded Adaptation Strategy for Vision-Language Models: Building a Unique Case-Set for Screening Mammograms for Residents Training
Aisha Urooj Khan, John W. Garrett, Tyler J. Bradshaw, Lonie Salkowski, Jiwoong Jason Jeong, Amara Tariq, Imon Banerjee |
MICCAI (12) | 6 |
| 2023 | Graph convolutional network-based fusion model to predict risk of hospital acquired infectionsabstractOBJECTIVE: Hospital acquired infections (HAIs) are one of the top 10 leading causes of death within the United States. While current standard of HAI risk prediction utilizes only a narrow set of predefined clinical variables, we propose a graph convolutional neural network (GNN)-based model which incorporates a wide variety of clinical features. MATERIALS AND METHODS: Our GNN-based model defines patients' similarity based on comprehensive clinical history and demographics and predicts all types of HAI rather than focusing on a single subtype. An HAI model was trained on 38 327 unique hospitalizations while a distinct model for surgical site infection (SSI) prediction was trained on 18 609 hospitalization. Both models were tested internally and externally on a geographically disparate site with varying infection rates. RESULTS: The proposed approach outperformed all baselines (single-modality models and length-of-stay [LoS]) with achieved area under the receiver operating characteristics of 0.86 [0.84-0.88] and 0.79 [0.75-0.83] (HAI), and 0.79 [0.75-0.83] and 0.76 [0.71-0.76] (SSI) for internal and external testing. Cost-effective analysis shows that the GNN modeling dominated the standard LoS model strategy on the basis of lower mean costs ($1651 vs $1915). DISCUSSION: The proposed HAI risk prediction model can estimate individualized risk of infection for patient by taking into account not only the patient's clinical features, but also clinical features of similar patients as indicated by edges of the patients' graph. CONCLUSIONS: The proposed model could allow prevention or earlier detection of HAI, which in turn could decrease hospital LoS and associated mortality, and ultimately reduce the healthcare cost. Amara Tariq, Lin Lancaster, Praneetha Elugunti, Eric Siebeneck, Katherine Noe, Bijan Borah, James Moriarty, Imon Banerjee, Bhavik N. Patel |
J. Am. Medical Informatics Assoc. | 1 |
| 2023 | Predicting 30-Day All-Cause Hospital Readmission Using Multimodal Spatiotemporal Graph Neural NetworksabstractReduction in 30-day readmission rate is an important quality factor for hospitals as it can reduce the overall cost of care and improve patient post-discharge outcomes. While deep-learning-based studies have shown promising empirical results, several limitations exist in prior models for hospital readmission prediction, such as: (a) only patients with certain conditions are considered, (b) do not leverage data temporality, (c) individual admissions are assumed independent of each other, which ignores patient similarity, (d) limited to single modality or single center data. In this study, we propose a multimodal, spatiotemporal graph neural network (MM-STGNN) for prediction of 30-day all-cause hospital readmission, which fuses in-patient multimodal, longitudinal data and models patient similarity using a graph. Using longitudinal chest radiographs and electronic health records from two independent centers, we show that MM-STGNN achieved an area under the receiver operating characteristic curve (AUROC) of 0.79 on both datasets. Furthermore, MM-STGNN significantly outperformed the current clinical reference standard, LACE+ (AUROC = 0.61), on the internal dataset. For subset populations of patients with heart disease, our model significantly outperformed baselines, such as gradient-boosting and Long Short-Term Memory models (e.g., AUROC improved by 3.7 points in patients with heart disease). Qualitative interpretability analysis indicated that while patients' primary diagnoses were not explicitly used to train the model, features crucial for model prediction may reflect patients' diagnoses. Our model could be utilized as an additional clinical decision aid during discharge disposition and triaging high-risk patients for closer post-discharge follow-up for potential preventive measures. Siyi Tang, Amara Tariq, Jared Dunnmon, Praneetha Elugunti, Daniel L. Rubin, Bhavik N. Patel, Imon Banerjee |
IEEE J. Biomed. Health Informatics | 2 |
| 2022 | Graph-based Fusion Modeling and Explanation for Disease Trajectory Prediction
Amara Tariq, Siyi Tang, Hifza Sakhi, Leo A. Celi, Janice M. Newsome, Daniel L. Rubin, Hari Trivedi, Judy Gichoya, Bhavik N. Patel, Imon Banerjee |
AMIA | 1 |
| 2022 | Bridging the Gap between Structured and Free-form Radiology Reporting: A Case-study on Coronary CT AngiographyabstractFree-form radiology reports associated with coronary computed tomography angiography (CCTA) include nuanced and complicated linguistics to report cardiovascular disease. Standardization and interpretation of such reports is crucial for clinical use of CCTA. Coronary Artery Disease Reporting and Data System (CAD-RADS) has been proposed to achieve such standardization by implementing a strict template-based report writing and assignment of a score between 0 and 5 indicating the severity of coronary artery lesions. Even after its introduction, free-form unstructured report writing remains popular among radiologists. In this work, we present our attempts at bridging the gap between structured and unstructured reporting by natural language processing. We present machine learning models that while being trained only on structured reports, can predict CAD-RADS scores by analysis of free-text of unstructured radiology reports. The best model achieves 98% accuracy on structured reports and 92% 1-margin accuracy (difference of \le 1 in the predicted and the actual scores) for free-form unstructured reports. Our model also performs well under very difficult circumstances including nuanced and widely varying terminology used for reporting cardiovascular functions and diseases, scarcity of labeled data for training our model, and uneven class label distribution. Amara Tariq, Marly van Assen, Carlo Nicola De Cecco, Imon Banerjee |
ACM Trans. Comput. Heal. | 1 |
| 2022 | A Socio-Technical Approach for Resilient Connected Transportation Systems in Smart CitiesabstractConnected transportation systems in smart cities can be regarded as Cyber-Physical-Social Systems (CPSSs) or Socio-Technical Systems (STSs) due to the presence of connectivity and complex interaction between human and technological infrastructure. Safety and security are essential requirements for such transportation CPSS. Specifically, resilience against cyber-attacks is crucial to address the vulnerabilities of Information and Communication Technologies (ICT) within the infrastructure. In this work, we explore a cyber-attack detection paradigm that leverages the socio-technical nature of such transportation CPSS. Essentially, we exploit the redundancies between physical signals (received from vehicle and infrastructure sensors) and social signals (received from consumers’ mobile devices and social media) to detect cyber-attack occurrences. The proposed scheme is developed combining system and control theoretic tools and natural language processing techniques. A case study of a vehicular platoon is considered on which extensive simulation studies have been performed. Such simulation results under different types of cyber-attacks illustrate the promising nature of the proposed approach. Tanushree Roy, Amara Tariq, Satadru Dey |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2021 | Query bot for retrieving patients' clinical history: A COVID-19 use-case
Yibo Wang 0001, Amara Tariq, Fiza Khan, Judy Gichoya, Hari Trivedi, Imon Banerjee |
J. Biomed. Informatics | 2 |
| 2020 | Domain Adaptation For Lane Marking: An Unsupervised ApproachabstractA major roadblock for designing deep learning based supervised solutions to any problem is the requirement of huge amount of labelled data for training. Domain adaptation techniques are designed to alleviate this problem. The aim of domain adaptation is to ensure that a model trained over available labelled data of source domain generalizes well over unlabelled data of target domain. Data distribution among source and target domains are different but related. Various techniques have been successfully applied to adapt deep learning based image classifiers from source to target domain data. In this paper, we expand the scope of domain adaptation by designing it for a much more complex image processing system, i.e., lane marking that involves instance segmentation. Automatic system to mark lanes in images is a crucial component of self-driving cars. Such systems have been trained over data from very specific domains such as highway traffic images. It is crucial for lane marking systems to be able to adapt to different domains as traffic scenarios vary across countries, road types, etc. We designed a lane marking model that successfully generalizes from highway-traffic images to images of traffic data from Lahore (provincial capital of Pakistan). Our system shows that domain adaptation significantly improves the performance of lane marking system for the unlabelled data of the target domain. Ammar Saqib, Sarah Sajid, Sheikh Mahad Arif, Amara Tariq, Nazim Ashraf |
ICIP | 4 |
| 2018 | Designing a symmetric classifier for image annotation using multi-layer sparse coding
Amara Tariq, Hassan Foroosh |
Image Vis. Comput. | 1 |
| 2017 | NELasso: Group-Sparse Modeling for Characterizing Relations Among Named Entities in News ArticlesabstractNamed entities such as people, locations, and organizations play a vital role in characterizing online content. They often reflect information of interest and are frequently used in search queries. Although named entities can be detected reliably from textual content, extracting relations among them is more challenging, yet useful in various applications (e.g., news recommending systems). In this paper, we present a novel model and system for learning semantic relations among named entities from collections of news articles. We model each named entity occurrence with sparse structured logistic regression, and consider the words (predictors) to be grouped based on background semantics. This sparse group LASSO approach forces the weights of word groups that do not influence the prediction towards zero. The resulting sparse structure is utilized for defining the type and strength of relations. Our unsupervised system yields a named entities' network where each relation is typed, quantified, and characterized in context. These relations are the key to understanding news material over time and customizing newsfeeds for readers. Extensive evaluation of our system on articles from TIME magazine and BBC News shows that the learned relations correlate with static semantic relatedness measures like WLM, and capture the evolving relationships among named entities over time. Amara Tariq, Asim Karim, Hassan Foroosh |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2017 | A Context-Driven Extractive Framework for Generating Realistic Image DescriptionsabstractAutomatic image annotation methods are extremely beneficial for image search, retrieval, and organization systems. The lack of strict correlation between semantic concepts and visual features, referred to as the semantic gap, is a huge challenge for annotation systems. In this paper, we propose an image annotation model that incorporates contextual cues collected from sources both intrinsic and extrinsic to images, to bridge the semantic gap. The main focus of this paper is a large real-world data set of news images that we collected. Unlike standard image annotation benchmark data sets, our data set does not require human annotators to generate artificial ground truth descriptions after data collection, since our images already include contextually meaningful and real-world captions written by journalists. We thoroughly study the nature of image descriptions in this real-world data set. News image captions describe both visual contents and the contexts of images. Auxiliary information sources are also available with such images in the form of news article and metadata (e.g., keywords and categories). The proposed framework extracts contextual-cues from available sources of different data modalities and transforms them into a common representation space, i.e., the probability space. Predicted annotations are later transformed into sentence-like captions through an extractive framework applied over news articles. Our context-driven framework outperforms the state of the art on the collected data set of approximately 20 000 items, as well as on a previously available smaller news images data set. Amara Tariq, Hassan Foroosh |
IEEE Trans. Image Process. | 1 |
| 2015 | Feature-independent context estimation for automatic image annotationabstractAutomatic image annotation is a highly valuable tool for image search, retrieval and archival systems. In the absence of an annotation tool, such systems have to rely on either users' input or large amount of text on the webpage of the image, to acquire its textual description. Users may provide insufficient/noisy tags and all the text on the webpage may not be a description or an explanation of the accompanying image. Therefore, it is of extreme importance to develop efficient tools for automatic annotation of images with correct and sufficient tags. The context of the image plays a significant role in this process, along with the content of the image. A suitable quantification of the context of the image may reduce the semantic gap between visual features and appropriate textual description of the image. In this paper, we present an unsupervised feature-independent quantification of the context of the image through tensor decomposition. We incorporate the estimated context as prior knowledge in the process of automatic image annotation. Evaluation of the predicted annotations provides evidence of the effectiveness of our feature-independent context estimation method. Amara Tariq, Hassan Foroosh |
CVPR | 1 |
| 2015 | T-clustering: Image clustering by tensor decompositionabstractImage clustering is an important tool for organizing evergrowing image repositories for efficient search and retrieval. A variety of clustering algorithms have been employed to cluster images. In this paper, we present a clustering algorithm, named T-Clustering, especially tailored to suit image collections. T-Clustering is based on tensor decomposition and takes into account the spatial configuration of images. This algorithm is non-parametric and works very well with raw images, thus alleviating the need for transformation of images in any feature domain. Our experiments prove that this algorithm outperforms well-known non-parametric clustering algorithms for a variety of image collections. Amara Tariq, Hassan Foroosh |
ICIP | 1 |
| 2014 | Scene-based automatic image annotationabstractImage search and retrieval systems depend heavily on availability of descriptive textual annotations with images, to match them with textual queries of users. In most cases, such systems have to rely on users to provide tags or keywords with images. Users may add insufficient or noisy tags. A system to automatically generate descriptive tags for images can be extremely helpful for search and retrieval systems. Automatic image annotation has been explored widely in both image and text processing research communities. In this paper, we present a novel approach to tackle this problem by incorporating contextual information provided by scene analysis of image. Image can be represented by features which indicate type of scene shown in the image, instead of representing individual objects or local characteristics of that image. We have used such features to provide context in the process of predicting tags for images. Amara Tariq, Hassan Foroosh |
ICIP | 1 |
| 2011 | Fast supervised feature extraction by term discrimination information poolingabstractDimensionality reduction (DR) through feature extraction (FE) is desirable for efficient and effective processing of text documents. Many of the techniques for text FE produce features that are not readily interpretable and require super-linear computation time. In this paper, we present a fast supervised DR/FE technique, named FEDIP, that is motivated by the notion of relatedness of terms to topics or contexts. This relatedness is quantified by using the discrimination information provided by a term for a topic in a labeled document collection. Features are constructed by pooling the discrimination information of highly related terms for each topic. FEDIP's time complexity is linear in the size of the vocabulary and document collection. FEDIP is evaluated for document classification with SVM and naive Bayes classifiers on six text data sets. The results show that FEDIP produces low-dimension feature spaces that yield higher classification accuracy when compared with LDA and LSI. FEDIP is also found to be significantly faster than the other techniques on our evaluation data sets. Amara Tariq, Asim Karim |
CIKM | 1 |