Zheng Zheng 0005

dblp:35/2837-5 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
4since 2021 · last 2025
0000-0002-6413-4126ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 7 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Radgen: A Cross-Modal Fusion System for Automated Radiology Report Generation
abstract
Automating radiology report generation can significantly reduce the workload of radiologists while improving the accuracy and consistency of clinical documentation. However, achieving optimal alignment between visual and textual representations in medical imaging remains a challenge. To address this, we demonstrate RadGen, a cross-modal fusion based system for automated medical report generation. RadGen uses MedCLIP as both a vision extractor and a retrieval mechanism to enhance the integration of imaging and textual data. By extracting features from retrieved reports and medical images through an attentionbased extraction module and integrating them with a fusion module, our system improves the coherence, accuracy, and clinical relevance of generated reports.
Qianhao Han, Daniel Ding, Zengchang Qin, Zheng Zheng 0005
CBMS5
2024 Integrating MedCLIP and Cross-Modal Fusion for Automatic Radiology Report Generation
abstract
Automating radiology report generation can significantly reduce the workload of radiologists and enhance the accuracy, consistency, and efficiency of clinical documentation. We propose a novel cross-modal framework that uses MedCLIP as both a vision extractor and a retrieval mechanism to improve the process of medical report generation. By extracting retrieved report features and image features through an attention-based extract module, and integrating them with a fusion module, our method improves the coherence and clinical relevance of generated reports. Experimental results on the widely used IU-Xray dataset demonstrate the effectiveness of our approach, showing improvements over commonly used methods in both report quality and relevance. Additionally, ablation studies provide further validation of the framework, highlighting the importance of accurate report retrieval and feature integration in generating comprehensive medical reports.
Qianhao Han, Zengchang Qin, Zheng Zheng 0005
IEEE Big Data4
2022 Confidence Bounded Replica Currency Estimation
abstract
Replicas of the same data item often exhibit varying consistency levels when executing read and write requests due to system availability and network limitations. When one or more replicas respond to a query, estimating the currency (or staleness) of the returned data item (without accessing the other replicas) is essential for applications requiring timely data. Depending on how confident the estimation is, the query may dynamically decide to return the retrieved replicas, or wait for the remaining replicas to respond. The replica currency estimation is expected to be accurate and extremely time efficient without introducing large overhead during query processing. In this paper, we provide theoretical bounds on the confidence of replica currency estimation. Our system computes with a minimum probability p, whether the retrieved replicas are current or stale. Using this confidence-bounded replica currency estimation, we implement a novel DYNAMIC read consistency level in the open-source, NoSQL database, Cassandra. Experiments show that the proposed replica currency estimation is intuitive and efficient. In most tested scenarios, with various query loads and cluster configurations, we show our estimations with confidence levels of at least 0.99 while keeping query latency low (close to reading ONE replica). Moreover, the overheads introduced due to estimation scoring and training are low, incurring only 0.76% to 1.17% of the query processing and replica synchronization time costs, respectively.
Yu Sun 0027, Zheng Zheng 0005, Shaoxu Song, Fei Chiang
SIGMOD Conference2
2021 Contextual Data Cleaning with Ontology FDs
abstract
Functional Dependencies (FDs) define attribute relationships based on syntactic equality, and, when used in data cleaning, they erroneously label syntactically different but semantically equivalent values as errors. We motivate the need to include context in data cleaning in order to account for the subjective nature of data quality. We enhance dependency-based data cleaning with Ontology Functional Dependencies (OFDs), which express semantic attribute relationships such as synonyms and is-a hierarchies defined by an ontology. We study the data and ontology repair problem for a set of OFDs, and propose an algorithm that finds the best ontological interpretation of the data that minimizes the number of repairs.
Zheng Zheng 0005
SIGMOD Conference1
2019 CurrentClean: Interactive Change Exploration and Cleaning of Stale Data
abstract
Enterprises often assume their data is up-to-date, where the presence of a timestamp in the recent past qualifies the data as current. However, entities modeled in the data experience varying rates of change that influence data currency. We argue that data currency is a relative notion based on individual spatio-temporal update patterns, and these patterns can be learned and predicted. We develop CurrentClean, a probabilistic system for identifying and cleaning stale values, and enables a user to interactively explore change in her data. Our system provides a Web-based user-interface, and a backend infrastructure that learns update correlations among cell values in a database to infer and repair stale values. Our demonstration provides two motivating scenarios that highlight change exploration, and cleaning features using clinical, and sensor data from a data centre enterprise.
Zheng Zheng 0005, Tri Minh Quach, Ziyi Jin, Fei Chiang, Mostafa Milani
CIKM1
2019 CurrentClean: Spatio-Temporal Cleaning of Stale Data
abstract
Data currency is imperative towards achieving up-to-date and accurate data analysis. Data is considered current if changes in real world entities are reflected in the database. When this does not occur, stale data arises. Identifying and repairing stale data goes beyond simply having timestamps. Individual entities each have their own update patterns in both space and time. These update patterns can be learned and predicted given available query logs. In this paper, we present CurrentClean, a probabilistic system for identifying and cleaning stale values. We introduce a spatio-temporal probabilistic model that captures the database update patterns to infer stale values, and propose a set of inference rules that model spatio-temporal update patterns commonly seen in real data. We recommend repairs to clean stale values by learning from past update values over cells. Our evaluation shows CurrentClean's effectiveness to identify stale values over real data, and achieves improved error detection and repair accuracy over state-of-the-art techniques.
Mostafa Milani, Zheng Zheng 0005, Fei Chiang
ICDE2
2018 FastOFD: Contextual Data Cleaning with Ontology Functional Dependencies
Zheng Zheng 0005, Morteza Alipour Langouri, Ian Currie, Fei Chiang, Lukasz Golab, Jarek Szlichta
EDBT1
2016 PARC: Privacy-Aware Data Cleaning
abstract
Poor data quality has become a persistent challenge for organizations as data continues to grow in complexity and size. Existing data cleaning solutions focus on identifying repairs to the data to minimize either a cost function or the number of updates. These techniques, however, fail to consider underlying data privacy requirements that exist in many real data sets containing sensitive and personal information. In this demonstration, we present PARC, a Privacy-AwaRe data Cleaning system that corrects data inconsistencies w.r.t. a set of FDs, and limits the disclosure of sensitive values during the cleaning process. The system core contains modules that evaluate three key metrics during the repair search, and solves a multi-objective optimization problem to identify repairs that balance the privacy vs. utility tradeoff. This demonstration will enable users to understand: (1) the characteristics of a privacy-preserving data repair; (2) how to customize data cleaning and data privacy requirements using two real datasets; and (3) the distinctions among the repair recommendations via visualization summaries.
Dejun Huang, Dhruv Gairola, Yu Huang 0019, Zheng Zheng 0005, Fei Chiang
CIKM4