Siying Chen

dblp:55/10674 · DBLP profile ↗
← Back
14ranked-venue papers
5as first author
11since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 8 · 3 first-author · 5 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 4 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 UniClean: A Multi-Signal Fusion Pipeline for Optimizing Data Cleaning Workflow
abstract
Data quality issues are prevalent in information systems, making data cleaning a complex and time-consuming task, particularly with large-scale datasets and the lack of standardized automated processes. Existing cleaning pipelines often lack automated schemes to guide the execution of cleaning algorithms and the sequence of error corrections, limiting their practicality in real-world big data applications. To address the growing demand for advanced cleaning tools driven by the complexity of data activities, we propose the UniClean framework for on-demand big data cleaning. UniClean employs a unified cleaning operation (Uniop) from multiple cleaners to optimize the data cleaning workflow. It integrates a cleaning parameter generation pipeline, a cleaning parameter selection pipeline, and a module for cleaning process preparation and optimization, covering the entire workflow from cleaner modeling and data preparation to optimal cleaning operation generation. UniClean provides an adaptive (data-driven cleaning workflow generation) and flexible (multi-signal extension system) solution to meet the urgent need for high-quality data in today's data-driven decision-making environments. We demonstrate how UniClean effectively addresses the challenges of big data cleaning across diverse information system landscapes.
Xiaoou Ding, Zekai Qian, Hongzhi Wang 0001, Siying Chen, Hongbin Su
ICDE5
2025 CBAClean:A Comprehensive System for Recommending Data Cleaning Solutions Through Cost-Benefit Analysis in Data Quality Management
abstract
The scale of data analysis tasks have increased, highlighting the critical importance of data quality. Data quality assessment and repair have become pivotal in data preparation. Despite the availability of numerous algorithms for data cleaning, these often focus on optimizing efficiency and minimizing labor costs, neglecting the explicit relationship between data quality management costs and benefits. This omission can lead to the failure of promising data analysis solutions. To address this, we propose CBAClean, a comprehensive system that integrates cost-benefit analysis into data cleaning. CBAClean aims to assist users in quantifying the costs of data quality management and providing optimal data cleaning solutions tailored to their needs. Key features include task-centered multi-perspective data quality assessment, a comprehensive data quality repair operator library, fine-grained human role division for effective cost control, and recommendation of optimal data cleaning solutions based on cost-benefit calculations. By incorporating cost-benefit analysis, CBAClean enhances the practical application of data quality management on real-world data governance platforms.
Xiaoou Ding, Hongbin Su, Zekai Qian, Wenxuan Cui, Siying Chen, Zheng Liang 0002, Chen Wang 0018, Hongzhi Wang 0001
ICDE5
2025 UniClean: A Scalable Data Cleaning Solution for Mixed Errors based on Unified Cleaners and Optimized Cleaning Workflow
abstract
Data cleaning is an essential technique to enhance data quality. Despite the proposal of various algorithms with different cleaning strategies, current automated cleaning technologies still fall short of practical requirements when dealing with large-scale data containing mixed errors. This paper presents UniClean to efficiently solve the mixed error cleaning problem with three key technical contributions. (1) A unified construction and extension method for cleaners, enabling cleaning methods to easily utilize various cleaners to perform cleaning tasks. (2) Three optimization strategies to achieve efficiency-oriented cleaning preparation. (3) A cleaning algorithm based on an optimized cleaning process to effectively clean mixed errors. UniClean achieves a time complexity of O (| D error | 4 · | Op | + |D| · | D error |), significantly enhancing scalability. Experiments on public and large-scale enterprise datasets demonstrate that UniClean achieves over 40% improvement across five metrics, compared to five state-of-the-art cleaning methods, and delivers more than 30% gains in F1 and REDR on complex datasets, while completing the cleaning process within hours even for millions of records.
Xiaoou Ding, Zekai Qian, Hongzhi Wang 0001, Siying Chen, Yafeng Tang, Hongbin Su, Chen Wang 0018
Proc. VLDB Endow.4
2025 A layered stock prediction model based on novel feature selection and model parameter optimization
Siying Chen, Guoqiang Tang
J. Supercomput.1
2024 Hyper-parameter Recommendation for Truth Discovery
Siying Chen, Xiaoou Ding, Zheng Liang 0002, Yafeng Tang, Hongzhi Wang 0001
DASFAA (3)1
2024 Uncertainty of Pure Rotational Raman-Rayleigh LiDAR for Temperature Measurement in Middle-to-Upper Atmosphere: Evaluation Method
abstract
This article constitutes the second installment in a series dedicated to exploring the uncertainties associated with pure rotational Raman–Rayleigh temperature measurement LiDARs in the middle-to-upper atmosphere (20–90 km). It presents uncertainty evaluation methods aimed at addressing the challenge of uncertainty assessment. These methods leverage both the Guide to the Expression of Uncertainty in Measurement (GUM) method and the Monte Carlo method (MCM) to ascertain the uncertainty of air density and temperature derived from LiDAR signals. Within the framework of these evaluation methods, various uncertainty sources are considered, encompassing saturation correction, photon noise, background noise, reference temperature, and two special uncertainty sources: atmospheric transmittance correction in Rayleigh LiDAR and the calibration process in Raman LiDAR. An uncertainty evaluation example in accordance with the proposed evaluation method is provided by employing raw signals from full-chain simulation system introduced in another series paper. In this example, we compare the evaluation results of the GUM method with those derived from the MCM. Our findings reveal consistent trends in different uncertainty components between the two methods. In addition, the uncertainty caused by atmospheric transmittance correction is involved with detailed considerations of correlation and iterative methods. The uncertainty resulting from the calibration process in Raman LiDAR, attributed to its nonlinear characteristics, can only be evaluated using the MCM. These findings affirm the applicability of the proposed uncertainty evaluation methods of pure rotational Raman–Rayleigh LiDAR. This research enhances our understanding of the sources of uncertainty in LiDAR systems and provides a convenient approach for evaluating the data quality of temperature LiDAR.
Rongzheng Cao, Siying Chen, Wangshu Tan, He Chen 0005, Pan Guo, Yinghong Yu, Huiyun Wu
IEEE Trans. Geosci. Remote. Sens.2
2024 Uncertainty of Pure Rotational Raman-Rayleigh LiDAR for Temperature Measurement in Middle-to-Upper Atmosphere: Simulation Method
abstract
This article initiates a series focusing on the uncertainties associated with temperature measurements using pure rotational Raman-Rayleigh LiDAR in the mid-to-upper atmosphere (20–90 km). We introduce a comprehensive simulation system designed for temperature measurement using pure rotational Raman-Rayleigh LiDARs. This simulation system considers hardware parameter fluctuations, atmospheric parameter variations, and detailed retrieval algorithms during the temperature detection process. Using the Monte Carlo method (MCM), the system achieves, for the first time, a simulation of temperature measurement uncertainties throughout the entire LiDAR measurement process. This article provides an illustrative example of applying the uncertainty simulation system to a prototype. In this example, photon noise (PN), the reference temperature (RT), the laser wavelength (LW), and saturation correction (SC) are significant sources of uncertainty in Raman LiDAR, whereas in Rayleigh LiDAR, PN, SC, and the RT play major roles. Additionally, sensitivity experiments of the uncertainty components are carried out to analyze the linearity of the uncertainty propagation in the two LiDAR systems. Then, two special cases of uncertainty coupling are discussed, which reveals the complexity of uncertainty propagation in LiDARs. The simulation method proposed in this article aids in identifying key sources of measurement uncertainty during the entire measurement process, offering insights for hardware selection and system design in LiDAR development.
Siying Chen, Rongzheng Cao, Wangshu Tan, He Chen 0005, Pan Guo, Yinghong Yu, Shusen Yao
IEEE Trans. Geosci. Remote. Sens.1
2023 Improved Aerosol Lidar Ratio Profile by Introducing Pseudo-Constant
abstract
Aerosol lidar ratio (LR) is important for the inversion of aerosol optical characteristics. Single aerosol LR value adopted in the inversion not only causes difficulties in improving aerosol inversion accuracy but also hinders the analysis of aerosol types. In this article, a pseudo-constant method with certain universality is proposed for inverting aerosol LR profiles. Simulation results show a decrease in the mean absolute error (MAE) of the aerosol LR from 3.57% (iteration method) and 9.96% (Fernald method) to 1.82% (pseudo-constant method). In the near-surface region (< 1.5 km), the MAE of aerosol LR decreases from 4.97% (iteration) and 21.66% (Fernald) to 1.54% (pseudo-constant). Moreover, the inversion accuracy of aerosol LR obtained from pseudo-constant method is less influenced by the precision of the pre-estimated aerosol LR value than that from iteration method. Experimentally, the aerosol LR and aerosol extinction coefficient (AEC) profiles obtained from the pseudo-constant method are more accurate than those from the iteration and Fernald methods. In addition, both the simulations and experiments indicate that the pseudo-constant method has higher efficiency; the processing time is shortened from approximately 5 s/signal (iteration) to approximately 0.1 s/signal. Finally, aerosol types can be distinguished to a certain degree by combining the improved aerosol LR profile.
Hongzhu Ji, Siying Chen, Pan Guo, Jingxi He, Yuefeng Zhao, Xinye Fan
IEEE Trans. Geosci. Remote. Sens.2
2022 Reinforcement Learning Based Diagnosis and Prediction for COVID-19 by Optimizing a Mixed Cost Function From CT Images
abstract
A novel coronavirus disease (COVID-19) is a pandemic disease has caused 4 million deaths and more than 200 million infections worldwide (as of August 4, 2021). Rapid and accurate diagnosis of COVID-19 infection is critical to controlling the spread of the epidemic. In order to quickly and efficiently detect COVID-19 and reduce the threat of COVID-19 to human survival, we have firstly proposed a detection framework based on reinforcement learning for COVID-19 diagnosis, which constructs a mixed loss function that can integrate the advantages of multiple loss functions. This paper uses the accuracy of the validation set as the reward value, and obtains the initial model for the next epoch by searching the model corresponding to the maximum reward value in each epoch. We also have proposed a prediction framework that integrates multiple detection frameworks using parameter sharing to predict the progression of patients' disease without additional training. This paper also constructed a higher-quality version of the CT image dataset containing 247 cases screened by professional physicians, and obtained more excellent results on this dataset. Meanwhile, we used the other two COVID-19 datasets as external verifications, and still achieved a high accuracy rate without additional training. Finally, the experimental results show that our classification accuracy can reach 98.31%, and the precision, sensitivity, specificity, and AUC (Area Under Curve) are 98.82%, 97.99%, 98.67%, and 0.989, respectively. The accuracy of external verification can reach 93.34% and 91.05%. What's more, the accuracy of our prediction framework is 91.54%. A large number of experiments demonstrate that our proposed method is effective and robust for COVID-19 detection and prediction.
Siying Chen, Minghui Liu 0002, Jiali Deng, Tianshu Xie, Libo Xie, Hai-gang Gong, Lifeng Xu, Hong Pu, Ming Liu 0002
IEEE J. Biomed. Health Informatics1
2021 Chasing Play on TikTok from Populations with Disabilities to Inspire Playful and Inclusive Technology Design
abstract
There is an open call for technology to be more playful [5, 79] and for tech design to be more inclusive of people with disabilities [80]. In the era of COVID19, it is often unsafe for the public in general and people with disabilities, in particular, to engage in in-person design exercises using traditional methods. This presents a missed opportunity as these populations are already sharing playful content rich with tacit design knowledge that can be used to inspire the design of playful everyday technology. We present our process of scraping play potentials [4] from TikTok from content creators with disabilities to generate design concepts that may inspire future technology design. We share 7 emerging themes from the scraped content, a catalog of design concepts that may inspire designers, and discuss the relevance of the emerging themes and possible implications for the design concepts.
Jared Duval, Ferran Altarriba Bertran, Siying Chen, Melissa Chu, Divya Subramonian, Austin Wang, Geoffrey Xiang, Sri Hastuti Kurniawan, Katherine Isbister
CHI3
2021 HSGACN: Hyperspectral Image Classification Algorithm Based on Graph Convolutional Network
abstract
Convolutional neural network has been widely used in hyperspectral image classification. Compared with the early machine learning method, it has made great progress. However, the convolution kernel used in hyperspectral image classification ignores the intrinsic relationship among spatial pixels when extracting spectral features, which will lead to poor contour and very small false prediction in the classification results. Besides, The hyperspectral data can only be labeled by experts, which requires a lot of labor and material resources. In order to improve the classification accuracy of hyperspectral images and reduce the dependence on labeled samples, this paper proposed a hyperspectral image classification algorithm based on graph neural network. Through the characteristics of inherent points and edges in the graph, the spatial information and spectral information of hyperspectral images are fused. The features of unlabeled samples are used to participate in the training to improve the effect of classification model.
Yi Xiao 0007, Siying Chen, Zongyao Yin, Ruiqing Yan, Xianchuan Yu
IGARSS2
2020 Unbalanced Geologic Body Classification of Hyperspectral Data Based on Squeeze and Excitation Networks at Tianshan Area
abstract
Hyperspectral data contains abundant spectral do main information, which is of great significance to classification of objects. However, due to the lack of labeled data, it is difficult to get an acceptable result by just using the small number of labeled data. We propose a semi-supervised classification model based on convolutional neural network and introduce the attention mechanism to balance the sample weight. After the convolution of the multi-layer network, more information is concentrated on the channels, so we use the Squeeze-and-Excitation block, which can adaptively recalibrates the channel-wise characteristic response by explicitly modelling the inter-channel dependencies. At the same time, we used focal loss to reduce the problem of poor training caused by uneven samples. We test our model on hyperspectral data at Tianshan area. From the result, we can find that our method can get a great result on the mineral classification task, which can be used for making geological map.
Ying Cao 0009, Yasmine Medjadba, Yuntao Wang 0006, RunCheng Jiao, Siying Chen, Xianchuan Yu
IGARSS9
2013 Development of Standardized Patient Scenarios for Usability Testing of Medication Alerts
Brittany Melton, Jeffery R. Spina, Alan J. Zillich, Jason J. Saleem, Michael Weiner 0002, Scott A. Russell, Siying Chen, Alissa L. Russ-Jara
AMIA7
2012 Boresight Calibration of Airborne LiDAR System Without Ground Control Points
abstract
This letter proposes a new method for boresight misalignment calibration of the charge-coupled device (CCD) camera which is one component of an airborne light detection and ranging (LiDAR) system without ground control points (GCPs). In the calibration, tie points in overlapping areas are first selected, and then, a multibaseline forward intersection is used for calculating object coordinates of these points. In the intersection, exterior elements of the CCD camera are obtained directly from positioning and orientation system (POS) data of the LiDAR system, which are error contaminated mainly due to the unparallel relation between the frameworks of the inertial measurement unit of the POS and the CCD camera. Elevation values of the ground points are then refined by those obtained from LiDAR point clouds by interpolation, which can be considered to be more accurate than those obtained by multibaseline forward intersection. Through projecting the ground points with refined elevation values into the image space by collinear equations and minimizing distances between the image points selected manually and those projected from ground points, the boresight misalignment is removed effectively. Therefore, the proposed method without GCPs in the whole process is more flexible than other traditional photogrammetric ways.
Siying Chen, Hongchao Ma, Yinchao Zhang, Jixian Xu, He Chen 0005
IEEE Geosci. Remote. Sens. Lett.1