Che-Lun Hung

dblp:80/6405 · DBLP profile ↗
← Back
3ranked-venue papers in the field
0as first author
3since 2021 · last 2025
0000-0002-8906-9367ORCID · reported

Domains — venue-derived; a paper can count in several

Big Data, Cloud & Distributed Data Systems · 3
YearPublicationVenuePosition
2025 From Genetic Reads to Information Granules: Scalable Big NGS Data Cleaning with Apache Pig
Bozena Malysiak-Mrozek, Tomasz Sitek, Vaidy S. Sunderam, Boleslaw Pochopien, Krzysztof Tokarz, Che-Lun Hung, Dariusz Mrozek
IEEE Big Data6
2024 Fuzzy Querying in the Cloud-based Environment for Data Stream-driven Predictive Maintenance in AGV-enabled Smart Factories
abstract
Fuzzy data processing enables data enrichment and increases data interpretation in industrial environments. In the cloud-based IoT data ingestion pipelines, fuzzy data processing can be implemented in several locations, closer to the IoT events gateways, stream processors, or the persistence layer before the data is visualized. Since Automated Guided Vehicles (AGV)-enabled manufacturing can produce vast amounts of data, the decision on the placement of the fuzzy data processing can be important for secondary processes performed on the enriched data, like the predictive maintenance inferencing. In this paper, we analyze two locations of fuzzy data processing in the cloud-based environment built for monitoring AGVs in smart factories - by formulating fuzzy queries against data streams on stream processing units and data at rest in a database. The querying scenarios cover fuzzy filtering with simple and complex criteria, fuzzy filtering through assignment to a linguistic variable, and joining data streams by representing joining attributes as fuzzy numbers. The experimental results show that querying the data stream can be more efficient and profitable in the scalable environment of many AGVs. However, the enrichment provided for the data at rest is also beneficial when gathering data for building future predictive maintenance models.
Bozena Malysiak-Mrozek, Dominik Romanów, Piotr Grzesik, Pawel Benecki, Alexandre Niyomugaba, Theodore Habimana, Daniel Kostrzewa, Krzysztof Tokarz, Che-Lun Hung, Dariusz Mrozek
IEEE Big Data9
2024 Decoding the Granular Puzzle of Macromolecules: Efficient 3D Protein Structure Alignment in the Age of Big Data with Apache Spark
abstract
Proteins are complex biological information granules that play a crucial role in various cellular processes within living organisms. Processing 3D protein structures, which are the most informative from the biological point of view, is both intricate and time-consuming. In particular, performing 3D protein structure searches against large protein datasets involves identifying similarities and conducting structural alignments across numerous molecules (granules). This task demands advanced methods for matching identical and similar regions within protein structures and substantial computational resources to handle large collections of macromolecular data efficiently. In this paper, we present our parallel implementation of scalable 3D structural alignment on the Apache Spark big data platform. We describe a customized approach that leverages Spark data transformations within the data processing pipeline for the alignment process. Our experimental results demonstrate that this solution, tightly integrated with the Spark processing model, is both efficient and scalable, even with the increasing volume of protein structure data.
Bozena Malysiak-Mrozek, Paulina Pawlowicz, Vaidy S. Sunderam, Che-Lun Hung, Andrzej Kwiecien, Dariusz Mrozek
IEEE Big Data4