VLDB 2026 Research / reviewers in the wild / expert
Hong Min
dblp:47/1976
· DBLP profile ↗
26ranked-venue papers
6as first author
9since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 4 since 2021Computer networks · 4 · 2 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A joint optimization of resource allocation management and multi-task offloading in high-mobility vehicular multi-access edge computing networks
Hong Min, Amir Masoud Rahmani, Payam Ghaderkourehpaz, Komeil Moghaddasi, Mehdi Hosseinzadeh 0001 |
Ad Hoc Networks | 1 |
| 2024 | Serving Deep Learning Models from Relational Databases
Lixi Zhou, Kanchan Chowdhury, Saif Masood, Alexandre E. Eichenberger, Hong Min, Alex Sim, Kesheng Wu, Binhang Yuan, Jia Zou 0001 |
EDBT | 6 |
| 2024 | Advancing medical data classification through federated learning and blockchain incentive mechanism: implications for modern software systems and applicationsabstractAbstract The key issue of medical data is patient information sensitivity and dataset finiteness, which need to guarantee high-efficient training. Besides, the current convolutional neural network has a low image classification and poor robustness concerning antagonistic samples. A lack of scalability in healthcare federated learning and incentive mechanism hinders the attraction of ample high-quality datasets. This paper proposes a Federated Learning Incentive Mechanism for Medical Data Classification (FedIn-MC). It realizes a collaborative model training of multi-party medical institutions through the combination of federated learning and blockchain. There is a marked improvement to the model’s robustness through a combination of the distance loss function and the prototype loss regulation. In addition, this incentive mechanism of blockchain in the project is applied to calculate client contribution values and encourage healthcare institutions to active training model participation. Simulation results verify an accomplishment of a multi-party training. With regard to image classifications, this framework also has a higher classification accuracy and stronger robustness concerning invisible class samples. Hong Min, Xin Su 0002 |
J. Supercomput. | 3 |
| 2023 | A Comparison of End-to-End Decision Forest Inference PipelinesabstractDecision forest, including RandomForest, XGBoost, and LightGBM, dominates the machine learning tasks over tabular data. Recently, several frameworks were developed for decision forest inference, such as ONNX, TreeLite from Amazon, TensorFlow Decision Forest from Google, HummingBird from Microsoft, Nvidia FIL, and lleaves. While these frameworks are fully optimized for inference computations, they are all decoupled with databases and general data management frameworks, which leads to cross-system performance overheads. We first provided a DICT model to understand the performance gaps between decoupled and in-database inference. We further identified that for in-database inference, in addition to the popular UDF-centric representation that encapsulates the ML into one User Defined Function (UDF), there also exists a relation-centric representation that breaks down the decision forest inference into several fine-grained SQL operations. The relation-centric representation can achieve significantly better performance for large models. We optimized both implementations and conducted a comprehensive benchmark to compare these two implementations to the aforementioned decoupled inference pipelines and existing in-database inference pipelines such as Spark-SQL and PostgresML. The evaluation results validated the DICT model and demonstrated the superior performance of our in-database inference design compared to the baselines. Saif Masood, Mahidhar Reddy Dwarampudi, Venkatesh Gunda, Hong Min, Lei Yu 0002, Soham Nag, Jia Zou 0001 |
SoCC | 5 |
| 2023 | Privacy-Preserving Redaction of Diagnosis Data through Source Code AnalysisabstractProtecting sensitive information in diagnostic data such as logs, is a critical concern in the industrial software diagnosis and debugging process. While there are many tools developed to automatically redact the logs for identifying and removing sensitive information, they have severe limitations which can cause either over redaction and loss of critical diagnostic information (false positives), or disclosure of sensitive information (false negatives), or both. To address the problem, in this paper, we argue for a source code analysis approach for log redaction. To identify a log message containing sensitive information, our method locates the corresponding log statement in the source code with logger code augmentation, and checks if the log statement outputs data from sensitive sources by using the data flow graph built from the source code. Appropriate redaction rules are further applied depending on the sensitiveness of the data sources to preserve the privacy information in the logs. We conducted experimental evaluation and comparison with other popular baselines. The results demonstrate that our approach can significantly improve the detection precision of the sensitive information and reduce both false positives and negatives. Lixi Zhou, Lei Yu 0002, Jia Zou 0001, Hong Min |
SSDBM | 4 |
| 2022 | Serving Deep Learning Models with Deduplication from Relational DatabasesabstractServing deep learning models from relational databases brings significant benefits. First, features extracted from databases do not need to be transferred to any decoupled deep learning systems for inferences, and thus the system management overhead can be significantly reduced. Second, in a relational database, data management along the storage hierarchy is fully integrated with query processing, and thus it can continue model serving even if the working set size exceeds the available memory. Applying model deduplication can greatly reduce the storage space, memory footprint, cache misses, and inference latency. However, existing data deduplication techniques are not applicable to the deep learning model serving applications in relational databases. They do not consider the impacts on model inference accuracy as well as the inconsistency between tensor blocks and database pages. This work proposed synergistic storage optimization techniques for duplication detection, page packing, and caching, to enhance database systems for model serving. Evaluation results show that our proposed techniques significantly improved the storage efficiency and the model inference latency, and outperformed existing deep learning frameworks in targeting scenarios. Lixi Zhou, Amitabh Das, Hong Min, Lei Yu 0002, Jia Zou 0001 |
Proc. VLDB Endow. | 4 |
| 2021 | Knowledge & Learning-based Adaptable System for Sensitive Information Identification and HandlingabstractDiagnostic data such as logs and memory dumps from production systems are often shared with development teams to do root cause analysis of system crashes. Invariably such diagnostic data contains sensitive information and sharing it can lead to data leaks. To handle this problem we present Knowledge and Learning-based Adaptable System for Sensitive InFormation Identification and Handling (KLASSIFI) which is an end to end system capable of identifying and redacting sensitive information present in diagnostic data. KLASSIFI is highly customizable, allowing it to be used for various different business use cases by simply changing the configuration. KLASSIFI ensures that the output file is useful by retaining the metadata which is used by various debugging tools. Various optimizations have been done to improve the performance of KLASSIFI. Empirical evaluation of KLASSIFI shows that it is able to process large files (128 GB) in 84 minutes and its performance scales linearly with varying factors. This points to practicability of KLASSIFI. Akshar Kaul, Manish Kesarwani, Hong Min, Qi Zhang 0009 |
CLOUD | 3 |
| 2021 | Automated Data Science for Relational DataabstractFeature engineering is a crucial but tedious task that requires up to 80% of the total time in data science projects. A significant challenge is when data consists of tables from different data sources, thus data scientists need to wisely aggregate and join tables while performing feature engineering task. In this work, we demonstrate a novel system called OneBM (One Button Machine), that enables data scientists to increase their efficiency with automated feature engineering for relational data. OneBM takes as input a relational dataset with multiple tables and its entity relation diagram (ERD) which can be declared with a novel, easy-to-use drag-and-drop graphical user interface. The system then automatically identifies and executes relevant joins and aggregates in the data, and generates new features with a rich set of transformations for various types of data including but not limited to time-series, sequences, number sets and itemsets, etc. The generated features then can be used by automated model selection and hyper-parameter optimization algorithms to complete a fully end-to-end automated data science (or AutoDS) workflow. A follow-up user evaluation illustrated how data scientists can perform multi-table feature engineering tasks in minutes using our system, compared to repeatedly coding SQL-like queries to transform and aggregate relational data requiring weeks of manual labor for comparable performance. In the live demos we plan to show two use cases with real-world datasets (video demos are available at the links in the footnote): sale prediction1and call center user experience2. Pre-registered partcipants can play with these use-cases and the given datasets via Watson Studio on the cloud. Hoang Thanh Lam, Beat Buesser, Hong Min, Tran Ngoc Minh, Martin Wistuba, Udayan Khurana, Gregory Bramble, Theodoros Salonidis, Dakuo Wang, Horst Samulowitz |
ICDE | 3 |
| 2021 | Key node selection based on a genetic algorithm for fast patching in social networksabstractSummary Online social network users provide considerable amounts of personal information and share this information with friends without space‐time limitations. The tight connectivity among users of social networks causes the rapid spreading of information. Given the popularity of social networking sites, there is a high probability of attacks. Worms target popular users with interesting information to infect them, as their higher reputations have more power in social networks. Therefore, timely patch propagation schemes must be able to inhibit the activity of worms. To improve the patch propagation speed, it is important to select key nodes that are the starting points of the patch process. In this paper, we proposed a key node selection scheme based on a genetic algorithm to find the most significant contribution nodes of patch propagation. We modeled the usage patterns of an online social network user and simulated the proposed scheme with data from this user. Simulation results show that the proposed scheme propagates patches more rapidly than existing schemes. Bongjae Kim, Jinman Jung, Junyoung Heo, Hong Min |
Concurr. Comput. Pract. Exp. | 4 |
| 2020 | Optimal job partitioning and allocation for vehicular cloud computing
Taesik Kim, Hong Min, Eunsoo Choi, Jinman Jung |
Future Gener. Comput. Syst. | 2 |
| 2018 | Vehicular datacenter modeling for cloud computing: Considering capacity and leave rate of vehicles
Taesik Kim, Hong Min, Jinman Jung |
Future Gener. Comput. Syst. | 2 |
| 2018 | Pattern Matching Based Sensor Identification Layer for an Android PlatformabstractAs sensor‐related technologies have been developed, smartphones obtain more information from internal and external sensors. This interaction accelerates the development of applications in the Internet of Things environment. Due to many attributes that may vary the quality of the IoT system, sensor manufacturers provide their own data format and application even if there is a well‐defined standard, such as ISO/IEEE 11073 for personal health devices. In this paper, we propose a client‐server‐based sensor adaptation layer for an Android platform to improve interoperability among nonstandard sensors. Interoperability is an important quality aspect for the IoT that may have a strong impact on the system especially when the sensors are coming from different sources. Here, the server compares profiles that have clues to identify the sensor device with a data packet stream based on a modified Boyer‐Moore‐Horspool algorithm. Our matching model considers features of the sensor data packet. To verify the operability, we have implemented a prototype of this proposed system. The evaluation results show that the start and end pattern of the data packet are more efficient when the length of the data packet is longer. Hong Min, Taesik Kim, Junyoung Heo, Tomás Cerný, Sriram Sankaran, Bestoun S. Ahmed, Jinman Jung |
Wirel. Commun. Mob. Comput. | 1 |
| 2015 | Octopus: Hybrid Big Data Integration EngineabstractNowadays large enterprises maintain a huge amount of data in multiple backend systems including traditional database systems and recently popular big data systems. In an example of telecom providers, the key business data (e.g., billing information) is maintained in database systems whereas the huge signaling log data is on HDFS with Hive. How to integrate such data and provide a consolidate query and analytic becomes a challenging task. Neither traditional database warehouse nor recent Big Data system (e.g. Apache Spark and Hadoop) can fully leverage the power of each backend system. In this paper, we build a hybrid data processing engine, called Octopus, to fully integrate backend systems. Given the backend systems, data is distributed at multiple locations. Octopus focuses on the optimization of the amount of data movement. To this end, Octopus proposes a technique of query pushdown for such optimization. A proof-of-concept prototype of Octopus successfully verifies that Octopus can achieve much faster running time than Spark. For example, Octopus outperforms the recent Spark version 1.4.0 by 5.25 X faster running time to process an aggregation query. Weixiong Rao, Hong Min, Gong Su |
CloudCom | 4 |
| 2014 | Inter-Data-Center Large-Scale Database Replication Optimization - A Workload Driven Partitioning Approach
Hong Min, Jie Huang 0032, Serge Bourbonnais, Miao Zheng, Gene Fuh |
DEXA (2) | 1 |
| 2013 | Accelerating Join Operation for Relational Databases with FPGAsabstractIn this paper, we investigate the use of field programmable gate arrays (FPGAs) to accelerate relational joins. Relational join is one of the most CPU-intensive, yet commonly used, database operations. Hashing can be used to reduce the time complexity from quadratic (naïve) to linear time. However, doing so can introduce false positives to the results which must be resolved. We present a hash-join engine on FPGA that performs hashing, conflict resolution, and joining on a PCIe-attached system, achieving greater than 11x speedup over software. Robert J. Halstead, Bharat Sukhwani, Hong Min, Mathew Thoennes, Parijat Dube, Sameh W. Asaad, Balakrishna Iyer |
FCCM | 3 |
| 2013 | Large Payload Streaming Database Sort and Projection on FPGAsabstractIn recent years, real-time analytics has seen widespread adoption in the business world. While it provides useful business insights and improved market responsiveness, it also adds a computational burden to traditional online transaction processing (OLTP) systems. Analytics queries involve complex database operations such as sort, aggregation, and join that consume significant computational resources, and, when executed on the same system, may affect the performance of OLTP queries. In this paper, we try to address this issue by accelerating two such database operations, namely, projection and sort, using a field programmable gate array (FPGA). Our prototype is implemented on an Alter a Stratix V FPGA and achieves an order of magnitude speedup in the sort operation compared to baseline software. Furthermore, our prototype implements projection in parallel with other query operations on FPGA, thus completely eliminating the cost of projection without consuming any extra cycles on the FPGA. FPGA accelerated sort and projection have been integrated with our previous work on accelerating other query operations [1], making our analytics acceleration prototype on FPGA applicable to a wider variety of queries. Bharat Sukhwani, Mathew Thoennes, Hong Min, Parijat Dube, Bernard Brezzo, Sameh W. Asaad, Donna Dillenberger |
SBAC-PAD | 3 |
| 2012 | Database analytics acceleration using FPGAsabstractBusiness growth and technology advancements have resulted in growing amounts of enterprise data. To gain valuable business insight and competitive advantage, businesses demand the capability of performing real-time analytics on such data. This, however, involves expensive query operations that are very time consuming on traditional CPUs. Additionally, in traditional database management systems (DBMS), the CPU resources are dedicated to mission-critical transactional workloads. Offloading expensive analytics query operations to a co-processor can allow efficient execution of analytics workloads in parallel with transactional workloads. Bharat Sukhwani, Hong Min, Mathew Thoennes, Parijat Dube, Balakrishna Iyer, Bernard Brezzo, Donna Dillenberger, Sameh W. Asaad |
PACT | 2 |
| 2010 | Improving In-memory Column-Store Database Predicate Evaluation Performance on Multi-core SystemsabstractThe ability to analyze a large volume of data for the purpose of business intelligence has led to various innovations in database technology. One example is the increased interest of using column-oriented data layout to address query performance in analytical and warehousing workloads. As system architectures move towards multi-core designs, it is important to address optimizing performance for these workloads on these platforms. In this paper we present SPHINX, an architecture that utilizes multi-core systems for search-based predicate evaluation operations in analytical query workloads against in-memory column store. We discuss the natural parallelism of predicate evaluations and various bottlenecks that impact search performance. We present several performance improvement techniques and apply a scan sharing technique based on cache reuse efficiency to further improve the performance. We demonstrate the performance benefits of our scan sharing scheduler over other scheduling approaches in a workload of mixed search queries. Hong Min, Hubertus Franke |
SBAC-PAD | 1 |
| 2008 | A Module Management Scheme for Dynamic Reconfiguration
Hong Min, Junyoung Heo, Yookun Cho, Kahyun Lee, Jaegi Son, Byunghun Song |
ICCSA (1) | 1 |
| 2008 | SensorMaker: A Wireless Sensor Network Simulator for Scalable and Fine-Grained Instrumentation
Sangho Yi, Hong Min, Yookun Cho, Jiman Hong |
ICCSA (1) | 2 |
| 2008 | Energy-Efficient Data Aggregation Protocol for Location-Aware Wireless Sensor NetworksabstractThe energy is the most critical resource in wireless sensor networks. Many energy efficient techniques have been studied especially in communication protocol. Among these protocols, PEGASIS proposed by Lindsey is the most superior interms of energy efficiency. In PEGASIS, all sensor nodes aggregate and transmit data along the single chain. However, this single chain causes some problems such as delay, unexpected long transmission and non-directional transmission to the sink. To resolve delay problem, PEGASIS cuts the chain into several chains. In this way, the delay can be decreased but there is new problem, wireless interference. In this paper, we propose a new chain construction algorithm to resolve above problems. The proposed algorithm minimizes the delay without the wireless interferences and maximizes the energy efficiency by removing the non-directional transmission when transmitting packets to the sink node on wireless sensor networks. Our simulation results show that our proposed algorithm can minimize the delay without expense of energy consumption. Hong Min, Sangho Yi, Junyoung Heo, Yookun Cho, Jiman Hong |
ISPA | 1 |
| 2008 | Molecule: An adaptive dynamic reconfiguration scheme for sensor operating systems
Sangho Yi, Hong Min, Yookun Cho, Jiman Hong |
Comput. Commun. | 2 |
| 2008 | Adaptive Multilevel Code Update Protocol for Real-Time Sensor Operating SystemsabstractIn wireless sensor networks each sensor node has very limited resources, and it is very difficult to find and collect them. For this reason, updating or adding programs in sensor nodes must be performed via a communication channel at run-time. Many code update protocols have been developed for sensor networks, ranging from function-level update to full-image replacement. However, they provide only a fixed level of code update protocols. These protocols require manual selection of an appropriate protocol because they do not consider a cost analysis of the update protocols. In addition, they do not consider real-time response while updating codes. In this paper, we present an adaptive multilevel code update protocol (AMCUP) for real-time sensor operating systems. AMCUP enables energy-efficient code update via support for multilevel protocols (i.e., full-image, module-level, function-level, and instruction-level). It adaptively selects a protocol which meets deadline of applications and consumes less energy based on a cost analysis of several protocols. Our simulation and experimental results show that AMCUP can reduce energy consumption and execution time compared with existing single-level code update protocols while meeting deadline of the running applications. Sangho Yi, Hong Min, Yookun Cho, Jiman Hong |
IEEE Trans. Ind. Informatics | 2 |
| 2006 | Adaptive Load Balancing Mechanism for Server Cluster
Geunyoung Park, Boncheol Gu, Junyoung Heo, Sangho Yi, Jungkyu Han, Hong Min, Xuefeng Piao, Yookun Cho, Chang-Won Park, Ha Joong Chung, Bongkyu Lee |
ICCSA (4) | 7 |
| 2006 | Performance Analysis of Task Schedulers in Operating Systems for Wireless Sensor Networks
Sangho Yi, Hong Min, Junyoung Heo, Boncheol Gu, Yookun Cho, Jiman Hong, Jin Won Kim, Kwangyong Lee, Seung-Min Park 0001 |
ICCSA (4) | 2 |
| 2006 | XMAS: An eXtraordinary Memory Allocation Scheme for Resource-Constrained Sensor Operating Systems
Sangho Yi, Hong Min, Junyoung Heo, Boncheol Gu, Yookun Cho, Jiman Hong, Hyukjun Oh, Byunghun Song |
MSN | 2 |