VLDB 2026 Research / reviewers in the wild / expert
Yong Shi 0002
dblp:84/5467-2
· DBLP profile ↗
10ranked-venue papers in the field
4as first author
4since 2021 · last 2024
0000-0002-3980-1425ORCID · conflict
Domains — venue-derived; a paper can count in several
Big Data, Cloud & Distributed Data Systems · 5Database Systems & Data Management · 3 (3 first)Data Mining & Knowledge Discovery · 2 (1 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Enhancing Contextual Understanding in Knowledge Graphs: Integration of Quantum Natural Language Processing with Neo4j LLM Knowledge GraphabstractTraditional Knowledge Graphs (KGs), such as Neo4j, face challenges in managing high-dimensional relationships and capturing semantic nuances due to their deterministic nature. Quantum Natural Language Processing (QNLP) introduces probabilistic reasoning into the KG context. This integration leverages quantum principles, such as superposition, which allows relationships to exist in multiple states simultaneously, and entanglement, where the state of one entity dynamically influences the state of another. This quantum-based probabilistic reasoning provides a richer, more flexible representation of connections, moving beyond binary relationships to model the nuances and variability of real-world interactions. Our research demonstrates that QNLP enhances Neo4j’s ability to analyze context-rich data, improving tasks like entity extraction and knowledge inference. By modeling relationship states probabilistically, QNLP addresses limitations in traditional methods, providing nuanced insights and enabling more advanced, context-aware NLP applications. Suman Bharti, Dan Chia-Tien Lo, Yong Shi 0002 |
IEEE Big Data | 3 |
| 2024 | Practical Considerations of Fully Homomorphic Encryption in Privacy-Preserving Machine LearningabstractMachine learning has been successfully applied to big data analytics across various disciplines. However, as data is collected from diverse sectors, much of it is private and confidential. At the same time, one of the major challenges in machine learning is the slow training speed of large models, which often requires high-performance servers or cloud services. To protect data privacy while still allowing model training on such servers, privacy-preserving machine learning using Fully Homomorphic Encryption (FHE) has gained significant attention. However, its widespread adoption is hindered by performance degradation. This paper presents our experiments on training models over encrypted data using FHE. The results show that while FHE ensures privacy, it can significantly degrade performance, requiring complex tuning to optimize. Dan Chia-Tien Lo, Yong Shi 0002, Hossain Shahriar, Bobin Deng, Xinyue Zhang 0001, Mei-Lan Chen |
IEEE Big Data | 2 |
| 2023 | Deep Machine Learning on Segmenting and Classifying Crop Images Taken by Unmanned Aerial VehicleabstractIn the realm of precision agriculture, a crucial element involves the precise quantification or estimation of seedlings, fruits, and other agricultural produce on expansive multi-acre farms at various stages of cultivation. With the advent of unmanned aerial vehicles (UAVs), capturing images of watermelon fields has become a straightforward task. These images can be subsequently processed, segmented, and categorized to determine the total count of watermelons. Currently, conventional methods are employed to address this challenge, but they have their limitations. The field has benefited from the evolution of machine learning, which has the potential to streamline the process. Nevertheless, the training phase is intricate, and achieving a valuable model can be demanding. This research delves into an examination and presentation of the existing pre-trained models for image processing in this context. Dan Chia-Tien Lo, Bobin Deng, Yong Shi 0002 |
IEEE Big Data | 3 |
| 2023 | Design and Implementation of an ERC-20 Smart Contract on the Ethereum BlockchainabstractAs technology continues to develop, there is a growing need to find sustainable solutions in all industries, including cryptocurrency. Due to the high energy consumption that cryptocurrencies are known for, there have been efforts to reduce waste consumption and in turn minimize the carbon footprint. We support the trends for creating an environment-friendly crypto token using the ERC-20 standard on the Ethereum blockchain. We outline the various aspects that make a token more sustainable and highlight the potential benefits of such tokens. Our proposal involves the design of a smart contract that incorporates eco-friendly features such as lower energy consumption, carbon offsetting, and more efficient methods or algorithms. We also discuss the importance of transparency and accountability in the design and implementation of such tokens. This paper discusses not only the practical tools and steps necessary in creating a crypto token but also highlights the challenges associated with creating a more sustainable token. Joshua Priest, Cameron Cooper, Savvy Lovell, Yong Shi 0002, Dan Chia-Tien Lo |
IEEE Big Data | 4 |
| 2019 | Big Data Analysis on Social NetworkingabstractSocial networking media, such as Twitters, Facebook, and Chinese Weibo, has become a major means for people to deliberately express their ideas, thoughts, and views about everything. A huge amount of posts in these social media are issued and viewed by the general public daily. Evidently the social networking media can directly affect people's perceptions on a specific topic. Those data can be used to valuable information that will help organizations to understand what thetrends or sentiments are. Most of the research efforts in social networking data analytics are conducted in English-based social networking media. Research on Chinese social networking media receives relatively little attention. In this paper, we examine the key problems in this field, focus particularly on the characteristics of new vocabulary, emotion expressions, and hierarchical structures in Chinese “Weibo”. Associated theoretical and technological methods to address these problems are reviewed and discussed. Zhengwu Sun, Dan Chia-Tien Lo, Yong Shi 0002 |
IEEE BigData | 3 |
| 2011 | COID: A cluster-outlier iterative detection approach to multi-dimensional data analysis
Yong Shi 0002, Li Zhang 0008 |
Knowl. Inf. Syst. | 1 |
| 2005 | Towards Exploring Interactive Relationship between Clusters and Outliers in Multi-Dimensional Data AnalysisabstractNowadays many data mining algorithms focus on clustering methods. There are also a lot of approaches designed for outlier detection. We observe that, in many situations, clusters and outliers are concepts whose meanings are inseparable to each other, especially for those data sets with noise. Thus, it is necessary to treat clusters and outliers as concepts of the same importance in data analysis. In this paper, we present a cluster-outlier iterative detection algorithm, tending to detect the clusters and outliers in another perspective for noisy data sets. In this algorithm, clusters are detected and adjusted according to the intra-relationship within clusters and the inter-relationship between clusters and outliers, and vice versa. The adjustment and modification of the clusters and outliers are performed iteratively until a certain termination condition is reached. This data processing algorithm can be applied in many fields such as pattern recognition, data clustering and signal processing. Experimental results demonstrate the advantages of our approach. Yong Shi 0002, Aidong Zhang 0001 |
ICDE | 1 |
| 2004 | A Shrinking-Based Dimension Reduction Approach for Multi-Dimensional Data Analysis
Yong Shi 0002, Aidong Zhang 0001 |
SSDBM | 1 |
| 2003 | A Shrinking-Based Approach for Multi-Dimensional Data Analysis
Yong Shi 0002, Yuqing Song 0002, Aidong Zhang 0001 |
VLDB | 1 |
| 2002 | VizCluster: An Interactive Visualization Approach to Cluster Analysis and Its Application on Microarray DataabstractVisualization enables us to find structures, features, patterns and relationship in a dataset by presenting the data in various graphical forms with possible interactions. Recent development of DNA microarray technology can be used to measure the expression levels of thousands of genes simultaneously. It has already had a significant impact on the field of bioinformatics, requiring innovative techniques to efficiently and effectively extract, analysis and visualize these fast growing data. In this paper, we present VizCluster, an interactive visualization approach to cluster analysis, and its application on microarray data. VizCluster combines the merits of both high dimensional scatter-plot and parallel coordinates. Integrated with useful features, it can give a simple, fast, intuitive and yet powerful view of the data set. VizCluster supports three major analyzing modes: cluster/class discovery, class prediction, and class assessment. Its primary applications are the classification of samples on microarray datasets. The experiments are based on gene expression data from a study of multiple sclerosis and leukemia patients. Li Zhang 0008, Chun Tang, Yong Shi 0002, Yuqing Song 0002, Aidong Zhang 0001, Murali Ramanathan |
SDM | 3 |