EDBT 2026 Demo / reviewers in the wild / expert
Yufeng Ge
dblp:186/9845
· DBLP profile ↗
5ranked-venue papers in the field
0as first author
3since 2021 · last 2024
—ORCID · conflict
Domains — venue-derived; a paper can count in several
Big Data, Cloud & Distributed Data Systems · 5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Building Multi-Agent Copilot towards Autonomous Agricultural Data Management and AnalysisabstractThe ubiquity of sensors and IoT devices has led to an explosion in data availability in modern agriculture. The large volume and heterogeneity of the data, together with the complexity of data processing requirements, pose huge obstacles for achieving the principles of Findable, Accessible, Interoperable, and Reusable (FAIR). Current data management and analysis paradigms are to a large extent traditional, in which data collecting, curating, integration, loading, storing, sharing and analyzing still involve too much human effort and know-how. The experts, researchers and the farm operators need to understand the data and the whole process of data management pipeline to make full use of the data. The essential problem of the traditional paradigm is the lack of a layer of orchestrational intelligence which can understand, organize and coordinate the data processing utilities to maximize data management and analysis outcome. The emerging reasoning and tool mastering abilities of large language models (LLM) make it a potentially good fit to this position, which helps a shift from the traditional user-driven paradigm to AI-driven paradigm. In this paper, we propose and explore the idea of a LLM based copilot for autonomous agricultural data management and analysis. Based on our previously developed platform of Agricultural Data Management and Analytics (ADMA), we build a proof-of-concept multi-agent system called ADMA Copilot, which can understand user’s intent, makes plans for data processing pipeline and accomplishes tasks automatically, in which three agents: a LLM based controller, an input formatter and an output formatter collaborate together. Different from existing LLM based solutions, by defining a meta-program graph, our work decouples control flow and data flow to enhance the predictability of the behavior of the agents. Experiments demonstrate the intelligence, autonomy, efficacy, efficiency, extensibility, flexibility and privacy of our system. Comparison is also made between ours and existing systems to show the superiority and potential of our system. Yu Pan 0007, Jianxin Sun 0001, Hongfeng Yu 0001, Joe Luck, Geng Bai, Nipuna Chamara, Yufeng Ge, Tala Awada |
IEEE Big Data | 7 |
| 2023 | Transforming Agriculture with Intelligent Data Management and InsightsabstractModern agriculture faces grand challenges to meet increased demands for food, fuel, feed, and fiber with population growth under the constraints of climate change and dwindling natural resources. Data innovation is urgently required to secure and improve the productivity, sustainability, and resilience of our agroecosystems. As various sensors and Internet of Things (IoT) instrumentation become more available, affordable, reliable, and stable, it has become possible to conduct data collection, integration, and analysis at multiple temporal and spatial scales, in real-time, and with high resolutions. At the same time, the sheer amount of data poses a great challenge to data storage and analysis, and the de facto data management and analysis practices adopted by scientists have become increasingly inefficient. Additionally, the data generated from different disciplines, such as genomics, phenomics, environment, agronomy, and socioeconomic, can be highly heterogeneous. That is, datasets across disciplines often do not share the same ontology, modality, or format. All of the above make it necessary to design a new data management infrastructure that implements the principles of Findable, Accessible, Interoperable, and Reusable (FAIR). In this paper, we propose Agriculture Data Management and Analytics (ADMA), which satisfies the FAIR principles. Our new data management infrastructure is intelligent by supporting semantic data management across disciplines, interactive by providing various data management/analysis portals such as web GUI, command line, and API, scalable by utilizing the power of high-performance computing (HPC), extensible by allowing users to load their own data analysis tools, trackable by keeping track of different operations on each file, and open by using a rich set of mature open source technologies. Yu Pan 0007, Jianxin Sun 0001, Hongfeng Yu 0001, Geng Bai, Yufeng Ge, Joe Luck, Tala Awada |
IEEE Big Data | 5 |
| 2023 | Visualization of 3D Hyperspectral Soil Mapping Data via Autoencoder-based ClusteringabstractSoil measurement and evaluation are crucial to various aspects of agriculture, including agricultural productivity, nutrient management, water management, and pH Regulation. Hyperspectral imaging is an advanced technique used to capture and analyze a wide range of light wavelengths (or spectral bands) across the electromagnetic spectrum. Hyperspectral imaging in soil research involves the use of this advanced imaging technique to analyze the spectral properties of soils. It allows researchers to capture detailed information about the composition, texture, and conditions of soil across a wide range of wavelengths in the electromagnetic spectrum. This in-depth spectral analysis provides valuable insights for studying soil health, nutrient content, moisture levels, and other critical parameters. However, existing hyperspectral analysis of soil relies on using imaging systems to exclusively capture information from the soil surface. This yields a two-dimensional image in which each pixel represents a spectrum vector. In this paper, we provide a new 3D hyperspectral data capturing features deep into the soil where each voxel represents a spectrum vector. For effective analysis of this type of new hyperspectral data, we develop a 3D visualization tool to not only directly visualize individual spectrum of the soil volume but also provide a way to cluster such high dimensional data leveraging a deep learning-based method through autoencoder. Jianxin Sun 0001, Xinyan Xie, Yu Pan 0007, Yakub Islamov, Yufeng Ge, Hongfeng Yu 0001 |
IEEE Big Data | 5 |
| 2019 | Spatial-Temporal Scientific Data Clustering via Deep Convolutional Neural NetworkabstractWe explore the usage of deep convolutional neural network for clustering the time steps of a spatial-temporal scientific dataset. Our approach first takes the scientific dataset as training data and trains a deep convolutional autoencoder. A low-dimensional feature space or latent space can be extracted by inferencing the encoding part of the network. As a result, each time step is transformed into a feature descriptor that can be compared with each other in the feature space. In this way, we can cluster time steps according to their feature descriptors, and each group of time steps has a similar characterization. We demonstrate the effectiveness of our approach using a real-world simulation dataset of water contamination. Multiple variables and their combinations of this dataset are fed into our approach. The trained network enables the clustering of the time steps and facilitates scientists to examine their large spatial-temporal datasets. Jianxin Sun 0001, Chunxia Wu, Yufeng Ge, Yusong Li, Hongfeng Yu 0001 |
IEEE BigData | 3 |
| 2018 | 3D Reconstruction of Plant Leaves for High-Throughput PhenotypingabstractGenerating 3D digital representations of plants is indispensable for researchers to gain a detailed understanding of plant dynamics. Emerging high-throughput plant phenotyping techniques can capture plant point clouds that, however, often contain imperfections and make it a changeling task to generate accurate 3D reconstructions. We present an end-to-end pipeline to reconstruct surfaces from point clouds of maize and rice plants. In particular, we propose a two-step clustering approach to accurately segment the points of each individual plant component according to maize and rice properties. We further employ surface fitting and edge fitting to ensure the smoothness of resulting surfaces. Realistic visualization results are obtained through post-processing, including texturing and lighting. Our experimental study has explored the parameter space and demonstrated the effectiveness of our pipeline for high-throughput plant phenotyping. Feiyu Zhu 0001, Suresh Thapa, Tiao Gao, Yufeng Ge, Harkamal Walia, Hongfeng Yu 0001 |
IEEE BigData | 4 |