Diya Li

dblp:239/2197 · DBLP profile ↗
← Back
10ranked-venue papers
6as first author
6since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 5 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 3 · 3 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 first-authorComputer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2024 Estimating Agreement by Chance for Sequence Annotation
abstract
In the field of natural language processing, correction of performance assessment for chance agreement plays a crucial role in evaluating the reliability of annotations.However, there is a notable dearth of research focusing on chance correction for assessing the reliability of sequence annotation tasks, despite their widespread prevalence in the field.To address this gap, this paper introduces a novel model for generating random annotations, which serves as the foundation for estimating chance agreement in sequence annotation tasks.Utilizing the proposed randomization model and a related comparison approach, we successfully derive the analytical form of the distribution, enabling the computation of the probable location of each annotated text segment and subsequent chance agreement estimation.Through a combination simulation and corpus-based evaluation, we successfully assess its applicability and validate its accuracy and efficacy.
Diya Li, Carolyn P. Rosé, Ao Yuan, Chunxiao Zhou
ACL (1)1
2024 Causality-based Big Data Analysis of Information Cascade in Communication Networks
abstract
The phenomenon of information cascades in communication networks has been widely studied, with earlier research primarily using statistical models for analysis. Recently, machine learning methods combined with big data analysis have proven to be powerful tools for understanding network behavior. However, previous research efforts have largely overlooked the relationships among various influential factors. In this paper, we extend machine learning based data analysis methods by introducing the concept of causality, termed the Causality-based Big Data Analysis (CBDA) method, which incorporates factors such as fake agents, welfare effects, and trusted versus untrusted prior knowledge to examine their impact on the outcomes of erroneous information cascades. Numerical results demonstrate that each factor significantly influences the outcome, generally increasing the likelihood of its preferred result. This is achieved by clarifying how the weight assigned to each influential factor affects the behavior of subsequent agents. The results of our proposed CBDA method show that trusted prior knowledge significantly reduces the probability of erroneous information cascades, offering superior explainability and effectiveness compared to existing methods.
Yuming Han, Diya Li, Jianbo Du
GLOBECOM3
2024 Assessing Urban Safety: A Digital Twin Approach Using Streetview and Large Language Models
abstract
This study explores a novel approach to reevaluating urban safety using Vision Large Language Models (VLLMs) integrated with digital twin technology. Our methodology involves randomly selecting street views across various U.S. cities and employing VLLMs to detect and analyze street safety. We incorporate existing user-reported data through an API to validate our findings. Preliminary results indicate that this new approach significantly enhances the accuracy and reliability of urban safety assessments. The integration of VLLMs with digital twin frameworks presents a promising avenue for urban planners and policymakers to achieve more dynamic and real-time insights into city safety, ultimately contributing to smarter and safer urban environments. Our findings suggest that this method holds substantial potential for broader applications in digital twin initiatives, facilitating more informed decision-making processes.
Yuhan Cheng, Zhengcong Yin, Diya Li, Zhuoying Li
VTC Fall3
2024 A reinforcement learning-based routing algorithm for large street networks
abstract
Evacuation planning and emergency routing systems are crucial in saving lives during disasters. Traditional emergency routing systems, despite their best efforts, often struggle to accurately capture the dynamic nature of flood conditions, road closures, and other real-time changes inherent in urban disaster logistics. This paper introduces the ReinforceRouting model, a novel approach to optimizing evacuation routes using reinforcement learning (RL). The model incorporates a unique RL environment that considers multiple criteria, such as traffic conditions, hazardous situations, and the availability of safe routes. The RL agent in this model learns optimal actions through interaction with the environment, receiving feedback in the form of rewards or penalties. The ReinforceRouting model excels in executing prompt and accurate route planning on large road networks, outperforming traditional RL algorithms and shortest-path-based algorithms. A higher safety score and episode reward of the model are demonstrated when compared to these classical methods. This innovative approach to disaster evacuation planning offers a promising avenue for enhancing the efficiency, safety, and reliability of emergency responses in dynamic urban environments.
Diya Li, Zhe Zhang 0001, Bahareh Alizadeh Kharazi, Nick G. Duffield, Michelle A. Meyer, Courtney M. Thompson, Huilin Gao, Amir H. Behzadan
Int. J. Geogr. Inf. Sci.1
2023 Health-guided recipe recommendation over knowledge graphs
Diya Li, Mohammed J. Zaki, Ching-Hua Chen
J. Web Semant.1
2022 Human-centered flood mapping and intelligent routing through augmenting flood gauge data with crowdsourced street photos
abstract
The number and intensity of flood events have been on the rise in many regions of the world. In some parts of the U.S., for example, almost all residential properties, transportation networks, and major infrastructure (e.g., hospitals, airports, power stations) are at risk of failure caused by floods. The vulnerability to flooding, particularly in coastal areas and among marginalized populations is expected to increase as the climate continues to change, thus necessitating more effective flood management practices that consider various data modalities and innovative approaches to monitor and communicate flood risk. Research points to the importance of reliable information about the movement of floodwater as a critical decision-making parameter in flood evacuation and emergency response. Existing flood mapping systems, however, rely on sparsely installed flood gauges that lack sufficient spatial granularity for precise characterization of flood risk in populated urban areas. In this paper, we introduce a floodwater depth estimation methodology that augments flood gauge data with user-contributed photos of flooded streets to reliably estimate the depth of floodwater and provide ad-hoc, risk-informed route optimization. The performance of the developed technique is evaluated in Houston, Texas, that experienced urban floods during the 2017 Hurricane Harvey. A subset of 20 user-contributed flood photos in combination with gauge readings taken at the same time is used to create a flood inundation map of the experiment area. Results show that augmenting flood gauge data with crowdsourced photos of flooded streets leads to shorter travel time and distance while avoiding flood-inundated areas.
Bahareh Alizadeh Kharazi, Diya Li, Julia Hillin, Michelle A. Meyer, Courtney M. Thompson, Zhe Zhang 0001, Amir H. Behzadan
Adv. Eng. Informatics2
2020 RECIPTOR: An Effective Pretrained Model for Recipe Representation Learning
abstract
Recipe representation plays an important role in food computing for perception, recognition, recommendation and other applications. Learning pretrained recipe embeddings is a challenging task, as there is a lack of high quality annotated food datasets. In this paper, we provide a joint approach for learning effective pretrained recipe embeddings using both the ingredients and cooking instructions. We present RECIPTOR, a novel set transformer-based joint model to learn recipe representations, that preserves permutation-invariance for the ingredient set and uses a novel knowledge graph (KG) derived triplet sampling approach to optimize the learned embeddings so that related recipes are closer in the latent semantic space. The embeddings are further jointly optimized by combining similarity among cooking instructions with a KG based triplet loss. We experimentally show that RECIPTOR's recipe embeddings outperform state-of-the-art baselines on two newly designed downstream classification tasks by a wide margin.
Diya Li, Mohammed J. Zaki
KDD1
2020 Low-cost, bottom-up measures for evaluating search result diversification
Zhicheng Dou, Diya Li, Ji-Rong Wen, Tetsuya Sakai
Inf. Retr. J.3
2019 Global Vision-Based Reconstruction of Three-Dimensional Road Surfaces Using Adaptive Extended Kalman Filter
abstract
This paper presents a vision-based technique and a system developed for the global reconstruction of three-dimensional (3-D) road surfaces. Using the system, the technique globally reconstructs 3-D road surfaces by estimating the global camera pose using the Adaptive Extended Kalman Filter (AEKF) and integrating it with existing local road surface reconstruction techniques. The AEKF adaptively updates the covariance of uncertainties such that the estimation works well even in environments with varying uncertainties. Numerical results show the efficacy of the proposed technique over the Extended Kalman Filter (EKF)-based technique by 50% in accuracy, and the on-road test has demonstrated the ability of the proposed technique for the real-world global 3-D road surface reconstruction.
Diya Li, Tomonari Furukawa
ICRA1
2018 AEKF-Based 3-D Localization of Road Surface Images with Sparse Low-Accuracy GPS Data
abstract
This paper presents a technique for localizing road surface images acquired by a downward-facing monocular camera on a vehicle with sparse low-accuracy Global Positioning System (GPS) readings. Images are collected by reading vehicle speed through on-board diagnostics (OBD) such that distance between two neighboring images is constant. The images are then stitched to create the road surface of arbitrary length. Lastly, the three-dimensional (3-D) road surface is created and globally corrected by using the GPS and the elevation map as well as the Adaptive Extended Kalman Filter (AEKF). The advantage of this technique is the possible deployment of a sparse low-accuracy GPS due to the use of the adaptive version of Extended Kalman Filter (EKF). The proposed technique was used for localization of local roads and highways of 6.9 km total length in Blacksburg, VA. The results of the localization show the reconstructed 3-D differ from the satellite imagery data only by 7.97%.
Diya Li, Yazhe Hu, Tomonari Furukawa
VTC Fall1