Yihe Zhang 0001

dblp:237/7632-1 · DBLP profile ↗
← Back
12ranked-venue papers
8as first author
7since 2021 · last 2025
0009-0009-7739-3870ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 7 · 6 first-author · 3 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 3 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Regional Weather Variable Predictions by Machine Learning With Near-Surface Observational and Atmospheric Numerical Data
abstract
Accurate and timely regional weather prediction is vital for sectors dependent on weather-related decisions. Traditional prediction methods, based on atmospheric equations, often struggle with coarse temporal resolutions and inaccuracies. This article presents a novel machine learning (ML) model, called Micro-Macro (MiMa), that integrates both near-surface observational data from Kentucky Mesonet stations (collected every 5 min, known as Micro data) and hourly atmospheric numerical outputs (termed as Macro data) for fine-resolution weather forecasting. The MiMa model employs an encoder-decoder transformer structure, with two encoders for processing multivariate data from both datasets and a decoder for forecasting weather variables over short time horizons. Each instance of the MiMa model, called a modelet, predicts the values of a specific weather parameter at an individual mesonet station. The approach is extended with Regional MiMa (Re-MiMa) modelets, which are designed to predict weather variables at ungauged locations by training on multivariate data from a few representative stations in a region, tagged with their elevations. Re-MiMa can provide highly accurate predictions across an entire region, even in areas without observational stations. Experimental results show that MiMa significantly outperforms current models, with Re-MiMa offering precise short-term forecasts for ungauged locations, marking a significant advancement in weather forecasting accuracy and applicability.
Yihe Zhang 0001, Bryce Turney, Purushottam Sigdel, Xu Yuan 0001, Eric Rappin, Adrian Lago, Sytske K. Kimball, Li Chen 0019, Paul J. Darby, Lu Peng 0001, Sercan Aygün, Yazhou Tu, M. Hassan Najafi, Nian-Feng Tzeng
IEEE Trans. Geosci. Remote. Sens.1
2024 An Open and Large-Scale Dataset for Multi-Modal Climate Change-aware Crop Yield Predictions
abstract
Precise crop yield predictions are of national importance for ensuring food security and sustainable agricultural practices. While AI-for-science approaches have exhibited promising achievements in solving many scientific problems such as drug discovery, precipitation nowcasting, etc., the development of deep learning models for predicting crop yields is constantly hindered by the lack of an open and large-scale deep learning-ready dataset with multiple modalities to accommodate sufficient information. To remedy this, we introduce the CropNet dataset, the first terabyte-sized, publicly available, and multi-modal dataset specifically targeting climate change-aware crop yield predictions for the contiguous United States (U.S.) continent at the county level. Our CropNet dataset is composed of three modalities of data, i.e., Sentinel-2 Imagery, WRF-HRRR Computed Dataset, and USDA Crop Dataset, for over 2200 U.S. counties spanning 6 years (2017-2022), expected to facilitate researchers in developing versatile deep learning models for timely and precisely predicting crop yields at the county-level, by accounting for the effects of both short-term growing season weather variations and long-term climate change on crop yields. Besides, we develop the CropNet package, offering three types of APIs, for facilitating researchers in downloading the CropNet data on the fly over the time and region of interest, and flexibly building their deep learning models for accurate crop yield predictions. Extensive experiments have been conducted on our CropNet dataset via employing various types of deep learning solutions, with the results validating the general applicability and the efficacy of the CropNet dataset in climate change-aware crop yield predictions. We have officially released our CropNet dataset on Hugging Face Datasets https://huggingface.co/datasets/CropNet/CropNet and our CropNet package on the Python Package Index (PyPI) https://pypi.org/project/cropnet. Code and tutorials are available at https://github.com/fudong03/CropNet.
Fudong Lin, Kaleb Guillot, Summer Crawford, Yihe Zhang 0001, Xu Yuan 0001, Nian-Feng Tzeng
KDD4
2023 Devils in Your Apps: Vulnerabilities and User Privacy Exposure in Mobile Notification Systems
abstract
Witnessing the blooming adoption of push notifications on mobile devices, this new message delivery paradigm has become pervasive in diverse applications. Accompanying with its broad adoption, the potential security risks and privacy exposure issues raise public concerns regarding its great social impacts. This paper conducts the first attempt to exploit the mobile notification ecosystem. By dissecting its structural elements and implementation process, a comprehensive vulnerability analysis is conducted towards the complete flow of mobile notification from platform enrollment to messaging. Meanwhile, for privacy exposure, we first examine the implementation of privacy policy compliance by proposing a three-level inspection approach to guide our analysis. Then, our top-down methods from documentation analysis, application network traffic study, to static analysis expose the illicit data collection behaviors in released applications. In addition, we uncover the potential privacy inference resulted from the notification monitoring. To support our analysis, we conduct empirical studies on 12 most popular notification platforms and perform static analysis over 30,000+ applications. We discover: 1) six platforms either provide ambiguous KEY naming rules or offer vulnerable messaging APIs; 2) privacy policy compliance implementations are either stagnated at the documentation stages (8 of 12 platforms) or never implemented in apps, resulting in billions of users suffering from privacy exposure; and 3) some apps can stealthily monitor notification messages delivering to other apps, potentially incurring user privacy inference risks. Our study raises the urgent demand for better regulations of mobile notification deployment.
Jiadong Lou, Yihe Zhang 0001, Xinghua Li 0001, Xu Yuan 0001, Ning Zhang 0017
DSN3
2023 MMST-ViT: Climate Change-aware Crop Yield Prediction via Multi-Modal Spatial-Temporal Vision Transformer
abstract
Precise crop yield prediction provides valuable information for agricultural planning and decision-making processes. However, timely predicting crop yields remains challenging as crop growth is sensitive to growing season weather variation and climate change. In this work, we develop a deep learning-based solution, namely Multi-Modal Spatial-Temporal Vision Transformer (MMST-ViT), for predicting crop yields at the county level across the United States, by considering the effects of short-term meteorological variations during the growing season and the long-term climate change on crops. Specifically, our MMST-ViT consists of a Multi-Modal Transformer, a Spatial Transformer, and a Temporal Transformer. The Multi-Modal Transformer leverages both visual remote sensing data and short-term meteorological data for modeling the effect of growing season weather variations on crop growth. The Spatial Transformer learns the high-resolution spatial dependency among counties for accurate agricultural tracking. The Temporal Transformer captures the long-range temporal dependency for learning the impact of long-term climate change on crops. Meanwhile, we also devise a novel multi-modal contrastive learning technique to pre-train our model without extensive human supervision. Hence, our MMST-ViT captures the impacts of both short-term weather variations and long-term climate change on crops by leveraging both satellite images and meteorological data. We have conducted extensive experiments on over 200 counties in the United States, with the experimental results exhibiting that our MMST-ViT outperforms its counterparts under three performance metrics of interest. Our dataset and code are available at https://github.com/fudong03/MMST-ViT.
Fudong Lin, Summer Crawford, Kaleb Guillot, Yihe Zhang 0001, Xu Yuan 0001, Li Chen 0019, Shelby Williams, Robert Minvielle, Xiangming Xiao, Drew Gholson, Nicolas Ashwell, Tri Setiyono, Brenda Tubana, Lu Peng 0001, Magdy A. Bayoumi, Nian-Feng Tzeng
ICCV4
2021 Platform-Oblivious Anti-Spam Gateway
abstract
This paper addresses a novel anti-spam gateway targeting multiple linguistic-based social platforms to expose the outlier property of their spam messages uniformly for effective detection. Instead of labeling ground truth datasets and extracting key features, which are labor-intensive and time-consuming, we start with coarsely mining seed corpora of spams and hams from the target data (aiming for spam classification), before reconstructing them as the reference. To catch each word’s rich information in the semantic and syntactic perspectives, we then leverage the natural language processing (NLP) model to embed each word into the high-dimensional vector space and use a neural network to train a spam word model. After that, each message is encoded by using the predicted spam scores from this model for all included stem words. The encoded messages are processed by the prominent outlier techniques to produce their respective scores, allowing us to rank them for making the outlier visible. Our solution is unsupervised, without relying on specifics of any platform or dataset, to be platform-oblivious. Through extensive experiments, our solution is demonstrated to expose spammers’ outlier characteristics effectively, outperform all examined unsupervised methods in almost all metrics, and may even better supervised counterparts.
Yihe Zhang 0001, Xu Yuan 0001, Nian-Feng Tzeng
ACSAC1
2021 Reverse Attack: Black-box Attacks on Collaborative Recommendation
abstract
Collaborative filtering (CF) recommender systems have been extensively developed and widely deployed in various social websites, promoting products or services to the users of interest. Meanwhile, work has been attempted at poisoning attacks to CF recommender systems for distorting the recommend results to reap commercial or personal gains stealthily. While existing poisoning attacks have demonstrated their effectiveness with the offline social datasets, they are impractical when applied to the real setting on online social websites. This paper develops a novel and practical poisoning attack solution toward the CF recommender systems without knowing involved specific algorithms nor historical social data information a priori. Instead of directly attacking the unknown recommender systems, our solution performs certain operations on the social websites to collect a set of sampling data for use in constructing a surrogate model for deeply learning the inherent recommendation patterns. This surrogate model can estimate the item proximities, learned by the recommender systems. By attacking the surrogate model, the corresponding solutions (for availability and target attacks) can be directly migrated to attack the original recommender systems. Extensive experiments validate the generated surrogate model's reproductive capability and demonstrate the effectiveness of our attack upon various CF recommender algorithms.
Yihe Zhang 0001, Xu Yuan 0001, Jin Li 0002, Jiadong Lou, Li Chen 0019, Nian-Feng Tzeng
CCS1
2021 Precise Weather Parameter Predictions for Target Regions via Neural Networks
Yihe Zhang 0001, Xu Yuan 0001, Sytske K. Kimball, Eric Rappin, Li Chen 0019, Paul J. Darby III, Tom Johnsten, Lu Peng 0001, Boisy Pitre, David M. Bourrie, Nian-Feng Tzeng
ECML/PKDD (5)1
2020 Towards Poisoning the Neural Collaborative Filtering-Based Recommender Systems
Yihe Zhang 0001, Jiadong Lou, Li Chen 0019, Xu Yuan 0001, Jin Li 0002, Tom Johnsten, Nian-Feng Tzeng
ESORICS (1)1
2020 Enabling Encrypted Boolean Queries in Geographically Distributed Databases
abstract
The persistent growth of big data applications has being raising new challenges in managing large volumes of datasets with high scalability, confidentiality protection, and flexible types of search queries. In this paper, we propose a secure design to disassemble the private dataset with the aim to store them across geographically distributed servers while supporting secure multi-client Boolean queries. In this design, the data owner encrypts the private database with the searchable index attributes. The encrypted dataset will be disassembled and distributed evenly across multiple servers by leveraging the property of a distributed index framework. By constructing an encryption structure, generating search tokens, and enabling parallel query, we show how the proposed design performs the secure while efficient Boolean search. These queries are not only limited to those initiated by the data owner but also can be extended to support multiple authorized clients, where each client is allowed to access a necessary part of the private database. In this stage, we advocate a non-interactive authorization scheme where data owner is not required to stay online to process the query request. Moreover, the query operation can be executed in parallel, which significantly improves the search efficiency. We formally characterize the leakage profile, which allow us to follow the existing security analysis method to demonstrate that our system can guarantee data confidentiality and query privacy. To validate our protocol, we implement a system prototype and evaluate the efficiency of our construction. Through experimental results, we demonstrate the effectiveness of our protocol in terms of data outsourcing time and Boolean query time.
Xu Yuan 0001, Xingliang Yuan, Yihe Zhang 0001, Baochun Li, Cong Wang 0001
IEEE Trans. Parallel Distributed Syst.3
2019 TweetScore: Scoring Tweets via Social Attribute Relationships for Twitter Spammer Detection
abstract
The spammers have been grossly detrimental since the inception of Twitter social networks and keep polluting social environments by hiding themselves among a large amount of normal users. In this paper, we aim to address two challenges existing in the spammer detection problem: 1) monitoring tweets that have a higher probability of including spam messages; 2) providing an accurate solution for spam classification. To address these two challenges, we first propose a pseudo-honeypot framework for efficient tweets monitoring and collection. By taking advantage of users' diversity and selecting normal users as the parasitic body, the pseudo-honeypot can harness normal users with features having much more potentials of attracting spammers. This lets the pseudo-honeypot collect tweets that are far more likely to include spam messages. Furthermore, we design a novel spam classification solution called TweetScore by exploring both the intrinsic attributes' and users' relationships in social networks. TweetScore quantifies such relationships into a vector of numerical values to represent each tweet's score, reflecting the associated user's behaviors. The neural network is then employed to take these vectors as input to classify spams and spammers. Through extensive experiments, we demonstrate the efficiency of the pseudo-honeypot system on spam monitoring and the accuracy of TweetScore on spam classification. Specifically, the spam and spammer ratios collected by our pseudo-honeypot system are four times as much as those of a non pseudo-honeypot counterpart while the TweetScore can achieve, on an average, 93.5% accuracy, 93.71% precision, and 1.52% false positive in online spam classification.
Yihe Zhang 0001, Hao Zhang 0023, Xu Yuan 0001, Nian-Feng Tzeng
AsiaCCS1
2019 Toward Efficient Spammers Gathering in Twitter Social Networks
abstract
This paper introduces a novel system, named pseudo-honeypot, for efficient spammers gathering. Different from the manual setup in the honeypot, the pseudo-honeypot takes advantage of Twitter users' diversity and selects accounts with the attributes of having the higher potentials of attracting spammers, as the parasitic bodies. By harnessing a set of normal accounts possessing these attributes and monitoring their streaming posts and behavioral patterns, the pseudo-honeypot can gather the tweets that are far more likely of including spammer activities, while removing the risks of being recognized by smart spammers. It substantially advances the honeypot-based solutions in attribute availability, deployment flexibility, network scalability, and system portability. We present the system design and implementation of pseudo-honeypot (including node selection, monitoring, feature extraction, and learning-based classification) in Twitter networks. Through experiments, we demonstrate its effectiveness in term of spammer gathering.
Yihe Zhang 0001, Hao Zhang 0023, Xu Yuan 0001
CODASPY1
2019 Pseudo-Honeypot: Toward Efficient and Scalable Spam Sniffer
abstract
Honeypot-based spammer gathering solutions usually lack attribute variability, deployment flexibility, and network scalability, deemed as their common drawbacks. This paper explores pseudo-honeypot, a novel honeypot-like system to overcome such drawbacks, for efficient and scalable spammer sniffing. The pseudo-honeypot takes advantage of user diversity and selects normal accounts, with attributes that have the higher potential of attracting spammers, as the parasitic bodies. By harnessing such category of users, pseudo-honeypot can monitor their streaming posts and behavioral patterns transparently. When compared with its traditional honeypot counterpart, the proposed solution offers the substantial advantages of attribute variability, deployment flexibility, network scalability, and system portability. Meanwhile, it offers a novel method to collect the social network dataset that has a higher probability of including spams and spammers, without being noticed by advanced spammers. We take the Twitter social network as an example to exhibit its system design, including pseudo-honeypot nodes selection, monitoring, feature extraction, ground truth labeling, and learning-based classification. Through experiments, we demonstrate the efficiency of pseudo-honeypot in terms of spams and spammers gathering. In particular, we confirm our solution can garner spammers at least 19 times faster than the state-of-the-art honeypot-based counterpart.
Yihe Zhang 0001, Hao Zhang 0023, Xu Yuan 0001, Nian-Feng Tzeng
DSN1