Haishan Wu

dblp:12/551 · DBLP profile ↗
← Back
14ranked-venue papers
1as first author
4since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 6 · 1 since 2021Artificial intelligence and machine learning · 4Applied, interdisciplinary, general and emerging computing · 4Software engineering, systems software and programming languages · 3 · 1 first-author · 2 since 2021Security and privacy · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Systems, architecture and hardware · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2025 Enhancing Meme Token Market Transparency: A Multi-Dimensional Entity-Linked Address Analysis for Liquidity Risk Evaluation
Frank Fan, Haishan Wu, Xueyan Tang
ICBC4
2025 Detecting Sybil Addresses in Blockchain Airdrops: A Subgraph-based Feature Propagation and Fusion Approach
Frank Fan, Haishan Wu, Xueyan Tang
ICBC4
2024 A Solution to ACMMM 2024 on Artificial Intelligence Generated Image Detection
ShiHang Li, Haishan Wu
ACM Multimedia2
2022 A Unified Framework for User Identification Across Online and Offline Data
abstract
User identification across multiple datasets has a wide range of applications and there has been an increasing set of research works on this topic during recent years. However, most of existing works focus on user identification with a single input data type, e.g., (I) identifying a user across multiple social networks with online data and (II) detecting a single user from heterogeneous trajectory datasets with offline data. Different from previous works, in this paper, we propose a framework on user identification between online and offline datasets. We build connections between these two types of data by a mapping from IP addresses to physical locations. To solve this problem, we propose a novel framework consisting of three steps. First, we use a clustering method based on locations of IP addresses to map IP addresses into specific physical location distributions. Second, we propose a novel pairwise index to reduce space cost and running time for computing the co-occurrence. Lastly, we apply a learning-to-rank method to merge the effect of multiple features we get in the first two steps. Based on our framework, we design experiments to demonstrate the efficiency (in time and space) of our framework, together with the precision and recall of our approach compared to other methods.
Tianyi Hao 0001, Yunsheng Cheng, Longbo Huang, Haishan Wu
IEEE Trans. Knowl. Data Eng.5
2018 Intent-Aware Audience Targeting for Ride-Hailing Service
Yuan Xia, Jingbo Zhou 0003, Jingjia Cao, Haishan Wu, Hui Xiong 0001
ECML/PKDD (3)7
2017 Lesion Detection and Grading of Diabetic Retinopathy via Two-Stages Deep Convolutional Neural Networks
Yehui Yang, Wensi Li, Haishan Wu, Wensheng Zhang 0002
MICCAI (3)4
2017 Hierarchical Contextual Attention Recurrent Neural Network for Map Query Suggestion
abstract
The query logs from an on-line map query system provide rich cues to understand the behaviors of human crowds. With the growing ability of collecting large scale query logs, the query suggestion has been a topic of recent interest. In general, query suggestion aims at recommending a list of relevant queries w.r.t. users’ inputs via an appropriate learning of crowds’ query logs. In this paper, we are particularly interested in map query suggestions (e.g., the predictions of location-related queries) and propose a novel modelHierarchical Contextual Attention Recurrent Neural Network(HCAR-NN) for map query suggestion in an encoding-decoding manner. Given crowds map query logs, our proposed HCAR-NN not only learns the local temporal correlation among map queries in a query session (e.g., queries in a short-term interval are relevant to accomplish a search mission), but also captures the global longer range contextual dependencies among map query sessions in query logs (e.g., how a sequence of queries within a short-term interval has an influence on another sequence of queries). We evaluate our approach over millions of queries from a commercial search engine (i.e.,Baidu Map). Experimental results show that the proposed approach provides significant performance improvements over the competitive existing methods in terms of classical metrics (i.e.,Recall@KandMRR) as well as the prediction of crowds’ search missions.
Jun Song 0004, Jun Xiao 0001, Fei Wu 0001, Haishan Wu, Tong Zhang 0001, Zhongfei Zhang, Wenwu Zhu 0001
IEEE Trans. Knowl. Data Eng.4
2016 User identification in cyber-physical space: a case study on mobile query logs and trajectories
abstract
User identification across domains draws lots of research effort in recent years. Although most of existing works focus on user identification in a single space, in this paper, we first try to identify users by fusing their activities in cyber space and physical space, which helps us obtain a comprehensive understanding about users' online behaviours as well as offline visitation. Out profound insight to tackle this problem is that we can build a connection between the cyber space and the physical space with the stable location distribution of IP addresses. Thus, we propose a novel framework for user identification in cyber-physical space, which consists of three key steps: 1) modeling the location distribution of each IP address; 2) computing the co-occurrence with an inverted index to reduce the space and time cost; and 3) a learning-to-rank tactic to fuse user's features shared in both spaces to improve the accuracy. We conduct experiments to identify individual users from mobile query logs (generated in cyber space) and trajectory data (generated in physical space) to demonstrate the efficiency and effectiveness of our framework.
Tianyi Hao 0001, Yunsheng Cheng, Longbo Huang, Haishan Wu
SIGSPATIAL/GIS5
2016 Demand driven store site selection via multiple spatial-temporal data
abstract
Choosing a good location when opening a new store is crucial for the future success of a business. Traditional methods include offline manual survey, analytic models based on census data, which are either unable to adapt to the dynamic market or very time consuming. The rapid increase of the availability of big data from various types of mobile devices, such as online query data and offline positioning data, provides us with the possibility to develop automatic and accurate data- driven prediction models for business store site selection. In this paper, we propose a Demand Driven Store Site Selection (DD3S) framework for business store site selection by mining search query data from Baidu Maps. DD3S first detects the spatial-temporal distributions of customer demands on different business services via query data from Baidu Maps, the largest online map search engine in China, and detects the gaps between demand and supply. Then we determine candidate locations via clustering such gaps. In the final stage, we solve the location optimization problem by predicting and ranking the number of customers. We not only deploy supervised regression models to predict the number of customers, but also use learning-to-rank model to directly rank the locations. We evaluate our framework on various types of businesses in real-world cases, and the experiment results demonstrate the effectiveness of our methods. DD3S as the core function for store site selection has already been implemented as a core component of our business analytics platform and could be potentially used by chain store merchants on Baidu Nuomi.
Mengwen Xu, Zhengwei Wu, Jingbo Zhou 0003, Jian Li 0015, Haishan Wu
SIGSPATIAL/GIS6
2016 Automatic user identification method across heterogeneous mobility data sources
abstract
With the ubiquity of location based services and applications, large volume of mobility data has been generated routinely, usually from heterogeneous data sources, such as different GPS-embedded devices, mobile apps or location based service providers. In this paper, we investigate efficient ways of identifying users across such heterogeneous data sources. We present a MapReduce-based framework called Automatic User Identification (AUI) which is easy to deploy and can scale to very large data set. Our framework is based on a novel similarity measure called the signal based similarity (SIG) which measures the similarity of users' trajectories gathered from different data sources, typically with very different sampling rates and noise patterns. We conduct extensive experimental evaluations, which show that our framework outperforms the existing methods significantly. Our study on one hand provides an effective approach for the mobility data integration problem on large scale data sets, i.e., combining the mobility data sets from different sources in order to enhance the data quality. On the other hand, our study provides an in-depth investigation for the widely studied human mobility uniqueness problem under heterogeneous data sources.
Wei Cao 0007, Zhengwei Wu, Dong Wang 0037, Jian Li 0015, Haishan Wu
ICDE5
2013 A Self-Optimizing Workload Management Solution for Cloud Applications
abstract
Given the dynamic nature of the cloud, resulting from mapping virtual to physical resources, changes in the usage pattern of resources, migration of virtual resources and the dynamic nature of the applications themselves, the bottleneck resource in a given application changes over time. Promptly identifying the bottleneck of cloud application and consequently taking corrective actions (e.g. admission control) are essential requirements for cloud application performance management. The traditional threshold based bottleneck detection technology, which adopts a pre-defined target performance measure (e.g. response time, CPU utilization, etc.), requires a good understanding of the application. It is difficult to identify which performance measures need to be monitored and how to set accurate threshold values for them. The commonly used technique of model-based workload management also faces a big challenge in modeling the highly dynamic, cloud application behavior. In this paper, we propose a self-optimizing application workload management solution for cloud applications which adapts well to the cloud dynamics. It utilizes a target-less bottleneck detection mechanism, without the need to define target thresholds. It also contains a model-free controller for workload management, thus avoiding the complexity of dynamically changing the model as the cloud environment changes. We believe that this is the first time such a design principle to cloud application performance management is introduced. The validity and efficiency of this solution have been verified by a real-case study on an IBM cloud platform, using the RUBiS web application benchmark.
Haishan Wu, Asser N. Tantawi
ICWS1
2013 Towards transparent and distributed workload management for large scale web servers
Shengzhi Zhang, Haishan Wu, Athanasios V. Vasilakos, Peng Liu 0005
Future Gener. Comput. Syst.3
2011 Distributed workload and response time management for web applications
Shengzhi Zhang, Haishan Wu, Bo Yang 0013, Peng Liu 0005, Athanasios V. Vasilakos
CNSM2
1995 Computer Aided Simulation System for Orthognathic Surgery
abstract
The easy simulation of orthognathic surgery with high accuracy and good reliability, and the visual explanation of the prediction of surgery to patients with marillo-mandibular deformities remains a key focus in oral and maxillofacial surgery. A Computer-Aided Simulation System for Orthognathic Surgery (CASSOS) was developed by means of the technique of digital image processing, with which a detailed analysis could be performed automatically, including 71 measurements of distance, degree and ratio for frontal cephalograms and 68 measurements for lateral ones, with computer-aided diagnosis. A quantitative surgical simulation for the prediction of operations could be satisfactorily carried out for clinical use. All data and images were printed by Laserjet and color videoprinters. All procedures were finished in 20 minutes.>
Jiong Xia, Feihu Qi, Wenhua Yuan, Weiliu Qiu, Yuanliang Huang, Guofang Shen, Haishan Wu
CBMS9