Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Haozhou Wang

dblp:131/4030 · DBLP profile ↗
← Back
13ranked-venue papers
4as first author
5since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 7 · 2 first-author · 1 since 2021Artificial intelligence and machine learning · 3 · 3 first-author · 1 since 2021Systems, architecture and hardware · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
5 papers
Spatial and temporal data management · 27% Database system architecture and tuning · 20% Distributed and cloud data management · 20%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
GPUs and heterogeneous computing · 78% Hardware accelerators and domain-specific architectures · 12% Storage systems · 10%
Network and information security
1 paper
Cryptographic primitives and cryptanalysis · 100%
Artificial intelligence
1 paper
Representation and self-supervised learning · 100%

Topics — the 18 heaviest of 20, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Cryptographic primitives and cryptanalysis › public-key cryptography › elliptic curve
elliptic-curve arithmetic
0.912025
gECC: A GPU-based high-throughput framework for Elliptic Curve Cryptography · ACM Trans. Archit. Code Optim. 2025
Cryptographic primitives and cryptanalysis › public-key cryptography
elliptic curve cryptography
0.912025
gECC: A GPU-based high-throughput framework for Elliptic Curve Cryptography · ACM Trans. Archit. Code Optim. 2025
GPUs and heterogeneous computing › GPU computing › GPGPU acceleration
GPU-accelerated cryptography
0.912025
gECC: A GPU-based high-throughput framework for Elliptic Curve Cryptography · ACM Trans. Archit. Code Optim. 2025
GPUs and heterogeneous computing
GPU computing
0.912025
gECC: A GPU-based high-throughput framework for Elliptic Curve Cryptography · ACM Trans. Archit. Code Optim. 2025
Distributed and cloud data management
distributed query processing
0.512021
Greenplum: A Hybrid Database for Transactional and Analytical Workloads · SIGMOD Conference 2021
Database system architecture and tuning
hybrid transactional and analytical processing
0.512021
Greenplum: A Hybrid Database for Transactional and Analytical Workloads · SIGMOD Conference 2021
Spatial and temporal data management
trajectory data management
0.422015
Calibrating trajectory data for spatio-temporal similarity analysis · VLDB J. 2015
Calibrating trajectory data for similarity-based analysis · SIGMOD Conference 2013
Machine learning › Representation and self-supervised learning › word representation › word embedding
cross-lingual word embedding
0.412019
Weakly-Supervised Concept-based Adversarial Learning for Cross-lingual Word Embeddings · EMNLP/IJCNLP (1) 2019
Hardware accelerators and domain-specific architectures
cryptographic accelerator
0.312025
gECC: A GPU-based high-throughput framework for Elliptic Curve Cryptography · ACM Trans. Archit. Code Optim. 2025
Indexing and storage engines
column store
0.212015
SharkDB: An In-Memory Storage System for Massive Trajectory Data · SIGMOD Conference 2015
Data mining
similarity analysis
0.212015
Calibrating trajectory data for spatio-temporal similarity analysis · VLDB J. 2015
Spatial and temporal data management
spatio-temporal similarity
0.212015
Calibrating trajectory data for spatio-temporal similarity analysis · VLDB J. 2015
Storage systems › storage architecture
in-memory storage
0.212015
SharkDB: An In-Memory Storage System for Massive Trajectory Data · SIGMOD Conference 2015
Recommender systems › domain-specific recommendation
route recommendation
0.212014
A crowd-based route recommendation system-CrowdPlanner · ICDE 2014
Transaction processing and concurrency control › distributed commit protocols
two-phase commit
0.112021
Greenplum: A Hybrid Database for Transactional and Analytical Workloads · SIGMOD Conference 2021
Machine learning › Representation and self-supervised learning › word representation
word embedding
0.112019
Weakly-Supervised Concept-based Adversarial Learning for Cross-lingual Word Embeddings · EMNLP/IJCNLP (1) 2019
Spatial and temporal data management
trajectory data
0.112015
SharkDB: An In-Memory Storage System for Massive Trajectory Data · SIGMOD Conference 2015
Collaborative and social computing
crowdsourcing
0.112014
A crowd-based route recommendation system-CrowdPlanner · ICDE 2014

Methods — techniques the papers use, named apart from their topics

montgomery's trick · 1.7modular arithmetic optimization · 1.7batch execution · 1.7resource group scheduling · 0.5one-phase commit · 0.5global deadlock detection · 0.5multicore parallelization · 0.4compression · 0.4weakly supervised learning · 0.4adversarial learning · 0.4task generation · 0.4machine learning · 0.2geometry-based calibration · 0.2
YearPublicationVenuePosition
2026 DODA: Adapting Object Detectors to Dynamic Agricultural Environments in Real-Time with Diffusion
Shuai Xiang, Pieter M. Blok, James Burridge, Haozhou Wang, Wei Guo 0002
WACV4
2025 gECC: A GPU-based high-throughput framework for Elliptic Curve Cryptography
abstract
Elliptic Curve Cryptography (ECC) is an encryption method that provides security comparable to traditional techniques like Rivest–Shamir–Adleman (RSA) but with lower computational complexity and smaller key sizes, making it a competitive option for applications such as blockchain, secure multi-party computation, and database security. However, the throughput of ECC is still hindered by the significant performance overhead associated with elliptic curve (EC) operations, which can affect their efficiency in real-world scenarios. This article presents gECC , a versatile framework for ECC optimized for GPU architectures, specifically engineered to achieve high-throughput performance in EC operations. To maximize throughput, gECC incorporates batch-based execution of EC operations and microarchitecture-level optimization of modular arithmetic. It employs Montgomery’s trick [ 40 ] to enable batch EC computation and incorporates novel computation parallelization and memory management techniques to maximize the computation parallelism and minimize the access overhead of GPU global memory. Furthermore, we analyze the primary bottleneck in modular multiplication by investigating how the user codes of modular multiplication are compiled into hardware instructions and what these instructions’ issuance rates are. We identify that the efficiency of modular multiplication is highly dependent on the number of Integer Multiply-Add (IMAD) instructions. To eliminate this bottleneck, we propose novel techniques to minimize the number of IMAD instructions by leveraging predicate registers to pass the carry information and using addition and subtraction instructions (IADD3) to replace IMAD instructions. Our experimental results show that, for ECDSA and ECDH, the two commonly used ECC algorithms, gECC can achieve performance improvements of 5.56 × and 4.94 ×, respectively, compared to the state-of-the-art GPU-based system. In a real-world blockchain application, we can achieve performance improvements of 1.56 ×, compared to the state-of-the-art CPU-based system. gECC is completely and freely available at https://github.com/CGCL-codes/gECC .
Qian Xiong, Weiliang Ma, Xuanhua Shi, Yongluan Zhou, Hai Jin 0001, Haozhou Wang, Zhengru Wang
ACM Trans. Archit. Code Optim.7
2025 Corrigendum: gECC: A GPU-based high-throughput framework for Elliptic Curve Cryptography
abstract
This is a corrigendum for the article “gECC: A GPU-based high-throughput framework for Elliptic Curve Cryptography” published in ACM Trans. Arch. Code Optim. 22, 3, Article 84 (September 2025), 27 pages.
Qian Xiong, Weiliang Ma, Xuanhua Shi, Yongluan Zhou, Hai Jin 0001, Haozhou Wang, Zhengru Wang
ACM Trans. Archit. Code Optim.7
2021 Multi-Adversarial Learning for Cross-Lingual Word Embeddings
abstract
Generative adversarial networks (GANs) have succeeded in inducing cross-lingual word embeddings -maps of matching words across languages-without supervision.Despite these successes, GANs' performance for the difficult case of distant languages is still not satisfactory.These limitations have been explained by GANs' incorrect assumption that source and target embedding spaces are related by a single linear mapping and are approximately isomorphic.We assume instead that, especially across distant languages, the mapping is only piece-wise linear, and propose a multi-adversarial learning method.This novel method induces the seed cross-lingual dictionary through multiple mappings, each induced to fit the mapping for one subspace.Our experiments on unsupervised bilingual lexicon induction and cross-lingual document classification show that this method improves performance over previous single-mapping methods, especially for distant languages.
Haozhou Wang, James Henderson 0001, Paola Merlo
NAACL-HLT1
2021 Greenplum: A Hybrid Database for Transactional and Analytical Workloads
abstract
Demand for enterprise data warehouse solutions to support real-time Online Transaction Processing (OLTP) queries as well as long-running Online Analytical Processing (OLAP) workloads is growing. Greenplum database is traditionally known as an OLAP data warehouse system with limited ability to process OLTP workloads. In this paper, we augment Greenplum into a hybrid system to serve both OLTP and OLAP workloads. The challenge we address here is to achieve this goal while maintaining the ACID properties with minimal performance overhead. In this effort, we identify the engineering and performance bottlenecks such as the under-performing restrictive locking and the two-phase commit protocol. Next we solve the resource contention issues between transactional and analytical queries. We propose a global deadlock detector to increase the concurrency of query processing. When transactions that update data are guaranteed to reside on exactly one segment we introduce one-phase commit to speed up query processing. Our resource group model introduces the capability to separate OLAP and OLTP workloads into more suitable query processing mode. Our experimental evaluation on the TPC-B and CH-benCHmark benchmarks demonstrates the effectiveness of our approach in boosting the OLTP performance without sacrificing the OLAP performance.
Zhenghua Lyu, Huan Hubert Zhang, Haozhou Wang, Jinbao Chen, Asim Praveen, Xiaoming Gao, Alexandra Wang, Wen Lin 0004, Ashwin Agrawal, Jesse Zhang, Venkatesh Raghavan
SIGMOD Conference5
2019 Weakly-Supervised Concept-based Adversarial Learning for Cross-lingual Word Embeddings
abstract
Haozhou Wang, James Henderson, Paola Merlo. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Haozhou Wang, James Henderson 0001, Paola Merlo
EMNLP/IJCNLP (1)1
2018 SharkDB: an in-memory column-oriented storage for trajectory analysis
Bolong Zheng, Haozhou Wang, Kai Zheng 0001, Han Su 0001, Kuien Liu, Shuo Shang
World Wide Web2
2015 SharkDB: An In-Memory Storage System for Massive Trajectory Data
abstract
An increasing amount of motion history data, which is called trajectory, is being collected from different sources such as GPS-enabled mobile devices, surveillance cameras and social networks. However it is hard to store and manage trajectory data in traditional database systems, since its variable lengths and asynchronous sampling rates do not fit disk-based and tuple-oriented structures, which are the fundamental structures of traditional database systems. We implement a novel trajectory storage system that is motivated by the success of column store and recent development of in-memory based databases. In this storage design, we try to explore the potential opportunities, which can boost the performance of query processing for trajectory data. To achieve this, we partition the trajectories into frames as column-oriented storage in order to store the sample points of a moving object, which are aligned by the time interval, within the main memory. Furthermore, the frames can be highly compressed and well structured to increase the memory utilization ratio and reduce the CPU-cache missing. It is also easier for parallelizing data processing on the multi-core server since the frames are mutually independent.
Haozhou Wang, Kai Zheng 0001, Xiaofang Zhou 0001, Shazia Sadiq
SIGMOD Conference1
2015 Calibrating trajectory data for spatio-temporal similarity analysis
Han Su 0001, Kai Zheng 0001, Jiamin Huang, Haozhou Wang, Xiaofang Zhou 0001
VLDB J.4
2014 SharkDB: An In-Memory Column-Oriented Trajectory Storage
abstract
The last decade has witnessed the prevalence of sensor and GPS technologies that produce a high volume of trajectory data representing the motion history of moving objects. However some characteristics of trajectories such as variable lengths and asynchronous sampling rates make it difficult to fit into traditional database systems that are disk-based and tuple-oriented. Motivated by the success of column store and recent development of in-memory databases, we try to explore the potential opportunities of boosting the performance of trajectory data processing by designing a novel trajectory storage within main memory. In contrast to most existing trajectory indexing methods that keep consecutive samples of the same trajectory in the same disk page, we partition the database into frames in which the positions of all moving objects at the same time instant are stored together and aligned in main memory. We found this column-wise storage to be surprisingly well suited for in-memory computing since most frames can be stored in highly compressed form, which is pivotal for increasing the memory throughput and reducing CPU-cache miss. The independence between frames also makes them natural working units when parallelizing data processing on a multi-core environment. Lastly we run a variety of common trajectory queries on both real and synthetic datasets in order to demonstrate advantages and study the limitations of our proposed storage.
Haozhou Wang, Kai Zheng 0001, Jiajie Xu 0001, Bolong Zheng, Xiaofang Zhou 0001, Shazia Sadiq
CIKM1
2014 A crowd-based route recommendation system-CrowdPlanner
abstract
Route recommendation service has become a big business in industry since traveling is now an important part of our daily life. We can travel to unknown places by simply typing in our destination and then following recommendation service's guidance, that a pleasant trip desires them to provide a good route. However, previous research shows that even the routes recommended by the big-thumb service providers can deviate significantly from the routes travelled by experienced drivers since the many latent factors affect drivers' preferences and it is hard for a single route recommendation algorithm to model all of them. In this demo we will present the CrowPlanner system to leverage crowds' knowledge to improve the recommendation quality. It requests human workers to evaluate candidates routes recommended by different sources and methods, and determines the best route based on the feedbacks of these workers. In this demo, we first introduce the core component of our system for smart question generation, and then show several real route recommendation cases and the feedback of users.
Han Su 0001, Kai Zheng 0001, Jiamin Huang, Haozhou Wang, Xiaofang Zhou 0001
ICDE5
2014 Cost-Efficient Spatial Network Partitioning for Distance-Based Query Processing
abstract
The efficiency of spatial query processing is crucial for many applications such as location-based services. In spatial networks, queries like k-NN queries are all based on network distance evaluation. Classic solutions for these queries rely on network expansion and are not efficient enough for large networks. Some approaches have improved the query efficiency but brought considerable space cost for index. To address these problems, we propose a hierarchical graph partitioning based index named Partition Tree. It organizes the vertices of a spatial network into a hierarchy through a series of graph partitioning processes. Meanwhile precomputed distances are associated with this hierarchy to facilitate efficient query processing. Inspired by the observation that queries are usually invoked around objects of interest, we propose a query-oriented optimization on top of the Partition Tree. It uses a cost model to evaluate the influence of the object distribution and partitioning topology on the query efficiency. Then a cost-efficient graph partitioning method is developed based on this cost model. Experimental results on real datasets demonstrate that our proposed index and algorithms have superior performance over the state-of-the-art approaches and are scalable to large spatial networks.
Kai Zheng 0001, Hoyoung Jeung, Haozhou Wang, Bolong Zheng, Xiaofang Zhou 0001
MDM (1)4
2013 Calibrating trajectory data for similarity-based analysis
abstract
Due to the prevalence of GPS-enabled devices and wireless communications technologies, spatial trajectories that describe the movement history of moving objects are being generated and accumulated at an unprecedented pace. Trajectory data in a database are intrinsically heterogeneous, as they represent discrete approximations of original continuous paths derived using different sampling strategies and different sampling rates. Such heterogeneity can have a negative impact on the effectiveness of trajectory similarity measures, which are the basis of many crucial trajectory processing tasks. In this paper, we pioneer a systematic approach to trajectory calibration that is a process to transform a heterogeneous trajectory dataset to one with (almost) unified sampling strategies. Specifically, we propose an anchor-based calibration system that aligns trajectories to a set of anchor points, which are fixed locations independent of trajectory data. After examining four different types of anchor points for the purpose of building a stable reference system, we propose a geometry-based calibration approach that considers the spatial relationship between anchor points and trajectories. Then a more advanced model-based calibration method is presented, which exploits the power of machine learning techniques to train inference models from historical trajectory data to improve calibration effectiveness. Finally, we conduct extensive experiments using real trajectory datasets to demonstrate the effectiveness and efficiency of the proposed calibration system.
Han Su 0001, Kai Zheng 0001, Haozhou Wang, Jiamin Huang, Xiaofang Zhou 0001
SIGMOD Conference3