VLDB 2026 Research / reviewers in the wild / expert
Haoxiang Zhang 0003
dblp:07/1059-3
· DBLP profile ↗
7ranked-venue papers
2as first author
7since 2021 · last 2024
0009-0007-6939-6591ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Sampling Methods for Inner Product SketchingabstractRecently, Bessa et al. (PODS 2023) showed that sketches based on coordinated weighted sampling theoretically and empirically outperform popular linear sketching methods like Johnson-Lindentrauss projection and CountSketch for the ubiquitous problem of inner product estimation. We further develop this finding by introducing and analyzing two alternative sampling-based methods. In contrast to the computationally expensive algorithm in Bessa et al., our methods run in linear time (to compute the sketch) and perform better in practice, significantly beating linear sketching on a variety of tasks. For example, they provide state-of-the-art results for estimating the correlation between columns in unjoined tables, a problem that we show how to reduce to inner product estimation in a black-box way. While based on known sampling techniques (threshold and priority sampling) we introduce significant new theoretical analysis to prove approximation guarantees for our methods. Majid Daliri, Juliana Freire, Christopher Musco, Aécio S. R. Santos, Haoxiang Zhang 0003 |
Proc. VLDB Endow. | 5 |
| 2023 | Time-Aware POI Recommendation Based on Multi-Grained Location GroupingabstractThe task of point-of-interest (POI) recommendation aims to recommend locations to users in location-based applications. Among them, the task of time-aware POI recommendation aims to capture the user’s preferences that change dynamically over time, so as to make more accurate recommendations to users at a specific time. While existing works take into account the spatial, temporal and category context of POIs, they cannot capture user preferences that are more fine-grained than the category granularity. Additionally, RNN-based methods suffer from the problem of long-term dependency when capturing a user’s check-in patterns. To address these challenges, we propose a novel model with POI multi-grained grouping method which captures the user’s co-visit patterns and weekly patterns, to obtain finer-grained POI groups. The model also utilizes the transformer model to capture the user’s check-in preference patterns. We evaluate our model on two real-world datasets, and the experimental results demonstrate the effectiveness of our proposed model. Haoxiang Zhang 0003, Wenchao Bai, Jingyi Ding, Jiahui Jin 0001 |
CSCWD | 1 |
| 2023 | Weighted Minwise Hashing Beats Linear Sketching for Inner Product EstimationabstractWe present a new approach for independently computing compact sketches that can be used to approximate the inner product between pairs of high-dimensional vectors. Based on the Weighted MinHash algorithm, our approach admits strong accuracy guarantees that improve on the guarantees of popular linear sketching approaches for inner product estimation, such as CountSketch and Johnson-Lindenstrauss projection. Specifically, while our method exactly matches linear sketching for dense vectors, it yields significantly lower error for sparse vectors with limited overlap between non-zero entries. Such vectors arise in many applications involving sparse data, as well as in increasingly popular dataset search applications, where inner products are used to estimate data covariance, conditional means, and other quantities involving columns in unjoined tables. We complement our theoretical results by showing that our approach empirically outperforms existing linear sketches and unweighted hashing-based sketches for sparse vectors. Aline Bessa, Majid Daliri, Juliana Freire, Cameron Musco, Christopher Musco, Aécio S. R. Santos, Haoxiang Zhang 0003 |
PODS | 7 |
| 2023 | CitySpec with shield: A secure intelligent assistant for requirement formalization
Zirong Chen, Isaac Li, Haoxiang Zhang 0003, Sarah Masud Preum, John A. Stankovic, Meiyi Ma |
Pervasive Mob. Comput. | 3 |
| 2022 | CitySpec: An Intelligent Assistant System for Requirement Specification in Smart CitiesabstractAn increasing number of monitoring systems have been developed in smart cities to ensure that a city's real-time operations satisfy safety and performance requirements. However, many existing city requirements are written in English with missing, inaccurate, or ambiguous information. There is a high demand for assisting city policy makers in converting human-specified requirements to machine-understandable formal specifications for monitoring systems. To tackle this limitation, we build CitySpec, the first intelligent assistant system for requirement specification in smart cities. To create CitySpec, we first collect over 1,500 real-world city requirements across different domains from over 100 cities and extract city-specific knowledge to generate a dataset of city vocabulary with 3,061 words. We also build a translation model and enhance it through requirement synthesis and develop a novel online learning framework with validation under uncertainty. The evaluation results on real-world city requirements show that CitySpec increases the sentence-level accuracy of requirement specification from 59.02 % to 86.64 %, and has strong adaptability to a new city and a new domain (e.g., F1 score for requirements in Seattle increases from 77.6 % to 93.75% with online learning). Zirong Chen, Isaac Li, Haoxiang Zhang 0003, Sarah Masud Preum, John A. Stankovic, Meiyi Ma |
SMARTCOMP | 3 |
| 2022 | An Intelligent Assistant for Converting City Requirements to Formal SpecificationabstractAs more and more monitoring systems have been deployed to smart cities, there comes a higher demand for converting new human-specified requirements to machine-understandable formal specifications automatically. However, these human-specific requirements are often written in English and bring missing, inaccurate, or ambiguous information. In this paper, we present City Spec [1], an intelligent assistant system for requirement specification in smart cities. CitySpec not only helps overcome the language differences brought by English requirements and formal specifications, but also offers solutions to those missing, inaccurate, or ambiguous information. The goal of this paper is to demonstrate how CitySpec works. Specifically, we present three demos: (1) interactive completion of requirements in CitySpec; (2) human-in-the-loop correction while CitySepc encounters exceptions; (3) online learning in CitySpec. Zirong Chen, Isaac Li, Haoxiang Zhang 0003, Sarah Masud Preum, John A. Stankovic, Meiyi Ma |
SMARTCOMP | 3 |
| 2021 | DSDD: Domain-Specific Dataset Discovery on the WebabstractWith the push for transparency and open data, many datasets and data repositories are becoming available on the Web. This opens new opportunities for data-driven exploration, from empowering analysts to answer new questions and obtain insights to improving predictive models through data augmentation. But as datasets are spread over a plethora of Web sites, finding data that are relevant for a given task is difficult. In this paper, we take a first step towards the construction of domain-specific data lakes. We propose an end-to-end dataset discovery system, targeted at domain experts, which given a small set of keywords, automatically finds potentially relevant datasets on the Web. The system makes use of search engines to hop across Web sites, uses online learning to incrementally build a model to recognize sites that contain datasets, utilizes a set of discovery actions to broaden the search, and applies a multi-armed bandit based algorithm to balance the trade-offs of different discovery actions. We report the results of an extensive experimental evaluation over multiple domains, and demonstrate that our strategy is effective and outperforms state-of-the-art content discovery methods. Haoxiang Zhang 0003, Aécio S. R. Santos, Juliana Freire |
CIKM | 1 |