EDBT 2026 Demo / reviewers in the wild / expert
Boxiang Dong
dblp:11/9085
· DBLP profile ↗
26ranked-venue papers
10as first author
4since 2021 · last 2021
0000-0001-9520-1494ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 16 · 7 first-author · 3 since 2021Artificial intelligence and machine learning · 11 · 6 first-author · 2 since 2021Security and privacy · 4 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 4 · 1 first-authorSystems, architecture and hardware · 2Computer networks · 2Software engineering, systems software and programming languages · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | VDGAN: A Collaborative Filtering Framework Based on Variational Denoising with GANsabstractGenerative Adversarial Networks (GANs) effectively capture the true posterior distribution. When applied to Collaborative Filtering (CF), GANs can generate a recommendation list through implicit feedback. However, the discriminators in the existing GANs-based CF methods are not utilized fully, and the generators perform poorly on sparse data mining. In this paper, we propose an improved collaborative filtering framework based on variational denoising for GANs (VDGAN). Specifically, VDGAN integrates the variational encoder and the self-attention mechanism into the GANs. By using the positive-negative sampling mechanism to add specific noise to the input data, the variational encoder obtains a robust feature matrix and improves the sparse data processing capability of the generator. In VDGAN, the denoising generator reconstructs the user-items interaction matrix through the feature matrix. And the discriminator is composed of the self-attention mechanism to obtain the explicit features of user preferences, which extends the ability of the discriminator. Furthermore, reinforcement learning replaces the traditional objective function of GANs, which better optimizes the generator and further improves the recommendation accuracy of the model. From our comprehensive experiments on three real-world datasets, we demonstrate that the performance of VDGAN significantly outperforms the state-of-the-art methods based on GANs and Auto-Encoders. Weifeng Sun 0002, Shumiao Yu, Boxiang Dong |
IJCNN | 3 |
| 2021 | VeriDL: Integrity Verification of Outsourced Deep Learning Services
Boxiang Dong, Bo Zhang 0051, Wendy Hui Wang |
ECML/PKDD (2) | 1 |
| 2021 | CorrectMR: Authentication of Distributed SQL Execution on MapReduceabstractIn this paper, we consider the SQL Selection-GroupBy-Aggregation (SGA) query evaluation on an untrusted MapReduce system in which mappers and reducers may return incorrect results. We design CorrectMR, a system that supports efficient verification of result correctness for both intermediate and final results of SGA queries. CorrectMR includes the design of Pedersen Merkle R-tree (PMR-tree), a new authenticated data structure (ADS). To enable efficient verification, CorrectMR includes a distributed ADS construction mechanism that allows mappers/reducers to construct PMR-trees in parallel without a centralized party. CorrectMR provides the following verification functionality: (1) correctness verification of PMR-trees by replication; (2) correctness verification of intermediate (final, resp.) query results by constructing local (global, resp.) PMR-trees and verification objects. Our experimental results demonstrate the efficiency and effectiveness of CorrectMR. Bo Zhang 0051, Boxiang Dong, Wendy Hui Wang |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2021 | Integrity Authentication for SQL Query Evaluation on Outsourced Databases: A SurveyabstractSpurred by the development of cloud computing, there has been considerable recent interest in the Database-as-a-Service (DaaS) paradigm. Users lacking in expertise or computational resources can outsource their data and database management needs to a third-party service provider. Outsourcing, however, raises an important issue of result integrity: how can the client verify with lightweight overhead that the query results returned by the service provider are correct (i.e., the same as the results of query execution locally)? This survey focuses on categorizing and reviewing the progress on the current approaches for result integrity of SQL query evaluation in the DaaS model. The survey also includes some potential future research directions for result integrity verification of the outsourced computations. Bo Zhang 0051, Boxiang Dong, Wendy Hui Wang |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2020 | 2D-ATT: Causal Inference for Mobile Game Organic Installs with 2-Dimensional Attentional Neural NetworkabstractIn the mobile gaming industry, organic installs refer to downloads that cannot be attributed to any advertising channel and thus do not introduce upfront user acquisition (UA) cost. Understanding the causal factors on organic installs is of vital importance for a game's ecosystem, as such knowledge can help bring in more organic users, who tend to be more loyal and active. A major challenge in discovering the causal effects is the potential temporal lag between an UA operation and the growth in organic installs. In this paper, we solve the problem by using a deep attentional neural network to analyze multivariate time series data. The core of our design is a novel attention mechanism, namely 2D-ATT, that can learn the contribution of each feature to the target at different levels of temporal delay. Our experiments on a series of synthetic datasets show that 2D-ATT outperforms existing approaches for discovering complex causal effects. We also use 2D-ATT to analyze a real-world mobile game dataset collected by Jam City, a video game company based in California. Our discoveries provide valuable insights to UA operations. Boxiang Dong, Hui Bill Li, Yang Ryan Wang, Rami Safadi |
IEEE BigData | 1 |
| 2020 | LSOMP: Large Scale Ordinance Mining PortalabstractWe propose a novel scalable Web portal called LSOMP (Large Scale Ordinance Mining Portal) to analyze ordinances and their tweets (of the order of thousands and millions). It entails commonsense knowledge (CSK) and natural language processing (NLP), disseminating ordinance-tweet mining results via interactive graphics and Question Answering (QA). Matthew Kowalski, Aparna S. Varde, Boxiang Dong |
IEEE BigData | 4 |
| 2020 | AuthPDB: Authentication of Probabilistic Queries on Outsourced Uncertain DataabstractQuery processing over uncertain data has gained much attention recently. Due to the high computational complexity of query evaluation on uncertain data, the data owner can outsource her data to a server that provides query evaluation as a service. However, a dishonest server may return cheap (and incorrect) query answers, hoping that the client who has weak computational power cannot catch the incorrect results. To address the integrity issue, in this paper, we design AuthPDB, a framework that supports efficient authentication of query evaluation for both all-answer and top-k queries on outsourced probabilistic databases. Our empirical results on real-world datasets demonstrate the effectiveness and efficiency of AuthPDB. Bo Zhang 0051, Boxiang Dong, Haipei Sun, Wendy Hui Wang |
CODASPY | 2 |
| 2020 | A Novel Collaborative Filtering Framework Based on Variational Self-Attention GANabstractIt is difficult for users to find the required information promptly in the massive data. The collaborative filtering is an effective way to help users get the proper information. To achieve better performance, an improved framework based on variational Generative Adversarial Networks with self-attention (VGCF) is proposed. In VGCF, self-attention mechanism and Variational Autoencoders are combined to form self-attention variational encoder, which improves the ability of acquiring explicit and implicit features of sparse data and obtains the personal preference and the correlation between users. By utilizing Generative Adversarial Networks for prediction, the generative model uses the compression matrix obtained by self-attention variational encoder to generate the predicted user-items interaction matrix. The discriminative model provides a better approximation for the posterior and maximum-likelihood assignment, which makes the generated result closer to the real data distribution. Finally, we show that the performance of VGCF is significantly better than the state-of-the-art recommendation methods on several real-world datasets. Weifeng Sun 0002, Shumiao Yu, Boxiang Dong |
GLOBECOM | 4 |
| 2020 | Towards Optimal System Deployment for Edge Computing: A Preliminary StudyabstractIn this preliminary study, we consider the server allocation problem for edge computing system deployment. Our goal is to minimize the average turnaround time of application requests/tasks, generated by all mobile devices/users in a geographical region. We consider two approaches for edge cloud deployment: the flat deployment, where all edge clouds co-locate with the base stations, and the hierarchical deployment, where edge clouds can also co-locate with other system components besides the base stations. In the flat deployment, we demonstrate that the allocation of edge cloud servers should be balanced across all the base stations, if the application request arrival rates at the base stations are equal to each other. We also show that the hierarchical deployment approach has great potentials in minimizing the system's average turnaround time. We conduct various simulation studies using the CloudSim Plus platform to verify our theoretical results. The collective findings trough theoretical analysis and simulation results will provide useful guidance in practical edge computing system deployment. Dawei Li 0002, Chigozie Asikaburu, Boxiang Dong, Huan Zhou 0002, Sadoon Azizi |
ICCCN | 3 |
| 2020 | Flat and hierarchical system deployment for edge computing systems
En Wang, Dawei Li 0002, Boxiang Dong, Huan Zhou 0002, Michelle Zhu |
Future Gener. Comput. Syst. | 3 |
| 2020 | Authenticated Outlier Mining for Outsourced DatabasesabstractThe Data-Mining-as-a-Service (DMaS) paradigm is becoming the focus of research, as it allows the data owner (client) who lacks expertise and/or computational resources to outsource their data and mining needs to a third-party service provider (server). Outsourcing, however, raises some issues about result integrity: how could the client verify the mining results returned by the server are both sound and complete? In this paper, we focus on outlier mining, an important mining task. Previous verification techniques use an authenticated data structure (ADS) for correctness authentication, which may incur much space and communication cost. In this paper, we propose a novel solution that returns a probabilistic result integrity guarantee with much cheaper verification cost. The key idea is to insert a set of artificial records (ARs) into the dataset, from which it constructs a set of artificial outliers (AOs) and artificial non-outliers (ANOs). The AOs and ANOs are used by the client to detect any incomplete and/or incorrect mining results with a probabilistic guarantee. The main challenge that we address is how to construct ARs so that they do not change the (non-)outlierness of original records, while guaranteeing that the client can identify ANOs and AOs without executing mining. Furthermore, we build a strategic game and show that a Nash equilibrium exists only when the server returns correct outliers. Our implementation and experiments demonstrate that our verification solution is efficient and lightweight. Boxiang Dong, Wendy Hui Wang, Anna Monreale, Dino Pedreschi, Fosca Giannotti, Wenge Guo |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2019 | MACCA: A SDN Based Collaborative Classification Algorithm for QoS Guaranteed Transmission on IoT
Weifeng Sun 0002, Zun Wang 0005, Boxiang Dong |
ADMA | 4 |
| 2018 | Pragmatics and Semantics to Connect Specific Local Laws with Public ReactionsabstractThis paper proposes an approach called TOLCS (Tweet Ordinance Linkage by Commonsense and Semantics) which helps connect specific ordinances or local laws to pertinent tweets expressing public reactions on them. TOLCS incorporates pragmatic aspects by commonsense knowledge, and semantic aspects by domain knowledge along with text similarity methods. It uses a blocking mechanism to reduce sample space for efficiently processing big data on ordinances and tweets. Manish Puri, Aparna S. Varde, Boxiang Dong |
IEEE BigData | 3 |
| 2018 | Truth Inference on Sparse Crowdsourcing Data with Local Differential PrivacyabstractCrowdsourcing is a new problem-solving paradigm for tasks that are difficult for computers but easy for humans. Since the answers collected from the recruited participants (workers) may contain sensitive information, crowdsourcing raises serious privacy concerns. In this paper, we investigate the problem of protecting user privacy under local differential privacy (LDP), where individual workers randomize their answers independently and send the perturbed answers to the task requester. The utility goal is to ensure high accuracy of the inferred true answers (i.e., truth) from the perturbed data. One of the challenges of LDP perturbation is the sparsity of worker answers (i.e., each worker only answers a small number of tasks). Simple extension of existing approaches (e.g., Laplace perturbation and randomized response) may incur large errors in truth inference on sparse data. Thus we design a new matrix factorization (MF) algorithm under LDP that addresses the trade-off between privacy and utility (i.e., accuracy of truth inference). We prove that our MF algorithm can provide both LDP guarantee and small error of truth inference, regardless of the sparsity of worker answers. We perform extensive experiments on real-world and synthetic datasets and demonstrate that the MF algorithm performs better than the existing LDP algorithms on sparse crowdsourcing data. Haipei Sun, Boxiang Dong, Wendy Hui Wang, Ting Yu 0001, Zhan Qin |
IEEE BigData | 2 |
| 2018 | Sensitive Task Assignments in Crowdsourcing Markets with Colluding WorkersabstractCrowdsourcing has raised several security concerns. One of the concerns is how to assign sensitive tasks in the crowdsourcing market, especially when there are colluding participants in crowdsourcing. In this paper, we consider adversarial colluding participants who intend to extract sensitive data by exchanging information. We design a 3-step sensitive task assignment method: (1) the collusion estimation step that quantifies the workers' pairwise collusion probability by estimating answer truth based on their responses; (2) the worker selection step that executes a heuristic sampling-based approach to select the fewest workers whose collusion probability satisfies the given security requirement; and (3) the task partitioning step that splits the sensitive information among the selected workers. We perform an extensive set of experiments on both real-world and synthetic datasets. The results demonstrate the accuracy and efficiency of our method. Haipei Sun, Boxiang Dong, Bo Zhang 0051, Wendy Hui Wang, Murat Kantarcioglu |
ICDE | 2 |
| 2018 | AssureMR: Verifiable SQL Execution on MapReduceabstractWe design AssureMR, a system that supports efficient verification of SQL Selection-GroupBy-Aggregation (SGA) query evaluation on an untrusted MapReduce system. AssureMR does not rely on a centralized trusted party to construct the authentication data structure (ADS). Instead, AssureMR allows the untrusted mappers/reducers to construct ADS. AssureMR provides the following verification functionality: (1) correctness verification of ADS; (2) correctness verification of intermediate query results by individual mapper; and (3) correctness verification of final query results by reducers. Our experimental results demonstrate the efficiency and effectiveness of AssureMR. Bo Zhang 0051, Boxiang Dong, Wendy Hui Wang |
ICDE | 2 |
| 2018 | Secure partial encryption with adversarial functional dependency constraints in the database-as-a-service model
Boxiang Dong, Wendy Hui Wang |
Data Knowl. Eng. | 1 |
| 2018 | Cost-efficient Data Acquisition on Online Data Marketplaces for Correlation AnalysisabstractIncentivized by the enormous economic profits, the data marketplace platform has been proliferated recently. In this paper, we consider the data marketplace setting where a data shopper would like to buy data instances from the data marketplace for correlation analysis of certain attributes. We assume that the data in the marketplace is dirty and not free. The goal is to find the data instances from a large number of datasets in the marketplace whose join result not only is of high-quality and rich join informativeness, but also delivers the best correlation between the requested attributes. To achieve this goal, we design DANCE, a middleware that provides the desired data acquisition service. DANCE consists of two phases: (1) In the off-line phase, it constructs a two-layer join graph from samples. The join graph includes the information of the datasets in the marketplace at both schema and instance levels; (2) In the online phase, it searches for the data instances that satisfy the constraints of data quality, budget, and join informativeness, while maximizing the correlation of source and target attribute sets. We prove that the complexity of the search problem is NP-hard, and design a heuristic algorithm based on Markov chain Monte Carlo (MCMC). Experiment results on two benchmark and one real datasets demonstrate the efficiency and effectiveness of our heuristic data acquisition algorithm. Haipei Sun, Boxiang Dong, Wendy Hui Wang |
Proc. VLDB Endow. | 3 |
| 2017 | Efficient Discovery of Abnormal Event Sequences in Enterprise Security SystemsabstractIntrusion detection system (IDS) is an important part of enterprise security system architecture. In particular, anomaly-based IDS has been widely applied to detect single abnormal process events that deviate from the majority. However, intrusion activity usually consists of a series of low-level heterogeneous events. The gap between low-level process events and high-level intrusion activities makes it particularly challenging to identify process events that are truly involved in a real malicious activity, and especially considering the massive 'noisy' events filling the event sequences. Hence, the existing work that focus on detecting single events can hardly achieve high detection accuracy. In this work, we formulate a novel problem in intrusion detection - suspicious event sequence discovery, and propose GID, an efficient graph-based intrusion detection technique that can identify abnormal event sequences from massive heterogeneous process traces with high accuracy. We fully implement GID and deploy it into a real-world enterprise security system, and it greatly helps detect the advanced threats and optimize the incident response. Executing GID on both static and streaming data shows that GID is efficient (processes about 2 million records per minute) and accurate for intrusion detection. Boxiang Dong, Zhengzhang Chen, Wendy Hui Wang, Lu-An Tang, Kai Zhang 0001, Zhichun Li |
CIKM | 1 |
| 2017 | Budget-Constrained Result Integrity Verification of Outsourced Data Mining Computations
Bo Zhang 0051, Boxiang Dong, Wendy Hui Wang |
DBSec | 2 |
| 2017 | Pairwise Ranking Aggregation by Non-interactive Crowdsourcing with Budget ConstraintsabstractCrowdsourced ranking algorithms ask the crowd to compare the objects and infer the full ranking based on the crowdsourced pairwise comparison results. In this paper, we consider the setting in which the task requester is equipped with a limited budget that can afford only a small number of pairwise comparisons. To make the problem more complicated, the crowd may return noisy comparison answers. We propose an approach to obtain a good-quality full ranking from a small number of pairwise preferences in two steps, namely task assignment and result inference. In the task assignment step, we generate pairwise comparison tasks that produce a full ranking with high probability. In the result inference step, based on the transitive property of pairwise comparisons and truth discovery, we design an efficient heuristic algorithm to find the best full ranking from the potentially conflictive pairwise preferences. The experiment results demonstrate the effectiveness and efficiency of our approach. Changjiang Cai, Haipei Sun, Boxiang Dong, Bo Zhang 0051, Ting Wang 0006, Wendy Hui Wang |
ICDCS | 3 |
| 2017 | Frequency-Hiding Dependency-Preserving Encryption for Outsourced DatabasesabstractThe cloud paradigm enables users to outsource their data to computationally powerful third-party service providers for data management. Many data management tasks rely on the data dependency in the outsourced data. This raises an important issue of how the data owner can protect the sensitive information in the outsourced data while preserving the data dependency. In this paper, we consider functional dependency (FD), an important type of data dependency. Although simple deterministic encryption schemes can preserve FDs, they may be vulnerable against the frequency analysis attack. We design a frequency hiding, FD-preserving probabilistic encryption scheme, named F2, that enables the service provider to discover the FDs from the encrypted dataset. We consider two attacks, namely the frequency analysis (FA) attack and the FD-preserving chosen plaintext attack (FCPA), and show that the F2 encryption scheme can defend against both attacks with formal provable guarantee. Our empirical study demonstrates the efficiency and effectiveness of F2, as well as its security against both FA and FCPA attacks. Boxiang Dong, Wendy Hui Wang |
ICDE | 1 |
| 2016 | Trust-but-Verify: Verifying Result Correctness of Outsourced Frequent Itemset Mining in Data-Mining-As-a-Service ParadigmabstractCloud computing is popularizing the computing paradigm in which data is outsourced to a third-party service provider (server) for data mining. Outsourcing, however, raises a serious security issue: how can the client of weak computational power verify that the server returned correct mining result? In this paper, we focus on the specific task of frequent itemset mining. We consider the server that is potentially untrusted and tries to escape from verification by using its prior knowledge of the outsourced data. We propose efficient probabilistic and deterministic verification approaches to check whether the server has returned correct and complete frequent itemsets. Our probabilistic approach can catch incorrect results with high probability, while our deterministic approach measures the result correctness with 100 percent certainty. We also design efficient verification methods for both cases that the data and the mining setup are updated. We demonstrate the effectiveness and efficiency of our methods using an extensive set of empirical results on real datasets. Boxiang Dong, Wendy Hui Wang |
IEEE Trans. Serv. Comput. | 1 |
| 2014 | PraDa: Privacy-preserving Data-Deduplication-as-a-ServiceabstractThe data-cleaning-as-a-service (DCaS) paradigm enables users to outsource their data and data cleaning needs to computationally powerful third-party service providers. It raises several security issues. One of the issues is how the client can protect the private information in the outsourced data. In this paper, we focus on data deduplication as the main data cleaning task, and design two efficient privacy-preserving data-deduplication methods for the DCaS paradigm. We analyze the robustness of our two methods against the attacks that exploit the auxiliary frequency distribution and the knowledge of the encoding algorithms. Our empirical study demonstrates the efficiency and effectiveness of our privacy preserving approaches. Boxiang Dong, Wendy Hui Wang |
CIKM | 1 |
| 2013 | Result Integrity Verification of Outsourced Frequent Itemset Mining
Boxiang Dong, Wendy Hui Wang |
DBSec | 1 |
| 2013 | Integrity Verification of Outsourced Frequent Itemset Mining with Deterministic GuaranteeabstractIn this paper, we focus on the problem of result integrity verification for outsourcing of frequent item set mining. We design efficient cryptographic approaches that verify whether the returned frequent item set mining results are correct and complete with deterministic guarantee. The key of our solution is that the service provider constructs cryptographic proofs of the mining results. Both correctness and completeness of the mining results are measured against the proofs. We optimize the verification by minimizing the number of proofs. Our empirical study demonstrates the efficiency and effectiveness of the verification approaches. Boxiang Dong, Wendy Hui Wang |
ICDM | 1 |