VLDB 2026 Research / reviewers in the wild / expert
Xikui Wang
dblp:91/8426
· DBLP profile ↗
10ranked-venue papers
4as first author
3since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 7 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Activating big data: Optimizing subscription-driven analytics
Shahrzad Haji Amin Shirazi, Xikui Wang, Michael J. Carey 0001, Vassilis J. Tsotras |
Inf. Syst. | 2 |
| 2025 | Optimizing Big Active Data Management Systems
Shahrzad Haji Amin Shirazi, Xikui Wang, Michael J. Carey 0001, Vassilis J. Tsotras |
DOLAP | 2 |
| 2022 | Subscribing to big data at scaleabstractAbstract Today, data is being actively generated by a variety of devices, services, and applications. Such data is important not only for the information that it contains, but also for its relationships to other data and to interested users. Most existing Big Data systems focus on passively answering queries from users, rather than actively collecting data, processing it, and serving it to users. To satisfy both passive and active requests at scale, application developers need either to heavily customize an existing passive Big Data system or to glue one together with systems like Streaming Engines and Pub-sub services. Either choice requires significant effort and incurs additional overhead. In this paper, we present the BAD (Big Active Data) system as an end-to-end, out-of-the-box solution for this challenge. It is designed to preserve the merits of passive Big Data systems and introduces new features for actively serving Big Data to users at scale. We show the design and implementation of the BAD system, demonstrate how BAD facilitates providing both passive and active data services, investigate the BAD system’s performance at scale, and illustrate the complexities that would result from instead providing BAD-like services with a “glued” system. Xikui Wang, Michael J. Carey 0001, Vassilis J. Tsotras |
Distributed Parallel Databases | 1 |
| 2020 | Bridging BAD Islands: Declarative Data Sharing at ScaleabstractIn many Big Data applications today, information needs to be actively shared between systems managed by different organizations. To enable sharing Big Data at scale, developers would have to create dedicated server programs and glue together multiple Big Data systems for scalability. Developing and managing such glued data sharing services requires a significant amount of work from developers. In our prior work, we developed a Big Active Data (BAD) system for enabling Big Data subscriptions and analytics with millions of subscribers. Based on that, we introduce a new mechanism for enabling the sharing of Big Data at scale declaratively so that developers can easily create and provide data sharing services using declarative statements and can benefit from an underlying scalable infrastructure. We show our implementation on top of the BAD system, explain the data sharing data flow among multiple systems, and present a prototype system with experimental results. Xikui Wang, Michael J. Carey 0001, Vassilis J. Tsotras |
IEEE BigData | 1 |
| 2020 | BAD to the bone: Big Active Data at its core
Steven Jacobs, Xikui Wang, Michael J. Carey 0001, Vassilis J. Tsotras, Md. Yusuf Sarwar Uddin |
VLDB J. | 2 |
| 2019 | An IDEA: An Ingestion Framework for Data Enrichment in AsterixDBabstractBig Data today is being generated at an unprecedented rate from various sources such as sensors, applications, and devices, and it often needs to be enriched based on other reference information to support complex analytical queries. Depending on the use case, the enrichment operations can be compiled code, declarative queries, or machine learning models with different complexities. For enrichments that will be frequently used in the future, it can be advantageous to push their computation into the ingestion pipeline so that they can be stored (and queried) together with the data. In some cases, the referenced information may change over time, so the ingestion pipeline should be able to adapt to such changes to guarantee the currency and/or correctness of the enrichment results. In this paper, we present a new data ingestion framework that supports data ingestion at scale, enrichments requiring complex operations, and adaptiveness to reference data changes. We explain how this framework has been built on top of Apache AsterixDB and investigate its performance at scale under various workloads. Xikui Wang, Michael J. Carey 0001 |
Proc. VLDB Endow. | 1 |
| 2017 | A BAD Demonstration: Towards Big Active DataabstractNearly all of today's Big Data systems are passive in nature. We demonstrate our Big Active Data ("BAD") system, a scalable system that continuously and reliably captures Big Data and facilitates the timely and automatic delivery of new information to a large population of interested users as well as supporting analyses of historical information. We built our BAD project by extending an existing scalable, open-source BDMS (AsterixDB [1]) in this active direction. In this demonstration, we allow our audience to participate in an emergency notification application built on top of our BAD platform, and highlight its capabilities. Steven Jacobs, Md. Yusuf Sarwar Uddin, Michael J. Carey 0001, Vagelis Hristidis, Vassilis J. Tsotras, Nalini Venkatasubramanian, Syed Safir, Purvi Kaul, Xikui Wang, Mohiuddin Abdul Qader |
Proc. VLDB Endow. | 10 |
| 2016 | Query-Driven Approach to Face Clustering and TaggingabstractIn the era of big data, a traditional offline setting to processing image data is simply not tenable. We simply do not have the computational power to process every image with every possible tag; moreover, we will not have the manpower to clean up the potentially noisy results. In this paper, we introduce a query-driven approach to visual tagging, focusing on the application of face tagging and clustering. We integrate active learning with query-driven probabilistic databases. Rather than asking a user to provide manual labels so as to minimize the uncertainty of labels (face tags) across the entire data set, we ask the user to provide labels that minimize the uncertainty of his/her query result (e.g., "How many times did Bob and Jim appear together?"). We use a data-driven Gaussian process model of facial appearance to write the probabilistic estimates of facial identity into a probabilistic database, which can then support inference through query answering. Importantly, the database is augmented with contextual constraints (faces in the same image cannot be the same identity, while faces in the same track must be identical). Experiments on the real-world photo collections demonstrate the effectiveness of the proposed method. Liyan Zhang 0001, Xikui Wang, Dmitri V. Kalashnikov, Sharad Mehrotra, Deva Ramanan |
IEEE Trans. Image Process. | 2 |
| 2013 | Cross-media topic mining on wikipediaabstractAs a collaborative wiki-based encyclopedia, Wikipedia provides a huge amount of articles of various categories. In addition to their text corpus, Wikipedia also contains plenty of images which makes the articles more intuitive for readers to understand. To better organize these visual and textual data, one promising area of research is to jointly model the embedding topics across multi-modal data (i.e, cross-media) from Wikipedia. In this work, we propose to learn the projection matrices that map the data from heterogeneous feature spaces into a unified latent topic space. Different from previous approaches, by imposing the l1 regularizers to the projection matrices, only a small number of relevant visual/textual words are associated with each topic, which makes our model more interpretable and robust. Furthermore, the correlations of Wikipedia data in different modalities are explicitly considered in our model. The effectiveness of the proposed topic extraction algorithm is verified by several experiments conducted on real Wikipedia datasets. Xikui Wang, Yang Liu 0098, Fei Wu 0001 |
ACM Multimedia | 1 |
| 2013 | Integration of multi-feature fusion and dictionary learning for face recognition
Xikui Wang, Shu Kong |
Image Vis. Comput. | 2 |