Wei Zhao 0001

dblp:z/WeiZhao-1 · DBLP profile ↗
← Back
18ranked-venue papers in the field
0as first author
5since 2021 · last 2025
0000-0002-6268-2559ORCID · conflict

Domains — venue-derived; a paper can count in several

Database Systems & Data Management · 8Data Mining & Knowledge Discovery · 4Big Data, Cloud & Distributed Data Systems · 2Other / Interdisciplinary · 2Information Retrieval & Web Search · 1Knowledge Engineering, Semantic Web & Information Systems · 1
YearPublicationVenuePosition
2025 LOVO: Efficient Complex Object Query in Large-Scale Video Datasets
abstract
The widespread deployment of cameras has led to an exponential increase in video data, creating vast opportunities for applications such as traffic management and crime surveillance. However, querying specific objects from large-scale video datasets presents challenges, including (1) processing massive and continuously growing data volumes, (2) supporting complex query requirements, and (3) ensuring low-latency execution. Existing video analysis methods struggle with either limited adaptability to unseen object classes or suffer from high query latency. In this paper, we present LOVO, a novel system designed to efficiently handle compLex Object queries in large-scale VideO datasets. Agnostic to user queries, LOVO performs one-time feature extraction using pre-trained visual encoders, generating compact visual embeddings for key frames to build an efficient index. These visual embeddings, along with associated bounding boxes, are organized in an inverted multi-index structure within a vector database, which supports queries for any objects. During the query phase, LOVO transforms object queries to query embeddings and conducts fast approximate nearest-neighbor searches on the visual embeddings. Finally, a cross-modal rerank is performed to refine the results by fusing visual features with detailed textual features. Evaluation on real-world video datasets demonstrates that LOVO outperforms existing methods in handling complex queries, with near-optimal query accuracy and up to 85x lower search latency, while significantly reducing index construction costs. This system redefines the state-of-theart object query approaches in video analysis, setting a new benchmark for complex object queries with a novel, scalable, and efficient approach that excels in dynamic environments.
Yuxin Liu 0007, Yuezhang Peng, Hefeng Zhou, Jiong Lou, Chentao Wu, Wei Zhao 0001, Jie Li 0002
ICDE8
2023 Privacy Data Diffusion Modeling and Preserving in Online Social Network
abstract
With the ubiquity of social media, privacy leakage has become a urgent problemfor social media managers. Studying how the privacy information diffuses through social media has attracted much attention. As a prerequisite, modeling privacy information diffusion is important research. Current approaches for modeling information diffusion are not available for privacy information since they did not consider the propagation features of privacy information in social media. Thispaper discusses the problem of modeling privacy information in social media and its challenges. We first analyse the information diffusion paths in the basic parameters of complex network and the high-order structures. We find that the privacy information is different in propagation features and the size of star structures. Second, a new information diffusion model is illustrated to simulate the diffusion process of information in social media by considering the following three parameters: 1) the probability of users receiving this message, 2) the probability that users have a tendency to forward this message and 3) the interest the users hold for this message. Finally, a block mechanism is designed to congest the diffusion of privacy information in social media.
Xiangyu Hu 0006, Tianqing Zhu, Xuemeng Zhai, Hengming Wang, Wanlei Zhou 0001, Wei Zhao 0001
IEEE Trans. Knowl. Data Eng.6
2023 Privacy Data Propagation and Preservation in Social Media: A Real-World Case Study
abstract
Social media has become a ubiquitous tool for spreading news, messages, and generally allowing for communication between individuals. Hence, studying how our privacy information might also spread across social media is important research. To date, many studies have used information diffusion models to simulate and then examine how information flows through social networks. But these models are theoretical, and newsworthy information may not behave in the same way as privacy information, raising the question: Are the observed phenomena indicative of real privacy propagation? To explore this question, we assembled a dataset from Twitter comprising propagated information flows for both private and normal information. We then built a graph convolutional network to trace and classify differences in the way each type of information spreads throughout the platform. The results reveal that there are indeed key differences in the diffusion processes of the two types of information. More importantly, we design privacy-preserving methods to reduce the privacy propagation in social media.
Xiangyu Hu 0006, Tianqing Zhu, Xuemeng Zhai, Wanlei Zhou 0001, Wei Zhao 0001
IEEE Trans. Knowl. Data Eng.5
2022 Distributed adaptive fuzzy control for multi-agent systems with full state constraints and unmeasured states
Yuzhen Ma, Yan-Jun Liu 0003, Wei Zhao 0001, Jie Lan, Tongyu Xu, Lei Liu 0006
Inf. Sci.3
2021 Intervening Coupling Diffusion of Competitive Information in Online Social Networks
abstract
The vigorously rising of social media brings a new opportunity for information diffusion in online social networks. However, the existing models of information diffusion only consider the single information, such as rumor. What's more, most of intervention frameworks are modeled under the ideal circumstances without reality constraints. In this article, we propose a novel model of competitive information coupling diffusion to describe the complex process of information diffusion in online social networks. Especially, in order to intervene the process of competitive information coupling diffusion, we introduce three intervention strategies and propose an intervention framework. More importantly, we take the dynamic constraints into consideration such as the budget of intervention and current state of the system, and further propose the constrained intervention model. To reduce the system loss, we establish an optimal control problem with constraints to achieve the optimal allocation of intervention strategies over time and minimize the total loss. We theoretically prove the existence and uniqueness of the optimal solution of the problem, and derive the optimal control solution. Through the experiments, we verify the effectiveness of the model and analyze the efficiency of different intervention strategies about competitive information coupling diffusion with or without constraints, respectively. The results show that the collaborative intervention strategies can effectively impact the process of diffusion and get the minimum system loss. This article provides high realistic significance to the commercial marketing in online social networks.
Pengfei Wan 0002, Xiaoming Wang 0001, Xinyan Wang 0001, Liang Wang 0014, Yaguang Lin, Wei Zhao 0001
IEEE Trans. Knowl. Data Eng.6
2019 Time-Sync Video Tag Extraction Using Semantic Association Graph
abstract
Time-sync comments (TSCs) reveal a new way of extracting the online video tags. However, such TSCs have lots of noises due to users’ diverse comments, introducing great challenges for accurate and fast video tag extractions. In this article, we propose an unsupervised video tag extraction algorithm named Semantic Weight-Inverse Document Frequency (SW-IDF). Specifically, we first generate corresponding semantic association graph (SAG) using semantic similarities and timestamps of the TSCs. Second, we propose two graph cluster algorithms, i.e., dialogue-based algorithm and topic center-based algorithm, to deal with the videos with different density of comments. Third, we design a graph iteration algorithm to assign the weight to each comment based on the degrees of the clustered subgraphs, which can differentiate the meaningful comments from the noises. Finally, we gain the weight of each word by combining Semantic Weight (SW) and Inverse Document Frequency (IDF). In this way, the video tags are extracted automatically in an unsupervised way. Extensive experiments have shown that SW-IDF (dialogue-based algorithm) achieves 0.4210 F1-score and 0.4932 MAP (Mean Average Precision) in high-density comments, 0.4267 F1-score and 0.3623 MAP in low-density comments; while SW-IDF (topic center-based algorithm) achieves 0.4444 F1-score and 0.5122 MAP in high-density comments, 0.4207 F1-score and 0.3522 MAP in low-density comments. It has a better performance than the state-of-the-art unsupervised algorithms in both F1-score and MAP.
Wenmian Yang, Kun Wang 0005, Na Ruan, Wenyuan Gao, Weijia Jia 0001, Wei Zhao 0001, Yunyong Zhang
ACM Trans. Knowl. Discov. Data6
2011 Privacy-Preserving OLAP: An Information-Theoretic Approach
abstract
We address issues related to the protection of private information in Online Analytical Processing (OLAP) systems, where a major privacy concern is the adversarial inference of private information from OLAP query answers. Most previous work on privacy-preserving OLAP focuses on a single aggregate function and/or addresses only exact disclosure, which eliminates from consideration an important class of privacy breaches where partial information, but not exact values, of private data is disclosed (i.e., partial disclosure). We address privacy protection against both exact and partial disclosure in OLAP systems with mixed aggregate functions. In particular, we propose an information-theoretic inference control approach that supports a combination of common aggregate functions (e.g., COUNT, SUM, MIN, MAX, and MEDIAN) and guarantees the level of privacy disclosure not to exceed thresholds predetermined by the data owners. We demonstrate that our approach is efficient and can be implemented in existing OLAP systems with little modification. It also satisfies the simulatable auditing model and leaks no private information through query rejections. Through performance analysis, we show that compared with previous approaches, our approach provides more effective privacy protection while maintaining a higher level of query-answer availability.
Nan Zhang 0004, Wei Zhao 0001
IEEE Trans. Knowl. Data Eng.2
2008 Privacy Protection Against Malicious Adversaries in Distributed Information Sharing Systems
abstract
We address issues related to sharing information in a distributed system consisting of autonomous entities, each of which holds a private database. We consider threats from malicious adversaries that can deviate from the designated protocol and change their input databases. We classify malicious adversaries into two widely existing subclasses, namely weakly and strongly malicious adversaries, and propose protocols that can effectively and efficiently protect privacy against malicious adversaries.
Nan Zhang 0004, Wei Zhao 0001
IEEE Trans. Knowl. Data Eng.2
2006 A Real-Time and Reliable Approach to Detecting Traffic Variations at Abnormally High and Low Rates
Ming Li 0002, Shengquan Wang, Wei Zhao 0001
ATC3
2005 On Multiterminal Source Code Design
abstract
Multiterminal (MT) source coding refers to separate lossy encoding and joint decoding of multiple correlated sources. This paper presents two practical MT coding schemes under the same general framework of Slepian-Wolf coded quantization (SWCQ) for both direct and indirect quadratic Gaussian MT source coding problems with two encoders. The first asymmetric SWCQ scheme relies on quantization and Wyner-Ziv coding, and is implemented via source-splitting to achieve any point on the inner sum-rate bound for both direct and indirect MT coding problems. In the second symmetric SWCQ scheme, the two quantization outputs are compressed using multilevel symmetric Slepian-Wolf coding. This scheme is conceptually simpler and can potentially achieve most of the points on the inner sum-rate bound. Our practical designs employ trellis coded quantization, LDPC code based asymmetric Slepian-Wolf code, and arithmetic code and turbo code based symmetric Slepian-Wolf code. Simulation results show a gap of only 0.24-0.29 bit per sample away from the inner sum-rate bound for both direct and indirect MT coding problems.
Yang Yang 0003, Vladimir Stankovic 0001, Zixiang Xiong, Wei Zhao 0001
DCC4
2005 A new scheme on privacy-preserving data classification
abstract
We address privacy-preserving classification problem in a distributed system. Randomization has been the approach proposed to preserve privacy in such scenario. However, this approach is now proven to be insecure as it has been discovered that some privacy intrusion techniques can be used to reconstruct private information from the randomized data tuples. We introduce an algebraic-technique-based scheme. Compared to the randomization approach, our new scheme can build classifiers more accurately but disclose less private information. Furthermore, our new scheme can be readily integrated as a middleware with existing systems.
Nan Zhang 0004, Shengquan Wang, Wei Zhao 0001
KDD3
2005 Performance Measurements for Privacy Preserving Data Mining
Nan Zhang 0004, Wei Zhao 0001, Jianer Chen
PAKDD2
2005 Distributed Privacy Preserving Information Sharing
Nan Zhang 0004, Wei Zhao 0001
VLDB2
2004 Asymmetric Code Design for Remote Multiterminal Source Coding
abstract
Asymmetric code design for remote multiterminal source coding in the quadratic Gaussian case is presented in this paper. For remote multiterminal source coding of X, to achieve the minimum sum rate of the two independent encoders subject to a fidelity criterion d, theoretical bounds were derived independently. The main idea is to quantize the first observation Y/sub 1/ and apply Wyner-Ziv coding on Y/sub 2/ by using the quantized version of Y/sub 1/ as side information in an efficient asymmetric coding scheme. The practical code design gives results that are very close to the sum-rate bound.
Yang Yang 0003, Vladimir Stankovic 0001, Zixiang Xiong, Wei Zhao 0001
Data Compression Conference4
2004 Cardinality-based inference control in OLAP systems: an information theoretic approach
abstract
We address the inference control problem in data cubes with some data known to users through external knowledge. The goal of inference controls is to prevent exact values of sensitive data from being inferred through answers to online analytical processing (OLAP) queries. We present an information theoretic approach for cardinality-based inference control, which simply counts the number of cells that all queries have covered thus far to determine whether a new query should be answered. Compared to previous approaches in sum-only data cubes, our new approach has a more general framework (applies to MIN, MAX and SUM) and is more effective.
Nan Zhang 0004, Wei Zhao 0001, Jianer Chen
DOLAP2
2004 A New Scheme on Privacy Preserving Association Rule Mining
Nan Zhang 0004, Shengquan Wang, Wei Zhao 0001
PKDD3
2000 A Whole Correlation Structure of Asymptotically Self-Similar Traffic in Communication Networks
abstract
Recent experimental research has revealed that the nature of WWW traffic is self-similarity (M.E. Crovella and A. Bestavros, 1997). That is, the behavior of WWW traffic is well modeled by second-order self-similar processes with long-range dependence. A closed form of autocorrelation functions about asymptotically self-similar processes is presented. The verification shows that this form is best for real traffic data on an Ethernet.
Ming Li 0002, Weijia Jia 0001, Wei Zhao 0001
WISE3
2000 Statistical delay analysis on an ATM switch with self-similar input traffic
Joseph Kee-Yin Ng, Shibin Song, Wei Zhao 0001
Inf. Process. Lett.3