Yijian Bai

dblp:30/3122 · DBLP profile ↗
← Back
7ranked-venue papers
4as first author
1since 2021 · last 2021
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 6 · 4 first-author · 1 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
3 papers
Data mining · 65% Data stream processing · 35%
Network and information security
1 paper
Privacy and data protection · 77% Security and privacy of machine learning · 23%

Topics — the 9 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data mining
clustering
0.512021
Clustering for Private Interest-based Advertising · KDD 2021
Privacy and data protection › web privacy › online advertising privacy
privacy-preserving advertising
0.512021
Clustering for Private Interest-based Advertising · KDD 2021
Security and privacy of machine learning
federated learning
0.112021
Clustering for Private Interest-based Advertising · KDD 2021
Data stream processing
continuous query processing
0.122007
RFID Data Processing with a Data Stream Query Language · ICDE 2007
Optimizing Timestamp Management in Data Stream Management Systems · ICDE 2007
Data stream processing
complex event processing
0.112007
RFID Data Processing with a Data Stream Query Language · ICDE 2007
Data stream processing
stream query languages
0.112007
RFID Data Processing with a Data Stream Query Language · ICDE 2007
Data mining › temporal data mining
temporal event detection
0.112007
RFID Data Processing with a Data Stream Query Language · ICDE 2007
Data stream processing
operator scheduling
0.012007
Optimizing Timestamp Management in Data Stream Management Systems · ICDE 2007
Internet of things and sensor networks › RFID systems › RFID data management
RFID data processing
0.012007
RFID Data Processing with a Data Stream Query Language · ICDE 2007

Methods — techniques the papers use, named apart from their topics

privacy guarantees · 1.0hashing · 1.0clustering · 1.0temporal operators · 0.1sliding-window constructs · 0.1on-demand punctuation · 0.1heartbeat tuples · 0.1
YearPublicationVenuePosition
2021 Clustering for Private Interest-based Advertising
abstract
We study the problem of designing privacy-enhanced solutions for interest-based advertisement (IBA). IBA is a key component of the online ads ecosystem and provides a better ad experience to users. Indeed, IBA enables advertisers to show users impressions that are relevant to them. Nevertheless, the current way ad tech companies achieve this is by building detailed interest profiles for individual users. In this work we ask whether such fine grained personalization is required, and present mechanisms that achieve competitive performance while giving privacy guarantees to the end users. More precisely we present the first detailed exploration of how to implement Chrome's Federated Learning of Cohorts (FLoC) API. We define the privacy properties required for the API and evaluate multiple hashing and clustering algorithms discussing the trade-offs between utility, privacy, and ease of implementation.
Alessandro Epasto, Andrés Muñoz Medina, Steven Avery, Yijian Bai, Róbert Busa-Fekete, CJ Carey, David Guthrie, Subham Ghosh, James Ioannidis, Junyi Jiao, Jakub Lacki, Arne Mauser, Brian Milch, Vahab S. Mirrokni, Deepak Ravichandran, Max Spero, Yunting Sun, Umar Syed, Sergei Vassilvitskii
KDD4
2007 Optimizing Timestamp Management in Data Stream Management Systems
abstract
It has long been recognized that multi-stream operators, such as union and join, often have to wait idly in a temporarily blocked state, as a result of skews between the timestamps of their input streams. It has been shown that the injection of heartbeat information through punctuation tuples can alleviate this problem. In this paper, we propose and investigate more effective solutions that use timestamps generated on-demand to reactivate idle-waiting operators. We thus introduce a simple execution model that efficiently supports on-demand punctuation. Experiments show that response time and memory usage are reduced substantially by this approach.
Yijian Bai, Hetal Thakkar, Haixun Wang, Carlo Zaniolo
ICDE1
2007 RFID Data Processing with a Data Stream Query Language
abstract
RFID technology provides significant advantages over traditional object-tracking technologies and is increasingly adopted and deployed in real applications. RFID applications generate large volume of streaming data, which have to be automatically filtered, processed, and transformed into semantic data, and integrated into business applications. Indeed, RFID data are highly temporal, and RFID observations form complex temporal event patterns which can be very different for various RFID applications. Thus, it is desirable to have a general RFID data processing framework with a powerful language, for the end users to express a variety of queries on RFID data streams, as well as detecting complex events patterns. While data stream management systems (DSMSs) are emerging for optimized stream data processing, they usually lack the language construct support for temporal event detection. In this paper, we discuss a stream query language to provide comprehensive temporal event detection, through temporal operators and extension of sliding-window constructs. With the integration of temporal event detection, a DSMS has the capability to serve as a powerful system for RFID data processing.
Yijian Bai, Fusheng Wang 0001, Peiya Liu, Carlo Zaniolo, Shaorong Liu
ICDE1
2007 A System for Technology Based Assessment of Language and Literacy in Young Children: the Role of Multiple Information Sources
abstract
This paper describes the design and realization of an automatic system for assessing and evaluating the language and literacy skills of young children. This system was developed in the context of the TBALL (technology based assessment of language and literacy) project and aims at automatically assessing the English literacy skills of both native talkers of American English and Mexican-American children in grades K-2. The automatic assessments were carried out employing appropriate speech recognition and understanding techniques. In this paper, we describe the system focusing on the role of the multiple sources of information at our disposal. We present the content of the assessment system, discuss some issues in creating a child-friendly interface, and how to provide a suitable feedback to the teachers. In addition, we will discuss the different assessment modules and the different algorithms used for speech analysis.
Abeer Alwan, Yijian Bai, Matthew Black, Larry Casey, Matteo Gerosa, Margaret Heritage, Markus Iseli, Abe Kazemzadeh, Sungbok Lee, Shri Narayanan, Patti Price, Joseph Tepperman, Shizhen Wang
MMSP2
2007 Load Shedding in Classifying Multi-Source Streaming Data: A Bayes Risk Approach
abstract
Monitoring multiple streaming sources for collective decision making presents several challenges. First, streaming data are often of large volume, fast speed, and highly bursty nature. Second, it is impossible to offload classification decisions to individual data sources, each of which lacks full knowledge for the decision making. Hence, the central classifier responsible for decision making may be frequently overloaded. In this paper, we study intelligent load shedding for classifying multi-source data. We aim at maximizing classification quality under resource (CPU and bandwidth) constraints. We use a Markov model to predict the distribution of feature values over time. Then, leveraging Bayesian decision theory, we use Bayes risk analysis to model the variances among different data sources in their contributions to the classification quality. We adopt an Expected Observational Risk criterion to quantify the loss of classification quality due to load shedding, and propose a Best Feature First (BFF) algorithm that greedily minimizes such risk. The effectiveness of the approach proposed is confirmed by experiments.
Yijian Bai, Haixun Wang, Carlo Zaniolo
SDM1
2006 A data stream language and system designed for power and extensibility
abstract
By providing an integrated and optimized support for user-defined aggregates (UDAs), data stream management systems (DSMS) can achieve superior power and generality while preserving compatibility with current SQL standards. This is demonstrated by the Stream Mill system that, through is Expressive Stream Language (ESL), efficiently supports a wide range of applications - including very advanced ones such as data stream mining, streaming XML processing, time-series queries, and RFID event processing. ESL supports physical and logical windows (with optional slides and tumbles) on both built-in aggregates and UDAs, using a simple framework that applies uniformly to both aggregate functions written in an external procedural languages and those natively written in ESL. The constructs introduced in ESL extend the power and generality of DSMS, and are conducive to UDA-specific optimization and efficient execution as demonstrated by several experiments.
Yijian Bai, Hetal Thakkar, Haixun Wang, Chang Luo, Carlo Zaniolo
CIKM1
2006 Bridging Physical and Virtual Worlds: Complex Event Processing for RFID Data Streams
Fusheng Wang 0001, Shaorong Liu, Peiya Liu, Yijian Bai
EDBT4