VLDB 2026 Research / reviewers in the wild / expert
Chaoqi Yang
dblp:224/2555
· DBLP profile ↗
24ranked-venue papers
12as first author
18since 2021 · last 2025
0000-0002-5017-6114ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 11 first-author · 14 since 2021Databases, data management, data science and information retrieval · 11 · 5 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Computer networks · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | RHealth: A R Toolkit for Deep Learning in HealthcareabstractMachine learning for electronic health records (EHR) is advancing rapidly and already underpins risk stratification, readmission and mortality prediction, and decision support, yet reliable translation still stalls on fragmented data pipelines, inconsistent medical-code handling, and hard-to-reproduce eval-uation-barriers that especially hinder R-centric clinical teams. Despite impressive methodological gains in temporal modeling, attention mechanisms, and strong classical baselines, most turnkey toolchains live in Python; as a result, many healthcare researchers and clinical data scientists working in R lack a single, integrated path from raw multi-table EHR to calibrated, auditable models. We address this gap with RHealth, an open-source, R-native toolkit that plays the role of an end-to-end conductor: from data harmonization and medical-code normal-ization to task specification, model training, and standardized reporting. Concretely, RHealth provides adapters for widely used public datasets (e.g., MIMIC-III/IV, eICU), utilities to traverse and map ICD-9/10 and CCS codes, task templates for common outcomes (mortality, 30-day readmission, length of stay), and a modeling stack that-at this development stage-supports standard recurrent baselines (e.g., RNN) and offers an extensible interface for user-defined architectures under active development, all evaluated with reproducible splits, AUROC/AUPRC, and cali-bration diagnostics. By packaging the full pipeline-from data to evaluation granularity-into modular, composable components, RHealth lowers the entry barrier for R users, reduces “glue code,” and promotes transparent, people-centric experimentation that can also serve as a trustworthy upstream substrate for LLM-enabled applications. To our knowledge, it is among the first comprehensive, integrated deep-learning toolkits for EHR in the R ecosystem. The code and documentation will be released after the double-blind review process. Ji Song, Zhixia Ren, Zhenbang Wu, John Wu, Chaoqi Yang, Yinghao Zhu, Wen Tang 0001, Jimeng Sun 0001, Ewen M. Harrison, Liantao Ma |
BIBM | 5 |
| 2024 | Topic-Oriented Open Relation Extraction with A Priori Seed GenerationabstractThe field of open relation extraction (ORE) has recently observed significant advancement thanks to the growing capability of large language models (LLMs).Nevertheless, challenges persist when ORE is performed on specific topics.Existing methods give suboptimal results in five dimensions: factualness, topic relevance, informativeness, coverage, and uniformity.To improve topic-oriented ORE, we propose a zero-shot approach called Pri-ORE: Open Relation Extraction with a Priori seed generation.PriORE leverages the builtin knowledge of LLMs to maintain a dynamic seed relation dictionary for the topic.The dictionary is initialized by seed relations generated from topic-relevant entity types and expanded during contextualized ORE.PriORE then reduces the randomness in generative ORE by converting it to a more robust relation classification task.Experiments show the approach empowers better topic-oriented control over the generated relations and thus improves ORE performance along the five dimensions, especially on specialized and narrow topics. Dimension Extracted Relation Explanation Factualness gives toWrong given text and topic Relevance is exposed to Correct given text but not directly relevant to topic Informativeness form compound with Correct but conceptually too general given text and topic Coverage form powdery magnesium oxide Correct but too specific, covering too few instances Uniformity {be oxidized by, give electrons to, . . .} Multiple correct expressions extracted for the same relation Linyi Ding, Jinfeng Xiao, Sizhe Zhou, Chaoqi Yang, Jiawei Han 0001 |
EMNLP | 4 |
| 2024 | Fine-grained Control of Generative Data Augmentation in IoT SensingabstractInternet of Things (IoT) sensing models often suffer from overfitting due to data distribution shifts between training dataset and real-world scenarios. To address this, data augmentation techniques have been adopted to enhance model robustness by bolstering the diversity of synthetic samples within a defined vicinity of existing samples. This paper introduces a novel paradigm of data augmentation for IoT sensing signals by adding fine-grained control to generative models. We define a metric space with statistical metrics that capture the essential features of the short-time Fourier transformed (STFT) spectrograms of IoT sensing signals. These metrics serve as strong conditions for a generative model, enabling us to tailor the spectrogram characteristics in the time-frequency domain according to specific application needs. Furthermore, we propose a set of data augmentation techniques within this metric space to create new data samples. Our method is evaluated across various generative models, datasets, and downstream IoT sensing models. The results demonstrate that our approach surpasses the conventional transformation-based data augmentation techniques and prior generative data augmentation models. Tianshi Wang 0002, Qikai Yang, Ruijie Wang 0004, Dachun Sun, Jinyang Li 0004, Yizhuo Chen, Yigong Hu, Chaoqi Yang, Tomoyoshi Kimura, Denizhan Kara, Tarek F. Abdelzaher |
NeurIPS | 8 |
| 2023 | ManyDG: Many-domain Generalization for Healthcare Applications
Chaoqi Yang, M. Brandon Westover, Jimeng Sun 0001 |
ICLR | 1 |
| 2023 | PyHealth: A Deep Learning Toolkit for Healthcare ApplicationsabstractDeep learning (DL) has emerged as a promising tool in healthcare applications. However, the reproducibility of many studies in this field is limited by the lack of accessible code implementations and standard benchmarks. To address the issue, we create PyHealth, a comprehensive library to build, deploy, and validate DL pipelines for healthcare applications. PyHealth supports various data modalities, including electronic health records (EHRs), physiological signals, medical images, and clinical text. It offers various advanced DL models and maintains comprehensive medical knowledge systems. The library is designed to support both DL researchers and clinical data scientists. Upon the time of writing, PyHealth has received 633 stars, 130 forks, and 15k+ downloads in total on GitHub. Chaoqi Yang, Zhenbang Wu, Patrick Jiang, Zhen Lin 0001, Benjamin P. Danek, Jimeng Sun 0001 |
KDD | 1 |
| 2023 | BIOT: Biosignal Transformer for Cross-data Learning in the WildabstractBiological signals, such as electroencephalograms (EEG), play a crucial role in numerous clinical applications, exhibiting diverse data formats and quality profiles. Current deep learning models for biosignals (based on CNN, RNN, and Transformers) are typically specialized for specific datasets and clinical settings, limiting their broader applicability. This paper explores the development of a flexible biosignal encoder architecture that can enable pre-training on multiple datasets and fine-tuned on downstream biosignal tasks with different formats.
To overcome the unique challenges associated with biosignals of various formats, such as mismatched channels, variable sample lengths, and prevalent missing val- ues, we propose Biosignal Transformer (BIOT). The proposed BIOT model can enable cross-data learning with mismatched channels, variable lengths, and missing values by tokenizing different biosignals into unified "sentences" structure. Specifically, we tokenize each channel separately into fixed-length segments containing local signal features and then rearrange the segments to form a long "sentence". Channel embeddings and relative position embeddings are added to each segment (viewed as "token") to preserve spatio-temporal features.
The BIOT model is versatile and applicable to various biosignal learning settings across different datasets, including joint pre-training for larger models. Comprehensive evaluations on EEG, electrocardiogram (ECG), and human activity sensory signals demonstrate that BIOT outperforms robust baselines in common settings and facilitates learning across multiple datasets with different formats. Using CHB-MIT seizure detection task as an example, our vanilla BIOT model shows 3% improvement over baselines in balanced accuracy, and the pre-trained BIOT models (optimized from other data sources) can further bring up to 4% improvements. Our repository is public at https://github.com/ycq091044/BIOT. Chaoqi Yang, M. Brandon Westover, Jimeng Sun 0001 |
NeurIPS | 1 |
| 2023 | Multi-faceted analysis and prediction for the outbreak of pediatric respiratory syncytial virusabstractOBJECTIVES: Respiratory syncytial virus (RSV) is a significant cause of pediatric hospitalizations. This article aims to utilize multisource data and leverage the tensor methods to uncover distinct RSV geographic clusters and develop an accurate RSV prediction model for future seasons. MATERIALS AND METHODS: This study utilizes 5-year RSV data from sources, including medical claims, CDC surveillance data, and Google search trends. We conduct spatiotemporal tensor analysis and prediction for pediatric RSV in the United States by designing (i) a nonnegative tensor factorization model for pediatric RSV diseases and location clustering; (ii) and a recurrent neural network tensor regression model for county-level trend prediction using the disease and location features. RESULTS: We identify a clustering hierarchy of pediatric diseases: Three common geographic clusters of RSV outbreaks were identified from independent sources, showing an annual RSV trend shifting across different US regions, from the South and Southeast regions to the Central and Northeast regions and then to the West and Northwest regions, while precipitation and temperature were found as correlative factors with the coefficient of determination R2≈0.5, respectively. Our regression model accurately predicted the 2022-2023 RSV season at the county level, achieving R2≈0.3 mean absolute error MAE < 0.4 and a Pearson correlation greater than 0.75, which significantly outperforms the baselines with P-values <.05. CONCLUSION: Our proposed framework provides a thorough analysis of RSV disease in the United States, which enables healthcare providers to better prepare for potential outbreaks, anticipate increased demand for services and supplies, and save more lives with timely interventions. Chaoqi Yang, Lucas Glass, Adam R. Cross, Jimeng Sun 0001 |
J. Am. Medical Informatics Assoc. | 1 |
| 2022 | Semi-supervised Hypergraph Node Classification on Hypergraph Line ExpansionabstractPrevious hypergraph expansions are solely carried out on either vertex level or hyperedge level, thereby missing the symmetric nature of data co-occurrence, and resulting in information loss. To address the problem, this paper treats vertices and hyperedges equally and proposes a new hypergraph expansion named the line expansion(LE) for hypergraphs learning. The new expansion bijectively induces a homogeneous structure from the hypergraph by modeling vertex-hyperedge pairs. Our proposal essentially reduces the hypergraph to a simple graph, which enables the existing graph learning algorithms to work seamlessly with the higher-order structure. We further prove that our line expansion is a unifying framework over various hypergraph expansions. We evaluate the proposed LE on five hypergraph datasets in terms of the hypergraph node classification task. The results show that our method could achieve at least 2% accuracy improvement over the best baseline consistently. Chaoqi Yang, Ruijie Wang 0004, Shuochao Yao, Tarek F. Abdelzaher |
CIKM | 1 |
| 2022 | GOCPT: Generalized Online Canonical Polyadic Tensor Factorization and CompletionabstractLow-rank tensor factorization or completion is well-studied and applied in various online settings, such as online tensor factorization (where the temporal mode grows) and online tensor completion (where incomplete slices arrive gradually). However, in many real-world settings, tensors may have more complex evolving patterns: (i) one or more modes can grow; (ii) missing entries may be filled; (iii) existing tensor elements can change. Existing methods cannot support such complex scenarios. To fill the gap, this paper proposes a Generalized Online Canonical Polyadic (CP) Tensor factorization and completion framework (named GOCPT) for this general setting, where we maintain the CP structure of such dynamic tensors during the evolution. We show that existing online tensor factorization and completion setups can be unified under the GOCPT framework. Furthermore, we propose a variant, named GOCPTE, to deal with cases where historical tensor elements are unavailable (e.g., privacy protection), which achieves similar fitness as GOCPT but with much less computational cost. Experimental results demonstrate that our GOCPT can improve fitness by up to 2.8% on the JHU Covid data and 9.2% on a proprietary patient claim dataset over baselines. Our variant GOCPTE shows up to 1.2% and 5.5% fitness improvement on two datasets with about 20% speedup compared to the best model. Chaoqi Yang, Cheng Qian 0001, Jimeng Sun 0001 |
IJCAI | 1 |
| 2022 | KLAttack: Towards Adversarial Attack and Defense on Neural Dependency Parsing ModelsabstractAlthough neural language models achieve great performance on many Natural Language Processing tasks, they suffer from various adversarial attacks. Previous works mainly focus on semantic adversarial examples, which have similar semantics to the original sentences, while syntactic adversarial attacks against the dependency parsing task are still in an early stage of research. In this paper, we propose a novel method KLAttack, crafting word-level adversarial examples to attack neural-network-based dependency parsing models. Specifically, we retrieve the class probabilities from the victim dependency parsing model and compute the KL divergence by masking every word in a sentence. Then we use pre-trained language models and reference parsers to generate candidates for substitution. Experiments on the English Penn Treebank (PTB) dataset show that our method improves the attack success rate against Deep Biaffine Parser by up to 13.04% compared with previous related studies. Based on KLAttack, we further propose Syntax-Aware Transformer for Input Reconstruction, a denoiser to recover the original sentences from the adversarial examples. Trained adversarially with successfully attacked sentences from KLAttack, we enhance the robustness of the dependency parsing models by concatenating the denoiser ahead of the victim models. Yutao Luo, Menghua Lu, Chaoqi Yang, Gongshen Liu, Shi-Lin Wang |
IJCNN | 3 |
| 2022 | SECT: A Successively Conditional Transformer for Controllable Paraphrase GenerationabstractParaphrase generation has consistently been a challenging area in the field of NLP. Despite the considerable achievements made by previous work, existing methods lack a flexible way to include multiple controllable attributes to enhance the diversity of paraphrased sentences. To overcome this challenge, we propose a Successively Conditional Transformer (SECT) to tackle this task. SECT is based on a combination of Conditional Variational AutoEncoder (CVAE) and Transformer framework to generate diversified words. More specifically, our SECT deploys multi-head attention and memory gate mechanism to keep the interaction between each of the attributes and the corresponding encoder layer hidden state. To address the problem of absorbing flexible attributes, we apply a successive structure to our SECT, which enables the framework to couple the CVAE latent variables with the encoder layer hidden states progressively. In addition, our SECT is trained by minimizing a tailor-designed loss for producing paraphrased sentences as required. Finally, we conduct extensive experiments to substantiate the validity and effectiveness of our proposed model. The results show that SECT significantly outperforms the existing state-of-the-art approaches and generates more diverse paraphrased sentences. Tang Xue, Yuran Zhao, Chaoqi Yang, Gongshen Liu |
IJCNN | 3 |
| 2022 | ATD: Augmenting CP Tensor Decomposition by Self SupervisionabstractTensor decompositions are powerful tools for dimensionality reduction and feature interpretation of multidimensional data such as signals. Existing tensor decomposition objectives (e.g., Frobenius norm) are designed for fitting raw data under statistical assumptions, which may not align with downstream classification tasks. In practice, raw input tensor can contain irrelevant information while data augmentation techniques may be used to smooth out class-irrelevant noise in samples. This paper addresses the above challenges by proposing augmented tensor decomposition (ATD), which effectively incorporates data augmentations and self-supervised learning (SSL) to boost downstream classification. To address the non-convexity of the new augmented objective, we develop an iterative method that enables the optimization to follow an alternating least squares (ALS) fashion. We evaluate our proposed ATD on multiple datasets. It can achieve 0.8%~2.5% accuracy gain over tensor-based baselines. Also, our ATD model shows comparable or better performance (e.g., up to 15% in accuracy) over self-supervised and autoencoder baselines while using less than 5% of learnable parameters of these baseline models. Chaoqi Yang, Cheng Qian 0001, Cao Xiao, M. Brandon Westover, Edgar Solomonik, Jimeng Sun 0001 |
NeurIPS | 1 |
| 2022 | Multi-Objective Actor-Critics for Real-Time Bidding in Display Advertising
Haolin Zhou, Chaoqi Yang, Xiaofeng Gao 0001, Gongshen Liu, Guihai Chen |
ECML/PKDD (4) | 2 |
| 2022 | Using Survival Theory in Early Pattern Detection for Viral CascadesabstractIn recent years, social networks have developed rapidly and become an indispensable part of people’s everyday life. Many models try to predict whether some reshare cascades are going to be popular or not, but most of their performances are limited due to the lack of cascades’ information in the early stage. In this paper, we proposeEarly Pattern detection model for Outbreak Cascades(in abbreviation, EPOC) inspired by the survival theory. We use three features to predict cascades’ virality: retweet sequence, follower number sequence, and timestamps of the first tweet which includes both the static and dynamic characteristics of cascades. Utilizing the theory that distributions of both viral and non-viral cascades are Gaussian, we get the boundary between these two kinds of cascades with sufficient proof to testify its rationality. To detect the virality more precisely and earlier, based on hazard functions in the survival theory, we propose two different hazard ceilings to capture the bursting of the cascades. We also provide a series of numerical experiments to analyze impacts of different factors to performance of our model measured by three practical metrics. The results shows that our model could stably outperforms several state-of-art baselines. Xiaofeng Gao 0001, Xiaosong Jia, Chaoqi Yang, Guihai Chen |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2021 | Change Matters: Medication Change Prediction with Recurrent Residual NetworksabstractDeep learning is revolutionizing predictive healthcare, including recommending medications to patients with complex health conditions. Existing approaches focus on predicting all medications for the current visit, which often overlaps with medications from previous visits. A more clinically relevant task is to identify medication changes. In this paper, we propose a new recurrent residual networks, named MICRON, for medication change prediction. MICRON takes the changes in patient health records as input and learns to update a hid- den medication vector and the medication set recurrently with a reconstruction design. The medication vector is like the memory cell that encodes longitudinal information of medications. Unlike traditional methods that require the entire patient history for prediction, MICRON has a residual-based inference that allows for sequential updating based only on new patient features (e.g., new diagnoses in the recent visit), which is efficient. We evaluated MICRON on real inpatient and outpatient datasets. MICRON achieves 3.5% and 7.8% relative improvements over the best baseline in F1 score, respectively. MICRON also requires fewer parameters, which significantly reduces the training time to 38.3s per epoch with 1.5× speed-up. Chaoqi Yang, Cao Xiao, Lucas Glass, Jimeng Sun 0001 |
IJCAI | 1 |
| 2021 | SafeDrug: Dual Molecular Graph Encoders for Recommending Effective and Safe Drug CombinationsabstractMedication recommendation is an essential task of AI for healthcare. Existing works focused on recommending drug combinations for patients with complex health conditions solely based on their electronic health records. Thus, they have the following limitations: (1) some important data such as drug molecule structures have not been utilized in the recommendation process. (2) drug-drug interactions (DDI) are modeled implicitly, which can lead to sub-optimal results. To address these limitations, we propose a DDI-controllable drug recommendation model named SafeDrug to leverage drugs’ molecule structures and model DDIs explicitly. SafeDrug is equipped with a global message passing neural network (MPNN) module and a local bipartite learning module to fully encode the connectivity and functionality of drug molecules. SafeDrug also has a controllable loss function to control DDI level in the recommended drug combinations effectively. On a benchmark dataset, our SafeDrug is relatively shown to reduce DDI by 19.43% and improves 2.88% on Jaccard similarity between recommended and actually prescribed drug combinations over previous approaches. Moreover, SafeDrug also requires much fewer parameters than previous deep learning based approaches, leading to faster training by about 14% and around 2× speed-up in inference. Chaoqi Yang, Cao Xiao, Fenglong Ma, Lucas Glass, Jimeng Sun 0001 |
IJCAI | 1 |
| 2021 | MTC: Multiresolution Tensor Completion from Partial and Coarse ObservationsabstractExisting tensor completion formulation mostly relies on partial observations from a single tensor. However, tensors extracted from real-world data often are more complex due to: (i) Partial observation: Only a small subset of tensor elements are available. (ii) Coarse observation: Some tensor modes only present coarse and aggregated patterns (e.g., monthly summary instead of daily reports). In this paper, we are given a subset of the tensor and some aggregated/coarse observations (along one or more modes) and seek to recover the original fine-granular tensor with low-rank factorization. We formulate a coupled tensor completion problem and propose an efficient Multi-resolution Tensor Completion model (MTC) to solve the problem. Our MTC model explores tensor mode properties and leverages the hierarchy of resolutions to recursively initialize an optimization setup, and optimizes on the coupled system using alternating least squares. MTC ensures low computational and space complexity. We evaluate our model on two COVID-19 related spatio-temporal tensors. The experiments show that MTC could provide 65.20% and 75.79% percentage of fitness (PoF) in tensor completion with only 5% fine granular observations, which is 27.96% relative improvement over the best baseline. To evaluate the learned low-rank factors, we also design a tensor prediction task for daily and cumulative disease case predictions, where MTC achieves 50% in PoF and 30% relative improvements over the best baseline. Chaoqi Yang, Cao Xiao, Cheng Qian 0001, Edgar Solomonik, Jimeng Sun 0001 |
KDD | 1 |
| 2021 | Ranking User-Generated Content via Multi-Relational Graph ConvolutionabstractThe quality variance in user-generated content is a major bottleneck to serving communities on online platforms. Current content ranking methods primarily evaluate text and non-textual content features of each user post in isolation. In this paper, we demonstrate the utility of considering the implicit and explicit relational aspects across user content to assess their quality. First, we develop a modular platform-agnostic framework to represent the contrastive (or competing) and similarity-based relational aspects of user-generated content via independently induced content graphs. Second, we develop two complementary graph convolutional operators that enable feature contrast for competing content and feature smoothing/sharing for similar content. Depending on the edge semantics of each content graph, we embed its nodes via one of the above two mechanisms. We also show that our contrastive operator creates discriminative magnification across the embeddings of competing posts. Third, we show a surprising result-applying classical boosting techniques to combine final-layer embeddings across the content graphs significantly outperforms the typical stacking, fusion, or neighborhood embedding aggregation methods in graph convolutional architectures. We exhaustively validate our method via accepted answer prediction over fifty diverse Stack-Exchange (https://stackexchange.com/) websites with consistent relative gains of over 5% accuracy over state-of-the-art neural, multi-relational and textual baselines. Kanika Narang, Adit Krishnan, Junting Wang 0001, Chaoqi Yang, Hari Sundaram, Carolyn Sutter |
SIGIR | 4 |
| 2020 | Hierarchical Overlapping Belief Estimation by Structured Matrix FactorizationabstractMuch work on social media opinion polarization focuses on a flat categorization of stances (or orthogonal beliefs) of different communities from media traces. We extend in this work in two important respects. First, we detect not only points of disagreement between communities, but also points of agreement. In other words, we estimate community beliefs in the presence of overlap. Second, in lieu of flat categorization, we consider hierarchical belief estimation, where communities might be hierarchically divided. For example, two opposing parties might disagree on core issues, but within a party, despite agreement on fundamentals, disagreement might occur on further details. We call the resulting combined problem a hierarchical overlapping belief estimation problem. To solve it, this paper develops a new class of unsupervised Non-negative Matrix Factorization (NMF) algorithms, we call Belief Structured Matrix Factorization (BSMF). Our proposed unsupervised algorithm captures both the latent belief intersections and dissimilarities, as well as hierarchical structure. We discuss properties of the algorithm and evaluate it on both synthetic and real-world datasets. In the synthetic dataset, our model reduces error by 40%. In real Twitter traces, it improves accuracy by around 10%. The model also achieves 96.08% self-consistency in a sanity check. Chaoqi Yang, Jinyang Li 0004, Ruijie Wang 0004, Shuochao Yao, Huajie Shao, Dongxin Liu, Shengzhong Liu, Tianshi Wang 0002, Tarek F. Abdelzaher |
ASONAM | 1 |
| 2020 | Misinformation Detection and Adversarial Attack Cost Analysis in Directional Social NetworksabstractThis paper develops a novel detection system of possibly fake accounts on public social media, called FADE, that uses features based on group behaviors to identify suspicious groups. The work is motivated by the prospect of mitigating misinformation campaigns on social media, where malicious entities on directional social networks pose as credible sources and coordinate the spreading of highly corroborated false information. Instead of account-level detection, this paper aims to detect the very group activity that underlies misinformation campaigns; namely, the coordinated spreading of messages to boost (misinformation) visibility. The existing group detection methods group users into two clusters (fake or not) and directly produce clusters of fake accounts. Conversely, we group users into many clusters based on information propagation patterns and user features and then classify them. The benefit of multiple clusters is that we can detect suspicious behavior more easily from cluster-wide statistics. In order to improve clustering accuracy, we analyze and select the most important features for clustering based on Bayesian optimization instead of using all the features. Accordingly, similarity metrics are defined that allow clustering of individually plausible accounts in a manner that enables one to detect suspicious clusters of activity. Cluster-level features are then used to decide if the cluster is benign. We further explore the cost of adversarial attacks on our detection model. Evaluation results on Twitter data sets demonstrate that our proposed approach outperforms state-of-the-art baselines in detecting accounts created for information manipulation campaigns. In addition, we show that the cost of subverting detection (without reducing the effectiveness of the attacker's campaign) is high. Huajie Shao, Shuochao Yao, Andong Jing, Shengzhong Liu, Dongxin Liu, Tianshi Wang 0002, Jinyang Li 0004, Chaoqi Yang, Ruijie Wang 0004, Tarek F. Abdelzaher |
ICCCN | 8 |
| 2019 | Reinforcement Learning with Sequential Information Clustering in Real-Time BiddingabstractDisplay advertising is a billion dollar business which is the primary income of many companies. In this scenario, real-time bidding optimization is one of the most important problems, where the bids of ads for each impression are determined by an intelligent policy such that some global key performance indicators are optimized. Due to the highly dynamic bidding environment, many recent works try to use reinforcement learning algorithms to train the bidding agents. However, as the probability of the occurrence of a particular state is typically low and the state representation in current work lacks sequential information, the convergence speed and performance of deep reinforcement algorithms are disappointing. To tackle these two challenges in the real-time bidding scenario, we propose ClusterA3C, a novel Advantage Asynchronous Actor-Critic (A3C) variant integrated with a sequential information extraction scheme and a clustering based state aggregation scheme. We conduct extensive experiments to validate the proposed scheme on a real-world commercial dataset. Experimental results show that the proposed scheme outperforms the state of the art methods in terms of either performance or convergence speed. Chaoqi Yang, Xiaofeng Gao 0001, Liubin Wang, Guihai Chen |
CIKM | 2 |
| 2018 | Adversarial Training Model Unifying Feature Driven and Point Process Perspectives for Event Popularity PredictionabstractThis paper targets a general popularity prediction problem for event sequence, which has recently gained great attention due to its extensive applications in various domains. Feature driven method and point process method are two basic thinking paradigms to tackle the prediction problem, but both of them suffer from limitations. In this paper, we propose PreNets unifying the two thinking paradigms in an adversarial manner. On one side, feature driven model acts like a 'critic' who aims to discriminate the predicted popularity from the real one based on a set of temporal features from the sequence. On the other side, point process model acts like an 'interpreter' who recognizes the dynamic patterns in sequence to generate a predicted popularity that can fool the 'critic'. Through a Wasserstein learning based two-player game, the training loss of the 'critic' guides the 'interpreter' to better exploit the sequence patterns and enhance prediction, while the 'interpreter' pushes the 'critic' to select effective early features that helps discrimination. This mechanism enables the framework to absorb the advantages of both feature driven and point process methods. Empirical results show that PreNets achieves significant MAPE improvement for both Twitter cascade and Amazon review prediction. Qitian Wu, Chaoqi Yang, Xiaofeng Gao 0001, Paul Weng, Guihai Chen |
CIKM | 2 |
| 2018 | EPOC: A Survival Perspective Early Pattern Detection Model for Outbreak Cascades
Chaoqi Yang, Qitian Wu, Xiaofeng Gao 0001, Guihai Chen |
DEXA (1) | 1 |
| 2018 | EPAB: Early Pattern Aware Bayesian Model for Social Content Popularity PredictionabstractThe boom of information technology enables social platforms (like Twitter) to disseminate social content (like news) in an unprecedented rate, which makes early-stage prediction for social content popularity of great practical significance. However, most existing studies assume a long-term observation before prediction and suffer from limited precision for early-stage prediction due to insufficient observation. In this paper, we take a fresh perspective, and propose a novel early pattern aware Bayesian model. The early pattern representation, which stands for early time series normalized on future popularity, can address what we call early-stage indistinctiveness challenge. Then we use an expressive evolving function to fit the time series and estimate three interpretable coefficients characterizing temporal effect of observed series on future evolution. Furthermore, Bayesian network is leveraged to model the probabilistic relations among features, early indicators and early patterns. Experiments on three real-world social platforms (Twitter, Weibo and WeChat) show that under different evaluation metrics, our model outperforms other methods in early-stage prediction and possesses low sensitivity to observation time. Qitian Wu, Chaoqi Yang, Xiaofeng Gao 0001, Guihai Chen |
ICDM | 2 |