EDBT 2026 Demo / reviewers in the wild / expert
Yichao Zhou 0001
dblp:146/9862-1
· DBLP profile ↗
18ranked-venue papers
9as first author
9since 2021 · last 2025
0009-0003-8632-446XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 7 first-author · 7 since 2021Databases, data management, data science and information retrieval · 11 · 5 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SUMIE: A Synthetic Benchmark for Incremental Entity SummarizationabstractNo existing dataset adequately tests how well language models can incrementally update entity summaries – a crucial ability as these models rapidly advance. The Incremental Entity Summarization (IES) task is vital for maintaining accurate, up-to-date knowledge. To address this, we introduce , a fully synthetic dataset designed to expose real-world IES challenges. This dataset addresses issues like incorrect entity association and incomplete information, capturing real-world complexity by generating diverse attributes, summaries, and unstructured paragraphs with 99% alignment accuracy between generated summaries and paragraphs. Extensive experiments demonstrate the dataset’s difficulty – state-of-the-art LLMs struggle to update summaries with an F1 higher than 80.4%. We will open-source the benchmark and the evaluation metrics to help the community make progress on IES tasks. Eunjeong Hwang, Yichao Zhou 0001, Beliz Gunel, James B. Wendt, Sandeep Tata |
COLING | 2 |
| 2024 | FieldSwap: Data Augmentation for Effective Form-Like Document ExtractionabstractExtracting structured data from visually rich documents like invoices, receipts, financial statements, and tax forms is key to automating many business workflows. However, building extraction models in this domain often demands a large collection of high-quality training examples. To address this challenge, we introduce FieldSwap, a novel data augmentation technique specifically designed for such extraction problems. FieldSwap generates synthetic training examples by replacing key phrases indicative of one field with those corresponding to another. Our experiments on five diverse datasets demonstrate that incorporating FieldSwap-augmented data into the training process can enhance model performance by 1–11 F1 points, particularly when dealing with limited training data (10–100 documents). Additionally, we propose algorithms for automatically inferring key phrases from the training data. Our findings indicate that FieldSwap is effective regardless of whether key phrases are manually provided by human experts or inferred automatically. Jing Xie 0002, James B. Wendt, Yichao Zhou 0001, Seth Ebner, Sandeep Tata |
ICDE | 3 |
| 2023 | Selective Labeling: How to Radically Lower Data-Labeling Costs for Document Extraction ModelsabstractBuilding automatic extraction models for visually rich documents like invoices, receipts, bills, tax forms, etc. has received significant attention lately.A key bottleneck in developing extraction models for new document types is the cost of acquiring the several thousand highquality labeled documents that are needed to train a model with acceptable accuracy.In this paper, we propose selective labeling as a solution to this problem.The key insight is to simplify the labeling task to provide "yes/no" labels for candidate extractions predicted by a model trained on partially labeled documents.We combine this with a custom active learning strategy to find the predictions that the model is most uncertain about.We show through experiments on document types drawn from 3 different domains that selective labeling can reduce the cost of acquiring labeled data by 10× with a negligible loss in accuracy. Yichao Zhou 0001, James B. Wendt, Navneet Potti, Jing Xie 0002, Sandeep Tata |
EMNLP | 1 |
| 2023 | VRDU: A Benchmark for Visually-rich Document UnderstandingabstractUnderstanding visually-rich business documents to extract structured data and automate business workflows has been receiving attention both in academia and industry. Although recent multi-modal language models have achieved impressive results, we find that existing benchmarks do not reflect the complexity of real documents seen in industry. In this work, we identify the desiderata for a more comprehensive benchmark and propose one we call Visually Rich Document Understanding (VRDU). VRDU contains two datasets that represent several challenges: rich schema including diverse data types as well as hierarchical entities, complex templates including tables and multi-column layouts, and diversity of different layouts (templates) within a single document type. We design few-shot and conventional experiment settings along with a carefully designed matching algorithm to evaluate extraction results. We report the performance of strong baselines and offer three observations: (1) generalizing to new document templates is still very challenging, (2) few-shot performance has a lot of headroom, and (3) models struggle with hierarchical fields such as line-items in an invoice. We plan to open source the benchmark and the evaluation toolkit. We hope this helps the community make progress on these challenging tasks in extracting structured data from visually rich documents. Zilong Wang 0002, Yichao Zhou 0001, Wei Wei 0019, Chen-Yu Lee, Sandeep Tata |
KDD | 2 |
| 2022 | ReLiable: Offline Reinforcement Learning for Tactical Strategies in Professional Basketball GamesabstractProfessional basketball provides an intriguing example of a dynamic spatio-temporal game that incorporates both hidden strategy policies and situational decision making. During a game, the coaches and players are assumed to follow a general game plan, but players are also forced to make spur-of-the-moment decisions based on immediate conditions on the court. However, because it is challenging to process heterogeneous signals on the court and the space of potential actions and outcomes is massive, it is hard for players to find an optimal strategy on the fly given a short amount of time to observe conditions and take action. In this work, we present ReLiable (ReinforcemEnt Learning In bAsketBaLl gamEs). Specifically, we investigate the possibility of using reinforcement learning (RL) to guide player decisions. We train an offline deep Q-network (DQN) on historical National Basketball Association (NBA) game data from 2015-2016. The data include play-by-play and player movement sensor data. We apply our trained agent to games that it has not seen. Our method is able to propose potentially smarter tactical strategies, compared with replay gameplay data, producing expected final game scores comparable to elite NBA teams. Our approach can be useful for learning strategy policies from other game-like domains characterized by competing groups and sequential spatio-temporal event data. Xiusi Chen, Jyun-Yu Jiang, Yichao Zhou 0001, Mingyan Liu, P. Jeffrey Brantingham, Wei Wang 0010 |
CIKM | 4 |
| 2022 | Learning Transferable Node Representations for Attribute Extraction from Web DocumentsabstractGiven a web page, extracting an object along with various attributes of interest (e.g. price, publisher, author, and genre for a book) can facilitate a variety of downstream applications such as large-scale knowledge base construction, e-commerce product search, and personalized recommendation. Prior approaches have either relied on computationally expensive visual feature engineering or required large amounts of training data to get to an acceptable precision. In this paper, we propose a novel method, LeArNing TransfErable node RepresentatioNs for Attribute Extraction (LANTERN), to tackle the problem. We model the problem as a tree node tagging task. The key insight is to learn a contextual representation for each node in the DOM tree where the context explicitly takes into account the tree structure of the neighborhood around the node. Experiments on the SWDE public dataset show that LANTERN outperforms the previous state-of-the-art (SOTA) by 1.44% (F1 score) with a dramatically simpler model architecture. Furthermore, we report that utilizing data from a different domain (for instance, using training data about web pages with cars to extract book objects) is surprisingly useful and helps beat the SOTA by a further 1.37%. Yichao Zhou 0001, Ying Sheng 0002, Nguyen Vo, Nicholas Gerard Edmonds, Sandeep Tata |
WSDM | 1 |
| 2021 | Clinical Temporal Relation Extraction with Probabilistic Soft Logic Regularization and Global InferenceabstractThere has been a steady need in the medical community to precisely extract the temporal relations between clinical events. In particular, temporal information can facilitate a variety of downstream applications such as case report retrieval and medical question answering. Existing methods either require expensive feature engineering or are incapable of modeling the global relational dependencies among the events. In this paper, we propose a novel method, Clinical Temporal ReLation Exaction with Probabilistic Soft Logic Regularization and Global Inference (CTRL-PG) to tackle the problem at the document level. Extensive experiments on two benchmark datasets, I2B2-2012 and TB-Dense, demonstrate that CTRL-PG significantly outperforms baseline methods for temporal relation extraction. Yichao Zhou 0001, Rujun Han, J. Harry Caufield, Kai-Wei Chang 0001, Yizhou Sun, Peipei Ping, Wei Wang 0010 |
AAAI | 1 |
| 2021 | #StayHome or #Marathon?: Social Media Enhanced Pandemic Surveillance on Spatial-temporal Dynamic GraphsabstractCOVID-19 has caused lasting damage to almost every domain in public health, society, and economy. To monitor the pandemic trend, existing studies rely on the aggregation of traditional statistical models and epidemic spread theory. In other words, historical statistics of COVID-19, as well as the population mobility data, become the essential knowledge for monitoring the pandemic trend. However, these solutions can barely provide precise prediction and satisfactory explanations on the long-term disease surveillance while the ubiquitous social media resources can be the key enabler for solving this problem. For example, serious discussions may occur on social media before and after some breaking events take place. To take advantage of the social media data, we propose a novel framework, Social Media enhAnced pandemic suRveillance Technique (SMART), which is composed of two modules: (i) information extraction module to construct heterogeneous knowledge graphs based on the extracted events and relationships among them; (ii) time series prediction module to provide both short-term and long-term forecasts of the confirmed cases and fatality at the state-level in the United States and to discover risk factors for COVID-19 interventions. Extensive experiments show that our method largely outperforms the state-of-the-art baselines by 7.3% and 7.4% in confirmed case/fatality prediction, respectively. Yichao Zhou 0001, Jyun-Yu Jiang, Xiusi Chen, Wei Wang 0010 |
CIKM | 1 |
| 2021 | CREATe: Clinical Report Extraction and Annotation TechnologyabstractClinical case reports are written descriptions of the unique aspects of a particular clinical case, playing an essential role in sharing clinical experiences about atypical disease phenotypes and new therapies. However, to our knowledge, there has been no attempt to develop an end-to-end system to annotate, index, or otherwise curate these reports. In this paper, we propose a novel computational resource platform, CREATe, for extracting, indexing, and querying the contents of clinical case reports. CREATe fosters an environment of sustainable resource support and discovery, enabling researchers to overcome the challenges of information science. An online video of the demonstration can be viewed at https://youtu.be/Q8owBQYTjDc. Yichao Zhou 0001, Bowen Zhang 0002, J. Harry Caufield, Kai-Wei Chang 0001, Yizhou Sun, Peipei Ping, Wei Wang 0010 |
ICDE | 1 |
| 2020 | "The Boating Store Had Its Best Sail Ever": Pronunciation-attentive Contextualized Pun RecognitionabstractHumor plays an important role in human languages and it is essential to model humor when building intelligence systems. Among different forms of humor, puns perform wordplay for humorous effects by employing words with double entendre and high phonetic similarity. However, identifying and modeling puns are challenging as puns usually involved implicit semantic or phonological tricks. In this paper, we propose Pronunciation-attentive Contextualized Pun Recognition (PCPR) to perceive human humor, detect if a sentence contains puns and locate them in the sentence. PCPR derives contextualized representation for each word in a sentence by capturing the association between the surrounding context and its corresponding phonetic symbols. Extensive experiments are conducted on two benchmark datasets. Results demonstrate that the proposed approach significantly outperforms the state-of-the-art methods in pun detection and location tasks. In-depth analyses verify the effectiveness and robustness of PCPR. Yichao Zhou 0001, Jyun-Yu Jiang, Jieyu Zhao 0001, Kai-Wei Chang 0001, Wei Wang 0010 |
ACL | 1 |
| 2020 | Learning to Create Better Ads: Generation and Ranking Approaches for Ad Creative RefinementabstractIn the online advertising industry, the process of designing an ad creative i.e., ad text and image) requires manual labor. Typically, each advertiser launches multiple creatives via online A/B tests to infer effective creatives for the target audience, that are then refined further in an iterative fashion. Due to the manual nature of this process, it is time-consuming to learn, refine, and deploy the modified creatives. Since major ad platforms typically run A/B tests for multiple advertisers in parallel, we explore the possibility of collaboratively learning ad creative refinement via A/B tests of multiple advertisers. In particular, given an input ad creative, we study approaches to refine the given ad text and image by: (i) generating new ad text, (ii) recommending keyphrases for new ad text, and (iii) recommending image tags (objects in the image) to select new ad image. Based on A/B tests conducted by multiple advertisers, we form pairwise examples of inferior and superior ad creatives and use such pairs to train models for the above tasks. For generating new ad text, we demonstrate the efficacy of an encoder-decoder architecture with copy mechanism, which allows some words from the (inferior) input text to be copied to the output while incorporating new words associated with higher click-through-rate. For the keyphrase and image tag recommendation task, we demonstrate the efficacy of a deep relevance matching model, as well as the relative robustness of ranking approaches compared to ad text generation in cold-start scenarios with unseen advertisers. We also share broadly applicable insights from our experiments using data from the Yahoo Gemini ad platform. Shaunak Mishra, Manisha Verma, Yichao Zhou 0001, Kapil Thadani, Wei Wang 0010 |
CIKM | 3 |
| 2020 | Domain Knowledge Empowered Structured Neural Net for End-to-End Event Temporal Relation ExtractionabstractExtracting event temporal relations is a critical task for information extraction and plays an important role in natural language understanding.Prior systems leverage deep learning and pre-trained language models to improve the performance of the task.However, these systems often suffer from two shortcomings: 1) when performing maximum a posteriori (MAP) inference based on neural models, previous systems only used structured knowledge that is assumed to be absolutely correct, i.e., hard constraints; 2) biased predictions on dominant temporal relations when training with a limited amount of data.To address these issues, we propose a framework that enhances deep neural network with distributional constraints constructed by probabilistic domain knowledge.We solve the constrained inference problem via Lagrangian Relaxation and apply it to end-to-end event temporal relation extraction tasks.Experimental results show our framework is able to improve the baseline neural network models with strong statistical significance on two widely used datasets in news and clinical domains. Rujun Han, Yichao Zhou 0001, Nanyun Peng 0001 |
EMNLP (1) | 2 |
| 2020 | Social Media User Geolocation via Hybrid AttentionabstractDetermining user geolocation is vital to various real-world applications on the internet, such as online marketing and event detection. To identify the geolocations of users, their behaviors on social media like published posts and social interactions can be strong evidence. However, most of the existing social media based approaches individually learn from text contexts and social networks. This separation can not only lead to sub-optimal performance but also ignore the distinct importance of two resources for different users. To address this challenge, we propose a novel end-to-end framework, Hybrid-attentive User Geolocation (HUG), to jointly model post texts and user interactions in social media. The hybrid attention mechanism is introduced to automatically determine the importance of texts and social networks for each user while social media posts and interactions are modeled by a graph attention network and a language attention network. Extensive experiments conducted on three benchmark geolocation datasets using Twitter data demonstrate that HUG significantly outperforms competitive baseline methods. The in-depth analysis also indicates the robustness and interpretability of HUG. Cheng Zheng 0004, Jyun-Yu Jiang, Yichao Zhou 0001, Sean D. Young, Wei Wang 0010 |
SIGIR | 3 |
| 2020 | Recommending Themes for Ad Creative Design via Visual-Linguistic RepresentationsabstractThere is a perennial need in the online advertising industry to refresh ad creatives, i.e., images and text used for enticing online users towards a brand. Such refreshes are required to reduce the likelihood of ad fatigue among online users, and to incorporate insights from other successful campaigns in related product categories. Given a brand, to come up with themes for a new ad is a painstaking and time consuming process for creative strategists. Strategists typically draw inspiration from the images and text used for past ad campaigns, as well as world knowledge on the brands. To automatically infer ad themes via such multimodal sources of information in past ad campaigns, we propose a theme (keyphrase) recommender system for ad creative strategists. The theme recommender is based on aggregating results from a visual question answering (VQA) task, which ingests the following: (i) ad images, (ii) text associated with the ads as well as Wikipedia pages on the brands in the ads, and (iii) questions around the ad. We leverage transformer based cross-modality encoders to train visual-linguistic representations for our VQA task. We study two formulations for the VQA task along the lines of classification and ranking; via experiments on a public dataset, we show that cross-modal representations lead to significantly better classification accuracy and ranking precision-recall metrics. Cross-modal representations show better performance compared to separate image and text representations. In addition, the use of multimodal information shows a significant lift over using only textual or visual information. Yichao Zhou 0001, Shaunak Mishra, Manisha Verma, Narayan L. Bhamidipati, Wei Wang 0010 |
WWW | 1 |
| 2019 | Learning to Discriminate Perturbations for Blocking Adversarial Attacks in Text ClassificationabstractYichao Zhou, Jyun-Yu Jiang, Kai-Wei Chang, Wei Wang. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Yichao Zhou 0001, Jyun-Yu Jiang, Kai-Wei Chang 0001, Wei Wang 0010 |
EMNLP/IJCNLP (1) | 1 |
| 2019 | Understanding Consumer Journey using Attention based Recurrent Neural NetworksabstractPaths of online users towards a purchase event (conversion) can be very complex, and guiding them through their journey is an integral part of online advertising. Studies in marketing indicate that a conversion event is typically preceded by one or more purchase funnel stages, viz., unaware, aware, interest, consideration, and intent. Intuitively, some online activities, including web searches, site visits and ad interactions, can serve as markers for the user's funnel stage. Identifying such markers can potentially refine conversion prediction, guide the design of ad creatives (text and images), and lead to higher ad effectiveness. We explore this hypothesis through a set of experiments designed for two tasks: (i) conversion prediction given a user's activity trail, and (ii) funnel stage specific targeting and creatives. To address challenges in the two tasks, we propose an attention based recurrent neural network (RNN) which ingests a user activity trail, and predicts the user's conversion probability along with attention weights for each activity (analogous to its position in the funnel). Specifically, we propose novel attention mechanisms, which maintain a global weight for each activity across all user trails, and also indicate the activity's funnel stage. Use of the proposed attention mechanisms for the first task of conversion prediction shows significant AUC lifts of 0.9% on a public dataset (RecSys 2015 challenge), and up to 3.6% on three proprietary datasets from a major advertising platform (Yahoo Gemini). To address the second task, the activity weights from the proposed mechanisms are used to automatically assign users to funnel stages via a scalable scoring method. Offline evaluation shows that such activity weights are more aligned with editorially tagged activity-funnel stages compared to weights from existing attention mechanisms and simpler conversion models like logistic regression. In addition, results of online ad campaigns in Yahoo Gemini with funnel specific user targeting and ad creatives show strong performance lifts further validating the connection across online activities, purchase funnel stages, stage-specific custom creatives, and conversions. Yichao Zhou 0001, Shaunak Mishra, Jelena Gligorijevic, Tarun Bhatia, Narayan L. Bhamidipati |
KDD | 1 |
| 2018 | Learning Gender-Neutral Word EmbeddingsabstractWord embedding models have become a fundamental component in a wide range of Natural Language Processing (NLP) applications.However, embeddings trained on human-generated corpora have been demonstrated to inherit strong gender stereotypes that reflect social constructs.To address this concern, in this paper, we propose a novel training procedure for learning gender-neutral word embeddings.Our approach aims to preserve gender information in certain dimensions of word vectors while compelling other dimensions to be free of gender influence.Based on the proposed method, we generate a Gender-Neutral variant of GloVe (GN-GloVe).Quantitative and qualitative experiments demonstrate that GN-GloVe successfully isolates gender information without sacrificing the functionality of the embedding model. Jieyu Zhao 0001, Yichao Zhou 0001, Zeyu Li 0001, Wei Wang 0010, Kai-Wei Chang 0001 |
EMNLP | 2 |
| 2017 | AZTEC: A Cloud-based Computational Platform to Integrate Biomedical ResourcesabstractOmics phenotyping has become increasingly recognizedin our path to Precision Medicine. A major computationalchallenge of our investigator community is to identify the necessarydata analytical tools to process multi-omics data. To aidnavigation of the analytical tools fragmented across the web, wehave created a novel computational resource platform, AZTEC,that empowers users to simultaneously search a diverse arrayof digital resources including databases, standalone software,web services, publications, and large libraries composed ofmany interrelated functions. AZTEC fosters an environment ofsustainable resource support and discovery, enabling researchersto overcome the challenges of information science. An online videoof the demonstration can be viewed at https://www.youtube.com/watch?v=FIxvlof6Kbg. Patrick Tan, Yichao Zhou 0001, Xinxin Huang, Giuseppe M. Mazzeo, Chelsea J.-T. Ju |
ICDE | 2 |