VLDB 2026 Research / reviewers in the wild / expert
Zihan Wang 0001
dblp:152/5077-1
· DBLP profile ↗
16ranked-venue papers
4as first author
13since 2021 · last 2025
0000-0002-3147-4642ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 4 first-author · 12 since 2021Databases, data management, data science and information retrieval · 4 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Direct Prompt Optimization with Continuous RepresentationsabstractPrompt optimization for language models faces challenges due to the large discrete search space, the reliance on continuous gradient updates, and the need to round continuous representations into discrete prompts, which causes inflexibility and instability.Existing methods attempt to address these by constraining the search space and adopting greedy, incremental improvements, but they often fail to fully leverage historical gradient information.In this paper, we model the prompt optimization problem by the probability distribution of the prompt and present a novel approach that integrates greedy strategies into optimization with continuous representations.This approach can exploit historical gradient information to address the instability caused by rounding in existing methods.Our study indicates that using continuous representations can improve prompt optimization performance on both text classification and attack tasks, as well as models, including GPT-2, OPT, Vicuna, and LLaMA-2, and also be adaptable to models of different sizes. Yangkun Wang, Zihan Wang 0001, Jingbo Shang |
ACL (1) | 2 |
| 2024 | Learn from Failure: Fine-tuning LLMs with Trial-and-Error Data for Intuitionistic Propositional Logic ProvingabstractChenyang An, Zhibo Chen, Qihao Ye, Emily First, Letian Peng, Jiayun Zhang, Zihan Wang, Sorin Lerner, Jingbo Shang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Chenyang An, Zhibo Chen 0009, Qihao Ye, Emily First, Letian Peng, Jiayun Zhang, Zihan Wang 0001, Sorin Lerner, Jingbo Shang |
ACL (1) | 7 |
| 2024 | Answer is All You Need: Instruction-following Text Embedding via Answering the QuestionabstractLetian Peng, Yuwei Zhang, Zilong Wang, Jayanth Srinivasa, Gaowen Liu, Zihan Wang, Jingbo Shang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Letian Peng, Yuwei Zhang 0001, Zilong Wang 0002, Jayanth Srinivasa, Gaowen Liu, Zihan Wang 0001, Jingbo Shang |
ACL (1) | 6 |
| 2024 | Text Grafting: Near-Distribution Weak Supervision for Minority Classes in Text ClassificationabstractFor extremely weak-supervised text classification, pioneer research generates pseudo labels by mining texts similar to the class names from the raw corpus, which may end up with very limited or even no samples for the minority classes.Recent works have started to generate the relevant texts by prompting LLMs using the class names or definitions; however, there is a high risk that LLMs cannot generate indistribution (i.e., similar to the corpus where the text classifier will be applied) data, leading to ungeneralizable classifiers.In this paper, we combine the advantages of these two approaches and propose to bridge the gap via a novel framework, text grafting, which aims to obtain clean and near-distribution weak supervision for minority classes.Specifically, we first use LLM-based logits to mine masked templates from the raw corpus, which have a high potential for data synthesis into the target minority class.Then, the templates are filled by state-of-the-art LLMs to synthesize neardistribution texts falling into minority classes.Text grafting shows significant improvement over direct mining or synthesis on minority classes.We also use analysis and case studies to comprehend the property of text grafting. Letian Peng, Yi Gu 0002, Chengyu Dong, Zihan Wang 0001, Jingbo Shang |
EMNLP | 4 |
| 2024 | How Few Davids Improve One Goliath: Federated Learning in Resource-Skewed Edge Computing EnvironmentsabstractReal-world deployment of federated learning requires orchestrating clients with widely varied compute resources, from strong enterprise-grade devices in data centers to weak mobile and Web-of-Things devices. Prior works have attempted to downscale large models for weak devices and aggregate shared parts among heterogeneous models. A typical architectural assumption is that there are equally many strong and weak devices. In reality, however, we often encounter resource skew where a few (1 or 2) strong devices hold substantial data resources, alongside many weak devices. This poses challenges-the unshared portion of the large model rarely receives updates or gains benefits from weak collaborators. Jiayun Zhang, Shuheng Li, Haiyu Huang 0003, Zihan Wang 0001, Xiaohan Fu, Dezhi Hong, Rajesh K. Gupta 0001, Jingbo Shang |
WWW | 4 |
| 2023 | WOT-Class: Weakly Supervised Open-world Text ClassificationabstractState-of-the-art weakly supervised text classification methods, while significantly reduced the required human supervision, still requires the supervision to cover all the classes of interest. This is never easy to meet in practice when human explore new, large corpora without complete pictures. In this paper, we work on a novel yet important problem of weakly supervised open-world text classification, where supervision is only needed for a few examples from a few known classes and the machine should handle both known and unknown classes in test time. General open-world classification has been studied mostly using image classification; however, existing methods typically assume the availability of sufficient known-class supervision and strong unknown-class prior knowledge (e.g., the number and/or data distribution). We propose a novel framework øur that lifts those strong assumptions. Specifically, it follows an iterative process of (a) clustering text to new classes, (b) mining and ranking indicative words for each class, and (c) merging redundant classes by using the overlapped indicative words as a bridge. Extensive experiments on 7 popular text classification datasets demonstrate that øur outperforms strong baselines consistently with a large margin, attaining 23.33% greater average absolute macro-F1 over existing approaches across all datasets. Such competent accuracy illuminates the practical potential of further reducing human effort for text classification. Tianle Wang 0003, Zihan Wang 0001, Weitang Liu, Jingbo Shang |
CIKM | 2 |
| 2023 | ClusterLLM: Large Language Models as a Guide for Text ClusteringabstractWe introduce CLUSTERLLM, a novel text clustering framework that leverages feedback from an instruction-tuned large language model, such as ChatGPT.Compared with traditional unsupervised methods that builds upon "small" embedders, CLUSTERLLM exhibits two intriguing advantages: (1) it enjoys the emergent capability of LLM even if its embeddings are inaccessible; and (2) it understands the user's preference on clustering through textual instruction and/or a few annotated data.First, we prompt ChatGPT for insights on clustering perspective by constructing hard triplet questions , where A, B and C are similar data points that belong to different clusters according to small embedder.We empirically show that this strategy is both effective for fine-tuning small embedder and cost-efficient to query ChatGPT.Second, we prompt ChatGPT for helps on clustering granularity by carefully designed pairwise questions , and tune the granularity from cluster hierarchies that is the most consistent with the ChatGPT answers.Extensive experiments on 14 datasets show that CLUSTERLLM consistently improves clustering quality, at an average cost of ∼$0.6 1 per dataset.The code will be available at https: //github.com/zhang-yu-wei/ClusterLLM. Yuwei Zhang 0001, Zihan Wang 0001, Jingbo Shang |
EMNLP | 2 |
| 2023 | Debiasing Made State-of-the-art: Revisiting the Simple Seed-based Weak Supervision for Text ClassificationabstractRecent advances in weakly supervised text classification mostly focus on designing sophisticated methods to turn high-level human heuristics into quality pseudo-labels.In this paper, we revisit the seed matching-based method, which is arguably the simplest way to generate pseudo-labels, and show that its power was greatly underestimated.We show that the limited performance of seed matching is largely due to the label bias injected by the simple seed-match rule, which prevents the classifier from learning reliable confidence for selecting high-quality pseudo-labels.Interestingly, simply deleting the seed words present in the matched input texts can mitigate the label bias and help learn better confidence.Subsequently, the performance achieved by seed matching can be improved significantly, making it on par with or even better than the state-of-theart.Furthermore, to handle the case when the seed words are not made known, we propose to simply delete the word tokens in the input text randomly with a high deletion ratio.Remarkably, seed matching equipped with this random deletion method can often achieve even better performance than that with seed deletion.We refer to our method as SimSeed, which is publicly available 1 . Chengyu Dong, Zihan Wang 0001, Jingbo Shang |
EMNLP | 2 |
| 2023 | Goal-Driven Explainable Clustering via Language DescriptionsabstractUnsupervised clustering is widely used to explore large corpora, but existing formulations neither consider the users' goals nor explain clusters' meanings.We propose a new task formulation, "Goal-Driven Clustering with Explanations" (GOALEX), which represents both the goal and the explanations as free-form language descriptions.For example, to categorize the errors made by a summarization system, the input to GOALEX is a corpus of annotatorwritten comments for system-generated summaries and a goal "cluster the comments based on why the annotators think the summary is imperfect.";the outputs are text clusters each with an explanation ("this cluster mentions that the summary misses important context information."),which relates to the goal and accurately explains which comments should (not) belong to a cluster.To tackle GOALEX, we prompt a language model with "[corpus subset] + [goal] + Brainstorm a list of explanations each representing a cluster.";then we classify whether each sample belongs to a cluster based on its explanation; finally, we use integer linear programming to select a subset of candidate clusters to cover most samples while minimizing overlaps.Under both automatic and human evaluation on corpora with or without labels, our method produces more accurate and goalrelated explanations than prior methods. Zihan Wang 0001, Jingbo Shang, Ruiqi Zhong |
EMNLP | 1 |
| 2022 | WeDef: Weakly Supervised Backdoor Defense for Text ClassificationabstractExisting backdoor defense methods are only effective for limited trigger types.To defend different trigger types at once, we start from the class-irrelevant nature of the poisoning process and propose a novel weakly supervised backdoor defense framework WeDef.Recent advances in weak supervision make it possible to train a reasonably accurate text classifier using only a small number of user-provided, class-indicative seed words.Such seed words shall be considered independent of the triggers.Therefore, a weakly supervised text classifier trained by only the poisoned documents without their labels will likely have no backdoor.Inspired by this observation, in WeDef, we define the reliability of samples based on whether the predictions of the weak classifier agree with their labels in the poisoned training set.We further improve the results through a two-phase sanitization: (1) iteratively refine the weak classifier based on the reliable samples and (2) train a binary poison classifier by distinguishing the most unreliable samples from the most reliable samples.Finally, we train the sanitized model on the samples that the poison classifier predicts as benign.Extensive experiments show that WeDef is effective against popular trigger-based attacks (e.g., words, sentences, and paraphrases), outperforming existing defense methods. Lesheng Jin, Zihan Wang 0001, Jingbo Shang |
EMNLP | 2 |
| 2021 | "Average" Approximates "First Principal Component"? An Empirical Analysis on Representations from Neural Language ModelsabstractContextualized representations based on neural language models have furthered the state of the art in various NLP tasks.Despite its great success, the nature of such representations remains a mystery.In this paper, we present an empirical property of these representations-"average" ≈ "first principal component".Specifically, experiments show that the average of these representations shares almost the same direction as the first principal component of the matrix whose columns are these representations.We believe this explains why the average representation is always a simple yet strong baseline.Our further examinations show that this property also holds in more challenging scenarios, for example, when the representations are from a model right after its random initialization.Therefore, we conjecture that this property is intrinsic to the distribution of representations and not necessarily related to the input structure.We realize that these representations empirically follow a normal distribution for each dimension, and by assuming this is true, we demonstrate that the empirical property can be in fact derived mathematically. Zihan Wang 0001, Chengyu Dong, Jingbo Shang |
EMNLP (1) | 1 |
| 2021 | UCPhrase: Unsupervised Context-aware Quality Phrase TaggingabstractIdentifying and understanding quality phrases from context is a fundamental task in text mining. The most challenging part of this task arguably lies in uncommon, emerging, and domain-specific phrases. The infrequent nature of these phrases significantly hurts the performance of phrase mining methods that rely on sufficient phrase occurrences in the input corpus. Context-aware tagging models, though not restricted by frequency, heavily rely on domain experts for either massive sentence-level gold labels or handcrafted gazetteers. In this work, we propose UCPhrase, a novel unsupervised context-aware quality phrase tagger. Specifically, we induce high-quality phrase spans as silver labels from consistently co-occurring word sequences within each document. Compared with typical context-agnostic distant supervision based on existing knowledge bases (KBs), our silver labels root deeply in the input domain and context, thus having unique advantages in preserving contextual completeness and capturing emerging, out-of-KB phrases. Training a conventional neural tagger based on silver labels usually faces the risk of overfitting phrase surface names. Alternatively, we observe that the contextualized attention maps generated from a transformer-based neural language model effectively reveal the connections between words in a surface-agnostic way. Therefore, we pair such attention maps with the silver labels to train a lightweight span prediction model, which can be applied to new input to recognize (unseen) quality phrases regardless of their surface names or frequency. Thorough experiments on various tasks and datasets, including corpus-level phrase ranking, document-level keyphrase extraction, and sentence-level phrase tagging, demonstrate the superiority of our design over state-of-the-art pre-trained, unsupervised, and distantly supervised methods. Xiaotao Gu, Zihan Wang 0001, Zhenyu Bi, Yu Meng 0001, Jiawei Han 0001, Jingbo Shang |
KDD | 2 |
| 2021 | X-Class: Text Classification with Extremely Weak SupervisionabstractIn this paper, we explore text classification with extremely weak supervision, i.e., only relying on the surface text of class names.This is a more challenging setting than the seed-driven weak supervision, which allows a few seed words per class.We opt to attack this problem from a representation learning perspective-ideal document representations should lead to nearly the same results between clustering and the desired classification.In particular, one can classify the same corpus differently (e.g., based on topics and locations), so document representations should be adaptive to the given class names.We propose a novel framework X-Class to realize the adaptive representations.Specifically, we first estimate class representations by incrementally adding the most similar word to each class until inconsistency arises.Following a tailored mixture of class attention mechanisms, we obtain the document representation via a weighted average of contextualized word representations.With the prior of each document assigned to its nearest class, we then cluster and align the documents to classes.Finally, we pick the most confident documents from each cluster to train a text classifier.Extensive experiments demonstrate that X-Class can rival and even outperform seed-driven weakly supervised methods on 7 benchmark datasets. Zihan Wang 0001, Dheeraj Mekala, Jingbo Shang |
NAACL-HLT | 1 |
| 2020 | Cross-Lingual Ability of Multilingual BERT: An Empirical Study
Karthikeyan K, Zihan Wang 0001, Stephen Mayhew 0001, Dan Roth 0001 |
ICLR | 2 |
| 2020 | Discriminative Topic Mining via Category-Name Guided Text EmbeddingabstractMining a set of meaningful and distinctive topics automatically from massive text corpora has broad applications. Existing topic models, however, typically work in a purely unsupervised way, which often generate topics that do not fit users’ particular needs and yield suboptimal performance on downstream tasks. We propose a new task, discriminative topic mining, which leverages a set of user-provided category names to mine discriminative topics from text corpora. This new task not only helps a user understand clearly and distinctively the topics he/she is most interested in, but also benefits directly keyword-driven classification tasks. We develop CatE, a novel category-name guided text embedding method for discriminative topic mining, which effectively leverages minimal user guidance to learn a discriminative embedding space and discover category representative terms in an iterative manner. We conduct a comprehensive set of experiments to show that CatE mines high-quality set of topics guided by category names only, and benefits a variety of downstream applications including weakly-supervised classification and lexical entailment direction identification. Yu Meng 0001, Jiaxin Huang 0001, Guangyuan Wang, Zihan Wang 0001, Chao Zhang 0014, Yu Zhang 0044, Jiawei Han 0001 |
WWW | 4 |
| 2019 | CrossWeigh: Training Named Entity Tagger from Imperfect AnnotationsabstractZihan Wang, Jingbo Shang, Liyuan Liu, Lihao Lu, Jiacheng Liu, Jiawei Han. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Zihan Wang 0001, Jingbo Shang, Lihao Lu, Jiacheng Liu 0004, Jiawei Han 0001 |
EMNLP/IJCNLP (1) | 1 |