Haohan Wang

dblp:132/4066 · DBLP profile ↗
← Back
72ranked-venue papers
23as first author
51since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 40 · 8 first-author · 35 since 2021Applied, interdisciplinary, general and emerging computing · 24 · 10 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 20 · 6 first-author · 16 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-author · 6 since 2021Computer networks · 3 · 3 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Query-Efficient Domain Knowledge Stealing Against Large Language Models
abstract
Large language models (LLMs) concentrate substantial knowledge in specialized domains due to extensive pretraining and instruction tuning, and they are now central to commercial and scientific practice. Yet access is usually limited to costly, rate-limited interfaces, which motivates methods that can extract targeted domain knowledge with minimal querying effort. A further challenge is that the target domain may be unknown in advance, so naive or generic prompts waste queries and fail to expose the underlying concepts and relations that structure the domain. In this work, we introduce a query-efficient approach for domain-specific knowledge stealing from black-box language models. Rather than issuing random questions or generic templates, our framework performs self-directed exploration that lets the model find the direction and mine domain knowledge by itself. Starting from a small and diverse seed, it discovers salient domain entities and induces their relations through structured question families that elicit definitional, functional, and compositional information. A feedback-driven controller analyzes the errors and uncertainty of the extracted surrogate model and uses this signal to refine subsequent queries, all without relying on prior domain knowledge or external resources. We evaluate the method in two expert-centric settings, medicine and finance, and observe consistently better performance while requiring significantly fewer queries.
Zhengao Li, Xiaopeng Yuan, Bolin Shen, Kien Le, Haohan Wang, Xugui Zhou, Shangqian Gao, Yushun Dong
AAAI5
2026 Learning to Conceal Risk: Controllable Multi-turn Red Teaming for LLMs in the Financial Domain
abstract
Large Language Models (LLMs) are increasingly deployed in finance, where unsafe behavior can lead to serious compliance and regulatory risks.However, most red-teaming research focuses on overtly harmful content and overlooks attacks that appear legitimate on the surface yet induce compliance-violating responses.We address this gap by introducing a controllable black-box multi-turn riskconcealed red-teaming framework (CoRT) that progressively conceals surface-level risk while exploiting non-compliant behaviors.CoRT contains two key components: (i) a Risk Concealment Attacker (RCA) that generates multi-turn prompts via iterative refinement, and (ii) a Risk Concealment Controller (RCC) that predicts a turn-level Risk Concealment Score (RCS) to steer RCA's follow-up style.We also build a domain-specific benchmark, FIN-Bench, with 522 instructions spanning six financial risk categories.Experiments on nine widely used LLMs show that CoRT (RCA) achieves 93.19% average attack success rate (ASR), and CoRT (RCA+RCC) further improves the average ASR to 95.00%.
Haibo Jin, Wenbin Zhang 0002, Haohan Wang, Jun Zhuang 0004
ACL (1)4
2026 "Capture Your Experience at This Moment": Collecting Concurrent User Experience Data in Immersive Virtual Environment
abstract
Environmental User Experience (UX) data collection is essential for user research, enabling evidence-based design decisions. However, traditional retrospective methods like micro-phenomenological interviews suffer from recall inaccuracies and memory distortions. Concurrent UX data collection methods with environmental contexts are promising but lack in-depth investigation. To examine this potential, we conducted a formative study with 34 participants, identifying design goals such as natural interaction, in-situ annotation, and spatial-temporal coupling. We developed JourneyCapturer, an interactive tool that fulfills these goals to integrate concurrent annotation within Immersive Virtual Environment (IVE), enabling real-time UX data capture within contextual scenarios. Using a mixed-method design, we comparatively evaluated concurrent IVE annotations, retrospective interviews, and the combined method with 20 participants, demonstrating how JourneyCapturer improves UX collection processes and outcomes. Our findings suggest that a consciously proactive concurrent IVE method with a first-person perspective advances UX research, offering implications for expert collaboration, multi-modal analytics, and IVE-based field studies.
Henry Been-Lirn Duh, Yihan Mei, Haohan Wang, Zhibin Zhou 0002
CHI5
2026 Channel Knowledge Map-aided Hierarchical Beam Training for Massive MIMO Systems
Haohan Wang, Xu Shi 0002, Yashuai Cao, Hengyu Zhang 0003, Jintao Wang 0001
ICC1
2026 BeamCKM: A Framework of Channel Knowledge Map Construction for Multi-Antenna Systems
Haohan Wang, Xu Shi 0002, Hengyu Zhang 0003, Yashuai Cao, Sufang Yang, Jintao Wang 0001, Kaibin Huang
IEEE Trans. Wirel. Commun.1
2025 Towards Adversarially Robust Dataset Distillation by Curvature Regularization
abstract
Dataset distillation (DD) allows datasets to be distilled to fractions of their original size while preserving the rich distributional information so that models trained on the distilled datasets can achieve a comparable accuracy while saving significant computational loads. Recent research in this area has been focusing on improving the accuracy of models trained on distilled datasets. In this paper, we aim to explore a new perspective of DD. We study how to embed adversarial robustness in distilled datasets, so that models trained on these datasets maintain the high accuracy and meanwhile acquire better adversarial robustness. We propose a new method that achieves this goal by incorporating curvature regularization into the distillation process with much less computational overhead than standard adversarial training. Extensive empirical experiments suggest that our method not only outperforms standard adversarial training on both accuracy and robustness with less computation overhead but is also capable of generating robust distilled datasets that can withstand various adversarial attacks.
Eric Xue 0002, Yijiang Li, Haoyang Liu 0001, Peiran Wang, Haohan Wang
AAAI6
2025 Beamforming-Codebook-Aware Channel Knowledge Map Construction for Multi-Antenna Systems
abstract
Channel knowledge map (CKM) has emerged as a crucial technology for next-generation communication, enabling the construction of high-fidelity mappings between spatial environments and channel parameters via electromagnetic information analysis. Traditional CKM construction methods like ray tracing are computationally intensive. Recent studies utilizing neural networks (NNs) have achieved efficient CKM generation with reduced computational complexity and real-time processing capabilities. Nevertheless, existing research predominantly focuses on single-antenna systems, failing to address the beamforming requirements inherent to MIMO configurations. Given that appropriate precoding vector selection in MIMO systems can substantially enhance user communication rates, this paper presents a TransUNet-based framework for constructing CKM, which effectively incorporates discrete Fourier transform (DFT) precoding vectors. The proposed architecture combines a UNet backbone for multiscale feature extraction with a Transformer module to capture global dependencies among encoded linear vectors. Experimental results demonstrate that the proposed method outperforms state-of-the-art (SOTA) deep learning (DL) approaches, yielding a 17% improvement in RMSE compared to RadioWNet. The code is publicly accessible at https://github.com/github-whh/TransUNet.
Haohan Wang, Xu Shi 0002, Hengyu Zhang 0003, Yashuai Cao, Jintao Wang 0001
GLOBECOM1
2025 Generate E-commerce Product Background by Integrating Category Commonality and Personalized Style
abstract
The state-of-the-art methods for e-commerce product background generation suffer from the inefficiency of designing product-wise prompts when scaling up the production, as well as the ineffectiveness of describing fine-grained styles when customizing personalized backgrounds for some specific brands. To address these obstacles, we integrate the category commonality and personalized style into diffusion models. Concretely, we propose a Category-Wise Generator to enable large-scale background generation with only one model for the first time. A unique identifier in the prompt is assigned to each category, whose attention is located on the background by a mask-guided cross attention layer to learn the category-wise style. Furthermore, for products with specific and fine-grained requirements in layout, elements, etc, a Personality-Wise Generator is devised to learn such personalized style directly from a reference image to resolve textual ambiguities, and is trained in a self-supervised manner for more efficient training data usage. To advance research in this field, the first large-scale e-commerce product background generation dataset BG60k is constructed, which covers more than 60k product images from over 2k categories. Experiments demonstrate that our method could generate high-quality backgrounds for different categories, and maintain the personalized background style of reference images. BG60k will be available at https://github.com/Whileherham/BG60k.
Haohan Wang, Jingjing Lv, Junjie Shen 0008, Zhangang Lin, Jingping Shao
ICASSP1
2025 Customizing Domain Adapters for Domain Generalization
Yuyang Ji, Zeyi Huang, Haohan Wang, Yong Jae Lee
ICCV3
2025 Dataset Distillation via the Wasserstein Metric
abstract
Dataset Distillation (DD) aims to generate a compact synthetic dataset that enables models to achieve performance comparable to training on the full large dataset, significantly reducing computational costs. Drawing from optimal transport theory, we introduce WMDD (Wasserstein Metric-based Dataset Distillation), a straightforward yet powerful method that employs the Wasserstein metric to enhance distribution matching. We compute the Wasserstein barycenter of features from a pretrained classifier to capture essential characteristics of the original data distribution. By optimizing synthetic data to align with this barycenter in feature space and leveraging per-class BatchNorm statistics to preserve intra-class variations, WMDD maintains the efficiency of distribution matching approaches while achieving state-of-the-art results across various high-resolution datasets. Our extensive experiments demonstrate WMDD's effectiveness and adaptability, highlighting its potential for advancing machine learning applications at scale.
Haoyang Liu 0001, Yijiang Li, Tiancheng Xing, Peiran Wang, Vibhu Dalal, Luwei Li, Jingrui He, Haohan Wang
ICCV8
2025 Improving Noise Efficiency in Privacy-Preserving Dataset Distillation
abstract
Modern machine learning models heavily rely on large datasets that often include sensitive and private information, raising serious privacy concerns. Differentially private (DP) data generation offers a solution by creating synthetic datasets that limit the leakage of private information within a predefined privacy budget; however, it requires a substantial amount of data to achieve performance comparable to models trained on the original data. To mitigate the significant expense incurred with synthetic data generation, Dataset Distillation (DD) stands out for its remarkable training and storage efficiency. This efficiency is particularly advantageous when integrated with DP mechanisms, curating compact yet informative synthetic datasets without compromising privacy. However, current state-of-the-art private DD methods suffer from a synchronized sampling-optimization process and the dependency on noisy training signals from randomly initialized networks. This results in the inefficient utilization of private information due to the addition of excessive noise. To address these issues, we introduce a novel framework that decouples sampling from optimization for better convergence and improves signal quality by mitigating the impact of DP noise through matching in an informative subspace. On CIFAR-10, our method achieves a \textbf{10.0\%} improvement with 50 images per class and \textbf{8.3\%} increase with just \textbf{one-fifth} the distilled set size of previous state-of-the-art methods, demonstrating significant potential to advance privacy-preserving DD.
Runkai Zheng, Vishnu Asutosh Dasu, Yinong Wang 0001, Haohan Wang, Fernando De la Torre
ICCV4
2025 Examining Alignment of Large Language Models through Representative Heuristics: the case of political stereotypes
abstract
Examining the alignment of large language models (LLMs) has become increasingly important, e.g., when LLMs fail to operate as intended. This study examines the alignment of LLMs with human values for the domain of politics. Prior research has shown that LLM-generated outputs can include political leanings and mimic the stances of political parties on various issues. However, the extent and conditions under which LLMs deviate from empirical positions are insufficiently examined. To address this gap, we analyze the factors that contribute to LLMs' deviations from empirical positions on political issues, aiming to quantify these deviations and identify the conditions that cause them. Drawing on findings from cognitive science about representativeness heuristics, i.e., situations where humans lean on representative attributes of a target group in a way that leads to exaggerated beliefs, we scrutinize LLM responses through this heuristics' lens. We conduct experiments to determine how LLMs inflate predictions about political parties, which results in stereotyping. We find that while LLMs can mimic certain political parties' positions, they often exaggerate these positions more than human survey respondents do. Also, LLMs tend to overemphasize representativeness more than humans. This study highlights the susceptibility of LLMs to representativeness heuristics, suggesting a potential vulnerability of LLMs that facilitates political stereotyping. We also test prompt-based mitigation strategies, finding that strategies that can mitigate representative heuristics in humans are also effective in reducing the influence of representativeness on LLM-generated responses.
Sullam Jeoung, Yubin Ge, Haohan Wang, Jana Diesner
ICLR3
2025 Revolve: Optimizing AI Systems by Tracking Response Evolution in Textual Optimization
abstract
Recent advancements in large language models (LLMs) have significantly enhanced the ability of LLM-based systems to perform complex tasks through natural language processing and tool interaction. However, optimizing these LLM-based systems for specific tasks remains challenging, often requiring manual interventions like prompt engineering and hyperparameter tuning. Existing automatic optimization methods, such as textual feedback-based techniques (*e.g.*, TextGrad), tend to focus on immediate feedback, analogous to using immediate derivatives in traditional numerical gradient descent. However, relying solely on such feedback can be limited when the adjustments made in response to this feedback are either too small or fluctuate irregularly, potentially slowing down or even stalling the optimization process. In this paper, we introduce $\textbf{REVOLVE}$, an optimization method that tracks how $\textbf{R}$esponses $\textbf{EVOLVE}$ across iterations in LLM systems. By focusing on the evolution of responses over time, REVOLVE enables more stable and effective optimization by making thoughtful, progressive adjustments at each step. Experiments across three tasks demonstrate the adaptability and efficiency of our proposal. Beyond its practical contributions, REVOLVE highlights a promising direction, where the rich knowledge from established optimization principles can be leveraged to enhance LLM systems, which paves the way for further advancements in this hybrid domain. Code is available at: https://llm-revolve.netlify.app.
Peiyan Zhang, Haibo Jin, Leyang Hu, Xinnuo Li, Liying Kang, Yangqiu Song, Haohan Wang
ICML8
2025 Evaluating the Inductive Abilities of Large Language Models: Why Chain-of-Thought Reasoning Sometimes Hurts More Than Helps
abstract
Large Language Models (LLMs) have shown remarkable progress across domains, yet their ability to perform inductive reasoning—inferring latent rules from sparse examples—remains limited. It is often assumed that chain-of-thought (CoT) prompting, as used in Large Reasoning Models (LRMs), enhances such reasoning. We investigate this assumption with creating four controlled, diagnostic game-based tasks—chess, Texas Hold’em, dice games, and blackjack—with hidden human-defined rules. We find that CoT reasoning can degrade inductive performance, with LRMs often underperforming their non-reasoning counterparts. To explain this, we present a theoretical framework that reveals how reasoning steps can amplify error through three failure modes: incorrect sub-task decomposition, incorrect sub-task solving, and incorrect final answer summarization. Based on our theoretical and empirical analysis, we introduce structured interventions that adapt CoT generation according to our identified failure types. These interventions improve inductive accuracy without retraining. Our findings suggest that effective (CoT) reasoning depends not only on taking more steps but also on ensuring those steps are well-structured.
Haibo Jin, Peiyan Zhang, Haohan Wang
NeurIPS4
2025 An Investigation on LLMs' Visual Understanding Ability Using SVG for Image-Text Bridging
abstract
Large language models (LLMs) have made significant advancements in natural language understanding. However, through that enormous semantic representation that the LLM has learnt, is it somehow possible for it to understand images as well? This work investigates this question. To enable the LLM to process images, we convert them into a representation given by Scalable Vector Graphics (SVG). To study what the LLM can do with this XML-based textual description of images, we test the LLM on three broad computer vision tasks: (i) visual reasoning and question answering, (ii) image classification under distribution shift, few-shot learning, and (iii) generating new images using visual prompting. Even though we do not naturally associate LLMs with any visual understanding capabilities, our results indicate that the LLM can often do a decent job in many of these tasks, potentially opening new avenues for research into LLMs' ability to understand image data. Our code, data, and models can be found here https://github.com/mu-cai/svg-11m.
Mu Cai, Zeyi Huang, Utkarsh Ojha, Haohan Wang, Yong Jae Lee
WACV5
2025 CTR-Driven Advertising Image Generation with Multimodal Large Language Models
abstract
In web data, advertising images are crucial for capturing user attention and improving advertising effectiveness. Most existing methods generate background for products primarily focus on the aesthetic quality, which may fail to achieve satisfactory online performance. To address this limitation, we explore the use of Multimodal Large Language Models (MLLMs) for generating advertising images by optimizing for Click-Through Rate (CTR) as the primary objective. Firstly, we build targeted pre-training tasks, and leverage a large-scale e-commerce multimodal dataset to equip MLLMs with initial capabilities for advertising image generation tasks. To further improve the CTR of generated images, we propose a novel reward model to fine-tune pre-trained MLLMs through Reinforcement Learning (RL), which can jointly utilize multimodal features and accurately reflect user click preferences. Meanwhile, a product-centric preference optimization strategy is developed to ensure that the generated background content aligns with the product characteristics after fine-tuning, enhancing the overall relevance and effectiveness of the advertising images. Extensive experiments have demonstrated that our method achieves state-of-the-art performance in both online and offline metrics. Our code and pre-trained models are publicly available at: https://github.com/Chenguoz/CAIG.
Xingye Chen, Zhenbang Du, Yanyin Chen, Haohan Wang, Linkai Liu 0002, Jinyuan Zhao, Jingjing Lv, Junjie Shen 0008, Zhangang Lin, Jingping Shao, Yuanjie Shao, Xinge You, Changxin Gao, Nong Sang
WWW6
2025 Re-assessing accuracy degradation: a framework for understanding DNN behavior on similar-but-non-identical test datasets
abstract
Abstract Deep Neural Networks (DNNs) often demonstrate remarkable performance when evaluated on the test dataset used during model creation. However, their ability to generalize effectively when deployed is crucial, especially in critical applications. One approach to assess the generalization capability of a DNN model is to evaluate its performance on replicated test datasets, which are created by closely following the same methodology and procedures used to generate the original test dataset. Our investigation focuses on the performance degradation of pre-trained DNN models in multi-class classification tasks when evaluated on these replicated datasets; this performance degradation has not been entirely explained by generalization shortcomings or dataset disparities. To address this, we introduce a new evaluation framework that leverages uncertainty estimates generated by the models studied. This framework is designed to isolate the impact of variations in the evaluated test datasets and assess DNNs based on the consistency of their confidence in their predictions. By employing this framework, we can determine whether an observed performance drop is primarily caused by model inadequacy or other factors. We applied our framework to analyze 564 pre-trained DNN models across the CIFAR-10 and ImageNet benchmarks, along with their replicated versions. Contrary to common assumptions about model inadequacy, our results indicate a substantial reduction in the performance gap between the original and replicated datasets when accounting for model uncertainty. This suggests a previously unrecognized adaptability of models to minor dataset variations. Our findings emphasize the importance of understanding dataset intricacies and adopting more nuanced evaluation methods when assessing DNN model performance. This research contributes to the development of more robust and reliable DNN models, especially in critical applications where generalization performance is of utmost importance. The code to reproduce our experiments will be available at https://github.com/esla/Reassessing_DNN_Accuracy .
Esla Timothy Anzaku, Haohan Wang, Ajiboye Babalola, Arnout Van Messem, Wesley De Neve
Mach. Learn.2
2025 Advancing Session-Based Recommendations with Atten-Mixer+: Dynamic and Adaptive Multi-Level Intent Mining
abstract
Session-Based Recommendation (SBR) systems, traditionally reliant on complex Graph Neural Networks (GNNs), often face challenges with marginal performance improvements despite increased model complexity. In this article, we dissect the classical GNN-based SBR models and empirically find that the sophisticated GNN propagations might be redundant, given the readout module plays a significant role in GNN-based models. Based on this observation, we introduce Atten-Mixer+, an advanced iteration of our previously developed Multi-Level Attention Mixture Network (Atten-Mixer). Atten-Mixer+ forgoes GNN propagation in favor of a dynamic and adaptive readout process, tailored to the unique characteristics of each session. Different from the vanilla version, Atten-Mixer+ features the Adaptive Intent Scaler (AIS) layer, which dynamically determines the depth of multi-level user intent analysis and a soft allocation approach for generating user intent queries across entire user interaction sequences. This innovative design allows Atten-Mixer+ to capture a nuanced and comprehensive understanding of user behaviors, overcoming the limitations of fixed-length analysis. Empirical evaluations on benchmark datasets highlight Atten-Mixer+’s superior efficiency and effectiveness, marking a significant step forward in the predictive accuracy of SBR systems.
Peiyan Zhang, Jiayan Guo, Chaozhuo Li, Liying Kang, Jae Boum Kim, Jie Xu 0015, Xi Zhang 0008, Yan Zhang 0117, Haohan Wang, Sung Hun Kim 0003
ACM Trans. Intell. Syst. Technol.9
2024 DiffSim: Aligning Diffusion Model and Molecular Dynamics Simulation for Blind Docking
abstract
Predicting the ligand’s binding conformation within a target protein is a pivotal step in drug discovery. Despite its speed, molecular docking is ill-suited for blind docking where the pocket is unknown. Recently, deep generative models, especially diffusion models, have been proposed for accurate blind docking. However, it is found that while deep generative models excel in locating the pocket, they still lag behind traditional methods in terms of conformation generation. In this study, we introduce a blind docking approach named DiffSim to seamlessly integrate the diffusion model with molecular dynamics (MD) simulation. By aligning reverse diffusion sampling with MD simulation trajectories, DiffSim aim to generate ligand conformations informed by MD-modelled protein-ligand interactions. We demonstrate that the diffusion model can essentially be a coarse-grained simulator for MD simulation. Empirical results demonstrate the effectiveness of the approach and highlight the potential of combining physics-informed MD simulation with deep learning models in drug discovery.
Yanlin Fei, Zhiling Zhou, Haohan Wang
BIBM4
2024 A Quantitative Approach for Evaluating Disease Focus and Interpretability of Deep Learning Models for Alzheimer's Disease Classification
abstract
Deep learning (DL) models have shown significant potential in Alzheimer’s Disease (AD) classification. However, understanding and interpreting these models remains challenging, which hinders the adoption of these models in clinical practice. Techniques such as saliency maps have been proven effective in providing visual and empirical clues about how these models work, but there still remains a gap in understanding which specific brain regions DL models focus on and whether these brain regions are pathologically associated with AD.To bridge such gap, in this study, we developed a quantitative disease-focusing strategy to first enhance the interpretability of DL models using saliency maps and brain segmentations; then we propose a disease-focus (DF) score that quantifies how much a DL model focuses on brain areas relevant to AD pathology based on clinically known MRI-based pathological regions of AD. Using this strategy, we compared several state-of-the-art DL models, including a baseline 3D ResNet model, a pretrained MedicalNet model, and a MedicalNet with data augmentation to classify patients with AD vs. cognitive normal patients using MRI data; then we evaluated these models in terms of their abilities to focus on disease-relevant regions. Our results show interesting disease-focusing patterns with different models, particularly characteristic patterns with the pretrained models and data augmentation, and also provide insight into their classification performance. These results suggest that the approach we developed for quantitatively assessing the abilities of DL models to focus on disease-relevant regions may help improve the interpretability of these models for AD classification and facilitate their adoption for AD diagnosis in clinical practice. The code is publicly available at https://github.com/Liang-lt/ADNI.
Thomas Yu Chow Tam, Litian Liang, Haohan Wang
BIBM4
2024 Trustworthy and Responsible AI for Information and Knowledge Management System
abstract
The way research and business manage and utilize knowledge is undergoing a significant transformation, driven by Artificial Intelligence (AI). Deep learning and machine learning are emerging as powerful tools for optimizing knowledge management systems, leading to more informed and productive development. AI offers unique solutions for organizations struggling with information overload and inefficient knowledge transfer. These AI models can significantly improve data management and utilization. Imagine an AI-powered system that streamlines onboarding processes, provides precise answers to various queries, and even captures the valuable tacit knowledge (implicit skills and expertise) often residing within individuals. AI bridges the gap between explicit knowledge (easily documented information) and tacit knowledge, fostering a more comprehensive and accessible knowledge base. However, such AI systems solicit trustworthy and responsible approaches to mitigate potential misuse and malfunction. In this workshop, we aim to gather researchers and engineers from academia and industry to discuss the latest advances in trustworthy and responsible AI solutions for information and knowledge management systems.
Huaming Chen, Jun Zhuang 0004, Yu Yao 0005, Wei Jin 0009, Haohan Wang, Yong Xie 0002, Chihung Chi, Kim-Kwang Raymond Choo
CIKM5
2024 EditShield: Protecting Unauthorized Image Editing by Instruction-Guided Diffusion Models
Ruoxi Chen, Haibo Jin, Yixin Liu 0002, Jinyin Chen, Haohan Wang, Lichao Sun 0001
ECCV (63)5
2024 Towards Reliable Advertising Image Generation Using Human Feedback
Zhenbang Du, Haohan Wang, Jingsen Wang, Jingjing Lv, Xin Zhu 0008, Junsheng Jin, Junjie Shen 0008, Zhangang Lin, Jingping Shao
ECCV (20)3
2024 CatchBackdoor: Backdoor Detection via Critical Trojan Neural Path Fuzzing
Haibo Jin, Ruoxi Chen, Jinyin Chen, Haibin Zheng, Haohan Wang
ECCV (47)6
2024 Simple Unsupervised Knowledge Distillation With Space Similarity
Haohan Wang
ECCV (2)2
2024 MemeCLIP: Leveraging CLIP Representations for Multimodal Meme Classification
abstract
The complexity of text-embedded images presents a formidable challenge in machine learning given the need for multimodal understanding of multiple aspects of expression conveyed by them.While previous research in multimodal analysis has primarily focused on singular aspects such as hate speech and its subclasses, this study expands this focus to encompass multiple aspects of linguistics: hate, targets of hate, stance, and humor.We introduce a novel dataset PrideMM comprising 5,063 text-embedded images associated with the LGBTQ+ Pride movement, thereby addressing a serious gap in existing resources.We conduct extensive experimentation on PrideMM by using unimodal and multimodal baseline methods to establish benchmarks for each task.Additionally, we propose a novel framework MemeCLIP for efficient downstream learning while preserving the knowledge of the pre-trained CLIP model.The results of our experiments show that Meme-CLIP achieves superior performance compared to previously proposed frameworks on two real-world datasets.We further compare the performance of MemeCLIP and zero-shot GPT-4 on the hate classification task.Finally, we discuss the shortcomings of our model by qualitatively analyzing misclassified samples.
Siddhant Bikram Shah, Shuvam Shiwakoti, Maheep Chaudhary, Haohan Wang
EMNLP4
2024 Foundation Model-oriented Robustness: Robust Image Model Evaluation with Pretrained Models
abstract
Machine learning has demonstrated remarkable performance over finite datasets, yet whether the scores over the fixed benchmarks can sufficiently indicate the model’s performance in the real world is still in discussion. In reality, an ideal robust model will probably behave similarly to the oracle (e.g., the human users), thus a good evaluation protocol is probably to evaluate the models’ behaviors in comparison to the oracle. In this paper, we introduce a new robustness measurement that directly measures the image classification model’s performance compared with a surrogate oracle (i.e., a zoo of foundation models). Besides, we design a simple method that can accomplish the evaluation beyond the scope of the benchmarks. Our method extends the image datasets with new samples that are sufficiently perturbed to be distinct from the ones in the original sets, but are still bounded within the same image-label structure the original test image represents, constrained by a zoo of foundation models pretrained with a large amount of samples. As a result, our new method will offer us a new way to evaluate the models’ robustness performance, free of limitations of fixed benchmarks or constrained perturbations, although scoped by the power of the oracle. In addition to the evaluation results, we also leverage our generated data to understand the behaviors of the model and our new evaluation strategies.
Peiyan Zhang, Haoyang Liu 0001, Chaozhuo Li, Xing Xie 0001, Sunghun Kim 0001, Haohan Wang
ICLR6
2024 Language Agent Tree Search Unifies Reasoning, Acting, and Planning in Language Models
abstract
While language models (LMs) have shown potential across a range of decision-making tasks, their reliance on simple acting processes limits their broad deployment as autonomous agents. In this paper, we introduce Language Agent Tree Search (LATS) -- the first general framework that synergizes the capabilities of LMs in reasoning, acting, and planning. By leveraging the in-context learning ability of LMs, we integrate Monte Carlo Tree Search into LATS to enable LMs as agents, along with LM-powered value functions and self-reflections for proficient exploration and enhanced decision-making. A key feature of our approach is the incorporation of an environment for external feedback, which offers a more deliberate and adaptive problem-solving mechanism that surpasses the constraints of existing techniques. Our experimental evaluation across diverse domains, including programming, interactive question-answering (QA), web navigation, and math, validates the effectiveness and generality of LATS in decision-making while maintaining competitive or improved reasoning performance. Notably, LATS achieves state-of-the-art pass@1 accuracy (92.7%) for programming on HumanEval with GPT-4 and demonstrates gradient-free performance (average score of 75.9) comparable to gradient-based fine-tuning for web navigation on WebShop with GPT-3.5. Code can be found at https://github.com/lapisrocks/LanguageAgentTreeSearch
Andy Zhou, Michal Shlapentokh-Rothman, Haohan Wang, Yu-Xiong Wang
ICML4
2024 A Lightweight Network for Radar Specific Emitter Identification via Differential Constellation Figure
abstract
Radar specific emitter identification aims to recognize individual radar emitters based on subtle differences in transmitted signals. This paper proposes using differential constellation figure (DCF) features extracted from received radar signals for this task. A compact deep neural network called AL-ResNet classifies the DCF features to identify individual radar emitters. Experiments on simulated radar signals show the proposed DCF method achieving 91.76 % accuracy in recognizing 6 radar emitters using ResNet50 architecture, outperforming techniques using other signal features. Nevertheless, experiments on the lightweight model ALResNet shows efficient deployment of the utilized network.
Haohan Wang, Zongyong Cui, Zongjie Cao
IGARSS1
2024 Jailbreaking Large Language Models Against Moderation Guardrails via Cipher Characters
abstract
Large Language Models (LLMs) are typically harmless but remain vulnerable to carefully crafted prompts known as ``jailbreaks'', which can bypass protective measures and induce harmful behavior. Recent advancements in LLMs have incorporated moderation guardrails that can filter outputs, which trigger processing errors for certain malicious questions. Existing red-teaming benchmarks often neglect to include questions that trigger moderation guardrails, making it difficult to evaluate jailbreak effectiveness. To address this issue, we introduce JAMBench, a harmful behavior benchmark designed to trigger and evaluate moderation guardrails. JAMBench involves 160 manually crafted instructions covering four major risk categories at multiple severity levels. Furthermore, we propose a jailbreak method, JAM (Jailbreak Against Moderation), designed to attack moderation guardrails using jailbreak prefixes to bypass input-level filters and a fine-tuned shadow model functionally equivalent to the guardrail model to generate cipher characters to bypass output-level filters. Our extensive experiments on four LLMs demonstrate that JAM achieves higher jailbreak success ($\sim$ $\times$ 19.88) and lower filtered-out rates ($\sim$ $\times$ 1/6) than baselines.
Haibo Jin, Andy Zhou, Joe D. Menke, Haohan Wang
NeurIPS4
2024 Robust Prompt Optimization for Defending Language Models Against Jailbreaking Attacks
abstract
Despite advances in AI alignment, large language models (LLMs) remain vulnerable to adversarial attacks or jailbreaking, in which adversaries can modify prompts to induce unwanted behavior. While some defenses have been proposed, they have not been adapted to newly proposed attacks and more challenging threat models. To address this, we propose an optimization-based objective for defending LLMs against jailbreaking attacks and an algorithm, Robust Prompt Optimization (RPO), to create robust system-level defenses. Our approach directly incorporates the adversary into the defensive objective and optimizes a lightweight and transferable suffix, enabling RPO to adapt to worst-case adaptive attacks. Our theoretical and experimental results show improved robustness to both jailbreaks seen during optimization and unknown jailbreaks, reducing the attack success rate (ASR) on GPT-4 to 6% and Llama-2 to 0% on JailbreakBench, setting the state-of-the-art.
Andy Zhou, Haohan Wang
NeurIPS3
2024 SAR Incremental Automatic Target Recognition Based on Mutual Information Maximization
abstract
To enable the synthetic aperture radar (SAR) automatic target recognition (ATR) system to continuously adapt to new recognition scenarios, it is necessary to equip the system with the ability to quickly update models. However, when these models learn new tasks, the knowledge of old tasks is quickly forgotten, a phenomenon known as catastrophic forgetting. The reason for catastrophic forgetting is that the model does not use the features of old tasks sufficiently. In this letter, an exemplar-free class incremental learning based on maximizing mutual information (CIL-MMI) is proposed to solve this problem. To effectively use the extracted features, CIL-MMI actively clusters features to maximize the mutual information (MI) between features and corresponding labels. The proposed method successfully avoids the distribution overlap caused by the small interclass differences and large intraclass variances inherent in SAR images. Experiments on the Moving and Stationary Target Acquisition and Recognition (MSTAR) dataset indicate that the proposed method outperforms state-of-the-art approaches, demonstrating improvements of 5.41%, 1.93%, and 2.47% at incremental steps 1, 2, and 3, respectively.
Bin Li 0102, Zongyong Cui, Haohan Wang, Yijie Deng, Jizhen Ma, Jianyu Yang 0001, Zongjie Cao
IEEE Geosci. Remote. Sens. Lett.3
2024 BadLabel: A Robust Perspective on Evaluating and Enhancing Label-Noise Learning
abstract
Label-noise learning (LNL) aims to increase the model's generalization given training data with noisy labels. To facilitate practical LNL algorithms, researchers have proposed different label noise types, ranging from class-conditional to instance-dependent noises. In this paper, we introduce a novel label noise type called BadLabel, which can significantly degrade the performance of existing LNL algorithms by a large margin. BadLabel is crafted based on the label-flipping attack against standard classification, where specific samples are selected and their labels are flipped to other labels so that the loss values of clean and noisy labels become indistinguishable. To address the challenge posed by BadLabel, we further propose a robust LNL method that perturbs the labels in an adversarial manner at each epoch to make the loss values of clean and noisy labels again distinguishable. Once we select a small set of (mostly) clean labeled data, we can apply the techniques of semi-supervised learning to train the model accurately. Empirically, our experimental results demonstrate that existing LNL algorithms are vulnerable to the newly introduced BadLabel noise type, while our proposed robust LNL method can effectively improve the generalization performance of the model under various types of label noise. The new dataset of noisy labels and the source codes of robust LNL algorithms are available at https://github.com/zjfheart/BadLabels.
Jingfeng Zhang, Haohan Wang, Bo Han 0003, Tongliang Liu, Lei Liu 0003, Masashi Sugiyama
IEEE Trans. Pattern Anal. Mach. Intell.3
2023 Calibrated Teacher for Sparsely Annotated Object Detection
abstract
Fully supervised object detection requires training images in which all instances are annotated. This is actually impractical due to the high labor and time costs and the unavoidable missing annotations. As a result, the incomplete annotation in each image could provide misleading supervision and harm the training. Recent works on sparsely annotated object detection alleviate this problem by generating pseudo labels for the missing annotations. Such a mechanism is sensitive to the threshold of the pseudo label score. However, the effective threshold is different in different training stages and among different object detectors. Therefore, the current methods with fixed thresholds have sub-optimal performance, and are difficult to be applied to other detectors. In order to resolve this obstacle, we propose a Calibrated Teacher, of which the confidence estimation of the prediction is well calibrated to match its real precision. In this way, different detectors in different training stages would share a similar distribution of the output confidence, so that multiple detectors could share the same fixed threshold and achieve better performance. Furthermore, we present a simple but effective Focal IoU Weight (FIoU) for the classification loss. FIoU aims at reducing the loss weight of false negative samples caused by the missing annotation, and thus works as the complement of the teacher-student paradigm. Extensive experiments show that our methods set new state-of-the-art under all different sparse settings in COCO. Code will be available at https://github.com/Whileherham/CalibratedTeacher.
Haohan Wang, Liang Liu 0007, Boshen Zhang, Jiangning Zhang, Wuhao Zhang, Zhenye Gan, Yabiao Wang, Chengjie Wang 0001, Haoqian Wang
AAAI1
2023 Toward Robust Diagnosis: A Contour Attention Preserving Adversarial Defense for COVID-19 Detection
abstract
As the COVID-19 pandemic puts pressure on healthcare systems worldwide, the computed tomography image based AI diagnostic system has become a sustainable solution for early diagnosis. However, the model-wise vulnerability under adversarial perturbation hinders its deployment in practical situation. The existing adversarial training strategies are difficult to generalized into medical imaging field challenged by complex medical texture features. To overcome this challenge, we propose a Contour Attention Preserving (CAP) method based on lung cavity edge extraction. The contour prior features are injected to attention layer via a parameter regularization and we optimize the robust empirical risk with hybrid distance metric. We then introduce a new cross-nation CT scan dataset to evaluate the generalization capability of the adversarial robustness under distribution shift. Experimental results indicate that the proposed method achieves state-of-the-art performance in multiple adversarial defense and generalization tasks. The code and dataset are available at https://github.com/Quinn777/CAP.
Kun Xiang, Jinwen She, Haohan Wang, Shiqi Deng, Shancheng Jiang
AAAI5
2023 Self-learning for Annotating Website Privacy Policies at Scale
abstract
With the increasing importance of user data privacy, it is crucial for individuals to understand how companies handle their information. While considerable research has been conducted on automatically identifying privacy-related information in policies, the lack of high-quality annotated training data in this domain remains a significant challenge. Manual annotation of privacy policies is a demanding and time-consuming task that requires domain knowledge. To address this issue, we propose a semi-supervised method, specifically an iterative self-learning approach, to augment the limited training dataset and improve classification performance. Our approach leverages two state-of-the-art models, BERT and XLNet, and involves automatic labelling of data and model retraining with pseudo-labels. We evaluated our approach on the OPP-115 corpora and observed a 10% improvement in the macro F-1 score for BERT, demonstrating the effectiveness of self-learning. This is the first attempt to automatically annotate privacy policies using a self-learning method without requiring additional annotations, offering a promising solution to the challenge of training data scarcity in this domain.
Shufan Ming, Haohan Wang
COMPSAC2
2023 A Sentence Speaks a Thousand Images: Domain Generalization through Distilling CLIP with Language Guidance
abstract
Domain generalization studies the problem of training a model with samples from several domains (or distributions) and then testing the model with samples from a new, unseen domain. In this paper, we propose a novel approach for domain generalization that leverages recent advances in large vision-language models, specifically a CLIP teacher model, to train a smaller model that generalizes to unseen domains. The key technical contribution is a new type of regularization that requires the student’s learned image representations to be close to the teacher’s learned text representations obtained from encoding the corresponding text descriptions of images. We introduce two designs of the loss function, absolute and relative distance, which provide specific guidance on how the training process of the student model should be regularized. We evaluate our proposed method, dubbed RISE (Regularized Invariance with Semantic Embeddings), on various benchmark datasets, and show that it outperforms several state-of-the-art domain generalization methods. To our knowledge, our work is the first to leverage knowledge distillation using a large vision-language model for domain generalization. By incorporating text-based information, RISE improves the generalization capability of machine learning models.
Zeyi Huang, Andy Zhou, Zijian Lin, Mu Cai, Haohan Wang, Yong Jae Lee
ICCV5
2023 Optimizing the Collaboration Structure in Cross-Silo Federated Learning
abstract
In federated learning (FL), multiple clients collaborate to train machine learning models together while keeping their data decentralized. Through utilizing more training data, FL suffers from the potential negative transfer problem: the global FL model may even perform worse than the models trained with local data only. In this paper, we propose FedCollab, a novel FL framework that alleviates negative transfer by clustering clients into non-overlapping coalitions based on their distribution distances and data quantities. As a result, each client only collaborates with the clients having similar data distributions, and tends to collaborate with more clients when it has less data. We evaluate our framework with a variety of datasets, models, and types of non-IIDness. Our results demonstrate that FedCollab effectively mitigates negative transfer across a wide range of FL algorithms and consistently outperforms other clustered FL algorithms.
Haohan Wang, Jun Wu 0019, Jingrui He
ICML2
2023 Trustworthy Machine Learning: Robustness, Generalization, and Interpretability
abstract
Machine learning is becoming increasingly important in today's world. Beyond its powerful performances, there has been an emerging concern about the trustworthiness of machine learning, including but not limited to: robustness to malicious attacks, generalization to unseen datasets, and interpretability to explain its outputs. Such concerns are even more urgent in some safety-critical applications such as medical diagnosis and autonomous driving. Trustworthy machine learning (TrustML) aims to tackle these challenges from the perspectives of theory, algorithm, and applications. In this tutorial, we will give a comprehensive introduction to the recent advance of trustworthy machine learning in robustness, generalization, and interpretability. We will cover their problem formulation, related research, popular algorithms, and successful applications. Additionally, we will also introduce some potential challenges for future research. We do hope that this tutorial will not only serve as a platform to understand TrustML, but also raise the awareness of everyone for more trustworthy applications.
Jindong Wang 0001, Haoliang Li, Haohan Wang, Sinno Jialin Pan, Xing Xie 0001
KDD3
2023 Adaptive Test-Time Personalization for Federated Learning
abstract
Personalized federated learning algorithms have shown promising results in adapting models to various distribution shifts. However, most of these methods require labeled data on testing clients for personalization, which is usually unavailable in real-world scenarios. In this paper, we introduce a novel setting called test-time personalized federated learning (TTPFL), where clients locally adapt a global model in an unsupervised way without relying on any labeled data during test-time. While traditional test-time adaptation (TTA) can be used in this scenario, most of them inherently assume training data come from a single domain, while they come from multiple clients (source domains) with different distributions. Overlooking these domain interrelationships can result in suboptimal generalization. Moreover, most TTA algorithms are designed for a specific kind of distribution shift and lack the flexibility to handle multiple kinds of distribution shifts in FL. In this paper, we find that this lack of flexibility partially results from their pre-defining which modules to adapt in the model. To tackle this challenge, we propose a novel algorithm called ATP to adaptively learns the adaptation rates for each module in the model from distribution shifts among source domains. Theoretical analysis proves the strong generalization of ATP. Extensive experiments demonstrate its superiority in handling various distribution shifts including label shift, image corruptions, and domain shift, outperforming existing TTA methods across multiple datasets and model architectures. Our code is available at https://github.com/baowenxuan/ATP.
Tianxin Wei, Haohan Wang, Jingrui He
NeurIPS3
2023 Distilling Out-of-Distribution Robustness from Vision-Language Foundation Models
abstract
We propose a conceptually simple and lightweight framework for improving the robustness of vision models through the combination of knowledge distillation and data augmentation. We address the conjecture that larger models do not make for better teachers by showing strong gains in out-of-distribution robustness when distilling from pretrained foundation models. Following this finding, we propose Discrete Adversarial Distillation (DAD), which leverages a robust teacher to generate adversarial examples and a VQGAN to discretize them, creating more informative samples than standard data augmentation techniques. We provide a theoretical framework for the use of a robust teacher in the knowledge distillation with data augmentation setting and demonstrate strong gains in out-of-distribution robustness and clean accuracy across different student architectures. Notably, our method adds minor computational overhead compared to similar techniques and can be easily combined with other data augmentations for further improvements.
Andy Zhou, Jindong Wang 0001, Yu-Xiong Wang, Haohan Wang
NeurIPS4
2023 Efficiently Leveraging Multi-level User Intent for Session-based Recommendation via Atten-Mixer Network
abstract
Session-based recommendation (SBR) aims to predict the user's next action based on short and dynamic sessions. Recently, there has been an increasing interest in utilizing various elaborately designed graph neural networks (GNNs) to capture the pair-wise relationships among items, seemingly suggesting the design of more complicated models is the panacea for improving the empirical performance. However, these models achieve relatively marginal improvements with exponential growth in model complexity. In this paper, we dissect the classical GNN-based SBR models and empirically find that some sophisticated GNN propagations are redundant, given the readout module plays a significant role in GNN-based models. Based on this observation, we intuitively propose to remove the GNN propagation part, while the readout module will take on more responsibility in the model reasoning process. To this end, we propose the Multi-Level Attention Mixture Network (Atten-Mixer), which leverages both concept-view and instance-view readouts to achieve multi-level reasoning over item transitions. As simply enumerating all possible high-level concepts is infeasible for large real-world recommender systems, we further incorporate SBR-related inductive biases, i.e., local invariance and inherent priority to prune the search space. Experiments on three benchmarks demonstrate the effectiveness and efficiency of our proposal. We also have already launched the proposed techniques to a large-scale e-commercial online service since April 2021, with significant improvements of top-tier business metrics demonstrated in the online experiments on live traffic.
Peiyan Zhang, Jiayan Guo, Chaozhuo Li, Yueqi Xie, Jae Boum Kim, Yan Zhang 0117, Xing Xie 0001, Haohan Wang, Sunghun Kim 0001
WSDM8
2022 The Two Dimensions of Worst-case Training and Their Integrated Effect for Out-of-domain Generalization
abstract
Training with an emphasis on “hard-to-learn” components of the data has been proven as an effective method to improve the generalization of machine learning models, especially in the settings where robustness (e.g., generalization across distributions) is valued. Existing literature discussing this “hard-to-learn” concept are mainly expanded either along the dimension of the samples or the dimension of the features. In this paper, we aim to introduce a simple view merging these two dimensions, leading to a new, simple yet effective, heuristic to train machine learning models by emphasizing the worst-cases on both the sample and the feature dimensions. We name our method W2D following the concept of “Worst-case along Two Dimensions”. We validate the idea and demonstrate its empirical strength over standard benchmarks.
Zeyi Huang, Haohan Wang, Dong Huang 0007, Yong Jae Lee, Eric P. Xing
CVPR2
2022 Iterative Few-shot Semantic Segmentation from Image Label Text
abstract
Few-shot semantic segmentation aims to learn to segment unseen class objects with the guidance of only a few support images. Most previous methods rely on the pixel-level label of support images. In this paper, we focus on a more challenging setting, in which only the image-level labels are available. We propose a general framework to firstly generate coarse masks with the help of the powerful vision-language model CLIP, and then iteratively and mutually refine the mask predictions of support and query images. Extensive experiments on PASCAL-5i and COCO-20i datasets demonstrate that our method not only outperforms the state-of-the-art weakly supervised approaches by a significant margin, but also achieves comparable or better results to recent supervised methods. Moreover, our method owns an excellent generalization ability for the images in the wild and uncommon classes. Code will be available at https://github.com/Whileherham/IMR-HSNet.
Haohan Wang, Liang Liu 0007, Wuhao Zhang, Jiangning Zhang, Zhenye Gan, Yabiao Wang, Chengjie Wang 0001, Haoqian Wang
IJCAI1
2022 Toward Learning Robust and Invariant Representations with Alignment Regularization and Data Augmentation
abstract
Data augmentation has been proven to be an effective technique for developing machine learning models that are robust to known classes of distributional shifts (e.g., rotations of images), and alignment regularization is a technique often used together with data augmentation to further help the model learn representations invariant to the shifts used to augment the data. In this paper, motivated by a proliferation of options of alignment regularizations, we seek to evaluate the performances of several popular design choices along the dimensions of robustness and invariance, for which we introduce a new test procedure. Our synthetic experiment results speak to the benefits of squared ℓ2 norm regularization. Further, we also formally analyze the behavior of alignment regularization to complement our empirical study under assumptions we consider realistic. Finally, we test this simple technique we identify (worst-case data augmentation with squared ℓ2 norm alignment regularization) and show that the benefits of this method outrun those of the specially designed methods. We also release a software package in both TensorFlow and PyTorch for users to use the method with a couple of lines at https://github.com/jyanln/AlignReg.
Haohan Wang, Zeyi Huang, Xindi Wu, Eric P. Xing
KDD1
2022 Measure and Improve Robustness in NLP Models: A Survey
abstract
As NLP models achieved state-of-the-art performances over benchmarks and gained wide applications, it has been increasingly important to ensure the safe deployment of these models in the real world, e.g., making sure the models are robust against unseen or challenging scenarios.Despite robustness being an increasingly studied topic, it has been separately explored in applications like vision and NLP, with various definitions, evaluation and mitigation strategies in multiple lines of research.In this paper, we aim to provide a unifying survey of how to define, measure and improve robustness in NLP.We first connect multiple definitions of robustness, then unify various lines of work on identifying robustness failures and evaluating models' robustness.Correspondingly, we present mitigation strategies that are data-driven, model-driven, and inductive-prior-based, with a more systematic view of how to effectively improve robustness in NLP models.Finally, we conclude by outlining open challenges and future directions to motivate further research in this area.
Xuezhi Wang 0002, Haohan Wang, Diyi Yang
NAACL-HLT2
2022 Gene Set Priorization Guided by Regulatory Networks with p-values through Kernel Mixed Model
Haohan Wang, Oscar L. Lopez, Wei Wu 0023, Eric P. Xing
RECOMB1
2022 Toward learning human-aligned cross-domain robust models by countering misaligned features
abstract
Machine learning has demonstrated remarkable prediction accuracy over i.i.d data, but the accuracy often drops when tested with data from another distribution. In this paper, we aim to offer another view of this problem in a perspective assuming the reason behind this accuracy drop is the reliance of models on the features that are not aligned well with how a data annotator considers similar across these two datasets. We refer to these features as misaligned features. We extend the conventional generalization error bound to a new one for this setup with the knowledge of how the misaligned features are associated with the label. Our analysis offers a set of techniques for this problem, and these techniques are naturally linked to many previous methods in robust machine learning literature. We also compared the empirical strength of these methods demonstrated the performance when these previous techniques are combined, with implementation available.
Haohan Wang, Zeyi Huang, Hanlin Zhang 0002, Yong Jae Lee, Eric P. Xing
UAI1
2021 Robust Contrastive Learning Using Negative Samples with Diminished Semantics
abstract
Unsupervised learning has recently made exceptional progress because of the development of more effective contrastive learning methods. However, CNNs are prone to depend on low-level features that humans deem non-semantic. This dependency has been conjectured to induce a lack of robustness to image perturbations or domain shift. In this paper, we show that by generating carefully designed negative samples, contrastive learning can learn more robust representations with less dependence on such features. Contrastive learning utilizes positive pairs which preserve semantic information while perturbing superficial features in the training images. Similarly, we propose to generate negative samples in a reversed way, where only the superfluous instead of the semantic features are preserved. We develop two methods, texture-based and patch-based augmentations, to generate negative samples. These samples achieve better generalization, especially under out-of-domain settings. We also analyze our method and the generated texture-based samples, showing that texture features are indispensable in classifying particular ImageNet classes and especially finer classes. We also show that the model bias between texture and shape features favors them differently under different test settings.
Songwei Ge, Shlok Kumar Mishra, Chun-Liang Li, Haohan Wang, David Jacobs 0001
NeurIPS4
2021 Active learning to classify macromolecular structures in situ for less supervision in cryo-electron tomography
abstract
MOTIVATION: Cryo-Electron Tomography (cryo-ET) is a 3D bioimaging tool that visualizes the structural and spatial organization of macromolecules at a near-native state in single cells, which has broad applications in life science. However, the systematic structural recognition and recovery of macromolecules captured by cryo-ET are difficult due to high structural complexity and imaging limits. Deep learning-based subtomogram classification has played critical roles for such tasks. As supervised approaches, however, their performance relies on sufficient and laborious annotation on a large training dataset. RESULTS: To alleviate this major labeling burden, we proposed a Hybrid Active Learning (HAL) framework for querying subtomograms for labeling from a large unlabeled subtomogram pool. Firstly, HAL adopts uncertainty sampling to select the subtomograms that have the most uncertain predictions. This strategy enforces the model to be aware of the inductive bias during classification and subtomogram selection, which satisfies the discriminativeness principle in AL literature. Moreover, to mitigate the sampling bias caused by such strategy, a discriminator is introduced to judge if a certain subtomogram is labeled or unlabeled and subsequently the model queries the subtomogram that have higher probabilities to be unlabeled. Such query strategy encourages to match the data distribution between the labeled and unlabeled subtomogram samples, which essentially encodes the representativeness criterion into the subtomogram selection process. Additionally, HAL introduces a subset sampling strategy to improve the diversity of the query set, so that the information overlap is decreased between the queried batches and the algorithmic efficiency is improved. Our experiments on subtomogram classification tasks using both simulated and real data demonstrate that we can achieve comparable testing performance (on average only 3% accuracy drop) by using less than 30% of the labeled subtomograms, which shows a very promising result for subtomogram classification task with limited labeling resources. AVAILABILITY AND IMPLEMENTATION: https://github.com/xulabs/aitom. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Xuefeng Du, Haohan Wang, Zhenxi Zhu, Yi-Wei Chang, Jing Zhang 0062, Eric P. Xing, Min Xu 0009
Bioinform.2
2021 Coupled mixed model for joint genetic analysis of complex disorders with two independently collected data sets
abstract
BACKGROUND: In the last decade, Genome-wide Association studies (GWASs) have contributed to decoding the human genome by uncovering many genetic variations associated with various diseases. Many follow-up investigations involve joint analysis of multiple independently generated GWAS data sets. While most of the computational approaches developed for joint analysis are based on summary statistics, the joint analysis based on individual-level data with consideration of confounding factors remains to be a challenge. RESULTS: In this study, we propose a method, called Coupled Mixed Model (CMM), that enables a joint GWAS analysis on two independently collected sets of GWAS data with different phenotypes. The CMM method does not require the data sets to have the same phenotypes as it aims to infer the unknown phenotypes using a set of multivariate sparse mixed models. Moreover, CMM addresses the confounding variables due to population stratification, family structures, and cryptic relatedness, as well as those arising during data collection such as batch effects that frequently appear in joint genetic studies. We evaluate the performance of CMM using simulation experiments. In real data analysis, we illustrate the utility of CMM by an application to evaluating common genetic associations for Alzheimer's disease and substance use disorder using datasets independently collected for the two complex human disorders. Comparison of the results with those from previous experiments and analyses supports the utility of our method and provides new insights into the diseases. The software is available at https://github.com/HaohanWang/CMM .
Haohan Wang, Fen Pei, Michael M. Vanyukov, Ivet Bahar, Wei Wu 0023, Eric P. Xing
BMC Bioinform.1
2020 High-Frequency Component Helps Explain the Generalization of Convolutional Neural Networks
abstract
We investigate the relationship between the frequency spectrum of image data and the generalization behavior of convolutional neural networks (CNN). We first notice CNN's ability in capturing the high-frequency components of images. These high-frequency components are almost imperceptible to a human. Thus the observation leads to multiple hypotheses that are related to the generalization behaviors of CNN, including a potential explanation for adversarial examples, a discussion of CNN's trade-off between robustness and accuracy, and some evidence in understanding training heuristics.
Haohan Wang, Xindi Wu, Zeyi Huang, Eric P. Xing
CVPR1
2020 Self-challenging Improves Cross-Domain Generalization
Zeyi Huang, Haohan Wang, Eric P. Xing, Dong Huang 0007
ECCV (2)2
2020 Supervised Adversarial Alignment of Single-Cell RNA-seq Data
Songwei Ge, Haohan Wang, Amir Alavi, Eric P. Xing, Ziv Bar-Joseph
RECOMB2
2020 Poly(A)-DG: A deep-learning-based domain generalization method to identify cross-species Poly(A) signal without prior knowledge from target species
abstract
In eukaryotes, polyadenylation (poly(A)) is an essential process during mRNA maturation. Identifying the cis-determinants of poly(A) signal (PAS) on the DNA sequence is the key to understand the mechanism of translation regulation and mRNA metabolism. Although machine learning methods were widely used in computationally identifying PAS, the need for tremendous amounts of annotation data hinder applications of existing methods in species without experimental data on PAS. Therefore, cross-species PAS identification, which enables the possibility to predict PAS from untrained species, naturally becomes a promising direction. In our works, we propose a novel deep learning method named Poly(A)-DG for cross-species PAS identification. Poly(A)-DG consists of a Convolution Neural Network-Multilayer Perceptron (CNN-MLP) network and a domain generalization technique. It learns PAS patterns from the training species and identifies PAS in target species without re-training. To test our method, we use four species and build cross-species training sets with two of them and evaluate the performance of the remaining ones. Moreover, we test our method against insufficient data and imbalanced data issues and demonstrate that Poly(A)-DG not only outperforms state-of-the-art methods but also maintains relatively high accuracy when it comes to a smaller or imbalanced training set.
Yumin Zheng, Haohan Wang, Yang Zhang 0042, Xin Gao 0001, Eric P. Xing, Min Xu 0009
PLoS Comput. Biol.2
2019 What if We Simply Swap the Two Text Fragments? A Straightforward yet Effective Way to Test the Robustness of Methods to Confounding Signals in Nature Language Inference Tasks
abstract
Nature language inference (NLI) task is a predictive task of determining the inference relationship of a pair of natural language sentences. With the increasing popularity of NLI, many state-of-the-art predictive models have been proposed with impressive performances. However, several works have noticed the statistical irregularities in the collected NLI data set that may result in an over-estimated performance of these models and proposed remedies. In this paper, we further investigate the statistical irregularities, what we refer as confounding factors, of the NLI data sets. With the belief that some NLI labels should preserve under swapping operations, we propose a simple yet effective way (swapping the two text fragments) of evaluating the NLI predictive models that naturally mitigate the observed problems. Further, we continue to train the predictive models with our swapping manner and propose to use the deviation of the model’s evaluation performances under different percentages of training text fragments to be swapped to describe the robustness of a predictive model. Our evaluation metrics leads to some interesting understandings of recent published NLI methods. Finally, we also apply the swapping operation on NLI models to see the effectiveness of this straightforward method in mitigating the confounding factor problems in training generic sentence embeddings for other NLP transfer tasks.
Haohan Wang, Da Sun, Eric P. Xing
AAAI1
2019 Graph-structured Sparse Mixed Models for Genetic Association with Confounding Factors Correction
abstract
Genome-Wide Association Study (GWAS) plays an essential role in understanding human genetics. While various methods have been introduced to increase the signals of GWAS with consideration of population stratification, polygenicity, or pleiotropy. There seems no existing methods that can consdier these three different aspects of genetic association studies together. In this paper, we introduce a new set of models that can utilize the relatedness of available phenotypes to help improve the signals regarding pleiotropy, calculate multivariate coefficients corresponds to polygenicity, and correct population stratification through modelling random effects. We first propose the sparse graph-structured linear mixed model (sGLMM). Then the tree-guided sparse linear mixed model (TgSLMM) has further put forward to explore how specifically clusters are. Our method turns out to outperform other existing approaches after simulation experiments and be capable of exploring the correct genetic association and scales to the large dataset like human genome. Further, we validate, compare and use the effectiveness of both sGLMM and TgSLMM in the real-world genomic dataset on Human Alzheimer's Disease discovered by our model, and justify a few of the most important genetic loci. Overlapping SNPs implies that association between each pair of traits follows only one path and is acyclic, which more corresponds to tree structure.
Haohan Wang, Changpeng Lu, Wei Wu 0023, Eric P. Xing
BIBM1
2019 Deep Inductive Matrix Completion for Biomedical Interaction Prediction
abstract
In many real tasks, side information in addition to the observed entries is available in the matrix completion problem. To make good use of this information, an inductive approach to matrix completion was proposed where the matrix entries are modeled as a bilinear function of real-valued vectors associated with the rows and the columns. However, it is not effective in handling data of nonlinear structures. In this paper, we propose a novel model called Deep Inductive Matrix Completion (DIMC) for nonlinear inductive matrix completion, which consists of two deep-structure neural networks to extract latent features from high-dimensional known side vectors, and then to predict their relationships using the latent features. In DIMC, the parameters of the neural networks are alternatively optimized to minimize the reconstruction error. Then the missing entries can be readily recovered with the side vectors of rows and columns. We compare DIMC with state-of-the-art methods of linear and nonlinear matrix completion in the tasks of drug repositioning, gene-disease and miRNA-disease association prediction. The experimental results verified that DIMC is capable to provide higher accuracy than existing methods and is applicable to predict inductively on new row-column interactions with auxiliary side information. In addition, we discuss the effects of alternating training frequency on the performance of DIMC and how we can utilize such property to implement a GPU-based parallel computing algorithm that significantly shortens the training time.
Haohan Wang, Yibing Wei, Mengxin Cao, Ming Xu 0008, Wei Wu 0023, Eric P. Xing
BIBM1
2019 Regularized Adversarial Training (RAT) for Robust Cellular Electron Cryo Tomograms Classification
abstract
Cellular Electron Cryo Tomography (CECT) 3D imaging has permitted biomedical community to study macromolecule structures inside single cells with deep learning approaches. Many deep learning-based methods have since been developed to classify macromolecule structures from tomograms with high accuracy. However, several recent studies have demonstrated the lack of robustness in these models against often-imperceptible, designed changes of input. Therefore, making existing subtomogram-classification models robust remains a serious challenge. In this paper, we study the robustness of the state-of-the-art subtomogram classifier on CECT images and propose a method called Regularized Adversarial Training (RAT) to defend the classifier against a wide range of designed threats. Our results show that RAT improves robustness for CECT image classification over the previous methods.
Xindi Wu, Yijun Mao, Haohan Wang, Xin Gao 0001, Eric P. Xing, Min Xu 0009
BIBM3
2019 Learning Robust Representations by Projecting Superficial Statistics Out
Haohan Wang, Zexue He, Zachary C. Lipton, Eric P. Xing
ICLR1
2019 Learning Robust Global Representations by Penalizing Local Predictive Power
abstract
Despite their renowned in-domain predictive power, convolutional neural networks are known to rely more on high-frequency patterns that humans deem superficial than on low-frequency patterns that agree better with intuitions about what constitutes category membership. This paper proposes a method for training robust convolutional networks by penalizing the predictive power of the local representations learned by earlier layers. Intuitively, our networks are forced to discard predictive signals such as color and texture that can be gleaned from local receptive fields and to rely instead on the global structures of the image. Across a battery of synthetic and benchmark domain adaptation tasks, our method confers improved generalization out of the domain. Additionally, to evaluate cross-domain transfer, we introduce ImageNet-Sketch, a new dataset consisting of sketch-like images that matches the ImageNet classification validation set in scale and dimension.
Haohan Wang, Songwei Ge, Zachary C. Lipton, Eric P. Xing
NeurIPS1
2019 Precision Lasso: accounting for correlations and linear dependencies in high-dimensional genomic data
abstract
MOTIVATION: Association studies to discover links between genetic markers and phenotypes are central to bioinformatics. Methods of regularized regression, such as variants of the Lasso, are popular for this task. Despite the good predictive performance of these methods in the average case, they suffer from unstable selections of correlated variables and inconsistent selections of linearly dependent variables. Unfortunately, as we demonstrate empirically, such problematic situations of correlated and linearly dependent variables often exist in genomic datasets and lead to under-performance of classical methods of variable selection. RESULTS: To address these challenges, we propose the Precision Lasso. Precision Lasso is a Lasso variant that promotes sparse variable selection by regularization governed by the covariance and inverse covariance matrices of explanatory variables. We illustrate its capacity for stable and consistent variable selection in simulated data with highly correlated and linearly dependent variables. We then demonstrate the effectiveness of the Precision Lasso to select meaningful variables from transcriptomic profiles of breast cancer patients. Our results indicate that in settings with correlated and linearly dependent variables, the Precision Lasso outperforms popular methods of variable selection such as the Lasso, the Elastic Net and Minimax Concave Penalty (MCP) regression. AVAILABILITY AND IMPLEMENTATION: Software is available at https://github.com/HaohanWang/thePrecisionLasso. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Haohan Wang, Benjamin J. Lengerich, Bryon Aragam, Eric P. Xing
Bioinform.1
2019 Deep mixed model for marginal epistasis detection and population stratification correction in genome-wide association studies
abstract
BACKGROUND: Genome-wide Association Studies (GWAS) have contributed to unraveling associations between genetic variants in the human genome and complex traits for more than a decade. While many works have been invented as follow-ups to detect interactions between SNPs, epistasis are still yet to be modeled and discovered more thoroughly. RESULTS: In this paper, following the previous study of detecting marginal epistasis signals, and motivated by the universal approximation power of deep learning, we propose a neural network method that can potentially model arbitrary interactions between SNPs in genetic association studies as an extension to the mixed models in correcting confounding factors. Our method, namely Deep Mixed Model, consists of two components: 1) a confounding factor correction component, which is a large-kernel convolution neural network that focuses on calibrating the residual phenotypes by removing factors such as population stratification, and 2) a fixed-effect estimation component, which mainly consists of an Long-short Term Memory (LSTM) model that estimates the association effect size of SNPs with the residual phenotype. CONCLUSIONS: After validating the performance of our method using simulation experiments, we further apply it to Alzheimer's disease data sets. Our results help gain some explorative understandings of the genetic architecture of Alzheimer's disease.
Haohan Wang, Tianwei Yue, Wei Wu 0023, Eric P. Xing
BMC Bioinform.1
2019 Discovery of Critical Nodes in Road Networks Through Mining From Vehicle Trajectories
abstract
Road networks are extremely vulnerable to cascading failure caused by traffic accidents or anomalous events. Therefore, accurate identification of critical nodes, whose failure may cause a dramatic reduction in the road network transmission efficiency, is of great significance to traffic management and control schemes. However, none of the existing approaches can locate city-wide critical nodes in real road networks. In this paper, we propose a novel data-driven framework to rank node importance through mining from comprehensive vehicle trajectory data, instead of analysis solely on the topology of the road network. In this framework, we introduce a trip network modeled by a tripartite graph to characterize the dynamics of the road network. Furthermore, we present two algorithms, integrating the origin-destination entropy with flow (ODEF) algorithm and the crossroad-rank (CRRank) algorithm, to better exploit the information included in the tripartite graph and to effectively assess the node importance. ODEF absorbs the idea of the information entropy to evaluate the centrality of a node and to calculate its importance rating by integrating its centrality with the traffic flow. CRRank is a ranking algorithm based on eigenvector centrality that captures the mutual reinforcing relationships among the OD-pair, path, and intersection. In addition to the factors considered in ODEF, CRRank considers the irreplaceability of a node and the spatial relationships between neighboring nodes. We conduct a synthetic experiment and a real case study based on a real-world dataset of taxi trajectories. Experiments verify the utility of the proposed algorithms.
Ming Xu 0008, Yunpeng Xiao 0001, Haohan Wang, Dongmei Hu
IEEE Trans. Intell. Transp. Syst.5
2019 Anomaly Detection in Road Networks Using Sliding-Window Tensor Factorization
abstract
Anomaly detection in road networks is vital for traffic management and emergency response. However, existing approaches do not directly address multiple anomaly types. We propose a tensor-based spatio-temporal model for detecting multiple types of anomalies in road networks. First, we represent network traffic data as a 3rd-order tensor. Next, we acquire spatial and multi-scale temporal patterns of traffic variations via a novel, computationally efficient tensor factorization algorithm: sliding window tensor factorization. Then, from the factorization results, we can identify different anomaly types by measuring deviations from different spatial and temporal patterns. Finally, we discover path-level anomalies by formulating anomalous path inference as a linear program that solves for the best matched paths of anomalous links. We evaluate the proposed methods via both synthetic experiments and case studies based on a real-world vehicle trajectory dataset, demonstrating advantages of our approach over baselines.
Ming Xu 0008, Haohan Wang, Mengxin Cao
IEEE Trans. Intell. Transp. Syst.3
2018 Heterogeneous Hi-C Data Super-resolution with a Conditional Generative Adversarial Network
Haohan Wang
BIBM3
2018 3-HBP: A Three-Level Hidden Bayesian Link Prediction Model in Social Networks
abstract
In social networks, link establishment among the users is affected by complex factors. In this paper, we try to investigate the internal and external factors that affect the formation of links and propose a three-level hidden Bayesian link prediction model by integrating the user behavior as well as user relationships to link prediction. First, based on the user multiple interest characteristics, a latent Dirichlet allocation (LDA) traditional text modeling method is applied into user behavior modeling. Taking the advantage of LDA topic model in dealing with the problem of polysemy and synonym, we can mine user latent interest distribution and analyze the effects of internal driving factors. Second, owing to the power-law characteristics of user behavior, LDA is improved by Gaussian weighting. In this way, the negative impact of the interest distribution to the high-frequency users can be reduced and the expression ability of interests can be enhanced. Furthermore, taking the impact of common neighbor dependencies in link establishment, the model can be extended with hidden naive Bayesian algorithm. By quantifying the dependencies between common neighbors, we can analyze the effects of external driving factors and combine internal driving factors to link prediction. Experimental results indicate that the model can not only mine user latent interest distribution but also can improve the performance of link prediction effectively.
Yunpeng Xiao 0001, Haohan Wang, Ming Xu 0008, Yanbing Liu 0004
IEEE Trans. Comput. Soc. Syst.3
2017 Variable selection in heterogeneous datasets: A truncated-rank sparse linear mixed model with applications to genome-wide association studies
abstract
A fundamental and important challenge in modern datasets of ever increasing dimensionality is variable selection, which has taken on renewed interest recently due to the growth of biological and medical datasets with complex, non-i.i.d. structures. Naïvely applying classical variable selection methods such as the Lasso to such datasets may lead to a large number of false discoveries. Motivated by genome-wide association studies in genetics, we study the problem of variable selection for datasets arising from multiple subpopulations, when this underlying population structure is unknown to the researcher. We propose a unified framework for sparse variable selection that adaptively corrects for population structure via a low-rank linear mixed model. Most importantly, the proposed method does not require prior knowledge of individual relationships in the data and adaptively selects a covariance structure of the correct complexity. Through extensive experiments, we illustrate the effectiveness of this framework over existing methods. Further, we test our method on three different genomic datasets from plants, mice, and humans, and discuss the knowledge we discover with our model.
Haohan Wang, Bryon Aragam, Eric P. Xing
BIBM1
2017 Multiplex confounding factor correction for genomic association mapping with squared sparse linear mixed model
abstract
Genome-wide Association Study has presented a promising way to understand the association between human genomes and complex traits. Many simple polymorphic loci have been shown to explain a significant fraction of phenotypic variability. However, challenges remain in the non-triviality of explaining complex traits associated with multifactorial genetic loci, especially considering the confounding factors caused by population structure, family structure, and cryptic relatedness. In this paper, we propose a Squared-LMM (LMM2) model, aiming to jointly correct population and genetic confounding factors. We offer two strategies of utilizing LMM2for association mapping: 1) It serves as an extension of univariate LMM, which could effectively correct population structure, but consider each SNP in isolation. 2) It is integrated with the multivariate regression model to discover association relationship between complex traits and multifactorial genetic loci. We refer to this second model as sparse Squared-LMM (sLMM2). Further, we extend LMM2/sLMM2by raising the power of our squared model to the LMMn/sLMMnmodel. We demonstrate the practical use of our model with synthetic phenotypic variants generated from genetic loci of Arabidopsis Thaliana. The experiment shows that our method achieves a more accurate and significant prediction on the association relationship between traits and loci. We also evaluate our models on collected phenotypes and genotypes with the number of candidate genes that the models could discover. The results suggest the potential and promising usage of our method in genome-wide association studies.
Haohan Wang, Xiang Liu 0017, Yunpeng Xiao 0001, Ming Xu 0008, Eric P. Xing
BIBM1
2017 Select-additive learning: Improving generalization in multimodal sentiment analysis
abstract
Multimodal sentiment analysis is drawing an increasing amount of attention these days. It enables mining of opinions in video reviews which are now available aplenty on online platforms. However, multimodal sentiment analysis has only a few high-quality data sets annotated for training machine learning algorithms. These limited resources restrict the generalizability of models, where, for example, the unique characteristics of a few speakers (e.g., wearing glasses) may become a confounding factor for the sentiment classification task. In this paper, we propose a Select-Additive Learning (SAL) procedure that improves the generalizability of trained neural networks for multimodal sentiment analysis. In our experiments, we show that our SAL approach improves prediction accuracy significantly in all three modalities (verbal, acoustic, visual), as well as in their fusion. Our results show that SAL, even when trained on one dataset, achieves good generalization across two new test datasets.
Haohan Wang, Aaksha Meghawat, Louis-Philippe Morency, Eric P. Xing
ICME1
2016 Multiple confounders correction with regularized linear mixed effect models, with application in biological processes
abstract
In this paper, we inspect the performance of regularized linear mixed effect models, as an extension of linear mixed effect model, when multiple confounding factors coexist. We first review its parameter estimation algorithms before we introduce three different methods for multiple confounding factors correction, namely concatenation, sequence, and interpolation. Then we investigate the performance on variable selection task and predictive task on three different data sets, synthetic data set, semi-empirical synthetic data set based on genome sequences and brain wave data set connecting to confused mental states. Our results suggest that sequence multiple confounding factors corrections behave the best when different confounders contribute equally to response variables. On the other hand, when various confounders affect the response variable unevenly, results mainly rely on the degree of how the major confounder is corrected.
Haohan Wang
BIBM1
2015 Learning structure in gene expression data using deep architectures, with an application to gene clustering
abstract
Genes play a central role in all biological processes. DNA microarray technology has made it possible to study the expression behavior of thousands of genes in one go. Often, gene expression data is used to generate features for supervised and unsupervised learning tasks. At the same time, advances in the field of deep learning have made available a plethora of architectures. In this paper, we use deep architectures pre-trained in an unsupervised manner using denoising autoencoders as a preprocessing step for a popular unsupervised learning task. Denoising autoencoders (DA) can be used to learn a compact representation of input, and have been used to generate features for further supervised learning tasks. We propose that our deep architectures can be treated as empirical versions of Deep Belief Networks (DBNs). We use our deep architectures to regenerate gene expression time series data for two different data sets. We test our hypothesis on two popular datasets for the unsupervised learning task of clustering and find promising improvements in performance.
Haohan Wang, Madhavi Ganapathiraju
BIBM2