EDBT 2026 Demo / reviewers in the wild / expert
Shaohan Huang
dblp:176/0380
· DBLP profile ↗
73ranked-venue papers
13as first author
55since 2021 · last 2026
0000-0003-4324-6337ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 54 · 4 first-author · 43 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 5 since 2021Computer networks · 6 · 2 first-author · 2 since 2021Systems, architecture and hardware · 5 · 2 first-author · 5 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Reasoning with Exploration: An Entropy PerspectiveabstractBalancing exploration and exploitation is a central goal in reinforcement learning (RL). Despite recent advances in enhancing language model (LM) reasoning, most methods lean toward exploitation, and increasingly encounter performance plateaus. In this work, we revisit entropy -- a signal of exploration in RL -- and examine its relationship to exploratory reasoning in LMs. Through empirical analysis, we uncover positive correlations between high-entropy regions and three types of exploratory reasoning actions: (1) pivotal tokens that determine or connect logical steps, (2) reflective actions such as self-verification and correction, and (3) rare behaviors under-explored by the base LMs. Motivated by this, we introduce a minimal modification to standard RL with only one line of code: augmenting the advantage function with an entropy-based term. Unlike traditional maximum-entropy methods which encourage exploration by promoting uncertainty, we encourage exploration by promoting deeper and longer reasoning chains. Notably, our method achieves significant gains on the Pass@K metric -- an upper-bound estimator of LM reasoning capabilities -- even when evaluated with extremely large K values, pushing the boundaries of LM reasoning. Daixuan Cheng, Shaohan Huang, Xuekai Zhu, Bo Dai 0026, Wayne Xin Zhao, Furu Wei |
AAAI | 2 |
| 2026 | VFA: Empowering Multilingual MLLMs via Vision-Free AdaptationabstractYixia Li, Yaqing Shi, Zhiwen Ruan, Dongdong Zhang, Lingjie Jiang, Shaohan Huang, Yun Chen, Guanhua Chen, Furu Wei. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yixia Li, Yaqing Shi, Zhiwen Ruan, Lingjie Jiang, Shaohan Huang, Furu Wei |
ACL (1) | 6 |
| 2026 | SPPO: Sequence-Level PPO for Long-Horizon Reasoning TasksabstractTianyi Wang, Yixia Li, Long Li, Yibiao Chen, Shaohan Huang, Yun Chen, Peng Li, Yang Liu, Guanhua Chen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yixia Li, Yibiao Chen, Shaohan Huang, Yun Chen 0007, Peng Li 0030, Yang Liu 0005, Guanhua Chen 0001 |
ACL (1) | 5 |
| 2026 | Towards Stable and Effective Reinforcement Learning for Mixture-of-ExpertsabstractDi Zhang, Xun Wu, Shaohan Huang, Lingjie Jiang, Yaru Hao, Li Dong, Zewen Chi, Zhifang Sui, Furu Wei. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Shaohan Huang, Lingjie Jiang, Yaru Hao, Li Dong 0004, Zewen Chi, Zhifang Sui, Furu Wei |
ACL (1) | 3 |
| 2025 | NL2Lean: Translating Natural Language into Lean 4 through Multi-Aspect Reinforcement LearningabstractYue Fang, Shaohan Huang, Xin Yu, Haizhen Huang, Zihan Zhang, Weiwei Deng, Furu Wei, Feng Sun, Qi Zhang, Zhi Jin. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Yue Fang 0001, Shaohan Huang, Haizhen Huang, Furu Wei, Feng Sun 0008, Qi Zhang 0066, Zhi Jin 0001 |
EMNLP | 2 |
| 2025 | Textual Aesthetics in Large Language ModelsabstractImage aesthetics is a crucial metric in the field of image generation.However, textual aesthetics has not been sufficiently explored.With the widespread application of large language models (LLMs), previous work has primarily focused on the correctness of content and the helpfulness of responses.Nonetheless, providing responses with textual aesthetics is also an important factor for LLMs, which can offer a cleaner layout and ensure greater consistency and coherence in content.In this work, we introduce a pipeline for aesthetics polishing and help construct a textual aesthetics dataset named TEXAES.We propose a textual aesthetics-powered fine-tuning method based on direct preference optimization, termed TAPO, which leverages textual aesthetics without compromising content correctness.Additionally, we develop two evaluation methods for textual aesthetics based on text and image analysis, respectively.Our experiments demonstrate that using textual aesthetics data and employing the TAPO fine-tuning method not only improves aesthetic scores but also enhances performance on general evaluation datasets such as AlpacalEval and Arena-hard.Our code and data are available at https://github.com/ JackLingjie/ Lingjie Jiang, Shaohan Huang, Furu Wei |
EMNLP | 2 |
| 2025 | Rethinking DPO-Style Diffusion Aligning Frameworks
Shaohan Huang, Lingjie Jiang, Furu Wei |
ICCV | 2 |
| 2025 | Learning to Follow Domain-specific Instruction with Verifiable RewardsabstractIn this paper, we address the challenge of enabling large language models (LLMs) to effectively follow domain-specific instructions, a critical requirement for their successful deployment across various industries. We propose a novel pipeline for constructing verifiable instructions tailored to specific domains. This pipeline consists of three key stages: the creation of meta-requirement templates, the generation of custom instructions using GPT-4 with seed prompts, and manual refinement to ensure clarity, precision, and relevance. A unique aspect of our approach is the incorporation of verifiability into the instruction-following tuning process. Specifically, we design a verified reward mechanism within the Direct Preference Optimization (DPO) framework. This mechanism leverages the ability to automatically verify whether the generated responses adhere to the given instructions. By integrating this verified reward, we enable more effective alignment of LLM behavior with domain-specific requirements, ensuring higher reliability and consistency in outputs. Our study also explores various strategies to enhance the instruction-following capabilities of LLMs, with a focus on fine-tuning methodologies and data augmentation techniques. We provide a comprehensive analysis of domain-specific requirements to better understand how LLMs can be adapted for practical, real-world applications. The efficacy of our approach is empirically validated on GPT-4 and the LLaMA2 series. Notably, the LLaMA-7B model demonstrates a significant performance improvement of over 19% compared to zero-shot settings, underscoring the effectiveness of our methods. This work contributes to the field by bridging the gap between the general capabilities of LLMs and the nuanced demands of domain-specific instruction following. Our findings pave the way for more reliable and adaptable LLM applications across diverse industries. Shaohan Huang, Yi Liu 0013, Jiaxing Qi, Hailong Yang 0002, Zhongzhi Luan, Depei Qian 0001 |
IJCNN | 1 |
| 2025 | LogReader: General-Purpose Log Analysis via Open-Source Large Language ModelsabstractLogs play a critical role in recording system behavior. The increasing volume of log data from software-intensive systems requires automated analysis. Researchers have proposed several approaches to automatically analyze logs, including log compression, log parsing, anomaly detection, log question and answering, and log summary. However, previous methods focused on a single task and lacked a general-purpose log analysis capability, which is critical for maintaining high system availability and reliability. This paper explores the potential of open-source Large Language Models (LLMs) as a general-purpose log analysis tool. To do this, we constructed task-oriented prompt datasets according to the characteristics of different tasks. Then, we presented a general-purpose log analysis system called LogReader powered by LLMs with a hybrid instruction tuning strategy. We compared LogReader powered by five different open-source LLMs, and the extensive evaluations demonstrate the potential of LogReader in terms of accuracy, speed, and generalization. Our work systematically explores the potential of LLMs to develop a general log analysis system, contributing to the integration of LLMs into log analysis and improving the efficiency of system maintenance and debugging. Jiaxing Qi, Zhongzhi Luan, Shaohan Huang, Hailong Yang 0002, Depei Qian 0001 |
IJCNN | 3 |
| 2025 | LogMoE: Lightweight Expert Mixture for Cross-System Log Anomaly DetectionabstractRobust anomaly detection in system logs plays a crucial role in maintaining stable and reliable software operations. However, existing methods often struggle to accommodate evolving log formats and distributional shifts across systems, as they heavily rely on large volumes of labeled data, log parsing, and predefined event templates. To address these challenges, we propose LogMoE, a scalable and parsing-free log anomaly detection framework. LogMoE utilizes labeled logs from multiple mature systems to train a set of lightweight expert models, which are integrated via a gating mechanism within a Mixture-of-Experts (MoE) architecture. This design enables LogMoE to generalize effectively to previously unseen target systems. By eliminating the need for log parsing, our approach remains robust against the heterogeneity of log formats and syntactic structures. We conduct extensive evaluations on eight log datasets under varying generalization scenarios: single-system, homogeneous-system, and heterogeneous-system. Experimental results demonstrate that LogMoE consistently achieves robust generalization, particularly under conditions with scarce labeled data in the target system. As such, LogMoE provides a scalable, parsing-free, and generalization-capable solution tailored for complex and continuously evolving software system environments, positioning it as a future-ready approach to log anomaly detection. Jiaxing Qi, Zhongzhi Luan, Shaohan Huang, Carol J. Fung, Aibin Wang, Hailong Yang 0002, Depei Qian 0001 |
ASE | 3 |
| 2025 | Reward Reasoning ModelsabstractReward models play a critical role in guiding large language models toward outputs that align with human expectations. However, an open challenge remains in effectively utilizing test-time compute to enhance reward model performance. In this work, we introduce Reward Reasoning Models (RRMs), which are specifically designed to execute a deliberate reasoning process before generating final rewards. Through chain-of-thought reasoning, RRMs leverage additional test-time compute for complex queries where appropriate rewards are not immediately apparent. To develop RRMs, we implement a reinforcement learning framework that fosters self-evolved reward reasoning capabilities without requiring explicit reasoning traces as training data. Experimental results demonstrate that RRMs achieve superior performance on reward modeling benchmarks across diverse domains. Notably, we show that RRMs can adaptively exploit test-time compute to further improve reward accuracy. The pretrained models are available at https://huggingface.co/Reward-Reasoning. Zewen Chi, Li Dong 0004, Qingxiu Dong, Shaohan Huang, Furu Wei |
NeurIPS | 6 |
| 2025 | Think Only When You Need with Large Hybrid-Reasoning ModelsabstractRecent Large Reasoning Models (LRMs) have shown substantially improved reasoning capabilities over traditional Large Language Models (LLMs) by incorporating extended thinking processes prior to producing final responses. However, excessively lengthy thinking introduces substantial overhead in terms of token consumption and latency, which is unnecessary for simple queries. In this work, we introduce Large Hybrid-Reasoning Models (LHRMs), the first kind of model capable of adaptively determining whether to perform reasoning based on the contextual information of user queries. To achieve this, we propose a two-stage training pipeline comprising Hybrid Fine-Tuning (HFT) as a cold start, followed by online reinforcement learning with the proposed Hybrid Group Policy Optimization (HGPO) to implicitly learn to select the appropriate reasoning mode. Furthermore, we introduce a metric called Hybrid Accuracy to quantitatively assess the model’s capability for hybrid reasoning. Extensive experimental results show that LHRMs can adaptively perform hybrid reasoning on queries of varying difficulty and type. It outperforms existing LRMs and LLMs in reasoning and general capabilities while significantly improving efficiency. Together, our work advocates for a reconsideration of the appropriate use of extended reasoning processes and provides a solid starting point for building hybrid reasoning systems. Lingjie Jiang, Shaohan Huang, Qingxiu Dong, Zewen Chi, Li Dong 0004, Xingxing Zhang 0002, Tengchao Lv, Lei Cui 0001, Furu Wei |
NeurIPS | 3 |
| 2025 | BitNet: 1-bit Pre-training for Large Language ModelsabstractThe increasing size of large language models (LLMs) has posed challenges for deployment and raised concerns about environmental impact due to high energy consumption. Previous research typically applies quantization after pre-training. While these methods avoid the need for model retraining, they often cause notable accuracy loss at extremely low bit-widths. In this work, we explore the feasibility and scalability of 1-bit pre-training. We introduce BitNet b1 and BitNet b1.58, the scalable and stable 1-bit Transformer architecture designed for LLMs. Specifically, we introduce BitLinear as a drop-in replacement of the nn.Linear layer in order to train 1-bit weights from scratch. Experimental results show that BitNet b1 achieves competitive performance, compared to state-of-the-art 8-bit quantization methods and FP16 Transformer baselines. With the ternary weight, BitNet b1.58 matches the half-precision Transformer LLM with the same model size and training tokens in terms of both perplexity and end-task performance, while being significantly more cost-effective in terms of latency, memory, throughput, and energy consumption. More profoundly, BitNet defines a new scaling law and recipe for training new generations of LLMs that are both high-performance and cost-effective. It enables a new computation paradigm and opens the door for designing specific hardware optimized for 1-bit LLMs. Hongyu Wang 0009, Shuming Ma, Lingxiao Ma, Lei Wang 0222, Wenhui Wang 0003, Li Dong 0004, Shaohan Huang, Huaijie Wang, Jilong Xue, Yi Wu 0013, Furu Wei |
J. Mach. Learn. Res. | 7 |
| 2024 | Text Diffusion with Reinforced ConditioningabstractDiffusion models have demonstrated exceptional capability in generating high-quality images, videos, and audio. Due to their adaptiveness in iterative refinement, they provide a strong potential for achieving better non-autoregressive sequence generation. However, existing text diffusion models still fall short in their performance due to a challenge in handling the discreteness of language. This paper thoroughly analyzes text diffusion models and uncovers two significant limitations: degradation of self-conditioning during training and misalignment between training and sampling. Motivated by our findings, we propose a novel Text Diffusion model called TReC, which mitigates the degradation with Reinforced Conditioning and the misalignment by Time-Aware Variance Scaling. Our extensive experiments demonstrate the competitiveness of TReC against autoregressive, non-autoregressive, and diffusion baselines. Moreover, qualitative analysis shows its advanced ability to fully utilize the diffusion process in refining samples. Yuxuan Liu 0011, Tianchi Yang, Shaohan Huang, Haizhen Huang, Furu Wei, Feng Sun 0008, Qi Zhang 0066 |
AAAI | 3 |
| 2024 | HD-Eval: Aligning Large Language Model Evaluators Through Hierarchical Criteria DecompositionabstractYuxuan Liu, Tianchi Yang, Shaohan Huang, Zihan Zhang, Haizhen Huang, Furu Wei, Weiwei Deng, Feng Sun, Qi Zhang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Yuxuan Liu 0011, Tianchi Yang, Shaohan Huang, Haizhen Huang, Furu Wei, Feng Sun 0008, Qi Zhang 0066 |
ACL (1) | 3 |
| 2024 | Calibrating LLM-Based EvaluatorabstractRecent advancements in large language models (LLMs) and their emergent capabilities make LLM a promising reference-free evaluator on the quality of natural language generation, and a competent alternative to human evaluation. However, hindered by the closed-source or high computational demand to host and tune, there is a lack of practice to further calibrate an off-the-shelf LLM-based evaluator towards better human alignment. In this work, we propose AutoCalibrate, a multi-stage, gradient-free approach to automatically calibrate and align an LLM-based evaluator toward human preference. Instead of explicitly modeling human preferences, we first implicitly encompass them within a set of human labels. Then, an initial set of scoring criteria is drafted by the language model itself, leveraging in-context learning on different few-shot examples. To further calibrate this set of criteria, we select the best performers and re-draft them with self-refinement. Our experiments on multiple text quality evaluation datasets illustrate a significant improvement in correlation with expert evaluation through calibration. Our comprehensive qualitative analysis conveys insightful intuitions and observations on the essence of effective scoring criteria. Yuxuan Liu 0011, Tianchi Yang, Shaohan Huang, Haizhen Huang, Furu Wei, Feng Sun 0008, Qi Zhang 0066 |
LREC/COLING | 3 |
| 2024 | Instruction Pre-Training: Language Models are Supervised Multitask LearnersabstractUnsupervised multitask pre-training has been the critical method behind the recent success of language models (LMs).However, supervised multitask learning still holds significant promise, as scaling it in the post-training stage trends towards better generalization.In this paper, we explore supervised multitask pretraining by proposing Instruction Pre-Training, a framework that scalably augments massive raw corpora with instruction-response pairs to pre-train LMs.The instruction-response pairs are generated by an efficient instruction synthesizer built on open-source models.In our experiments, we synthesize 200M instruction-response pairs covering 40+ task categories to verify the effectiveness of Instruction Pre-Training.In pre-training from scratch, Instruction Pre-Training not only consistently enhances pre-trained base models but also benefits more from further instruction tuning.In continual pre-training, Instruction Pre-Training enables Llama3-8B to be comparable to or even outperform Llama3-70B.Our model, code, and data are available at https://github.com/microsoft/LMOps.Ins: When is the finale of season 7? Let's think step by step. Daixuan Cheng, Yuxian Gu, Shaohan Huang, Junyu Bi, Minlie Huang, Furu Wei |
EMNLP | 3 |
| 2024 | Semantic-Aware Log Understanding and AnalysisabstractThe exponential growth in system complexity and the corresponding surge in log data volume necessitate advanced log analysis techniques for efficient system management and anomaly detection. Traditional log understanding and analysis methods often fail to capture the rich semantic context inherent in log messages, leading to suboptimal monitoring and diagnostic capabilities. This paper aims to bridge the semantic gap by integrating cutting-edge semantic technologies into the log analysis pipeline. We leverage natural language processing, information retrieval, and large language models to enrich log data with semantic information, facilitating a deeper understanding of log messages. Our methodology enhances anomaly detection accuracy by utilizing hierarchical contextual information and pre-training technology, and refining log-based QA processes by log retrieval and log reader. Preliminary results demonstrate a significant improvement in identifying and diagnosing system anomalies, as well as in the automated answering log questions. This research not only presents a breakthrough in log data analysis but also sets the stage for future advancements in intelligent system monitoring and proactive fault resolution. Through this semantic-aware approach, we envision a new paradigm in log analysis that transcends traditional machine learning methods, offering a more robust and intuitive understanding of system behaviors and states. Shaohan Huang, Zhongzhi Luan |
HPDC | 1 |
| 2024 | Adapting Large Language Models via Reading ComprehensionabstractWe explore how continued pre-training on domain-specific corpora influences large language models, revealing that training on the raw corpora endows the model with domain knowledge, but drastically hurts its prompting ability for question answering. Taken inspiration from human learning via reading comprehension--practice after reading improves the ability to answer questions based on the learned knowledge--we propose a simple method for transforming raw corpora into reading comprehension texts. Each raw text is enriched with a series of tasks related to its content. Our method, highly scalable and applicable to any pre-training corpora, consistently enhances performance across various tasks in three different domains: biomedicine, finance, and law. Notably, our 7B language model achieves competitive performance with domain-specific models of much larger scales, such as BloombergGPT-50B. Furthermore, we demonstrate that domain-specific reading comprehension texts can improve the model's performance even on general benchmarks, showing the potential to develop a general model across even more domains. Our model, code, and data are available at https://github.com/microsoft/LMOps. Daixuan Cheng, Shaohan Huang, Furu Wei |
ICLR | 2 |
| 2024 | Kosmos-G: Generating Images in Context with Multimodal Large Language ModelsabstractRecent advancements in subject-driven image generation have made significant strides. However, current methods still fall short in diverse application scenarios, as they require test-time tuning and cannot accept interleaved multi-image and text input. These limitations keep them far from the ultimate goal of "image as a foreign language in image generation." This paper presents Kosmos-G, a model that leverages the advanced multimodal perception capabilities of Multimodal Large Language Models (MLLMs) to tackle the aforementioned challenge. Our approach aligns the output space of MLLM with CLIP using the textual modality as an anchor and performs compositional instruction tuning on curated data. Kosmos-G demonstrates an impressive capability of zero-shot subject-driven generation with interleaved multi-image and text input. Notably, the score distillation instruction tuning requires no modifications to the image decoder. This allows for a seamless substitution of CLIP and effortless integration with a myriad of U-Net techniques ranging from fine-grained controls to personalized image decoder variants. We posit Kosmos-G as an initial attempt towards the goal of "image as a foreign language in image generation." Xichen Pan, Li Dong 0004, Shaohan Huang, Zhiliang Peng, Wenhu Chen, Furu Wei |
ICLR | 3 |
| 2024 | Grounding Multimodal Large Language Models to the WorldabstractWe introduce Kosmos-2, a Multimodal Large Language Model (MLLM), enabling new capabilities of perceiving object descriptions (e.g., bounding boxes) and grounding text to the visual world. Specifically, we represent text spans (i.e., referring expressions and noun phrases) as links in Markdown, i.e., [text span](bounding boxes), where object descriptions are sequences of location tokens. To train the model, we construct a large-scale dataset about grounded image-text pairs (GrIT) together with multimodal corpora. In addition to the existing capabilities of MLLMs (e.g., perceiving general modalities, following instructions, and performing in-context learning), Kosmos-2 integrates the grounding capability to downstream applications, while maintaining the conventional capabilities of MLLMs (e.g., perceiving general modalities, following instructions, and performing in-context learning). Kosmos-2 is evaluated on a wide range of tasks, including (i) multimodal grounding, such as referring expression comprehension and phrase grounding, (ii) multimodal referring, such as referring expression generation, (iii) perception-language tasks, and (iv) language understanding and generation. This study sheds a light on the big convergence of language, multimodal perception, and world modeling, which is a key step toward artificial general intelligence. Code can be found in [https://aka.ms/kosmos-2](https://aka.ms/kosmos-2). Zhiliang Peng, Wenhui Wang 0003, Li Dong 0004, Yaru Hao, Shaohan Huang, Shuming Ma, Qixiang Ye, Furu Wei |
ICLR | 5 |
| 2024 | Mixture of LoRA ExpertsabstractLoRA has gained widespread acceptance in the fine-tuning of large pre-trained models to cater to a diverse array of downstream tasks, showcasing notable effectiveness and efficiency, thereby solidifying its position as one of the most prevalent fine-tuning techniques. Due to the modular nature of LoRA's plug-and-play plugins, researchers have delved into the amalgamation of multiple LoRAs to empower models to excel across various downstream tasks. Nonetheless, extant approaches for LoRA fusion grapple with inherent challenges. Direct arithmetic merging may result in the loss of the original pre-trained model's generative capabilities or the distinct identity of LoRAs, thereby yielding suboptimal outcomes. On the other hand, Reference tuning-based fusion exhibits limitations concerning the requisite flexibility for the effective combination of multiple LoRAs. In response to these challenges, this paper introduces the Mixture of LoRA Experts (MoLE) approach, which harnesses hierarchical control and unfettered branch selection. The MoLE approach not only achieves superior LoRA fusion performance in comparison to direct arithmetic merging but also retains the crucial flexibility for combining LoRAs effectively. Extensive experimental evaluations conducted in both the Natural Language Processing (NLP) and Vision \& Language (V\&L) domains substantiate the efficacy of MoLE. Shaohan Huang, Furu Wei |
ICLR | 2 |
| 2024 | KOSMOS-E : Learning to Follow Instruction for Robotic GraspingabstractTuning on instruction-following data has been shown to enhance the capabilities and controllability of language models, but the idea is less explored in the robotic field. In this work, we introduce KOSMOS-E, a Multimodal Large Language Model (MLLM) that leverages instruction-following robotic grasping data to enhance capabilities for precise and intricate robotic grasping maneuvers. To achieve this, we craft a large-scale instruction-following robotic grasping dataset, termed INSTRUCT-GRASP, primarily comprising two aspects: (i) grasp a single object following varying levels of granularity descriptions, e.g., different angles and aspects, and (ii) grasp a specific object within a multi-object environment following specific attributes, e.g., color and shape. Extensive experiments show the effectiveness of KOSMOS-E on robotic grasping tasks across a variety of environments. Shaohan Huang, Li Dong 0004, Wenhui Wang 0003, Shuming Ma, Furu Wei |
IROS | 3 |
| 2024 | You Only Cache Once: Decoder-Decoder Architectures for Language ModelsabstractWe introduce a decoder-decoder architecture, YOCO, for large language models, which only caches key-value pairs once. It consists of two components, i.e., a cross-decoder stacked upon a self-decoder. The self-decoder efficiently encodes global key-value (KV) caches that are reused by the cross-decoder via cross-attention. The overall model behaves like a decoder-only Transformer, although YOCO only caches once. The design substantially reduces GPU memory demands, yet retains global attention capability. Additionally, the computation flow enables prefilling to early exit without changing the final output, thereby significantly speeding up the prefill stage. Experimental results demonstrate that YOCO achieves favorable performance compared to Transformer in various settings of scaling up model size and number of training tokens. We also extend YOCO to 1M context length with near-perfect needle retrieval accuracy. The profiling results show that YOCO improves inference memory, prefill latency, and throughput by orders of magnitude across context lengths and model sizes. Yutao Sun, Li Dong 0004, Shaohan Huang, Wenhui Wang 0003, Shuming Ma, Quanlu Zhang, Jianyong Wang 0001, Furu Wei |
NeurIPS | 4 |
| 2024 | Multi-Head Mixture-of-ExpertsabstractSparse Mixtures of Experts (SMoE) scales model capacity without significant increases in computational costs. However, it exhibits the low expert activation issue, i.e., only a small subset of experts are activated for optimization, leading to suboptimal performance and limiting its effectiveness in learning a larger number of experts in complex tasks. In this paper, we propose Multi-Head Mixture-of-Experts (MH-MoE). MH-MoE split each input token into multiple sub-tokens, then these sub-tokens are assigned to and processed by a diverse set of experts in parallel, and seamlessly reintegrated into the original token form. The above operations enables MH-MoE to significantly enhance expert activation while collectively attend to information from various representation spaces within different experts to deepen context understanding. Besides, it's worth noting that our MH-MoE is straightforward to implement and decouples from other SMoE frameworks, making it easy to integrate with these frameworks for enhanced performance. Extensive experimental results across different parameter scales (300M to 7B) and three pre-training tasks—English-focused language modeling, multi-lingual language modeling and masked multi-modality modeling—along with multiple downstream validation tasks, demonstrate the effectiveness of MH-MoE. Shaohan Huang, Wenhui Wang 0003, Shuming Ma, Li Dong 0004, Furu Wei |
NeurIPS | 2 |
| 2024 | Multimodal Large Language Models Make Text-to-Image Generative Models Align BetterabstractRecent studies have demonstrated the exceptional potentials of leveraging human preference datasets to refine text-to-image generative models, enhancing the alignment between generated images and textual prompts. Despite these advances, current human preference datasets are either prohibitively expensive to construct or suffer from a lack of diversity in preference dimensions, resulting in limited applicability for instruction tuning in open-source text-to-image generative models and hinder further exploration. To address these challenges and promote the alignment of generative models through instruction tuning, we leverage multimodal large language models to create VisionPrefer, a high-quality and fine-grained preference dataset that captures multiple preference aspects. We aggregate feedback from AI annotators across four aspects: prompt-following, aesthetic, fidelity, and harmlessness to construct VisionPrefer. To validate the effectiveness of VisionPrefer, we train a reward model VP-Score over VisionPrefer to guide the training of text-to-image generative models and the preference prediction accuracy of VP-Score is comparable to human annotators. Furthermore, we use two reinforcement learning methods to supervised fine-tune generative models to evaluate the performance of VisionPrefer, and extensive experimental results demonstrate that VisionPrefer significantly improves text-image alignment in compositional image generation across diverse aspects, e.g., aesthetic, and generalizes better than previous human-preference metrics across various image distributions. Moreover, VisionPrefer indicates that the integration of AI-generated synthetic data as a supervisory signal is a promising avenue for achieving improved alignment with human preferences in vision generative models. Shaohan Huang, Guolong Wang 0001, Furu Wei |
NeurIPS | 2 |
| 2024 | Boosting Text-to-Video Generative Model with MLLMs FeedbackabstractRecent advancements in text-to-video generative models, such as Sora, have showcased impressive capabilities. These models have attracted significant interest for their potential applications. However, they often rely on extensive datasets of variable quality, which can result in generated videos that lack aesthetic appeal and do not accurately reflect the input text prompts. A promising approach to mitigate these issues is to leverage Reinforcement Learning from Human Feedback (RLHF), which aims to align the outputs of text-to-video generative with human preferences. However, the considerable costs associated with manual annotation have led to a scarcity of comprehensive preference datasets. In response to this challenge, our study begins by investigating the efficacy of Multimodal Large Language Models (MLLMs) generated annotations in capturing video preferences, discovering a high degree of concordance with human judgments. Building upon this finding, we utilize MLLMs to perform fine-grained video preference annotations across two dimensions, resulting in the creation of VideoPrefer, which includes 135,000 preference annotations. Utilizing this dataset, we introduce VideoRM, the first general-purpose reward model tailored for video preference in the text-to-video domain. Our comprehensive experiments confirm the effectiveness of both VideoPrefer and VideoRM, representing a significant step forward in the field. Shaohan Huang, Guolong Wang 0001, Furu Wei |
NeurIPS | 2 |
| 2024 | Gloss: Guiding Large Language Models to Answer Questions from System LogsabstractSystem logs contain valuable information and they have emerged as one of the most crucial data sources for system monitoring aimed at enhancing service quality. IT support teams and system administrators are in dire need of an intelligent log-based QA system to help them quickly identify, diagnose, and resolve issues. In this paper, we propose a novel method for constructing log-based question-answering (QA) data using large language models, addressing challenges associated with limited dataset size and diversity in existing log-based QA systems. Our pipeline consists of three steps: generating questions, answering log questions, and refining question-answer pairs. The purpose of the generating questions is to create a diverse set of log-related queries that cover a wide range of potential issues. The second step, answering log questions, aims to extract relevant information from the logs to address the generated questions. This step ensures accurate and context-aware responses. Refining question-answer pairs is intended to improve the overall quality and consistency of the generated log-based QA data. We present a case study using ChatGPT to generate a new dataset, LogQuAD, containing over 28,000 question-answer pairs derived from more than 31,000 raw logs, representing a significant increase compared to existing datasets like LogQA. In our experimental setting, we sample half of the data as the training set and use memory-effect fine-tuning to fine-tune the model, named Gloss. Experimental results show that our method can generate high-quality log-based QA data, leading to improved performance of log-based QA models. Notably, our fine-tuned 7B model outperforms the LLaMA-65B model. This approach can potentially save valuable time for IT support teams and system administrators, enabling proactive problem resolution and optimal system performance. Shaohan Huang, Yi Liu 0013, Jiaxing Qi, Jing Shang 0001, Zhiwen Xiao, Carol J. Fung, Hailong Yang 0002, Zhongzhi Luan, Depei Qian 0001 |
SANER | 1 |
| 2024 | DeepNet: Scaling Transformers to 1,000 LayersabstractIn this paper, we propose a simple yet effective method to stabilize extremely deep Transformers. Specifically, we introduce a new normalization function (DeepNorm) to modify the residual connection in Transformer, accompanying with theoretically derived initialization. In-depth theoretical analysis shows that model updates can be bounded in a stable way. The proposed method combines the best of two worlds, i.e., good performance of Post-LN and stable training of Pre-LN, makingDeepNorma preferred alternative. We successfully scale Transformers up to 1,000 layers (i.e., 2,500 attention and feed-forward network sublayers) without difficulty, which is one order of magnitude deeper than previous deep Transformers. Extensive experiments demonstrate thatDeepNethas superior performance across various benchmarks, including machine translation, language modeling (i.e., BERT, GPT) and vision pre-training (i.e., BEiT). Remarkably, on a multilingual benchmark with 7,482 translation directions, our 200-layer model with 3.2B parameters significantly outperforms the 48-layer state-of-the-art model with 12B parameters by 5 BLEU points, which indicates a promising scaling direction. Our code is available athttps://aka.ms/torchscale. Hongyu Wang 0009, Shuming Ma, Li Dong 0004, Shaohan Huang, Dongdong Zhang 0001, Furu Wei |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | LogSay: An Efficient Comprehension System for Log Numerical ReasoningabstractWith the growth of smart systems and applications, high volume logs are generated that record important data for system maintenance. System developers are usually required to analyze logs to track the status of the system or applications. Therefore, it is essential to find the answers in large-scale logs when they have some questions. In this work, we design a multi-step“Retriever-Reader”question-answering system, namely LogSay, which aims at predicting answers accurately and efficiently. Our system can not only answers simple questions, such as a segment log or span, but also can answer complex logical questions through numerical reasoning. LogSay has two key components:Log RetrieverandLog Reasoner, and we designed five operators to implement them.Log Retrieveraims at retrieving some relevant logs based on a question. Then,Log Reasonerperforms numerical reasoning to infer the final answer. In addition, due to the lack of available question-answering datasets for system logs, we constructed question-answering datasets based on three public log datasets and will make them publicly available. Our evaluation results show that LogSay outperforms the state-of-the-art works in terms of accuracy and efficiency. Jiaxing Qi, Zhongzhi Luan, Shaohan Huang, Carol J. Fung, Hailong Yang 0002 |
IEEE Trans. Computers | 3 |
| 2024 | SpikeLog: Log-Based Anomaly Detection via Potential-Assisted Spiking Neuron NetworkabstractThe increasing volume and complexity of log data generated by modern systems have made it challenging to analyze and extract useful insights manually. To address this problem, many machine learning methods have been proposed for log-based anomaly detection. However, most of these methods lack interpretability, and their underlying premises do not always reflect real scenarios. In this paper, we consider a more reasonable premise scenario where a large number of logs are unlabeled, while only a small number of anomalous logs are labeled. Moreover, a small proportion of anomaly contamination may be present. To handle this practical scenario, we propose a novel hybrid potential-assisted framework (SpikeLog) using the membrane potential of spiking neurons. SpikeLog adopts a weakly supervised approach to train an anomaly score model, which effectively utilizes a limited number of labeled anomalies alongside abundant unlabeled logs while ensuring computational efficiency without compromising accuracy. Extensive experiments have demonstrated that SpikeLog outperforms baseline methods in terms of performance, robustness, interpretability, and energy consumption. Jiaxing Qi, Zhongzhi Luan, Shaohan Huang, Carol J. Fung, Hailong Yang 0002, Depei Qian 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2023 | MoEC: Mixture of Expert ClustersabstractSparsely Mixture of Experts (MoE) has received great interest due to its promising scaling capability with affordable computational overhead. MoE models convert dense layers into sparse experts, and utilize a gated routing network to make experts conditionally activated. However, as the number of experts grows, MoE with outrageous parameters suffers from overfitting and sparse data allocation. Such problems are especially severe on tasks with limited data, thus hindering the progress towards improving performance by scaling up. We verify that there exists a performance upper bound of scaling up sparse MoE. In this work, we propose Mixture of Expert Clusters — a general approach to enable expert layers to learn more diverse and appropriate knowledge by imposing variance-based constraints on the routing stage. Given this, we could further propose a cluster-level expert dropout strategy specifically designed for the expert cluster structure. Our experiments reveal that MoEC could improve performance on machine translation and natural language understanding tasks. MoEC plays a positive role in mitigating overfitting and sparse data allocation problems, thus fully releasing the potential of large-scale sparse models. Shaohan Huang, Furu Wei |
AAAI | 2 |
| 2023 | Dual-Alignment Pre-training for Cross-lingual Sentence EmbeddingabstractZiheng Li, Shaohan Huang, Zihan Zhang, Zhi-Hong Deng, Qiang Lou, Haizhen Huang, Jian Jiao, Furu Wei, Weiwei Deng, Qi Zhang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Ziheng Li 0003, Shaohan Huang, Zhi-Hong Deng 0001, Qiang Lou, Haizhen Huang, Jian Jiao 0007, Furu Wei, Qi Zhang 0066 |
ACL (1) | 2 |
| 2023 | Beyond English-Centric Bitexts for Better Multilingual Language Representation LearningabstractBarun Patra, Saksham Singhal, Shaohan Huang, Zewen Chi, Li Dong, Furu Wei, Vishrav Chaudhary, Xia Song. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Barun Patra, Saksham Singhal, Shaohan Huang, Zewen Chi, Li Dong 0004, Furu Wei, Vishrav Chaudhary |
ACL (1) | 3 |
| 2023 | A Length-Extrapolatable TransformerabstractYutao Sun, Li Dong, Barun Patra, Shuming Ma, Shaohan Huang, Alon Benhaim, Vishrav Chaudhary, Xia Song, Furu Wei. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Yutao Sun, Li Dong 0004, Barun Patra, Shuming Ma, Shaohan Huang, Alon Benhaim, Vishrav Chaudhary, Furu Wei |
ACL (1) | 5 |
| 2023 | GanLM: Encoder-Decoder Pre-training with an Auxiliary DiscriminatorabstractJian Yang, Shuming Ma, Li Dong, Shaohan Huang, Haoyang Huang, Yuwei Yin, Dongdong Zhang, Liqun Yang, Furu Wei, Zhoujun Li. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Jian Yang 0030, Shuming Ma, Li Dong 0004, Shaohan Huang, Haoyang Huang, Yuwei Yin, Dongdong Zhang 0001, Liqun Yang, Furu Wei, Zhoujun Li 0001 |
ACL (1) | 4 |
| 2023 | UPRISE: Universal Prompt Retrieval for Improving Zero-Shot EvaluationabstractDaixuan Cheng, Shaohan Huang, Junyu Bi, Yuefeng Zhan, Jianfeng Liu, Yujing Wang, Hao Sun, Furu Wei, Weiwei Deng, Qi Zhang. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Daixuan Cheng, Shaohan Huang, Junyu Bi, Yuefeng Zhan, Yujing Wang 0002, Hao Sun 0015, Furu Wei, Qi Zhang 0066 |
EMNLP | 2 |
| 2023 | Democratizing Reasoning Ability: Tailored Learning from Large Language ModelabstractZhaoyang Wang, Shaohan Huang, Yuxuan Liu, Jiahai Wang, Minghui Song, Zihan Zhang, Haizhen Huang, Furu Wei, Weiwei Deng, Feng Sun, Qi Zhang. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Shaohan Huang, Yuxuan Liu 0011, Jiahai Wang, Minghui Song, Haizhen Huang, Furu Wei, Feng Sun 0008, Qi Zhang 0066 |
EMNLP | 2 |
| 2023 | Magneto: A Foundation TransformerabstractA big convergence of model architectures across language, vision, speech, and multimodal is emerging. However, under the same name ”Transformers”, the above areas use different implementations for better performance, e.g., Post-LayerNorm for BERT, and Pre-LayerNorm for GPT and vision Transformers. We call for the development of Foundation Transformer for true general-purpose modeling, which serves as a go-to architecture for various tasks and modalities with guaranteed training stability. In this work, we introduce a Transformer variant, named Magneto, to fulfill the goal. Specifically, we propose Sub-LayerNorm for good expressivity, and the initialization strategy theoretically derived from DeepNet for stable scaling up. Extensive experiments demonstrate its superior performance and better stability than the de facto Transformer variants designed for various applications, including language modeling (i.e., BERT, and GPT), machine translation, vision pretraining (i.e., BEiT), speech recognition, and multimodal pretraining (i.e., BEiT-3). Hongyu Wang 0009, Shuming Ma, Shaohan Huang, Li Dong 0004, Wenhui Wang 0003, Zhiliang Peng, Payal Bajaj, Saksham Singhal, Alon Benhaim, Barun Patra, Zhun Liu, Vishrav Chaudhary, Furu Wei |
ICML | 3 |
| 2023 | Language Is Not All You Need: Aligning Perception with Language ModelsabstractA big convergence of language, multimodal perception, action, and world modeling is a key step toward artificial general intelligence. In this work, we introduce KOSMOS-1, a Multimodal Large Language Model (MLLM) that can perceive general modalities, learn in context (i.e., few-shot), and follow instructions (i.e., zero-shot). Specifically, we train KOSMOS-1 from scratch on web-scale multi-modal corpora, including arbitrarily interleaved text and images, image-caption pairs, and text data. We evaluate various settings, including zero-shot, few-shot, and multimodal chain-of-thought prompting, on a wide range of tasks without any gradient updates or finetuning. Experimental results show that KOSMOS-1 achieves impressive performance on (i) language understanding, generation, and even OCR-free NLP (directly fed with document images), (ii) perception-language tasks, including multimodal dialogue, image captioning, visual question answering, and (iii) vision tasks, such as image recognition with descriptions (specifying classification via text instructions). We also show that MLLMs can benefit from cross-modal transfer, i.e., transfer knowledge from language to multimodal, and from multimodal to language. In addition, we introduce a dataset of Raven IQ test, which diagnoses the nonverbal reasoning capability of MLLMs. Shaohan Huang, Li Dong 0004, Wenhui Wang 0003, Yaru Hao, Saksham Singhal, Shuming Ma, Tengchao Lv, Lei Cui 0001, Owais Khan Mohammed, Barun Patra, Kriti Aggarwal, Zewen Chi, Johan Bjorck, Vishrav Chaudhary, Subhojit Som, Furu Wei |
NeurIPS | 1 |
| 2023 | Improving Log-Based Anomaly Detection by Pre-Training Hierarchical TransformersabstractPre-trained models, such as BERT, have resulted in significant pre-trained models, such as BERT, have resulted in significant improvements in many natural language processing (NLP) applications. However, due to differences in word distribution and domain data distribution, applying NLP advancements to log analysis directly faces some performance challenges. This paper studies how to adapt the recently introduced pre-trained language model BERT for log analysis. In this work, we propose a pre-trained log representation model with hierarchical bidirectional encoder transformers (namely, HilBERT). Unlike previous work, which used raw text as pre-training data, we parse logs into templates before using the log templates to pre-train HilBERT. We also design a hierarchical transformers model to capture log template sequence-level information. We use log-based anomaly detection for downstream tasks and fine-tune our model with different log data. Our experiments demonstrate that HilBERT outperforms other baseline techniques on unstable log data. While BERT obtains performance comparable to that of previous state-of-the-art models, HilBERT can significantly address the problem of log instability and achieve accurate and robust results. Shaohan Huang, Yi Liu 0013, Carol J. Fung, Hailong Yang 0002, Zhongzhi Luan |
IEEE Trans. Computers | 1 |
| 2023 | LogEncoder: Log-Based Contrastive Representation Learning for Anomaly DetectionabstractIn recent years, cloud computing centers have grown rapidly in size. Analyzing system logs is an important way for the quality of service monitoring. However, systems produce massive amounts of logs, and it is impractical to analyze them manually. Automatic and accurate log analysis to detect abnormal events in systems has become extremely important. However, due to the nature of the log analysis problem, such as discrete property, class imbalance, and quality of log, log-based anomaly detection remains a difficult problem. To address these challenges, we propose LogEncoder, a framework of log sequence encoding for semi-supervised anomaly detection. LogEncoder utilizes a pre-trained model to obtain a semantic vector for each log event. To separate normal and abnormal log event sequences and preserve their contextual information, we integrate one-class and contrastive learning objectives training into the representation model. Finally, we propose two methods, one for offline and one for online, to detect system anomalies. Compared to six state-of-the-art baselines on three benchmark datasets, LogEncoder outperforms five unsupervised and semi-supervised methods, and the performance is comparable to the supervised method LogRobust. Jiaxing Qi, Zhongzhi Luan, Shaohan Huang, Carol J. Fung, Hailong Yang 0002, Hanlu Li, Danfeng Zhu, Depei Qian 0001 |
IEEE Trans. Netw. Serv. Manag. | 3 |
| 2022 | XLM-E: Cross-lingual Language Model Pre-training via ELECTRAabstractZewen Chi, Shaohan Huang, Li Dong, Shuming Ma, Bo Zheng, Saksham Singhal, Payal Bajaj, Xia Song, Xian-Ling Mao, Heyan Huang, Furu Wei. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Zewen Chi, Shaohan Huang, Li Dong 0004, Shuming Ma, Bo Zheng 0010, Saksham Singhal, Payal Bajaj, Xianling Mao, Heyan Huang, Furu Wei |
ACL (1) | 2 |
| 2022 | Black-box Attacks to Log-based Anomaly DetectionabstractAnomaly detection is the key to Quality of Service (QoS) in many modern systems. Logs, which record the runtime information of system, are widely used for anomaly detection. The security of the log-based anomaly detection has not been well investigated. In this paper, we conduct an empirical study on black-box attacks on log-based anomaly detection. We investigate eight different methods on log attacking and compare their performance on various log parsing methods and log anomaly detection models. We propose a method to evaluate the imperceptibility of log attacking methods. In our experiments, we evaluate the performance on the attack methods on two real log datasets. The results of our experiments show that LogBug outperforms the others in almost all situations. We also compare the imperceptibility of various attack methods and find a trade-off between performance and imperceptibility, where better attack performance means worse imperceptibility. To the best of our knowledge, this is the first work to investigate and compare the attack models on log-based anomaly detection. Shaohan Huang, Yi Liu 0013, Carol J. Fung, Hailong Yang 0002, Zhongzhi Luan |
CNSM | 1 |
| 2022 | PromptBERT: Improving BERT Sentence Embeddings with PromptsabstractTing Jiang, Jian Jiao, Shaohan Huang, Zihan Zhang, Deqing Wang, Fuzhen Zhuang, Furu Wei, Haizhen Huang, Denvy Deng, Qi Zhang. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Jian Jiao 0007, Shaohan Huang, Fuzhen Zhuang, Furu Wei, Haizhen Huang, Denvy Deng, Qi Zhang 0066 |
EMNLP | 3 |
| 2022 | Learning Music Sequence Representation From Text SupervisionabstractMusic representation learning is notoriously difficult for its complex human-related concepts contained in the sequence of numerical signals. To excavate better MUsic SEquence Representation from labeled audio, we propose a novel text-supervision pre-training method, namely MUSER. MUSER adopts an audio-spectrum-text tri-modal contrastive learning framework, where the text input could be any form of meta-data with the help of text templates while the spectrum is derived from an audio sequence. Our experiments reveal that MUSER could be more flexibly adapted to downstream tasks compared with the current data-hungry pre-training method, and it only requires 0.056% of pre-training data to achieve the state-of-the-art performance. Tianyu Chen 0017, Shuai Zhang 0026, Shaohan Huang, Haoyi Zhou, Jianxin Li 0002 |
ICASSP | 4 |
| 2022 | On the Representation Collapse of Sparse Mixture of ExpertsabstractSparse mixture of experts provides larger model capacity while requiring a constant computational overhead. It employs the routing mechanism to distribute input tokens to the best-matched experts according to their hidden representations. However, learning such a routing mechanism encourages token clustering around expert centroids, implying a trend toward representation collapse. In this work, we propose to estimate the routing scores between tokens and experts on a low-dimensional hypersphere. We conduct extensive experiments on cross-lingual language model pre-training and fine-tuning on downstream tasks. Experimental results across seven multilingual benchmarks show that our method achieves consistent gains. We also present a comprehensive analysis on the representation and routing behaviors of our models. Our method alleviates the representation collapse issue and achieves more consistent routing than the baseline mixture-of-experts methods. Zewen Chi, Li Dong 0004, Shaohan Huang, Damai Dai, Shuming Ma, Barun Patra, Saksham Singhal, Payal Bajaj, Xianling Mao, Heyan Huang, Furu Wei |
NeurIPS | 3 |
| 2022 | Kformer: Knowledge Injection in Transformer Feed-Forward Layers
Yunzhi Yao, Shaohan Huang, Li Dong 0004, Furu Wei, Huajun Chen, Ningyu Zhang 0001 |
NLPCC (1) | 2 |
| 2022 | Adanomaly: Adaptive Anomaly Detection for System Logs with Adversarial LearningabstractLogs are commonly used to record the running status of application service systems. Log-based anomaly detection in the system can significantly improve the quality of system services by avoiding catastrophic failures. However, existing log-based anomaly detection methods do not consider class imbalance, which is a common challenge in anomaly detection. In addition, existing methods require hyperparameters in the detection stage, which negatively impacts the accuracy of detection. In this paper, we propose a novel log-based anomaly detection method named Adanomaly, which uses the BiGAN model to extract features and use the ensemble method to detect anomalies. Experimental demonstrate that Adanomaly can detect system abnormalities efficiently, and outperform recall and accuracy compared to other methods. Jiaxing Qi, Zhongzhi Luan, Shaohan Huang, Carol J. Fung, Hailong Yang 0002, Depei Qian 0001 |
NOMS | 3 |
| 2022 | REVAL: Recommend Which Variables to Log With Pretrained Model and Graph Neural NetworkabstractVariable logging plays a vital role in software service management. Developers usually print a set of selected variables in logs to record software system status. Due to the lack of strict logging instructions and domain-specific knowledge, it is challenging for developers to decide which variables to log. Therefore, a technology that enables developers to log high- quality log variables is desirable. There are two reasons that make such a technology feasible. First, there exists semantic relevance between logged variables and other code statements. Second, the structural relationship between variables helps technology learn more information. In this paper, we propose a novel method to recommend variables to log — given a code snippet that needs to be followed by a logging statement, our method will tag every token in this code snippet to indicate whether it should be logged. Our method utilizes a pre-trained model to encode semantic information and a graph neural network to encode graph structure information. Given a code snippet without logging statements, our method first extracts graph structure information by graph neural network, then fuses the graph structure information with semantic information extracted by the pre-trained model to recommend logging variables. We use nine open-source projects’ java files to evaluate our method. The experimental results demonstrate that our method outperforms other baseline methods in terms of Hits@1, MRR, and MAP, which indicate that the quality of the first recommended variable and all recommended variables is superior to other baseline models. Moreover this benefits from encoding better semantic information and incorporating graph structure information. Shaozhi Dai, Zhongzhi Luan, Shaohan Huang, Carol J. Fung, Hailong Yang 0002, Depei Qian 0001 |
IEEE Trans. Netw. Serv. Manag. | 3 |
| 2021 | Improving Pretrained Cross-Lingual Language Models via Self-Labeled Word AlignmentabstractZewen Chi, Li Dong, Bo Zheng, Shaohan Huang, Xian-Ling Mao, Heyan Huang, Furu Wei. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Zewen Chi, Li Dong 0004, Bo Zheng 0010, Shaohan Huang, Xianling Mao, Heyan Huang, Furu Wei |
ACL/IJCNLP (1) | 4 |
| 2021 | Consistency Regularization for Cross-Lingual Fine-TuningabstractBo Zheng, Li Dong, Shaohan Huang, Wenhui Wang, Zewen Chi, Saksham Singhal, Wanxiang Che, Ting Liu, Xia Song, Furu Wei. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Bo Zheng 0010, Li Dong 0004, Shaohan Huang, Wenhui Wang 0003, Zewen Chi, Saksham Singhal, Wanxiang Che, Ting Liu 0001, Furu Wei |
ACL/IJCNLP (1) | 3 |
| 2021 | PriPro: Towards Effective Privacy Protection on Edge-Cloud System running DNN InferenceabstractThe huge computation demand for deep learning models and limited computation resources on the edge devices calls for the cooperation between the edge device and cloud service. On a typical edge-cloud system accommodating DNN inference, a deep model is split into two partial models running on the edge device and the cloud service, respectively. The two partial models collaborate closely to satisfy the DNN inference requested by the user. However, user's privacy is vulnerable when transferring the intermediate results generated by the partial model at edge device to cloud service. Existing research works rely on metrics that are either impractical or insufficient to measure the effectiveness of privacy protection methods in the above scenario, especially from a single input aspect. In this paper, we first thoroughly analyze the state-of-the-art methods and drawbacks of existing methods from the aspects of both evaluation metrics and proposed techniques. Then, we propose a new metric system, including privacy accuracy (PA) and privacy index (PI), that can accurately measure the effectiveness of privacy protection methods. Furthermore, we propose PriPro, a privacy protection method that can dynamically inject noise to the intermediate results at various layers regarding the input features through the self-attention mechanism. The experiment results demonstrate our method outperforms existing methods for protecting user privacy on deep models such as AlexNet, VGG, and ResNet. Ruiyuan Gao 0001, Hailong Yang 0002, Shaohan Huang, Ming Dun, Mingzhen Li 0001, Zerong Luan, Zhongzhi Luan, Depei Qian 0001 |
CCGRID | 3 |
| 2021 | mT6: Multilingual Pretrained Text-to-Text Transformer with Translation PairsabstractZewen Chi, Li Dong, Shuming Ma, Shaohan Huang, Saksham Singhal, Xian-Ling Mao, Heyan Huang, Xia Song, Furu Wei. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021. Zewen Chi, Li Dong 0004, Shuming Ma, Shaohan Huang, Saksham Singhal, Xianling Mao, Heyan Huang, Furu Wei |
EMNLP (1) | 4 |
| 2021 | Allocating Large Vocabulary Capacity for Cross-Lingual Language Model Pre-TrainingabstractCompared to monolingual models, crosslingual models usually require a more expressive vocabulary to represent all languages adequately.We find that many languages are under-represented in recent cross-lingual language models due to the limited vocabulary capacity.To this end, we propose an algorithm VOCAP to determine the desired vocabulary capacity of each language.However, increasing the vocabulary size significantly slows down the pre-training speed.In order to address the issues, we propose k-NN-based target sampling to accelerate the expensive softmax.Our experiments show that the multilingual vocabulary learned with VOCAP benefits cross-lingual language model pre-training.Moreover, k-NN-based target sampling mitigates the side-effects of increasing the vocabulary size while achieving comparable performance and faster pre-training speed.The code and the pretrained multilingual vocabularies are available at https://github. com/bozheng-hit/VoCapXLM. Bo Zheng 0010, Li Dong 0004, Shaohan Huang, Saksham Singhal, Wanxiang Che, Ting Liu 0001, Furu Wei |
EMNLP (1) | 3 |
| 2020 | Transfer Log-based Anomaly Detection with Pseudo LabelsabstractLog-based anomaly detection is an important task for service management and system maintenance. Although anomaly labels are valuable to learn anomaly detection model, they are difficult to collect due to their rarity. To tackle this problem, existing methods employ domain adaptation algorithms to transfer anomaly detectors from labeled source domain to unlabeled target domain. However, most of those methods focus on key performance indicator anomaly detection. The semantic information in logs plays an important role in log-based anomaly detection. Therefore, adaptation methods need to consider how to transfer the semantic information in logs. In this paper, we propose a simple and effective adaptation method to transfer log-based anomaly detection model with pseudo labels. In our work, we first train a detection model with labeled samples as a pseudo-label annotator. Then we use it to assign pseudo-labels to unlabeled samples and train anomaly detectors as if they are true labels. Both models share the same feature extraction part, which can help model to transfer the semantic information in logs. We evaluated our proposed method on three log datasets. Our experimental results demonstrate that our method has outperformed other baseline methods. Shaohan Huang, Yi Liu 0013, Carol J. Fung, Hailong Yang 0002, Zhongzhi Luan |
CNSM | 1 |
| 2020 | Unsupervised Fine-tuning for Text ClusteringabstractFine-tuning with pre-trained language models (e.g.BERT) has achieved great success in many language understanding tasks in supervised settings (e.g.text classification).However, relatively little work has been focused on applying pre-trained models in unsupervised settings, such as text clustering.In this paper, we propose a novel method to fine-tune pre-trained models unsupervisedly for text clustering, which simultaneously learns text representations and cluster assignments using a clustering oriented loss.Experiments on three text clustering datasets (namely TREC-6, Yelp, and DBpedia) show that our model outperforms the baseline methods and achieves stateof-the-art results. Shaohan Huang, Furu Wei, Lei Cui 0001, Xingxing Zhang 0002, Ming Zhou 0001 |
COLING | 1 |
| 2020 | DocBank: A Benchmark Dataset for Document Layout AnalysisabstractDocument layout analysis usually relies on computer vision models to understand documents while ignoring textual information that is vital to capture.Meanwhile, high quality labeled datasets with both visual and textual information are still insufficient.In this paper, we present DocBank, a benchmark dataset that contains 500K document pages with fine-grained tokenlevel annotations for document layout analysis.DocBank is constructed using a simple yet effective way with weak supervision from the L A T E X documents available on the arXiv.com.With DocBank, models from different modalities can be compared fairly and multi-modal approaches will be further investigated and boost the performance of document layout analysis.We build several strong baselines and manually split train/dev/test sets for evaluation.Experiment results show that models trained on DocBank accurately recognize the layout information for a variety of documents.The DocBank dataset is publicly available at https: //github.com/doc-analysis/DocBank. Minghao Li 0004, Yiheng Xu, Lei Cui 0001, Shaohan Huang, Furu Wei, Zhoujun Li 0001, Ming Zhou 0001 |
COLING | 4 |
| 2020 | Language Generation with Multi-Hop Reasoning on Commonsense Knowledge GraphabstractDespite the success of generative pre-trained language models on a series of text generation tasks, they still suffer in cases where reasoning over underlying commonsense knowledge is required during generation.Existing approaches that integrate commonsense knowledge into generative pre-trained language models simply transfer relational knowledge by post-training on individual knowledge triples while ignoring rich connections within the knowledge graph.We argue that exploiting both the structural and semantic information of the knowledge graph facilitates commonsenseaware text generation.In this paper, we propose Generation with Multi-Hop Reasoning Flow (GRF) that enables pre-trained models with dynamic multi-hop reasoning on multirelational paths extracted from the external commonsense knowledge graph.We empirically show that our model outperforms existing baselines on three text generation tasks that require reasoning over commonsense knowledge.We also demonstrate the effectiveness of the dynamic multi-hop reasoning module with reasoning paths inferred by the model that provide rationale to the generation. 1 Haozhe Ji, Pei Ke, Shaohan Huang, Furu Wei, Xiaoyan Zhu 0001, Minlie Huang |
EMNLP (1) | 3 |
| 2020 | LayoutLM: Pre-training of Text and Layout for Document Image UnderstandingabstractPre-training techniques have been verified successfully in a variety of NLP tasks in recent years. Despite the widespread use of pre-training models for NLP applications, they almost exclusively focus on text-level manipulation, while neglecting layout and style information that is vital for document image understanding. In this paper, we propose the LayoutLM to jointly model interactions between text and layout information across scanned document images, which is beneficial for a great number of real-world document image understanding tasks such as information extraction from scanned documents. Furthermore, we also leverage image features to incorporate words' visual information into LayoutLM. To the best of our knowledge, this is the first time that text and layout are jointly learned in a single framework for document-level pre-training. It achieves new state-of-the-art results in several downstream tasks, including form understanding (from 70.72 to 79.27), receipt understanding (from 94.02 to 95.24) and document image classification (from 93.07 to 94.42). The code and pre-trained LayoutLM models are publicly available at https://aka.ms/layoutlm. Yiheng Xu, Minghao Li 0004, Lei Cui 0001, Shaohan Huang, Furu Wei, Ming Zhou 0001 |
KDD | 4 |
| 2020 | TableBank: Table Benchmark for Image-based Table Detection and RecognitionabstractWe present TableBank, a new image-based table detection and recognition dataset built with novel weak supervision from Word and Latex documents on the internet. Existing research for image-based table detection and recognition usually fine-tunes pre-trained models on out-of-domain data with a few thousand human-labeled examples, which is difficult to generalize on real-world applications. With TableBank that contains 417K high quality labeled tables, we build several strong baselines using state-of-the-art models with deep neural networks. We make TableBank publicly available and hope it will empower more deep learning approaches in the table detection and recognition task. The dataset and models can be downloaded from https://github.com/doc-analysis/TableBank. Minghao Li 0004, Lei Cui 0001, Shaohan Huang, Furu Wei, Ming Zhou 0001, Zhoujun Li 0001 |
LREC | 3 |
| 2020 | Paddy: An Event Log Parsing Approach using Dynamic DictionaryabstractLarge enterprise systems often produce a large volume of event logs, and event log parsing is an important log management task. The goal of log parsing is to construct log templates from log messages and convert raw log messages into structured log messages. A log parser can help engineers monitor their systems and detect anomalous behaviors and errors. Most existing log parsing methods focus on offline methods, which require all log data to be available before parsing. In addition, the massive volume of log messages makes the process complex and time-consuming. In this paper, we propose Paddy, an online event log parsing method. Paddy uses a dynamic dictionary structure to build an inverted index, which can search the template candidates efficiently with a high rate of recall. The use of Jaccard similarity and length feature to rank candidates can improve parsing precision. We evaluated our proposed method on 16 real log datasets from various sources including distributed systems, supercomputers, operating systems, mobile systems, and standalone software. Our experimental results demonstrate that Paddy achieves the highest accuracy on eight data sets out of sixteen datasets compared to other baseline methods. We also evaluated the robustness and runtime efficiency of the methods and the experimental results show that our method Paddy achieves superior stableness and is scalable with a large volume of log messages. Shaohan Huang, Yi Liu 0013, Carol J. Fung, Hailong Yang 0002, Zhongzhi Luan |
NOMS | 1 |
| 2020 | A Joint Sentence Scoring and Selection Framework for Neural Extractive Document SummarizationabstractExtractive document summarization methods aim to extract important sentences to form a summary. Previous works perform this task by first scoring all sentences in the document then selecting most informative ones; while we propose to jointly learn the two steps with a novel end-to-end neural network framework. Specifically, the sentences in the input document are represented as real-valued vectors through a neural document encoder. Then the method builds the output summary by extracting important sentences one by one. Different from previous works, the proposed joint sentence scoring and selection framework directly predicts the relative sentence importance score according to both sentence content and previously selected sentences. We evaluate the proposed framework with two realizations: a hierarchical recurrent neural network based model; and a pre-training based model that uses BERT as the document encoder. Experiments on two datasets show that the proposed joint framework outperforms the state-of-the-art extractive summarization models which treat sentence scoring and selection as two subtasks. Qingyu Zhou, Nan Yang 0002, Furu Wei, Shaohan Huang, Ming Zhou 0001, Tiejun Zhao |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2020 | HitAnomaly: Hierarchical Transformers for Anomaly Detection in System LogabstractEnterprise systems often produce a large volume of logs to record runtime status and events. Anomaly detection from system logs is crucial for service management and system maintenance. Most existing log-based anomaly detection methods use log event indexes parsed from log data to detect anomalies. Those methods cannot handle unseen log templates and lead to inaccurate anomaly detection. Some recent studies focused on the semantics of log templates but ignored the information of parameter values. Therefore, their approaches failed to address the abnormal logs caused by parameter values. In this article, we propose HitAnomaly, a log-based anomaly detection model utilizing a hierarchical transformer structure to model both log template sequences and parameter values. We designed a log sequence encoder and a parameter value encoder to obtain their representations correspondingly. We then use an attention mechanism as our final classification model. In this way, HitAnomaly is able to capture the semantic information in both log template sequence and parameter values and handle various types of anomalies. We evaluated our proposed method on three log datasets. Our experimental results demonstrate that HitAnomaly has outperformed other existing log-based anomaly detection methods. We also assess the robustness of our proposed model on unstable log data. Shaohan Huang, Yi Liu 0013, Carol J. Fung, Hailong Yang 0002, Zhongzhi Luan |
IEEE Trans. Netw. Serv. Manag. | 1 |
| 2019 | Response Generation by Context-Aware Prototype EditingabstractOpen domain response generation has achieved remarkable progress in recent years, but sometimes yields short and uninformative responses. We propose a new paradigm, prototypethen-edit for response generation, that first retrieves a prototype response from a pre-defined index and then edits the prototype response according to the differences between the prototype context and current context. Our motivation is that the retrieved prototype provides a good start-point for generation because it is grammatical and informative, and the post-editing process further improves the relevance and coherence of the prototype. In practice, we design a contextaware editing model that is built upon an encoder-decoder framework augmented with an editing vector. We first generate an edit vector by considering lexical differences between a prototype context and current context. After that, the edit vector and the prototype response representation are fed to a decoder to generate a new response. Experiment results on a large scale dataset demonstrate that our new paradigm significantly increases the relevance, diversity and originality of generation results, compared to traditional generative models. Furthermore, our model outperforms retrieval-based methods in terms of relevance and originality. Yu Wu 0012, Furu Wei, Shaohan Huang, Yunli Wang, Zhoujun Li 0001, Ming Zhou 0001 |
AAAI | 3 |
| 2019 | Dictionary-Guided Editing Networks for Paraphrase GenerationabstractAn intuitive way for a human to write paraphrase sentences is to replace words or phrases in the original sentence with their corresponding synonyms and make necessary changes to ensure the new sentences are fluent and grammatically correct. We propose a novel approach to modeling the process with dictionary-guided editing networks which effectively conduct rewriting on the source sentence to generate paraphrase sentences. It jointly learns the selection of the appropriate word level and phrase level paraphrase pairs in the context of the original sentence from an off-the-shelf dictionary as well as the generation of fluent natural language sentences. Specifically, the system retrieves a set of word level and phrase level paraphrase pairs derived from the Paraphrase Database (PPDB) for the original sentence, which is used to guide the decision of which the words might be deleted or inserted with the soft attention mechanism under the sequence-to-sequence framework. We conduct experiments on two benchmark datasets for paraphrase generation, namely the MSCOCO and Quora dataset. The automatic evaluation results demonstrate that our dictionary-guided editing networks outperforms the baseline methods. On human evaluation, results indicate that the generated paraphrases are grammatically correct and relevant to the input sentence. Shaohan Huang, Yu Wu 0012, Furu Wei, Zhongzhi Luan |
AAAI | 1 |
| 2019 | Neural Melody Composition from Lyrics
Hangbo Bao, Shaohan Huang, Furu Wei, Lei Cui 0001, Yu Wu 0012, Chuanqi Tan, Ming Zhou 0001 |
NLPCC (1) | 2 |
| 2018 | Neural Document Summarization by Jointly Learning to Score and Select SentencesabstractSentence scoring and sentence selection are two main steps in extractive document summarization systems.However, previous works treat them as two separated subtasks.In this paper, we present a novel end-to-end neural network framework for extractive document summarization by jointly learning to score and select sentences.It first reads the document sentences with a hierarchical encoder to obtain the representation of sentences.Then it builds the output summary by extracting sentences one by one.Different from previous methods, our approach integrates the selection strategy into the scoring model, which directly predicts the relative importance given previously selected sentences.Experiments on the CNN/Daily Mail dataset show that the proposed framework significantly outperforms the state-of-the-art extractive summarization models. Qingyu Zhou, Nan Yang 0002, Furu Wei, Shaohan Huang, Ming Zhou 0001, Tiejun Zhao |
ACL (1) | 4 |
| 2017 | Arena: Adaptive real-time update anomaly prediction in cloud systemsabstractIn current cloud systems, their monitoring relies strongly on rule-based and supervised-learning-based detection methods for anomaly detection. These methods require either some knowledge provided by an expert system or monitoring data to be labeled as a training set. In practice, the systems behavior changes over time. It is difficult to adjust the rules or re-train detection model for these methods. In this paper, we present an Adaptive REal-time update uNsupervised Anomaly prediction system (Arena) for cloud systems. Arena uses a clustering technique based on a density spatial clustering algorithm to identify clusters and outliers. We propose two prediction strategies to improve the ability to predict anomaly and a real-time update strategy by adding new monitoring points into Arenas model. To improve the prediction efficiency and reduce the scale of the model, we adopt a pruning method to remove redundant points. The anomaly data used in the experiments was collected from the Yahoo Lab and the component based system of enterprise T. The experimental results show that our proposed methods can achieve high prediction accuracy compared to existing methods. Realtime update strategy can improve the prediction performance. The pruning method can further reduce the scale of the model and demonstrates the prediction efficiency. Shaohan Huang, Carol J. Fung, Shupeng Zhang, Guang Wei, Zhongzhi Luan, Depei Qian 0001 |
CNSM | 1 |
| 2017 | Learning to Generate Product Reviews from AttributesabstractLi Dong, Shaohan Huang, Furu Wei, Mirella Lapata, Ming Zhou, Ke Xu. Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 1, Long Papers. 2017. Li Dong 0004, Shaohan Huang, Furu Wei, Mirella Lapata, Ming Zhou 0001, Ke Xu 0001 |
EACL (1) | 2 |
| 2017 | PSOM: Periodic Self-Organizing Maps for unsupervised anomaly detection in periodic time seriesabstractNowadays, systems providing user-oriented services often demonstrate periodic patterns due to the repetitive behaviors from people's daily routines. The monitoring data of such systems are time series of observations that record observed system status at sampled times during each day. The periodic feature and multidimensional character of such monitoring data can be well utilized by anomaly detection algorithms to enhance their detection capability. The data periodicity can be used to provide proactive anomaly prediction capability and the correlation among multidimensional series can provide more accurate results than processing the observations separately. However, existing anomaly detection methods only handle one dimensional series and do not consider the data periodicity. In addition, they often require sufficient labelled data to train the models before they can be used. In this paper, we present an unsupervised anomaly detection algorithm called Periodic Self-Organizing Maps (PSOM) to detect anomalies in periodic time series. PSOMs can be used to detect anomalies in multidimensional periodic series as well as one dimensional periodic series and aperiodic series. Our real data evaluation shows that the PSOM outperforms other supervised methods such as SARIMA and Holt-Winters method. Shupeng Zhang, Carol J. Fung, Shaohan Huang, Zhongzhi Luan, Depei Qian 0001 |
IWQoS | 3 |
| 2016 | Using recurrent neural networks toward black-box system anomaly predictionabstractComponent based enterprise systems are becoming extremely complex in which the availability and usability are influenced intensively by the system's anomalies. Anomaly prediction is highly important for ensuring a system's stability, which aims at preventing anomaly from occurring through pre-failure warning. However, due to the system's complex nature and the noise from monitoring, capturing pre-failure symptoms is a challenging problem. In this paper, we present a sequential and an averaged recurrent neural networks (RNN) models for distributed systems and component based systems. Specifically, we use cycle representation to capture cyclical system behaviors, which can be used to improve prediction accuracy. The anomaly data used in the experiments is collected from RUBis, IBM System S, and the component based system of enterprise T. The experimental results show that our proposed methods can achieve high prediction accuracy with satisfying lead time. Our recurrent neural networks model also demonstrates time efficiency for monitoring large-scale systems. Shaohan Huang, Carol J. Fung, Polo Pei, Zhongzhi Luan, Depei Qian 0001 |
IWQoS | 1 |
| 2015 | A methodology for root-cause analysis in component based systemsabstractIn component based enterprise systems, anomaly detectors are commonly deployed on application-level components, but not on lower-level functional components. When anomaly alarms are triggered, system managers are expected to handle them in a timely manner to avoid cascading failures. Excessive large volume of anomaly alarms makes them impractical to handle manually. Most existing root cause analysis methods are based on the assumption that all components are monitored and analysis are performed based on the time correlation of the generated alarms. However, full monitoring coverage may not be practical due to cost and complexity. In this paper, we present RCSF, a root cause analysis method that targets at systems where only application-level components are monitored by anomaly detectors. The method analyzes the components performance log on functional components and seek for most probable fault propagation sequences based on anomaly analysis. We evaluate the RCSF method based on real enterprise system data and compare it with some baseline methods. Experimental results show that our proposed method can effectively anchor the root causes of failures by providing a short list of most probable causes, and the performance is significantly improved compared to the baseline methods. Carol J. Fung, Polo Pei, Shaohan Huang, Zhongzhi Luan, Depei Qian 0001 |
IWQoS | 5 |