Songfang Huang

dblp:05/4919 · DBLP profile ↗
← Back
65ranked-venue papers
8as first author
36since 2021 · last 2026
0000-0001-8084-0904ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 55 · 6 first-author · 31 since 2021Graphics, computer vision, multimedia, augmented reality and games · 21 · 7 first-author · 8 since 2021Databases, data management, data science and information retrieval · 4 · 2 since 2021Systems, architecture and hardware · 2 · 1 since 2021Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 TTF-VLA: Temporal Token Fusion via Pixel-Attention Integration for Vision-Language-Action Models
abstract
Vision-Language-Action (VLA) models process visual inputs independently at each timestep, discarding valuable temporal information inherent in robotic manipulation tasks. This frame-by-frame processing makes models vulnerable to visual noise while ignoring the substantial coherence between consecutive frames in manipulation sequences. We propose Temporal Token Fusion (TTF), a training-free approach that intelligently integrates historical and current visual representations to enhance VLA inference quality. Our method employs dual-dimension detection combining efficient grayscale pixel difference analysis with attention-based semantic relevance assessment, enabling selective temporal token fusion through hard fusion strategies and keyframe anchoring to prevent error accumulation. Comprehensive experiments across LIBERO, SimplerEnv, and real robot tasks demonstrate consistent improvements: 4.0 percentage points average on LIBERO (72.4% vs 68.4% baseline), cross-environment validation on SimplerEnv (4.8% relative improvement), and 8.7% relative improvement on real robot tasks. Our approach proves model-agnostic, working across OpenVLA and VLA-Cache architectures. Notably, TTF reveals that selective Query matrix reuse in attention mechanisms enhances rather than compromises performance, suggesting promising directions for direct KQV matrix reuse strategies that achieve computational acceleration while improving task success rates.
Chengxuan Li, Zhimu Zhou, Shixin Wu, Songfang Huang, Huiling Duan
AAAI6
2024 Harder Task Needs More Experts: Dynamic Routing in MoE Models
abstract
Quzhe Huang, Zhenwei An, Nan Zhuang, Mingxu Tao, Chen Zhang, Yang Jin, Kun Xu, Kun Xu, Liwei Chen, Songfang Huang, Yansong Feng. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Quzhe Huang, Zhenwei An, Nan Zhuang, Mingxu Tao, Chen Zhang 0019, Kun Xu 0005, Songfang Huang, Yansong Feng 0002
ACL (1)9
2024 Synergetic Event Understanding: A Collaborative Approach to Cross-Document Event Coreference Resolution with Large Language Models
abstract
Cross-document event coreference resolution (CDECR) involves clustering event mentions across multiple documents that refer to the same real-world events.Existing approaches utilize fine-tuning of small language models (SLMs) like BERT to address the compatibility among the contexts of event mentions.However, due to the complexity and diversity of contexts, these models are prone to learning simple co-occurrences.Recently, large language models (LLMs) like ChatGPT have demonstrated impressive contextual understanding, yet they encounter challenges in adapting to specific information extraction (IE) tasks.In this paper, we propose a collaborative approach for CDECR, leveraging the capabilities of both a universally capable LLM and a task-specific SLM.The collaborative strategy begins with the LLM accurately and comprehensively summarizing events through prompting.Then, the SLM refines its learning of event representations based on these insights during fine-tuning.Experimental results demonstrate that our approach surpasses the performance of both the large and small language models individually, forming a complementary advantage.Across various datasets, our approach achieves stateof-the-art performance, underscoring its effectiveness in diverse scenarios.
Qingkai Min, Qipeng Guo, Xiangkun Hu, Songfang Huang
ACL (1)4
2024 Text Diffusion Model with Encoder-Decoder Transformers for Sequence-to-Sequence Generation
abstract
Hongyi Yuan, Zheng Yuan, Chuanqi Tan, Fei Huang, Songfang Huang. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Hongyi Yuan, Chuanqi Tan, Songfang Huang
NAACL-HLT5
2023 Vision Language Pre-training by Contrastive Learning with Cross-Modal Similarity Regulation
abstract
Cross-modal contrastive learning in vision language pretraining (VLP) faces the challenge of (partial) false negatives.In this paper, we study this problem from the perspective of Mutual Information (MI) optimization.It is common sense that InfoNCE loss used in contrastive learning will maximize the lower bound of MI between anchors and their positives, while we theoretically prove that MI involving negatives also matters when noises commonly exist.Guided by a more general lower bound form for optimization, we propose a contrastive learning strategy regulated by progressively refined cross-modal similarity, to more accurately optimize MI between an image/text anchor and its negative texts/images instead of improperly minimizing it.Our method performs competitively on four downstream cross-modal tasks and systematically balances the beneficial and harmful effects of (partial) false negative samples under theoretical guidance.
Chaoya Jiang, Wei Ye 0004, Haiyang Xu 0001, Songfang Huang, Fei Huang 0002, Shikun Zhang
ACL (1)4
2023 Transforming Visual Scene Graphs to Image Captions
abstract
Xu Yang, Jiawei Peng, Zihua Wang, Haiyang Xu, Qinghao Ye, Chenliang Li, Songfang Huang, Fei Huang, Zhangzikang Li, Yu Zhang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Xu Yang 0021, Jiawei Peng 0001, Zihua Wang, Haiyang Xu 0001, Qinghao Ye, Chenliang Li 0003, Songfang Huang, Fei Huang 0002, Zhangzikang Li, Yu Zhang 0004
ACL (1)7
2023 HyPe: Better Pre-trained Language Model Fine-tuning with Hidden Representation Perturbation
abstract
Language models with the Transformers structure have shown great performance in natural language processing.However, there still poses problems when fine-tuning pre-trained language models on downstream tasks, such as over-fitting or representation collapse.In this work, we propose HyPe, a simple yet effective fine-tuning technique to alleviate such problems by perturbing hidden representations of Transformers layers.Unlike previous works that only add noise to inputs or parameters, we argue that the hidden representations of Transformers layers convey more diverse and meaningful language information.Therefore, making the Transformers layers more robust to hidden representation perturbations can further benefit the fine-tuning of PLMs en bloc.We conduct extensive experiments and analyses on GLUE and other natural language inference datasets.Results demonstrate that HyPe outperforms vanilla fine-tuning and enhances generalization of hidden representations from different layers.In addition, HyPe acquires negligible computational overheads, and is better than and compatible with previous state-of-theart fine-tuning techniques.Codes are released at https://github.com/Yuanhy1997/HyPe.
Hongyi Yuan, Zheng Yuan 0002, Chuanqi Tan, Fei Huang 0002, Songfang Huang
ACL (1)5
2023 BUS : Efficient and Effective Vision-language Pre-training with Bottom-Up Patch Summarization
abstract
Vision Transformer (ViT) based Vision-Language Pre-training (VLP) models have demonstrated impressive performance in various tasks. However, the lengthy visual token sequences fed into ViT can lead to training inefficiency and ineffectiveness. Existing efforts address the challenge by either bottom-level patch extraction in the ViT backbone or top-level patch abstraction outside, not balancing training efficiency and effectiveness well. Inspired by text summarization in natural language processing, we propose a Bottom-Up Patch Summarization approach named BUS, coordinating bottom-level extraction and top-level abstraction to learn a concise summary of lengthy visual token sequences efficiently. Specifically, We incorporate a Text-Semantics-Aware Patch Selector (TSPS) into the ViT backbone to perform a coarse-grained visual token extraction and then attach a flexible Transformer-based Patch Abstraction Decoder (PAD) upon the backbone for top-level visual abstraction. This bottom-up collaboration enables our BUS to yield high training efficiency while maintaining or even improving effectiveness. We evaluate our approach on various visual-language understanding and generation tasks and show competitive downstream task performance while boosting the training efficiency by 50%. Additionally, our model achieves state-of-the-art performance on many downstream tasks by increasing input image resolution without increasing computational costs over baselines.
Chaoya Jiang, Haiyang Xu 0001, Wei Ye 0004, Qinghao Ye, Chenliang Li 0003, Ming Yan 0008, Bin Bi, Shikun Zhang, Fei Huang 0002, Songfang Huang
ICCV10
2023 Learning Trajectory-Word Alignments for Video-Language Tasks
abstract
In a video, an object usually appears as the trajectory, i.e., it spans over a few spatial but longer temporal patches, that contains abundant spatiotemporal contexts. However, modern Video-Language BERTs (VDL-BERTs) neglect this trajectory characteristic that they usually follow image-language BERTs (IL-BERTs) to deploy the patch-to-word (P2W) attention that may over-exploit trivial spatial contexts and neglect significant temporal contexts. To amend this, we propose a novel TW-BERT to learn Trajectory-Word alignment by a newly designed trajectory-to-word (T2W) attention for solving video-language tasks. Moreover, previous VDL-BERTs usually uniformly sample a few frames into the model while different trajectories have diverse graininess, i.e., some trajectories span longer frames and some span shorter, and using a few frames will lose certain useful temporal contexts. However, simply sampling more frames will also make pre-training infeasible due to the largely increased training burdens. To alleviate the problem, during the fine-tuning stage, we insert a novel Hierarchical Frame-Selector (HFS) module into the video encoder. HFS gradually selects the suitable frames conditioned on the text context for the later cross-modal encoder to learn better trajectory-word alignments. By the proposed T2W attention and HFS, our TW-BERT achieves SOTA performances on text-to-video retrieval tasks, and comparable performances on video question-answering tasks with some VDL-BERTs trained on much more data. The code will be available in the supplementary material.
Xu Yang 0021, Zhangzikang Li, Haiyang Xu 0001, Hanwang Zhang, Qinghao Ye, Chenliang Li 0003, Ming Yan 0008, Yu Zhang 0004, Fei Huang 0002, Songfang Huang
ICCV10
2023 mPLUG-2: A Modularized Multi-modal Foundation Model Across Text, Image and Video
abstract
Recent years have witnessed a big convergence of language, vision, and multi-modal pretraining. In this work, we present mPLUG-2, a new unified paradigm with modularized design for multi-modal pretraining, which can benefit from modality collaboration while addressing the problem of modality entanglement. In contrast to predominant paradigms of solely relying on sequence-to-sequence generation or encoder-based instance discrimination, mPLUG-2 introduces a multi-module composition network by sharing common universal modules for modality collaboration and disentangling different modality modules to deal with modality entanglement. It is flexible to select different modules for different understanding and generation tasks across all modalities including text, image, and video. Empirical study shows that mPLUG-2 achieves state-of-the-art or competitive results on a broad range of over 30 downstream tasks, spanning multi-modal tasks of image-text and video-text understanding and generation, and uni-modal tasks of text-only, image-only, and video-only understanding. Notably, mPLUG-2 shows new state-of-the-art results of 48.0 top-1 accuracy and 80.3 CIDEr on the challenging MSRVTT video QA and video caption tasks with a far smaller model size and data scale. It also demonstrates strong zero-shot transferability on vision-language and video-language tasks. Code and models will be released in https://github.com/X-PLUG/mPLUG-2.
Haiyang Xu 0001, Qinghao Ye, Ming Yan 0008, Yaya Shi, Jiabo Ye, Yuanhong Xu, Chenliang Li 0003, Bin Bi, Qi Qian 0001, Wei Wang 0225, Guohai Xu, Ji Zhang 0011, Songfang Huang, Fei Huang 0002, Jingren Zhou 0001
ICML13
2023 RAMM: Retrieval-augmented Biomedical Visual Question Answering with Multi-modal Pre-training
abstract
Vision-and-language multi-modal pretraining and fine-tuning have shown great success in visual question answering (VQA). Compared to general domain VQA, the performance of biomedical VQA suffers from limited data. In this paper, we propose a retrieval-augmented pretrain-and-finetune paradigm named RAMM for biomedical VQA to overcome the data limitation issue. Specifically, we collect a new biomedical dataset named PMCPM which offers patient-based image-text pairs containing diverse patient situations from PubMed. Then, we pretrain the biomedical multi-modal model to learn visual and textual representation for image-text pairs and align these representations with image-text contrastive objective (ITC). Finally, we propose a retrieval-augmented method to better use the limited data. We propose to retrieve similar image-text pairs based on ITC from pretraining datasets and introduce a novel retrieval-attention module to fuse the representation of the image and the question with the retrieved images and texts. Experiments demonstrate that our retrieval-augmented pretrain-and-finetune paradigm obtains state-of-the-art performance on Med-VQA2019, Med-VQA2021, VQARAD, and SLAKE datasets. Further analysis shows that the proposed RAMM and PMCPM can enhance biomedical VQA performance compared with previous resources and methods. The pre-trained models and codes are published at https://github.com/GanjinZero/RAMM.
Zheng Yuan 0005, Qiao Jin 0001, Chuanqi Tan, Zhengyun Zhao, Hongyi Yuan, Fei Huang 0002, Songfang Huang
ACM Multimedia7
2023 RRHF: Rank Responses to Align Language Models with Human Feedback
abstract
Reinforcement Learning from Human Feedback (RLHF) facilitates the alignment of large language models with human preferences, significantly enhancing the quality of interactions between humans and models. InstructGPT implements RLHF through several stages, including Supervised Fine-Tuning (SFT), reward model training, and Proximal Policy Optimization (PPO). However, PPO is sensitive to hyperparameters and requires multiple models in its standard implementation, making it hard to train and scale up to larger parameter counts. In contrast, we propose a novel learning paradigm called RRHF, which scores sampled responses from different sources via a logarithm of conditional probabilities and learns to align these probabilities with human preferences through ranking loss. RRHF can leverage sampled responses from various sources including the model responses from itself, other large language model responses, and human expert responses to learn to rank them. RRHF only needs 1 to 2 models during tuning and can efficiently align language models with human preferences robustly without complex hyperparameter tuning. Additionally, RRHF can be considered an extension of SFT and reward model training while being simpler than PPO in terms of coding, model counts, and hyperparameters. We evaluate RRHF on the Helpful and Harmless dataset, demonstrating comparable alignment performance with PPO by reward model score and human labeling. Extensive experiments show that the performance of RRHF is highly related to sampling quality which suggests RRHF is a best-of-$n$ learner.
Hongyi Yuan, Zheng Yuan 0002, Chuanqi Tan, Wei Wang 0225, Songfang Huang
NeurIPS5
2023 Making Pre-trained Language Models End-to-end Few-shot Learners with Contrastive Prompt Tuning
abstract
Pre-trained Language Models (PLMs) have achieved remarkable performance for various language understanding tasks in IR systems, which require the fine-tuning process based on labeled training data. For low-resource scenarios, prompt-based learning for PLMs exploits prompts as task guidance and turns downstream tasks into masked language problems for effective few-shot fine-tuning. In most existing approaches, the high performance of prompt-based learning heavily relies on handcrafted prompts and verbalizers, which may limit the application of such approaches in real-world scenarios. To solve this issue, we present CP-Tuning, an end-to-end Contrastive Prompt Tuning framework for fine-tuning PLMs without any manual engineering of task-specific prompts and verbalizers. It is integrated with the task-invariant continuous prompt encoding technique with fully trainable prompt parameters. We further propose the pair-wise cost-sensitive contrastive learning procedure to optimize the model in order to achieve verbalizer-free class mapping and enhance the task-invariance of prompts. It explicitly learns to distinguish different classes and makes the decision boundary smoother by assigning different costs to easy and hard cases. Experiments over a variety of language understanding tasks and different PLMs show that CP-Tuning outperforms state-of-the-art methods.
Ziyun Xu, Chengyu Wang 0001, Minghui Qiu, Fuli Luo, Runxin Xu, Songfang Huang, Jun Huang 0007
WSDM6
2023 LOGEN: Few-Shot Logical Knowledge-Conditioned Text Generation With Self-Training
abstract
Natural language generation from structured data mainly focuses on surface-level descriptions, suffering from uncontrollable content selection and low fidelity. Previous works leverage logical forms to facilitate logical knowledge-conditioned text generation. Though achieving remarkable progress, they are data-hungry, which makes the adoption for real-world applications challenging with limited data. To this end, this paper proposes a unified framework for logical knowledge-conditioned text generation in the few-shot setting. With only a few seeds logical forms (e.g., 20/100 shot), our approach leverages self-training and samples pseudo logical forms based on content and structure consistency. Experimental results demonstrate that our approach can obtain better few-shot performance than baselines.
Shumin Deng, Hongbin Ye, Chuanqi Tan, Mosha Chen, Songfang Huang, Fei Huang 0002, Huajun Chen, Ningyu Zhang 0001
IEEE ACM Trans. Audio Speech Lang. Process.6
2023 Achieving Human Parity on Visual Question Answering
abstract
The Visual Question Answering (VQA) task utilizes both visual image and language analysis to answer a textual question with respect to an image. It has been a popular research topic with an increasing number of real-world applications in the last decade. This paper introduces a novel hierarchical integration of vision and language AliceMind-MMU (ALIbaba’s Collection of Encoder-decoders from Machine IntelligeNce lab of Damo academy - MultiMedia Understanding) , which leads to similar or even slightly better results than a human being does on VQA. A hierarchical framework is designed to tackle the practical problems of VQA in a cascade manner including: (1) diverse visual semantics learning for comprehensive image content understanding; (2) enhanced multi-modal pre-training with modality adaptive attention; and (3) a knowledge-guided model integration with three specialized expert modules for the complex VQA task. Treating different types of visual questions with corresponding expertise needed plays an important role in boosting the performance of our VQA architecture up to the human level. An extensive set of experiments and analysis are conducted to demonstrate the effectiveness of the new research work.
Ming Yan 0008, Haiyang Xu 0001, Chenliang Li 0003, Bin Bi, Wei Wang 0225, Ji Zhang 0011, Songfang Huang, Fei Huang 0002, Luo Si, Rong Jin 0001
ACM Trans. Inf. Syst.9
2022 From Dense to Sparse: Contrastive Pruning for Better Pre-trained Language Model Compression
abstract
Pre-trained Language Models (PLMs) have achieved great success in various Natural Language Processing (NLP) tasks under the pre-training and fine-tuning paradigm. With large quantities of parameters, PLMs are computation-intensive and resource-hungry. Hence, model pruning has been introduced to compress large-scale PLMs. However, most prior approaches only consider task-specific knowledge towards downstream tasks, but ignore the essential task-agnostic knowledge during pruning, which may cause catastrophic forgetting problem and lead to poor generalization ability. To maintain both task-agnostic and task-specific knowledge in our pruned model, we propose ContrAstive Pruning (CAP) under the paradigm of pre-training and fine-tuning. It is designed as a general framework, compatible with both structured and unstructured pruning. Unified in contrastive learn- ing, CAP enables the pruned model to learn from the pre-trained model for task-agnostic knowledge, and fine-tuned model for task-specific knowledge. Besides, to better retain the performance of the pruned model, the snapshots (i.e., the intermediate models at each pruning iteration) also serve as effective supervisions for pruning. Our extensive experiments show that adopting CAP consistently yields significant improvements, especially in extremely high sparsity scenarios. With only 3% model parameters reserved (i.e., 97% sparsity), CAP successfully achieves 99.2% and 96.3% of the original BERT performance in QQP and MNLI tasks. In addition, our probing experiments demonstrate that the model pruned by CAP tends to achieve better generalization ability.
Runxin Xu, Fuli Luo, Chengyu Wang 0001, Baobao Chang, Jun Huang 0007, Songfang Huang, Fei Huang 0002
AAAI6
2022 Probing Structured Pruning on Multilingual Pre-trained Models: Settings, Algorithms, and Efficiency
abstract
Structured pruning has been extensively studied on monolingual pre-trained language models and is yet to be fully evaluated on their multilingual counterparts.This work investigates three aspects of structured pruning on multilingual pre-trained language models: settings, algorithms, and efficiency.Experiments on nine downstream tasks show several counterintuitive phenomena: for settings, individually pruning for each language does not induce a better result; for algorithms, the simplest method performs the best; for efficiency, a fast model does not imply that it is also small.To facilitate the comparison on all sparsity levels, we present Dynamic Sparsification, a simple approach that allows training the model once and adapting to different model sizes at inference.We hope this work fills the gap in the study of structured pruning on multilingual pre-trained models and sheds light on future research.
Yanyang Li, Fuli Luo, Runxin Xu, Songfang Huang, Fei Huang 0002, Liwei Wang 0009
ACL (1)4
2022 TRIPS: Efficient Vision-and-Language Pre-training with Text-Relevant Image Patch Selection
abstract
Vision Transformers (ViTs) have been widely used in large-scale Vision and Language Pretraining (VLP) models.Though previous VLP works have proved the effectiveness of ViTs, they still suffer from computational efficiency brought by the long visual sequence.To tackle this problem, in this paper, we propose an efficient vision-and-language pre-training model with Text-Relevant Image Patch Selection, namely TRIPS, which reduces the visual sequence progressively with a text-guided patchselection layer in the visual backbone for efficient training and inference.The patchselection layer can dynamically compute textdependent visual attention to identify the attentive image tokens with text guidance and fuse inattentive ones in an end-to-end manner.Meanwhile, TRIPS does not introduce extra parameters to ViTs.Experimental results on a variety of popular benchmark datasets demonstrate that TRIPS gain a speedup of 40% over previous similar VLP models, yet with competitive or better downstream task performance.
Chaoya Jiang, Haiyang Xu 0001, Chenliang Li 0003, Ming Yan 0008, Wei Ye 0004, Shikun Zhang, Bin Bi, Songfang Huang
EMNLP8
2022 mPLUG: Effective and Efficient Vision-Language Learning by Cross-modal Skip-connections
abstract
Chenliang Li, Haiyang Xu, Junfeng Tian, Wei Wang, Ming Yan, Bin Bi, Jiabo Ye, He Chen, Guohai Xu, Zheng Cao, Ji Zhang, Songfang Huang, Fei Huang, Jingren Zhou, Luo Si. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022.
Chenliang Li 0003, Haiyang Xu 0001, Wei Wang 0225, Ming Yan 0008, Bin Bi, Jiabo Ye, Guohai Xu, Zheng Cao 0003, Ji Zhang 0011, Songfang Huang, Fei Huang 0002, Jingren Zhou 0001, Luo Si
EMNLP12
2022 SpanProto: A Two-stage Span-based Prototypical Network for Few-shot Named Entity Recognition
abstract
Few-shot Named Entity Recognition (NER) aims to identify named entities with very little annotated data.Previous methods solve this problem based on token-wise classification, which ignores the information of entity boundaries, and inevitably the performance is affected by the massive non-entity tokens.To this end, we propose a seminal span-based prototypical network (SpanProto) that tackles few-shot NER via a two-stage approach, including span extraction and mention classification.In the span extraction stage, we transform the sequential tags into a global boundary matrix, enabling the model to focus on the explicit boundary information.For mention classification, we leverage prototypical learning to capture the semantic representations for each labeled span and make the model better adapt to novel-class entities.To further improve the model performance, we split out the false positives generated by the span extractor but not labeled in the current episode set, and then present a margin-based loss to separate them from each prototype region.Experiments over multiple benchmarks demonstrate that our model outperforms strong baselines by a large margin. 1
Jianing Wang 0002, Chengyu Wang 0001, Chuanqi Tan, Minghui Qiu, Songfang Huang, Jun Huang 0007, Ming Gao 0001
EMNLP5
2022 Parameter-Efficient Sparsity for Large Language Models Fine-Tuning
abstract
With the dramatically increased number of parameters in language models, sparsity methods have received ever-increasing research focus to compress and accelerate the models. While most research focuses on how to accurately retain appropriate weights while maintaining the performance of the compressed model, there are challenges in the computational overhead and memory footprint of sparse training when compressing large-scale language models. To address this problem, we propose a Parameter-efficient Sparse Training (PST) method to reduce the number of trainable parameters during sparse-aware training in downstream tasks. Specifically, we first combine the data-free and data-driven criteria to efficiently and accurately measure the importance of weights. Then we investigate the intrinsic redundancy of data-driven weight importance and derive two obvious characteristics i.e. low-rankness and structuredness. Based on that, two groups of small matrices are introduced to compute the data-driven importance of weights, instead of using the original large importance score matrix, which therefore makes the sparse training resource-efficient and parameter-efficient. Experiments with diverse networks (i.e. BERT, RoBERTa and GPT-2) on dozens of datasets demonstrate PST performs on par or better than previous sparsity methods, despite only training a small number of parameters. For instance, compared with previous sparsity methods, our PST only requires 1.5% trainable parameters to achieve comparable performance on BERT.
Fuli Luo, Chuanqi Tan, Songfang Huang
IJCAI5
2022 STRONGHOLD: Fast and Affordable Billion-Scale Deep Learning Model Training
abstract
Deep neural networks (DNNs) with billion-scale parameters have demonstrated impressive performance in solving many tasks. Unfortunately, training a billion-scale DNN is out of the reach of many data scientists because it requires high-performance GPU servers that are too expensive to purchase and maintain. We present STRONGHOLD, a novel approach for enabling large DNN model training with no change to the user code. STRONGHOLD scales up the largest trainable model size by dynamically offloading data to the CPU RAM and enabling the use of secondary storage. It automatically determines the minimum amount of data to be kept in the GPU memory to minimize GPU memory usage. Compared to state-of-the-art offloading-based solutions, STRONGHOLD improves the trainable model size by 1.9x~6. Sx on a 32GB V100 GPU, with 1.2x~3.7x improvement on the training throughput. It has been deployed into production to successfully support large-scale DNN training.
Wei Wang 0225, Shenghao Qiu, Renyu Yang, Songfang Huang, Jie Xu 0007, Zheng Wang 0001
SC5
2021 Nested Named Entity Recognition with Partially-Observed TreeCRFs
abstract
Named entity recognition (NER) is a well-studied task in natural language processing. However, the widely-used sequence labeling framework is difficult to detect entities with nested structures. In this work, we view nested NER as constituency parsing with partially-observed trees and model it with partially-observed TreeCRFs. Specifically, we view all labeled entity spans as observed nodes in a constituency tree, and other spans as latent nodes. With the TreeCRF we achieve a uniform way to jointly model the observed and the latent nodes. To compute the probability of partial trees with partial marginalization, we propose a variant of the Inside algorithm, the Masked Inside algorithm, that supports different inference operations for different nodes (evaluation for the observed, marginalization for the latent, and rejection for nodes incompatible with the observed) with efficient parallelized implementation, thus significantly speeding up training and inference. Experiments show that our approach achieves the state-of-the-art (SOTA) F1 scores on the ACE2004, ACE2005 dataset, and shows comparable performance to SOTA models on the GENIA dataset. We release the code at https://github.com/FranxYao/Partially-Observed-TreeCRFs.
Chuanqi Tan, Mosha Chen, Songfang Huang, Fei Huang 0002
AAAI4
2021 A Unified Pretraining Framework for Passage Ranking and Expansion
abstract
Pretrained language models have recently advanced a wide range of natural language processing tasks. Nowadays, the application of pretrained language models to IR tasks has also achieved impressive results. Typical methods either directly apply a pretrained model to improve the re-ranking stage, or use it to conduct passage expansion and term weighting for first-stage retrieval. We observe that the passage ranking and passage expansion tasks share certain inherent relations, and can benefit from each other. Therefore, in this paper, we propose a general pretraining framework to enhance both tasks with Unified Encoder-Decoder networks (UED). The overall ranking framework consists of two parts in a cascade manner: (1) passage expansion with a pretraining-based query generation method; (2) re-ranking of passage candidates from a traditional retrieval method with a pretrained transformer encoder. Both the two parts are based on the same pretrained UED model, where we jointly train the passage ranking and query generation tasks for further improving the full ranking pipeline. An extensive set of experiments have been conducted on two large-scale passage retrieval datasets to demonstrate the state-of-the-art results of the proposed framework in both the first-stage retrieval and the final re-ranking. In addition, we successfully deploy the framework to our online production system, which can stably serve industrial applications with a request volume of up to 100 QPS in less than 300ms.
Ming Yan 0008, Chenliang Li 0003, Bin Bi, Wei Wang 0225, Songfang Huang
AAAI5
2021 StructuralLM: Structural Pre-training for Form Understanding
abstract
Chenliang Li, Bin Bi, Ming Yan, Wei Wang, Songfang Huang, Fei Huang, Luo Si. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Chenliang Li 0003, Bin Bi, Ming Yan 0008, Wei Wang 0225, Songfang Huang, Fei Huang 0002, Luo Si
ACL/IJCNLP (1)5
2021 VECO: Variable and Flexible Cross-lingual Pre-training for Language Understanding and Generation
abstract
Fuli Luo, Wei Wang, Jiahao Liu, Yijia Liu, Bin Bi, Songfang Huang, Fei Huang, Luo Si. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Fuli Luo, Wei Wang 0225, Bin Bi, Songfang Huang, Fei Huang 0002, Luo Si
ACL/IJCNLP (1)6
2021 E2E-VLP: End-to-End Vision-Language Pre-training Enhanced by Visual Learning
abstract
Haiyang Xu, Ming Yan, Chenliang Li, Bin Bi, Songfang Huang, Wenming Xiao, Fei Huang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Haiyang Xu 0001, Ming Yan 0008, Chenliang Li 0003, Bin Bi, Songfang Huang, Wenming Xiao, Fei Huang 0002
ACL/IJCNLP (1)5
2021 Rethinking Denoised Auto-Encoding in Language Pre-Training
abstract
Pre-trained self-supervised models such as BERT have achieved striking success in learning sequence representations, especially for natural language processing.These models typically corrupt the given sequences with certain types of noise, such as masking, shuffling, or substitution, and then try to recover the original input.However, such pre-training approaches are prone to learning representations that are covariant with the noise, leading to the discrepancy between the pre-training and finetuning stage.To remedy this, we present Con-trAstive Pre-Training (CAPT) to learn noise invariant sequence representations.The proposed CAPT encourages the consistency between representations of the original sequence and its corrupted version via unsupervised instance-wise training signals.In this way, it not only alleviates the pretrain-finetune discrepancy induced by the noise of pre-training, but also aids the pre-trained model in better capturing global semantics of the input via more effective sentence-level supervision.Different from most prior work that focuses on a particular modality, comprehensive empirical evidence on 11 natural language understanding and cross-modal tasks illustrates that CAPT is applicable for both language and vision-language tasks, and obtains surprisingly consistent improvement, including 0.6% absolute gain on GLUE benchmarks and 0.8% absolute increment on NLVR 2 . * Equal Contribution.Models Noise types BERT (Devlin et al., 2019) Mask tokens SpanBERT (Joshi et al., 2019) Mask spans RoBERTa (Liu et al., 2019) Mask token XLNet (Yang et al., 2019) Shuffle token ELECTRA (Clark et al., 2019) Replace tokens StructBERT (Wang et al., 2019b) Mask + Shuffle tokens BART (Lewis et al., 2019) Mask + Shuffle + Replace.UNITER (Chen et al., 2019) Mask tokens/regions LXMERT (Tan and Bansal, 2019) Mask tokens/regions
Fuli Luo, Xuancheng Ren, Xu Sun 0001, Songfang Huang, Fei Huang 0002
EMNLP (1)6
2021 Raise a Child in Large Language Model: Towards Effective and Generalizable Fine-tuning
abstract
Recent pretrained language models extend from millions to billions of parameters.Thus the need to fine-tune an extremely large pretrained model with a limited training corpus arises in various downstream tasks.In this paper, we propose a straightforward yet effective fine-tuning technique, CHILD-TUNING, which updates a subset of parameters (called child network) of large pretrained models via strategically masking out the gradients of the non-child network during the backward process.Experiments on various downstream tasks in GLUE benchmark show that CHILD-TUNING consistently outperforms the vanilla fine-tuning by 1.5 ∼ 8.6 average score among four different pretrained models, and surpasses the prior fine-tuning techniques by 0.6 ∼ 1.3 points.Furthermore, empirical results on domain transfer and task transfer show that CHILD-TUNING can obtain better generalization performance by large margins.
Runxin Xu, Fuli Luo, Chuanqi Tan, Baobao Chang, Songfang Huang, Fei Huang 0002
EMNLP (1)6
2021 MELR: Meta-Learning via Modeling Episode-Level Relationships for Few-Shot Learning
Nanyi Fei, Zhiwu Lu 0001, Tao Xiang 0002, Songfang Huang
ICLR4
2021 IEPT: Instance-Level and Episode-Level Pretext Tasks for Few-Shot Learning
Manli Zhang, Zhiwu Lu 0001, Tao Xiang 0002, Mingyu Ding, Songfang Huang
ICLR6
2021 Lattice-BERT: Leveraging Multi-Granularity Representations in Chinese Pre-trained Language Models
abstract
Yuxuan Lai, Yijia Liu, Yansong Feng, Songfang Huang, Dongyan Zhao. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Yuxuan Lai, Yansong Feng 0002, Songfang Huang, Dongyan Zhao 0001
NAACL-HLT4
2021 Noisy-Labeled NER with Confidence Estimation
abstract
Kun Liu, Yao Fu, Chuanqi Tan, Mosha Chen, Ningyu Zhang, Songfang Huang, Sheng Gao. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Chuanqi Tan, Mosha Chen, Ningyu Zhang 0001, Songfang Huang
NAACL-HLT6
2021 Automated conformance testing for JavaScript engines via deep compiler fuzzing
abstract
JavaScript (JS) is a popular, platform-independent programming language. To ensure the interoperability of JS programs across different platforms, the implementation of a JS engine should conform to the ECMAScript standard. However, doing so is challenging as there are many subtle definitions of API behaviors, and the definitions keep evolving.
Guixin Ye, Zhanyong Tang, Shin Hwei Tan, Songfang Huang, Dingyi Fang, Lizhong Bian, Zheng Wang 0001
PLDI4
2021 Contrastive Information Extraction With Generative Transformer
abstract
Information extraction tasks such as triple extraction and event extraction are of great importance for natural language processing and knowledge graph construction. In this paper, we revisit the end-to-end information extraction task for sequence generation. Since generative information extraction may struggle to capture long-term dependencies and generate unfaithful triples, we introduce a novel model, contrastive information extraction with a generative transformer. Specifically, we introduce a single shared transformer module for an encoder-decoder-based generation. To generate faithful results, we propose a novel triplet contrastive training object. Moreover, we introduce two mechanisms to further improve model performance (i.e., batch-wise dynamic attention-masking and triple-wise calibration). Experimental results on five datasets (i.e., NYT, WebNLG, MIE, ACE-2005, and MUC-4) show that our approach achieves better performance than baselines.
Ningyu Zhang 0001, Hongbin Ye, Shumin Deng, Chuanqi Tan, Mosha Chen, Songfang Huang, Fei Huang 0002, Huajun Chen
IEEE ACM Trans. Audio Speech Lang. Process.6
2021 Combining Graph-Based Learning With Automated Data Collection for Code Vulnerability Detection
abstract
This paper presents FUNDED (Flow-sensitive vUl-Nerability coDE Detection), a novel learning framework for building vulnerability detection models. Funded leverages the advances in graph neural networks (GNNs) to develop a novel graph-based learning method to capture and reason about the program's control, data, and call dependencies. Unlike prior work that treats the program as a sequential sequence or an untyped graph, Funded learns and operates on a graph representation of the program source code, in which individual statements are connected to other statements through relational edges. By capturing the program syntax, semantics and flows, Funded finds better code representation for the downstream software vulnerability detection task. To provide sufficient training data to build an effective deep learning model, we combine probabilistic learning and statistical assessments to automatically gather high-quality training samples from open-source projects. This provides many real-life vulnerable code training samples to complement the limited vulnerable code samples available in standard vulnerability databases. We apply Funded to identify software vulnerabilities at the function level from program source code. We evaluate Funded on large real-world datasets with programs written in C, Java, Swift and Php, and compare it against six state-of-the-art code vulnerability detection models. Experimental results show that Funded significantly outperforms alternative approaches across evaluation settings.
Huanting Wang, Guixin Ye, Zhanyong Tang, Shin Hwei Tan, Songfang Huang, Dingyi Fang, Yansong Feng 0002, Lizhong Bian, Zheng Wang 0001
IEEE Trans. Inf. Forensics Secur.5
2020 Deep Program Structure Modeling Through Multi-Relational Graph-based Learning
abstract
Deep learning is emerging as a promising technique for building predictive models to support code-related tasks like performance optimization and code vulnerability detection. One of the critical aspects of building a successful predictive model is having the right representation to characterize the model input for the given task. Existing approaches in the area typically treat the program structure as a sequential sequence but fail to capitalize on the rich semantics of data and control flow information, for which graphs are a proven representation structure.
Guixin Ye, Zhanyong Tang, Huanting Wang, Dingyi Fang, Jianbin Fang, Songfang Huang, Zheng Wang 0001
PACT6
2020 PALM: Pre-training an Autoencoding&Autoregressive Language Model for Context-conditioned Generation
abstract
Self-supervised pre-training, such as BERT (Devlin et al., 2018), MASS (Song et al., 2019) and BART (Lewis et al., 2019), has emerged as a powerful technique for natural language understanding and generation.Existing pre-training techniques employ autoencoding and/or autoregressive objectives to train Transformer-based models by recovering original word tokens from corrupted text with some masked tokens.The training goals of existing techniques are often inconsistent with the goals of many language generation tasks, such as generative question answering and conversational response generation, for producing new text given context.This work presents PALM with a novel scheme that jointly pre-trains an autoencoding and autoregressive language model on a large unlabeled corpus, specifically designed for generating new text conditioned on context.The new scheme alleviates the mismatch introduced by the existing denoising scheme between pre-training and fine-tuning where generation is more than reconstructing original text.An extensive set of experiments show that PALM achieves new state-of-theart results on a variety of language generation benchmarks covering generative question answering (Rank 1 on the official MARCO leaderboard), abstractive summarization on CNN/DailyMail as well as Gigaword, question generation on SQuAD, and conversational response generation on Cornell Movie Dialogues.
Bin Bi, Chenliang Li 0003, Chen Wu 0006, Ming Yan 0008, Wei Wang 0225, Songfang Huang, Fei Huang 0002, Luo Si
EMNLP (1)6
2020 Predicting Clinical Trial Results by Implicit Evidence Integration
abstract
Clinical trials provide essential guidance for practicing Evidence-Based Medicine, though often accompanying with unendurable costs and risks.To optimize the design of clinical trials, we introduce a novel Clinical Trial Result Prediction (CTRP) task.In the CTRP framework, a model takes a PICO-formatted clinical trial proposal with its background as input and predicts the result, i.e. how the Intervention group compares with the Comparison group in terms of the measured Outcome in the studied Population.While structured clinical evidence is prohibitively expensive for manual collection, we exploit large-scale unstructured sentences from medical literature that implicitly contain PICOs and results as evidence.Specifically, we pre-train a model to predict the disentangled results from such implicit evidence and fine-tune the model with limited data on the downstream datasets.Experiments on the benchmark Evidence Integration dataset show that the proposed model outperforms the baselines by large margins, e.g., with a 10.7% relative gain over BioBERT in macro-F1.Moreover, the performance improvement is also validated on another dataset composed of clinical trials related to COVID-19.
Qiao Jin 0001, Chuanqi Tan, Mosha Chen, Xiaozhong Liu 0001, Songfang Huang
EMNLP (1)5
2019 Cross-language document summarization via extraction and ranking of multiple summaries
Xiaojun Wan 0001, Fuli Luo, Songfang Huang, Jin-ge Yao
Knowl. Inf. Syst.4
2018 Marrying Up Regular Expressions with Neural Networks: A Case Study for Spoken Language Understanding
abstract
The success of many natural language processing (NLP) tasks is bound by the number and quality of annotated data, but there is often a shortage of such training data.In this paper, we ask the question: "Can we combine a neural network (NN) with regular expressions (RE) to improve supervised learning for NLP?".In answer, we develop novel methods to exploit the rich expressiveness of REs at different levels within a NN, showing that the combination significantly enhances the learning effectiveness when a small number of training examples are available.We evaluate our approach by applying it to spoken language understanding for intent detection and slot filling.Experimental results show that our approach is highly effective in exploiting the available training data, giving a clear boost to the RE-unaware NN.
Bingfeng Luo, Yansong Feng 0002, Zheng Wang 0001, Songfang Huang, Rui Yan 0001, Dongyan Zhao 0001
ACL (1)4
2018 Encoding implicit relation requirements for relation extraction: A joint inference approach
Yansong Feng 0002, Songfang Huang, Bingfeng Luo, Dongyan Zhao 0001
Artif. Intell.3
2017 FeaBoost: Joint Feature and Label Refinement for Semantic Segmentation
abstract
We propose a novel approach, called FeaBoost, to image semantic segmentation with only image-level labels taken as weakly-supervised constraints. Our approach is motivated from two evidences: 1) each superpixel can be represented as a linear combination of basic components (e.g., predefined classes); 2) visually similar superpixels have high probability to share the same set of labels, i.e., they tend to have common combination of predefined classes. By taking these two evidences into consideration, semantic segmentation is formulated as joint feature and label refinement over superpixels. Furthermore, we develop an efficient FeaBoost algorithm to solve such optimization problem. Extensive experiments on the MSRC and LabelMe datasets demonstrate the superior performance of our FeaBoost approach in comparison with the state-of-the-art methods, especially when noisy labels are provided for semantic segmentation.
Yulei Niu, Zhiwu Lu 0001, Songfang Huang, Xin Gao 0001, Ji-Rong Wen
AAAI3
2017 Learning with Noise: Enhance Distantly Supervised Relation Extraction with Dynamic Transition Matrix
abstract
Bingfeng Luo, Yansong Feng, Zheng Wang, Zhanxing Zhu, Songfang Huang, Rui Yan, Dongyan Zhao. Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2017.
Bingfeng Luo, Yansong Feng 0002, Zheng Wang 0001, Zhanxing Zhu, Songfang Huang, Rui Yan 0001, Dongyan Zhao 0001
ACL (1)5
2016 Question Answering on Freebase via Relation Extraction and Textual Evidence
abstract
Existing knowledge-based question answering systems often rely on small annotated training data.While shallow methods like relation extraction are robust to data scarcity, they are less expressive than the deep meaning representation methods like semantic parsing, thereby failing at answering questions involving multiple constraints.Here we alleviate this problem by empowering a relation extraction method with additional evidence from Wikipedia.We first present a neural network based relation extractor to retrieve the candidate answers from Freebase, and then infer over Wikipedia to validate these answers.Experiments on the WebQuestions question answering dataset show that our method achieves an F 1 of 53.3%, a substantial improvement over the state-of-the-art.
Kun Xu 0005, Siva Reddy, Yansong Feng 0002, Songfang Huang, Dongyan Zhao 0001
ACL (1)4
2016 Sentiment Domain Adaptation with Multi-Level Contextual Sentiment Knowledge
abstract
Sentiment domain adaptation is widely studied to tackle the domain-dependence problem in sentiment analysis field. Existing domain adaptation methods usually train a sentiment classifier in a source domain and adapt it to the target domain using transfer learning techniques. However, when the sentiment feature distributions of the source and target domains are significantly different, the adaptation performance will heavily decline. In this paper, we propose a new sentiment domain adaptation approach by adapting the sentiment knowledge in general-purpose sentiment lexicons to a specific domain. Since the general sentiment words of general-purpose sentiment lexicons usually convey consistent sentiments in different domains, they have better generalization performance than the sentiment classifier trained in a source domain. In addition, we propose to extract various kinds of contextual sentiment knowledge from massive unlabeled samples in target domain and formulate them as sentiment relations among sentiment expressions. It can propagate the sentiment information in general sentiment words to massive domain-specific sentiment expressions. Besides, we propose a unified framework to incorporate these different kinds of sentiment knowledge and learn an accurate domain-specific sentiment classifier for target domain. Moreover, we propose an efficient optimization algorithm to solve the model of our approach. Extensive experiments on benchmark datasets validate the effectiveness and efficiency of our approach.
Fangzhao Wu, Sixing Wu, Yongfeng Huang 0001, Songfang Huang, Yong Qin 0001
CIKM4
2016 Hybrid Question Answering over Knowledge Base and Free Text
abstract
Recent trend in question answering (QA) systems focuses on using structured knowledge bases (KBs) to find answers. While these systems are able to provide more precise answers than information retrieval (IR) based QA systems, the natural incompleteness of KB inevitably limits the question scope that the system can answer. In this paper, we present a hybrid question answering (hybrid-QA) system which exploits both structured knowledge base and free text to answer a question. The main challenge is to recognize the meaning of a question using these two resources, i.e., structured KB and free text. To address this, we map relational phrases to KB predicates and textual relations simultaneously, and further develop an integer linear program (ILP) model to infer on these candidates and provide a globally optimal solution. Experiments on benchmark datasets show that our system can benefit from both structured KB and free text, outperforming the state-of-the-art systems.
Kun Xu 0005, Yansong Feng 0002, Songfang Huang, Dongyan Zhao 0001
COLING3
2016 Segmentation with Selectively Propagated Constraints
Peng Han 0005, Guangzhen Liu, Songfang Huang, Wenwu Yuan, Zhiwu Lu 0001
ICONIP (2)3
2015 Noise-Robust Semi-Supervised Learning by Large-Scale Sparse Coding
abstract
This paper presents a large-scale sparse coding algorithm to deal with the challenging problem of noise-robust semi-supervised learning over very large data with only few noisy initial labels. By giving an L1-norm formulation of Laplacian regularization directly based upon the manifold structure of the data, we transform noise-robust semi-supervised learning into a generalized sparse coding problem so that noise reduction can be imposed upon the noisy initial labels. Furthermore, to keep the scalability of noise-robust semi-supervised learning over very large data, we make use of both nonlinear approximation and dimension reduction techniques to solve this generalized sparse coding problem in linear time and space complexity. Finally, we evaluate the proposed algorithm in the challenging task of large-scale semi-supervised image classification with only few noisy initial labels. The experimental results on several benchmark image datasets show the promising performance of the proposed algorithm.
Zhiwu Lu 0001, Xin Gao 0001, Liwei Wang 0001, Ji-Rong Wen, Songfang Huang
AAAI5
2015 What Is the Longest River in the USA? Semantic Parsing for Aggregation Questions
abstract
Answering natural language questions against structured knowledge bases (KB) has been attracting increasing attention in both IR and NLP communities. The task involves two main challenges: recognizing the questions' meanings, which are then grounded to a given KB. Targeting simple factoid questions, many existing open domain semantic parsers jointly solve these two subtasks, but are usually expensive in complexity and resources.In this paper, we propose a simple pipeline framework to efficiently answer more complicated questions, especially those implying aggregation operations, e.g., argmax, argmin.We first develop a transition-based parsing model to recognize the KB-independent meaning representation of the user's intention inherent in the question. Secondly, we apply a probabilistic model to map the meaning representation, including those aggregation functions, to a structured query.The experimental results showed that our method can better understand aggregation questions, outperforming the state-of-the-art methods on the Free917 dataset while still maintaining promising performance on a more challenging dataset, WebQuestions, without extra training.
Kun Xu 0005, Sheng Zhang 0012, Yansong Feng 0002, Songfang Huang, Dongyan Zhao 0001
AAAI4
2015 Semantic Relation Classification via Convolutional Neural Networks with Simple Negative Sampling
abstract
Syntactic features play an essential role in identifying relationship in a sentence.Previous neural network models directly work on raw word sequences or constituent parse trees, thus often suffer from irrelevant information introduced when subjects and objects are in a long distance.In this paper, we propose to learn more robust relation representations from shortest dependency paths through a convolution neural network.We further take the relation directionality into account and propose a straightforward negative sampling strategy to improve the assignment of subjects and objects.Experimental results show that our method outperforms the state-of-theart approaches on the SemEval-2010 Task 8 dataset.
Kun Xu 0005, Yansong Feng 0002, Songfang Huang, Dongyan Zhao 0001
EMNLP3
2015 Social Image Parsing by Cross-Modal Data Refinement
Zhiwu Lu 0001, Xin Gao 0001, Songfang Huang, Liwei Wang 0001, Ji-Rong Wen
IJCAI3
2015 Weakly Supervised Matrix Factorization for Noisily Tagged Image Parsing
Yulei Niu, Zhiwu Lu 0001, Songfang Huang, Peng Han 0005, Ji-Rong Wen
IJCAI3
2014 Encoding Relation Requirements for Relation Extraction via Joint Inference
abstract
Most existing relation extraction models make predictions for each entity pair locally and individually, while ignoring implicit global clues available in the knowledge base, sometimes leading to conflicts among local predictions from different entity pairs.In this paper, we propose a joint inference framework that utilizes these global clues to resolve disagreements among local predictions.We exploit two kinds of clues to generate constraints which can capture the implicit type and cardinality requirements of a relation.Experimental results on three datasets, in both English and Chinese, show that our framework outperforms the state-of-theart relation extraction models when such clues are applicable to the datasets.And, we find that the clues learnt automatically from existing knowledge bases perform comparably to those refined by human.
Yansong Feng 0002, Songfang Huang, Yong Qin 0001, Dongyan Zhao 0001
ACL (1)3
2014 Joint Inference for Knowledge Base Population
abstract
Populating Knowledge Base (KB) with new knowledge facts from reliable text resources usually consists of linking name mentions to KB entities and identifying relationship between entity pairs.However, the task often suffers from errors propagating from upstream entity linkers to downstream relation extractors.In this paper, we propose a novel joint inference framework to allow interactions between the two subtasks and find an optimal assignment by addressing the coherence among preliminary local predictions: whether the types of entities meet the expectations of relations explicitly or implicitly, and whether the local predictions are globally compatible.We further measure the confidence of the extracted triples by looking at the details of the complete extraction process.Experiments show that the proposed framework can significantly reduce the error propagations thus obtain more reliable facts, and outperforms competitive baselines with state-of-the-art relation extraction models.
Yansong Feng 0002, Jinghui Mo, Songfang Huang, Dongyan Zhao 0001
EMNLP4
2014 Community-based matrix factorization for scalable music recommendation on smartphones
abstract
Mobile karaoke has attracted more attention as a popular mobile entertainment and social network platform, where music recommendations are highly desired to improve its user experiences. Traditional music recommendation methods suffer from the data sparsity issue and usually ignore the social interactions among users. In this paper, we propose a novel parallel community-based matrix factorization method which exploits implicit user behavior data to model user preferences from both social level, via community detection, and individual level. Both offline evaluation on a real dataset from Changba and online traffic investigations show the effectiveness of our method.
Jinghui Mo, Yansong Feng 0002, Aixia Jia, Songfang Huang, Yong Qin 0001, Dongyan Zhao 0001
ICME4
2013 The IBM speech-to-speech translation system for smartphone: Improvements for resource-constrained tasks
Bowen Zhou 0002, Songfang Huang, Martin Cmejrek, Wei Zhang 0057, Jia Cui, Bing Xiang, Gregg Daggett, Upendra V. Chaudhari, Sameer Maskey, Etienne Marcheret
Comput. Speech Lang.3
2011 Using Features from Topic Models to Alleviate Over-Generation in Hierarchical Phrase-Based Translation
Songfang Huang
INTERSPEECH1
2011 An Empirical Study on Improving Hierarchical Phrase-Based Translation Using Alignment Features
Songfang Huang
INTERSPEECH1
2010 Power law discounting for n-gram language models
abstract
We present an approximation to the Bayesian hierarchical Pitman-Yor process language model which maintains the power law distribution over word tokens, while not requiring a computationally expensive approximate inference process. This approximation, which we term power law discounting, has a similar computational complexity to interpolated and modified Kneser-Ney smoothing. We performed experiments on meeting transcription using the NIST RT06s evaluation data and the AMI corpus, with a vocabulary of 50,000 words and a language model training set of up to 211 million words. Our results indicate that power law discounting results in statistically significant reductions in perplexity and word error rate compared to both interpolated and modified Kneser-Ney smoothing, while producing similar results to the hierarchical Pitman-Yor process language model.
Songfang Huang, Steve Renals
ICASSP1
2010 Hierarchical Bayesian Language Models for Conversational Speech Recognition
abstract
Traditional$n$-gram language models are widely used in state-of-the-art large vocabulary speech recognition systems. This simple model suffers from some limitations, such as overfitting of maximum-likelihood estimation and the lack of rich contextual knowledge sources. In this paper, we exploit a hierarchical Bayesian interpretation for language modeling, based on a nonparametric prior called Pitman–Yor process. This offers a principled approach to language model smoothing, embedding the power-law distribution for natural language. Experiments on the recognition of conversational speech in multiparty meetings demonstrate that by using hierarchical Bayesian language models, we are able to achieve significant reductions in perplexity and word error rate.
Songfang Huang, Steve Renals
IEEE Trans. Speech Audio Process.1
2009 An EM algorithm for SCFG in formal syntax-based translation
abstract
In this paper, we investigate the use of bilingual parsing on parallel corpora to better estimate the rule parameters in a formal syntax-based machine translation system, which are normally estimated from the inaccurate heuristics. We use an Expectation-Maximization (EM) algorithm to re-estimate the parameters of synchronous context-free grammar (SCFG) rules according to the derivation knowledge from parallel corpora based on maximum likelihood principle, rather than using only the heuristic information. The proposed algorithm produces significantly better BLEU scores than a state-of-the-art formal syntax-based machine translation system on the IWSLT 2006 Chinese to English task.
Songfang Huang
ICASSP1
2009 A parallel training algorithm for hierarchical pitman-yor process language models
abstract
The Hierarchical Pitman Yor Process Language Model (HPYLM) is a Bayesian language model based on a nonparametric prior, the Pitman-Yor Process. It has been demonstrated, both theoretically and practically, that the HPYLM can provide better smoothing for language modeling, compared with state-of-the-art approaches such as interpolated Kneser-Ney and modified Kneser-Ney smoothing. However, estimation of Bayesian language models is expensive in terms of both computation time and memory; the inference is approximate and requires a number of iterations to converge. In this paper, we present a parallel training algorithm for the HPYLM, which enables the approach to be applied in the context of automatic speech recognition, using large training corpora with large vocabularies. We demonstrate the effectiveness of the proposed algorithm by estimating language models from corpora for meeting transcription containing over 200 million words, and observe significant reductions in perplexity and word error rate. Index Terms: language model, Pitman-Yor processes, hierarchical Bayesian models, parallel training, meetings
Songfang Huang, Steve Renals
INTERSPEECH1
2008 Unsupervised language model adaptation based on topic and role information in multiparty meetings
abstract
We continue our previous work on the modeling of topic and role information from multiparty meetings using a hierarchical Dirichlet process (HDP), in the context of language model adaptation. In this paper we focus on three problems: 1) an empirical analysis of the HDP as a nonparametric topic model; 2) the mismatch problem of vocabularies of the baseline n-gram model and the HDP; and 3) an automatic speech recognition experiment to further verify the effectiveness of our adaptation framework. Experiments on a large meeting corpus of more than 70 hours speech data show consistent and significant improvements in terms of word error rate for language model adaptation based on the topic and role information. Index Terms: language model, adaptation, topic model, hierarchical Dirichlet process, participant role
Songfang Huang, Steve Renals
INTERSPEECH1
2007 Hierarchical Pitman-Yor language models for ASR in meetings
abstract
In this paper we investigate the application of a hierarchical Bayesian language model (LM) based on the Pitman-Yor process for automatic speech recognition (ASR) of multiparty meetings. The hierarchical Pitman-Yor language model (HPYLM) provides a Bayesian interpretation of LM smoothing. An approximation to the HPYLM recovers the exact formulation of the interpolated Kneser-Ney smoothing method in n-gram models. This paper focuses on the application and scalability of HPYLM on a practical large vocabulary ASR system. Experimental results on NIST RT06s evaluation meeting data verify that HPYLM is a competitive and promising language modeling technique, which consistently performs better than interpolated Kneser-Ney and modified Kneser-Ney n-gram LMs in terms of both perplexity and word error rate.
Songfang Huang, Steve Renals
ASRU1