VLDB 2026 Research / reviewers in the wild / expert
Jie Fu 0001
dblp:16/7565-1
· DBLP profile ↗
55ranked-venue papers
1as first author
39since 2021 · last 2025
0000-0002-4494-843XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 52 · 1 first-author · 37 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | OpenCoder: The Open Cookbook for Top-Tier Code Large Language ModelsabstractSiming Huang, Tianhao Cheng, Jason Klein Liu, Weidi Xu, Jiaran Hao, Liuyihan Song, Yang Xu, Jian Yang, Jiaheng Liu, Chenchen Zhang, Linzheng Chai, Ruifeng Yuan, Xianzhen Luo, Qiufeng Wang, YuanTao Fan, Qingfu Zhu, Zhaoxiang Zhang, Yang Gao, Jie Fu, Qian Liu, Houyi Li, Ge Zhang, Yuan Qi, Xu Yinghui, Wei Chu, Zili Wang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Siming Huang, Tianhao Cheng, Jason Klein Liu, Weidi Xu, Jiaran Hao, Liuyihan Song, Jian Yang 0030, Linzheng Chai, Ruifeng Yuan, Xianzhen Luo, YuanTao Fan, Qingfu Zhu, Zhaoxiang Zhang 0001, Yang Gao 0021, Jie Fu 0001, Qian Liu 0033, Houyi Li, Ge Zhang 0009, Yuan Qi 0001 |
ACL (1) | 19 |
| 2025 | PopAlign: Diversifying Contrasting Patterns for a More Comprehensive AlignmentabstractZekun Moore Wang, Shenzhi Wang, King Zhu, Jiaheng Liu, Ke Xu, Jie Fu, Wangchunshu Zhou, Wenhao Huang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Zekun Moore Wang, Shenzhi Wang, King Zhu, Ke Xu 0001, Jie Fu 0001, Wangchunshu Zhou, Wenhao Huang 0001 |
ACL (1) | 6 |
| 2025 | Finite State Automata Inside Transformers with Chain-of-Thought: A Mechanistic Study on State TrackingabstractChain-of-thought (CoT) significantly enhances the performance of large language models (LLMs) across a wide range of tasks, and prior research shows that CoT can theoretically increase expressiveness. However, there is limited mechanistic understanding of the algorithms that Transformer+CoT can learn. Our key contributions are: (1) We evaluate the state tracking capabilities of Transformer+CoT and its variants, confirming the effectiveness of CoT. (2) Next, we identify the circuit (a subset of model components, responsible for tracking the world state), indicating that late-layer MLP neurons play a key role. We propose two metrics, compression and distinction, and show that the neuron sets for each state achieve nearly 100% accuracy, providing evidence of an implicit finite state automaton (FSA) embedded within the model. (3) Additionally, we explore three challenging settings: skipping intermediate steps, introducing data noises, and testing length generalization. Our results demonstrate that Transformer+CoT learns robust algorithms (FSAs), highlighting its resilience in challenging scenarios. Our code is available at https://github.com/IvanChangPKU/FSA. Yifan Zhang 0004, Wenyu Du, Dongming Jin, Jie Fu 0001, Zhi Jin 0001 |
ACL (1) | 4 |
| 2025 | MIO: A Foundation Model on Multimodal TokensabstractZekun Moore Wang, King Zhu, Chunpu Xu, Wangchunshu Zhou, Jiaheng Liu, Yibo Zhang, Jessie Wang, Ning Shi, Siyu Li, Yizhi Li, Haoran Que, Zhaoxiang Zhang, Yuanxing Zhang, Ge Zhang, Ke Xu, Jie Fu, Wenhao Huang. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Zekun Moore Wang, King Zhu, Chunpu Xu, Wangchunshu Zhou, Jessie Jiashuo Wang, Ning Shi, Haoran Que, Zhaoxiang Zhang 0001, Yuanxing Zhang, Ge Zhang 0009, Ke Xu 0001, Jie Fu 0001, Wenhao Huang 0001 |
EMNLP | 16 |
| 2025 | MAP: Low-compute Model Merging with Amortized Pareto Fronts via Quadratic ApproximationabstractModel merging has emerged as an effective approach to combining multiple single-task models into a multitask model. This process typically involves computing a weighted average of the model parameters without additional training. Existing model-merging methods focus on improving average task accuracy. However, interference and conflicts between the objectives of different tasks can lead to trade-offs during the merging process. In real-world applications, a set of solutions with various trade-offs can be more informative, helping practitioners make decisions based on diverse preferences. In this paper, we introduce a novel and low-compute algorithm, Model Merging with Amortized Pareto Front (MAP). MAP efficiently identifies a Pareto set of scaling coefficients for merging multiple models, reflecting the trade-offs involved. It amortizes the substantial computational cost of evaluations needed to estimate the Pareto front by using quadratic approximation surrogate models derived from a preselected set of scaling coefficients. Experimental results on vision and natural language processing tasks demonstrate that MAP can accurately identify the Pareto front, providing practitioners with flexible solutions to balance competing task objectives. We also introduce Bayesian MAP for scenarios with a relatively low number of tasks and Nested MAP for situations with a high number of tasks, further reducing the computational cost of evaluation. Zhiqi Bu, Suyuchen Wang, Jie Fu 0001, Yonghui Wu 0001, Jiang Bian 0002, Yong Chen 0016, Yoshua Bengio |
ICLR | 6 |
| 2025 | Layerwise Recurrent Router for Mixture-of-ExpertsabstractThe scaling of large language models (LLMs) has revolutionized their capabilities in various tasks, yet this growth must be matched with efficient computational strategies.
The Mixture-of-Experts (MoE) architecture stands out for its ability to scale model size without significantly increasing training costs.
Despite their advantages, current MoE models often display parameter inefficiency.
For instance, a pre-trained MoE-based LLM with 52 billion parameters might perform comparably to a standard model with 6.7 billion.
Being a crucial part of MoE,
current routers in different layers independently assign tokens without leveraging historical routing information, potentially leading to suboptimal token-expert combinations and the parameter inefficiency problem.
To alleviate this issue, we introduce the Layerwise Recurrent Router for Mixture-of-Experts (RMoE).
RMoE leverages a Gated Recurrent Unit (GRU) to establish dependencies between routing decisions across consecutive layers.
Such layerwise recurrence can be efficiently parallelly computed for input tokens and introduces negotiable costs.
Our extensive empirical evaluations demonstrate that RMoE-based language models consistently outperform a spectrum of baseline models.
Furthermore, RMoE integrates a novel computation stage orthogonal to existing methods, allowing seamless compatibility with other MoE architectures.
Our analyses attribute RMoE's gains to its effective cross-layer information sharing, which also improves expert selection and diversity. Zihan Qiu, Shuang Cheng, Yizhi Zhou, Ivan Titov 0001, Jie Fu 0001 |
ICLR | 7 |
| 2025 | MuPT: A Generative Symbolic Music Pretrained TransformerabstractIn this paper, we explore the application of Large Language Models (LLMs) to the pre-training of music. While the prevalent use of MIDI in music modeling is well-established, our findings suggest that LLMs are inherently more compatible with ABC Notation, which aligns more closely with their design and strengths, thereby enhancing the model's performance in musical composition.
To address the challenges associated with misaligned measures from different tracks during generation, we propose the development of a $\underline{S}$ynchronized $\underline{M}$ulti-$\underline{T}$rack ABC Notation ($\textbf{SMT-ABC Notation}$), which aims to preserve coherence across multiple musical tracks.
Our contributions include a series of models capable of handling up to 8192 tokens, covering 90\% of the symbolic music data in our training set. Furthermore, we explore the implications of the $\underline{S}$ymbolic $\underline{M}$usic $\underline{S}$caling Law ($\textbf{SMS Law}$) on model performance. The results indicate a promising research direction in music generation, offering extensive resources for further research through our open-source contributions. Xingwei Qu, Yuelin Bai, Yinghao Ma, Ziya Zhou, Ka Man Lo, Ruibin Yuan, Lejun Min, Xueling Liu 0001, Xeron Du, Shuyue Guo, Yiming Liang, Shangda Wu, Junting Zhou, Tianyu Zheng, Ziyang Ma 0001, Fengze Han, Wei Xue 0002, Gus Xia, Emmanouil Benetos, Xiang Yue, Chenghua Lin 0002, Xu Tan 0003, Wenhao Huang 0001, Jie Fu 0001, Ge Zhang 0009 |
ICLR | 27 |
| 2025 | VCR: A Task for Pixel-Level Complex Reasoning in Vision Language Models via Restoring Occluded TextabstractWe introduce Visual Caption Restoration (VCR), a novel vision-language task that challenges models to accurately restore partially obscured texts using pixel-level hints within images through complex reasoning. This task stems from the observation that text embedded in images intrinsically differs from common visual elements and text due to the need to align the modalities of vision, text, and text embedded in images. While many works incorporate text into images for visual question answering, they mostly rely on OCR or masked language modeling, reducing the task to text-based processing. However, text-based processing becomes ineffective in VCR as accurate text restoration depends on the combined information from provided images, context, and subtle cues from the tiny, exposed areas of masked texts. We develop a pipeline to generate synthetic images for the VCR task using image-caption pairs, with adjustable caption visibility to control the task difficulty. With this pipeline, we construct VCR-WIKI for VCR using Wikipedia images with captions, including 2.11M English and 346K Chinese training entities, plus 5K validation and 5K test entities in both languages, each in easy and hard configurations. We also make a hidden test set, VCR-HIDDEN, to avoid potential overfitting on VCR-WIKI. Our results reveal that current vision-language models significantly lag behind human performance in the VCR task, and merely fine-tuning the models on our dataset does not lead to notable improvements. We release VCR-WIKI and the data construction code to facilitate future research. Suyuchen Wang, Ge Zhang 0009, Perouz Taslakian, Sai Rajeswar, Jie Fu 0001, Bang Liu 0003, Yoshua Bengio |
ICLR | 7 |
| 2025 | Enhancing Language Model Hypernetworks with Restart: A Study on OptimizationabstractYihan Zhang, Jie Fu, Rongrong Ji, Jie Chen. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Jie Fu 0001, Rongrong Ji, Jie Chen 0006 |
NAACL (Long Papers) | 2 |
| 2025 | Thinker: Learning to Think Fast and SlowabstractRecent studies show that the reasoning capabilities of Large Language Models (LLMs) can be improved by applying Reinforcement Learning (RL) to question-answering (QA) tasks in areas such as math and coding. With a long context length, LLMs may learn to perform search, as indicated by the self-correction behavior observed in DeepSeek R1. However, this search behavior is often imprecise and lacks confidence, resulting in long, redundant responses and highlighting deficiencies in intuition and verification. Inspired by the Dual Process Theory in psychology, we introduce a simple modification to the QA task that includes four stages: Fast Thinking, where the LLM must answer within a strict token budget; Verification, where the model evaluates its initial response; Slow Thinking, where it refines the initial response with more deliberation; and Summarization, where it distills the refinement from the previous stage into precise steps. Our proposed task improves average accuracy from 25.6% to 27.3% for Qwen2.5-1.5B, and from 45.9% to 51.0% for DeepSeek-R1-Qwen-1.5B. Notably, for Qwen2.5-1.5B, the Fast Thinking mode alone achieves 25.2% accuracy using fewer than 1000 tokens, demonstrating substantial inference efficiency gains. These findings suggest that intuition and deliberative reasoning are distinct, complementary systems benefiting from targeted training. Additionally, we have open-sourced both the trained models and the source code. Stephen Chung, Wenyu Du, Jie Fu 0001 |
NeurIPS | 3 |
| 2025 | Predicting protein stability changes upon mutations with dual-view ensemble learning from single sequenceabstractPredicting the protein stability changes upon mutations is one of the effective ways to improve the efficiency of protein engineering. Here, we propose a dual-view ensemble learning-based framework, DVE-stability, for mutation-induced protein stability change prediction from single sequence. DVE-stability integrates the global and local dependencies of mutations to capture the intramolecular interactions from two views through ensemble learning, in which a structural microenvironment simulation module is designed to indirectly introduce the information of structural microenvironment at the sequence level. DVE-stability achieved state-of-the-art prediction performance on seven single-point mutation benchmark datasets, and comprehensively surpassed other methods on five of them. Furthermore, DVE-stability outperformed other methods comprehensively through zero-shot inference on multiple-point mutation prediction task, demonstrating superior model generalizability to capture the epistasis of multiple-point mutations. More importantly, DVE-stability exhibited superior generalization performance in predicting rare beneficial mutations that are crucial for practical protein directed evolution scenarios. In addition, DVE-stability identified important intramolecular interactions via attention scores, demonstrating interpretable. Overall, DVE-stability provides a flexible and efficient tool for mutation-induced protein stability change prediction in an interpretable ensemble learning manner. Zhiwei Nie, Yutian Liu 0004, Xiansong Huang, Peng Yang 0001, Zigang Li, Jie Fu 0001, Zhixiang Ren, Jie Chen 0001 |
Briefings Bioinform. | 10 |
| 2024 | Scalable Geometric Fracture Assembly via Co-creation Space among AssemblersabstractGeometric fracture assembly presents a challenging practical task in archaeology and 3D computer vision. Previous methods have focused solely on assembling fragments based on semantic information, which has limited the quantity of objects that can be effectively assembled. Therefore, there is a need to develop a scalable framework for geometric fracture assembly without relying on semantic information. To improve the effectiveness of assembling geometric fractures without semantic information, we propose a co-creation space comprising several assemblers capable of gradually and unambiguously assembling fractures. Additionally, we introduce a novel loss function, i.e., the geometric-based collision loss, to address collision issues during the fracture assembly process and enhance the results. Our framework exhibits better performance on both PartNet and Breaking Bad datasets compared to existing state-of-the-art frameworks. Extensive experiments and quantitative comparisons demonstrate the effectiveness of our proposed framework, which features linear computational complexity, enhanced abstraction, and improved generalization. Our code is publicly available at https://github.com/Ruiyuan-Zhang/CCS. Ruiyuan Zhang, Zexi Li 0001, Hao Dong 0003, Jie Fu 0001, Chao Wu 0001 |
AAAI | 5 |
| 2024 | AnyGPT: Unified Multimodal LLM with Discrete Sequence ModelingabstractJun Zhan, Junqi Dai, Jiasheng Ye, Yunhua Zhou, Dong Zhang, Zhigeng Liu, Xin Zhang, Ruibin Yuan, Ge Zhang, Linyang Li, Hang Yan, Jie Fu, Tao Gui, Tianxiang Sun, Yu-Gang Jiang, Xipeng Qiu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Junqi Dai, Jiasheng Ye, Yunhua Zhou, Zhigeng Liu, Ruibin Yuan, Ge Zhang 0009, Linyang Li, Hang Yan 0001, Jie Fu 0001, Tao Gui, Tianxiang Sun, Yu-Gang Jiang 0001, Xipeng Qiu |
ACL (1) | 12 |
| 2024 | HyperMoE: Towards Better Mixture of Experts via Transferring Among ExpertsabstractThe Mixture of Experts (MoE) for language models has been proven effective in augmenting the capacity of models by dynamically routing each input token to a specific subset of experts for processing.Despite the success, most existing methods face a challenge for balance between sparsity and the availability of expert knowledge: enhancing performance through increased use of expert knowledge often results in diminishing sparsity during expert selection.To mitigate this contradiction, we propose HyperMoE, a novel MoE framework built upon Hypernetworks.This framework integrates the computational processes of MoE with the concept of knowledge transferring in multi-task learning.Specific modules generated based on the information of unselected experts serve as supplementary information, which allows the knowledge of experts not selected to be used while maintaining selection sparsity.Our comprehensive empirical evaluations across multiple datasets and backbones establish that HyperMoE significantly outperforms existing MoE methods under identical conditions concerning the number of experts. Zihan Qiu, Huijia Wu, Zhaofeng He 0001, Jie Fu 0001 |
ACL (1) | 6 |
| 2024 | CMDAG: A Chinese Metaphor Dataset with Annotated Grounds as CoT for Boosting Metaphor GenerationabstractMetaphor is a prominent linguistic device in human language and literature, as they add color, imagery, and emphasis to enhance effective communication. This paper introduces a large-scale high quality annotated Chinese Metaphor Corpus, which comprises around 28K sentences drawn from a diverse range of Chinese literary sources, such as poems, prose, song lyrics, etc. To ensure the accuracy and consistency of our annotations, we introduce a comprehensive set of guidelines. These guidelines address the facets of metaphor annotation, including identifying tenors, vehicles, and grounds to handling the complexities of similes, personifications, juxtapositions, and hyperboles. Breaking tradition, our approach to metaphor generation emphasizes tenors and their distinct features rather than the conventional combination of tenors and vehicles. By integrating “ground” as a CoT (Chain of Thoughts) input, we are able to generate metaphors that resonate more with real-world intuition. We test generative models such as Belle, Baichuan, and Chinese-alpaca-33B using our annotated corpus. These models are able to generate creative and fluent metaphor sentences more frequently induced by selected samples from our dataset, demonstrating the value of our corpus for Chinese metaphor research. Yujie Shao, Xinrong Yao, Xingwei Qu, Chenghua Lin 0002, Shi Wang 0002, Wenhao Huang 0001, Ge Zhang 0009, Jie Fu 0001 |
LREC/COLING | 8 |
| 2024 | UniIR: Training and Benchmarking Universal Multimodal Information Retrievers
Cong Wei 0001, Yang Chen 0065, Hexiang Hu, Ge Zhang 0009, Jie Fu 0001, Alan Ritter, Wenhu Chen |
ECCV (87) | 6 |
| 2024 | ChatEval: Towards Better LLM-based Evaluators through Multi-Agent DebateabstractText evaluation has historically posed significant challenges, often demanding substantial labor and time cost. With the emergence of large language models (LLMs), researchers have explored LLMs' potential as alternatives for human evaluation. While these single-agent-based approaches show promise, experimental results suggest that further advancements are needed to bridge the gap between their current effectiveness and human-level evaluation quality.
Recognizing that best practices of human evaluation processes often involve multiple human annotators collaborating in the evaluation, we resort to a multi-agent debate framework, moving beyond single-agent prompting strategies.
In this paper, we construct a multi-agent referee team called $\textbf{ChatEval}$ to autonomously discuss and evaluate the quality of different texts.
Our experiments on two benchmarks illustrate that ChatEval delivers superior accuracy and correlation in alignment with human assessment. Furthermore, we find that the diverse role prompts (different personas) are essential in the multi-agent debate process; that is, utilizing the same role description in the prompts can lead to a degradation in performance. Our qualitative analysis also shows that ChatEval transcends mere textual scoring, offering a human-mimicking evaluation process for reliable assessments. Chi-Min Chan, Weize Chen, Yusheng Su, Jianxuan Yu, Wei Xue 0002, Shanghang Zhang, Jie Fu 0001, Zhiyuan Liu 0001 |
ICLR | 7 |
| 2024 | MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised TrainingabstractSelf-supervised learning (SSL) has recently emerged as a promising paradigm for training generalisable models on large-scale data in the fields of vision, text, and speech.
Although SSL has been proven effective in speech and audio, its application to music audio has yet to be thoroughly explored. This is partially due to the distinctive challenges associated with modelling musical knowledge, particularly tonal and pitched characteristics of music.
To address this research gap, we propose an acoustic **M**usic und**ER**standing model with large-scale self-supervised **T**raining (**MERT**), which incorporates teacher models to provide pseudo labels in the masked language modelling (MLM) style acoustic pre-training.
In our exploration, we identified an effective combination of teacher models, which outperforms conventional speech and audio approaches in terms of performance.
This combination includes an acoustic teacher based on Residual Vector Quantization - Variational AutoEncoder (RVQ-VAE) and a musical teacher based on the Constant-Q Transform (CQT).
Furthermore, we explore a wide range of settings to overcome the instability in acoustic language model pre-training, which allows our designed paradigm to scale from 95M to 330M parameters.
Experimental results indicate that our model can generalise and perform well on 14 music understanding tasks and attain state-of-the-art (SOTA) overall scores. Ruibin Yuan, Ge Zhang 0009, Yinghao Ma, Xingran Chen, Hanzhi Yin, Chenghao Xiao, Chenghua Lin 0002, Anton Ragni, Emmanouil Benetos, Norbert Gyenge, Roger B. Dannenberg, Ruibo Liu, Wenhu Chen, Gus Xia, Yemin Shi 0001, Wenhao Huang 0001, Yike Guo, Jie Fu 0001 |
ICLR | 20 |
| 2024 | Massive Editing for Large Language Models via Meta LearningabstractWhile large language models (LLMs) have enabled learning knowledge from the pre-training corpora, the acquired knowledge may be fundamentally incorrect or outdated over time, which necessitates rectifying the knowledge of the language model (LM) after the training. A promising approach involves employing a hyper-network to generate parameter shift, whereas existing hyper-networks suffer from inferior scalability in synchronous editing operation amount (Hase et al., 2023b; Huang et al., 2023). For instance, Mitchell et al. (2022) mimics gradient accumulation to sum the parameter shifts together, which lacks statistical significance and is prone to cancellation effect. To mitigate the problem, we propose the MAssive Language Model Editing Network (MALMEN), which formulates the parameter shift aggregation as the least square problem, subsequently updating the LM parameter using the normal equation. To accommodate editing multiple facts simultaneously with limited memory budgets, we separate the computation on the hyper-network and LM, enabling arbitrary batch size on both neural networks. Our method is evaluated by editing up to thousands of facts on LMs with different architectures, i.e., BERT-base, GPT-2, and GPT-J (6B), across various knowledge-intensive NLP tasks, i.e., closed book fact-checking and question answering. Remarkably, MALMEN is capable of editing hundreds of times more facts than MEND (Mitchell et al., 2022) with the identical hyper-network architecture and outperforms editor specifically designed for GPT, i.e., MEMIT (Meng et al., 2023). Chenmien Tan, Ge Zhang 0009, Jie Fu 0001 |
ICLR | 3 |
| 2024 | Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMsabstractEmpowering large language models (LLMs) to accurately express confidence in their answers is essential for reliable and trustworthy decision-making. Previous confidence elicitation methods, which primarily rely on *white-box access* to internal model information or model fine-tuning, have become less suitable for LLMs, especially closed-source commercial APIs. This leads to a growing need to explore the untapped area of *black-box* approaches for LLM uncertainty estimation. To better break down the problem, we define a systematic framework with three components: *prompting* strategies for eliciting verbalized confidence, *sampling* methods for generating multiple responses, and *aggregation* techniques for computing consistency. We then benchmark these methods on two key tasks—confidence calibration and failure prediction—across five types of datasets (e.g., commonsense and arithmetic reasoning) and five widely-used LLMs including GPT-4 and LLaMA 2 Chat. Our analysis uncovers several key insights: 1) LLMs, when verbalizing their confidence, tend to be *overconfident*, potentially imitating human patterns of expressing confidence. 2) As model capability scales up, both calibration and failure prediction performance improve, yet still far from ideal performance.
3) Employing our proposed strategies, such as human-inspired prompts, consistency among multiple responses, and better aggregation strategies can help mitigate this overconfidence from various perspectives.
4) Comparisons with white-box methods indicate that while white-box methods perform better, the gap is narrow, e.g., 0.522 to 0.605 in AUROC. Despite these advancements, none of these techniques consistently outperform others, and all investigated methods struggle in challenging tasks, such as those requiring professional knowledge, indicating significant scope for improvement. We believe this study can serve as a strong baseline and provide insights for eliciting confidence in black-box LLMs. The code is publicly available at https://github.com/MiaoXiong2320/llm-uncertainty. Miao Xiong, Xinyang Lu, Jie Fu 0001, Junxian He, Bryan Hooi |
ICLR | 5 |
| 2024 | Think Before You Act: Decision Transformers with Working MemoryabstractDecision Transformer-based decision-making agents have shown the ability to generalize across multiple tasks. However, their performance relies on massive data and computation. We argue that this inefficiency stems from the forgetting phenomenon, in which a model memorizes its behaviors in parameters throughout training. As a result, training on a new task may deteriorate the model’s performance on previous tasks. In contrast to LLMs’ implicit memory mechanism, the human brain utilizes distributed memory storage, which helps manage and organize multiple skills efficiently, mitigating the forgetting phenomenon. Inspired by this, we propose a working memory module to store, blend, and retrieve information for different downstream tasks. Evaluation results show that the proposed method improves training efficiency and generalization in Atari games and Meta-World object manipulation tasks. Moreover, we demonstrate that memory fine-tuning further enhances the adaptability of the proposed architecture. Jikun Kang, Romain Laroche, Xingdi Yuan, Adam Trischler, Xue (Steve) Liu, Jie Fu 0001 |
ICML | 6 |
| 2024 | AutoAgents: A Framework for Automatic Agent Generation
Siwei Dong, Yu Shu, Ge Zhang 0009, Jaward Sesay, Börje Karlsson 0001, Jie Fu 0001, Yemin Shi 0001 |
IJCAI | 7 |
| 2024 | ReForm-Eval: Evaluating Large Vision Language Models via Unified Re-Formulation of Task-Oriented BenchmarksabstractRecent years have witnessed remarkable progress in the development of large vision-language models (LVLMs). Benefiting from the strong language backbones and efficient cross-modal alignment strategies, LVLMs exhibit surprising capabilities to perceive visual signals and perform visually grounded reasoning. However, the capabilities of LVLMs have not been comprehensively and quantitatively evaluated. Most existing multi-modal benchmarks require task-oriented input-output formats, posing great challenges to automatically assess the free-form text output of LVLMs. To effectively leverage the annotations available and reduce the manual efforts required for constructing new benchmarks, we propose to re-formulate existing benchmarks into unified LVLM-compatible formats. Through systematic data collection and reformulation, we present ReForm-Eval benchmark, offering substantial data for evaluating various capabilities of LVLMs. Through extensive experiments and analysis in ReForm-Eval, we demonstrate the comprehensiveness and reliability of ReForm-Eval in assessing various LVLMs. Our benchmark and evaluation framework is now available at https://github.com/FudanDISC/ReForm-Eval Mengfei Du, Qingwen Liu 0002, Binhao Wu, Jiwen Zhang, Chengxing Zhou, Zhihao Fan, Jie Fu 0001, Jingjing Chen 0001, Zhongyu Wei, Xuanjing Huang 0001 |
ACM Multimedia | 9 |
| 2024 | Unlocking Emergent Modularity in Large Language ModelsabstractZihan Qiu, Zeyu Huang, Jie Fu. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Zihan Qiu, Jie Fu 0001 |
NAACL-HLT | 3 |
| 2024 | Stacking Your Transformers: A Closer Look at Model Growth for Efficient LLM Pre-TrainingabstractLLMs are computationally expensive to pre-train due to their large scale.
Model growth emerges as a promising approach by leveraging smaller models to accelerate the training of larger ones.
However, the viability of these model growth methods in efficient LLM pre-training remains underexplored.
This work identifies three critical $\underline{\textit{O}}$bstacles: ($\textit{O}$1) lack of comprehensive evaluation, ($\textit{O}$2) untested viability for scaling, and ($\textit{O}$3) lack of empirical guidelines.
To tackle $\textit{O}$1, we summarize existing approaches into four atomic growth operators and systematically evaluate them in a standardized LLM pre-training setting.
Our findings reveal that a depthwise stacking operator, called $G_{\text{stack}}$, exhibits remarkable acceleration in training, leading to decreased loss and improved overall performance on eight standard NLP benchmarks compared to strong baselines.
Motivated by these promising results, we conduct extensive experiments to delve deeper into $G_{\text{stack}}$ to address $\textit{O}$2 and $\textit{O}$3.
For $\textit{O}$2 (untested scalability), our study shows that $G_{\text{stack}}$ is scalable and consistently performs well, with experiments up to 7B LLMs after growth and pre-training LLMs with 750B tokens.
For example, compared to a conventionally trained 7B model using 300B tokens, our $G_{\text{stack}}$ model converges to the same loss with 194B tokens, resulting in a 54.6\% speedup.
We further address $\textit{O}$3 (lack of empirical guidelines) by formalizing guidelines to determine growth timing and growth factor for $G_{\text{stack}}$, making it practical in general LLM pre-training.
We also provide in-depth discussions and comprehensive ablation studies of $G_{\text{stack}}$.
Our code and pre-trained model are available at https://llm-stacking.github.io/. Wenyu Du, Tongxu Luo, Zihan Qiu, Yikang Shen, Reynold Cheng, Yike Guo, Jie Fu 0001 |
NeurIPS | 8 |
| 2024 | D-CPT Law: Domain-specific Continual Pre-Training Scaling Law for Large Language ModelsabstractContinual Pre-Training (CPT) on Large Language Models (LLMs) has been widely used to expand the model’s fundamental understanding of specific downstream domains (e.g., math and code). For the CPT on domain-specific LLMs, one important question is how to choose the optimal mixture ratio between the general-corpus (e.g., Dolma, Slim-pajama) and the downstream domain-corpus. Existing methods usually adopt laborious human efforts by grid-searching on a set of mixture ratios, which require high GPU training consumption costs. Besides, we cannot guarantee the selected ratio is optimal for the specific domain. To address the limitations of existing methods, inspired by the Scaling Law for performance prediction, we propose to investigate the Scaling Law of the Domain-specific Continual Pre-Training (D-CPT Law) to decide the optimal mixture ratio with acceptable training costs for LLMs of different sizes. Specifically, by fitting the D-CPT Law, we can easily predict the general and downstream performance of arbitrary mixture ratios, model sizes, and dataset sizes using small-scale training costs on limited experiments. Moreover, we also extend our standard D-CPT Law on cross-domain settings and propose the Cross-Domain D-CPT Law to predict the D-CPT law of target domains, where very small training costs (about 1\% of the normal training costs) are needed for the target domains. Comprehensive experimental results on six downstream domains demonstrate the effectiveness and generalizability of our proposed D-CPT Law and Cross-Domain D-CPT Law. Haoran Que, Ge Zhang 0009, Xingwei Qu, Yinghao Ma, Feiyu Duan, Zhiqi Bai, Jiakai Wang, Yuanxing Zhang, Xu Tan 0003, Jie Fu 0001, Jiamang Wang, Lin Qu, Wenbo Su, Bo Zheng 0007 |
NeurIPS | 12 |
| 2024 | Exploring Clean Label Backdoor Attacks and Defense in Language ModelsabstractDespite being widely applied, pre-trained language models have been proven vulnerable to backdoor attacks. Backdoor attacks are designed to introduce targeted vulnerabilities into models by poisoning a subset of training samples through trigger injection and label modification. Traditional textual backdoor attacks suffer several flaws: the triggers lead to abnormal natural language expressions, and poisoned sample labels are mistakenly labeled. These flaws reduce the stealthiness of the attack and can be easily detected by defense models. In this study, we introduce Cbat, a novel and efficient method to perform clean-label backdoor attack with text style, which does not require external trigger, and the poisoned samples are correctly labeled. Specifically, we develop a sentence rewriting model by leveraging the powerful few-shot learning capability of prompt tuning to generate clean label poisoned samples. Cbat then injects text style as an abstract trigger into the victim model through poisoned samples. We also introduce an algorithm for defending against backdoor attacks, named CbatD, which effectively erases the poisoned samples by locating the lowest training loss and calculating feature relevance. The experiments on text classification tasks demonstrate that our Cbat and CbatD show overall competitive performance in textual backdoor attack and defense. It is noteworthy that Cbat attained leading results in the clean-label backdoor attack benchmark without triggers. Shuai Zhao 0007, Anh Tuan Luu, Jie Fu 0001, Jinming Wen, Weiqi Luo 0002 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2023 | Unifying Discrete and Continuous Representations for Unsupervised Paraphrase GenerationabstractMingfeng Xue, Dayiheng Liu, Wenqiang Lei, Jie Fu, Jian Lan, Mei Li, Baosong Yang, Jun Xie, Yidan Zhang, Dezhong Peng, Jiancheng Lv. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Mingfeng Xue, Dayiheng Liu, Wenqiang Lei, Jie Fu 0001, Baosong Yang, Yidan Zhang 0004, Dezhong Peng, Jiancheng Lv 0001 |
EMNLP | 4 |
| 2023 | Prototype-based HyperAdapter for Sample-Efficient Multi-task TuningabstractParameter-efficient fine-tuning (PEFT) has shown its effectiveness in adapting the pretrained language models to downstream tasks while only updating a small number of parameters.Despite the success, most existing methods independently adapt to each task without considering knowledge transfer between tasks and are limited to low-data regimes.To overcome this issue, we propose Prototype-based HyperAdapter (PHA), a novel framework built on the adapter-tuning and hypernetwork.It introduces an instance-dense retriever and a prototypical hypernetwork to generate the conditional modules in a sample-efficient manner.This leads to comparable performance improvements against existing PEFT methods on multi-task learning and few-shot transfer learning.More importantly, when the available data size gets smaller, our method outperforms other strong baselines by a large margin.Based on our extensive empirical experiments across various datasets, we demonstrate that PHA strikes a better trade-off between trainable parameters, accuracy on stream tasks, and sample efficiency.Our code is publicly available at https://github.com/Bumble666/PHA Jie Fu 0001, Zhaofeng He 0001 |
EMNLP | 2 |
| 2023 | Prompt as Triggers for Backdoor Attack: Examining the Vulnerability in Language ModelsabstractThe prompt-based learning paradigm, which bridges the gap between pre-training and finetuning, achieves state-of-the-art performance on several NLP tasks, particularly in few-shot settings.Despite being widely applied, promptbased learning is vulnerable to backdoor attacks.Textual backdoor attacks are designed to introduce targeted vulnerabilities into models by poisoning a subset of training samples through trigger injection and label modification.However, they suffer from flaws such as abnormal natural language expressions resulting from the trigger and incorrect labeling of poisoned samples.In this study, we propose ProAttack, a novel and efficient method for performing clean-label backdoor attacks based on the prompt, which uses the prompt itself as a trigger.Our method does not require external triggers and ensures correct labeling of poisoned samples, improving the stealthy nature of the backdoor attack.With extensive experiments on rich-resource and few-shot text classification tasks, we empirically validate ProAttack's competitive performance in textual backdoor attacks.Notably, in the rich-resource setting, ProAttack achieves state-of-the-art attack success rates in the clean-label backdoor attack benchmark without external triggers 1 . Shuai Zhao 0007, Jinming Wen, Anh Tuan Luu, Jie Fu 0001 |
EMNLP | 5 |
| 2023 | Med-UniC: Unifying Cross-Lingual Medical Vision-Language Pre-Training by Diminishing BiasabstractThe scarcity of data presents a critical obstacle to the efficacy of medical vision-language pre-training (VLP). A potential solution lies in the combination of datasets from various language communities.
Nevertheless, the main challenge stems from the complexity of integrating diverse syntax and semantics, language-specific medical terminology, and culture-specific implicit knowledge. Therefore, one crucial aspect to consider is the presence of community bias caused by different languages.
This paper presents a novel framework named Unifying Cross-Lingual Medical Vision-Language Pre-Training (\textbf{Med-UniC}), designed to integrate multi-modal medical data from the two most prevalent languages, English and Spanish.
Specifically, we propose \textbf{C}ross-lingual \textbf{T}ext Alignment \textbf{R}egularization (\textbf{CTR}) to explicitly unify cross-lingual semantic representations of medical reports originating from diverse language communities.
\textbf{CTR} is optimized through latent language disentanglement, rendering our optimization objective to not depend on negative samples, thereby significantly mitigating the bias from determining positive-negative sample pairs within analogous medical reports. Furthermore, it ensures that the cross-lingual representation is not biased toward any specific language community.
\textbf{Med-UniC} reaches superior performance across 5 medical image tasks and 10 datasets encompassing over 30 diseases, offering a versatile framework for unifying multi-modal medical data within diverse linguistic communities.
The experimental outcomes highlight the presence of community bias in cross-lingual VLP. Reducing this bias enhances the performance not only in vision-language tasks but also in uni-modal visual tasks. Zhongwei Wan, Che Liu 0002, Mi Zhang 0002, Jie Fu 0001, Benyou Wang, Sibo Cheng, Lei Ma 0008, César Quilodrán Casas, Rossella Arcucci |
NeurIPS | 4 |
| 2023 | MARBLE: Music Audio Representation Benchmark for Universal EvaluationabstractIn the era of extensive intersection between art and Artificial Intelligence (AI), such as image generation and fiction co-creation, AI for music remains relatively nascent, particularly in music understanding. This is evident in the limited work on deep music representations, the scarcity of large-scale datasets, and the absence of a universal and community-driven benchmark. To address this issue, we introduce the Music Audio Representation Benchmark for universaL Evaluation, termed MARBLE. It aims to provide a benchmark for various Music Information Retrieval (MIR) tasks by defining a comprehensive taxonomy with four hierarchy levels, including acoustic, performance, score, and high-level description. We then establish a unified protocol based on 18 tasks on 12 public-available datasets, providing a fair and standard assessment of representations of all open-sourced pre-trained models developed on music recordings as baselines. Besides, MARBLE offers an easy-to-use, extendable, and reproducible suite for the community, with a clear statement on copyright issues on datasets. Results suggest recently proposed large-scale pre-trained musical language models perform the best in most tasks, with room for further improvement. The leaderboard and toolkit repository are published to promote future music AI research. Ruibin Yuan, Yinghao Ma, Ge Zhang 0009, Xingran Chen, Hanzhi Yin, Le Zhuo, Zeyue Tian, Binyue Deng, Ningzhi Wang, Chenghua Lin 0002, Emmanouil Benetos, Anton Ragni, Norbert Gyenge, Roger B. Dannenberg, Wenhu Chen, Gus Xia, Wei Xue 0002, Shi Wang 0002, Ruibo Liu, Yike Guo, Jie Fu 0001 |
NeurIPS | 25 |
| 2023 | GIMLET: A Unified Graph-Text Model for Instruction-Based Molecule Zero-Shot LearningabstractMolecule property prediction has gained significant attention in recent years. The main bottleneck is the label insufficiency caused by expensive lab experiments. In order to alleviate this issue and to better leverage textual knowledge for tasks, this study investigates the feasibility of employing natural language instructions to accomplish molecule-related tasks in a zero-shot setting. We discover that existing molecule-text models perform poorly in this setting due to inadequate treatment of instructions and limited capacity for graphs.
To overcome these issues, we propose GIMLET, which unifies language models for both graph and text data. By adopting generalized position embedding, our model is extended to encode both graph structures and instruction text without additional graph encoding modules. GIMLET also decouples encoding of the graph from tasks instructions in the attention mechanism, enhancing the generalization of graph features across novel tasks. We construct a dataset consisting of more than two thousand molecule tasks with corresponding instructions derived from task descriptions. We pretrain GIMLET on the molecule tasks along with instructions, enabling the model to transfer effectively to a broad range of tasks. Experimental results demonstrate that GIMLET significantly outperforms molecule-text baselines in instruction-based zero-shot learning, even achieving closed results to supervised GNN models on tasks such as toxcast and muv. Haiteng Zhao, Shengchao Liu, Hannan Xu, Jie Fu 0001, Zhi-Hong Deng 0001, Lingpeng Kong, Qi Liu 0049 |
NeurIPS | 5 |
| 2022 | Unifying Likelihood-free Inference with Black-box Optimization and Beyond
Dinghuai Zhang, Jie Fu 0001, Yoshua Bengio, Aaron C. Courville |
ICLR | 2 |
| 2022 | Biological Sequence Design with GFlowNetsabstractDesign of de novo biological sequences with desired properties, like protein and DNA sequences, often involves an active loop with several rounds of molecule ideation and expensive wet-lab evaluations. These experiments can consist of multiple stages, with increasing levels of precision and cost of evaluation, where candidates are filtered. This makes the diversity of proposed candidates a key consideration in the ideation phase. In this work, we propose an active learning algorithm leveraging epistemic uncertainty estimation and the recently proposed GFlowNets as a generator of diverse candidate solutions, with the objective to obtain a diverse batch of useful (as defined by some utility function, for example, the predicted anti-microbial activity of a peptide) and informative candidates after each round. We also propose a scheme to incorporate existing labeled datasets of candidates, in addition to a reward function, to speed up learning in GFlowNets. We present empirical results on several biological sequence design tasks, and we find that our method generates more diverse and novel batches with high scoring candidates compared to existing approaches. Moksh Jain, Emmanuel Bengio, Alex Hernández-García, Jarrid Rector-Brooks, Bonaventure F. P. Dossou, Chanakya Ajit Ekbote, Jie Fu 0001, Michael Kilgour, Dinghuai Zhang, Lena Simine, Yoshua Bengio |
ICML | 7 |
| 2022 | MentalBERT: Publicly Available Pretrained Language Models for Mental HealthcareabstractMental health is a critical issue in modern society, and mental disorders could sometimes turn to suicidal ideation without adequate treatment. Early detection of mental disorders and suicidal ideation from social content provides a potential way for effective social intervention. Recent advances in pretrained contextualized language representations have promoted the development of several domainspecific pretrained models and facilitated several downstream applications. However, there are no existing pretrained language models for mental healthcare. This paper trains and release two pretrained masked language models, i.e., MentalBERT and MentalRoBERTa, to benefit machine learning for the mental healthcare research community. Besides, we evaluate our trained domain-specific models and several variants of pretrained language models on several mental disorder detection benchmarks and demonstrate that language representations pretrained in the target domain improve the performance of mental health detection tasks. Shaoxiong Ji, Luna Ansari, Jie Fu 0001, Prayag Tiwari, Erik Cambria |
LREC | 4 |
| 2022 | Bidirectional Learning for Offline Infinite-width Model-based OptimizationabstractIn offline model-based optimization, we strive to maximize a black-box objective function by only leveraging a static dataset of designs and their scores. This problem setting arises in numerous fields including the design of materials, robots, DNAs, proteins, etc. Recent approaches train a deep neural network (DNN) model on the static dataset to act as a proxy function, and then perform gradient ascent on the existing designs to obtain potentially high-scoring designs. This methodology frequently suffers from the out-of-distribution problem where the proxy function often returns adversarial designs. To mitigate this problem, we propose $\textit{\textbf{B}i\textbf{D}irectional learning for offline \textbf{I}nfinite-width model-based optimization}~(\textbf{BDI})$. BDI consists of two mappings: the forward mapping leverages the static dataset to predict the scores of the high-scoring designs, and the backward mapping leverages the high-scoring designs to predict the scores of the static dataset. The backward mapping, neglected in previous work, can distill more information of the static dataset into the high-scoring designs, which effectively mitigates the out-of-distribution problem. Yet, for a finite-width DNN model, the loss function of the backward mapping is intractable and only has an approximate form, which leads to a significant deterioration of the design quality. We thus adopt an infinite-width DNN model and propose to employ the corresponding neural tangent kernel to yield a closed-form loss for more accurate design updates. Experiments on various tasks verify the effectiveness of BDI. The code is available [here](https://github.com/GGchen1997/BDI). Can Chen 0005, Yingxue Zhang 0001, Jie Fu 0001, Xue (Steve) Liu, Mark Coates |
NeurIPS | 3 |
| 2021 | FloW: A Dataset and Benchmark for Floating Waste Detection in Inland WatersabstractMarine debris is severely threatening the marine lives and causing sustained pollution to the whole ecosystem. To prevent the wastes from getting into the ocean, it is helpful to clean up the floating wastes in inland waters using the autonomous cleaning devices like unmanned surface vehicles. The cleaning efficiency relies on a high-accurate and robust object detection system. However, the small size of the target, the strong light reflection over water surface, and the reflection of other objects on bank-side all bring challenges to the vision-based object detection system. To promote the practical application for autonomous floating wastes cleaning, we present FloW†, the first dataset for floating waste detection in inland water areas. The dataset consists of an image sub-dataset FloW-Img and a multimodal sub-dataset FloW-RI which contains synchronized millimeter wave radar data and images. Accurate annotations for images and radar data are provided, supporting floating waste detection strategies based on image, radar data, and the fusion of two sensors. We perform several baseline experiments on our dataset, including vision-based and radar-based detection methods. The results show that, the detection accuracy is relatively low and floating waste detection still remains a challenging task. Yuwei Cheng, Jiannan Zhu, Mengxin Jiang, Jie Fu 0001, Changsong Pang, Kris Sankaran, Olawale Onabola, Dianbo Liu, Yoshua Bengio |
ICCV | 4 |
| 2021 | Beyond Fully-Connected Layers with Quaternions: Parameterization of Hypercomplex Multiplications with 1/n Parameters
Aston Zhang, Yi Tay, Shuai Zhang 0007, Alvin Chan, Anh Tuan Luu, Siu Cheung Hui, Jie Fu 0001 |
ICLR | 7 |
| 2020 | Revision in Continuous Space: Unsupervised Text Style Transfer without Adversarial LearningabstractTypical methods for unsupervised text style transfer often rely on two key ingredients: 1) seeking the explicit disentanglement of the content and the attributes, and 2) troublesome adversarial learning. In this paper, we show that neither of these components is indispensable. We propose a new framework that utilizes the gradients to revise the sentence in a continuous space during inference to achieve text style transfer. Our method consists of three key components: a variational auto-encoder (VAE), some attribute predictors (one for each attribute), and a content predictor. The VAE and the two types of predictors enable us to perform gradient-based optimization in the continuous space, which is mapped from sentences in a discrete space, to find the representation of a target sentence with the desired attributes and preserved content. Moreover, the proposed method naturally has the ability to simultaneously manipulate multiple fine-grained attributes, such as sentence length and the presence of specific words, when performing text style transfer tasks. Compared with previous adversarial learning based methods, the proposed method is more interpretable, controllable and easier to train. Extensive experimental studies on three popular text style transfer tasks show that the proposed method significantly outperforms five state-of-the-art methods. Dayiheng Liu, Jie Fu 0001, Yidan Zhang 0004, Christopher Joseph Pal, Jiancheng Lv 0001 |
AAAI | 2 |
| 2020 | RikiNet: Reading Wikipedia Pages for Natural Question AnsweringabstractReading long documents to answer opendomain questions remains challenging in natural language understanding.In this paper, we introduce a new model, called RikiNet, which reads Wikipedia pages for natural question answering.RikiNet contains a dynamic paragraph dual-attention reader and a multi-level cascaded answer predictor.The reader dynamically represents the document and question by utilizing a set of complementary attention mechanisms.The representations are then fed into the predictor to obtain the span of the short answer, the paragraph of the long answer, and the answer type in a cascaded manner.On the Natural Questions (NQ) dataset, a single RikiNet achieves 74.3 F1 and 57.9 F1 on longanswer and short-answer tasks.To our best knowledge, it is the first single model that outperforms the single human performance.Furthermore, an ensemble RikiNet obtains 76.1 F1 and 61.3 F1 on long-answer and shortanswer tasks, achieving the best performance on the official NQ leaderboard 1 . Dayiheng Liu, Yeyun Gong, Jie Fu 0001, Jiusheng Chen, Daxin Jiang, Jiancheng Lv 0001, Nan Duan 0001 |
ACL | 3 |
| 2020 | Would you Rather? A New Benchmark for Learning Machine Alignment with Cultural Values and Social PreferencesabstractUnderstanding human preferences, along with cultural and social nuances, lives at the heart of natural language understanding.Concretely, we present a new task and corpus for learning alignments between machine and human preferences.Our newly introduced problem is concerned with predicting the preferable options from two sentences describing scenarios that may involve social and cultural situations.Our problem is framed as a natural language inference task with crowd-sourced preference votes by human players, obtained from a gamified voting platform.We benchmark several state-of-the-art neural models, along with BERT and friends on this task.Our experimental results show that current state-ofthe-art NLP models still leave much room for improvement. Yi Tay, Donovan Ong, Jie Fu 0001, Alvin Chan, Nancy F. Chen, Anh Tuan Luu, Christopher Joseph Pal |
ACL | 3 |
| 2020 | Interactive Machine Comprehension with Information Seeking AgentsabstractExisting machine reading comprehension (MRC) models do not scale effectively to realworld applications like web-level information retrieval and question answering (QA).We argue that this stems from the nature of MRC datasets: most of these are static environments wherein the supporting documents and all necessary information are fully observed.In this paper, we propose a simple method that reframes existing MRC datasets as interactive, partially observable environments.Specifically, we "occlude" the majority of a document's text and add context-sensitive commands that reveal "glimpses" of the hidden text to a model.We repurpose SQuAD and NewsQA as an initial case study, and then show how the interactive corpora can be used to train a model that seeks relevant information through sequential decision making.We believe that this setting can contribute in scaling models to web-level QA scenarios.1 Xingdi Yuan, Jie Fu 0001, Marc-Alexandre Côté, Yi Tay, Christopher Joseph Pal, Adam Trischler |
ACL | 2 |
| 2020 | Tell Me How to Ask Again: Question Data Augmentation with Controllable Rewriting in Continuous SpaceabstractIn this paper, we propose a novel data augmentation method, referred to as Controllable Rewriting based Question Data Augmentation (CRQDA), for machine reading comprehension (MRC), question generation, and question-answering natural language inference tasks.We treat the question data augmentation task as a constrained question rewriting problem to generate context-relevant, high-quality, and diverse question data samples.CRQDA utilizes a Transformer autoencoder to map the original discrete question into a continuous embedding space.It then uses a pre-trained MRC model to revise the question representation iteratively with gradientbased optimization.Finally, the revised question representations are mapped back into the discrete space, which serve as additional question data.Comprehensive experiments on SQuAD 2.0, SQuAD 1.1 question generation, and QNLI tasks demonstrate the effectiveness of CRQDA 1 . Dayiheng Liu, Yeyun Gong, Jie Fu 0001, Jiusheng Chen, Jiancheng Lv 0001, Nan Duan 0001, Ming Zhou 0001 |
EMNLP (1) | 3 |
| 2020 | Diverse, Controllable, and Keyphrase-Aware: A Corpus and Method for News Multi-Headline GenerationabstractNews headline generation aims to produce a short sentence to attract readers to read the news.One news article often contains multiple keyphrases that are of interest to different users, which can naturally have multiple reasonable headlines.However, most existing methods focus on the single headline generation.In this paper, we propose generating multiple headlines with keyphrases of user interests, whose main idea is to generate multiple keyphrases of interest to users for the news first, and then generate multiple keyphrase-relevant headlines.We propose a multi-source Transformer decoder, which takes three sources as inputs: (a) keyphrase, (b) keyphrase-filtered article, and (c) original article to generate keyphrase-relevant, highquality, and diverse headlines.Furthermore, we propose a simple and effective method to mine the keyphrases of interest in the news article and build a first large-scale keyphraseaware news headline corpus, which contains over 180K aligned triples of news article, headline, keyphrase .Extensive experimental comparisons on the real-world dataset show that the proposed method achieves state-of-theart results in terms of quality and diversity 1 . Dayiheng Liu, Yeyun Gong, Jie Fu 0001, Daxin Jiang, Jiancheng Lv 0001, Nan Duan 0001 |
EMNLP (1) | 4 |
| 2019 | Learning Multi-Task Communication with Message Passing for Sequence LearningabstractWe present two architectures for multi-task learning with neural sequence models. Our approach allows the relationships between different tasks to be learned dynamically, rather than using an ad-hoc pre-defined structure as in previous work. We adopt the idea from message-passing graph neural networks, and propose a general graph multi-task learning framework in which different tasks can communicate with each other in an effective and interpretable way. We conduct extensive experiments in text classification and sequence labelling to evaluate our approach on multi-task learning and transfer learning. The empirical results show that our models not only outperform competitive baselines, but also learn interpretable and transferable patterns across tasks. Pengfei Liu 0003, Jie Fu 0001, Yue Dong 0002, Xipeng Qiu, Jackie Chi Kit Cheung |
AAAI | 2 |
| 2019 | TIGS: An Inference Algorithm for Text Infilling with Gradient SearchabstractText infilling is defined as a task for filling in the missing part of a sentence or paragraph, which is suitable for many real-world natural language generation scenarios.However, given a well-trained sequential generative model, generating missing symbols conditioned on the context is challenging for existing greedy approximate inference algorithms.In this paper, we propose an iterative inference algorithm based on gradient search, which is the first inference algorithm that can be broadly applied to any neural sequence generative models for text infilling tasks.We compare the proposed method with strong baselines on three text infilling tasks with various mask ratios and different mask strategies.The results show that our proposed method is effective and efficient for fill-in-the-blank tasks, consistently outperforming all baselines.1 Dayiheng Liu, Jie Fu 0001, Pengfei Liu 0003, Jiancheng Lv 0001 |
ACL (1) | 2 |
| 2019 | Simple and Effective Curriculum Pointer-Generator Networks for Reading Comprehension over Long NarrativesabstractYi Tay, Shuohang Wang, Anh Tuan Luu, Jie Fu, Minh C. Phan, Xingdi Yuan, Jinfeng Rao, Siu Cheung Hui, Aston Zhang. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019. Yi Tay, Shuohang Wang, Anh Tuan Luu, Jie Fu 0001, Minh C. Phan, Xingdi Yuan, Jinfeng Rao, Siu Cheung Hui, Aston Zhang |
ACL (1) | 4 |
| 2019 | Lightweight and Efficient Neural Natural Language Processing with Quaternion NetworksabstractMany state-of-the-art neural models for NLP are heavily parameterized and thus memory inefficient.This paper proposes a series of lightweight and memory efficient neural architectures for a potpourri of natural language processing (NLP) tasks.To this end, our models exploit computation using Quaternion algebra and hypercomplex spaces, enabling not only expressive inter-component interactions but also significantly (75%) reduced parameter size due to lesser degrees of freedom in the Hamilton product.We propose Quaternion variants of models, giving rise to new architectures such as the Quaternion attention Model and Quaternion Transformer.Extensive experiments on a battery of NLP tasks demonstrates the utility of proposed Quaternion-inspired models, enabling up to 75% reduction in parameter size without significant loss in performance. Yi Tay, Aston Zhang, Anh Tuan Luu, Jinfeng Rao, Shuai Zhang 0007, Shuohang Wang, Jie Fu 0001, Siu Cheung Hui |
ACL (1) | 7 |
| 2019 | Graph Neural Networks with Generated Parameters for Relation Extractionabstract10.18653/v1/P19-1128 Hao Zhu 0006, Yankai Lin 0001, Zhiyuan Liu 0001, Jie Fu 0001, Tat-Seng Chua, Maosong Sun 0001 |
ACL (1) | 4 |
| 2019 | Dataflow-Based Joint Quantization for Deep Neural NetworksabstractThis paper addresses a challenging problem - how to reduce energy consumption without incurring performance drop when deploying deep neural networks (DNNs) at the inference stage. In order to alleviate the computation and storage burdens, we propose a novel dataflow-based joint quantization approach with the hypothesis that a fewer number of quantization operations would incur less information loss and thus improve the final performance. It first introduces a quantization scheme with efficient bit-shifting and rounding operations to represent network parameters and activations in low precision. Then it re-structures the network architectures to form unified modules for optimization on the quantized model. Extensive experiments on ImageNet and KITTI validate the effectiveness of our model, demonstrating that state-of-the-art results for various tasks can be achieved by this quantized model. Besides, we designed and synthesized an RTL model to measure the hardware costs among various quantization methods. For each quantization operation, it reduces area cost by about 15 times and energy consumption by about 9 times, compared to a strong baseline. Xue Geng, Jie Fu 0001, Jie Lin 0001, Mohamed M. Sabry, Christopher Joseph Pal, Vijay Chandrasekhar 0001 |
DCC | 2 |
| 2019 | Interactive Language Learning by Question AnsweringabstractXingdi Yuan, Marc-Alexandre Côté, Jie Fu, Zhouhan Lin, Chris Pal, Yoshua Bengio, Adam Trischler. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Xingdi Yuan, Marc-Alexandre Côté, Jie Fu 0001, Zhouhan Lin, Christopher Joseph Pal, Yoshua Bengio, Adam Trischler |
EMNLP/IJCNLP (1) | 3 |
| 2019 | BFGAN: Backward and Forward Generative Adversarial Networks for Lexically Constrained Sentence GenerationabstractIncorporating prior knowledge like lexical constraints into the model's output to generate meaningful and coherent sentences has many applications in dialogue system, machine translation, image captioning, etc. However, existing auto-regressive models incrementally generate sentences from left to right via beam search, which makes it difficult to directly introduce lexical constraints into the generated sentences. In this paper, we propose a new algorithmic framework, dubbed BFGAN, to address this challenge. Specifically, we employ a backward generator and a forward generator to generate lexically constrained sentences together, and use a discriminator to guide the joint training of two generators by assigning them reward signals. Due to the difficulty of BFGAN training, we propose several training techniques to make the training process more stable and efficient. Our extensive experiments on three large-scale datasets with human evaluation demonstrate that BFGAN has significant improvements over previous methods. Dayiheng Liu, Jie Fu 0001, Qian Qu, Jiancheng Lv 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2016 | DrMAD: Distilling Reverse-Mode Automatic Differentiation for Optimizing Hyperparameters of Deep Neural Networks
Jie Fu 0001, Hongyin Luo, Jiashi Feng, Kian Hsiang Low, Tat-Seng Chua |
IJCAI | 1 |
| 2015 | AffectiveSpace 2: Enabling Affective Intuition for Concept-Level Sentiment AnalysisabstractPredicting the affective valence of unknown multi-word expressions is key for concept-level sentiment analysis. AffectiveSpace 2 is a vector space model, built by means of random projection, that allows for reasoning by analogy on natural language con- cepts. By reducing the dimensionality of affec- tive common-sense knowledge, the model allows semantic features associated with concepts to be generalized and, hence, allows concepts to be intu- itively clustered according to their semantic and affective relatedness. Such an affective intuition (so called because it does not rely on explicit fea- tures, but rather on implicit analogies) enables the inference of emotions and polarity conveyed by multi-word expressions, thus achieving efficient concept-level sentiment analysis. Erik Cambria, Jie Fu 0001, Federica Bisio, Soujanya Poria |
AAAI | 2 |