Junxian He

dblp:188/6127 · DBLP profile ↗
← Back
50ranked-venue papers
10as first author
39since 2021 · last 2026
0009-0007-9559-6941ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 48 · 10 first-author · 37 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 How Can Synthetic Data Improve Multilingual Language Model Pretraining? A Data Quality Perspective
abstract
Tongyao Zhu, Qian Liu, Chang Ma, Jinghan Zhang, Longxu Dou, Junxian He, Shiqi Chen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Tongyao Zhu, Qian Liu 0033, Jinghan Zhang 0006, Longxu Dou, Junxian He, Shiqi Chen 0002
ACL (1)6
2025 Revisiting Scaling Laws for Language Models: The Role of Data Quality and Training Strategies
abstract
Zhengyu Chen, Siqi Wang, Teng Xiao, Yudong Wang, Shiqi Chen, Xunliang Cai, Junxian He, Jingang Wang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Zhengyu Chen 0001, Teng Xiao, Shiqi Chen 0002, Junxian He, Jingang Wang
ACL (1)7
2025 OS-Genesis: Automating GUI Agent Trajectory Construction via Reverse Task Synthesis
abstract
Graphical User Interface (GUI) agents powered by Vision-Language Models (VLMs) have demonstrated human-like computer control capability. Despite their utility in advancing digital automation, a critical bottleneck persists: collecting high-quality trajectory data for training. Common practices for collecting such data rely on human supervision or synthetic data generation through executing pre-defined tasks, which are either resource-intensive or unable to guarantee data quality. Moreover, these methods suffer from limited data diversity and significant gaps between synthetic data and real-world environments. To address these challenges, we propose OS-Genesis, a novel GUI data synthesis pipeline that reverses the conventional trajectory collection process. Instead of relying on pre-defined tasks, OS-Genesis enables agents first to perceive environments and perform step-wise interactions, then retrospectively derive high-quality tasks to enable trajectory-level exploration. A trajectory reward model is then employed to ensure the quality of the generated trajectories. We demonstrate that training GUI agents with OS-Genesis significantly improves their performance on highly challenging online benchmarks. In-depth analysis further validates OS-Genesis's efficiency and its superior data quality and diversity compared to existing synthesis methods. Our codes, data, and checkpoints are available at OS-Genesis Homepage.
Qiushi Sun, Kanzhi Cheng, Zichen Ding 0002, Chuanyang Jin, Yian Wang 0003, Fangzhi Xu, Chengyou Jia, Zhoumianze Liu, Ben Kao, Guohao Li 0001, Junxian He, Yu Qiao 0001, Zhiyong Wu 0003
ACL (1)13
2025 Genius: A Generalizable and Purely Unsupervised Self-Training Framework For Advanced Reasoning
abstract
Fangzhi Xu, Hang Yan, Chang Ma, Haiteng Zhao, Qiushi Sun, Kanzhi Cheng, Junxian He, Jun Liu, Zhiyong Wu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Fangzhi Xu, Hang Yan 0010, Haiteng Zhao, Qiushi Sun, Kanzhi Cheng, Junxian He, Jun Liu 0002, Zhiyong Wu 0003
ACL (1)7
2025 Non-myopic Generation of Language Models for Reasoning and Planning
abstract
Large Language Models (LLMs) have demonstrated remarkable abilities in reasoning and planning. Despite their success in various domains, such as mathematical problem-solving and coding, LLMs face challenges in ensuring reliable and optimal planning due to the inherent myopic nature of autoregressive decoding. This paper revisits LLM reasoning from an optimal control perspective, proposing a novel method, Predictive-Decoding, that leverages Model Predictive Control to enhance planning accuracy. By reweighting LLM distributions based on foresight trajectories, Predictive-Decoding aims to mitigate early errors and promote non-myopic planning. Our experiments show significant improvements across a wide range of tasks in math, coding, and agent-based scenarios. Furthermore, Predictive-Decoding demonstrates computational efficiency, outperforming search baselines while utilizing inference compute more effectively. This study provides insights into optimizing LLM planning capabilities.
Haiteng Zhao, Junlei Zhang, Junxian He, Lingpeng Kong
ICLR4
2025 B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners
abstract
In the absence of extensive human-annotated data for complex reasoning tasks, self-improvement -- where models are trained on their own outputs -- has emerged as a primary method for enhancing performance. Recently, the approach to self-improvement has shifted toward a more dynamic, online fashion through iterative training processes. However, the critical factors underlying the mechanism of these self-improving methods remain poorly understood, such as under what conditions self-improvement is effective, and what are the bottlenecks in the current iterations. In this work, we identify and propose methods to monitor two pivotal factors in this iterative process: (1) the model's ability to explore and generate high-quality responses among multiple candidates (exploration); and (2) the reliability of external rewards in selecting the best responses from the generated outputs (exploitation). These factors are inherently moving targets throughout the self-improvement cycles, yet their dynamics are rarely discussed in prior research -- It remains unclear what impedes continual model enhancement after only a few iterations. Using mathematical reasoning as a case study, we begin with a quantitative analysis to track the dynamics of exploration and exploitation, discovering that a model's exploratory capabilities rapidly deteriorate over iterations, and the effectiveness of exploiting external rewards diminishes as well due to shifts in distribution from the original policy. Motivated by these findings, we introduce B-STaR, a Self-Taught Reasoning framework that autonomously adjusts configurations across iterations to Balance exploration and exploitation, thereby optimizing the self-teaching effectiveness based on the current policy model and available rewards. Our experiments in mathematical reasoning demonstrate that B-STaR not only enhances the model's exploratory capabilities throughout training but also achieves a more effective balance between exploration and exploitation, leading to superior performance. Crucially, this work deconstructs the opaque nature of self-training algorithms, elucidating the interpretable dynamics throughout the process and highlighting current limitations for future research to address.
Weihao Zeng 0003, Yuzhen Huang 0002, Zifei Shan, Junxian He
ICLR6
2025 Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging
abstract
Vision-Language Models (VLMs) combine visual perception with the general capabilities, such as reasoning, of Large Language Models (LLMs). However, the mechanisms by which these two abilities can be combined and contribute remain poorly understood. In this work, we explore to compose perception and reasoning through model merging that connects parameters of different models. Unlike previous works that often focus on merging models of the same kind, we propose merging models across modalities, enabling the incorporation of the reasoning capabilities of LLMs into VLMs. Through extensive experiments, we demonstrate that model merging offers a successful pathway to transfer reasoning abilities from LLMs to VLMs in a training-free manner. Moreover, we utilize the merged models to understand the internal mechanism of perception and reasoning and how merging affects it. We find that perception capabilities are predominantly encoded in the early layers of the model, whereas reasoning is largely facilitated by the middle-to-late layers. After merging, we observe that all layers begin to contribute to reasoning, whereas the distribution of perception abilities across layers remains largely unchanged. These observations shed light on the potential of model merging as a tool for multimodal integration and interpretation.
Shiqi Chen 0002, Jinghan Zhang 0006, Tongyao Zhu, Wei Liu 0131, Siyang Gao, Miao Xiong, Manling Li, Junxian He
ICML8
2025 Why Is Spatial Reasoning Hard for VLMs? An Attention Mechanism Perspective on Focus Areas
abstract
Large Vision Language Models (VLMs) have long struggled with spatial reasoning tasks. Surprisingly, even simple spatial reasoning tasks, such as recognizing “under” or “behind” relationships between only two objects, pose significant challenges for current VLMs. We believe it is crucial to use the lens of mechanism interpretability, opening up the model and diving into model’s internal states to examine the interactions between image and text tokens during spatial reasoning. Our analysis of attention behaviors reveals significant differences in how VLMs allocate attention to image versus text. By tracing the areas of images that receive the highest attention scores throughout intermediate layers, we observe a notable pattern: errors often coincide with attention being misdirected towards irrelevant objects within the image. Moreover, such attention patterns exhibit substantial differences between familiar (e.g., “on the left side of ”) and unfamiliar (e.g.,“in front of ”) spatial relationships. Motivated by these findings, we propose ADAPTVIS based on inference-time confidence scores to sharpen the attention on highly relevant regions when the model exhibits high confidence, while smoothing and broadening the attention window to consider a wider context when confidence is lower. This training-free decoding method shows significant improvement (e.g., up to a 50 absolute point improvement) on spatial reasoning benchmarks such as WhatsUp and VSR with negligible additional cost.
Shiqi Chen 0002, Tongyao Zhu, Ruochen Zhou, Jinghan Zhang 0006, Siyang Gao, Juan Carlos Niebles, Mor Geva, Junxian He, Jiajun Wu 0001, Manling Li
ICML8
2025 Diving into Self-Evolving Training for Multimodal Reasoning
abstract
Self-evolving training—where models iteratively learn from their own outputs—has emerged as a key approach for complex reasoning tasks, addressing the scarcity of high-quality chain-of-thought data. However, its effectiveness in multimodal reasoning, a domain more intricate than text-only reasoning, remains underexplored, and the understanding of critical factors in this training paradigm remains limited. Furthermore, a central challenge for this training method is performance saturation, which impedes further improvements and scalability. Inspired by reinforcement learning (RL), in this paper, we reframe self-evolving training for multimodal reasoning through the lens of RL, identifying three pivotal factors: $\textit{Training Method}$, $\textit{Reward Model}$, and $\textit{Prompt Variation}$. Through systematic analysis, we establish relatively optimal design principles that significantly enhance multimodal reasoning capabilities. Moreover, delving deeper into training dynamics, we uncover the roots of saturation and propose a new automatic balancing mechanism to mitigate this limitation. Building on these insights, we propose M-STaR (**M**ultimodal **S**elf-evolving **T**r**a**ining for **R**easoning), a framework that achieves consistent performance gains across models of varying sizes and diverse benchmarks. All resources will be made publicly available.
Wei Liu 0131, Yu Cheng 0001, Junxian He
ICML6
2025 CodeIO: Condensing Reasoning Patterns via Code Input-Output Prediction
abstract
Reasoning is a fundamental capability of Large Language Models. While prior research predominantly focuses on enhancing narrow skills like math or code generation, improving performance on many other reasoning tasks remains challenging due to sparse and fragmented training data. To address this issue, we propose CodeI/O, a novel approach that systematically condenses diverse reasoning patterns inherently embedded in contextually-grounded codes, through transforming the original code into a code input-output prediction format. By training models to predict inputs/outputs given code and test cases entirely in natural language as Chain-of-Thought (CoT) rationales, we expose them to universal reasoning primitives—like logic flow planning, state-space searching, decision tree traversal, and modular decomposition—while decoupling structured reasoning from code-specific syntax and preserving procedural rigor. Experimental results demonstrate CodeI/O leads to consistent improvements across symbolic, scientific, logic, math & numerical, and commonsense reasoning tasks. By matching the existing ground-truth outputs or re-executing the code with predicted inputs, we can verify each prediction and further enhance the CoTs through multi-turn revision, resulting in CodeI/O++ and achieving higher performance. Our data and models will be publicly available.
Daya Guo, Dejian Yang, Runxin Xu, Yu Wu 0024, Junxian He
ICML6
2025 Predictive Data Selection: The Data That Predicts Is the Data That Teaches
abstract
Language model pretraining involves training on extensive corpora, where data quality plays a pivotal role. In this work, we aim to directly estimate the contribution of data during pretraining and select pretraining data in an efficient manner. Specifically, we draw inspiration from recent findings showing that compression efficiency (i.e., normalized loss) of diverse models on certain text correlates strongly with their downstream performance, when the text domain aligns with the downstream benchmarks (Huang et al., 2024). Building on this observation, we hypothesize that data on which model losses are predictive of downstream abilities also contribute effectively to learning, which shares similar intuition with Thrush et al. (2024). To leverage this insight, we introduce predictive data selection (PreSelect), a lightweight and efficient data selection method that requires training and deploying only a fastText-based scorer. Through comprehensive experiments with 1B and 3B parameter models, we demonstrate that models trained on 30B tokens selected with PreSelect surpass the performance of the vanilla baseline trained on 300B tokens, achieving a 10x reduction in compute requirements. Furthermore, PreSelect significantly outperforms other competitive data selection baselines, such as DCLM and FineWeb-Edu on a scale of 3B models trained on 100B tokens. We open-source our trained data selection scorer along with the curated datasets at https://github.com/hkust-nlp/PreSelect.
KaShun Shum, Yuzhen Huang 0002, Hongjian Zou, Yixuan Liao, Xiaoxin Chen 0001, Qian Liu 0033, Junxian He
ICML8
2025 SynLogic: Synthesizing Verifiable Reasoning Data at Scale for Learning Logical Reasoning and Beyond
abstract
Recent advances such as OpenAI-o1 and DeepSeek R1 have demonstrated the potential of Reinforcement Learning (RL) to enhance reasoning abilities in Large Language Models (LLMs). While open-source replication efforts have primarily focused on mathematical and coding domains, methods and resources for developing general reasoning capabilities remain underexplored. This gap is partly due to the challenge of collecting diverse and verifiable reasoning data suitable for RL. We hypothesize that logical reasoning is critical for developing general reasoning capabilities, as logic forms a fundamental building block of reasoning. In this work, we present SynLogic, a data synthesis framework and dataset that generates diverse logical reasoning data at scale, encompassing 35 diverse logical reasoning tasks. The SynLogic approach enables controlled synthesis of data with adjustable difficulty and quantity. Importantly, all examples can be verified by simple rules, making them ideally suited for RL with verifiable rewards. In our experiments, we validate the effectiveness of RL training on the SynLogic dataset based on 7B and 32B models. SynLogic leads to state-of-the-art logical reasoning performance among open-source datasets, surpassing DeepSeek-R1-Distill-Qwen-32B by 6 points on BBEH. Furthermore, mixing SynLogic data with mathematical and coding tasks improves the training efficiency of these domains and significantly enhances reasoning generalization. Notably, our mixed training model outperforms DeepSeek-R1-Zero-Qwen-32B across multiple benchmarks. These findings position SynLogic as a valuable resource for advancing the broader reasoning capabilities of LLMs. We will open-source both the data synthesis pipeline and the SynLogic dataset.
Junteng Liu, Yuanxiang Fan, Zhuo Jiang, Yongyi Hu, Yiqi Shi, Shitong Weng, Aili Chen, Shiqi Chen 0002, Mozhi Zhang, Junxian He
NeurIPS13
2024 Prompt Optimization via Adversarial In-Context Learning
abstract
Xuan Long Do, Yiran Zhao, Hannah Brown, Yuxi Xie, James Xu Zhao, Nancy F. Chen, Kenji Kawaguchi, Michael Shieh, Junxian He. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Do Xuan Long, Yiran Zhao 0006, Hannah Brown, Yuxi Xie, James Xu Zhao, Nancy F. Chen, Kenji Kawaguchi, Michael Shieh, Junxian He
ACL (1)9
2024 On the Universal Truthfulness Hyperplane Inside LLMs
abstract
While large language models (LLMs) have demonstrated remarkable abilities across various fields, hallucination remains a significant challenge.Recent studies have explored hallucinations through the lens of internal representations, proposing mechanisms to decipher LLMs' adherence to facts.However, these approaches often fail to generalize to out-of-distribution data, leading to concerns about whether internal representation patterns reflect fundamental factual awareness, or only overfit spurious correlations on the specific datasets.In this work, we investigate whether a universal truthfulness hyperplane that distinguishes the model's factually correct and incorrect outputs exists within the model.To this end, we scale up the number of training datasets and conduct an extensive evaluation -we train the truthfulness hyperplane on a diverse collection of over 40 datasets and examine its cross-task, cross-domain, and in-domain generalization.Our results indicate that increasing the diversity of the training datasets significantly enhances the performance in all scenarios, while the volume of data samples plays a less critical role.This finding supports the optimistic hypothesis that a universal truthfulness hyperplane may indeed exist within the model, offering promising directions for future research.Code is publicly available at https://github.com/hkust-nlp/ Universal_Truthfulness_Hyperplane.Tend to overfit
Junteng Liu, Shiqi Chen 0002, Yu Cheng 0001, Junxian He
EMNLP4
2024 Belief Revision: The Adaptability of Large Language Models Reasoning
abstract
The capability to reason from text is crucial for real-world NLP applications.Real-world scenarios often involve incomplete or evolving data.In response, individuals update their beliefs and understandings accordingly.However, most existing evaluations assume that language models (LMs) operate with consistent information.We introduce Belief-R 1 , a new dataset designed to test LMs' belief revision ability when presented with new evidence.Inspired by how humans suppress prior inferences, this task assesses LMs within the newly proposed delta reasoning (∆R) framework.Belief-R features sequences of premises designed to simulate scenarios where additional information could necessitate prior conclusions drawn by LMs.We evaluate ∼30 LMs across diverse prompting strategies and found that LMs generally struggle to appropriately revise their beliefs in response to new information.Further, models adept at updating often underperformed in scenarios without necessary updates, highlighting a critical trade-off.These insights underscore the importance of improving LMs' adaptiveness to changing information, a step toward more reliable AI systems.
Bryan Wilie, Samuel Cahyawijaya, Etsuko Ishii, Junxian He, Pascale Fung
EMNLP4
2024 What Makes Good Data for Alignment? A Comprehensive Study of Automatic Data Selection in Instruction Tuning
abstract
Instruction tuning is a standard technique employed to align large language models to end tasks and user preferences after the initial pretraining phase. Recent research indicates the critical role of data engineering in instruction tuning -- when appropriately selected, only limited data is necessary to achieve superior performance. However, we still lack a principled understanding of what makes good instruction tuning data for alignment, and how we should select data automatically and effectively. In this work, we delve deeply into automatic data selection strategies for alignment. We start with controlled studies to measure data across three dimensions: complexity, quality, and diversity, along which we examine existing methods and introduce novel techniques for enhanced data measurement. Subsequently, we propose a simple strategy to select data samples based on the measurement. We present Deita (short for Data-Efficient Instruction Tuning for Alignment), a series of models fine-tuned from LLaMA models using data samples automatically selected with our proposed approach. When assessed through both automatic metrics and human evaluation, Deita performs better or on par with the state-of-the-art open-source alignment models such as Vicuna and WizardLM with only 6K training data samples -- 10x less than the data used in the baselines. We anticipate this work to provide clear guidelines and tools on automatic data selection, aiding researchers and practitioners in achieving data-efficient alignment.
Wei Liu 0131, Weihao Zeng 0003, Keqing He 0001, Yong Jiang 0001, Junxian He
ICLR5
2024 Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs
abstract
Empowering large language models (LLMs) to accurately express confidence in their answers is essential for reliable and trustworthy decision-making. Previous confidence elicitation methods, which primarily rely on *white-box access* to internal model information or model fine-tuning, have become less suitable for LLMs, especially closed-source commercial APIs. This leads to a growing need to explore the untapped area of *black-box* approaches for LLM uncertainty estimation. To better break down the problem, we define a systematic framework with three components: *prompting* strategies for eliciting verbalized confidence, *sampling* methods for generating multiple responses, and *aggregation* techniques for computing consistency. We then benchmark these methods on two key tasks—confidence calibration and failure prediction—across five types of datasets (e.g., commonsense and arithmetic reasoning) and five widely-used LLMs including GPT-4 and LLaMA 2 Chat. Our analysis uncovers several key insights: 1) LLMs, when verbalizing their confidence, tend to be *overconfident*, potentially imitating human patterns of expressing confidence. 2) As model capability scales up, both calibration and failure prediction performance improve, yet still far from ideal performance. 3) Employing our proposed strategies, such as human-inspired prompts, consistency among multiple responses, and better aggregation strategies can help mitigate this overconfidence from various perspectives. 4) Comparisons with white-box methods indicate that while white-box methods perform better, the gap is narrow, e.g., 0.522 to 0.605 in AUROC. Despite these advancements, none of these techniques consistently outperform others, and all investigated methods struggle in challenging tasks, such as those requiring professional knowledge, indicating significant scope for improvement. We believe this study can serve as a strong baseline and provide insights for eliciting confidence in black-box LLMs. The code is publicly available at https://github.com/MiaoXiong2320/llm-uncertainty.
Miao Xiong, Xinyang Lu, Jie Fu 0001, Junxian He, Bryan Hooi
ICLR6
2024 In-Context Sharpness as Alerts: An Inner Representation Perspective for Hallucination Mitigation
abstract
Large language models (LLMs) frequently hallucinate, e.g., making factual errors, yet our understanding of why they make these errors remains limited. In this study, we aim to understand the underlying mechanisms of LLM hallucinations from the perspective of *inner representations*. We discover a pattern associated with hallucinations: correct generations tend to have *sharper* context activations in the hidden states of the in-context tokens, compared to that of the incorrect generations. Leveraging this signal, we propose an entropy-based metric to quantify the *sharpness* among the in-context hidden states and incorporate it into the decoding process, i.e, use the entropy value to adjust the next token prediction distribution to improve the factuality and overall quality of the generated text. Experiments on knowledge-seeking datasets (Natural Questions, HotpotQA, TriviaQA) and hallucination benchmark (TruthfulQA) demonstrate our consistent effectiveness, e.g., up to 8.6 absolute points on TruthfulQA. We believe this study can improve our understanding of hallucinations and serve as a practical solution for hallucination mitigation.
Shiqi Chen 0002, Miao Xiong, Junteng Liu, Zhengxuan Wu, Teng Xiao, Siyang Gao, Junxian He
ICML7
2024 Uncertainty of Thoughts: Uncertainty-Aware Planning Enhances Information Seeking in LLMs
abstract
In the face of uncertainty, the ability to *seek information* is of fundamental importance. In many practical applications, such as medical diagnosis and troubleshooting, the information needed to solve the task is not initially given, and has to be actively sought by asking follow-up questions (for example, a doctor asking a patient for more details about their symptoms). In this work, we introduce **Uncertainty of Thoughts (UoT)**, an algorithm to augment large language models with the ability to actively seek information by asking effective questions. UoT combines: 1. An *uncertainty-aware simulation approach* which enables the model to simulate possible future scenarios and how likely they are to occur, 2. *Uncertainty-based rewards* motivated by information gain which incentivizes the model to seek information, and 3. A *reward propagation scheme* to select the optimal question to ask in a way that maximizes the expected reward. In experiments on medical diagnosis, troubleshooting and the `20 Questions' game, UoT achieves an average performance improvement of 38.1% in the rate of successful task completion across multiple LLMs compared with direct prompting, and also improves efficiency (i.e., the number of questions needed to complete the task).
Chumin Liu, Xidong Feng, Yilun Zhao 0001, See-Kiong Ng, Anh Tuan Luu, Junxian He, Pang Wei W. Koh, Bryan Hooi
NeurIPS7
2024 AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents
abstract
Evaluating large language models (LLMs) as general-purpose agents is essential for understanding their capabilities and facilitating their integration into practical applications. However, the evaluation process presents substantial challenges. A primary obstacle is the benchmarking of agent performance across diverse scenarios within a unified framework, especially in maintaining partially-observable environments and ensuring multi-round interactions. Moreover, current evaluation frameworks mostly focus on the final success rate, revealing few insights during the process and failing to provide a deep understanding of the model abilities. To address these challenges, we introduce AgentBoard, a pioneering comprehensive benchmark and accompanied open-source evaluation framework tailored to analytical evaluation of LLM agents. AgentBoard offers a fine-grained progress rate metric that captures incremental advancements as well as a comprehensive evaluation toolkit that features easy assessment of agents for multi-faceted analysis through interactive visualization. This not only sheds light on the capabilities and limitations of LLM agents but also propels the interpretability of their performance to the forefront. Ultimately, AgentBoard serves as a significant step towards demystifying agent behaviors and accelerating the development of stronger LLM agents.
Junlei Zhang, Cheng Yang 0007, Yujiu Yang 0001, Yaohui Jin, Zhen-Zhong Lan, Lingpeng Kong, Junxian He
NeurIPS9
2024 DART-Math: Difficulty-Aware Rejection Tuning for Mathematical Problem-Solving
abstract
Solving mathematical problems requires advanced reasoning abilities and presents notable challenges for large language models. Previous works usually synthesize data from proprietary models to augment existing datasets, followed by instruction tuning to achieve top-tier results. However, our analysis of these datasets reveals severe biases towards easy queries, with frequent failures to generate any correct response for the most challenging queries. Hypothesizing that difficult queries are crucial to learning complex reasoning, we propose *Difficulty-Aware Rejection Tuning* (`DART`), a method that allocates difficult queries more trials during the synthesis phase, enabling more extensive training on difficult samples. Utilizing `DART`, we have created new datasets for mathematical problem-solving that focus more on difficult queries and are substantially smaller than previous ones. Remarkably, our synthesis process solely relies on a 7B-sized open-weight model, without reliance on the commonly used proprietary GPT-4. We fine-tune various base models on our datasets ranging from 7B to 70B in size, resulting in a series of strong models called `DART-Math`. In comprehensive in-domain and out-of-domain evaluation on 6 mathematical benchmarks, `DART-Math` outperforms vanilla rejection tuning significantly, being superior or comparable to previous arts, despite using much smaller datasets and no proprietary models. Furthermore, our results position our synthetic datasets as the most effective and cost-efficient publicly available resources for advancing mathematical problem-solving. Our datasets, models and code are publicly available at https://github.com/hkust-nlp/dart-math.
Yuxuan Tong, Ruidong Wu, Junxian He
NeurIPS5
2024 K2: A Foundation Language Model for Geoscience Knowledge Understanding and Utilization
abstract
Large language models (LLMs) have achieved great success in general domains of natural language processing. In this paper, we bring LLMs to the realm of geoscience with the objective of advancing research and applications in this field. To this end, we present the first-ever LLM in geoscience, K2, alongside a suite of resources developed to further promote LLM research within geoscience. For instance, we have curated the first geoscience instruction tuning dataset, GeoSignal, which aims to align LLM responses to geoscience-related user queries. Additionally, we have established the first geoscience benchmark, GeoBench, to evaluate LLMs in the context of geoscience. In this work, we experiment with a complete recipe to adapt a pre-trained general-domain LLM to the geoscience domain. Specifically, we further train the LLaMA-7B model on 5.5B tokens of geoscience text corpus, including over 1 million pieces of geoscience literature, and utilize GeoSignal's supervised data to fine-tune the model. Moreover, we share a protocol that can efficiently gather domain-specific data and construct domain-supervised data, even in situations where manpower is scarce. Meanwhile, we equip K2 with the abilities of using tools to be a naive geoscience aide. Experiments conducted on the GeoBench demonstrate the effectiveness of our approach and datasets on geoscience knowledge understanding and utilization.We open-source all the training data and K2 model checkpoints at https://github.com/davendw49/k2
Cheng Deng 0001, Tianhang Zhang, Zhongmou He, Qiyuan Chen 0002, Yi Xu 0004, Luoyi Fu, Weinan Zhang 0001, Xinbing Wang, Chenghu Zhou, Zhouhan Lin, Junxian He
WSDM12
2023 Simple Temporal Adaptation to Changing Label Sets: Hashtag Prediction via Dense KNN
abstract
User-generated social media data is constantly changing as new trends influence online discussion and personal information is deleted due to privacy concerns.However, traditional NLP models rely on fixed training datasets, which means they are unable to adapt to temporal change-both test distribution shift and deleted training data-without frequent, costly re-training.In this paper, we study temporal adaptation through the task of longitudinal hashtag prediction and propose a nonparametric dense retrieval technique, which does not require re-training, as a simple but effective solution.In experiments on a newly collected, publicly available, year-long Twitter dataset exhibiting temporal distribution shift, our method improves by 64% over the best static parametric baseline while avoiding costly gradient-based re-training.Our approach is also particularly well-suited to dynamically deleted user data in line with data privacy laws, with negligible computational cost/performance loss.
Niloofar Mireshghallah, Nikolai Vogler, Junxian He, Omar Florez, Ahmed El-Kishky, Taylor Berg-Kirkpatrick
EMNLP3
2023 Contrastive Learning of Sentence Embeddings from Scratch
abstract
Contrastive learning has been the dominant approach to train state-of-the-art sentence embeddings.Previous studies have typically learned sentence embeddings either through the use of human-annotated natural language inference (NLI) data or via large-scale unlabeled sentences in an unsupervised manner.However, even in the case of 1;unlabeled data, their acquisition presents challenges in certain domains due to various reasons.To address these issues, we present SynCSE, a contrastive learning framework that trains sentence embeddings with synthesized data.Specifically, we explore utilizing large language models to synthesize the required data samples for contrastive learning, including (1) producing positive and negative annotations given unlabeled sentences (SynCSE-partial), and (2) generating sentences along with their corresponding annotations from scratch (SynCSE-scratch).Experimental results on sentence similarity and reranking tasks indicate that both SynCSE-partial and SynCSE-scratch greatly outperform unsupervised baselines, and SynCSE-partial even achieves comparable performance to the supervised models in most settings.1 * Work done during Junlei's visit to HKUST.† Corresponding author. 1 Code and the synthesized datasets are available at https://github.com/hkust-nlp/SynCSE.I saw a sunset at the beach today.My city exploration led me to a beautiful building.
Junlei Zhang, Zhen-Zhong Lan, Junxian He
EMNLP3
2023 Mega: Moving Average Equipped Gated Attention
Xuezhe Ma, Chunting Zhou, Xiang Kong, Junxian He, Liangke Gui, Graham Neubig, Jonathan May, Luke Zettlemoyer
ICLR4
2023 FELM: Benchmarking Factuality Evaluation of Large Language Models
abstract
Assessing factuality of text generated by large language models (LLMs) is an emerging yet crucial research area, aimed at alerting users to potential errors and guiding the development of more reliable LLMs. Nonetheless, the evaluators assessing factuality necessitate suitable evaluation themselves to gauge progress and foster advancements. This direction remains under-explored, resulting in substantial impediments to the progress of factuality evaluators. To mitigate this issue, we introduce a benchmark for Factuality Evaluation of large Language Models, referred to as FELM. In this benchmark, we collect responses generated from LLMs and annotate factuality labels in a fine-grained manner. Contrary to previous studies that primarily concentrate on the factuality of world knowledge (e.g. information from Wikipedia), FELM focuses on factuality across diverse domains, spanning from world knowledge to math and reasoning. Our annotation is based on text segments, which can help pinpoint specific factual errors. The factuality annotations are further supplemented by predefined error types and reference links that either support or contradict the statement. In our experiments, we investigate the performance of several LLM-based factuality evaluators on FELM, including both vanilla LLMs and those augmented with retrieval mechanisms and chain-of-thought processes. Our findings reveal that while retrieval aids factuality evaluation, current LLMs are far from satisfactory to faithfully detect factual errors.
Shiqi Chen 0002, Yiran Zhao 0006, Jinghan Zhang 0006, I-Chun Chern, Siyang Gao, Pengfei Liu 0003, Junxian He
NeurIPS7
2023 C-Eval: A Multi-Level Multi-Discipline Chinese Evaluation Suite for Foundation Models
abstract
New NLP benchmarks are urgently needed to align with the rapid development of large language models (LLMs). We present C-Eval, the first comprehensive Chinese evaluation suite designed to assess advanced knowledge and reasoning abilities of foundation models in a Chinese context. C-Eval comprises multiple-choice questions across four difficulty levels: middle school, high school, college, and professional. The questions span 52 diverse disciplines, ranging from humanities to science and engineering. C-Eval is accompanied by C-Eval Hard, a subset of very challenging subjects in C-Eval that requires advanced reasoning abilities to solve. We conduct a comprehensive evaluation of the most advanced LLMs on C-Eval, including both English- and Chinese-oriented models. Results indicate that only GPT-4 could achieve an average accuracy of over 60%, suggesting that there is still significant room for improvement for current LLMs. We anticipate C-Eval will help analyze important strengths and shortcomings of foundation models, and foster their development and growth for Chinese users.
Yuzhen Huang 0002, Yuzhuo Bai, Junlei Zhang, Jinghan Zhang 0006, Tangjun Su, Junteng Liu, Chuancheng Lv, Jiayi Lei, Maosong Sun 0001, Junxian He
NeurIPS13
2023 Self-Evaluation Guided Beam Search for Reasoning
abstract
Breaking down a problem into intermediate steps has demonstrated impressive performance in Large Language Model (LLM) reasoning. However, the growth of the reasoning chain introduces uncertainty and error accumulation, making it challenging to elicit accurate final results. To tackle this challenge of uncertainty in multi-step reasoning, we introduce a stepwise self-evaluation mechanism to guide and calibrate the reasoning process of LLMs. We propose a decoding algorithm integrating the self-evaluation guidance via stochastic beam search. The self-evaluation guidance serves as a better-calibrated automatic criterion, facilitating an efficient search in the reasoning space and resulting in superior prediction quality. Stochastic beam search balances exploitation and exploration of the search space with temperature-controlled randomness. Our approach surpasses the corresponding Codex-backboned baselines in few-shot accuracy by $6.34$%, $9.56$%, and $5.46$% on the GSM8K, AQuA, and StrategyQA benchmarks, respectively. Experiment results with Llama-2 on arithmetic reasoning demonstrate the efficiency of our method in outperforming the baseline methods with comparable computational budgets. Further analysis in multi-step reasoning finds our self-evaluation guidance pinpoints logic failures and leads to higher consistency and robustness. Our code is publicly available at [https://guideddecoding.github.io/](https://guideddecoding.github.io/).
Yuxi Xie, Kenji Kawaguchi, Yiran Zhao 0006, James Xu Zhao, Min-Yen Kan, Junxian He, Qizhe Xie
NeurIPS6
2023 Composing Parameter-Efficient Modules with Arithmetic Operation
abstract
As an efficient alternative to conventional full fine-tuning, parameter-efficient fine-tuning (PEFT) is becoming the prevailing method to adapt pretrained language models. In PEFT, a lightweight module is learned on each dataset while the underlying pretrained language model remains unchanged, resulting in multiple compact modules representing diverse skills when applied to various domains and tasks. In this paper, we propose to compose these parameter-efficient modules through linear arithmetic operations in the weight space, thereby integrating different module capabilities. Specifically, we first define an addition and negation operator for the module, and then further compose these two basic operators to perform flexible arithmetic. Our approach requires no additional training and enables highly flexible module composition. We apply different arithmetic operations to compose the parameter-efficient modules for (1) distribution generalization, (2) multi-tasking, (3) detoxifying, and (4) domain transfer. Additionally, we extend our approach to detoxify Alpaca-LoRA, the latest instruction-tuned large language model based on LLaMA. Empirical results demonstrate that our approach produces new and effective parameter-efficient modules that significantly outperform existing ones across all settings.
Jinghan Zhang 0006, Shiqi Chen 0002, Junteng Liu, Junxian He
NeurIPS4
2023 Low-cost real-time VLSI system for high-accuracy optical flow estimation using biological motion features and random forests
Cong Shi 0003, Junxian He, Shrinivas J. Pundlik, Xichuan Zhou, Nanjian Wu, Gang Luo 0003
Sci. China Inf. Sci.2
2023 An 8-T Processing-in-Memory SRAM Cell-Based Pixel-Parallel Array Processor for Vision Chips
abstract
Vision chip is a high-speed image processing device, featuring a massively-parallel pixel-level processing element (PE) array to boost pixel processing speed. However, the collocated processing unit and fine-grained data memory unit inside each PE impose a huge requirement on memory access bandwidth as well as big area and energy consumption. To overcome this bottleneck, this paper proposes a full custom 8T SRAM-based Processing-in-Memory (PIM) architecture together with a multiplexer-based arithmetic-logic unit (mux-based ALU) to realize pixel-parallel array processor for energy-efficient vision chips. The proposed PIM architecture is constructed by embroidering each dual-port 8T SRAM cell with mux-based ALU, so as to form a PIM PE array. Each PIM PE holds a 130-bit 8T SRAM cell block embedding in-memory logic functions, of which 128-bit 8T SRAM cells serve as the PE memory, and 2-bit 8T SRAM cells act as a buffer register in the PE. A full custom physical layout of a$128\times128$prototyping PIM PE array is designed and evaluated using a 65 nm CMOS technology. The simulation results demonstrate that our proposed PIM PE architecture could operate under a 200 MHz clock frequency with a 1.0 V power supply, and reach a high energy efficiency of 512 GOPS/W and a high area efficiency of 29 GOPS/mm2.
Leyi Chen, Cong Shi 0003, Junxian He, Jianyi Yu, Haibing Wang, Nanjian Wu, Min Tian 0003
IEEE Trans. Circuits Syst. I Regul. Pap.3
2022 CTRLsum: Towards Generic Controllable Text Summarization
abstract
Current summarization systems yield generic summaries that are disconnected from users' preferences and expectations.To address this limitation, we present CTRLSUM, a generic framework to control generated summaries through a set of keywords.During training keywords are extracted automatically without requiring additional human annotations.At test time CTRLSUM features a control function to map control signal to keywords; through engineering the control function, the same trained model is able to be applied to control summaries on various dimensions, while neither affecting the model training process nor the pretrained models.We additionally explore the combination of keywords and text prompts for more control tasks.Experiments demonstrate the effectiveness of CTRLSUM on three domains of summarization datasets and five control tasks: (1) entity-centric and (2) length-controllable summarization, (3) contribution summarization on scientific papers, (4) invention purpose summarization on patent filings, and (5) question-guided summarization on news articles.Moreover, when used in a standard, unconstrained summarization setting, CTRLSUM is comparable or better than strong pretrained systems. 1
Junxian He, Wojciech Kryscinski, Bryan McCann, Nazneen Fatema Rajani, Caiming Xiong
EMNLP1
2022 Towards a Unified View of Parameter-Efficient Transfer Learning
Junxian He, Chunting Zhou, Xuezhe Ma, Taylor Berg-Kirkpatrick, Graham Neubig
ICLR1
2022 Capturing Structural Locality in Non-parametric Language Models
Frank F. Xu, Junxian He, Graham Neubig, Vincent J. Hellendoorn
ICLR2
2022 Neuro-Symbolic Language Modeling with Automaton-augmented Retrieval
abstract
Retrieval-based language models (R-LM) model the probability of natural language text by combining a standard language model (LM) with examples retrieved from an external datastore at test time. While effective, a major bottleneck of using these models in practice is the computationally costly datastore search, which can be performed as frequently as every time step. In this paper, we present RetoMaton - retrieval automaton - which approximates the datastore search, based on (1) saving pointers between consecutive datastore entries, and (2) clustering of entries into "states". This effectively results in a weighted finite automaton built on top of the datastore, instead of representing the datastore as a flat list. The creation of the automaton is unsupervised, and a RetoMaton can be constructed from any text collection: either the original training corpus or from another domain. Traversing this automaton at inference time, in parallel to the LM inference, reduces its perplexity by up to 1.85, or alternatively saves up to 83% of the nearest neighbor searches over $k$NN-LM (Khandelwal et al., 2020) without hurting perplexity. Our code and trained models are available at https://github.com/neulab/retomaton .
Uri Alon 0002, Frank F. Xu, Junxian He, Sudipta Sengupta, Dan Roth 0001, Graham Neubig
ICML3
2021 Dependency Induction Through the Lens of Visual Perception
abstract
Most previous work on grammar induction focuses on learning phrasal or dependency structure purely from text.However, because the signal provided by text alone is limited, recently introduced visually grounded syntax models make use of multimodal information leading to improved performance in constituency grammar induction.However, as compared to dependency grammars, constituency grammars do not provide a straightforward way to incorporate visual information without enforcing language-specific heuristics.In this paper, we propose an unsupervised grammar induction model that leverages word concreteness and a structural vision-based heuristic to jointly learn constituency-structure and dependency-structure grammars.Our experiments find that concreteness is a strong indicator for learning dependency grammars, improving the direct attachment score (DAS) by over 50% as compared to state-of-the-art models trained on pure text.Next, we propose an extension of our model that leverages both word concreteness and visual semantic role labels in constituency and dependency parsing.Our experiments show that the proposed extension outperforms the current state-of-the-art visually grounded models in constituency parsing even with a smaller grammar size. 1
Ruisi Su, Shruti Rijhwani, Hao Zhu 0011, Junxian He, Yonatan Bisk, Graham Neubig
CoNLL4
2021 The Source-Target Domain Mismatch Problem in Machine Translation
abstract
Jiajun Shen, Peng-Jen Chen, Matthew Le, Junxian He, Jiatao Gu, Myle Ott, Michael Auli, Marc’Aurelio Ranzato. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021.
Peng-Jen Chen, Matt Le 0001, Junxian He, Jiatao Gu, Myle Ott, Michael Auli, Marc'Aurelio Ranzato
EACL4
2021 Efficient Nearest Neighbor Language Models
abstract
Non-parametric neural language models (NLMs) learn predictive distributions of text utilizing an external datastore, which allows them to learn through explicitly memorizing the training datapoints.While effective, these models often require retrieval from a large datastore at test time, significantly increasing the inference overhead and thus limiting the deployment of non-parametric NLMs in practical applications.In this paper, we take the recently proposed k-nearest neighbors language model (Khandelwal et al., 2019) as an example, exploring methods to improve its efficiency along various dimensions.Experiments on the standard WikiText-103 benchmark and domain-adaptation datasets show that our methods are able to achieve up to a 6x speed-up in inference speed while retaining comparable performance.The empirical analysis we present may provide guidelines for future research seeking to develop or deploy more efficient non-parametric NLMs. 1
Junxian He, Graham Neubig, Taylor Berg-Kirkpatrick
EMNLP (1)1
2021 CompSNN: A lightweight spiking neural network based on spatiotemporally compressive spike features
Tengxiao Wang, Cong Shi 0003, Xichuan Zhou, Yingcheng Lin, Junxian He, Ping Gan, Ping Li 0042, Ying Wang 0001, Nanjian Wu, Gang Luo 0003
Neurocomputing5
2020 On the Sentence Embeddings from Pre-trained Language Models
abstract
Pre-trained contextual representations like BERT have achieved great success in natural language processing.However, the sentence embeddings from the pre-trained language models without fine-tuning have been found to poorly capture semantic meaning of sentences.In this paper, we argue that the semantic information in the BERT embeddings is not fully exploited.We first reveal the theoretical connection between the masked language model pre-training objective and the semantic similarity task theoretically, and then analyze the BERT sentence embeddings empirically.We find that BERT always induces a non-smooth anisotropic semantic space of sentences, which harms its performance of semantic similarity.To address this issue, we propose to transform the anisotropic sentence embedding distribution to a smooth and isotropic Gaussian distribution through normalizing flows that are learned with an unsupervised objective.Experimental results show that our proposed BERT-flow method obtains significant performance gains over the state-of-the-art sentence embeddings on a variety of semantic textual similarity tasks.The code is available at https://github.com/ bohanli/BERT-flow.
Hao Zhou 0012, Junxian He, Mingxuan Wang, Yiming Yang 0002, Lei Li 0005
EMNLP (1)3
2020 Revisiting Self-Training for Neural Sequence Generation
Junxian He, Jiatao Gu, Marc'Aurelio Ranzato
ICLR1
2020 A Probabilistic Formulation of Unsupervised Text Style Transfer
Junxian He, Xinyi Wang 0001, Graham Neubig, Taylor Berg-Kirkpatrick
ICLR1
2020 Learning Sparse Prototypes for Text Generation
abstract
Prototype-driven text generation uses non-parametric models that first choose from a library of sentence "prototypes" and then modify the prototype to generate the output text. While effective, these methods are inefficient at test time as a result of needing to store and index the entire training corpus. Further, existing methods often require heuristics to identify which prototypes to reference at training time. In this paper, we propose a novel generative model that automatically learns a sparse prototype support set that, nonetheless, achieves strong language modeling performance. This is achieved by (1) imposing a sparsity-inducing prior on the prototype selection distribution, and (2) utilizing amortized variational inference to learn a prototype retrieval function. In experiments, our model outperforms previous prototype-driven language models while achieving up to a 1000x memory reduction, as well as a 1000x speed-up at test time. More interestingly, we show that the learned prototypes are able to capture semantics and syntax at different granularity as we vary the sparsity of prototype selection, and that certain sentence attributes can be controlled by specifying the prototype for generation.
Junxian He, Taylor Berg-Kirkpatrick, Graham Neubig
NeurIPS1
2019 Cross-Lingual Syntactic Transfer through Unsupervised Adaptation of Invertible Projections
abstract
Cross-lingual transfer is an effective way to build syntactic analysis tools in low-resource languages.However, transfer is difficult when transferring to typologically distant languages, especially when neither annotated target data nor parallel corpora are available.In this paper, we focus on methods for cross-lingual transfer to distant languages and propose to learn a generative model with a structured prior that utilizes labeled source data and unlabeled target data jointly.The parameters of source model and target model are softly shared through a regularized log likelihood objective.An invertible projection is employed to learn a new interlingual latent embedding space that compensates for imperfect crosslingual word embedding input.We evaluate our method on two syntactic tasks: part-ofspeech (POS) tagging and dependency parsing.On the Universal Dependency Treebanks, we use English as the only source corpus and transfer to a wide range of target languages.On the 10 languages in this dataset that are distant from English, our method yields an average of 5.2% absolute improvement on POS tagging and 8.3% absolute improvement on dependency parsing over a direct transfer method using state-of-the-art discriminative models. 1 3 Following Ahmad et al. (2019), we use the offline pre-trained alignment matrix present in https://github.com/Babylonpartners/ fastText_multilingual, which contains alignment matrices for 78 languages, which also allows comparison with their numbers in Section 4.3.
Junxian He, Zhisong Zhang, Taylor Berg-Kirkpatrick, Graham Neubig
ACL (1)1
2019 Choosing Transfer Languages for Cross-Lingual Learning
abstract
Yu-Hsiang Lin, Chian-Yu Chen, Jean Lee, Zirui Li, Yuyan Zhang, Mengzhou Xia, Shruti Rijhwani, Junxian He, Zhisong Zhang, Xuezhe Ma, Antonios Anastasopoulos, Patrick Littell, Graham Neubig. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019.
Yu-Hsiang Lin, Chian-Yu Chen, Jean Lee, Mengzhou Xia, Shruti Rijhwani, Junxian He, Zhisong Zhang, Xuezhe Ma, Antonios Anastasopoulos, Patrick Littell, Graham Neubig
ACL (1)8
2019 A Surprisingly Effective Fix for Deep Latent Variable Modeling of Text
abstract
Bohan Li, Junxian He, Graham Neubig, Taylor Berg-Kirkpatrick, Yiming Yang. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Junxian He, Graham Neubig, Taylor Berg-Kirkpatrick, Yiming Yang 0002
EMNLP/IJCNLP (1)2
2019 Lagging Inference Networks and Posterior Collapse in Variational Autoencoders
Junxian He, Daniel Spokoyny, Graham Neubig, Taylor Berg-Kirkpatrick
ICLR (Poster)1
2018 StructVAE: Tree-structured Latent Variable Models for Semi-supervised Semantic Parsing
abstract
Semantic parsing is the task of transducing natural language (NL) utterances into formal meaning representations (MRs), commonly represented as tree structures.Annotating NL utterances with their corresponding MRs is expensive and timeconsuming, and thus the limited availability of labeled data often becomes the bottleneck of data-driven, supervised models.We introduce STRUCTVAE, a variational auto-encoding model for semisupervised semantic parsing, which learns both from limited amounts of parallel data, and readily-available unlabeled NL utterances.STRUCTVAE models latent MRs not observed in the unlabeled data as treestructured latent variables.Experiments on semantic parsing on the ATIS domain and Python code generation show that with extra unlabeled data, STRUCTVAE outperforms strong supervised models. 1
Chunting Zhou, Junxian He, Graham Neubig
ACL (1)3
2018 Unsupervised Learning of Syntactic Structure with Invertible Neural Projections
abstract
Unsupervised learning of syntactic structure is typically performed using generative models with discrete latent variables and multinomial parameters.In most cases, these models have not leveraged continuous word representations.In this work, we propose a novel generative model that jointly learns discrete syntactic structure and continuous word representations in an unsupervised fashion by cascading an invertible neural network with a structured generative prior.We show that the invertibility condition allows for efficient exact inference and marginal likelihood computation in our model so long as the prior is well-behaved.In experiments we instantiate our approach with both Markov and tree-structured priors, evaluating on two tasks: part-of-speech (POS) induction, and unsupervised dependency parsing without gold POS annotation.On the Penn Treebank, our Markov-structured model surpasses state-of-the-art results on POS induction.Similarly, we find that our tree-structured model achieves state-of-the-art performance on unsupervised dependency parsing for the difficult training condition where neither gold POS annotation nor punctuation-based constraints are available.
Junxian He, Graham Neubig, Taylor Berg-Kirkpatrick
EMNLP1
2017 Efficient Correlated Topic Modeling with Topic Embedding
abstract
Correlated topic modeling has been limited to small model and problem sizes due to their high computational cost and poor scaling. In this paper, we propose a new model which learns compact topic embeddings and captures topic correlations through the closeness between the topic vectors. Our method enables efficient inference in the low-dimensional embedding space, reducing previous cubic or quadratic time complexity to linear w.r.t the topic size. We further speedup variational inference with a fast sampler to exploit sparsity of topic occurrence. Extensive experiments show that our approach is capable of handling model and data scales which are several orders of magnitude larger than existing correlation results, without sacrificing modeling quality by providing competitive or superior performance in document classification and retrieval.
Junxian He, Zhiting Hu, Taylor Berg-Kirkpatrick, Eric P. Xing
KDD1