Jiacheng Ye

dblp:249/2706 · DBLP profile ↗
← Back
20ranked-venue papers
10as first author
16since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 19 · 9 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3Computer networks · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 G-LLaVA: Solving Geometric Problem with Multi-Modal Large Language Model
abstract
Large language models (LLMs) have shown remarkable proficiency in human-level reasoning and generation capabilities, which encourages extensive research on their application in mathematical problem solving. However, current work has been largely focused on text-based mathematical problems, with limited investigation in problems involving multi-modal geometric information. Addressing this gap, we aim to enable LLMs to solve geometric problems by understanding image input. We first identify the limitations of current Multimodal Large Language Models (MLLMs) in this area: they struggle to accurately comprehend basic geometric elements and their relationships. To address these challenges, we leverage the inherent attribute of logical structure compactness in geometric figures, utilizing text-only Large Language Models (LLMs) to curate a comprehensive multimodal geometry dataset. This dataset, named Geo170k, contains more than 170K geometric image-caption and question-answer pairs. Utilizing the Geo170k dataset, we introduce G-LLaVA, a model that demonstrates exceptional performance in solving geometric problems. It significantly outperforms GPT4-V on the geometry task of MathVista benchmark with only 7B parameters.
Jiahui Gao 0002, Renjie Pi, Jiacheng Ye, Wanjun Zhong, Yufei Wang 0005, Lanqing Hong, Jianhua Han, Hang Xu 0004, Zhenguo Li, Lingpeng Kong
ICLR4
2025 Scaling Diffusion Language Models via Adaptation from Autoregressive Models
abstract
Diffusion Language Models (DLMs) have emerged as a promising new paradigm for text generative modeling, potentially addressing limitations of autoregressive (AR) models. However, current DLMs have been studied at a smaller scale compared to their AR counterparts and lack fair comparison on language modeling benchmarks. Additionally, training diffusion models from scratch at scale remains challenging. Given the prevalence of open-source AR language models, we propose adapting these models to build text diffusion models. We demonstrate connections between AR and diffusion modeling objectives and introduce a simple continual pre-training approach for training diffusion models. Through systematic evaluation on language modeling, reasoning, and commonsense benchmarks, we show that we can convert AR models ranging from 127M to 7B parameters (GPT2 and LLaMA) into diffusion models DiffuGPT and DiffuLLaMA, using less than 200B tokens for training. Our experimental results reveal that these models outperform earlier DLMs and are competitive with their AR counterparts. We release a suite of DLMs (127M-355M-7B) capable of generating fluent text, performing in-context learning, filling in the middle without prompt re-ordering, and following instructions.
Shansan Gong, Shivam Agarwal, Yizhe Zhang 0002, Jiacheng Ye, Mukai Li, Chenxin An, Peilin Zhao, Wei Bi, Jiawei Han 0001, Hao Peng 0009, Lingpeng Kong
ICLR4
2025 Beyond Autoregression: Discrete Diffusion for Complex Reasoning and Planning
abstract
Autoregressive language models, despite their impressive capabilities, struggle with complex reasoning and long-term planning tasks. We introduce discrete diffusion models as a novel solution to these challenges. Through the lens of subgoal imbalance, we demonstrate how diffusion models effectively learn difficult subgoals that elude autoregressive approaches. We propose Multi-Granularity Diffusion Modeling (MGDM), which prioritizes subgoals based on difficulty during learning. On complex tasks like Countdown, Sudoku, and Boolean Satisfiability Problems, MGDM significantly outperforms autoregressive models without using search techniques. For instance, MGDM achieves 91.5\% and 100\% accuracy on Countdown and Sudoku, respectively, compared to 45.8\% and 20.7\% for autoregressive models. Our work highlights the potential of diffusion-based approaches in advancing AI capabilities for sophisticated language understanding and problem-solving tasks. All associated codes are available at \href{https://github.com/HKUNLP/diffusion-vs-ar}{https://github.com/HKUNLP/diffusion-vs-ar}.
Jiacheng Ye, Jiahui Gao 0002, Shansan Gong, Xin Jiang 0002, Zhenguo Li, Lingpeng Kong
ICLR1
2025 Implicit Search via Discrete Diffusion: A Study on Chess
abstract
In the post-AlphaGo era, there has been a renewed interest in search techniques such as Monte Carlo Tree Search (MCTS), particularly in their application to Large Language Models (LLMs). This renewed attention is driven by the recognition that current next-token prediction models often lack the ability for long-term planning. Is it possible to instill search-like abilities within the models to enhance their planning abilities without relying on explicit search? We propose DiffuSearch , a model that does \textit{implicit search} by looking into the future world via discrete diffusion modeling. We instantiate DiffuSearch on a classical board game, Chess, where explicit search is known to be essential. Through extensive controlled experiments, we show DiffuSearch outperforms both the searchless and explicit search-enhanced policies. Specifically, DiffuSearch outperforms the one-step policy by 19.2\% and the MCTS-enhanced policy by 14\% on action accuracy. Furthermore, DiffuSearch demonstrates a notable 30\% enhancement in puzzle-solving abilities compared to explicit search-based policies, along with a significant 540 Elo increase in game-playing strength assessment. These results indicate that implicit search via discrete diffusion is a viable alternative to explicit search over a one-step policy. All codes are publicly available at \href{https://github.com/HKUNLP/DiffuSearch}{https://github.com/HKUNLP/DiffuSearch}.
Jiacheng Ye, Jiahui Gao 0002, Zhiyong Wu 0003, Xin Jiang 0002, Zhenguo Li, Lingpeng Kong
ICLR1
2024 PRoLoRA: Partial Rotation Empowers More Parameter-Efficient LoRA
abstract
Sheng Wang, Boyang Xue, Jiacheng Ye, Jiyue Jiang, Liheng Chen, Lingpeng Kong, Chuan Wu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Boyang Xue, Jiacheng Ye, Jiyue Jiang, Lingpeng Kong
ACL (1)3
2024 Diffusion of Thought: Chain-of-Thought Reasoning in Diffusion Language Models
abstract
Recently, diffusion models have garnered significant interest in the field of text processing due to their many potential advantages compared to conventional autoregressive models. In this work, we propose Diffusion-of-Thought (DoT), a novel approach that integrates diffusion models with Chain-of-Thought, a well-established technique for improving the reasoning ability of autoregressive language models. In contrast to autoregressive language models that make decisions in a left-to-right, token-by-token manner, DoT allows reasoning steps to diffuse over time through a diffusion language model and offers greater flexibility in trading-off computation for reasoning performance. Our experimental results demonstrate the effectiveness of DoT in multi-digit multiplication, boolean logic, and grade school math problems. In addition to that, DoT showcases promising self-correction abilities and benefits from existing reasoning-enhancing techniques like self-consistency decoding. Our findings contribute to the understanding and development of reasoning with diffusion language models.
Jiacheng Ye, Shansan Gong, Jiahui Gao 0002, Xin Jiang 0002, Zhenguo Li, Wei Bi, Lingpeng Kong
NeurIPS1
2024 Leaky Cable Perimeter Intrusion Detection Based on Deep Reinforcement Learning
abstract
Many existing leaky cable intrusion monitoring methods and systems on the market are based on threshold detection, which can lead to misclassification and omission, limited adaptability to various installation environments, and low localization accuracy. In this article, we present a deep reinforcement learning-based approach for monitoring perimeter intrusion in leaky cables, leveraging the self-adaptation of the interaction between the reinforcement learning agent and its environment. This approach is a typical Internet of Things (IoT) application, integrating advanced machine learning techniques with IoT devices for enhanced security and adaptability. By converting threshold detection into an intelligent classification decision problem and devising a reward function tailored to the perimeter intrusion localization scenario, the learning agent can continually adapt to generate an optimal model suitable for dynamic IoT environments. To assess its feasibility and practicality, experiments in both simulated validation and field deployment environments are conducted. The experimental results indicate that the proposed model achieves 98.73% accuracy for the test set, 1-m accuracy in field localization, a false alarm rate of less than 1%, and robust adaptability to environmental noise, demonstrating its potential as an effective IoT solution for perimeter intrusion monitoring.
Jiacheng Ye, Junshi Lv, Gaoming Xu, Taijun Liu
IEEE Internet Things J.1
2023 Self-Adaptive In-Context Learning: An Information Compression Perspective for In-Context Example Selection and Ordering
abstract
Despite the impressive few-shot performance of in-context learning (ICL), it remains a common practice to randomly select examples to serve as the context.In this paper, we advocate self-adaptive in-context learning, a new principle for ICL, in which the self-adaption mechanism is introduced to help each input find an in-context example organization (i.e., selection and permutation) that can derive the correct output, thus maximizing performance.To validate the effectiveness of self-adaptive ICL, we propose a general select-then-rank framework and a set of novel selection and ranking algorithms.Upon extensive evaluation on eight different NLP datasets, our self-adaptive ICL method achieves a 40% relative improvement over the common practice setting.Further analysis reveals the great potential of selfadaptive ICL as a promising method to close the gap between ICL and finetuning.Our code will be released to facilitate future research.
Zhiyong Wu 0003, Yaoxiang Wang, Jiacheng Ye, Lingpeng Kong
ACL (1)3
2023 Generating Data for Symbolic Language with Large Language Models
abstract
While large language models (LLMs) bring not only performance but also complexity, recent work has started to turn LLMs into data generators rather than task inferencers, where another affordable task model is trained for efficient deployment and inference.However, such an approach has primarily been applied to natural language tasks, and has not yet been explored for symbolic language tasks with complex structured outputs (e.g., semantic parsing and code generation).In this paper, we propose SYMGEN which utilizes LLMs for generating various annotationexpensive symbolic language data.SYMGEN consists of an informative prompt to steer generation and an agreement-based verifier to improve data correctness.We conduct extensive experiments on six symbolic language tasks across various settings.Compared with the LLMs, we demonstrate the 1%-sized task model can achieve comparable or better performance, largely cutting inference and deployment costs.We also show that generated data with only a few human demonstrations can be as effective as over 10 times the amount of human-annotated data when training the task model, saving a considerable amount of annotation effort.SYMGEN takes a step toward data generation for annotation-expensive complex tasks, and we release the code at https://github.com/HKUNLP/SymGen.
Jiacheng Ye, Chengzu Li, Lingpeng Kong, Tao Yu 0009
EMNLP1
2023 Self-Guided Noise-Free Data Generation for Efficient Zero-Shot Learning
Jiahui Gao 0002, Renjie Pi, Hang Xu 0004, Jiacheng Ye, Zhiyong Wu 0003, Xiaodan Liang, Zhenguo Li, Lingpeng Kong
ICLR5
2023 Compositional Exemplars for In-context Learning
abstract
Large pretrained language models (LMs) have shown impressive In-Context Learning (ICL) ability, where the model learns to do an unseen task simply by conditioning on a prompt consisting of input-output examples as demonstration, without any parameter updates. The performance of ICL is highly dominated by the quality of the selected in-context examples. However, previous selection methods are mostly based on simple heuristics, leading to sub-optimal performance. In this work, we systematically formulate in-context example selection as a subset selection problem, and optimize it in an end-to-end fashion. We propose CEIL (Compositional Exemplars for In-context Learning), which is instantiated by Determinantal Point Processes (DPPs) to model the interaction between the given input and in-context examples, and optimized through carefully-designed contrastive learning to obtain preference from LMs. We validate CEIL on 12 classification and generation datasets from 7 distinct NLP tasks, including sentiment analysis, phraphrase detection, natural language inference, commonsense reasoning, open-domain question answering, code generation and semantic parsing. Extensive experiments demonstrate the effectiveness, transferability, compositionality of CEIL, shedding new lights on in-context leaning. Our code is released at https://github.com/HKUNLP/icl-ceil.
Jiacheng Ye, Zhiyong Wu 0003, Jiangtao Feng, Tao Yu 0009, Lingpeng Kong
ICML1
2022 ZeroGen: Efficient Zero-shot Learning via Dataset Generation
abstract
There is a growing interest in dataset generation recently due to the superior generative capacity of large pre-trained language models (PLMs).In this paper, we study a flexible and efficient zero-short learning method, ZEROGEN.Given a zero-shot task, we first generate a dataset from scratch using PLMs in an unsupervised manner.Then, we train a tiny task model (e.g., LSTM) under the supervision of the synthesized dataset.This approach allows highly efficient inference as the final task model only has orders of magnitude fewer parameters comparing to PLMs (e.g., GPT2-XL).Apart from being annotation-free and efficient, we argue that ZEROGEN can also provide useful insights from the perspective of datafree model-agnostic knowledge distillation, and unreferenced text generation evaluation.Experiments and analysis on different NLP tasks, namely, text classification, question answering, and natural language inference, show the effectiveness of ZEROGEN.
Jiacheng Ye, Jiahui Gao 0002, Qintong Li, Hang Xu 0004, Jiangtao Feng, Zhiyong Wu 0003, Tao Yu 0009, Lingpeng Kong
EMNLP1
2022 Uncertainty-Aware Sequence Labeling
abstract
Conditional random fields (CRFs) have been widely used for sequence labeling tasks in the field of natural language processing. However, how to model both local and global dependencies among labels is not well solved yet. In this study, we introduce a novel two-stage label decoding method to better model the short- and long-term label dependencies, while being much more computationally efficient with the use of graphics processing units (GPUs). A base model is first used to propose draft labels, and then a novel two-stream self-attention model makes refinements on these draft predictions based on long-range label dependencies. Besides, in order to mitigate the side effects of incorrect draft labels, Bayesian neural networks are used to indicate the labels with high probabilities of being wrong, which helps to mitigate the error propagation. Not only can our method model sentence-level label dependencies, but it is also easily extended to document-level sequence labeling by querying and storing a key-value memory matrix with label co-occurrence relationships. The experimental results on both sentence-level and document-level sequence labeling benchmarks show that the proposed method outperforms existing label decoding methods while taking advantage of parallel computations on GPUs.
Jiacheng Ye, Xiaoqing Zheng, Tao Gui, Qi Zhang 0001
IEEE ACM Trans. Audio Speech Lang. Process.1
2021 One2Set: Generating Diverse Keyphrases as a Set
abstract
Jiacheng Ye, Tao Gui, Yichao Luo, Yige Xu, Qi Zhang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Jiacheng Ye, Tao Gui, Yichao Luo, Yige Xu 0001, Qi Zhang 0001
ACL/IJCNLP (1)1
2021 Heterogeneous Graph Neural Networks for Keyphrase Generation
abstract
The encoder-decoder framework achieves stateof-the-art results in keyphrase generation (KG) tasks by predicting both present keyphrases that appear in the source document and absent keyphrases that do not.However, relying solely on the source document can result in generating uncontrollable and inaccurate absent keyphrases.To address these problems, we propose a novel graph-based method that can capture explicit knowledge from related references.Our model first retrieves some document-keyphrases pairs similar to the source document from a pre-defined index as references.Then a heterogeneous graph is constructed to capture relationships of different granularities between the source document and its references.To guide the decoding process, a hierarchical attention and copy mechanism is introduced, which directly copies appropriate words from both the source document and its references based on their relevance and significance.The experimental results on multiple KG benchmarks show that the proposed model achieves significant improvements against other baseline models, especially with regard to the absent keyphrase prediction.
Jiacheng Ye, Ruijian Cai, Tao Gui, Qi Zhang 0001
EMNLP (1)1
2021 A Faster Maximum Cardinality Matching Algorithm with Applications in Machine Learning
abstract
Maximum cardinality bipartite matching is an important graph optimization problem with several applications. For instance, maximum cardinality matching in a $\delta$-disc graph can be used in the computation of the bottleneck matching as well as the $\infty$-Wasserstein and the Lévy-Prokhorov distances between probability distributions. For any point sets $A, B \subset \mathbb{R}^2$, the $\delta$-disc graph is a bipartite graph formed by connecting every pair of points $(a,b) \in A\times B$ by an edge if the Euclidean distance between them is at most $\delta$. Using the classical Hopcroft-Karp algorithm, a maximum-cardinality matching on any $\delta$-disc graph can be found in $\tilde{O}(n^{3/2})$ time.~\footnote{We use $\tilde{O}(\cdot)$ to suppress poly-logarithmic terms in the complexity.} In this paper, we present a simplification of a recent algorithm (Lahn and Raghvendra, JoCG 2021) for the maximum cardinality matching problem and describe how a maximum cardinality matching in a $\delta$-disc graph can be computed asymptotically faster than $O(n^{3/2})$ time for any moderately dense point set. As applications, we show that if $A$ and $B$ are point sets drawn uniformly at random from a unit square, an exact bottleneck matching can be computed in $\tilde{O}(n^{4/3})$ time. On the other hand, experiments suggest that the Hopcroft-Karp algorithm seems to take roughly $\Theta (n^{3/2})$ time for this case. This translates to substantial improvements in execution time for larger inputs.
Nathaniel Lahn, Sharath Raghvendra, Jiacheng Ye
NeurIPS3
2020 Constructing Multiple Tasks for Augmentation: Improving Neural Image Classification with K-Means Features
abstract
Multi-task learning (MTL) has received considerable attention, and numerous deep learning applications benefit from MTL with multiple objectives. However, constructing multiple related tasks is difficult, and sometimes only a single task is available for training in a dataset. To tackle this problem, we explored the idea of using unsupervised clustering to construct a variety of auxiliary tasks from unlabeled data or existing labeled data. We found that some of these newly constructed tasks could exhibit semantic meanings corresponding to certain human-specific attributes, but some were non-ideal. In order to effectively reduce the impact of non-ideal auxiliary tasks on the main task, we further proposed a novel meta-learning-based multi-task learning approach, which trained the shared hidden layers on auxiliary tasks, while the meta-optimization objective was to minimize the loss on the main task, ensuring that the optimizing direction led to an improvement on the main task. Experimental results across five image datasets demonstrated that the proposed method significantly outperformed existing single task learning, semi-supervised learning, and some data augmentation methods, including an improvement of more than 9% on the Omniglot dataset.
Tao Gui, Lizhi Qing, Qi Zhang 0001, Jiacheng Ye, Hang Yan 0001, Zichu Fei, Xuanjing Huang 0001
AAAI4
2020 Uncertainty-Aware Label Refinement for Sequence Labeling
abstract
Conditional random fields (CRF) for label decoding has become ubiquitous in sequence labeling tasks.However, the local label dependencies and inefficient Viterbi decoding have always been a problem to be solved.In this work, we introduce a novel two-stage label decoding framework to model long-term label dependencies, while being much more computationally efficient.A base model first predicts draft labels, and then a novel twostream self-attention model makes refinements on these draft predictions based on longrange label dependencies, which can achieve parallel decoding for a faster prediction.In addition, in order to mitigate the side effects of incorrect draft labels, Bayesian neural networks are used to indicate the labels with a high probability of being wrong, which can greatly assist in preventing error propagation.The experimental results on three sequence labeling benchmarks demonstrated that the proposed method not only outperformed the CRF-based methods but also greatly accelerated the inference process.* Both authors contributed equally.
Tao Gui, Jiacheng Ye, Qi Zhang 0001, Zhengyan Li, Zichu Fei, Yeyun Gong, Xuanjing Huang 0001
EMNLP (1)2
2020 Leveraging Document-Level Label Consistency for Named Entity Recognition
abstract
Document-level label consistency is an effective indicator that different occurrences of a particular token sequence are very likely to have the same entity types. Previous work focused on better context representations and used the CRF for label decoding. However, CRF-based methods are inadequate for modeling document-level label consistency. This work introduces a novel two-stage label refinement approach to handle document-level label consistency, where a key-value memory network is first used to record draft labels predicted by the base model, and then a multi-channel Transformer makes refinements on these draft predictions based on the explicit co-occurrence relationship derived from the memory network. In addition, in order to mitigate the side effects of incorrect draft labels, Bayesian neural networks are used to indicate the labels with a high probability of being wrong, which can greatly assist in preventing the incorrect refinement of correct draft labels. The experimental results on three named entity recognition benchmarks demonstrated that the proposed method significantly outperformed the state-of-the-art methods.
Tao Gui, Jiacheng Ye, Qi Zhang 0001, Yaqian Zhou 0001, Yeyun Gong, Xuanjing Huang 0001
IJCAI2
2019 Re-ID Driven Localization Refinement for Person Search
abstract
Person search aims at localizing and identifying a query person from a gallery of uncropped scene images. Different from person re-identification (re-ID), its performance also depends on the localization accuracy of a pedestrian detector. The state-of-the-art methods train the detector individually, and the detected bounding boxes may be sub-optimal for the following re-ID task. To alleviate this issue, we propose a re-ID driven localization refinement framework for providing the refined detection boxes for person search. Specifically, we develop a differentiable ROI transform layer to effectively transform the bounding boxes from the original images. Thus, the box coordinates can be supervised by the re-ID training other than the original detection task. With this supervision, the detector can generate more reliable bounding boxes, and the downstream re-ID model can produce more discriminative embeddings based on the refined person localizations. Extensive experimental results on the widely used benchmarks demonstrate that our proposed method performs favorably against the state-of-the-art person search methods.
Chuchu Han, Jiacheng Ye, Yunshan Zhong, Xin Tan 0002, Chi Zhang 0026, Changxin Gao, Nong Sang
ICCV2