VLDB 2026 Research / reviewers in the wild / expert
Dejiao Zhang
dblp:131/6876
· DBLP profile ↗
16ranked-venue papers
6as first author
12since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 5 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Reasoning in Token Economies: Budget-Aware Evaluation of LLM Reasoning StrategiesabstractA diverse array of reasoning strategies has been proposed to elicit the capabilities of large language models.However, in this paper, we point out that traditional evaluations which focus solely on performance metrics miss a key factor: the increased effectiveness due to additional compute.By overlooking this aspect, a skewed view of strategy efficiency is often presented.This paper introduces a framework that incorporates the compute budget into the evaluation, providing a more informative comparison that takes into account both performance metrics and computational cost.In this budgetaware perspective, we find that complex reasoning strategies often don't surpass simpler baselines purely due to algorithmic ingenuity, but rather due to the larger computational resources allocated.When we provide a simple baseline like chain-of-thought self-consistency with comparable compute resources, it frequently outperforms reasoning strategies proposed in the literature.In this scale-aware perspective, we find that unlike self-consistency, certain strategies such as multi-agent debate or Reflexion can become worse if more compute budget is utilized. Siddhartha Jain 0001, Dejiao Zhang, Baishakhi Ray, Ben Athiwaratkun |
EMNLP | 3 |
| 2024 | Code Representation Learning at ScaleabstractRecent studies have shown that code language model at scale demonstrate significant performance gains on downstream tasks, i.e., code generation. However, most of the existing works on code representation learning train models at a hundred million parameter scale using very limited pretraining corpora. In this work, we fuel code representation learning with a vast amount of code data via a two-stage pretraining scheme. We first train the encoders via a mix that leverages both randomness in masking language modeling and implicit structure and semantic aspects of programming language. We then enhance the representations via contrastive learning with hard negative and hard positive constructed in an unsupervised manner. We establish an off-the-shelf encoder model that persistently outperforms the existing models on a wide variety of downstream tasks by large margins. To comprehend the factors contributing to successful code representation learning, we conduct detailed ablations and share our findings on (i) a customized and effective token-level denoising scheme for source code; (ii) the importance of hard negatives and hard positives; (iii) how the proposed bimodal contrastive learning boost the cross-lingual semantic search performance; and (iv) how the pretraining schemes decide the downstream task performance scales with the model size. Dejiao Zhang, Wasi Uddin Ahmad, Hantian Ding, Ramesh Nallapati, Dan Roth 0001, Xiaofei Ma 0001, Bing Xiang |
ICLR | 1 |
| 2024 | Repoformer: Selective Retrieval for Repository-Level Code CompletionabstractRecent advances in retrieval-augmented generation (RAG) have initiated a new era in repository-level code completion. However, the invariable use of retrieval in existing methods exposes issues in both efficiency and robustness, with a large proportion of the retrieved contexts proving unhelpful or harmful to code language models (code LMs). In this paper, we propose a selective RAG framework to avoid retrieval when unnecessary. To power this framework, we design a self-supervised learning approach to enable a code LM to accurately self-evaluate whether retrieval can improve its output quality and robustly leverage the potentially noisy retrieved contexts. Using this LM as both the selective RAG policy and the generation model, our framework achieves state-of-the-art repository-level code completion performance on diverse benchmarks including RepoEval, CrossCodeEval, and CrossCodeLongEval, a new long-form code completion benchmark. Meanwhile, our analyses show that selectively retrieving brings as much as 70% inference speedup in the online serving setting without harming the performance. We further demonstrate that our framework is able to accommodate different generation models, retrievers, and programming languages. These advancements position our framework as an important step towards more accurate and efficient repository-level code completion. Di Wu 0054, Wasi Uddin Ahmad, Dejiao Zhang, Murali Krishna Ramanathan, Xiaofei Ma 0001 |
ICML | 3 |
| 2023 | Multitask Pretraining with Structured Knowledge for Text-to-SQL GenerationabstractRobert Giaquinto, Dejiao Zhang, Benjamin Kleiner, Yang Li, Ming Tan, Parminder Bhatia, Ramesh Nallapati, Xiaofei Ma. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Robert Giaquinto, Dejiao Zhang, Benjamin Kleiner, Parminder Bhatia, Ramesh Nallapati, Xiaofei Ma 0001 |
ACL (1) | 2 |
| 2023 | ContraCLM: Contrastive Learning For Causal Language ModelabstractNihal Jain, Dejiao Zhang, Wasi Uddin Ahmad, Zijian Wang, Feng Nan, Xiaopeng Li, Ming Tan, Ramesh Nallapati, Baishakhi Ray, Parminder Bhatia, Xiaofei Ma, Bing Xiang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Nihal Jain, Dejiao Zhang, Wasi Uddin Ahmad, Zijian Wang 0002, Feng Nan, Xiaopeng Li 0002, Ramesh Nallapati, Baishakhi Ray, Parminder Bhatia, Xiaofei Ma 0001, Bing Xiang |
ACL (1) | 2 |
| 2022 | Learning Dialogue Representations from Consecutive UtterancesabstractZhihan Zhou, Dejiao Zhang, Wei Xiao, Nicholas Dingwall, Xiaofei Ma, Andrew Arnold, Bing Xiang. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Zhihan Zhou 0001, Dejiao Zhang, Wei Xiao 0001, Nicholas Dingwall, Xiaofei Ma 0001, Andrew O. Arnold, Bing Xiang |
NAACL-HLT | 2 |
| 2022 | Lifelong Pretraining: Continually Adapting Language Models to Emerging CorporaabstractXisen Jin, Dejiao Zhang, Henghui Zhu, Wei Xiao, Shang-Wen Li, Xiaokai Wei, Andrew Arnold, Xiang Ren. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Xisen Jin, Dejiao Zhang, Henghui Zhu, Wei Xiao 0001, Shang-Wen Li 0001, Xiaokai Wei, Andrew O. Arnold, Xiang Ren 0001 |
NAACL-HLT | 2 |
| 2021 | Answering Ambiguous Questions through Generative Evidence Fusion and Round-Trip PredictionabstractYifan Gao, Henghui Zhu, Patrick Ng, Cicero Nogueira dos Santos, Zhiguo Wang, Feng Nan, Dejiao Zhang, Ramesh Nallapati, Andrew O. Arnold, Bing Xiang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Yifan Gao 0001, Henghui Zhu, Patrick Ng, Cícero Nogueira dos Santos, Zhiguo Wang 0006, Feng Nan, Dejiao Zhang, Ramesh Nallapati, Andrew O. Arnold, Bing Xiang |
ACL/IJCNLP (1) | 7 |
| 2021 | Improving Factual Consistency of Abstractive Summarization via Question AnsweringabstractFeng Nan, Cicero Nogueira dos Santos, Henghui Zhu, Patrick Ng, Kathleen McKeown, Ramesh Nallapati, Dejiao Zhang, Zhiguo Wang, Andrew O. Arnold, Bing Xiang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Feng Nan, Cícero Nogueira dos Santos, Henghui Zhu, Patrick Ng, Kathy McKeown, Ramesh Nallapati, Dejiao Zhang, Zhiguo Wang 0006, Andrew O. Arnold, Bing Xiang |
ACL/IJCNLP (1) | 7 |
| 2021 | Entity-level Factual Consistency of Abstractive Text SummarizationabstractFeng Nan, Ramesh Nallapati, Zhiguo Wang, Cicero Nogueira dos Santos, Henghui Zhu, Dejiao Zhang, Kathleen McKeown, Bing Xiang. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021. Feng Nan, Ramesh Nallapati, Zhiguo Wang 0006, Cícero Nogueira dos Santos, Henghui Zhu, Dejiao Zhang, Kathy McKeown, Bing Xiang |
EACL | 6 |
| 2021 | Pairwise Supervised Contrastive Learning of Sentence RepresentationsabstractMany recent successes in sentence representation learning have been achieved by simply fine-tuning on the Natural Language Inference (NLI) datasets with triplet loss or siamese loss.Nevertheless, they share a common weakness: sentences in a contradiction pair are not necessarily from different semantic categories.Therefore, optimizing the semantic entailment and contradiction reasoning objective alone is inadequate to capture the high-level semantic structure.The drawback is compounded by the fact that the vanilla siamese or triplet losses only learn from individual sentence pairs or triplets, which often suffer from bad local optima.In this paper, we propose PairSupCon, an instance discrimination based approach aiming to bridge semantic entailment and contradiction understanding with high-level categorical concept encoding.We evaluate PairSupCon on various downstream tasks that involve understanding sentence semantics at different granularities.We outperform the previous state-of-theart method with 10%-13% averaged improvement on eight clustering tasks, and 5%-6% averaged improvement on seven semantic textual similarity (STS) tasks. Dejiao Zhang, Shang-Wen Li 0001, Wei Xiao 0001, Henghui Zhu, Ramesh Nallapati, Andrew O. Arnold, Bing Xiang |
EMNLP (1) | 1 |
| 2021 | Supporting Clustering with Contrastive LearningabstractDejiao Zhang, Feng Nan, Xiaokai Wei, Shang-Wen Li, Henghui Zhu, Kathleen McKeown, Ramesh Nallapati, Andrew O. Arnold, Bing Xiang. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Dejiao Zhang, Feng Nan, Xiaokai Wei, Shang-Wen Li 0001, Henghui Zhu, Kathy McKeown, Ramesh Nallapati, Andrew O. Arnold, Bing Xiang |
NAACL-HLT | 1 |
| 2018 | Learning to Share: simultaneous parameter tying and Sparsification in Deep Learning
Dejiao Zhang, Haozhu Wang, Mário A. T. Figueiredo, Laura Balzano |
ICLR (Poster) | 1 |
| 2017 | Matched subspace detection using compressively sampled dataabstractWe consider the problem of detecting whether a high dimensional signal lies in a given low dimensional subspace using only a few compressive measurements of it. By leveraging modern random matrix theory, we show that, even when we are short on information, a reliable detector can be constructed via a properly defined measure of energy of the signal outside the subspace. Our results extend those in [1] to a more general sampling framework. Moreover, the test statistic we define is much simpler than that required by [1], and it results in more efficient computation, which is crucial for high-dimensional data processing. Dejiao Zhang, Laura Balzano |
ICASSP | 1 |
| 2016 | Global Convergence of a Grassmannian Gradient Descent Algorithm for Subspace EstimationabstractIt has been observed in a variety of contexts that gradient descent methods have great success in solving low-rank matrix factorization problems, despite the relevant problem formulation being non-convex. We tackle a particular instance of this scenario, where we seek the d-dimensional subspace spanned by a streaming data matrix. We apply the natural first order incremental gradient descent method, constraining the gradient method to the Grassmannian. In this paper, we propose an adaptive step size scheme that is greedy for the noiseless case, that maximizes the improvement of our metric of convergence at each data index t, and yields an expected improvement for the noisy case. We show that, with noise-free data, this method converges from any random initialization to the global minimum of the problem. For noisy data, we provide the expected convergence rate of the proposed algorithm per iteration. Dejiao Zhang, Laura Balzano |
AISTATS | 1 |
| 2014 | Iterative Grassmannian optimization for robust image alignment
Jun He 0006, Dejiao Zhang, Laura Balzano |
Image Vis. Comput. | 2 |