EDBT 2026 Demo / reviewers in the wild / expert
Longhui Zhang
dblp:147/9111
· DBLP profile ↗
16ranked-venue papers
8as first author
11since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 4 first-author · 8 since 2021Databases, data management, data science and information retrieval · 8 · 5 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Noah: Tile-Level Interval Analysis for NPU Performance Modeling
Mengqi Fu, Rulan Yang, Mengrui Zhang, Qingyu Song 0002, Yuanxun Kang, Longhui Zhang, Qiao Xiang |
IWQoS | 7 |
| 2026 | Progressive Adaptation of Large Language Models for Multilingual Text RankingabstractDespite increasing research attention to text ranking, most studies focus on monolingual scenarios, with a particular emphasis on English-language contexts. This narrow focus limits the applicability of ranking models in cross-lingual contexts, such as ranking Chinese documents based on English queries. Recent advances in large language models (LLMs) have significantly reduced inter-language barriers through pre-training on extensive multilingual corpora, thus facilitating the study of multilingual text ranking (MTR). In this work, we explore the potential of LLMs in MTR tasks. Specifically, we first introduce an MTR benchmark encompassing both monolingual and cross-lingual scenarios. Then, we propose a two-stage training pipeline to alleviate the misalignment between LLMs and text ranking. Lastly, we adapt this training pipeline to multilingual scenarios from the perspective of training data and methods. Our experiments on the MTR benchmark demonstrate that the proposed multilingual two-stage training pipeline significantly improves LLM ranking performance in both monolingual and cross-lingual scenarios, particularly in out-domain settings. We complement these findings with a thorough analysis to deepen the understanding of our approach. Longhui Zhang, Yanzhao Zhang, Dingkun Long, Pengjun Xie, Meishan Zhang, Jing Li 0034, Min Zhang 0005 |
ACM Trans. Inf. Syst. | 1 |
| 2025 | Speed Up Your Code: Progressive Code Acceleration Through Bidirectional Tree EditingabstractLonghui Zhang, Jiahao Wang, Meishan Zhang, GaoXiong Cao, Ensheng Shi, Mayuchi Mayuchi, Jun Yu, Honghai Liu, Jing Li, Min Zhang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Longhui Zhang, Meishan Zhang, GaoXiong Cao, Ensheng Shi, Mayuchi Mayuchi, Jun Yu 0002, Honghai Liu 0001, Jing Li 0034, Min Zhang 0005 |
ACL (1) | 1 |
| 2025 | Function-to-Style Guidance of LLMs for Code TranslationabstractLarge language models (LLMs) have made significant strides in code translation tasks. However, ensuring both the correctness and readability of translated code remains a challenge, limiting their effective adoption in real-world software development. In this work, we propose F2STrans, a function-to-style guiding paradigm designed to progressively improve the performance of LLMs in code translation. Our approach comprises two key stages: (1) Functional learning, which optimizes translation correctness using high-quality source-target code pairs mined from online programming platforms, and (2) Style learning, which improves translation readability by incorporating both positive and negative style examples. Additionally, we introduce a novel code translation benchmark that includes up-to-date source code, extensive test cases, and manually annotated ground-truth translations, enabling comprehensive functional and stylistic evaluations. Experiments on both our new benchmark and existing datasets demonstrate that our approach significantly improves code translation performance. Notably, our approach enables Qwen-1.5B to outperform prompt-enhanced Qwen-32B and GPT-4 on average across 20 diverse code translation scenarios. Longhui Zhang, Bin Wang 0004, Hao Yang 0007, Meishan Zhang, Yu Li 0007, Jing Li 0034, Jun Yu 0002, Min Zhang 0005 |
ICML | 1 |
| 2024 | Chinese Sequence Labeling with Semi-Supervised Boundary-Aware Language Model Pre-trainingabstractChinese sequence labeling tasks are sensitive to word boundaries. Although pretrained language models (PLM) have achieved considerable success in these tasks, current PLMs rarely consider boundary information explicitly. An exception to this is BABERT, which incorporates unsupervised statistical boundary information into Chinese BERT’s pre-training objectives. Building upon this approach, we input supervised high-quality boundary information to enhance BABERT’s learning, developing a semi-supervised boundary-aware PLM. To assess PLMs’ ability to encode boundaries, we introduce a novel “Boundary Information Metric” that is both simple and effective. This metric allows comparison of different PLMs without task-specific fine-tuning. Experimental results on Chinese sequence labeling datasets demonstrate that the improved BABERT version outperforms the vanilla version, not only in these tasks but also in broader Chinese natural language understanding tasks. Additionally, our proposed metric offers a convenient and accurate means of evaluating PLMs’ boundary awareness. Longhui Zhang, Dingkun Long, Meishan Zhang, Yanzhao Zhang, Pengjun Xie, Min Zhang 0005 |
LREC/COLING | 1 |
| 2022 | Mining High-Value Patents Leveraging Massive Patent Data
Ruixiang Luo, Lijuan Weng, Junxiang Ji, Longbiao Chen, Longhui Zhang |
ICA3PP | 5 |
| 2022 | A Simple but Effective Bidirectional Framework for Relational Triple ExtractionabstractTagging based relational triple extraction methods are attracting growing research attention recently. However, most of these methods take a unidirectional extraction framework that first extracts all subjects and then extracts objects and relations simultaneously based on the subjects extracted. This framework has an obvious deficiency that it is too sensitive to the extraction results of subjects. To overcome this deficiency, we propose a bidirectional extraction framework based method that extracts triples based on the entity pairs extracted from two complementary directions. Concretely, we first extract all possible subject-object pairs from two paralleled directions. These two extraction directions are connected by a shared encoder component, thus the extraction features from one direction can flow to another direction and vice versa. By this way, the extractions of two directions can boost and complement each other. Next, we assign all possible relations for each entity pair by a biaffine model. During training, we observe that the share structure will lead to a convergence rate inconsistency issue which is harmful to performance. So we propose a share-aware learning mechanism to address it. We evaluate the proposed model on multiple benchmark datasets. Extensive experimental results show that the proposed model is very effective and it achieves state-of-the-art results on all of these datasets. Moreover, experiments show that both the proposed bidirectional extraction framework and the share-aware learning mechanism have good adaptability and can be used to improve the performance of other tagging based methods. The source code of our work is available at: https://github.com/neukg/BiRTE. Feiliang Ren, Longhui Zhang, Shujuan Yin, Shilei Liu, Bochao Li |
WSDM | 2 |
| 2021 | A Conditional Cascade Model for Relational Triple ExtractionabstractTagging based methods are one of the mainstream methods in relational triple extraction. However, most of them suffer from the class imbalance issue greatly. Here we propose a novel tagging based model that addresses this issue from following two aspects. First, at the model level, we propose a three-step extraction framework that can reduce the total number of samples greatly, which implicitly decreases the severity of the mentioned issue. Second, at the intra-model level, we propose a confidence threshold based cross entropy loss that can directly neglect some samples in the major classes. We evaluate the proposed model on NYT and WebNLG. Extensive experiments show that it can address the mentioned issue effectively and achieves state-of-the-art results on both datasets. The source code of our model is available at: https://github.com/neukg/ConCasRTE. Feiliang Ren, Longhui Zhang, Shujuan Yin, Shilei Liu, Bochao Li |
CIKM | 2 |
| 2021 | A Three-Stage Learning Framework for Low-Resource Knowledge-Grounded Dialogue GenerationabstractNeural conversation models have shown great potentials towards generating fluent and informative responses by introducing external background knowledge.Nevertheless, it is laborious to construct such knowledge-grounded dialogues, and existing models usually perform poorly when transfer to new domains with limited training samples.Therefore, building a knowledge-grounded dialogue system under the low-resource setting is a still crucial issue.In this paper, we propose a novel threestage learning framework based on weakly supervised learning which benefits from large scale ungrounded dialogues and unstructured knowledge base.To better cooperate with this framework, we devise a variant of Transformer with decoupled decoder which facilitates the disentangled learning of response generation and knowledge incorporation.Evaluation results on two benchmarks indicate that our approach can outperform other state-of-the-art methods with less training data, and even in zero-resource scenario, our approach still performs well. Shilei Liu, Bochao Li, Feiliang Ren, Longhui Zhang, Shujuan Yin |
EMNLP (1) | 5 |
| 2021 | A Novel Global Feature-Oriented Relational Triple Extraction Model based on Table FillingabstractTable filling based relational triple extraction methods are attracting growing research interests due to their promising performance and their abilities on extracting triples from complex sentences.However, this kind of methods are far from their full potential because most of them only focus on using local features but ignore the global associations of relations and of token pairs, which increases the possibility of overlooking some important information during triple extraction.To overcome this deficiency, we propose a global feature-oriented triple extraction model that makes full use of the mentioned two kinds of global associations.Specifically, we first generate a table feature for each relation.Then two kinds of global associations are mined from the generated table features.Next, the mined global associations are integrated into the table feature of each relation.This "generate-mine-integrate" process is performed multiple times so that the table feature of each relation is refined step by step.Finally, each relation's table is filled based on its refined table feature, and all triples linked to this relation are extracted based on its filled table.We evaluate the proposed model on three benchmark datasets.Experimental results show our model is effective and it achieves state-of-the-art results on all of these datasets.The source code of our work is available at: https://github.com/neukg/GRTE. Feiliang Ren, Longhui Zhang, Shujuan Yin, Shilei Liu, Bochao Li, Yaduo Liu |
EMNLP (1) | 2 |
| 2021 | An Effective System for Multi-format Information Extraction
Yaduo Liu, Longhui Zhang, Shujuan Yin, Feiliang Ren |
NLPCC (2) | 2 |
| 2018 | PatSearch: an integrated framework for patentability retrieval
Longhui Zhang, Li Zheng 0001, Lei Li 0001, Chao Shen 0005, Tao Li 0001 |
Knowl. Inf. Syst. | 1 |
| 2015 | PatentCom: A Comparative View of Patent Document RetrievalabstractPatent document retrieval, as a recall-orientated search task, does not allow missing relevant patent documents due to the great commercial value of patents and significant costs of processing a patent application or patent infringement case. Thus, it is important to retrieve all possible relevant documents rather than only a small subset of patents from the top ranked results. However, patents are often lengthy and rich in technical terms, and it often requires enormous human efforts to compare a given document with retrieved results. In this paper, we formulate the problem of comparing patent documents as a comparative summarization problem, and explore automatic strategies that generate comparative summaries to assist patent analysts in quickly reviewing any given patent document pairs. To this end, we present a novel approach, named PatentCom, which first extracts discriminative terms from each patent document, and then connects the dots on a term co-occurrence graph. In this way, we are able to comprehensively extract the gists of the two patent documents being compared, and meanwhile highlight their relationship in terms of commonalities and differences. Extensive quantitative analysis and case studies on real world patent documents demonstrate the effectiveness of our proposed approach. Longhui Zhang, Lei Li 0001, Chao Shen 0005, Tao Li 0001 |
SDM | 1 |
| 2014 | iMiner: Mining Inventory Data for Intelligent ManagementabstractInventory management refers to tracing inventory levels, orders and sales of a retailing business. In the current retailing market, a tremendous amount of data regarding stocked goods (items) in an inventory will be generated everyday. Due to the increasing volume of transaction data and the correlated relations of items, it is often a non-trivial task to efficiently and effectively manage stocked goods. In this demo, we present an intelligent system, called iMiner, to ease the management of enormous inventory data. We utilize distributed computing resources to process the huge volume of inventory data, and incorporate the latest advances of data mining technologies into the system to perform the tasks of inventory management, e.g., forecasting inventory, detecting abnormal items, and analyzing inventory aging. Since 2014, iMiner has been deployed as the major inventory management platform of ChangHong Electric Co., Ltd, one of the world's largest TV selling companies in China. Lei Li 0001, Chao Shen 0005, Li Zheng 0001, Yexi Jiang, Hongtai Li, Longhui Zhang, Chunqiu Zeng, Tao Li 0001 |
CIKM | 8 |
| 2014 | PatentDom: Analyzing Patent Relationships on Multi-View Patent GraphsabstractThe fast growth of technologies has driven the advancement of our society. It is often necessary to quickly grasp the linkage between different technologies in order to better understand the technical trend. The availability of huge volumes of granted patent documents provides a reasonable basis for analyzing the relationships between technologies. In this paper, we propose a unified framework, named PatentDom, to identify important patents related to key techniques from a large number of patent documents. The framework integrates different types of patent information, including patent content, citations of patents, and temporal relations, and provides a concise yet comprehensive technology summary. The identified key patents enable a variety of patent-related analytical applications, e.g., outlining the technology evolution of a particular domain, tracing a given technique to prior technologies, and mining the technical connection of two given patent documents. Empirical analysis and extensive case studies on a collection of US patent documents demonstrate the efficacy of our proposed framework. Longhui Zhang, Lei Li 0001, Tao Li 0001, Dingding Wang 0001 |
CIKM | 1 |
| 2014 | PatentLine: analyzing technology evolution on multi-view patent graphsabstractThe fast growth of technologies has driven the advancement of our society. It is often necessary to quickly grab the evolution of technologies in order to better understand the technology trend. The availability of huge volumes of granted patent documents provides a reasonable basis for analyzing technology evolution. In this paper, we propose a unified framework, named PatentLine, to generate a technology evolution tree for a given topic or a classification code related to granted patents. The framework integrates different types of patent information, including patent content, citations of patents, temporal relations, etc., and provides a concise yet comprehensive evolution summary. The generated summary enables a variety of patent-related analyses such as identifying relevant prior art and detecting technology gap. A case study on a collection of US patents demonstrates the efficacy of our proposed framework. Longhui Zhang, Lei Li 0001, Tao Li 0001, Qi Zhang 0001 |
SIGIR | 1 |