VLDB 2026 Research / reviewers in the wild / expert
Haojie Lu
dblp:314/3879
· DBLP profile ↗
9ranked-venue papers
2as first author
9since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | REFO: Reinforced Evolutionary Faithfulness Optimization for Large Language ModelsabstractDespite its success in enriching LLMs with external knowledge, RAG remains plagued by faithfulness hallucinations, where generated text contradicts the retrieved source information. Previous research on faithfulness hallucination in LLMs is frequently hindered by prohibitive manual annotation costs and a dependency on static datasets, which caps their performance and adaptability. Furthermore, these models lack a clear training mechanism to explicitly promote contextual focus. In this work, we propose a novel iterative self-evolution framework to enhance model faithfulness. This framework autonomously generates high-quality data and leverages it for the continuous self-optimization of the model, leading to significant improvements in faithfulness. Our experimental analysis reveals that improving model faithfulness encourages a closer alignment of the attention distribution with the given context. Based on this finding, we design an attention-based loss function to further promote this process. Experimental results show that our model achieves state-of-the-art faithfulness on a range of context-based question-answering datasets, marking a significant advancement over previous approaches. Xiaqiang Tang, Keyu Hu, Haojie Lu, Sihong Xie |
AAAI | 4 |
| 2026 | UTMD: An Unsupervised Transformer-Based Misbehavior Detection Method in IoV
Zhao Tian 0005, Haojie Lu, Wei She, Wei Liu 0043 |
ICIC (8) | 4 |
| 2026 | A Dynamic Reputation Framework Based on Deep Learning and Hybrid Blockchain for the Internet of Vehicles
Zhao Tian 0005, Haojie Lu, Wei Liu 0043, Wei She |
ICIC (2) | 4 |
| 2025 | LIO-DPC: Accurate and Fast LiDAR-Inertial Odometry with Dynamic Pose ChainabstractLiDAR-inertial odometry is widely used in robotics navigation, autonomous driving, and drone operation to provide precise, low-latency motion estimation. Filter-based methods are fast but suffer from significant cumulative errors. Graph optimization methods reduce cumulative errors through loop closure detection but are computationally expensive. In this work, we propose LIO-DPC, a framework that combines the benefits of the filter-based approach and graph-based approach. First, we propose a dynamic pose chain optimization method. It generates an initial pose chain using the fast filter. This is followed by applying computationally efficient local graph optimization to a set of local pose chains to generate refined relative poses, which are then used to update the motion estimation. Second, we propose a loop sparsification approach to select representative loops that are both temporally and spatially proximate, to reduce the computational complexity in graph optimization and minimize loop errors. Extensive experiments demonstrate that LIO-DPC achieves real-time performance and outperforms state-of-the-art methods in accuracy. Yuexin Mu, Ao Ren, Duo Liu 0002, Zihao Zhang 0002, Haojie Lu, Longyi Zhou, Huachen Tan, Kan Zhong, Yujuan Tan, Chaoxia Qin |
DAC | 5 |
| 2025 | SH-RAG: A Syntax-Based Hierarchical Retrieval-Augmented Generation Framework for Handling Syntactic Complexity in Literary Machine TranslationabstractWith the rapid advancements in machine translation (MT), large language models (LLMs) have demonstrated the ability to generate high-quality translations with relatively low computational resources compared to training neural machine translation (NMT) systems from scratch. Recent studies have shown significant improvements in translation quality by lever- aging in-context learning (ICL), which provides LLMs with context during the translation process. However, when it comes to literary texts, which are inherently informative in content, diverse in genres and complex in sentence structures, current LLMs and existing methodologies struggle to perform effectively. A key factor contributing to this suboptimal performance is the insufficient infusion of classical literary translation knowledge during the pre-training phase of LLMs, which hampers their effectiveness in literary translation tasks. In this paper, we propose a novel method based on syntactic hierarchical Retrieval Augmented Generation (RAG) to address the unique challenges of literary machine translation (LMT) for English-to-Chinese novel translations. Grounded in the theory of elements in fiction, we construct a comprehensive dataset that encompasses four distinct content types commonly found in literary texts. Our dataset consists of manually collected high-quality bilingual parallel translation pairs extracted from four novels, which are utilized for database construction, example retrieval, LLMs fine- tuning and testing. By establishing a maintainable database and designing a hierarchical retrieval system, our method selects semantically relevant examples from the source language to facilitate augmented generation by the LLMs. Through extensive experimentation, we demonstrate the effectiveness of our method in enhancing the translation quality of LLMs. The results indicate significant improvements, with our hierarchical RAG framework yielding an average increase of 25.75% in METEOR scores compared to LLaMA3-70B, 21.31% compared to Google Translate, 11.46% compared to GPT-3.5 and an improvement of more than 10% in GPT-4o scores over all metrics. These empirical findings substantiate the superiority of our approach, establishing its potential to advance the field of literary machine translation. Haojie Lu, Zhenjie Zhao |
IJCNN | 1 |
| 2024 | Rethinking Literary Plagiarism in LLMs through the Lens of Copyright Laws
Huachen Tan, Moming Duan, Duo Liu 0002, Haojie Lu, Yuexin Mu, Longyi Zhou, Ao Ren, Yujuan Tan, Kan Zhong |
ACML | 4 |
| 2024 | Attention and Frequency-Domain Fusion Network for Detecting Skeleton Edge Points in Ankle Joint Fracture DiagnosisabstractAnkle fractures are one of the most common types of fractures in clinical practice, and incorrect assessment of joint stability during surgery can easily lead to various ligament injuries and complications such as arthritis. Ultrasonic imaging is currently one of the least invasive and least burdensome methods for diagnosing ankle fractures in patients. However, previous studies mostly diagnosed joint fractures by segmenting the shape and position of the entire bone, while the latest clinical research found that joint stability can be determined by measuring the gap between two segments of bone. Therefore, we propose a new model that accurately detects skeletal edge points on both sides of the ankle joint, measures the length of the bone gap, and thus assists doctors in accurate diagnosis or treatment planning. Our model uses a Region Proposal Network and Transformer encoder as the basic backbone. It first filters the background to obtain rough interesting regions, then collectively models the internal attention between background pixels and target pixels to enhance the distinction between the skeletal main body and background. Additionally, it extends frequency domain recognition to strengthen the differentiation between skeletal main body and background, thereby promoting edge point detection. We use ankle joint ultrasonic images from real patients, annotated by bone disease experts. Experimental results demonstrate that our model can accurately locate the detailed positions of skeletal edge points, even in images containing complex body tissues, ensuring correct detection. Furthermore, In addition, our model can process 24 frames per second and has the ability to track frames in a video. Yaqi Tian, Haojie Lu, Zhe Zhao 0005, Fang Chen 0007 |
BIBM | 3 |
| 2023 | Leveraging trans-ethnic genetic risk scores to improve association power for complex traits in underrepresented populationsabstractTrans-ethnic genome-wide association studies have revealed that many loci identified in European populations can be reproducible in non-European populations, indicating widespread trans-ethnic genetic similarity. However, how to leverage such shared information more efficiently in association analysis is less investigated for traits in underrepresented populations. We here propose a statistical framework, trans-ethnic genetic risk score informed gene-based association mixed model (GAMM), by hierarchically modeling single-nucleotide polymorphism effects in the target population as a function of effects of the same trait in well-studied populations. GAMM powerfully integrates genetic similarity across distinct ancestral groups to enhance power in understudied populations, as confirmed by extensive simulations. We illustrate the usefulness of GAMM via the application to 13 blood cell traits (i.e. basophil count, eosinophil count, hematocrit, hemoglobin concentration, lymphocyte count, mean corpuscular hemoglobin, mean corpuscular hemoglobin concentration, mean corpuscular volume, monocyte count, neutrophil count, platelet count, red blood cell count and total white blood cell count) in Africans of the UK Biobank (n = 3204) while utilizing genetic overlap shared in Europeans (n = 746 667) and East Asians (n = 162 255). We discovered multiple new associated genes, which had otherwise been missed by existing methods, and revealed that the trans-ethnic information indirectly contributed much to the phenotypic variance. Overall, GAMM represents a flexible and powerful statistical framework of association analysis for complex traits in underrepresented populations by integrating trans-ethnic genetic similarity across well-studied populations, and helps attenuate health inequities in current genetics research for people of minority populations. Haojie Lu, Ping Zeng |
Briefings Bioinform. | 1 |
| 2022 | Identifying pleiotropic genes for complex phenotypes with summary statistics from a perspective of composite null hypothesis testingabstractPleiotropy has important implication on genetic connection among complex phenotypes and facilitates our understanding of disease etiology. Genome-wide association studies provide an unprecedented opportunity to detect pleiotropic associations; however, efficient pleiotropy test methods are still lacking. We here consider pleiotropy identification from a methodological perspective of high-dimensional composite null hypothesis and propose a powerful gene-based method called MAIUP. MAIUP is constructed based on the traditional intersection-union test with two sets of independent P-values as input and follows a novel idea that was originally proposed under the high-dimensional mediation analysis framework. The key improvement of MAIUP is that it takes the composite null nature of pleiotropy test into account by fitting a three-component mixture null distribution, which can ultimately generate well-calibrated P-values for effective control of family-wise error rate and false discover rate. Another attractive advantage of MAIUP is its ability to effectively address the issue of overlapping subjects commonly encountered in association studies. Simulation studies demonstrate that compared with other methods, only MAIUP can maintain correct type I error control and has higher power across a wide range of scenarios. We apply MAIUP to detect shared associated genes among 14 psychiatric disorders with summary statistics and discover many new pleiotropic genes that are otherwise not identified if failing to account for the issue of composite null hypothesis testing. Functional and enrichment analyses offer additional evidence supporting the validity of these identified pleiotropic genes associated with psychiatric disorders. Overall, MAIUP represents an efficient method for pleiotropy identification. Haojie Lu, Ping Zeng |
Briefings Bioinform. | 2 |