VLDB 2026 Research / reviewers in the wild / expert
Jinghui Lu
dblp:14/983
· DBLP profile ↗
18ranked-venue papers
7as first author
15since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 6 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Security and privacy · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MEML-GRPO: Heterogeneous Multi-Expert Mutual Learning for RLVR AdvancementabstractRecent advances demonstrate that reinforcement learning with verifiable rewards (RLVR) significantly enhances the reasoning capabilities of large language models (LLMs). However, standard RLVR faces challenges with reward sparsity, where zero rewards from consistently incorrect candidate answers provide no learning signal, particularly in challenging tasks. To address this,we propose Multi-Expert Mutual Learning GRPO (MEML-GRPO), an innovative framework that utilizes diverse expert prompts as system prompts to generate a broader range of responses, substantially increasing the likelihood of identifying correct solutions. Additionally, we introduce an inter-expert mutual learning mechanism that facilitates knowledge sharing and transfer among experts, further boosting the model’s performance through RLVR. Extensive experiments across multiple reasoning benchmarks show that MEML-GRPO delivers significant improvements, achieving an average performance gain of 4.89% with Qwen and 11.33% with Llama, effectively overcoming the core limitations of traditional RLVR methods. Weitao Jia, Jinghui Lu, Haiyang Yu 0004, Guozhi Tang, An-Lan Wang, Weijie Yin, Dingkang Yang, Yuxiang Nie, Bin Shan, Hao Feng 0009, Irene Li, Kun Yang 0010, Jingqun Tang, Teng Fu 0001, Changhong Jin, Xiaohui Lv, Can Huang 0002 |
AAAI | 2 |
| 2026 | ContourFD-Net: A Finite-Difference-Driven Contour Attention Network for Efficient Medical Image Segmentation on Edge Devices
Zhengbei Jin, Jinghui Lu, Beibei Jin |
IEEE Internet Things J. | 2 |
| 2025 | HealthGenie: A Knowledge-Driven LLM Framework for Tailored Dietary GuidanceabstractSeeking dietary guidance often requires navigating complex nutritional knowledge while considering individual health needs. To address this, we present HealthGenie, an interactive platform that leverages the interpretability of knowledge graphs (KGs) and the conversational power of large language models (LLMs) to deliver tailored dietary recommendations alongside integrated nutritional visualizations for fast, intuitive insights. Upon receiving a user query, HealthGenie performs intent refinement and maps user's needs to a curated nutritional knowledge graph. The system then retrieves and visualizes relevant subgraphs, while offering detailed, explainable recommendations. Users can interactively adjust preferences to further tailor results. A within-subject study and quantitative analysis show that HealthGenie reduces cognitive load and interaction effort while supporting personalized, health-aware decision-making. Xinjie Zhao 0004, Ding Xia, Zhongyi Zhou, Rui Yang 0016, Jinghui Lu, Chanjun Park, Irene Li |
CIKM | 6 |
| 2025 | WildDoc: How Far Are We from Achieving Comprehensive and Robust Document Understanding in the Wild?abstractAn-Lan Wang, Jingqun Tang, Lei Liao, Hao Feng, Qi Liu, Xiang Fei, Jinghui Lu, Han Wang, Hao Liu, Yuliang Liu, Xiang Bai, Can Huang. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. An-Lan Wang, Jingqun Tang, Hao Feng 0009, Jinghui Lu, Hao Liu 0003, Xiang Bai, Can Huang 0002 |
EMNLP | 7 |
| 2025 | MMLU-ProX: A Multilingual Benchmark for Advanced Large Language Model EvaluationabstractWeihao Xuan, Rui Yang, Heli Qi, Qingcheng Zeng, Yunze Xiao, Aosong Feng, Dairui Liu, Yun Xing, Junjue Wang, Fan Gao, Jinghui Lu, Yuang Jiang, Huitao Li, Xin Li, Kunyu Yu, Ruihai Dong, Shangding Gu, Yuekang Li, Xiaofei Xie, Felix Juefei-Xu, Foutse Khomh, Osamu Yoshie, Qingyu Chen, Douglas Teodoro, Nan Liu, Randy Goebel, Lei Ma, Edison Marrese-Taylor, Shijian Lu, Yusuke Iwasawa, Yutaka Matsuo, Irene Li. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Weihao Xuan, Rui Yang 0016, Heli Qi, Qingcheng Zeng, Yunze Xiao, Aosong Feng, Dairui Liu, Yun Xing 0001, Jinghui Lu, Yuang Jiang, Huitao Li, Xin Li 0079, Kunyu Yu, Ruihai Dong, Shangding Gu, Yuekang Li, Xiaofei Xie, Felix Juefei-Xu, Foutse Khomh, Osamu Yoshie, Qingyu Chen 0001, Douglas Teodoro, Nan Liu 0003, Randy Goebel, Lei Ma 0003, Edison Marrese-Taylor, Shijian Lu, Yusuke Iwasawa, Yutaka Matsuo, Irene Li |
EMNLP | 11 |
| 2025 | Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLMabstractThe application of Large Vision-Language Models (LVLMs) for analyzing images and videos is an exciting and rapidly evolving field. In recent years, we've seen significant growth in high-quality image-text datasets for fine-tuning image understanding, but there is still a lack of comparable datasets for videos. Additionally, many VideoLLMs are extensions of single-image VLMs, which may not efficiently handle the complexities of longer videos. In this study, we introduce a large-scale synthetic dataset created from proprietary models, using carefully designed prompts to tackle a wide range of questions. We also explore a dynamic visual token compression architecture that strikes a balance between computational efficiency and performance. Our proposed \model{} achieves state-of-the-art results across various video tasks and shows impressive generalization, setting new baselines in multi-image understanding. Notably, \model{} delivers an absolute improvement of 2.7\% over LLaVA-OneVision on VideoMME and 10.7\% on MuirBench. Codes are available at https://github.com/Hon-Wong/ByteVideoLLM Yuxiang Nie, Yongjie Ye, Haiyang Yu 0004, Jinghui Lu, Can Huang 0002 |
ICCV | 7 |
| 2024 | SDA: Simple Discrete Augmentation for Contrastive Sentence Representation LearningabstractContrastive learning has recently achieved compelling performance in unsupervised sentence representation. As an essential element, data augmentation protocols, however, have not been well explored. The pioneering work SimCSE resorting to a simple dropout mechanism (viewed as continuous augmentation) surprisingly dominates discrete augmentations such as cropping, word deletion, and synonym replacement as reported. To understand the underlying rationales, we revisit existing approaches and attempt to hypothesize the desiderata of reasonable data augmentation methods: balance of semantic consistency and expression diversity. We then develop three simple yet effective discrete sentence augmentation schemes: punctuation insertion, modal verbs, and double negation. They act as minimal noises at lexical level to produce diverse forms of sentences. Furthermore, standard negation is capitalized on to generate negative samples for alleviating feature suppression involved in contrastive learning. We experimented extensively with semantic textual similarity on diverse datasets. The results support the superiority of the proposed methods consistently. Our key code is available at https://github.com/Zhudongsheng75/SDA Dongsheng Zhu, Zhenyu Mao, Jinghui Lu, Fei Tan 0002 |
LREC/COLING | 3 |
| 2024 | VisLingInstruct: Elevating Zero-Shot Learning in Multi-Modal Language Models with Autonomous Instruction OptimizationabstractDongsheng Zhu, Daniel Tang, Weidong Han, Jinghui Lu, Yukun Zhao, Guoliang Xing, Junfeng Wang, Dawei Yin. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Dongsheng Zhu, Daniel Tang, Weidong Han 0002, Jinghui Lu, Yukun Zhao, Guoliang Xing, Junfeng Wang 0009, Dawei Yin 0001 |
NAACL-HLT | 4 |
| 2024 | PaDeLLM-NER: Parallel Decoding in Large Language Models for Named Entity RecognitionabstractIn this study, we aim to reduce generation latency for Named Entity Recognition (NER) with Large Language Models (LLMs). The main cause of high latency in LLMs is the sequential decoding process, which autoregressively generates all labels and mentions for NER, significantly increase the sequence length. To this end, we introduce Parallel Decoding in LLM for NE} (PaDeLLM-NER), a approach that integrates seamlessly into existing generative model frameworks without necessitating additional modules or architectural modifications. PaDeLLM-NER allows for the simultaneous decoding of all mentions, thereby reducing generation latency. Experiments reveal that PaDeLLM-NER significantly increases inference speed that is 1.76 to 10.22 times faster than the autoregressive approach for both English and Chinese. Simultaneously it maintains the quality of predictions as evidenced by the performance that is on par with the state-of-the-art across various datasets. All resources are available at https://github.com/GeorgeLuImmortal/PaDeLLM_NER. Jinghui Lu, Xuejing Liu, Brian Mac Namee, Can Huang 0002 |
NeurIPS | 1 |
| 2024 | From Liberty to 1984: A Methodology for Systematically Deteriorating LLM Outputs through Habituation TendenciesabstractLLM has profoundly impacted various aspects of society, including public cognition and political systems. LLM has similar characteristics to human brain and is particularly susceptibility to external influences. As the social influence of LLM continue to expand, there is a growing need to strengthen its safety to reduce potential risks. This study considers habituation as a fundamental flaw of LLM and conducts necessary analysis in the fields of cognitive neuroscience, philosophy, and computer science. The main findings show that that habituation causes the human brain and LLM to gradually lose sensitivity and tolerance to small changes, resulting in uncontrollable outputs. This root cause of this phenomenon lies in the scarcity of computing resources, which is an inevitable basic mechanism in complex situations, in either computer operation or biological brain function. This study hypothesizes that habituation can systematically induce malicious outputs in LLM. By analyzing common attacks on LLM, this study confirms that habituation is a key underlying principle that affects the effectiveness of these attacks. To further verify this hypothesis, two novel attack methods are proposed based on the analysis of habituation defects in LLM: the Repetition Exposure Effect Attack and the Progressive Sensitivity Reduction Attack. The Repetition Exposure Effect attack increased the average ASR (Attack Success Rate) from 0.1 to 0.5; The Progressive Sensitivity Reduction attack increased the average ASR from 0.2 to 0.4. These attacks show higher ASR in major LLM tests, further verifying habituation is a fundamental flaw of LLM. Subsequently, this study discusses targeted mitigation strategies for LLM caused by habituation. Finally, future research directions are explored, including evolving new attack methods based on habituation and proposing comprehensive and fundamental defense strategies to enhance the safety and stability of LLM. Huijun Chen, YinFeng Zheng, Jinghui Lu |
TrustCom | 6 |
| 2023 | PUnifiedNER: A Prompting-Based Unified NER System for Diverse DatasetsabstractMuch of named entity recognition (NER) research focuses on developing dataset-specific models based on data from the domain of interest, and a limited set of related entity types. This is frustrating as each new dataset requires a new model to be trained and stored. In this work, we present a ``versatile'' model---the Prompting-based Unified NER system (PUnifiedNER)---that works with data from different domains and can recognise up to 37 entity types simultaneously, and theoretically it could be as many as possible. By using prompt learning, PUnifiedNER is a novel approach that is able to jointly train across multiple corpora, implementing intelligent on-demand entity recognition. Experimental results show that PUnifiedNER leads to significant prediction benefits compared to dataset-specific models with impressively reduced model deployment costs. Furthermore, the performance of PUnifiedNER can achieve competitive or even better performance than state-of-the-art domain-specific methods for some datasets. We also perform comprehensive pilot and ablation studies to support in-depth analysis of each component in PUnifiedNER. Jinghui Lu, Brian Mac Namee, Fei Tan 0002 |
AAAI | 1 |
| 2023 | What Makes Pre-trained Language Models Better Zero-shot Learners?abstractCurrent methods for prompt learning in zeroshot scenarios widely rely on a development set with sufficient human-annotated data to select the best-performing prompt template a posteriori.This is not ideal because in a real-world zero-shot scenario of practical relevance, no labelled data is available.Thus, we propose a simple yet effective method for screening reasonable prompt templates in zero-shot text classification: Perplexity Selection (Perplection).We hypothesize that language discrepancy can be used to measure the efficacy of prompt templates, and thereby develop a substantiated perplexity-based scheme allowing for forecasting the performance of prompt templates in advance.Experiments show that our method leads to improved prediction performance in a realistic zero-shot setting, eliminating the need for any labelled examples.PPL Acc.(%) PPL Acc.(%) PPL Acc.(%) PPL Acc.(%) DOUBAN 24.61 57.12 40.93 50.98 28.80 56.68 71.01 51.31 WEIBO 19.78 61.79 30.37 51.16 22.34 58.35 44.45 50.92WAIMAI 16.44 67.80 23.34 53.15 19.68 69.72 36.07 48.49ECOMMERCE 14.07 73.12 18.45 55.68 16.88 67. Jinghui Lu, Dongsheng Zhu, Weidong Han 0002, Brian Mac Namee, Fei Tan 0002 |
ACL (1) | 1 |
| 2022 | A Rationale-Centric Framework for Human-in-the-loop Machine LearningabstractWe present a novel rationale-centric framework with human-in-the-loop -Rationales-centric Double-robustness Learning (RDL) -to boost model out-of-distribution performance in few-shot learning scenarios.By using static semi-factual generation and dynamic humanintervened correction, RDL exploits rationales (i.e.phrases that cause the prediction), human interventions and semi-factual augmentations to decouple spurious associations and bias models towards generally applicable underlying distributions, which enables fast and accurate generalisation.Experimental results show that RDL leads to significant prediction benefits on both in-distribution and out-of-distribution tests compared to many state-of-the-art benchmarks-especially for few-shot learning scenarios.We also perform extensive ablation studies to support in-depth analyses of each component in our framework. Jinghui Lu, Linyi Yang, Brian Mac Namee, Yue Zhang 0004 |
ACL (1) | 1 |
| 2022 | Android Mobile Terminal Security Assessment Based on Analytical Hierarchy Process (AHP)
Linghang Shi, Huijun Chen, Jinghui Lu |
ICDF2C | 4 |
| 2021 | A Sentence-Level Hierarchical BERT Model for Document Classification with Limited Labelled Data
Jinghui Lu, Maeve Henchion, Ivan Bacher, Brian Mac Namee |
DS | 1 |
| 2020 | Diverging Divergences: Examining Variants of Jensen Shannon Divergence for Corpus Comparison TasksabstractJensen-Shannon divergence (JSD) is a distribution similarity measurement widely used in natural language processing. In corpus comparison tasks, where keywords are extracted to reveal the divergence between different corpora (for example, social media posts from proponents of different views on a political issue), two variants of JSD have emerged in the literature. One of these uses a weighting based on the relative sizes of the corpora being compared. In this paper we argue that this weighting is unnecessary and, in fact, can lead to misleading results. We recommend that this weighted version is not used. We base this recommendation on an analysis of the JSD variants and experiments showing how they impact corpus comparison results as the relative sizes of the corpora being compared change. Jinghui Lu, Maeve Henchion, Brian Mac Namee |
LREC | 1 |
| 2011 | Fading Characteristics in the Railway Terrain CuttingsabstractA high performance wireless network is essential for for the railway communication and control systems. Research on the fading characteristics in railway environment is of great importance for the design of the railway wireless network. In this paper, measurements are taken in railway terrain cuttings area using track side base stations of the GSM-R network. The fitted path loss model, shadow fading, and dynamaic range of the small scale fading are obtained and compared to the results of viaduct scenario. The propagation environment of the terrain cuttings turns out to be worse than the viaduct area. The path loss exponent is found to be 4.3. The shadow loss can be reasonably described by a log-normal distribution. It is also found that the bridges over the cuttings can cause extra loss of about 5 dB. The dynamaic range of the small scale fading is from 27 dB to 40 dB with a mean value of about 33 dB. Jinghui Lu, Cesar Briso-Rodríguez |
VTC Spring | 1 |
| 2010 | A Reference Ontology based Approach for Service Oriented Ontology Management
Shuying Wang, Jinghui Lu, Miriam A. M. Capretz |
WEBIST (2) | 2 |