VLDB 2026 Research / reviewers in the wild / expert
Yunfei Lu
dblp:87/8501
· DBLP profile ↗
22ranked-venue papers
8as first author
17since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 4 first-author · 11 since 2021Databases, data management, data science and information retrieval · 7 · 5 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LangGPS: Language Separability Guided Data Pre-Selection for Joint Multilingual Instruction TuningabstractJoint multilingual instruction tuning is a widely adopted approach to improve the multilingual instruction-following ability and downstream performance of large language models (LLMs), but the resulting multilingual capability remains highly sensitive to the composition and selection of the training data. Existing selection methods, often based on features like text quality, diversity, or task relevance, typically overlook the intrinsic linguistic structure of multilingual data. In this paper, we propose LangGPS, a lightweight two-stage pre-selection framework guided by language separability—a signal that quantifies how well samples in different languages can be distinguished in the model’s representation space. LangGPS first filters training data based on separability scores and then refines the subset using existing selection methods. Extensive experiments across six benchmarks and 22 languages demonstrate that applying LangGPS on top of existing selection methods improves their effectiveness and generalizability in multilingual training, especially for understanding tasks and low-resource languages. Further analysis reveals that highly separable samples facilitate the formation of clearer language boundaries and support faster adaptation, while low-separability samples tend to function as bridges for cross-lingual alignment. Besides, we also find that language separability can serves as an effective signal for multilingual curriculum learning, where interleaving samples with diverse separability levels yields stable and generalizable gains. Together, we hope our work offers a new perspective on data utility in multilingual contexts and support the development of more linguistically informed LLMs. Yangfan Ye, Xiachong Feng, Lei Huang 0021, Weitao Ma, Qichen Hong, Yunfei Lu, Duyu Tang, Dandan Tu, Bing Qin 0001 |
AAAI | 7 |
| 2026 | MPR-GUI: Benchmarking and Enhancing Multilingual Perception and Reasoning in GUI AgentsabstractRuihan Chen, Qiming Li, Xiaocheng Feng, Weihong Zhong, Xiaoliang Yang, Yuxuan Gu, Zekun Zhou, Yunfei Lu, Haoyu Ren, Kun Chen, Dandan Tu, Bing Qin. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Ruihan Chen 0001, Weihong Zhong, Xiaoliang Yang, Yuxuan Gu 0004, Ze-kun Zhou, Yunfei Lu, Dandan Tu, Bing Qin 0001 |
ACL (1) | 8 |
| 2026 | Unlocking Multilingual Reasoning Capability of LLMs and LVLMs through Representation EngineeringabstractQiming Li, Xiaocheng Feng, Yixuan Ma, Ruihan Chen, Zihe Tong, Zekai Ye, Xiachong Feng, Libo Qin, Haoyu Ren, Kun Chen, Yunfei Lu, Dandan Tu, Bing Qin. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yixuan Ma, Ruihan Chen 0001, Zihe Tong, Zekai Ye, Xiachong Feng, Libo Qin 0001, Yunfei Lu, Dandan Tu, Bing Qin 0001 |
ACL (1) | 11 |
| 2026 | Culture-Aware Machine Translation in Large Language Models: Benchmarking and InvestigationabstractZekun Yuan, Yangfan Ye, Xiaocheng Feng, Baohang Li, Qichen Hong, Yunfei Lu, Dandan Tu, Bing Qin. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Zekun Yuan, Yangfan Ye, Baohang Li, Qichen Hong, Yunfei Lu, Dandan Tu, Bing Qin 0001 |
ACL (1) | 6 |
| 2025 | Enhancing Non-English Capabilities of English-Centric Large Language Models Through Deep Supervision Fine-TuningabstractLarge language models (LLMs) have demonstrated significant progress in multilingual language understanding and generation. However, due to the imbalance in training data, their capabilities in non-English languages are limited. Recent studies revealed the English-pivot multilingual mechanism of LLMs, where LLMs implicitly convert non-English queries into English ones at the bottom layers and adopt English for thinking at the middle layers. However, due to the absence of explicit supervision for cross-lingual alignment in the intermediate layers of LLMs, the internal representations during these stages may become inaccurate. In this work, we introduce a deep supervision fine-tuning method (DFT) that incorporates additional supervision in the internal layers of the model to guide its workflow. Specifically, we introduce two training objectives on different layers of LLMs: one at the bottom layers to constrain the conversion of the target language into English, and another at the middle layers to constrain reasoning in English. To effectively achieve the guiding purpose, we designed two types of supervision signals: logits and feature, which represent a stricter constraint and a relatively more relaxed guidance. Our method guides the model to not only consider the final generated result when processing non-English inputs but also ensure the accuracy of internal representations. We conducted extensive experiments on typical English-centric large models, LLaMA-2 and Gemma-2, and the results on multiple multilingual datasets show that our method significantly outperforms traditional fine-tuning methods. Wenshuai Huo, Yichong Huang, Chengpeng Fu, Baohang Li, Yangfan Ye, Zhirui Zhang, Dandan Tu, Duyu Tang, Yunfei Lu, Hui Wang 0030, Bing Qin 0001 |
AAAI | 10 |
| 2025 | CC-Tuning: A Cross-Lingual Connection Mechanism for Improving Joint Multilingual Supervised Fine-TuningabstractYangfan Ye, Xiaocheng Feng, Zekun Yuan, Xiachong Feng, Libo Qin, Lei Huang, Weitao Ma, Yichong Huang, Zhirui Zhang, Yunfei Lu, Xiaohui Yan, Duyu Tang, Dandan Tu, Bing Qin. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Yangfan Ye, Zekun Yuan, Xiachong Feng, Libo Qin 0001, Lei Huang 0021, Weitao Ma, Yichong Huang, Zhirui Zhang, Yunfei Lu, Duyu Tang, Dandan Tu, Bing Qin 0001 |
ACL (1) | 10 |
| 2025 | CLAIM: Mitigating Multilingual Object Hallucination in Large Vision-Language Models with Cross-Lingual Attention InterventionabstractLarge Vision-Language Models (LVLMs) have demonstrated impressive multimodal abilities but remain prone to multilingual object hallucination, with a higher likelihood of generating responses inconsistent with the visual input when utilizing queries in non-English languages compared to English. Most existing approaches to address these rely on pretraining or fine-tuning, which are resource-intensive. In this paper, inspired by observing the disparities in cross-modal attention patterns across languages, we propose Cross-Lingual Attention Intervention for Mitigating multilingual object hallucination (CLAIM) in LVLMs, a novel near training-free method by aligning attention patterns. CLAIM first identifies language-specific cross-modal attention heads, then estimates language shift vectors from English to the target language, and finally intervenes in the attention outputs during inference to facilitate cross-lingual visual perception capability alignment. Extensive experiments demonstrate that CLAIM achieves an average improvement of 13.56% (up to 30% in Spanish) on the POPE and 21.75% on the hallucination subsets of the MME benchmark across various languages. Further analysis reveals that multilingual attention divergence is most prominent in intermediate layers, highlighting their critical role in multilingual scenarios. Zekai Ye, Libo Qin 0001, Yichong Huang, Baohang Li, Kui Jiang, Yang Xiang 0003, Zhirui Zhang, Yunfei Lu, Duyu Tang, Dandan Tu, Bing Qin 0001 |
ACL (1) | 10 |
| 2025 | Enhancing Speech Large Language Models with Prompt-Aware Mixture of Audio EncodersabstractWeiqiao Shan, Yuang Li, Yuhao Zhang, Yingfeng Luo, Chen Xu, Xiaofeng Zhao, Long Meng, Yunfei Lu, Min Zhang, Hao Yang, Tong Xiao, JingBo Zhu. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Weiqiao Shan, Yuang Li, Yingfeng Luo, Chen Xu 0008, Long Meng, Yunfei Lu, Min Zhang 0042, Hao Yang 0006, Tong Xiao 0001 |
EMNLP | 8 |
| 2025 | MFTP: Multi-round Feedback for Dynamic Travel Itinerary Optimization
Yunfei Lu, Lingjiao Xu |
KSEM (2) | 1 |
| 2025 | Ensuring Context Completeness in Retrieval-Augmented Generation for Knowledge-Intensive Question-Answering
Lingjiao Xu, Yunfei Lu |
NLPCC (3) | 4 |
| 2025 | ViSNeRF: Efficient Multidimensional Neural Radiance Field Representation for Visualization Synthesis of Dynamic Volumetric ScenesabstractDomain scientists often face I/O and storage challenges when keeping raw data from large-scale simulations. Saving visualization images, albeit practical, is limited to preselected viewpoints, transfer functions, and simulation parameters. Recent advances in scientific visualization leverage deep learning techniques for visualization synthesis by offering effective ways to infer unseen visualizations when only image samples are given during training. However, due to the lack of 3D geometry awareness, existing methods typically require many training images and significant learning time to generate novel visualizations faithfully. To address these limitations, we propose ViSNeRF, a novel 3D-aware approach for visualization synthesis using neural radiance fields. Leveraging a multidimensional radiance field representation, ViSNeRF efficiently reconstructs visualizations of dynamic volumetric scenes from a sparse set of labeled image samples with flexible parameter exploration over transfer functions, isovalues, timesteps, or simulation parameters. Through qualitative and quantitative comparative evaluation, we demonstrate ViSNeRF’s superior performance over several representative baseline methods, positioning it as the state-of-the-art solution. The code is available at https://github.com/JCBreath/ViSNeRF. Siyuan Yao, Yunfei Lu, Chaoli Wang 0001 |
PacificVis | 2 |
| 2025 | Boosting Document Image Translation via Layout-Aware Semantic Paragraph Clustering
Yupu Liang, Yunfei Lu, Dandan Tu, Chengqing Zong, Yu Zhou 0001 |
PRCV (7) | 6 |
| 2025 | Boosting adversarial example detection via local histogram equalization and spectral feature analysis
Yunfei Lu, Chenxia Chang, Shaowen Yao 0001, Ahmed Zahir |
Vis. Comput. | 1 |
| 2024 | FCNR: Fast Compressive Neural Representation of Visualization ImagesabstractWe present FCNR, a fast compressive neural representation for tens of thousands of visualization images under varying viewpoints and timesteps. The existing NeRVI solution, albeit enjoying a high compression ratio, incurs slow speeds in encoding and decoding. Built on the recent advances in stereo image compression, FCNR assimilates stereo context modules and joint context transfer modules to compress image pairs. Our solution significantly improves encoding and decoding speed while maintaining high reconstruction quality and satisfying compression ratio. To demonstrate its effectiveness, we compare FCNR with state-of-the-art neural compression methods, including E-NeRV, HNeRV, NeRVI, and ECSIC. The source code can be found at https://github.com/YunfeiLu0112/FCNR. Yunfei Lu, Pengfei Gu, Chaoli Wang 0001 |
IEEE VIS | 1 |
| 2023 | Online Tensor Method for Moving Objective Detection with FMCW RadarabstractFrequency modulated continuous wave (FMCW) radar can precisely detect moving objects utilizing the Doppler information. However, only exploiting the Doppler information in one frame can usually lead to object false detection when static background has large radar cross section or the moving objective occludes some static background. In this paper, we investigate the moving objective detection problem with FMCW radar through utilizing the Doppler information in multiple frames to increase objective detection accuracy. To solve this problem, an online tensor robust principal component analysis (RPCA) algorithm is proposed with low hardware and computation complexity. The proposed algorithm can maintain the intrinsic tensor data structure. Experimental results show that the proposed algorithm can accurately detect the static background and moving object even for the case of occlusion or static object with large RCS. Yunfei Lu, Zhaoyang Zhang 0001, Xin Tong 0008, Zhaohui Yang 0001 |
VTC2023-Spring | 1 |
| 2023 | Exploring Fine-Grained In-Memory Database Performance for Modern CPUsabstractModern CPUs keep integrating more cores and large size cache, which is beneficial for in-memory databases to improve parallel processing power and cache locality. While state-of-the-art CPUs have diverse architectures and roadmaps such as large core count and large cache size (AMD x86), moderate core count and cache size (intel x86), large core count and moderate cache size (ARM), exploring in-memory databases performance characteristics for different CPU architectures is important for in-memory database designs and optimizations. In this paper, we develop a fine-grained in-memory database benchmark to evaluate the performance of each operator on different CPUs to explore how CPU hardware architectures influence performance. Different from well known conclusions that more cores and larger cache size can achieve higher performance, we find out that the micro cache architectures play an important role opposite to core count and cache size, the shared monolithic L3 cache with moderate size beats large disaggregated L3 cache. The experiments also show that predicting operator performance on different CPUs is difficult according to diverse CPU architectures and micro cache architectures, and different implementations of each operator are not always high or low with interleaved strong and weak performance regions influenced by CPU hardware architectures. Intel x86 CPUs represent cache-centric processor design, while AMD x86 and ARM CPUs represent computing-centric processor design, the OLAP benchmark experiments of SSB discover that OmniSciDB and OLAP Accelerator with vector-wise processing model performs well on intel x86 CPUs compared to AMD x86 CPUs and the JIT compliant based Hyper prefers to AMD x86 CPUs rather than intel x86 CPUs. The CPU roadmaps of increasing cores or improving cache locality should be considered for in-memory database algorithm design and platform selection. Zhuan Liu, Ruichen Han, Yu Zhang 0183, Tao Zhong 0001, Roman Dementiev, Yunfei Lu, Mingjian Que |
IEEE Trans. Parallel Distributed Syst. | 9 |
| 2022 | Uncovering the Heterogeneous Effects of Preference Diversity on User Activeness: A Dynamic Mixture ModelabstractPreference diversity arouses much research attention in recent years, as it is believed to be closely related to many profound problems such as user activeness in social media or recommendation systems. However, due to the lack of large-scale data with comprehensive user behavior log and accurate content labels, the real quantitative effect of preference diversity on user activeness is still largely unknown. This paper studies the heterogeneous effect of preference diversity on user activeness in social media. We examine large-scale real-world datasets collected from two of the most popular video-sharing social platforms in China, including the behavior logs of more than 787 thousand users and 1.95 million videos, with accurate content category information. We investigate the distribution and evolution of preference diversity, and find rich heterogeneity in the effect of preference diversity on the dynamic activeness. Furthermore, we discover the divergence of preference diversity mechanisms for the same user under different usage scenarios, such as active (where users actively seek information) and passive (where users passively receive information) modes. Unlike existing qualitative studies, we propose a universal mixture model with the capability of accurately fitting dynamic activeness curves while reflecting the heterogeneous patterns of preference diversity. To our best knowledge, this is the first quantitative model that incorporates the effect of preference diversity on user activeness. With the modeling parameters, we are able to make accurate churn and activeness predictions and provide decision support for increasing user activity through the intervention of diversity. Our findings and model comprehensively reveal the significance of preference diversity and provide potential implications for the design of future recommendation systems and social media. Yunfei Lu, Peng Cui 0001, Linyun Yu, Lei Li 0005, Wenwu Zhu 0001 |
KDD | 1 |
| 2020 | Exploring the collective human behavior in cascading systems: a comprehensive framework
Yunfei Lu, Linyun Yu, Tianyang Zhang 0001, Chengxi Zang, Peng Cui 0001, Chaoming Song, Wenwu Zhu 0001 |
Knowl. Inf. Syst. | 1 |
| 2019 | Uncovering the Co-driven Mechanism of Social and Content Links in User Churn PhenomenaabstractRecent years witness the merge of social networks and user-generated content (UGC) platforms. In these new platforms, users establish links to others not only driven by their social relationships in the physical world but also driven by the contents published by others. During this merging process, social networks gradually integrate both social and content links and become unprecedentedly complicated, with the motivation to exploit both the advantages of social viscosity and content attractiveness to reach the best customer retention situation. However, due to the lack of fine-grained data recording such merging phenomena, the co-driven mechanism of social and content links in churn remains unexplored. How do social and content factors jointly influence customers' churn? What is the best ratio of social and content links for retention? Is there a model to capture this co-driven mechanism in churn phenomena? In this paper, we collect a real-world dataset with more than 5.77 million users and 1.15 billion links, with each link being tagged as a social one or a content one. We find that both social and content links have a significant impact on users' churn and they work jointly as a complicated mixture effect. As a result, we propose a novel survival model, which incorporates both social and content factors, to predict churn probability over time. Our model successfully fits the churn distribution in reality and accurately predicts the churn rate of different subpopulations in the future. By analyzing the modeling parameters, we try to strike a balance between social-driven and content-driven links in a user's social network to reach the lowest churn rate. Our model and findings may have potential implications for the design of future social media. Yunfei Lu, Linyun Yu, Peng Cui 0001, Chengxi Zang, Renzhe Xu, Lei Li 0005, Wenwu Zhu 0001 |
KDD | 1 |
| 2018 | Collective Human Behavior in Cascading System: Discovery, Modeling and ApplicationsabstractThe collective behavior, describing spontaneously emerging social processes and events, is ubiquitous in both physical society and online social media. The knowledge of collective behavior is critical in understanding and predicting social movements, fads, riots and so on. However, detecting, quantifying and modeling the collective behavior in online social media at large scale are seldom unexplored. In this paper, we examine a real-world online social media with more than 1.7 million information spreading records, which explicitly document the detailed human behavior in this online information cascading system. We observe evident collective behavior in information cascading, and then propose metrics to quantify the collectivity. We find that previous information cascading models cannot capture the collective behavior in the real-world and thus never utilize it. Furthermore, we propose a generative framework with a latent user interest layer to capture the collective behavior in cascading system. Our framework achieves high accuracy in modeling the information cascades with respect to popularity, structure and collectivity. By leveraging the knowledge of collective behavior, our model shows the capability of making predictions without temporal features or early-stage information. Our framework can serve as a more generalized one in modeling cascading system, and, together with empirical discovery and applications, advance our understanding of human behavior. Yunfei Lu, Linyun Yu, Tianyang Zhang 0001, Chengxi Zang, Peng Cui 0001, Chaoming Song, Wenwu Zhu 0001 |
ICDM | 1 |
| 2017 | comeNgo: A Dynamic Model for Social Group EvolutionabstractHow do social groups, such as Facebook groups and Wechat groups, dynamically evolve over time? How do people join the social groups, uniformly or with burst? What is the pattern of people quitting from groups? Is there a simple universal model to depict the come-and-go patterns of various groups? In this article, we examine temporal evolution patterns of more than 100 thousands social groups with more than 10 million users. We surprisingly find that the evolution patterns of real social groups goes far beyond the classic dynamic models like SI and SIR. For example, we observe both diffusion and non-diffusion mechanism in the group joining process, and power-law decay in group quitting process, rather than exponential decay as expected in SIR model. Therefore, we propose a new modelcomeNgo, a concise yet flexible dynamic model for group evolution. Our model has the following advantages: (a) Unification power: it generalizes earlier theoretical models and different joining and quitting mechanisms we find from observation. (b) Succinctness and interpretability: it contains only six parameters with clear physical meanings. (c) Accuracy: it can capture various kinds of group evolution patterns preciously, and the goodness of fit increases by 58% over baseline. (d) Usefulness: it can be used in multiple application scenarios, such as forecasting and pattern discovery. Furthermore, our model can provide insights about different evolution patterns of social groups, and we also find that group structure and its evolution has notable relations with temporal patterns of group evolution. Tianyang Zhang 0001, Peng Cui 0001, Christos Faloutsos, Yunfei Lu, Wenwu Zhu 0001, Shiqiang Yang |
ACM Trans. Knowl. Discov. Data | 4 |
| 2016 | Come-and-Go Patterns of Group Evolution: A Dynamic ModelabstractHow do social groups, such as Facebook groups and Wechat groups, dynamically evolve over time? How do people join the social groups, uniformly or with burst? What is the pattern of people quitting from groups? Is there a simple universal model to depict the come-and-go patterns of various groups? Tianyang Zhang 0001, Peng Cui 0001, Christos Faloutsos, Yunfei Lu, Wenwu Zhu 0001, Shiqiang Yang |
KDD | 4 |