VLDB 2026 Research / reviewers in the wild / expert
Feihu Che
dblp:266/5664
· DBLP profile ↗
12ranked-venue papers
5as first author
11since 2021 · last 2026
0000-0001-7921-7154ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 4 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Two-Stage Regularization-Based Structured Pruning for LLMsabstractMingkuan Feng, Jinyang Wu, Siyuan Liu, Shuai Zhang, Hongjian Fang, Ruihan Jin, Feihu Che, Pengpeng Shao, Zhengqi Wen, Jianhua Tao. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Mingkuan Feng, Shuai Zhang 0014, Hongjian Fang, Ruihan Jin, Feihu Che, Pengpeng Shao, Zhengqi Wen, Jianhua Tao 0001 |
ACL (1) | 7 |
| 2026 | Beyond Examples: Towards Automated Thought-level In-Context Reasoning for Large Language ModelsabstractJinyang Wu, Mingkuan Feng, Shuai Zhang, Feihu Che, Zhengqi Wen, Chonghua Liao, Ling Yang, Haoran Luo, Zheng Lian, Jianhua Tao. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Mingkuan Feng, Shuai Zhang 0014, Feihu Che, Zhengqi Wen, Chonghua Liao, Zheng Lian 0004, Jianhua Tao 0001 |
ACL (1) | 4 |
| 2025 | Code-switching Mediated Sentence-level Semantic LearningabstractCode-switching is a linguistic phenomenon in which different languages are used interactively during conversation. It poses significant performance challenges to natural language processing (NLP) tasks due to the often monolingual nature of the underlying system. We focus on sentence-level semantic associations between the different code-switching expressions. And we propose an innovative task-free semantic learning method based on the semantic property. Specifically, there are many different ways of languages switching for a sentence with the same meaning. We refine this into a semantic computational method by designing the loss of semantic invariant constraint during the model optimization. In this work, we conduct thorough experiments on speech recognition, speech translation, and language modeling tasks. The experimental results fully demonstrate that the proposed method can widely improve the performance of code-switching related tasks. Shuai Zhang 0014, Jiangyan Yi, Zhengqi Wen, Jianhua Tao 0001, Feihu Che, Ruibo Fu |
AAAI | 5 |
| 2025 | Pandora's Box or Aladdin's Lamp: A Comprehensive Analysis Revealing the Role of RAG Noise in Large Language ModelsabstractRetrieval-Augmented Generation (RAG) has emerged as a key method to address hallucinations in large language models (LLMs).While recent research has extended RAG models to complex noisy scenarios, these explorations often confine themselves to limited noise types and presuppose that noise is inherently detrimental to LLMs, potentially deviating from real-world retrieval environments and restricting practical applicability.In this paper, we define seven distinct noise types from a linguistic perspective and establish a Noise RAG Benchmark (NoiserBench), a comprehensive evaluation framework encompassing multiple datasets and reasoning tasks.Through empirical evaluation of eight representative LLMs with diverse architectures and scales, we reveal that these noises can be further categorized into two practical groups: noise that is beneficial to LLMs (aka beneficial noise) and noise that is harmful to LLMs (aka harmful noise).While harmful noise generally impairs performance, beneficial noise may enhance several aspects of model capabilities and overall performance.Our analysis offers insights for developing robust RAG solutions and mitigating hallucinations across diverse retrieval scenarios.Code is available at Shuai Zhang 0014, Feihu Che, Mingkuan Feng, Pengpeng Shao, Jianhua Tao 0001 |
ACL (1) | 3 |
| 2025 | Hard or False: Keep the Balance for Negative Sampling in Knowledge GraphsabstractNegative sampling is an essential part in knowledge graph embedding, which offers significant advantages to numerous downstream related tasks. There are two kinds of important negatives: hard and false negatives. Hard negatives are the negatives which are difficult to distinguish from positive samples, while false negatives are positive samples which are mistakenly identified as negatives. Harnessing hard negatives effectively can make the model more discriminative, and reducing false negatives can avoid misleading the model during training. Therefore, the two kinds of negatives are essential in high-quality negative sampling. However, the present negative sampling methods face two shortcomings: 1.judging one negative is hard or false mainly relies on score functions; 2. difficulty in balancing the impact of hard and false negatives. In this paper, we absorb bigram language model and propose a novel criterion to help verify the negatives are hard or false, and discuss how to keep the balance between hard and false negatives. Experiments on four representative score functions and two public datasets demonstrate the effects of the proposed negative sampling method. Feihu Che, Jianhua Tao 0001, Qionghai Dai |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2024 | Multi-stage Vs Single-Stage: A Local Information Focused Approach for Overlapping Event Extraction
Shuaihu Han, Guohua Yang, Dawei Zhang 0001, Jianhua Tao 0001, Feihu Che |
ICANN (7) | 5 |
| 2024 | M2ixKG: Mixing for harder negative samples in knowledge graph
Feihu Che, Jianhua Tao 0001 |
Neural Networks | 1 |
| 2023 | Adaptive pseudo-Siamese policy network for temporal knowledge prediction
Pengpeng Shao, Feihu Che, Dawei Zhang 0001, Jianhua Tao 0001 |
Neural Networks | 3 |
| 2022 | Tucker decomposition-based temporal knowledge graph completion
Pengpeng Shao, Dawei Zhang 0001, Guohua Yang, Jianhua Tao 0001, Feihu Che |
Knowl. Based Syst. | 5 |
| 2021 | Self-supervised graph representation learning via bootstrapping
Feihu Che, Guohua Yang, Dawei Zhang 0001, Jianhua Tao 0001 |
Neurocomputing | 1 |
| 2021 | Multi-aspect self-supervised learning for heterogeneous information network
Feihu Che, Jianhua Tao 0001, Guohua Yang, Dawei Zhang 0001 |
Knowl. Based Syst. | 1 |
| 2020 | ParamE: Regarding Neural Network Parameters as Relation Embeddings for Knowledge Graph CompletionabstractWe study the task of learning entity and relation embeddings in knowledge graphs for predicting missing links. Previous translational models on link prediction make use of translational properties but lack enough expressiveness, while the convolution neural network based model (ConvE) takes advantage of the great nonlinearity fitting ability of neural networks but overlooks translational properties. In this paper, we propose a new knowledge graph embedding model called ParamE which can utilize the two advantages together. In ParamE, head entity embeddings, relation embeddings and tail entity embeddings are regarded as the input, parameters and output of a neural network respectively. Since parameters in networks are effective in converting input to output, taking neural network parameters as relation embeddings makes ParamE much more expressive and translational. In addition, the entity and relation embeddings in ParamE are from feature space and parameter space respectively, which is in line with the essence that entities and relations are supposed to be mapped into two different spaces. We evaluate the performances of ParamE on standard FB15k-237 and WN18RR datasets, and experiments show ParamE can significantly outperform existing state-of-the-art models, such as ConvE, SACN, RotatE and D4-STE/Gumbel. Feihu Che, Dawei Zhang 0001, Jianhua Tao 0001, Mingyue Niu, Bocheng Zhao |
AAAI | 1 |