VLDB 2026 Research / reviewers in the wild / expert
Junkai Li
dblp:239/2645
· DBLP profile ↗
13ranked-venue papers
2as first author
13since 2021 · last 2027
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 1 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Representation and self-supervised learning · 24% Vision and language · 22% Deep learning architectures and training · 16% | |
| Computer graphics and multimedia
1 paper |
Multimedia analysis and retrieval · 100% | |
| Human-computer interaction and pervasive computing
1 paper |
Wearable and physiological sensing · 100% |
Topics — the 11 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Representation and self-supervised learning
contrastive learning |
1.0 | 1 | 2026 | Chinese Two-part Allegorical Sayings Reading Comprehension: Exploration from Reasoning to Metaphor · AAAI 2026 |
Natural language and speech › Question answering and dialogue systems
machine reading comprehension |
1.0 | 1 | 2026 | Chinese Two-part Allegorical Sayings Reading Comprehension: Exploration from Reasoning to Metaphor · AAAI 2026 |
Natural language and speech › Information extraction and text analysis › natural language semantics › figurative language processing
metaphor understanding |
1.0 | 1 | 2026 | Chinese Two-part Allegorical Sayings Reading Comprehension: Exploration from Reasoning to Metaphor · AAAI 2026 |
Computer vision › Vision and language
multimodal reasoning |
1.0 | 1 | 2026 | MePe: Rethinking Multimodal Chinese Idiom Reading Comprehension from a Metaphorical Perspective · WWW 2026 |
Computer vision › Vision and language
multimodal understanding |
1.0 | 1 | 2026 | MePe: Rethinking Multimodal Chinese Idiom Reading Comprehension from a Metaphorical Perspective · WWW 2026 |
Machine learning › Representation and self-supervised learning › contrastive learning
multi-view contrastive learning |
1.0 | 1 | 2026 | Chinese Two-part Allegorical Sayings Reading Comprehension: Exploration from Reasoning to Metaphor · AAAI 2026 |
Multimedia analysis and retrieval › affective computing › sentiment analysis
multimodal sentiment analysis |
1.0 | 1 | 2026 | Beyond Words: Enhancing Desire, Emotion, and Sentiment Recognition with Non-Verbal Cues · WWW 2026 |
Natural language and speech › Language models and text generation
hallucination mitigation |
0.8 | 1 | 2024 | Citation-Enhanced Generation for LLM-based Chatbots · ACL (1) 2024 |
Natural language and speech › Language models and text generation
retrieval-augmented generation |
0.8 | 1 | 2024 | Citation-Enhanced Generation for LLM-based Chatbots · ACL (1) 2024 |
Machine learning › Deep learning architectures and training › transformer
spatio-temporal transformer |
0.8 | 1 | 2024 | Jointly Modeling Spatio-Temporal Features of Tactile Signals for Action Classification · AAAI 2024 |
Machine learning › Deep learning architectures and training
transformer |
0.8 | 1 | 2024 | Jointly Modeling Spatio-Temporal Features of Tactile Signals for Action Classification · AAAI 2024 |
Methods — techniques the papers use, named apart from their topics
temporal pretraining · 1.5spatial embeddings · 1.5reproducing kernel hilbert space · 1.0mixture of experts · 1.0maximum mean discrepancy · 1.0cross-projection · 1.0contrastive learning · 1.0temporal embeddings · 0.8temporal embedding · 0.8natural language inference · 0.8citation-enhanced generation · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | DFM-NC: Disentangling fine-grained multiplex non-literal cues for multimodal sentiment analysis
Tongguan Wang, Feiyue Xue, Junkai Li, Xiqiao Ba, Yang Xiao 0018, Ying Sha |
Expert Syst. Appl. | 4 |
| 2027 | SIPTrack: Reliability-aware identity prediction for sparse-interval pig multi-object tracking with a new benchmark
Feiyue Xue, Wangjun Huang, Junkai Li, Tongguan Wang, Huaiping Jin, Ying Sha |
Expert Syst. Appl. | 3 |
| 2026 | Chinese Two-part Allegorical Sayings Reading Comprehension: Exploration from Reasoning to MetaphorabstractThe Two-Part Allegorical Saying (TPAS) is a Chinese linguistic phenomenon with a riddle-explanation structure, and an important component of Chinese metaphors. Existing research has primarily used TPAS to assist other semantic tasks, but lacks in-depth exploration of its intrinsic mechanisms: semantic rhetoric, logical reasoning, and metaphorical expression. To address this gap, we construct the first Chinese TPAS Reading Comprehension dataset (CTRC), which contains 18,103 TPASs and 75,296 passages. We frame it as a cloze test where the model selects the most suitable TPAS from candidates to fill passage blanks. To tackle the challenges of this CTRC task, we propose a Multi-view TPAS Contrastive Learning Network (MTCLN). Firstly, the joint vector cross-projection module extracts the rhetorical features of TPAS, such as homophonic puns, through vector space mapping to mitigate the semantic deviations caused by rhetoric. Then, the softened contrastive learning module strengthens the modeling of TPAS logical reasoning through feature association. Finally, the multi-view feature fusion module integrates contextual semantics with diverse TPAS features to facilitate the understanding of metaphorical expressions. Experiments on the CTRC dataset demonstrate that MTCLN achieves an average accuracy of 67.47%, outperforming large language models by 25.48%. Dongyu Su, Yimin Xiao, Tongguan Wang, Feiyue Xue, Junkai Li, Ying Sha |
AAAI | 5 |
| 2026 | RISCTrade: A Dynamic Multi-Market Benchmark for Risk-Mandate Compliance of Large Language Models in Closed-Loop Trading
Wenliang Huang, Junkai Li, Zengyi Yu, Xiangjie Kong |
ICIC (7) | 2 |
| 2026 | Beyond Words: Enhancing Desire, Emotion, and Sentiment Recognition with Non-Verbal Cues
Tongguan Wang, Feiyue Xue, Junkai Li, Ying Sha |
WWW | 4 |
| 2026 | MePe: Rethinking Multimodal Chinese Idiom Reading Comprehension from a Metaphorical PerspectiveabstractThe multimodal Chinese idiom reading comprehension task aims to select the most appropriate idiom from a candidate list via the given text and image. This poses a significant challenge for the model to comprehend each Chinese idiom accurately. Existing multimodal Chinese idiom reading comprehension methods primarily focus on aligning contextual text and images, while overlooking two key attributes of Chinese idioms.(1) There is a discrepancy between the literal and metaphorical meanings of Chinese idioms. (2) The same Chinese idiom has different meanings in different scenarios, which requires targeted understanding by experts who specialize in different fields. To address the above challenges, we rethink the solution to the multimodal idiom reading comprehension task from a metaphorical perspective and propose a framework named MePe. Firstly, we propose a literal metaphorical semantic graph that systematically transforms the implicit discrepancy between the literal and metaphorical meanings of Chinese idioms into structured explicit relationships, thereby making metaphorical meanings more understandable. Then, we propose a mixture of idiom experts consisting of a literal idiom expert and a metaphorical idiom expert. Through division of labor and collaboration among these experts, we achieve an understanding of the dual meanings of Chinese idioms across different scenarios. Finally, we employ the maximum mean discrepancy to adjust the variance between the literal and metaphorical semantic features of Chinese idioms. By mapping these features onto a shared reproducing kernel Hilbert space, the model can better distinguish between the two based on contextual clues. Extensive experiments demonstrate that MePe achieves state-of-the-art performance on the MChIRC dataset. Tongguan Wang, Junkai Li, Feiyue Xue, Dongyu Su, Wangjun Huang, Ying Sha |
WWW | 2 |
| 2026 | SCA-Net: Semantic text-enhanced context-aware multimodal framework for fish feeding assessment in aquaculture
Junkai Li, Feiyue Xue, Tongguan Wang, Chunfang Wang, Zongyao Sha, Ying Sha |
Expert Syst. Appl. | 1 |
| 2026 | BIG-TM: Bridging Individual Guidance with Trifusion MoPoE for Chinese memes understanding
Tongguan Wang, Junkai Li, Feiyue Xue, Dongyu Su, Guixin Su, Xiaopeng Wen, Ying Sha |
Knowl. Based Syst. | 2 |
| 2026 | Improved Spontaneous EEG Signal Decoding Efficiency by Function Predefined Convolutional Neural NetworkabstractA spontaneous electroencephalogram (EEG)-based brain-computer interface (BCI) is an ideal form of brain-computer interaction. The classical decoding methods can achieve classification by using meaningful manual features, but their performance is poor. The neural network (NN) methods have significantly improved the performance, but their interpretability and computational efficiency are much lower than those of the classical methods. This is because NN abandons the strong a priori knowledge of neuroscience and completely relies on training to extract EEG features. How to integrate the characteristics of neural signals into the design of the basic operator of the NNs while retaining its learning ability is the focus of this work. In this work, we proposed a function predefined convolutional NN (FPCNN) to search for the best frequency points and channel weights to decode spontaneous EEG signals. Among the FPCNN, a novel function predefined convolutional (FPC) layer adopts a learnable way to search for the key spatial-frequency parameters of spontaneous EEG, making its parameters have clear physical meanings. Furthermore, a trainable quadrature detector (TQD) based on FPC was constructed, and the quadrature characteristic was utilized to ensure the capture of complex phase change signals. The core contribution of our method lies in the proposal of a novel NN operator for decoding spontaneous EEG, and a quadrature scheme for handling the phase changes of signals. The experimental results show that the proposed FPCNN significantly improves the performance by 2.09% ( ${}^{\ast } $ ), 3.08% ( ${}^{\ast } $ ), and 3.41% ( ${}^{\ast \ast }$ ), respectively, compared with the state-of-the-art (SOTA) methods on three spontaneous EEG datasets. Moreover, the training and testing time cost of FPCNN in a non-GPU environment only takes 67.96 and 19.36 s per epoch. Its savings in computing resources and time are very beneficial for EEG processing in diverse environments. In addition, visualization experiments demonstrated the interpretability and stability of the proposed FPCNN. The experimental results show that our method is efficient, stable, and interpretable. This work has effectively improved the decoding efficiency of spontaneous EEG signals and demonstrated the power of combining traditional signal processing methods with NNs. Boxun Fu, Fu Li 0002, Junkai Li, Youshuo Ji, Yang Li 0019, Yinghui Quan, Lijian Zhang, Guangming Shi |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2025 | RCLMuFN: Relational context learning and multiplex fusion network for multimodal sarcasm detection
Tongguan Wang, Junkai Li, Guixin Su, Yongcheng Zhang, Dongyu Su, Yuxue Hu, Ying Sha |
Knowl. Based Syst. | 2 |
| 2024 | Jointly Modeling Spatio-Temporal Features of Tactile Signals for Action ClassificationabstractTactile signals collected by wearable electronics are essential in modeling and understanding human behavior. One of the main applications of tactile signals is action classification, especially in healthcare and robotics. However, existing tactile classification methods fail to capture the spatial and temporal features of tactile signals simultaneously, which results in sub-optimal performances. In this paper, we design Spatio-Temporal Aware tactility Transformer (STAT) to utilize continuous tactile signals for action classification. We propose spatial and temporal embeddings along with a new temporal pretraining task in our model, which aims to enhance the transformer in modeling the spatio-temporal features of tactile signals. Specially, the designed temporal pretraining task is to differentiate the time order of tubelet inputs to model the temporal properties explicitly. Experimental results on a public action classification dataset demonstrate that our model outperforms state-of-the-art methods in all metrics. Jimmy Lin, Junkai Li, Jiasi Gao, Weizhi Ma |
AAAI | 2 |
| 2024 | Citation-Enhanced Generation for LLM-based ChatbotsabstractLarge language models (LLMs) exhibit powerful general intelligence across diverse scenarios, including their integration into chatbots.However, a vital challenge of LLMbased chatbots is that they may produce hallucinated content in responses, which significantly limits their applicability.Various efforts have been made to alleviate hallucination, such as retrieval augmented generation and reinforcement learning with human feedback, but most of them require additional training and data annotation.In this paper, we propose a novel post-hoc Citation-Enhanced Generation (CEG) approach combined with retrieval argumentation.Unlike previous studies that focus on preventing hallucinations during generation, our method addresses this issue in a post-hoc way.It incorporates a retrieval module to search for supporting documents relevant to the generated content, and employs a natural language inference-based citation generation module.Once the statements in the generated content lack of reference, our model can regenerate responses until all statements are supported by citations.Note that our method is a training-free plug-and-play plugin that is capable of various LLMs.Experiments on various hallucination-related datasets show our framework outperforms state-of-theart methods in both hallucination detection and response regeneration on three benchmarks. Junkai Li, Weizhi Ma |
ACL (1) | 2 |
| 2024 | Efficient Guided Query Network for Human-Object Interaction DetectionabstractRecently, Transformer-based one-stage methods have demonstrated excellent efficiency in Human-Object Interaction (HOI) tasks. However, these methods often utilize semantically ambiguous initial queries, thus constraining the model’s ability for set prediction. In addition, currently widely used HOI datasets suffer from long-tail distribution issues, so accurately identifying rare interaction categories remains challenging. To address these challenges, we propose an Efficient Guided Query Network (EGQ-Net). The network introduces a forward-guided relational queries approach, which accurately captures the triplets of interaction relationships by effectively integrating the initial queries predicted by the encoder and the output features of each decoder layer. Furthermore, we used the visual language pre-training models CLIP and BLIP2 to design interaction position query guidance and interaction content query guidance to achieve accurate recognition and localization of interactive areas by queries. Experimental results demonstrate that our proposed method achieves state-of-the-art performance on widely used HOI benchmarks (V-COCO and HICO-DET). Junkai Li, Huicheng Lai, Tongguan Wang, Hutuo Quan, Dongji Chen |
ICME | 1 |