EDBT 2026 Demo / reviewers in the wild / expert
Yong Dai 0001
dblp:147/7764-1
· DBLP profile ↗
24ranked-venue papers
2as first author
22since 2021 · last 2026
0000-0002-3041-5851ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 2 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 4 since 2021Security and privacy · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | E-ViC: Reasoning Beyond Text via Embodied Visual Chain for Spatial IntelligenceabstractJunbo Qi, Yi Zhang, Hanchu Ni, Che Liu, Zhimin Yao, Ruilin Yang, Xiancong Ren, Liangjian Wen, Wei Ge, Yuya Ieiri, Osamu Yoshie, Yong Dai, Xiaozhu Ju. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Junbo Qi, Yi Zhang 0001, Hanchu Ni, Che Liu 0002, Zhimin Yao, Ruilin Yang, Xiancong Ren, Liangjian Wen, Yuya Ieiri, Osamu Yoshie, Yong Dai 0001, Xiaozhu Ju |
ACL (1) | 12 |
| 2026 | GazeCLIP: Enhancing gaze estimation through text-guided multimodal learning
Jun Wang 0089, Hao Ruan, Liangjian Wen, Yong Dai 0001, Mingjie Wang 0002 |
Neurocomputing | 4 |
| 2026 | MRDNet: Multivariable Relational Decomposition Network for Multivariate Time Series Forecasting
Ao Hu, Liangjian Wen, Yong Dai 0001, Dongkai Wang, Jun Wang 0089, Jiang Duan |
Knowl. Based Syst. | 4 |
| 2026 | TimeCNN: Refining inscross-variable interaction on time point for time series forecasting
Ao Hu, Liangjian Wen, Yong Dai 0001, Shiyi Qi, Jun Wang 0089, Xun Zhou 0001, Dongkai Wang, Zenglin Xu, Jiang Duan |
Neural Networks | 3 |
| 2026 | FDNet: High-frequency disentanglement network with information-theoretic guidance for multivariate time series forecasting
Ao Hu, Liangjian Wen, Jiang Duan, Yong Dai 0001, Dongkai Wang, Shudong Huang, Jun Wang 0089, Zenglin Xu |
Pattern Recognit. | 4 |
| 2026 | VLCounting: Taming zero-shot counting via language-driven exemplar grounding
Mingjie Wang 0002, Yong Dai 0001, Eric Buys, Minglun Gong |
Pattern Recognit. | 3 |
| 2026 | TencentLLMEval: A Hierarchical Evaluation of Real-World Capabilities for Human-Aligned LLMsabstractLarge language models (LLMs) have shown impressive capabilities across various natural language tasks. However, evaluating their alignment with human preferences remains a challenge. To this end, we propose a comprehensive human evaluation framework to assess LLMs’ proficiency in following instructions on diverse real-world tasks. We construct a hierarchical task tree encompassing seven major areas covering over 200 categories and over 800 tasks, which covers diverse capabilities such as question answering, reasoning, multi-turn dialogue, and text generation, to evaluate LLMs in a comprehensive and in-depth manner. We also design detailed evaluation standards and processes to facilitate consistent, unbiased judgments from human evaluators. A test set of over 3,000 instances is released, spanning different difficulty levels and knowledge domains. Our work provides a standardized methodology to evaluate human alignment in LLMs for both English and Chinese. We also analyze the feasibility of automating parts of evaluation with a strong LLM (GPT-4). Our framework supports a thorough assessment of LLMs as they are integrated into real-world applications. We have made publicly available the task tree, TencentLLMEval dataset, and evaluation methodology which have been demonstrated as effective in assessing the performance of Tencent Hunyuan LLMs. By doing so, we aim to facilitate the benchmarking of advances in the development of safe and human-aligned LLMs. Shuyi Xie, Wenlin Yao, Yong Dai 0001, Zishan Xu, Fan Lin, Donglin Zhou, Lifeng Jin, Xinhua Feng, Pengzhi Wei, Zhichao Hu, Dong Yu 0001, Zhengyou Zhang |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2025 | NEXUS-O: An Omni-Perceptive and -Interactive Model for Language, Audio, and VisionabstractHuman beings perceive the real world through a spectrum of sensory modalities, encompassing auditory, visual, and linguistic faculties. This work proposes an industry-level omni-modal large language model (LLM) pipeline that integrates auditory, visual, and linguistic modalities to overcome challenges such as limited tri-modal datasets, high computational costs, and complex feature alignments. Our pipeline consists of three main components: First, a modular, end-to-end framework enabling flexible configuration of various encoder-LLM-decoder architectures. Second, a lightweight training strategy that pre-trains audio-language alignment on the state-of-the-art vision-language model Qwen2.5-VL, thus avoiding the costly pre-training of vision-specific modalities. Third, an audio synthesis pipeline that generates high-quality audio-text data from diverse real-world scenarios, supporting applications such as Automatic Speech Recognition and Speech-to-Speech chat. To this end, we introduce an industry-level omni-modal LLM, NEXUS-O. Extensive experiments validate the efficacy of our pipeline, yielding the following key findings: (1) In the visual understanding task, NEXUSO exhibits superior performance compared with its backbone model - Qwen2.5-VL-7B, validating the efficiency of our training strategy. (2) Within the English Spoken Question-Answering task, the model achieves better accuracy than the same-period competitor (i.e, MiniCPM-o2.6-7B) in the LLaMA Q. benchmark. (3) In our realworld ASR testset, NEXUS-O achieves outstanding performance, indicating its robustness in real scenarios. (4) In the Speech-to-Text Translation task, our model outperforms Qwen2-Audio-Instruct-7B. (5) In the Text-to-Speech task, based on pretrained vocoder (e.g., Fishspeech1.4 or CosyVoice2.0), NEXUS-O is comparable to its backbone vocoder on Seed-TTS benchmark. (6) An in-depth analysis of tri-modal alignment reveals that incorporating the audio modality enhances representational alignment between vision and language. Che Liu 0002, Yingji Zhang, Chenggong Gong, Yu Lu 0020, Shilin Zhou 0002, Ziliang Gan, Haipang Wu, Ji Liu 0003, André Freitas, Qifan Wang 0001, Zenglin Xu, Rongjunchen Zhang, Yong Dai 0001 |
ACM Multimedia | 16 |
| 2025 | MME-Finance: A Multimodal Finance Benchmark for Expert-level Understanding and ReasoningabstractTo date, there is a notable lack of rigorous benchmarks that assess Multimodal Large Language Models (MLLMs) within the financial domain, a field characterized by specialized financial charts and complex domain-specific expertise. To address this gap, we introduce MME-Finance, the first comprehensive bilingual multimodal benchmark tailored for financial analysis. MME-Finance comprises 4,751 meticulously curated samples, encompassing 2,274 open-ended questions, 2,000 binary-choice questions, and 477 multi-turn questions. To mitigate bias when LLMs act as judges, we also created an evaluation framework that strengthens alignment with human judgments by embedding visual context into the multimodal assessment pipeline. A comprehensive evaluation of 31 popular MLLMs has been conducted to assess their perception, reasoning, and cognitive capabilities. Gemini2.5Pro achieves highest accuracy of 79.28% and 85.71% on the open-ended questions and multi-turn questions, respectively. Among open-source models, InternVL3-78B attains 71.24 % accuracy on the open-ended question, whereas Qwen2.5-VL-72B achieves an F1 score of 88.73 % on the binary-choice question. The results indicate that state-of-the-art MLLMs demonstrate considerable overall competence, yet exhibit significant deficiencies in fine-grained visual perception and the understanding of domain-specific financial images. Source code is available at https://github.com/HiThink-Research/MME-Finance. Ziliang Gan, Haohan Li, Xueyuan Lin, Ji Liu 0003, Haipang Wu, Chaoyou Fu, Zenglin Xu, Rongjunchen Zhang, Yong Dai 0001 |
ACM Multimedia | 11 |
| 2025 | Distribution-Aligned Decoding for Efficient LLM Task AdaptationabstractAdapting billion-parameter language models to a downstream task is still costly, even with parameter-efficient fine-tuning (PEFT). We re-cast task adaptation as output-distribution alignment: the objective is to steer the output distribution toward the task distribution directly during decoding rather than indirectly through weight updates. Building on this view, we introduce Steering Vector Decoding (SVDecode), a lightweight, PEFT-compatible, and theoretically grounded method. We start with a short warm-start fine-tune and extract a task-aware steering vector from the Kullback-Leibler (KL) divergence gradient between the output distribution of the warm-started and pre-trained models. This steering vector is then used to guide the decoding process to steer the model's output distribution towards the task distribution. We theoretically prove that SVDecode is first-order equivalent to the gradient step of full fine-tuning and derive a globally optimal solution for the strength of the steering vector. Across three tasks and nine benchmarks, SVDecode paired with four standard PEFT methods improves multiple-choice accuracy by up to 5 percentage points and open-ended truthfulness by 2 percentage points, with similar gains (1-2 percentage points) on commonsense datasets without adding trainable parameters beyond the PEFT adapter. SVDecode thus offers a lightweight, theoretically grounded path to stronger task adaptation for large language models. Senkang Hu, Jinqi Jiang, Yihang Tao, Yong Dai 0001, Sam Kwong, Yuguang Fang |
NeurIPS | 6 |
| 2025 | InfMasking: Unleashing Synergistic Information by Contrastive Multimodal InteractionsabstractIn multimodal representation learning, synergistic interactions between modalities not only provide complementary information but also create unique outcomes through specific interaction patterns that no single modality could achieve alone. Existing methods may struggle to effectively capture the full spectrum of synergistic information, leading to suboptimal performance in tasks where such interactions are critical. This is particularly problematic because synergistic information constitutes the fundamental value proposition of multimodal representation. To address this challenge, we introduce InfMasking, a contrastive synergistic information extraction method designed to enhance synergistic information through an Infinite Masking strategy. InfMasking stochastically occludes most features from each modality during fusion, preserving only partial information to create representations with varied synergistic patterns. Unmasked fused representations are then aligned with masked ones through mutual information maximization to encode comprehensive synergistic information. This infinite masking strategy enables capturing richer interactions by exposing the model to diverse partial modality combinations during training. As computing mutual information estimates with infinite masking is computationally prohibitive, we derive an InfMasking loss to approximate this calculation. Through controlled experiments, we demonstrate that InfMasking effectively enhances synergistic information between modalities. In evaluations on large-scale real-world datasets, InfMasking achieves state-of-the-art performance across seven benchmarks. Code is released at https://github.com/brightest66/InfMasking. Liangjian Wen, Qun Dai, Jianzhuang Liu, Jiangtao Zheng, Yong Dai 0001, Dongkai Wang, Zhao Kang 0001, Jun Wang 0089, Zenglin Xu, Jiang Duan |
NeurIPS | 5 |
| 2025 | ConceptPsy: A comprehensive benchmark suite for hierarchical psychological concept understanding in LLMs
Junlei Zhang, Hongliang He 0002, Lizhi Ma, Nirui Song, Shuyuan He, Huachuan Qiu, Zhanchao Zhou, Anqi Li 0002, Yong Dai 0001, Renjun Xu, Zhen-Zhong Lan |
Neurocomputing | 10 |
| 2024 | WebVoyager: Building an End-to-End Web Agent with Large Multimodal ModelsabstractHongliang He, Wenlin Yao, Kaixin Ma, Wenhao Yu, Yong Dai, Hongming Zhang, Zhenzhong Lan, Dong Yu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Hongliang He 0002, Wenlin Yao, Kaixin Ma, Wenhao Yu 0002, Yong Dai 0001, Hongming Zhang 0009, Zhen-Zhong Lan, Dong Yu 0001 |
ACL (1) | 5 |
| 2024 | Chunk, Align, Select: A Simple Long-sequence Processing Method for TransformersabstractAlthough dominant in natural language processing, transformer-based models still struggle with long-sequence processing, due to the computational costs of their self-attention operations, which increase exponentially as the length of the input sequence grows.To address this challenge, we propose a Simple framework to enhance the long-content processing of off-the-shelf pre-trained transformers via three steps: Chunk, Align, and Select (SimCAS).More specifically, we first divide each longsequence input into a batch of chunks, then align the inter-chunk information during the encoding steps, and finally, select the most representative hidden states from the encoder for the decoding process.With our SimCAS, the computation and memory costs can be reduced to linear complexity.In experiments, we demonstrate the effectiveness of the proposed method on various real-world long-text summarization and reading comprehension tasks, in which SimCAS significantly outperforms prior longsequence processing baselines.The code is at https://github.com/xjw-nlp/SimCAS. Jiawen Xie, Pengyu Cheng, Yong Dai 0001 |
ACL (1) | 4 |
| 2024 | SkillNet-X: A Multilingual Multitask Model with Sparsely Activated SkillsabstractTraditional multitask learning methods typically can only leverage shared knowledge within specific tasks or languages, resulting in a loss of either cross-language or cross-task knowledge. This paper proposes a general multilingual multitask model, named SkillNet-X, which enables a single model to tackle many different tasks from different languages. To this end, we define several language-specific skills and task-specific skills, each of which corresponds to a skill module. SkillNet-X sparsely activates parts of the skill modules which are relevant to eitherthe target task or the target language. Acting as knowledge transit hubs, skill modules are capable of absorbing task-related knowledge and language-related knowledge consecutively. We evaluate SkillNet-X on eleven natural language understanding datasets in four languages. Results show that SkillNet-X performs better than task-specific and two multitask learning baselines.To investigate the generalization of our model, we conduct experiments on two new tasks and find that SkillNet-X significantly outperforms baselines. Zhangyin Feng, Yong Dai 0001, Fan Zhang 0092, Duyu Tang, Shuangzhi Wu, Bing Qin 0001, Yunbo Cao, Shuming Shi 0001 |
ICASSP | 2 |
| 2024 | Self-playing Adversarial Language Game Enhances LLM ReasoningabstractWe explore the potential of self-play training for large language models (LLMs) in a two-player adversarial language game called Adversarial Taboo. In this game, an attacker and a defender communicate around a target word only visible to the attacker. The attacker aims to induce the defender to speak the target word unconsciously, while the defender tries to infer the target word from the attacker's utterances. To win the game, both players must have sufficient knowledge about the target word and high-level reasoning ability to infer and express in this information-reserved conversation. Hence, we are curious about whether LLMs' reasoning ability can be further enhanced by Self-Playing this Adversarial language Game (SPAG). With this goal, we select several open-source LLMs and let each act as the attacker and play with a copy of itself as the defender on an extensive range of target words. Through reinforcement learning on the game outcomes, we observe that the LLMs' performances uniformly improve on a broad range of reasoning benchmarks. Furthermore, iteratively adopting this self-play process can continuously promote LLMs' reasoning abilities. The code is available at https://github.com/Linear95/SPAG. Pengyu Cheng, Tianhao Hu, Zhisong Zhang, Yong Dai 0001 |
NeurIPS | 5 |
| 2024 | IDGen: Item Discrimination Induced Prompt Generation for LLM EvaluationabstractAs Large Language Models (LLMs) become more capable of handling increasingly complex tasks, the evaluation set must keep pace with these advancements to ensure it remains sufficiently discriminative. Item Discrimination (ID) theory, which is widely used in educational assessment, measures the ability of individual test items to differentiate between high and low performers. Inspired by this theory, we propose an ID-induced prompt synthesis framework for evaluating LLMs so that the evaluation set continually updates and refines according to model abilities.
Our data synthesis framework prioritizes both breadth and specificity. It can generate prompts that comprehensively evaluate the capabilities of LLMs while revealing meaningful performance differences between models, allowing for effective discrimination of their relative strengths and weaknesses across various tasks and domains.
To produce high-quality data, we incorporate a self-correct mechanism into our generalization framework and develop two models to predict prompt discrimination and difficulty score to facilitate our data synthesis framework, contributing valuable tools to evaluation data synthesis research. We apply our generated data to evaluate five SOTA models. Our data achieves an average score of 51.92, accompanied by a variance of 10.06. By contrast, previous works (i.e., SELF-INSTRUCT and WizardLM) obtain an average score exceeding 67, with a variance below 3.2.
The results demonstrate that the data generated by our framework is more challenging and discriminative compared to previous works.
We will release a dataset of over 3,000 carefully crafted prompts to facilitate evaluation research of LLMs. Fan Lin, Shuyi Xie, Yong Dai 0001, Wenlin Yao, Tianjiao Lang, Yu Zhang 0004 |
NeurIPS | 3 |
| 2023 | MarkBERT: Marking Word Boundaries Improves Chinese BERT
Linyang Li, Yong Dai 0001, Duyu Tang, Xipeng Qiu, Shuming Shi 0001 |
NLPCC (1) | 2 |
| 2023 | Dynamic Mixture of Counter Network for Location-Agnostic Crowd CountingabstractCrowd counting has attracted increasing attentions in recent years due to its challenges and wide societal applications. Despite persevering efforts made by the research community, most of existing methods require a large amount of location-level annotations. Collecting such type of fine-granularity supervisory signals is extremely time-consuming and labour-intensive, thereby hindering the well generalization of these location-adherent models. To shun this drawback, several pioneering studies open a promising research direction of location-agonistic crowd counting. Albeit the noticeable efforts, they somewhat ignore the merits of diverse learning paradigms and the issue of intractable density shift. To ameliorate these issues, in this paper, a novel Dynamic Mixture of Counter Network (DMCNet) is proposed for location-agnostic crowd counting. Specifically, our DMCNet inherits the hybrid advantages of CNNs (e.g. locality-oriented and pyramidal property) and MLP-based structure (e.g. global receptive fields and light weight). Particularly, the dynamic counter predictor and the mixture of counter heads are delicately designed to hammer at combating huge density shift and overfitting. Extensive experiments demonstrate that our DMCNet attains state-of-the-art performance against existing location-agnostic approaches and performs on par with many conventional location-adherent ones. Mingjie Wang 0002, Hao Cai 0004, Yong Dai 0001, Minglun Gong |
WACV | 3 |
| 2022 | Exploring and Adapting Chinese GPT to Pinyin Input MethodabstractMinghuan Tan, Yong Dai, Duyu Tang, Zhangyin Feng, Guoping Huang, Jing Jiang, Jiwei Li, Shuming Shi. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Minghuan Tan, Yong Dai 0001, Duyu Tang, Zhangyin Feng, Guoping Huang, Jing Jiang 0001, Shuming Shi 0001 |
ACL (1) | 2 |
| 2022 | Graph Fusion Network for Text Classification
Yong Dai 0001, Linjun Shou, Ming Gong 0001, Xiaolin Xia, Zhao Kang 0001, Zenglin Xu, Daxin Jiang |
Knowl. Based Syst. | 1 |
| 2021 | Contextualize Knowledge Bases with Transformer for End-to-end Task-Oriented Dialogue SystemsabstractIncorporating knowledge bases (KB) into endto-end task-oriented dialogue systems is challenging, since it requires to properly represent the entity of KB, which is associated with its KB context and dialogue context.The existing works represent the entity with only perceiving a part of its KB context, which can lead to the less effective representation due to the information loss, and adversely favor KB reasoning and response generation.To tackle this issue, we explore to fully contextualize the entity representation by dynamically perceiving all the relevant entities and dialogue history.To achieve this, we propose a COntextaware Memory Enhanced Transformer framework (COMET), which treats the KB as a sequence and leverages a novel Memory Mask to enforce the entity to only focus on its relevant entities and dialogue history, while avoiding the distraction from the irrelevant entities.Through extensive experiments, we show that our COMET framework can achieve superior performance over the state of the arts. Yanjie Gou, Yinjie Lei, Lingqiao Liu, Yong Dai 0001, Chunxu Shen |
EMNLP (1) | 4 |
| 2020 | Adversarial Training Based Multi-Source Unsupervised Domain Adaptation for Sentiment AnalysisabstractMulti-source unsupervised domain adaptation (MS-UDA) for sentiment analysis (SA) aims to leverage useful information in multiple source domains to help do SA in an unlabeled target domain that has no supervised information. Existing algorithms of MS-UDA either only exploit the shared features, i.e., the domain-invariant information, or based on some weak assumption in NLP, e.g., smoothness assumption. To avoid these problems, we propose two transfer learning frameworks based on the multi-source domain adaptation methodology for SA by combining the source hypotheses to derive a good target hypothesis. The key feature of the first framework is a novel Weighting Scheme based Unsupervised Domain Adaptation framework ((WS-UDA), which combine the source classifiers to acquire pseudo labels for target instances directly. While the second framework is a Two-Stage Training based Unsupervised Domain Adaptation framework (2ST-UDA), which further exploits these pseudo labels to train a target private extractor. Importantly, the weights assigned to each source classifier are based on the relations between target instances and source domains, which measured by a discriminator through the adversarial training. Furthermore, through the same discriminator, we also fulfill the separation of shared features and private features.Experimental results on two SA datasets demonstrate the promising performance of our frameworks, which outperforms unsupervised state-of-the-art competitors. Yong Dai 0001, Xiancong Ren, Zenglin Xu |
AAAI | 1 |
| 2013 | Remote Sensing Images Super-resolution Based on Sparse Dictionaries and Residual DictionariesabstractIn this paper, a sensing image super-resolution (SR) reconstruction method is proposed. Sparse dictionary dealing with remote sensing image SR problem is introduced in this work. The sparse dictionary is based on a sparsity model where the dictionary atoms have sparse representation over a basic dictionary. The sparse dictionary consists of two parts: basic dictionary and atom representation matrix. The sparse dictionary leads to compact representation and it is both adaptive and efficient. Furthermore, compared with conventional SR methods, two dictionary pairs, i.e. primitive sparse dictionary pair and residual sparse dictionary pair, are proposed. The primitive sparse dictionary pair is learned to reconstruct initial high-resolution (HR) remote sensing image from a single low-resolution (LR) input. However, the initial HR remote sensing image loses some details compare with the corresponding original HR image completely. Therefore, residual sparse dictionary pair is learned to reconstruct residual information. The proposed method is tested on remote sensing images, and the experimental results indicate that the proposed algorithm can provide substantial improvement in resolution of remote sensing images, and the results are superior in quality to the results produced by other methods. Wei Wu 0002, Yong Dai 0001, Xiaomin Yang, Binyu Yan, Wei Lu 0021 |
DASC | 3 |