EDBT 2026 Demo / reviewers in the wild / expert
Yanyang Li
dblp:208/4741
· DBLP profile ↗
26ranked-venue papers
9as first author
19since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 22 · 8 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 3 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ConsistEAE: Enhancing low-resource event argument extraction with linguistically consistent demonstrations
Yikai Guo, Xuemeng Tian, Bin Ge 0006, Wenjun Ke 0002, Yanyang Li, Haoran Luo 0001 |
Neurocomputing | 8 |
| 2026 | EnsDiffAD: Ensemble Diffusion Models for Multivariate Time Series Anomaly Detection
Qian Ma 0003, Yanyang Li, Mei Bai, Xite Wang, Shikai Guo, Yu Gu 0002, Ge Yu 0001 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2025 | Learning to Reason from Feedback at Test-TimeabstractSolving complex tasks in a single attempt is challenging for large language models (LLMs). Iterative interaction with the environment and feedback is often required to achieve success, making effective feedback utilization a critical topic. Existing approaches either struggle with length generalization or rely on naive retries without leveraging prior information. In this paper, we introduce FTTT, a novel paradigm that formulates feedback utilization as an optimization problem at test time. Additionally, we propose a learnable test-time optimizer, OpTune, to effectively exploit feedback. Experiments on two LLMs across four reasoning datasets demonstrate that FTTT and OpTune achieve superior scalability and performance. Yanyang Li, Michael R. Lyu, Liwei Wang 0009 |
ACL (1) | 1 |
| 2025 | Learning from Videos for 3D World: Enhancing MLLMs with 3D Vision Geometry PriorsabstractPrevious research has investigated the application of Multimodal Large Language Models (MLLMs) in understanding 3D scenes by interpreting them as videos. These approaches generally depend on comprehensive 3D data inputs, such as point clouds or reconstructed Bird's-Eye View (BEV) maps. In our research, we advance this field by enhancing the capability of MLLMs to understand and reason in 3D spaces directly from video data, without the need for additional 3D input. We propose a novel and efficient method called the Video-3D Geometry Large Language Model (VG LLM). Our approach utilizes a 3D visual geometry encoder to extract 3D prior information from video sequences. This information is then integrated with visual tokens and input into the MLLM. Extensive experiments have shown that our method has achieved substantial improvements in various tasks related to 3D scene understanding and spatial reasoning, all directly learned from video sources. Impressively, our 4B model, which does not rely on explicit 3D data inputs, achieves competitive results compared to existing state-of-the-art methods, and even surpasses the Gemini-1.5-Pro in the VSI-Bench evaluations. Duo Zheng, Shijia Huang, Yanyang Li, Liwei Wang 0009 |
NeurIPS | 3 |
| 2025 | Connection Enhancement for Meta-Computing in IIoT: A Credit-Aware Cooperative D2D Transmission ApproachabstractUsers’ connection is pivotal for advancing meta-computing in Industrial Internet of Things (IIoT), where efficient data transmission can help overcome the barriers of computing resources distributed among billions of devices from diverse IIoT networks. Cooperative device-to-device (D2D) transmission, which is underlaid with cellular networks, plays an important role in enhancing the users’ connections for IIoT. However, since serving as a content provider in cooperative transmission needs consuming resources, IoT devices are typically willing to participate in cooperative transmission only when the content requesters are within the same IIoT networks. To encourage the cooperation across different IIoT networks, we in this article introduce social credit as a universal virtual concurrency to stimulate users’ willingness to assist content requesters via D2D transmission. Specifically, we first establish the relationship between social credit and data transmission, and model the process of credit acquisition and expenditure using a queue, which is initially unstable. Then, the inversely queuing technique is adopted to construct an equivalent stable queue, and a statistical credit guarantee mechanism is proposed to maintain cooperative transmission. Based on this mechanism, we explore the throughput maximization problem for cooperative D2D transmission in IIoT underlaying cellular networks, considering constraints on average and peak transmit power as well as interference to cellular communications, for obtaining the optimal credit-aware power control scheme. When multiple content providers are available, we prioritize the provider located at the shortest distance and examine the corresponding credit-aware power minimization problem, where the optimal power control scheme is obtained. Simulation results demonstrate that the proposal offers more stable credit guarantees than baseline methods, encouraging greater participation in cooperative transmission while achieving higher throughput and reducing transmit power consumption. Jianquan Wang 0002, Yuquan Xiao, Qinghe Du, Yanyang Li, Xuejie Zhu, Likang Zhang |
IEEE Internet Things J. | 4 |
| 2024 | Making Long-Context Language Models Better Multi-Hop ReasonersabstractRecent advancements in long-context modeling have enhanced language models (LMs) for complex tasks across multiple NLP applications.Despite this progress, we find that these models struggle with multi-hop reasoning and exhibit decreased performance in the presence of noisy contexts.In this paper, we introduce Reasoning with Attributions, a novel approach that prompts LMs to supply attributions for each assertion during their reasoning.We validate our approach through experiments on three multi-hop datasets, employing both proprietary and open-source models, and demonstrate its efficacy and resilience.Furthermore, we explore methods to augment reasoning capabilities via fine-tuning and offer an attribution-annotated dataset and a specialized training strategy.Our fine-tuned model achieves competitive performance on multi-hop reasoning benchmarks, closely paralleling proprietary LMs such as ChatGPT and Claude-instant 1 . Yanyang Li, Shuo Liang, Michael R. Lyu, Liwei Wang 0009 |
ACL (1) | 1 |
| 2024 | Analysis of the Discriminability of Three Types of CMR Image Features for CardiomyopathyabstractTo assess the efficacy of point cloud features, conventional image indices, and radiomics signatures from cardiovascular magnetic resonance (CMR) images, as well as their combinations in distinguishing hypertrophic cardiomyopathy (HCM) and dilated cardiomyopathy (DCM) patients and normal (NOR) (healthy) subjects. A total of 452 participants (142 HCM, 157 DCM, 153 NOR) from two public datasets were analyzed. Features were extracted from the left ventricle (LV), right ventricle (RV), and myocardium (MYO) in end-diastolic (ED) and end-systolic (ES) phases, including 78 point cloud, 84 radiomics, and 20 image indices. Feature selection and SVM classification were used to create discriminative signatures. Reproducibility was assessed with a 20% training set and full test set. The combined feature model with all three types yielded the highest accuracy (92.3%, AUC 0.977), followed by point cloud and the combination of point cloud and image features (90.2% accuracy). Selected features had high repeatability (ICC ≥ 0.85), offering a detailed multi-angle view of differences among HCM, DCM, and NOR. Shifeng Zhao, Yanyang Li, Yun Tian 0002, Jinxiao Xiao, Lianming Wu |
BIBM | 2 |
| 2023 | SPRING: Situated Conversation Agent Pretrained with Multimodal Questions from Incremental Layout GraphabstractExisting multimodal conversation agents have shown impressive abilities to locate absolute positions or retrieve attributes in simple scenarios, but they fail to perform well when complex relative positions and information alignments are involved, which poses a bottleneck in response quality. In this paper, we propose a Situated Conversation Agent Pretrained with Multimodal Questions from Incremental Layout Graph (SPRING) with abilities of reasoning multi-hops spatial relations and connecting them with visual attributes in crowded situated scenarios. Specifically, we design two types of Multimodal Question Answering (MQA) tasks to pretrain the agent. All QA pairs utilized during pretraining are generated from novel Increment Layout Graphs (ILG). QA pair difficulty labels automatically annotated by ILG are used to promote MQA-based Curriculum Learning. Experimental results verify the SPRING's effectiveness, showing that it significantly outperforms state-of-the-art approaches on both SIMMC 1.0 and SIMMC 2.0 datasets. We release our code and data at https://github.com/LYX0501/SPRING. Yuxing Long, Binyuan Hui, Fulong Ye, Yanyang Li, Zhuoxin Han, Caixia Yuan, Yongbin Li 0001, Xiaojie Wang 0006 |
AAAI | 4 |
| 2023 | MVP-Tuning: Multi-View Knowledge Retrieval with Prompt Tuning for Commonsense ReasoningabstractYongfeng Huang, Yanyang Li, Yichong Xu, Lin Zhang, Ruyi Gan, Jiaxing Zhang, Liwei Wang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Yongfeng Huang 0001, Yanyang Li, Yichong Xu, Ruyi Gan, Jiaxing Zhang 0001, Liwei Wang 0009 |
ACL (1) | 2 |
| 2023 | Learning Preference Model for LLMs via Automatic Preference Data GenerationabstractDespite the advanced capacities of the stateof-the-art large language models (LLMs), they suffer from issues of hallucination, stereotype, etc. Preference models play an important role in LLM alignment, yet training preference models predominantly rely on human-annotated data.This reliance limits their versatility and scalability.In this paper, we propose learning the preference model for LLMs via automatic preference data generation (AutoPM).Our approach involves both In-Breadth Data Generation, which elicits pairwise preference data from LLMs following the helpful-honestharmless (HHH) criteria, and In-Depth Data Generation, which enriches the dataset with responses spanning a wide quality range.With HHH-guided preference data, our approach simultaneously enables the LLMs to learn human preferences and align with human values.Quantitative assessments on five benchmark datasets demonstrate the reliability and potential of AutoPM, pointing out a more general and scalable way to improve LLM performance. Shijia Huang, Jianqiao Zhao, Yanyang Li, Liwei Wang 0009 |
EMNLP | 3 |
| 2023 | VL-PET: Vision-and-Language Parameter-Efficient Tuning via Granularity ControlabstractAs the model size of pre-trained language models (PLMs) grows rapidly, full fine-tuning becomes prohibitively expensive for model training and storage. In vision-and-language (VL), parameter-efficient tuning (PET) techniques are proposed to integrate modular modifications (e.g., Adapter and LoRA) into encoder-decoder PLMs. By tuning a small set of trainable parameters, these techniques perform on par with full fine-tuning. However, excessive modular modifications and neglecting the functionality gap between the encoders and decoders can lead to performance degradation, while existing PET techniques (e.g., VL-Adapter) overlook these critical issues. In this paper, we propose a Vision-and-Language Parameter-Efficient Tuning (VL-PET) framework to impose effective control over modular modifications via a novel granularity-controlled mechanism. Considering different granularity-controlled matrices generated by this mechanism, a variety of model-agnostic VL-PET modules can be instantiated from our framework for better efficiency and effective-ness trade-offs. We further propose lightweight PET module designs to enhance VL alignment and modeling for the encoders and maintain text generation for the decoders. Extensive experiments conducted on four image-text tasks and four video-text tasks demonstrate the efficiency, effectiveness and transferability of our VL-PET framework. In particular, our VL-PETlargewith lightweight PET module designs significantly outperforms VL-Adapter by 2.92% (3.41%) and LoRA by 3.37% (7.03%) with BART-base (T5-base) on image-text tasks. Furthermore, we validate the enhanced effect of employing our VL-PET designs on existing PET techniques, enabling them to achieve significant performance improvements. Our code is available at https://github.com/HenryHZY/VL-PET. Zi-Yuan Hu, Yanyang Li, Michael R. Lyu, Liwei Wang 0009 |
ICCV | 2 |
| 2023 | Knowledgeable Salient Span Mask for Enhancing Language Models as Knowledge Base
Cunxiang Wang, Fuli Luo, Yanyang Li, Runxin Xu, Fei Huang 0002, Yue Zhang 0004 |
NLPCC (2) | 3 |
| 2022 | Probing Structured Pruning on Multilingual Pre-trained Models: Settings, Algorithms, and EfficiencyabstractStructured pruning has been extensively studied on monolingual pre-trained language models and is yet to be fully evaluated on their multilingual counterparts.This work investigates three aspects of structured pruning on multilingual pre-trained language models: settings, algorithms, and efficiency.Experiments on nine downstream tasks show several counterintuitive phenomena: for settings, individually pruning for each language does not induce a better result; for algorithms, the simplest method performs the best; for efficiency, a fast model does not imply that it is also small.To facilitate the comparison on all sparsity levels, we present Dynamic Sparsification, a simple approach that allows training the model once and adapting to different model sizes at inference.We hope this work fills the gap in the study of structured pruning on multilingual pre-trained models and sheds light on future research. Yanyang Li, Fuli Luo, Runxin Xu, Songfang Huang, Fei Huang 0002, Liwei Wang 0009 |
ACL (1) | 1 |
| 2022 | Eliciting Knowledge from Large Pre-Trained Models for Unsupervised Knowledge-Grounded ConversationabstractRecent advances in large-scale pre-training provide large models with the potential to learn knowledge from the raw text.It is thus natural to ask whether it is possible to leverage these large models as knowledge bases for downstream tasks.In this work, we answer the aforementioned question in unsupervised knowledge-grounded conversation.We explore various methods that best elicit knowledge from large models.Our human study indicates that, though hallucinations exist, large models post the unique advantage of being able to output common sense and summarize facts that cannot be directly retrieved from the search engine.To better exploit such generated knowledge in dialogue generation, we treat the generated knowledge as a noisy knowledge source and propose the posterior-based reweighing as well as the noisy training strategy.Empirical results on two benchmarks show advantages over the state-of-the-art methods. Yanyang Li, Jianqiao Zhao, Michael R. Lyu, Liwei Wang 0009 |
EMNLP | 1 |
| 2022 | FlowEval: A Consensus-Based Dialogue Evaluation Framework Using Segment Act FlowsabstractDespite recent progress in open-domain dialogue evaluation, how to develop automatic metrics remains an open problem.We explore the potential of dialogue evaluation featuring dialog act information, which was hardly explicitly modeled in previous methods.However, defined at the utterance level in general, dialog act is of coarse granularity, as an utterance can contain multiple segments possessing different functions.Hence, we propose segment act, an extension of dialog act from utterance level to segment level, and crowdsource a largescale dataset for it.To utilize segment act flows, sequences of segment acts, for evaluation, we develop the first consensus-based dialogue evaluation framework, FlowEval.This framework provides a reference-free approach for dialog evaluation by finding pseudo-references.Extensive experiments against strong baselines on three benchmark datasets demonstrate the effectiveness and other desirable characteristics of our FlowEval, pointing out a potential path for better dialogue evaluation. Jianqiao Zhao, Yanyang Li, Wanyu Du, Yangfeng Ji, Dong Yu 0001, Michael R. Lyu, Liwei Wang 0009 |
EMNLP | 2 |
| 2021 | An Efficient Transformer Decoder with Compressed Sub-layersabstractThe large attention-based encoder-decoder network (Transformer) has become prevailing recently due to its effectiveness. But the high computation complexity of its decoder raises the inefficiency issue. By examining the mathematic formulation of the decoder, we show that under some mild conditions, the architecture could be simplified by compressing its sub-layers, the basic building block of Transformer, and achieves a higher parallelism. We thereby propose Compressed Attention Network, whose decoder layer consists of only one sub-layer instead of three. Extensive experiments on 14 WMT machine translation tasks show that our model is 1.42x faster with performance on par with a strong baseline. This strong baseline is already 2x faster than the widely used standard baseline without loss in performance. Yanyang Li, Tong Xiao 0001 |
AAAI | 1 |
| 2021 | Weight Distillation: Transferring the Knowledge in Neural Network ParametersabstractYe Lin, Yanyang Li, Ziyang Wang, Bei Li, Quan Du, Tong Xiao, Jingbo Zhu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Yanyang Li, Quan Du, Tong Xiao 0001 |
ACL/IJCNLP (1) | 2 |
| 2021 | Stacked Acoustic-and-Textual Encoding: Integrating the Pre-trained Models into Speech Translation EncodersabstractChen Xu, Bojie Hu, Yanyang Li, Yuhao Zhang, Shen Huang, Qi Ju, Tong Xiao, Jingbo Zhu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Chen Xu 0008, Bojie Hu, Yanyang Li, Shen Huang, Qi Ju 0002, Tong Xiao 0001 |
ACL/IJCNLP (1) | 3 |
| 2021 | A Cooperative Multi-RSU Caching Scheme in Vehicular Networks With Fountain CodesabstractIn this paper, we propose an edge caching scheme in vehicular networks, which allows multiple roadside units (RSUs) to collaboratively cache part of content in a distributed way and delivery coded packets by using fountain codes. The overlapping placement algorithm is proposed to increase the correlation of data packets cached in RSUs. When distributing content, the unequal probability selection algorithm of generating fountain codes is proposed to decrease the decoding overhead. The simulation results prove the scheme above that the transmission efficiency and the successful content acquisition rate can be improved than the traditional coded caching when the caching memory is limited. Yanyang Li, Qinghe Du |
PIMRC | 1 |
| 2020 | Neural Machine Translation with Joint RepresentationabstractThough early successes of Statistical Machine Translation (SMT) systems are attributed in part to the explicit modelling of the interaction between any two source and target units, e.g., alignment, the recent Neural Machine Translation (NMT) systems resort to the attention which partially encodes the interaction for efficiency. In this paper, we employ Joint Representation that fully accounts for each possible interaction. We sidestep the inefficiency issue by refining representations with the proposed efficient attention operation. The resulting Reformer models offer a new Sequence-to-Sequence modelling paradigm besides the Encoder-Decoder framework and outperform the Transformer baseline in either the small scale IWSLT14 German-English, English-German and IWSLT15 Vietnamese-English or the large scale NIST12 Chinese-English translation tasks by about 1 BLEU point. We also propose a systematic model scaling approach, allowing the Reformer model to beat the state-of-the-art Transformer in IWSLT14 German-English and NIST12 Chinese-English with about 50% fewer parameters. The code is publicly available at https://github.com/lyy1994/reformer. Yanyang Li, Qiang Wang 0050, Tong Xiao 0001, Tongran Liu |
AAAI | 1 |
| 2020 | A Simple and Effective Approach to Robust Unsupervised Bilingual Dictionary InductionabstractUnsupervised Bilingual Dictionary Induction methods based on the initialization and the selflearning have achieved great success in similar language pairs, e.g., English-Spanish.But they still fail and have an accuracy of 0% in many distant language pairs, e.g., English-Japanese.In this work, we show that this failure results from the gap between the actual initialization performance and the minimum initialization performance for the self-learning to succeed.We propose Iterative Dimension Reduction to bridge this gap.Our experiments show that this simple method does not hamper the performance of similar language pairs and achieves an accuracy of 13.64∼55.53%between English and four distant languages, i.e., Chinese, Japanese, Vietnamese and Thai. Yanyang Li, Yingfeng Luo, Quan Du, Huizhen Wang, Shujian Huang, Tong Xiao 0001 |
COLING | 1 |
| 2020 | Towards Fully 8-bit Integer Inference for the Transformer Modelabstract8-bit integer inference, as a promising direction in reducing both the latency and storage of deep neural networks, has made great progress recently. On the other hand, previous systems still rely on 32-bit floating point for certain functions in complex models (e.g., Softmax in Transformer), and make heavy use of quantization and de-quantization. In this work, we show that after a principled modification on the Transformer architecture, dubbed Integer Transformer, an (almost) fully 8-bit integer inference algorithm Scale Propagation could be derived. De-quantization is adopted when necessary, which makes the network more efficient. Our experiments on WMT16 En<->Ro, WMT14 En<->De and En->Fr translation tasks as well as the WikiText-103 language modelling task show that the fully 8-bit Transformer system achieves comparable performance with the floating point baseline but requires nearly 4x less memory footprint. Yanyang Li, Tengbo Liu, Tong Xiao 0001, Tongran Liu |
IJCAI | 2 |
| 2019 | Analysis of Back-Translation Methods for Low-Resource Neural Machine Translation
Nuo Xu 0010, Yinqiao Li, Chen Xu 0008, Yanyang Li, Tong Xiao 0001 |
NLPCC (2) | 4 |
| 2018 | Multi-layer Representation Fusion for Neural Machine TranslationabstractNeural machine translation systems require a number of stacked layers for deep models. But the prediction depends on the sentence representation of the top-most layer with no access to low-level representations. This makes it more difficult to train the model and poses a risk of information loss to prediction. In this paper, we propose a multi-layer representation fusion (MLRF) approach to fusing stacked layers. In particular, we design three fusion functions to learn a better representation from the stack. Experimental results show that our approach yields improvements of 0.92 and 0.56 BLEU points over the strong Transformer baseline on IWSLT German-English and NIST Chinese-English MT tasks respectively. The result is new state-of-the-art in German-English translation. Qiang Wang 0050, Fuxue Li, Tong Xiao 0001, Yanyang Li, Yinqiao Li |
COLING | 4 |
| 2017 | Application of Data Augmentation Methods to Unmanned Aerial Vehicle Monitoring System for Facial Camouflage Recognition
Yanyang Li, Sanqing Hu |
ICONIP (3) | 1 |
| 2017 | Recognition of Voluntary Blink and Bite Base on Single Forehead EMG
Shaokai Zhao, Yanyang Li, Sanqing Hu |
ICONIP (2) | 4 |