Yuanxin Liu

dblp:55/5877 · DBLP profile ↗
← Back
36ranked-venue papers
10as first author
21since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 23 · 7 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 1 since 2021Theory of computation · 3 · 1 first-authorComputer networks · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
YearPublicationVenuePosition
2026 TEMPLE: Incentivizing Temporal Understanding of Video Large Language Models via Progressive Pre-SFT Alignment
abstract
Video Large Language Models (Video LLMs) have achieved significant success by adopting the paradigm of large-scale pre-training followed by supervised fine-tuning (SFT). However, existing approaches struggle with temporal reasoning due to weak temporal correspondence in the data and over-reliance on the next-token prediction paradigm, which collectively result in the absence temporal supervision. To address these limitations, we propose TEMPLE (TEMporal Preference Learning), a systematic framework that enhances temporal reasoning capabilities through Direct Preference Optimization (DPO). To address temporal information scarcity in data, we introduce an automated pipeline for systematically constructing temporality-intensive preference pairs comprising three steps: selecting temporally rich videos, designing video-specific perturbation strategies, and evaluating model responses on clean and perturbed inputs. Complementing this data pipeline, we provide additional supervision signals via preference learning and propose a novel Progressive Pre-SFT Alignment strategy featuring two key innovations: a curriculum learning strategy which progressively increases perturbation difficulty to maximize data efficiency; and applying preference optimization before instruction tuning to incentivize fundamental temporal alignment. Extensive experiments demonstrate that our approach consistently improves Video LLM performance across multiple benchmarks with a relatively small set of self-generated DPO data. Our findings highlight TEMPLE as a scalable and efficient complement to SFT-based methods, paving the way for developing reliable Video LLMs.
Lei Li 0039, Kun Ouyang, Shuhuai Ren, Yuanxin Liu, Yuanxing Zhang, Lingpeng Kong, Qi Liu 0049, Xu Sun 0001
AAAI5
2026 Investigating Cross-Modal Skill Injection: Scenarios, Methods, and Hyperparameters
abstract
Zhiyu Xu, Lean Wang, Yuanxin Liu, Lei Li, Hao Zhou, Fandong Meng, Jie Zhou, Xu Sun. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Lean Wang, Yuanxin Liu, Lei Li 0009, Hao Zhou 0012, Fandong Meng, Jie Zhou 0016, Xu Sun 0001
ACL (1)3
2026 Secure federated learning with configurable reliability in heterogeneous edge computing
Yuanxin Liu, Yihuai Liang, Zhengchun Zhou
Comput. Networks1
2026 Knowledge-Driven and Relation-Aware Synergistic Learning for Drug Repositioning
abstract
As an effective and low-risk approach to identify new therapeutic pathways for existing drugs, drug repositioning has been extensively utilized to expedit drug discovery processes. However, current knowledge graph (KG)-based methodologies encounter several hurdles in this context. Firstly, most graph neural network (GNN)-based approaches fail to adequately capture the intricate relationships between drug-drug, drug-disease, or disease-disease. Secondly, the subtle synergistic mechanisms between drugs and diseases remain underexplored. Lastly, the training of knowledge graph embedding (KGE) methods is susceptible to noise, leading to unstable model optimization. To address these challenges, we intruduce KRANE, a knowledge-driven and relation-aware synergistic learning method for drug repositioning. KRANE addresses these issues through three innovative modules. Firstly, we design a relation-aware feature extractor (RAFE), which utilizes the contextual triples attention scores in KG to effectively integrate drug-related knowledge and enhance the representation of complex relational features. Secondly, we adopt a synergistic feature reconstruction module as a decoder to extract synergistic heterogeneous feature interactions between drugs and diseases from entity and relation representations. Finally, we propose a knowledge-regulated loss function to mitigate the impact of noise on model training. Experiments conducted on three publicly available datasets demonstrate that KRANE significantly outperforms existing methods.
Shilong Wang 0004, Yuanxin Liu, Xiaobo Li 0007, Hai Cui, Yi-Jia Zhang 0001
IEEE J. Biomed. Health Informatics2
2025 PunchBench: Benchmarking MLLMs in Multimodal Punchline Comprehension
abstract
Kun Ouyang, Yuanxin Liu, Shicheng Li, Yi Liu, Hao Zhou, Fandong Meng, Jie Zhou, Xu Sun. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Kun Ouyang, Yuanxin Liu, Hao Zhou 0012, Fandong Meng, Jie Zhou 0016, Xu Sun 0001
ACL (1)2
2025 RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction
abstract
Yuchi Wang, Yishuo Cai, Shuhuai Ren, Sihan Yang, Linli Yao, Yuanxin Liu, Yuanxing Zhang, Pengfei Wan, Xu Sun. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Yuchi Wang, Yishuo Cai, Shuhuai Ren, Linli Yao, Yuanxin Liu, Yuanxing Zhang, Pengfei Wan 0001, Xu Sun 0001
EMNLP6
2025 Temporal Reasoning Transfer from Text to Video
abstract
Video Large Language Models (Video LLMs) have shown promising capabilities in video comprehension, yet they struggle with tracking temporal changes and reasoning about temporal relationships. While previous research attributed this limitation to the ineffective temporal encoding of visual inputs, our diagnostic study reveals that video representations contain sufficient information for even small probing classifiers to achieve perfect accuracy. Surprisingly, we find that the key bottleneck in Video LLMs' temporal reasoning capability stems from the underlying LLM's inherent difficulty with temporal concepts, as evidenced by poor performance on textual temporal question-answering tasks. Building on this discovery, we introduce the Textual Temporal reasoning Transfer (T3). T3 synthesizes diverse temporal reasoning tasks in pure text format from existing image-text datasets, addressing the scarcity of video samples with complex temporal scenarios. Remarkably, without using any video data, T3 enhances LongVA-7B's temporal understanding, yielding a 5.3 absolute accuracy improvement on the challenging TempCompass benchmark, which enables our model to outperform ShareGPT4Video-8B trained on 28,000 video samples. Additionally, the enhanced LongVA-7B model achieves competitive performance on comprehensive video benchmarks. For example, it achieves a 49.7 accuracy on the Temporal Reasoning task of Video-MME, surpassing powerful large-scale models such as InternVL-Chat-V1.5-20B and VILA1.5-40B. Further analysis reveals a strong correlation between textual and video temporal task performance, validating the efficacy of transferring temporal reasoning abilities from text to video domains.
Lei Li 0039, Yuanxin Liu, Linli Yao, Peiyuan Zhang, Chenxin An, Lean Wang, Xu Sun 0001, Lingpeng Kong, Qi Liu 0049
ICLR2
2025 BRiTE: Bootstrapping Reinforced Thinking Process to Enhance Language Model Reasoning
abstract
Large Language Models (LLMs) have demonstrated remarkable capabilities in complex reasoning tasks, yet generating reliable reasoning processes remains a significant challenge. We present a unified probabilistic framework that formalizes LLM reasoning through a novel graphical model incorporating latent thinking processes and evaluation signals. Our framework addresses two critical questions: (1) how to generate high-quality reasoning processes during inference automatically, and (2) how to integrate these processes into post-training. We propose the Bootstrapping Reinforced Thinking Process (BRiTE) algorithm and demonstrate its theoretical convergence at a rate of $1/T$, where $T$ is the number of iterations. The algorithm operates in two steps. First, it generates high-quality rationales by approximating the desired posterior distribution using a reinforcement learning approach with a novel reward shaping mechanism. Second, it fine-tunes the base LLM by maximizing the joint probability of rationale generation with respect to LLM parameters. Empirical evaluation on GSM8K and MATH benchmarks demonstrates that our approach consistently improves performance across different model sizes without requiring human-annotated thinking processes, outperforming standard chain-of-thought prompting while enhancing existing post-training methods.
Han Zhong 0001, Yutong Yin, Shenao Zhang, Yuanxin Liu, Yifei Zuo, Boyi Liu 0001, Sirui Zheng, Hongyi Guo, Liwei Wang 0001, Mingyi Hong 0001, Zhaoran Wang 0001
ICML5
2025 TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos
abstract
The rapid growth of online video platforms, particularly live streaming services, has created an urgent need for real-time video understanding systems. These systems must process continuous video streams and respond to user queries instantaneously, presenting unique challenges for current Video Large Language Models (VideoLLMs). While existing VideoLLMs excel at processing complete videos, they face significant limitations in streaming scenarios due to their inability to handle dense, redundant frames efficiently. We introduce TimeChat-Online, a novel online VideoLLM that revolutionizes real-time video interaction. At its core lies our innovative Differential Token Drop (DTD) module, which addresses the fundamental challenge of visual redundancy in streaming videos. Drawing inspiration from human visual perception's Change Blindness phenomenon, DTD preserves meaningful temporal changes while filtering out static, redundant content between frames. Remarkably, our experiments demonstrate that DTD achieves an 82.8% reduction in video tokens while maintaining 98% performance on StreamingBench, revealing that over 80% of visual content in streaming videos is naturally redundant without requiring language guidance. To enable seamless real-time interaction, we present TimeChat-Online-139K, a comprehensive streaming video dataset featuring diverse interaction patterns including backward-tracing, current-perception, and future-responding scenarios. TimeChat-Online's unique Proactive Response capability, naturally achieved through continuous monitoring of video scene transitions via DTD, sets it apart from conventional approaches. Our extensive evaluation demonstrates TimeChat-Online's superior performance on streaming benchmarks (StreamingBench and OvOBench) and maintaining competitive results on long-form video tasks such as Video-MME and MLVU. Notably, when integrated with Qwen2.5VL-7B, DTD achieves a 5.7-point accuracy improvement on the challenging VideoMME subset containing videos of 30-60 minutes, while reducing video tokens by 84.6%. Project page: https://timechat-online.github.io.
Linli Yao, Yuancheng Wei, Lei Li 0039, Shuhuai Ren, Yuanxin Liu, Kun Ouyang, Lean Wang, Lingpeng Kong, Qi Liu 0049, Yuanxing Zhang, Xu Sun 0001
ACM Multimedia6
2025 UVE: Are MLLMs Unified Evaluators for AI-Generated Videos?
abstract
With the rapid growth of video generative models (VGMs), it is essential to develop reliable and comprehensive automatic metrics for AI-generated videos (AIGVs). Existing methods either use off-the-shelf models optimized for other tasks or rely on human assessment data to train specialized evaluators. These approaches are constrained to specific evaluation aspects and are difficult to scale with the increasing demands for finer-grained and more comprehensive evaluations. To address this issue, this work investigates the feasibility of using multimodal large language models (MLLMs) as a unified evaluator for AIGVs, leveraging their strong visual perception and language understanding capabilities. To evaluate the performance of automatic metrics in unified AIGV evaluation, we introduce a benchmark called UVE-Bench. UVE-Bench collects videos generated by state-of-the-art VGMs and provides pairwise human preference annotations across 15 evaluation aspects. Using UVE-Bench, we extensively evaluate 18 MLLMs. Our empirical results suggest that while advanced MLLMs (e.g., Qwen2VL-72B and InternVL2.5-78B) still lag behind human evaluators, they demonstrate promising ability in unified AIGV evaluation, significantly surpassing existing specialized evaluation methods. Additionally, we conduct an in-depth analysis of key design choices that impact the performance of MLLM-driven evaluators, offering valuable insights for future research on AIGV evaluation.
Yuanxin Liu, Shuhuai Ren, Jiacong Wang, Haoyuan Guo
NeurIPS1
2025 Dual-stage learning framework for underwater acoustic target recognition with cross-attention mechanism and audio-guided contrastive learning
Rongyao Zhao, Lyufang Zhao, Daihui Li, Yuanxin Liu, Tongsheng Shen
Neurocomputing6
2024 VITATECS: A Diagnostic Dataset for Temporal Concept Understanding of Video-Language Models
Lei Li 0039, Shuhuai Ren, Yuanxin Liu, Rundong Gao, Xu Sun 0001, Lu Hou 0002
ECCV (70)5
2023 Compressing and Debiasing Vision-Language Pre-Trained Models for Visual Question Answering
abstract
Despite the excellent performance of visionlanguage pre-trained models (VLPs) on conventional VQA task, they still suffer from two problems: First, VLPs tend to rely on language biases in datasets and fail to generalize to outof-distribution (OOD) data.Second, they are inefficient in terms of memory footprint and computation.Although promising progress has been made in both problems, most existing works tackle them independently.To facilitate the application of VLP to VQA tasks, it is imperative to jointly study VLP compression and OOD robustness, which, however, has not yet been explored.This paper investigates whether a VLP can be compressed and debiased simultaneously by searching sparse and robust subnetworks.To this end, we systematically study the design of a training and compression pipeline to search the subnetworks, as well as the assignment of sparsity to different modality-specific modules.Our experiments involve 3 VLPs, 2 compression methods, 4 training methods, 2 datasets and a range of sparsity levels.Our results show that there indeed exist sparse and robust subnetworks, which are competitive with the debiased full VLP and clearly outperform the debiasing SoTAs with fewer parameters on OOD datasets VQA-CP v2 and VQA-VS. 1
Qingyi Si, Yuanxin Liu, Zheng Lin 0001, Peng Fu 0008, Yanan Cao 0001, Weiping Wang 0005
EMNLP2
2023 FETV: A Benchmark for Fine-Grained Evaluation of Open-Domain Text-to-Video Generation
abstract
Recently, open-domain text-to-video (T2V) generation models have made remarkable progress. However, the promising results are mainly shown by the qualitative cases of generated videos, while the quantitative evaluation of T2V models still faces two critical problems. Firstly, existing studies lack fine-grained evaluation of T2V models on different categories of text prompts. Although some benchmarks have categorized the prompts, their categorization either only focuses on a single aspect or fails to consider the temporal information in video generation. Secondly, it is unclear whether the automatic evaluation metrics are consistent with human standards. To address these problems, we propose FETV, a benchmark for Fine-grained Evaluation of Text-to-Video generation. FETV is multi-aspect, categorizing the prompts based on three orthogonal aspects: the major content, the attributes to control and the prompt complexity. FETV is also temporal-aware, which introduces several temporal categories tailored for video generation. Based on FETV, we conduct comprehensive manual evaluations of four representative T2V models, revealing their pros and cons on different categories of prompts from different aspects. We also extend FETV as a testbed to evaluate the reliability of automatic T2V metrics. The multi-aspect categorization of FETV enables fine-grained analysis of the metrics' reliability in different scenarios. We find that existing automatic metrics (e.g., CLIPScore and FVD) correlate poorly with human evaluation. To address this problem, we explore several solutions to improve CLIPScore and FVD, and develop two automatic metrics that exhibit significant higher correlation with humans than existing metrics. Benchmark page: https://github.com/llyx97/FETV.
Yuanxin Liu, Lei Li 0039, Shuhuai Ren, Rundong Gao, Sishuo Chen, Xu Sun 0001, Lu Hou 0002
NeurIPS1
2022 COST-EFF: Collaborative Optimization of Spatial and Temporal Efficiency with Slenderized Multi-exit Language Models
abstract
Transformer-based pre-trained language models (PLMs) mostly suffer from excessive overhead despite their advanced capacity.For resource-constrained devices, there is an urgent need for a spatially and temporally efficient model which retains the major capacity of PLMs.However, existing statically compressed models are unaware of the diverse complexities between input instances, potentially resulting in redundancy and inadequacy for simple and complex inputs.Also, miniature models with early exiting encounter challenges in the trade-off between making predictions and serving the deeper layers.Motivated by such considerations, we propose a collaborative optimization for PLMs that integrates static model compression and dynamic inference acceleration.Specifically, the PLM is slenderized in width while the depth remains intact, complementing layer-wise early exiting to speed up inference dynamically.To address the trade-off of early exiting, we propose a joint training approach that calibrates slenderization and preserves contributive structures to each exit instead of only the final layer.Experiments are conducted on GLUE benchmark and the results verify the Pareto optimality of our approach at high compression and acceleration rate with 1/8 parameters and 1/19 FLOPs of BERT.
Bowen Shen, Zheng Lin 0001, Yuanxin Liu, Zhengxiao Liu, Lei Wang 0135, Weiping Wang 0005
EMNLP3
2022 Connecting Targets via Latent Topics And Contrastive Learning: A Unified Framework For Robust Zero-Shot and Few-Shot Stance Detection
abstract
Zero-shot and few-shot stance detection (ZFSD) aims to automatically identify the users’ stance toward a wide range of continuously emerging targets without or with limited labeled data. Previous works on in-target and cross-target stance detection typically focus on extremely limited targets, which is not applicable to the zero-shot and few-shot scenarios. Additionally, existing ZFSD models are not good at modeling the relationship between seen and unseen targets. In this paper, we propose a unified end-to-end framework with a discrete latent topic variable that implicitly establishes the connections between targets. Moreover, we apply supervised contrastive learning to enhance the generalization ability of the model. Comprehensive experiments on the ZFSD task verify the effectiveness and superiority of our proposed method.
Rui Liu 0032, Zheng Lin 0001, Peng Fu 0008, Yuanxin Liu, Weiping Wang 0005
ICASSP4
2022 Learning to Win Lottery Tickets in BERT Transfer via Task-agnostic Mask Training
abstract
Yuanxin Liu, Fandong Meng, Zheng Lin, Peng Fu, Yanan Cao, Weiping Wang, Jie Zhou. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Yuanxin Liu, Fandong Meng, Zheng Lin 0001, Peng Fu 0008, Yanan Cao 0001, Weiping Wang 0005, Jie Zhou 0016
NAACL-HLT1
2022 A Win-win Deal: Towards Sparse and Robust Pre-trained Language Models
abstract
Despite the remarkable success of pre-trained language models (PLMs), they still face two challenges: First, large-scale PLMs are inefficient in terms of memory footprint and computation. Second, on the downstream tasks, PLMs tend to rely on the dataset bias and struggle to generalize to out-of-distribution (OOD) data. In response to the efficiency problem, recent studies show that dense PLMs can be replaced with sparse subnetworks without hurting the performance. Such subnetworks can be found in three scenarios: 1) the fine-tuned PLMs, 2) the raw PLMs and then fine-tuned in isolation, and even inside 3) PLMs without any parameter fine-tuning. However, these results are only obtained in the in-distribution (ID) setting. In this paper, we extend the study on PLMs subnetworks to the OOD setting, investigating whether sparsity and robustness to dataset bias can be achieved simultaneously. To this end, we conduct extensive experiments with the pre-trained BERT model on three natural language understanding (NLU) tasks. Our results demonstrate that \textbf{sparse and robust subnetworks (SRNets) can consistently be found in BERT}, across the aforementioned three scenarios, using different training and compression methods. Furthermore, we explore the upper bound of SRNets using the OOD information and show that \textbf{there exist sparse and almost unbiased BERT subnetworks}. Finally, we present 1) an analytical study that provides insights on how to promote the efficiency of SRNets searching process and 2) a solution to improve subnetworks' performance at high sparsity. The code is available at \url{https://github.com/llyx97/sparse-and-robust-PLM}.
Yuanxin Liu, Fandong Meng, Zheng Lin 0001, Peng Fu 0008, Yanan Cao 0001, Weiping Wang 0005, Jie Zhou 0016
NeurIPS1
2021 ROSITA: Refined BERT cOmpreSsion with InTegrAted techniques
abstract
Pre-trained language models of the BERT family have defined the state-of-the-arts in a wide range of NLP tasks. However, the performance of BERT-based models is mainly driven by the enormous amount of parameters, which hinders their application to resource-limited scenarios. Faced with this problem, recent studies have been attempting to compress BERT into a small-scale model. However, most previous work primarily focuses on a single kind of compression technique, and few attention has been paid to the combination of different methods. When BERT is compressed with integrated techniques, a critical question is how to design the entire compression framework to obtain the optimal performance. In response to this question, we integrate three kinds of compression methods (weight pruning, low-rank factorization and knowledge distillation (KD)) and explore a range of designs concerning model architecture, KD strategy, pruning frequency and learning rate schedule. We find that a careful choice of the designs is crucial to the performance of the compressed model. Based on the empirical findings, our best compressed model, dubbed Refined BERT cOmpreSsion with InTegrAted techniques (ROSITA), is 7.5x smaller than BERT while maintains 98.5% of the performance on five tasks of the GLUE benchmark, outperforming the previous BERT compression methods with similar parameter budget.
Yuanxin Liu, Zheng Lin 0001, Fengcheng Yuan
AAAI1
2021 Marginal Utility Diminishes: Exploring the Minimum Knowledge for BERT Knowledge Distillation
abstract
Yuanxin Liu, Fandong Meng, Zheng Lin, Weiping Wang, Jie Zhou. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Yuanxin Liu, Fandong Meng, Zheng Lin 0001, Weiping Wang 0005, Jie Zhou 0016
ACL/IJCNLP (1)1
2021 Learning Class-Transductive Intent Representations for Zero-shot Intent Detection
abstract
Zero-shot intent detection (ZSID) aims to deal with the continuously emerging intents without annotated training data. However, existing ZSID systems suffer from two limitations: 1) They are not good at modeling the relationship between seen and unseen intents. 2) They cannot effectively recognize unseen intents under the generalized intent detection (GZSID) setting. A critical problem behind these limitations is that the representations of unseen intents cannot be learned in the training stage. To address this problem, we propose a novel framework that utilizes unseen class labels to learn Class-Transductive Intent Representations (CTIR). Specifically, we allow the model to predict unseen intents during training, with the corresponding label names serving as input utterances. On this basis, we introduce a multi-task learning objective, which encourages the model to learn the distinctions among intents, and a similarity scorer, which estimates the connections among intents more accurately. CTIR is easy to implement and can be integrated with existing ZSID and GZSID methods. Experiments on two real-world datasets show that CTIR brings considerable improvement to the baseline systems.
Qingyi Si, Yuanxin Liu, Peng Fu 0008, Zheng Lin 0001, Weiping Wang 0005
IJCAI2
2019 Generating Paraphrase with Topic as Prior Knowledge
abstract
Paraphrase generation can be modeled as a sequence-to-sequence (Seq2Seq) learning problem. Nonetheless, a typical Seq2Seq model is liable to convey the original meaning incorrectly, as the vectorial representation of the given sentence is sometimes inadequate in recapitulating complicated semantic. Naturally, paraphrases concern the same topic, which can serve as an auxiliary guidance to promote the preservation of source semantic. Moreover, some interesting words for restatements can be derived from the topical information. To exploit topic in paraphrase generation, we incorporate topic words into the Seq2Seq framework through a topic-aware input and a topic-biased generation distribution. Direct supervision signals are also introduced to help dealing with the topic information more accurately. Empirical studies on two benchmark datasets show that the proposed method significantly improves the basic Seq2Seq model, and it is comparable with the state-of-the-art systems.
Yuanxin Liu, Zheng Lin 0001, Qinyun Dai, Weiping Wang 0005
CIKM1
2019 Self-Adaptive Scaling for Learnable Residual Structure
abstract
Residual has been widely applied to build deep neural networks with enhanced feature propagation and improved accuracy.In the literature, multiple variants of residual structure are proposed.However, most of them are manually designed for particular tasks and datasets and the combination of existing residual structures has not been well studied.In this work, we propose the Self-Adaptive Scaling (SAS) approach that automatically learns the design of residual structure from data.The proposed approach makes the best of various residual structures, resulting in a general architecture covering several existing ones.In this manner, we construct a learnable residual structure which can be easily integrated into a wide range of residual-based models.We evaluate our approach on various tasks concerning different modalities, including machine translation (IWSLT-2015 EN-VI and WMT-2014 EN-DE, EN-FR), image classification (CIFAR-10 and CIFAR-100), and image captioning (MSCOCO).Empirical results show that the proposed approach consistently improves the residual-based models and exhibits desirable generalization ability.In particular, by incorporating the proposed approach to the Transformer model, we establish new state-of-thearts on the IWSLT-2015 EN-VI low-resource machine translation dataset.
Yuanxin Liu, Kai Lei
CoNLL3
2019 Ranking and Sampling in Open-Domain Question Answering
abstract
Yanfu Xu, Zheng Lin, Yuanxin Liu, Rui Liu, Weiping Wang, Dan Meng. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Yanfu Xu, Zheng Lin 0001, Yuanxin Liu, Rui Liu 0032, Weiping Wang 0005, Dan Meng 0002
EMNLP/IJCNLP (1)3
2019 Exploring and Distilling Cross-Modal Information for Image Captioning
abstract
Recently, attention-based encoder-decoder models have been used extensively in image captioning. Yet there is still great difficulty for the current methods to achieve deep image understanding. In this work, we argue that such understanding requires visual attention to correlated image regions and semantic attention to coherent attributes of interest. To perform effective attention, we explore image captioning from a cross-modal perspective and propose the Global-and-Local Information Exploring-and-Distilling approach that explores and distills the source information in vision and language. It globally provides the aspect vector, a spatial and relational representation of images based on caption contexts, through the extraction of salient region groupings and attribute collocations, and locally extracts the fine-grained regions and attributes in reference to the aspect vector for word selection. Our fully-attentive model achieves a CIDEr score of 129.3 in offline COCO evaluation with remarkable efficiency in terms of accuracy, speed, and parameter budget.
Xuancheng Ren, Yuanxin Liu, Kai Lei, Xu Sun 0001
IJCAI3
2019 Aligning Visual Regions and Textual Concepts for Semantic-Grounded Image Representations
abstract
In vision-and-language grounding problems, fine-grained representations of the image are considered to be of paramount importance. Most of the current systems incorporate visual features and textual concepts as a sketch of an image. However, plainly inferred representations are usually undesirable in that they are composed of separate components, the relations of which are elusive. In this work, we aim at representing an image with a set of integrated visual regions and corresponding textual concepts, reflecting certain semantics. To this end, we build the Mutual Iterative Attention (MIA) module, which integrates correlated visual features and textual concepts, respectively, by aligning the two modalities. We evaluate the proposed approach on two representative vision-and-language grounding tasks, i.e., image captioning and visual question answering. In both tasks, the semantic-grounded image representations consistently boost the performance of the baseline models under all metrics across the board. The results demonstrate that our approach is effective and generalizes well to a wide range of models for image-related applications. (The code is available at \url{https://github.com/fenglinliu98/MIA)
Yuanxin Liu, Xuancheng Ren, Xiaodong He 0001, Xu Sun 0001
NeurIPS2
2019 Efficient multivariate analysis algorithms for longitudinal genome-wide association studies
abstract
MOTIVATION: Current dynamic phenotyping system introduces time as an extra dimension to genome-wide association studies (GWAS), which helps to explore the mechanism of dynamical genetic control for complex longitudinal traits. However, existing methods for longitudinal GWAS either ignore the covariance among observations of different time points or encounter computational efficiency issues. RESULTS: We herein developed efficient genome-wide multivariate association algorithms for longitudinal data. In contrast to existing univariate linear mixed model analyses, the proposed method has improved statistic power for association detection and computational speed. In addition, the new method can analyze unbalanced longitudinal data with thousands of individuals and more than ten thousand records within a few hours. The corresponding time for balanced longitudinal data is just a few minutes. AVAILABILITY AND IMPLEMENTATION: A software package to implement the efficient algorithm named GMA (https://github.com/chaoning/GMA) is available freely for interested users in relevant fields. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Chao Ning 0004, Lei Zhou 0024, Julong Wei, Yuanxin Liu, Huimin Kang, Shizhong Xu, Jianfeng Liu 0003
Bioinform.5
2018 simNet: Stepwise Image-Topic Merging Network for Generating Detailed and Comprehensive Image Captions
abstract
The encode-decoder framework has shown recent success in image captioning.Visual attention, which is good at detailedness, and semantic attention, which is good at comprehensiveness, have been separately proposed to ground the caption on the image.In this paper, we propose the Stepwise Image-Topic Merging Network (simNet) that makes use of the two kinds of attention at the same time.At each time step when generating the caption, the decoder adaptively merges the attentive information in the extracted topics and the image according to the generated context, so that the visual information and the semantic information can be effectively combined.The proposed approach is evaluated on two benchmark datasets and reaches the state-of-the-art performances.1
Xuancheng Ren, Yuanxin Liu, Houfeng Wang, Xu Sun 0001
EMNLP3
2014 Deviation of land-use data from quickbird and TM remote sensing imagery: A case study of the Tanjiaying small watershed in the loess hilly-gully area of China
abstract
Land-use maps provide important data and basic information for accomplishing the optimal allocation of resources and for ensuring sustainable development. However, the land-use data interpreted from various sources of remotely sensed, low-resolution imagery are highly variable. Therefore, an analysis of these deviations in land-use data is imperative for correcting the accuracy of the land-use maps derived from such low-resolution, remote sensing imagery. With ArcGIS 9.3 software, we derived the land-use maps with different accuracies for the Tanjiaying small watershed in the loess hilly-gully area of China by interpreting Landsat 5 Thematic Mapper (TM) and QuickBird images acquired in 2010. We compared the maps derived from these two sources of imagery data and analyzed the resulting deviation distribution with respect to the slope, aspect, and elevation of the land.
Lina Zhong, Wenwu Zhao, Yuanxin Liu, Xiao Zhang 0013, Xuening Fang
IGARSS3
2007 Quadratic and cubic b-splines by generalizing higher-order voronoi diagrams
abstract
A long-standing problem in spline theory has been to generalize classic B-splines to the multivariate setting, and its full solution will have broad impact. We initiate a study of triangulations that generalize the duals of higher order Voronoi diagrams, and show that these can serve as a foundation for a family of multivariate splines that generalize the classic univariate B-splines. This paper focuseson Voronoi diagrams of orders two and three, which produce families of quadratic and cubic bivariate B-splines. We believe that these families are the most general bivariate B-splines to date and supportour belief by demonstrating that a classic quadratic box spline, the Zwart-Powell (ZP) element, is contained in our family. Our work is directly based on that of Neamtu, who established the fascinating connection between splines and higher order Voronoi diagrams.
Yuanxin Liu, Jack Snoeyink
SCG1
2006 Illustrating the streaming construction of 2D delaunay triangulations
abstract
No abstract available.
Martin Isenburg, Yuanxin Liu, Jonathan Richard Shewchuk, Jack Snoeyink
SCG2
2006 Streaming computation of Delaunay triangulations
abstract
We show how to greatly accelerate algorithms that compute Delaunay triangulations of huge, well-distributed point sets in 2D and 3D by exploiting the natural spatial coherence in a stream of points. We achieve large performance gains by introducing spatial finalization into point streams: we partition space into regions, and augment a stream of input points with finalization tags that indicate when a point is the last in its region. By extending an incremental algorithm for Delaunay triangulation to use finalization tags and produce streaming mesh output, we compute a billion-triangle terrain representation for the Neuse River system from 11.2 GB of LIDAR data in 48 minutes using only 70 MB of memory on a laptop with two hard drives. This is a factor of twelve faster than the previous fastest out-of-core Delaunay triangulation software.
Martin Isenburg, Yuanxin Liu, Jonathan Richard Shewchuk, Jack Snoeyink
ACM Trans. Graph.2
2004 Flooding Triangulated Terrain
Yuanxin Liu, Jack Snoeyink
SDH1
2004 Testing Homotopy for Paths in the Plane
Sergio Cabello, Yuanxin Liu, Andrea Mantler, Jack Snoeyink
Discret. Comput. Geom.2
2004 A viscous paint model for interactive applications
abstract
Abstract We present a viscous paint model for use in an interactive painting system based on the well‐known Stokes' equations for viscous flow. Our method is, to our knowledge, the first unconditionally stable numerical method that treats viscous fluid with a free surface boundary. We have also developed a real‐time implementation of the Kubelka‐Munk reflectance model for pigment mixing, compositing and rendering entirely on graphics hardware, using programmable fragment shading capabilities. We have integrated our paint model with a prototype painting system, which demonstrates the model's effectiveness in rendering viscous paint and capturing a thick,impasto‐like style of painting. Several users have tested our prototype system and were able to start creating original art work in an intuitive manner not possible with the existing techniques in commercial systems. Copyright © 2004 John Wiley & Sons, Ltd.
William V. Baxter III, Yuanxin Liu, Ming C. Lin
Comput. Animat. Virtual Worlds2
2002 Testing Homotopy for paths in the plane
abstract
In this paper we present an efficient algorithm to test if two given paths are homotopic; that is, whether they wind around obstacles in the plane in the same way. For simple paths specified by n line segments with obstacles described by n points, our algorithm runs in O(n log n) time, which we show is tight. For self-intersecting paths the problem is related to Hopcroft's problem.
Sergio Cabello, Yuanxin Liu, Andrea Mantler, Jack Snoeyink
SCG2