Qi Liu 0049

dblp:95/2446-49 · DBLP profile ↗
← Back
45ranked-venue papers
8as first author
28since 2021 · last 2026
0000-0003-4608-5778ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 38 · 8 first-author · 26 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 5 · 2 since 2021Computer networks · 1
YearPublicationVenuePosition
2026 TEMPLE: Incentivizing Temporal Understanding of Video Large Language Models via Progressive Pre-SFT Alignment
abstract
Video Large Language Models (Video LLMs) have achieved significant success by adopting the paradigm of large-scale pre-training followed by supervised fine-tuning (SFT). However, existing approaches struggle with temporal reasoning due to weak temporal correspondence in the data and over-reliance on the next-token prediction paradigm, which collectively result in the absence temporal supervision. To address these limitations, we propose TEMPLE (TEMporal Preference Learning), a systematic framework that enhances temporal reasoning capabilities through Direct Preference Optimization (DPO). To address temporal information scarcity in data, we introduce an automated pipeline for systematically constructing temporality-intensive preference pairs comprising three steps: selecting temporally rich videos, designing video-specific perturbation strategies, and evaluating model responses on clean and perturbed inputs. Complementing this data pipeline, we provide additional supervision signals via preference learning and propose a novel Progressive Pre-SFT Alignment strategy featuring two key innovations: a curriculum learning strategy which progressively increases perturbation difficulty to maximize data efficiency; and applying preference optimization before instruction tuning to incentivize fundamental temporal alignment. Extensive experiments demonstrate that our approach consistently improves Video LLM performance across multiple benchmarks with a relatively small set of self-generated DPO data. Our findings highlight TEMPLE as a scalable and efficient complement to SFT-based methods, paving the way for developing reliable Video LLMs.
Lei Li 0039, Kun Ouyang, Shuhuai Ren, Yuanxin Liu, Yuanxing Zhang, Lingpeng Kong, Qi Liu 0049, Xu Sun 0001
AAAI9
2026 EconProver: Towards More Economical Test-Time Scaling for Automated Theorem Proving
abstract
Mukai Li, Linfeng Song, Zhenwen Liang, Jiahao Xu, Shansan Gong, Qi Liu, Haitao Mi, Dong Yu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Mukai Li, Linfeng Song, Zhenwen Liang, Shansan Gong, Qi Liu 0049, Haitao Mi, Dong Yu 0001
ACL (1)6
2026 OS-Sentinel: Towards Safety-Enhanced Mobile GUI Agents via Hybrid Validation in Realistic Workflows
abstract
Qiushi Sun, Mukai Li, Zhoumianze Liu, Zhihui Xie, Fangzhi Xu, Zhangyue Yin, Kanzhi Cheng, Zehao Li, Zichen Ding, Qi Liu, Zhiyong Wu, Zhuosheng Zhang, Ben Kao, Lingpeng Kong. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Qiushi Sun, Mukai Li, Zhoumianze Liu, Zhihui Xie 0002, Fangzhi Xu, Zhangyue Yin, Kanzhi Cheng, Zichen Ding 0002, Qi Liu 0049, Zhiyong Wu 0003, Zhuosheng Zhang 0001, Ben Kao, Lingpeng Kong
ACL (1)10
2026 Unlock the Potential of Large Language Models for Predictive Tabular Tasks in Data Science With Table-Specific Pretraining
abstract
In data science, predictive tasks such as classification, regression, and missing value imputation are fundamental challenges in tabular data analysis. This research investigates the application of Large Language Models (LLMs) to these tasks. While LLMs excel in natural language understanding, their effectiveness on structured tabular data remains limited due to minimal exposure during pretraining. To address this gap, we construct a large-scale corpus of annotated tables and introduce a tailored pretraining framework. Our trained model achieves significant improvements over baselines, with an average gain of 8.9% in classification and 10.7% in regression tasks. We further evaluate its performance in zero-shot and few-shot prediction, as well as in-context learning scenarios. Extensive experiments demonstrate substantial gains over existing benchmarks, highlighting the potential of LLMs for tabular data processing. Additionally, we apply our approach across multiple open-source LLMs and demonstrate its generalizability. This work establishes a new benchmark for enhancing tabular intelligence through LLM-based pretraining.
Yazheng Yang, Yuqi Wang 0003, Yaxuan Li 0002, Sankalok Sen, Lei Li 0039, Qi Liu 0049
IEEE Trans. Knowl. Data Eng.7
2025 Design Choices for Extending the Context Length of Visual Language Models
abstract
Visual Language Models (VLMs) demonstrate impressive capabilities in processing multimodal inputs, yet applications such as visual agents, which require handling multiple images and high-resolution videos, demand enhanced long-range modeling.Moreover, existing opensource VLMs lack systematic exploration into extending their context length, and commercial models often provide limited details.To tackle this, we aim to establish an effective solution that enhances long context performance of VLMs while preserving their capacities in short context scenarios.Towards this goal, we make the best design choice through extensive experiment settings from data curation to context window extending and utilizing: ( 1) we analyze data sources and length distributions to construct ETVLM -a data recipe to balance the performance across scenarios; (2) we examine existing position extending methods, identify their limitations and propose M-RoPE++ as an enhanced approach; we also choose to solely instruction-tune the backbone with mixed-source data; (3) we discuss how to better utilize extended context windows and propose hybrid-resolution training.Built on the Qwen-VL series model, we propose GI-RAFFE, which is effectively extended to 128K lengths.Evaluated on extensive long context VLM benchmarks such as VideoMME and Viusal Haystacks, our GIRAFFE achieves stateof-the-art performance among similarly sized open-source long VLMs and is competitive with commercial model GPT-4V. 1
Mukai Li, Lei Li 0039, Shansan Gong, Qi Liu 0049
ACL (1)4
2025 VL-RewardBench: A Challenging Benchmark for Vision-Language Generative Reward Models
abstract
Vision-language generative reward models (VL-GenRMs) play a crucial role in aligning and evaluating multimodal AI systems, yet their own evaluation remains under-explored. Current assessment methods primarily rely on AI-annotated preference labels from traditional VL tasks, which can introduce biases and often fail to effectively challenge state-of-the-art models. To address these limitations, we introduce VL-RewardBench, a comprehensive benchmark spanning general multimodal queries, visual hallucination detection, and complex reasoning tasks. Through our AI-assisted annotation pipeline that combines sample selection with human verification, we curate 1,250 high-quality examples specifically designed to probe VL-GenRMs limitations. Comprehensive evaluation across 16 leading large vision-language models demonstrates VL-RewardBench’s effectiveness as a challenging testbed, where even GPT-4o achieves only 65.4% accuracy, and state-of-the-art open-source models such as Qwen2-VL-72B, struggle to surpass random-guessing. Importantly, performance on VL-RewardBench strongly correlates (Pearson’s r > 0.9) with MMMU-Pro accuracy using Best-of-N sampling with VL-GenRMs. Analysis experiments uncover three critical insights for improving VL-GenRMs: (i) models predominantly fail at basic visual perception tasks rather than reasoning tasks; (ii) inference-time scaling benefits vary dramatically by model capacity; and (iii) training VL-GenRMs to learn to judge substantially boosts judgment capability (+14.7% accuracy for a 7B VL-GenRM). We believe VL-RewardBench along with the experimental insights will become a valuable resource for advancing VL-GenRMs. Project page: https://vl-rewardbench.github.io.
Lei Li 0039, Yuancheng Wei, Zhihui Xie 0002, Xuqing Yang, Yifan Song 0002, Peiyi Wang, Chenxin An, Tianyu Liu 0001, Sujian Li, Bill Y. Lin, Lingpeng Kong, Qi Liu 0049
CVPR12
2025 Jailbreaking as a Reward Misspecification Problem
abstract
The widespread adoption of large language models (LLMs) has raised concerns about their safety and reliability, particularly regarding their vulnerability to adversarial attacks. In this paper, we propose a new perspective that attributes this vulnerability to reward misspecification during the alignment process. This misspecification occurs when the reward function fails to accurately capture the intended behavior, leading to misaligned model outputs. We introduce a metric ReGap to quantify the extent of reward misspecification and demonstrate its effectiveness and robustness in detecting harmful backdoor prompts. Building upon these insights, we present ReMiss, a system for automated red teaming that generates adversarial prompts in a reward-misspecified space. ReMiss achieves state-of-the-art attack success rates on the AdvBench benchmark against various target aligned LLMs while preserving the human readability of the generated prompts. Furthermore, these attacks on open-source models demonstrate high transferability to closed-source models like GPT-4o and out-of-distribution tasks from HarmBench. Detailed analysis highlights the unique advantages of the proposed reward misspecification objective compared to previous methods, offering new insights for improving LLM safety and robustness.
Zhihui Xie 0002, Jiahui Gao 0002, Lei Li 0039, Zhenguo Li, Qi Liu 0049, Lingpeng Kong
ICLR5
2025 Temporal Reasoning Transfer from Text to Video
abstract
Video Large Language Models (Video LLMs) have shown promising capabilities in video comprehension, yet they struggle with tracking temporal changes and reasoning about temporal relationships. While previous research attributed this limitation to the ineffective temporal encoding of visual inputs, our diagnostic study reveals that video representations contain sufficient information for even small probing classifiers to achieve perfect accuracy. Surprisingly, we find that the key bottleneck in Video LLMs' temporal reasoning capability stems from the underlying LLM's inherent difficulty with temporal concepts, as evidenced by poor performance on textual temporal question-answering tasks. Building on this discovery, we introduce the Textual Temporal reasoning Transfer (T3). T3 synthesizes diverse temporal reasoning tasks in pure text format from existing image-text datasets, addressing the scarcity of video samples with complex temporal scenarios. Remarkably, without using any video data, T3 enhances LongVA-7B's temporal understanding, yielding a 5.3 absolute accuracy improvement on the challenging TempCompass benchmark, which enables our model to outperform ShareGPT4Video-8B trained on 28,000 video samples. Additionally, the enhanced LongVA-7B model achieves competitive performance on comprehensive video benchmarks. For example, it achieves a 49.7 accuracy on the Temporal Reasoning task of Video-MME, surpassing powerful large-scale models such as InternVL-Chat-V1.5-20B and VILA1.5-40B. Further analysis reveals a strong correlation between textual and video temporal task performance, validating the efficacy of transferring temporal reasoning abilities from text to video domains.
Lei Li 0039, Yuanxin Liu, Linli Yao, Peiyuan Zhang, Chenxin An, Lean Wang, Xu Sun 0001, Lingpeng Kong, Qi Liu 0049
ICLR9
2025 TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos
abstract
The rapid growth of online video platforms, particularly live streaming services, has created an urgent need for real-time video understanding systems. These systems must process continuous video streams and respond to user queries instantaneously, presenting unique challenges for current Video Large Language Models (VideoLLMs). While existing VideoLLMs excel at processing complete videos, they face significant limitations in streaming scenarios due to their inability to handle dense, redundant frames efficiently. We introduce TimeChat-Online, a novel online VideoLLM that revolutionizes real-time video interaction. At its core lies our innovative Differential Token Drop (DTD) module, which addresses the fundamental challenge of visual redundancy in streaming videos. Drawing inspiration from human visual perception's Change Blindness phenomenon, DTD preserves meaningful temporal changes while filtering out static, redundant content between frames. Remarkably, our experiments demonstrate that DTD achieves an 82.8% reduction in video tokens while maintaining 98% performance on StreamingBench, revealing that over 80% of visual content in streaming videos is naturally redundant without requiring language guidance. To enable seamless real-time interaction, we present TimeChat-Online-139K, a comprehensive streaming video dataset featuring diverse interaction patterns including backward-tracing, current-perception, and future-responding scenarios. TimeChat-Online's unique Proactive Response capability, naturally achieved through continuous monitoring of video scene transitions via DTD, sets it apart from conventional approaches. Our extensive evaluation demonstrates TimeChat-Online's superior performance on streaming benchmarks (StreamingBench and OvOBench) and maintaining competitive results on long-form video tasks such as Video-MME and MLVU. Notably, when integrated with Qwen2.5VL-7B, DTD achieves a 5.7-point accuracy improvement on the challenging VideoMME subset containing videos of 30-60 minutes, while reducing video tokens by 84.6%. Project page: https://timechat-online.github.io.
Linli Yao, Yuancheng Wei, Lei Li 0039, Shuhuai Ren, Yuanxin Liu, Kun Ouyang, Lean Wang, Lingpeng Kong, Qi Liu 0049, Yuanxing Zhang, Xu Sun 0001
ACM Multimedia12
2025 ImgTrojan: Jailbreaking Vision-Language Models with ONE Image
abstract
Xijia Tao, Shuai Zhong, Lei Li, Qi Liu, Lingpeng Kong. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Xijia Tao, Shuai Zhong, Lei Li 0039, Qi Liu 0049, Lingpeng Kong
NAACL (Long Papers)4
2024 Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models
abstract
Large vision-language models (LVLMs) excel across diverse tasks involving concrete images from natural scenes.However, their ability to interpret abstract figures, such as geometry shapes and scientific plots, remains limited due to a scarcity of training datasets in scientific domains.To fill this gap, we introduce Multimodal ArXiv, consisting of ArXivCap and ArXivQA, for enhancing LVLMs scientific comprehension.ArXivCap is a figure-caption dataset comprising 6.4M images and 3.9M captions, sourced from 572K ArXiv papers spanning various scientific domains.Drawing from ArXivCap, we introduce ArXivQA, a questionanswering dataset generated by prompting GPT-4V based on scientific figures.ArXivQA greatly enhances open-sourced LVLMs' mathematical reasoning capabilities, achieving a 10.4% absolute accuracy gain on a multimodal mathematical reasoning benchmark.Furthermore, employing ArXivCap, we devise four vision-to-text tasks for benchmarking LVLMs.Evaluation results with state-of-the-art LVLMs underscore their struggle with the nuanced semantics of academic figures, while domainspecific training yields substantial performance gains.Our error analysis uncovers misinterpretations of visual context, recognition errors, and the production of overly simplified captions by current LVLMs, shedding light on future improvements.
Lei Li 0039, Yuqi Wang 0003, Runxin Xu, Peiyi Wang, Xiachong Feng, Lingpeng Kong, Qi Liu 0049
ACL (1)7
2024 Large Language Models are not Fair Evaluators
abstract
Peiyi Wang, Lei Li, Liang Chen, Zefan Cai, Dawei Zhu, Binghuai Lin, Yunbo Cao, Lingpeng Kong, Qi Liu, Tianyu Liu, Zhifang Sui. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Peiyi Wang, Lei Li 0039, Liang Chen 0024, Zefan Cai, Binghuai Lin, Yunbo Cao, Lingpeng Kong, Qi Liu 0049, Tianyu Liu 0001, Zhifang Sui
ACL (1)9
2024 VLFeedback: A Large-Scale AI Feedback Dataset for Large Vision-Language Models Alignment
abstract
Lei Li, Zhihui Xie, Mukai Li, Shunian Chen, Peiyi Wang, Liang Chen, Yazheng Yang, Benyou Wang, Lingpeng Kong, Qi Liu. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Lei Li 0039, Zhihui Xie 0002, Mukai Li, Shunian Chen, Peiyi Wang, Liang Chen 0024, Yazheng Yang, Benyou Wang, Lingpeng Kong, Qi Liu 0049
EMNLP10
2024 Retrieved Sequence Augmentation for Protein Representation Learning
abstract
Chang Ma, Haiteng Zhao, Lin Zheng, Jiayi Xin, Qintong Li, Lijun Wu, Zhihong Deng, Yang Young Lu, Qi Liu, Sheng Wang, Lingpeng Kong. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Haiteng Zhao, Jiayi Xin, Qintong Li, Zhi-Hong Deng 0001, Qi Liu 0049, Lingpeng Kong
EMNLP9
2024 UniTabE: A Universal Pretraining Protocol for Tabular Foundation Model in Data Science
abstract
Recent advancements in Natural Language Processing (NLP) have witnessed the groundbreaking impact of pretrained models, yielding impressive outcomes across various tasks. This study seeks to extend the power of pretraining methodologies to facilitating the prediction over tables in data science, a domain traditionally overlooked, yet inherently challenging due to the plethora of table schemas intrinsic to different tasks. The primary research questions underpinning this work revolve around the establishment of a universal pretraining protocol for tables with varied structures, the generalizability and transferability of learned knowledge across tasks, the adaptation to diverse downstream applications, and the incorporation of incremental columns over time. In response to these challenges, we introduce UniTabE, a straightforward yet effective method designed to process tables in a uniform manner, devoid of constraints imposed by specific table structures. UniTabE's core concept relies on representing each basic table element with a module, termed TabUnit. This is subsequently followed by a Transformer encoder to refine the representation. Moreover, our model is designed to facilitate pretraining and finetuning through the utilization of free-form prompts. In order to implement the pretraining phase, we curated an expansive tabular dataset comprising approximately 13 billion samples, meticulously gathered from the Kaggle platform. This research primarily centers on classification and regression tasks involving tabular data, and conducts rigorous experimental testing and analyses to validate the effectiveness of our methodology. The experimental results demonstrate UniTabE's superior performance against several baseline models across a multitude of benchmark datasets. This, therefore, underscores UniTabE's potential to significantly enhance the semantic representation of tabular data, thereby marking a significant stride for tabular data analysis.
Yazheng Yang, Yuqi Wang 0003, Guang Liu 0006, Ledell Wu, Qi Liu 0049
ICLR5
2023 An Empirical Study of Retrieval-Enhanced Graph Neural Networks
abstract
Graph Neural Networks (GNNs) are effective tools for graph representation learning. Most GNNs rely on a recursive neighborhood aggregation scheme, named message passing, thereby their theoretical expressive power is limited to the first-order Weisfeiler-Lehman test (1-WL). An effective approach to this challenge is to explicitly retrieve some annotated examples used to enhance GNN models. While retrieval-enhanced models have been proved to be effective in many language and vision domains, it remains an open question how effective retrieval-enhanced GNNs are when applied to graph datasets. Motivated by this, we want to explore how the retrieval idea can help augment the useful information learned in the graph neural networks, and we design a retrieval-enhanced scheme called GRAPHRETRIEVAL, which is agnostic to the choice of graph neural network models. In GRAPHRETRIEVAL, for each input graph, similar graphs together with their ground-true labels are retrieved from an existing database. Thus they can act as a potential enhancement to complete various graph property predictive tasks. We conduct comprehensive experiments over 13 datasets, and we observe that GRAPHRETRIEVAL is able to reach substantial improvements over existing GNNs. Moreover, our empirical study also illustrates that retrieval enhancement is a promising remedy for alleviating the long-tailed label distribution problem.
Dingmin Wang, Shengchao Liu, Hanchen Wang 0002, Bernardo Cuenca Grau, Linfeng Song, Jian Tang 0005, Qi Liu 0049
ECAI8
2023 Can Language Models Understand Physical Concepts?
abstract
Language models (LMs) gradually become general-purpose interfaces in the interactive and embodied world, where the understanding of physical concepts is an essential prerequisite.However, it is unclear whether LMs can understand physical concepts in the human world.To investigate this, we design a benchmark VEC that covers the tasks of (i) Visual concepts, such as the shape and material of objects, and (ii) Embodied Concepts, learned from the interaction with the world such as the temperature of objects.Our zero (few)-shot prompting results show that the understanding of certain visual concepts emerges as scaling up LMs, but there are still basic concepts to which the scaling law does not apply.For example, OPT-175B performs close to humans with a zero-shot accuracy of 85% on the material concept, yet behaves like random guessing on the mass concept.Instead, vision-augmented LMs such as CLIP and BLIP achieve a human-level understanding of embodied concepts.Analysis indicates that the rich semantics in visual representation can serve as a valuable source of embodied knowledge.Inspired by this, we propose a distillation method to transfer embodied knowledge from VLMs to LMs, achieving performance gain comparable with that by scaling up parameters of LMs 134×. 1 o 1 : This is a photo of the water.o 2 : This is a photo of a frying oil.Attribute: This is a photo of a cold object.
Lei Li 0039, Jingjing Xu 0001, Qingxiu Dong, Xu Sun 0001, Lingpeng Kong, Qi Liu 0049
EMNLP7
2023 MSSRNet: Manipulating Sequential Style Representation for Unsupervised Text Style Transfer
abstract
Unsupervised text style transfer task aims to rewrite a text into target style while preserving its main content. Traditional methods rely on the use of a fixed-sized vector to regulate text style, which is difficult to accurately convey the style strength for each individual token. In fact, each token of a text contains different style intensity and makes different contribution to the overall style. Our proposed method addresses this issue by assigning individual style vector to each token in a text, allowing for fine-grained control and manipulation of the style strength. Additionally, an adversarial training framework integrated with teacher-student learning is introduced to enhance training stability and reduce the complexity of high-dimensional optimization. The results of our experiments demonstrate the efficacy of our method in terms of clearly improved style transfer accuracy and content preservation in both two-style transfer and multi-style transfer settings.
Yazheng Yang, Zhou Zhao 0001, Qi Liu 0049
KDD3
2023 Evaluating Self-Supervised Learning for Molecular Graph Embeddings
abstract
Graph Self-Supervised Learning (GSSL) provides a robust pathway for acquiring embeddings without expert labelling, a capability that carries profound implications for molecular graphs due to the staggering number of potential molecules and the high cost of obtaining labels. However, GSSL methods are designed not for optimisation within a specific domain but rather for transferability across a variety of downstream tasks. This broad applicability complicates their evaluation. Addressing this challenge, we present "Molecular Graph Representation Evaluation" (MOLGRAPHEVAL), generating detailed profiles of molecular graph embeddings with interpretable and diversified attributes. MOLGRAPHEVAL offers a suite of probing tasks grouped into three categories: (i) generic graph, (ii) molecular substructure, and (iii) embedding space properties. By leveraging MOLGRAPHEVAL to benchmark existing GSSL methods against both current downstream datasets and our suite of tasks, we uncover significant inconsistencies between inferences drawn solely from existing datasets and those derived from more nuanced probing. These findings suggest that current evaluation methodologies fail to capture the entirety of the landscape.
Hanchen Wang 0002, Jean Kaddour, Shengchao Liu, Jian Tang 0005, Joan Lasenby, Qi Liu 0049
NeurIPS6
2023 GIMLET: A Unified Graph-Text Model for Instruction-Based Molecule Zero-Shot Learning
abstract
Molecule property prediction has gained significant attention in recent years. The main bottleneck is the label insufficiency caused by expensive lab experiments. In order to alleviate this issue and to better leverage textual knowledge for tasks, this study investigates the feasibility of employing natural language instructions to accomplish molecule-related tasks in a zero-shot setting. We discover that existing molecule-text models perform poorly in this setting due to inadequate treatment of instructions and limited capacity for graphs. To overcome these issues, we propose GIMLET, which unifies language models for both graph and text data. By adopting generalized position embedding, our model is extended to encode both graph structures and instruction text without additional graph encoding modules. GIMLET also decouples encoding of the graph from tasks instructions in the attention mechanism, enhancing the generalization of graph features across novel tasks. We construct a dataset consisting of more than two thousand molecule tasks with corresponding instructions derived from task descriptions. We pretrain GIMLET on the molecule tasks along with instructions, enabling the model to transfer effectively to a broad range of tasks. Experimental results demonstrate that GIMLET significantly outperforms molecule-text baselines in instruction-based zero-shot learning, even achieving closed results to supervised GNN models on tasks such as toxcast and muv.
Haiteng Zhao, Shengchao Liu, Hannan Xu, Jie Fu 0001, Zhi-Hong Deng 0001, Lingpeng Kong, Qi Liu 0049
NeurIPS8
2023 Search-engine-augmented dialogue response generation with cheaply supervised query production
Ante Wang, Linfeng Song, Qi Liu 0049, Haitao Mi, Longyue Wang, Zhaopeng Tu, Jinsong Su, Dong Yu 0001
Artif. Intell.3
2023 Investigating Pose Representations and Motion Contexts Modeling for 3D Motion Prediction
abstract
Predicting human motion from historical pose sequence is crucial for a machine to succeed in intelligent interactions with humans. One aspect that has been obviated so far, is the fact that how we represent the skeletal pose has a critical impact on the prediction results. Yet there is no effort that investigates across different pose representation schemes. We conduct an indepth study on various pose representations with a focus on their effects on the motion prediction task. Moreover, recent approaches build upon off-the-shelf RNN units for motion prediction. These approaches process input pose sequence sequentially and inherently have difficulties in capturing long-term dependencies. In this paper, we propose a novel RNN architecture termed AHMR (Attentive Hierarchical Motion Recurrent network) for motion prediction which simultaneously models local motion contexts and a global context. We further explore a geodesic loss and a forward kinematics loss for the motion prediction task, which have more geometric significance than the widely employed L2 loss. Interestingly, we applied our method to a range of articulate objects including human, fish, and mouse. Empirical results show that our approach outperforms the state-of-the-art methods in short-term prediction and achieves much enhanced long-term prediction proficiency, such as retaining natural human-like motions over 50 seconds predictions. Our codes are released.
Zhenguang Liu, Shuang Wu 0002, Shuyuan Jin, Shouling Ji, Qi Liu 0049, Shijian Lu, Li Cheng 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2022 Relational Memory-Augmented Language Models
abstract
Abstract We present a memory-augmented approach to condition an autoregressive language model on a knowledge graph. We represent the graph as a collection of relation triples and retrieve relevant relations for a given context to improve text generation. Experiments on WikiText-103, WMT19, and enwik8 English datasets demonstrate that our approach produces a better language model in terms of perplexity and bits per character. We also show that relational memory improves coherence, is complementary to token-based memory, and enables causal interventions. Our model provides a simple yet effective way to combine an autoregressive language model and a knowledge graph for more coherent and logical generation.
Qi Liu 0049, Dani Yogatama, Phil Blunsom
Trans. Assoc. Comput. Linguistics1
2021 Unsupervised Point Cloud Pre-training via Occlusion Completion
abstract
We describe a simple pre-training approach for point clouds. It works in three steps: 1. Mask all points occluded in a camera view; 2. Learn an encoder-decoder model to reconstruct the occluded points; 3. Use the encoder weights as initialisation for downstream point cloud tasks. We find that even when we pre-train on a single dataset (ModelNet40), this method improves accuracy across different datasets and encoders, on a wide range of downstream tasks. Specifically, we show that our method outperforms previous pre-training methods in object classification, and both part-based and semantic segmentation tasks. We study the pre-trained features and find that they lead to wide downstream minima, have high transformation invariance, and have activations that are highly correlated with part labels. Code and data are available at: https://github.com/hansen7/OcCo
Hanchen Wang 0002, Qi Liu 0049, Xiangyu Yue 0001, Joan Lasenby, Matt J. Kusner
ICCV2
2021 Counterfactual Data Augmentation for Neural Machine Translation
abstract
We propose a data augmentation method for neural machine translation.It works by interpreting language models and phrasal alignment causally.Specifically, it creates augmented parallel translation corpora by generating (path-specific) counterfactual aligned phrases.We generate these by sampling new source phrases from a masked language model, then sampling an aligned counterfactual target phrase by noting that a translation language model can be interpreted as a Gumbel-Max Structural Causal Model (Oberst and Sontag, 2019).Compared to previous work, our method takes both context and alignment into account to maintain the symmetry between source and target sequences.Experiments on IWSLT'15 English → Vietnamese, WMT'17 English → German, WMT'18 English → Turkish, and WMT'19 robust English → French show that the method can improve the performance of translation, backtranslation and translation robustness.
Qi Liu 0049, Matt J. Kusner, Phil Blunsom
NAACL-HLT1
2021 Fast and Scalable Dialogue State Tracking with Explicit Modular Decomposition
abstract
This is a repository copy of Fast and scalable dialogue state tracking with explicit modular decomposition.
Dingmin Wang, Chenghua Lin 0002, Qi Liu 0049, Kam-Fai Wong
NAACL-HLT3
2021 Causal Effect Inference for Structured Treatments
abstract
We address the estimation of conditional average treatment effects (CATEs) for structured treatments (e.g., graphs, images, texts). Given a weak condition on the effect, we propose the generalized Robinson decomposition, which (i) isolates the causal estimand (reducing regularization bias), (ii) allows one to plug in arbitrary models for learning, and (iii) possesses a quasi-oracle convergence guarantee under mild assumptions. In experiments with small-world and molecular graphs we demonstrate that our approach outperforms prior work in CATE estimation.
Jean Kaddour, Qi Liu 0049, Matt J. Kusner, Ricardo Silva 0001
NeurIPS3
2021 Pretraining the Noisy Channel Model for Task-Oriented Dialogue
abstract
Abstract Direct decoding for task-oriented dialogue is known to suffer from the explaining-away effect, manifested in models that prefer short and generic responses. Here we argue for the use of Bayes’ theorem to factorize the dialogue task into two models, the distribution of the context given the response, and the prior for the response itself. This approach, an instantiation of the noisy channel model, both mitigates the explaining-away effect and allows the principled incorporation of large pretrained models for the response prior. We present extensive experiments showing that a noisy channel model decodes better responses compared to direct decoding and that a two-stage pretraining strategy, employing both open-domain and task-oriented dialogue data, improves over randomly initialized models.
Qi Liu 0049, Lei Yu 0008, Laura Rimell, Phil Blunsom
Trans. Assoc. Comput. Linguistics1
2020 Multi-Task Self-Supervised Learning for Disfluency Detection
abstract
Most existing approaches to disfluency detection heavily rely on human-annotated data, which is expensive to obtain in practice. To tackle the training data bottleneck, we investigate methods for combining multiple self-supervised tasks-i.e., supervised tasks where data can be collected without manual labeling. First, we construct large-scale pseudo training data by randomly adding or deleting words from unlabeled news data, and propose two self-supervised pre-training tasks: (i) tagging task to detect the added noisy words. (ii) sentence classification to distinguish original sentences from grammatically-incorrect sentences. We then combine these two tasks to jointly train a network. The pre-trained network is then fine-tuned using human-annotated disfluency detection training data. Experimental results on the commonly used English Switchboard test set show that our approach can achieve competitive performance compared to the previous systems (trained using the full dataset) by using less than 1% (1000 sentences) of the training data. Our method trained on the full dataset significantly outperforms previous methods, reducing the error by 21% on English Switchboard.
Shaolei Wang, Wanxiang Che, Qi Liu 0049, Pengda Qin, Ting Liu 0001, William Yang Wang
AAAI3
2020 Smart Contract Vulnerability Detection using Graph Neural Network
abstract
The security problems of smart contracts have drawn extensive attention due to the enormous financial losses caused by vulnerabilities. Existing methods on smart contract vulnerability detection heavily rely on fixed expert rules, leading to low detection accuracy. In this paper, we explore using graph neural networks (GNNs) for smart contract vulnerability detection. Particularly, we construct a contract graph to represent both syntactic and semantic structures of a smart contract function. To highlight the major nodes, we design an elimination phase to normalize the graph. Then, we propose a degree-free graph convolutional neural network (DR-GCN) and a novel temporal message propagation network (TMP) to learn from the normalized graphs for vulnerability detection. Extensive experiments show that our proposed approach significantly outperforms state-of-the-art methods in detecting three different types of vulnerabilities.
Zhenguang Liu, Qi Liu 0049, Qinming He
IJCAI4
2019 Towards Natural and Accurate Future Motion Prediction of Humans and Animals
abstract
Anticipating the future motions of 3D articulate objects is challenging due to its non-linear and highly stochastic nature. Current approaches typically represent the skeleton of an articulate object as a set of 3D joints, which unfortunately ignores the relationship between joints, and fails to encode fine-grained anatomical constraints. Moreover, conventional recurrent neural networks, such as LSTM and GRU, are employed to model motion contexts, which inherently have difficulties in capturing long-term dependencies. To address these problems, we propose to explicitly encode anatomical constraints by modeling their skeletons with a Lie algebra representation. Importantly, a hierarchical recurrent network structure is developed to simultaneously encodes local contexts of individual frames and global contexts of the sequence. We proceed to explore the applications of our approach to several distinct quantities including human, fish, and mouse. Extensive experiments show that our approach achieves more natural and accurate predictions over state-of-the-art methods.
Zhenguang Liu, Shuang Wu 0002, Shuyuan Jin, Qi Liu 0049, Shijian Lu, Roger Zimmermann, Li Cheng 0001
CVPR4
2019 Quaternion Knowledge Graph Embeddings
abstract
In this work, we move beyond the traditional complex-valued representations, introducing more expressive hypercomplex representations to model entities and relations for knowledge graph embeddings. More specifically, quaternion embeddings, hypercomplex-valued embeddings with three imaginary components, are utilized to represent entities. Relations are modelled as rotations in the quaternion space. The advantages of the proposed approach are: (1) Latent inter-dependencies (between all components) are aptly captured with Hamilton product, encouraging a more compact interaction between entities and relations; (2) Quaternions enable expressive rotation in four-dimensional space and have more degree of freedom than rotation in complex plane; (3) The proposed framework is a generalization of ComplEx on hypercomplex space while offering better geometrical interpretations, concurrently satisfying the key desiderata of relational representation learning (i.e., modeling symmetry, anti-symmetry and inversion). Experimental results demonstrate that our method achieves state-of-the-art performance on four well-established knowledge graph completion benchmarks.
Shuai Zhang 0007, Yi Tay, Lina Yao 0001, Qi Liu 0049
NeurIPS4
2019 Hyperbolic Graph Neural Networks
abstract
Learning from graph-structured data is an important task in machine learning and artificial intelligence, for which Graph Neural Networks (GNNs) have shown great promise. Motivated by recent advances in geometric representation learning, we propose a novel GNN architecture for learning representations on Riemannian manifolds with differentiable exponential and logarithmic maps. We develop a scalable algorithm for modeling the structural properties of graphs, comparing Euclidean and hyperbolic geometry. In our experiments, we show that hyperbolic GNNs can lead to substantial improvements on various benchmark datasets.
Qi Liu 0049, Maximilian Nickel, Douwe Kiela
NeurIPS1
2019 Insertion-based Decoding with Automatically Inferred Generation Order
abstract
Conventional neural autoregressive decoding commonly assumes a fixed left-to-right generation order, which may be sub-optimal. In this work, we propose a novel decoding algorithm— InDIGO—which supports flexible sequence generation in arbitrary orders through insertion operations. We extend Transformer, a state-of-the-art sequence generation model, to efficiently implement the proposed approach, enabling it to be trained with either a pre-defined generation order or adaptive orders obtained from beam-search. Experiments on four real-world tasks, including word order recovery, machine translation, image caption, and code generation, demonstrate that our algorithm can generate sequences following arbitrary orders, while achieving competitive or even better performance compared with the conventional left-to-right generation. The generated sequences show that InDIGO adopts adaptive generation orders based on input information.
Jiatao Gu, Qi Liu 0049, Kyunghyun Cho
Trans. Assoc. Comput. Linguistics2
2018 Multi-Modal Multi-Task Learning for Automatic Dietary Assessment
abstract
We investigate the task of automatic dietary assessment: given meal images and descriptions uploaded by real users, our task is to automatically rate the meals and deliver advisory comments for improving users' diets. To address this practical yet challenging problem, which is multi-modal and multi-task in nature, an end-to-end neural model is proposed. In particular, comprehensive meal representations are obtained from images, descriptions and user information. We further introduce a novel memory network architecture to store meal representations and reason over the meal representations to support predictions. Results on a real-world dataset show that our method outperforms two strong image captioning baselines significantly.
Qi Liu 0049, Yue Zhang 0004, Zhenguang Liu, Ye Yuan 0001, Li Cheng 0001, Roger Zimmermann
AAAI1
2018 Sentence-State LSTM for Text Representation
abstract
Bi-directional LSTMs are a powerful tool for text representation.On the other hand, they have been shown to suffer various limitations due to their sequential nature.We investigate an alternative LSTM structure for encoding text, which consists of a parallel state for each word.Recurrent steps are used to perform local and global information exchange between words simultaneously, rather than incremental reading of a sequence of words.Results on various classification and sequence labelling benchmarks show that the proposed model has strong representation power, giving highly competitive performances compared to stacked BiLSTM models with similar parameter numbers.
Yue Zhang 0004, Qi Liu 0049, Linfeng Song
ACL (1)2
2018 Mining Evidences for Concept Stock Recommendation
abstract
We investigate the task of mining relevant stocks given a topic of concern on emerging capital markets, for which there is lack of structural understanding.Deep learning is leveraged to mine evidences from large scale textual data, which contain valuable market information.In particular, distributed word similarities trained over large scale raw texts are taken as a basis of relevance measuring, and deep reinforcement learning is leveraged to learn a strategy of topic expansion, given a small amount of manually labeled data from financial analysts.Results on two Chinese stock market datasets show that our method outperforms a strong baseline using information retrieval techniques.
Qi Liu 0049, Yue Zhang 0004
NAACL-HLT1
2018 Learning Domain Representation for Multi-Domain Sentiment Classification
abstract
Qi Liu, Yue Zhang, Jiangming Liu. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Qi Liu 0049, Yue Zhang 0004, Jiangming Liu
NAACL-HLT1
2018 Constrained Graph Variational Autoencoders for Molecule Design
abstract
Graphs are ubiquitous data structures for representing interactions between entities. With an emphasis on applications in chemistry, we explore the task of learning to generate graphs that conform to a distribution observed in training data. We propose a variational autoencoder model in which both encoder and decoder are graph-structured. Our decoder assumes a sequential ordering of graph extension steps and we discuss and analyze design choices that mitigate the potential downsides of this linearization. Experiments compare our approach with a wide range of baselines on the molecule generation task and show that our method is successful at matching the statistics of the original dataset on semantically important metrics. Furthermore, we show that by using appropriate shaping of the latent space, our model allows us to design molecules that are (locally) optimal in desired properties.
Qi Liu 0049, Miltiadis Allamanis, Marc Brockschmidt, Alexander L. Gaunt
NeurIPS1
2018 Toward Personalized Activity Level Prediction in Community Question Answering Websites
abstract
Community Question Answering (CQA) websites have become valuable knowledge repositories. Millions of internet users resort to CQA websites to seek answers to their encountered questions. CQA websites provide information far beyond a search on a site such as Google due to (1) the plethora of high-quality answers, and (2) the capabilities to post new questions toward the communities of domain experts. While most research efforts have been made to identify experts or to preliminarily detect potential experts of CQA websites, there has been a remarkable shift toward investigating how to keep the engagement of experts. Experts are usually the major contributors of high-quality answers and questions of CQA websites. Consequently, keeping the expert communities active is vital to improving the lifespan of these websites. In this article, we present an algorithm termed PALP to predict the activity level of expert users of CQA websites. To the best of our knowledge, PALP is the first approach to address a personalized activity level prediction model for CQA websites. Furthermore, it takes into consideration user behavior change over time and focuses specifically on expert users. Extensive experiments on the Stack Overflow website demonstrate the competitiveness of PALP over existing methods.
Zhenguang Liu, Yingjie Xia, Qi Liu 0049, Qinming He, Chao Zhang 0014, Roger Zimmermann
ACM Trans. Multim. Comput. Commun. Appl.3
2017 QALink: Enriching Text Documents with Relevant Q&A Site Contents
abstract
With rapid development of Q&A sites such as Quora and StackExchange, high quality question-answer pairs have been produced by users. These Q&A contents cover a wide range of topics, and they are useful for users to resolve queries and obtain new knowledge. Meanwhile, when people are reading digital documents, they may encounter reading problems such as lack of background information and unclear illustration of concepts. We believe that Q&A sites offer high-quality contents which can serve as rich supplements to digital documents. In this paper, we devise a rigorous formulation of the novel text enrichment problem, and design an end-to-end system named QALink which assigns the most relevant Q&A contents to the corresponding section of the document. We first present a new segmentation approach to model each document with a hierarchical structure. Based on the hierarchy, queries are constructed to retrieve and rank related question-answer pairs. Both syntactical and semantic features are adopted in our system. The empirical evaluation results indicate that QALink is able to effectively enrich text documents with relevant Q&A contents to help people better understand the documents.
Weilong Huang, Qi Liu 0049, Anthony K. H. Tung, Xiaoli Wang 0002, Jisong Yang
CIKM3
2017 EtherQL: A Query Layer for Blockchain System
Kai Zheng 0001, Ying Yan 0006, Qi Liu 0049, Xiaofang Zhou 0001
DASFAA (2)4
2017 Behavior pattern clustering in blockchain networks
Butian Huang, Zhenguang Liu, Jianhai Chen, Anan Liu, Qi Liu 0049, Qinming He
Multim. Tools Appl.5
2017 Fusion of Magnetic and Visual Sensors for Indoor Localization: Infrastructure-Free and More Effective
abstract
Accurate and infrastructure-free indoor positioning can be very useful in a variety of applications. However, most existing approaches (e.g., WiFi and infrared-based methods) for indoor localization heavily rely on infrastructure, which is neither scalable nor pervasively available. In this paper, we propose a novel indoor localization and tracking approach, termed VMag, that does not require any infrastructure assistance. The user can be localized while simply holding a smartphone. To the best of our knowledge, the proposed method is the first exploration of fusing geomagnetic and visual sensing for indoor localization. More specifically, we conduct an in-depth study on both the advantageous properties and the challenges in leveraging the geomagnetic field and visual images for indoor localization. Based on these studies, we design a context-aware particle filtering framework to track the user with the goal of maximizing the positioning accuracy. We also introduce a neural-network-based method to extract deep features for the purpose of indoor positioning. We have conducted extensive experiments on four different indoor settings including a laboratory, a garage, a canteen, and an office building. Experimental results demonstrate the superior performance of VMag over the state of the art with these four indoor settings.
Zhenguang Liu, Qi Liu 0049, Yifang Yin, Li Cheng 0001, Roger Zimmermann
IEEE Trans. Multim.3
2015 DocRicher: An Automatic Annotation System for Text Documents Using Social Media
abstract
We demonstrate a system, DocRicher, to enrich a text document with social media, that implicitly reference certain passages of it. The aim is to provide an automatic annotation interface to satisfy users' information need, without cumbersome queries to traditional search engines. The system consists of four components: text analysis, query construction, data assignment, and user feedback. Through text analysis, the system decomposes a text document into appropriate topical passages, of which each is represented using detected key phrases. By submitting combinations of these phrases as queries to social media systems, the relevant results are used to suggest new annotations, that are linked to the corresponding passages. We have built a user-friendly visualization tool for users to browse automatically recommended annotations on their reading documents. Users are either allowed to rate a recommended annotation by accepting it or not; or add a new annotation by manually highlighting texts and adding personal comments. Both these annotations are regarded as the ground truth to derive new queries for retrieving more relevant contents. We also apply data fusion to merge the query results from various contexts and retain most relevant ones.
Qi Liu 0049, Xiaoli Wang 0002, Anthony K. H. Tung, Shubham Goyal, Jisong Yang
SIGMOD Conference2