Jiaqi Liao

dblp:144/7470 · DBLP profile ↗
← Back
11ranked-venue papers
4as first author
10since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 3 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Theory of computation · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Isodiametric Inequality for Vector Spaces
abstract
Abstract. A theorem of Kleitman states that a collection of binary vectors with diameter [Formula: see text] has cardinality at most that of a Hamming ball of radius [Formula: see text]. In this paper, we give a [Formula: see text]-analogue of it.
Jiaqi Liao, Guiying Yan
SIAM J. Discret. Math.1
2025 LangBridge: Interpreting Image as a Combination of Language Embeddings
abstract
Recent years have witnessed remarkable advances in Large Vision-Language Models (LVLMs), which have achieved human-level performance across various complex vision-language tasks. Following LLaVA's paradigm, mainstream LVLMs typically employ a shallow MLP for visual-language alignment through a two-stage training process: pretraining for cross-modal alignment followed by instruction tuning. While this approach has proven effective, the underlying mechanisms of how MLPs bridge the modality gap remain poorly understood. Although some research has explored how LLMs process transformed visual tokens, few studies have investigated the fundamental alignment mechanism. Furthermore, the MLP adapter requires retraining whenever switching LLM backbones. To address these limitations, we first investigate the working principles of MLP adapters and discover that they learn to project visual embeddings into subspaces spanned by corresponding text embeddings progressively. Based on this insight, we propose LangBridge, a novel adapter that explicitly maps visual tokens to linear combinations of LLM vocabulary embeddings. This innovative design enables pretraining-free adapter transfer across different LLMs while maintaining performance. Our experimental results demonstrate that a LangBridge adapter pre-trained on Qwen2-0.5B can be directly applied to larger models such as LLaMA3-8B or Qwen2.5-14B while maintaining competitive performance. Overall, LangBridge enables interpretable vision-language alignment by grounding visual representations in LLM vocab embedding, while its plug-and-play design ensures efficient reuse across multiple LLMs with nearly no performance degradation. See our project page at https://curryx-001.github.io/LangBridge.github.io/
Jiaqi Liao, Yuwei Niu, Fanqing Meng, Hao Li 0069, Changyao Tian, Yinuo Du, Yuwen Xiong, Dianqi Li, Xizhou Zhu, Jifeng Dai, Yu Cheng 0001
ICCV1
2025 ImageGen-CoT: Enhancing Text-to-Image in-context Learning with Chain-of-Thought Reasoning
Jiaqi Liao, Zhengyuan Yang, Dianqi Li, Yu Cheng 0001
ICCV1
2025 MMIU: Multimodal Multi-image Understanding for Evaluating Large Vision-Language Models
abstract
The capability to process multiple images is crucial for Large Vision-Language Models (LVLMs) to develop a more thorough and nuanced understanding of a scene. Recent multi-image LVLMs have begun to address this need. However, their evaluation has not kept pace with their development. To fill this gap, we introduce the Multimodal Multi-image Understanding (MMIU) benchmark, a comprehensive evaluation suite designed to assess LVLMs across a wide range of multi-image tasks. MMIU encompasses 7 types of multi-image relationships, 52 tasks, 77K images, and 11K meticulously curated multiple-choice questions, making it the most extensive benchmark of its kind. Our evaluation of nearly 30 popular LVLMs, including both open-source and proprietary models, reveals significant challenges in multi-image comprehension, particularly in tasks involving spatial understanding. Even the most advanced models, such as GPT-4o, achieve only 55.7\% accuracy on MMIU. Through multi-faceted analytical experiments, we identify key performance gaps and limitations, providing valuable insights for future model and data improvements. We aim for MMIU to advance the frontier of LVLM research and development. We release the data and code at https://github.com/MMIUBenchmark/MMIU.
Fanqing Meng, Chuanhao Li 0001, Quanfeng Lu, Hao Tian 0006, Tianshuo Yang, Jiaqi Liao, Xizhou Zhu, Jifeng Dai, Yu Qiao 0001, Ping Luo 0002, Kaipeng Zhang, Wenqi Shao
ICLR7
2025 Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation
abstract
Text-to-video (T2V) models like Sora have made significant strides in visualizing complex prompts, which is increasingly viewed as a promising path towards constructing the universal world simulator. Cognitive psychologists believe that the foundation for achieving this goal is the ability to understand intuitive physics. However, the capacity of these models to accurately represent intuitive physics remains largely unexplored. To bridge this gap, we introduce PhyGenBench, a comprehensive \textbf{Phy}sics \textbf{Gen}eration \textbf{Ben}chmark designed to evaluate physical commonsense correctness in T2V generation. PhyGenBench comprises 160 carefully crafted prompts across 27 distinct physical laws, spanning four fundamental domains, which could comprehensively assesses models' understanding of physical commonsense. Alongside PhyGenBench, we propose a novel evaluation framework called PhyGenEval. This framework employs a hierarchical evaluation structure utilizing appropriate advanced vision-language models and large language models to assess physical commonsense. Through PhyGenBench and PhyGenEval, we can conduct large-scale automated assessments of T2V models' understanding of physical commonsense, which align closely with human feedback. Our evaluation results and in-depth analysis demonstrate that current models struggle to generate videos that comply with physical commonsense. Moreover, simply scaling up models or employing prompt engineering techniques is insufficient to fully address the challenges presented by PhyGenBench (e.g., dynamic scenarios). We hope this study will inspire the community to prioritize the learning of physical commonsense in these models beyond entertainment applications. We will release the data and codes at https://github.com/OpenGVLab/PhyGenBench
Fanqing Meng, Jiaqi Liao, Quanfeng Lu, Wenqi Shao, Kaipeng Zhang, Yu Cheng 0001, Dianqi Li, Ping Luo 0002
ICML2
2025 VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models
abstract
Recent advancements in text-to-video (T2V) diffusion models have enabled high-fidelity and realistic video synthesis. However, current T2V models often struggle to generate physically plausible content due to their limited inherent ability to accurately understand physics. We found that while the representations within T2V models possess some capacity for physics understanding, they lag significantly behind those from recent video self-supervised learning methods. To this end, we propose a novel framework called {VideoREPA}, which distills physics understanding capability from video understanding foundation models into T2V models by aligning token-level relations. This closes the physics understanding gap and enables more physics-plausible generation. Specifically, we introduce the {Token Relation Distillation (TRD) loss}, leveraging spatio-temporal alignment to provide soft guidance suitable for finetuning powerful pre-trained T2V models—a critical departure from prior representation alignment (REPA) methods. To our knowledge, VideoREPA is the first REPA method designed for finetuning T2V models and specifically for injecting physical knowledge. Empirical evaluations show that VideoREPA substantially enhances the physics commonsense of baseline method, CogVideoX, achieving significant improvement on relevant benchmarks and demonstrating a strong capacity for generating videos consistent with intuitive physics. Code and more video results are available at https://videorepa.github.io/.
Jiaqi Liao, Shaofeng Zhang, Fanqing Meng, Xiangpeng Wan, Junchi Yan, Yu Cheng 0001
NeurIPS2
2025 KTransformers: Unleashing the Full Potential of CPU/GPU Hybrid Inference for MoE Models
abstract
Due to the sparse nature of Mixture-of-Experts (MoE) models, they are particularly suitable for hybrid CPU/GPU inference, especially in low-concurrency scenarios. This hybrid approach leverages both the large, cost-effective memory capacity of CPU/DRAM and the high bandwidth of GPU/VRAM. However, existing hybrid solutions remain bottlenecked by CPU computation limits and CPU-GPU synchronization overheads, severely restricting their ability to efficiently run state-of-the-art large MoE models, such as the 671B DeepSeek-V3/R1.
Hongtao Chen, Weiyu Xie, Boxin Zhang, Jingqi Tang, Shaoyuan Chen, Ziwei Yuan, Chengyu Qiu, Yuening Zhu, Qingliang Ou, Jiaqi Liao, Xianglin Chen, Zhiyuan Ai, Yongwei Wu 0001
SOSP13
2025 Identifying the drivers of new product development performance change in China's high-tech industry: A two-stage production-theoretical decomposition analysis
Yu Yu 0011, Jiaqi Liao, Xianmei Wang
Expert Syst. Appl.2
2024 HardMo: A Large-Scale Hardcase Dataset for Motion Capture
abstract
Recent years have witnessed rapid progress in monoc-ular human mesh recovery. Despite their impressive performance on public benchmarks, existing methods are vulnerable to unusual poses, which prevents them from deploying to challenging scenarios such as dance and martial arts. This issue is mainly attributed to the domain gap induced by the data scarcity in relevant cases. Most existing datasets are captured in constrained scenarios and lack samples of such complex movements. For this reason, we propose a data collection pipeline comprising automatic crawling, precise annotation, and hardcase mining. Based on this pipeline, we establish a large dataset in a short time. The dataset, named HardMo, contains 7M images along with precise annotations covering 15 categories of dance and 14 categories of martial arts. Empirically, we find that the prediction failure in dance and martial arts is mainly characterized by the misalignment of hand-wrist and foot-ankle. To dig deeper into the two hardcases, we leverage the proposed automatic pipeline to filter collected data and construct two subsets named HardMo-Hand and HardMo-Foot. Extensive experiments demonstrate the effectiveness of the annotation pipeline and the data-driven solution to failure cases. Specifically, after being trained on HardMo, HMR, an early pioneering method, can even outperform the current state of the art, 4DHumans, on our benchmarks. Dataset will be publicly available at https://ljqnb.github.io/HardMo.github.io.
Jiaqi Liao, Chuanchen Luo, Yinuo Du, Yuxi Wang 0001, Xu-Cheng Yin, Man Zhang 0005, Zhaoxiang Zhang 0001, Junran Peng
CVPR1
2024 LaneDAG: Automatic HD Map Topology Generator Based on Geometry and Attention Fusion Mechanism
abstract
In high-definition maps (HD maps), the road lane centerline and lane topology graph play essential roles in navigation, planning, and decision-making. Existing research focusing on extracting physical infrastructure, such as lane boundaries, has made significant progress. But lane centerline detection and topology reasoning still remains challenging due to the severe overlapping centerlines and complicated topology. To tackle these challenges, we introduce an automatic lane topology extraction method for HD maps, termed LaneDAG, which extracts vectorized centerlines and their topology from prebuilt lane lines and road boundaries in HD maps. It formulates centerline extraction as a set prediction problem and lane topology prediction as a directed acyclic graph (DAG) construction problem. A novel mechanism that fusing geometric and attention-based features in the DAG is proposed to model the topological relationship between centerlines. Experiments conducted on the Argoverse 2 dataset demonstrate the proposed method’s superior performance compared to existing methods, showcasing its capability to extract lane centerlines and topology in HD maps automatically.
Peijin Jia, Tuopu Wen, Ziang Luo, Zheng Fu, Jiaqi Liao, Huixian Chen, Kun Jiang 0002, Mengmeng Yang 0001, Diange Yang
IV5
2015 Online-SVR for vehicular position prediction during GPS outages using low-cost INS
abstract
Vehicle position prediction has become more and more critical for most applications in intelligent transportation systems (ITS). Prediction based INS/GPS integration provides continuous and reliable navigation solution when compared to standalone Inertial Navigation System (INS) or Global Positioning System (GPS). Although there have been several research works for fusing INS and GPS data to bridge navigation during GPS outages, most of them are offline methods and do not consider sensors data fluctuation due to traffic incident, inclement weather conditions or rush hour. This paper proposes a supervised statistical learning technique called Online Support Vector Machine for Regression (OL-SVR) for the prediction of vehicle position. During GPS availability, the OL-SVR models INS errors by fusing the INS and GPS data; meanwhile during outages, the trained OL-SVR method is utilized to predict accurate vehicle position. The proposed method is compared with two well-known prediction techniques including Partial Least Squares Regression (PLSR) and Artificial Neural Network (ANN). Experiments conducted at rush hour on real urban roads and simulation results prove that OL-SVR is more efficient and accurate in position prediction than PLSR and ANN, achieving an accuracy improvement of 20.3%-64.8%.
Dong Wang 0016, Jiaqi Liao, Zhu Xiao, Xiaohong Li 0004, Vincent Havyarimana
PIMRC2