Yiming Liang

dblp:83/1858 · DBLP profile ↗
← Back
18ranked-venue papers
6as first author
16since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 5 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 SynthAgent: Adapting Web Agents with Synthetic Supervision
abstract
Zhaoyang Wang, Yiming Liang, Xuchao Zhang, Qianhui Wu, Siwei Han, Anson Bastos, Rujia Wang, Chetan Bansal, Baolin Peng, Jianfeng Gao, Saravan Rajmohan, Huaxiu Yao. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Zhaoyang Wang 0004, Yiming Liang, Xuchao Zhang, Qianhui Wu, Siwei Han, Anson Bastos, Rujia Wang, Chetan Bansal, Baolin Peng, Jianfeng Gao 0001, Saravan Rajmohan, Huaxiu Yao
ACL (1)2
2026 EfficientLLM: Unified Pruning-Aware Pretraining for Auto-Designed Compact Language Models
abstract
Xingrun Xing, Zheng Liu, Shitao Xiao, Boyan Gao, Yiming Liang, Haokun Lin, Xianlin Zeng, Guoqi Li, Jiajun Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Xingrun Xing, Shitao Xiao, Boyan Gao, Yiming Liang, Haokun Lin, Xianlin Zeng, Guoqi Li 0002
ACL (1)5
2026 Lexicalized Constituency Parsing for Middle Dutch: Low-resource Training and Cross-Domain Generalization
Yiming Liang
LREC1
2026 Dynamic weighted mixture-of-experts physics-informed learning for real-time microchannel cooling field reconstruction
Xingpu Feng, Yiming Liang, Sanli Liu, Simon Maher
Eng. Appl. Artif. Intell.2
2026 DiffuSteer: Conditional generative steering of large language models in activation space
Haoying Wang, Zhenyu Ding, Tengyue Xiao, Yiming Liang, Caigui Jiang, Ning Ding 0006
Neurocomputing4
2025 Can MLLMs Understand the Deep Implication Behind Chinese Images?
abstract
As the capabilities of Multimodal Large Language Models (MLLMs) improve, the need for higher-order evaluation of them is increasing. However, there is a lack of work evaluating MLLM for higher-order perception and understanding of Chinese visual content. To address this, we introduce the CII-Bench, which aims to assess MLLMs’ such capabilities for Chinese images. To ensure the authenticity of the Chinese context, images in CII-Bench are sourced from the Chinese Internet and manually reviewed, with corresponding answers also manually crafted. Additionally, CII-Bench incorporates images that represent Chinese traditional culture, such as famous Chinese traditional paintings, which can deeply reflect the model’s understanding of Chinese traditional culture. Through experiments on multiple MLLMs using CII-Bench, significant findings emerged. There is a large gap between MLLMs and humans in performance. The highest MLLM accuracy is 64.4%, while the human average is 78.2% and the peak is 81.0%. MLLMs perform poorly on traditional culture images, indicating limitations in understanding high-level semantics and lacking a deep knowledge base of Chinese traditional culture. Moreover, most models have higher accuracy when image emotion hints are added to the prompts. We believe CII-Bench will help MLLMs better understand Chinese semantics and specific images, and move forward the development of expert artificial general intelligence (AGI). Our project is publicly available at https://cii-bench.github.io.
Chenhao Zhang 0005, Yuelin Bai, Xeron Du, Jinchang Hou, Kaixin Deng, Guangzeng Han, Qinrui Li, Bingli Wang, Xingwei Qu, Qixuan Zhao, Yiming Liang, Feiteng Fang, Min Yang 0007, Wenhao Huang 0001, Chenghua Lin 0002, Ge Zhang 0009, Shiwen Ni
ACL (1)14
2025 HiMoR: Monocular Deformable Gaussian Reconstruction with Hierarchical Motion Representation
abstract
We present Hierarchical Motion Representation (HiMoR), a novel deformation representation for 3D Gaussian primitives capable of achieving high-quality monocular dynamic 3D reconstruction. The insight behind HiMoR is that motions in everyday scenes can be decomposed into coarser motions that serve as the foundation for finer details. Using a tree structure, HiMoR’s nodes represent different levels of motion detail, with shallower nodes modeling coarse motion for temporal smoothness and deeper nodes capturing finer motion. Additionally, our model uses a few shared motion bases to represent motions of different sets of nodes, aligning with the assumption that motion tends to be smooth and simple. This motion representation design provides Gaussians with a more structured deformation, maximizing the use of temporal relationships to tackle the challenging task of monocular dynamic 3D reconstruction. We also propose using a more reliable perceptual metric as an alternative, given that pixel-level metrics for evaluating monocular dynamic 3D reconstruction can sometimes fail to accurately reflect the true quality of reconstruction. Extensive experiments demonstrate our method’s efficacy in achieving superior novel view synthesis from challenging monocular videos with complex motions.
Yiming Liang, Tianhan Xu, Yuta Kikuchi
CVPR1
2025 MuPT: A Generative Symbolic Music Pretrained Transformer
abstract
In this paper, we explore the application of Large Language Models (LLMs) to the pre-training of music. While the prevalent use of MIDI in music modeling is well-established, our findings suggest that LLMs are inherently more compatible with ABC Notation, which aligns more closely with their design and strengths, thereby enhancing the model's performance in musical composition. To address the challenges associated with misaligned measures from different tracks during generation, we propose the development of a $\underline{S}$ynchronized $\underline{M}$ulti-$\underline{T}$rack ABC Notation ($\textbf{SMT-ABC Notation}$), which aims to preserve coherence across multiple musical tracks. Our contributions include a series of models capable of handling up to 8192 tokens, covering 90\% of the symbolic music data in our training set. Furthermore, we explore the implications of the $\underline{S}$ymbolic $\underline{M}$usic $\underline{S}$caling Law ($\textbf{SMS Law}$) on model performance. The results indicate a promising research direction in music generation, offering extensive resources for further research through our open-source contributions.
Xingwei Qu, Yuelin Bai, Yinghao Ma, Ziya Zhou, Ka Man Lo, Ruibin Yuan, Lejun Min, Xueling Liu 0001, Xeron Du, Shuyue Guo, Yiming Liang, Shangda Wu, Junting Zhou, Tianyu Zheng, Ziyang Ma 0001, Fengze Han, Wei Xue 0002, Gus Xia, Emmanouil Benetos, Xiang Yue, Chenghua Lin 0002, Xu Tan 0003, Wenhao Huang 0001, Jie Fu 0001, Ge Zhang 0009
ICLR13
2025 OmniBench: Towards The Future of Universal Omni-Language Models
abstract
Recent advancements in multimodal large language models (MLLMs) have focused on integrating multiple modalities, yet their ability to simultaneously process and reason across different inputs remains underexplored. We introduce OmniBench, a novel benchmark designed to evaluate models’ ability to recognize, interpret, and reason across visual, acoustic, and textual inputs simultaneously. We define language models capable of such tri-modal processing as omni-language models (OLMs). OmniBench features high-quality human annotations that require integrated understanding across all modalities. Our evaluation reveals that: i) open-source OLMs show significant limitations in instruction-following and reasoning in tri-modal contexts; and ii) most baseline models perform poorly (below 50% accuracy) even with textual alternatives to image/audio inputs. To address these limitations, we develop OmniInstruct, an 96K-sample instruction tuning dataset for training OLMs. We advocate for developing more robust tri-modal integration techniques and training strategies to enhance OLM performance. Codes and data could be found at https://m-a-p.ai/OmniBench/.
Ge Zhang 0009, Yinghao Ma, Ruibin Yuan, Kang Zhu, Hangyu Guo, Yiming Liang, Noah Wang, Jian Yang 0003, Siwei Wu, Xingwei Qu, Jinjie Shi, Xinyue Zhang 0005, Zhenzhu Yang, Yidan Wen, Yanghai Wang, Zhaoxiang Zhang 0001, Ruibo Liu, Emmanouil Benetos, Wenhao Huang 0001, Chenghua Lin 0002
NeurIPS7
2025 SimWorld: An Open-ended Simulator for Agents in Physical and Social Worlds
abstract
While LLM/VLM-powered AI agents have advanced rapidly in math, coding, and computer use, their applications in complex physical and social environments remain challenging. Building agents that can survive and thrive in the real world (e.g., by autonomously earning income) requires massive-scale interaction, reasoning, training, and evaluation across diverse scenarios. However, existing world simulators for such development fall short: they often rely on limited hand-crafted environments, simulate simplified game-like physics and social rules, and lack native support for LLM/VLM agents. We introduce SimWorld, a new simulator built on Unreal Engine 5, designed for developing and evaluating LLM/VLM agents in rich, real-world-like settings. SimWorld offers three core capabilities: (1) realistic, open-ended world simulation, including accurate physical and social dynamics and language-driven procedural environment generation; (2) rich interface for LLM/VLM agents, with multi-modal world inputs/feedback and open-vocabulary action outputs at varying levels of abstraction; and (3) diverse physical and social reasoning scenarios that are easily customizable by users. We demonstrate SimWorld by deploying frontier LLM agents (e.g., Gemini-2.5-Flash, Claude-3.5, GPT-4o, and DeepSeek-Prover-V2) on both short-horizon navigation tasks requiring grounded re-planning, and long-horizon multi-agent food delivery tasks involving strategic cooperation and competition. The results reveal distinct reasoning patterns and limitations across models. We open-source SimWorld and hope it becomes a foundational platform for advancing real-world agent intelligence across disciplines. Please refer to the project website for the most up-to-date information: http://simworld.org/.
Xiaokang Ye, Xuhong He, Yiming Liang, Yiqing Yang, Mrinaal Dogra, Xianrui Zhong, Eric Liu 0006, Kevin Benavente, Rajiv Mandya Nagaraju, Dhruv Vivek Sharma, Ziqiao Ma 0001, Tianmin Shu, Zhiting Hu, Lianhui Qin
NeurIPS5
2025 A Progressive Transformer for Unifying Binary Code Embedding and Knowledge Transfer
abstract
Language models have recently been applied to binary analysis tasks, such as function similarity detection and function signature recovery. These models typically employ a two-stage training process: pre-training via Masked Language Modeling (MLM) on machine code and fine-tuning for specific tasks. While MLM helps to understand binary code structures, it ignores essential code characteristics, including control and data flow, which negatively affect model generalization. Recent work leverages domain-specific features (e.g., control flow graphs and dynamic execution traces) in transformer-based approaches to improve binary code semantic understanding. This approach, however, involves complex feature engineering, a cumbersome and time-consuming process that can introduce predictive uncertainty when dealing with stripped or obfuscated code, which leads to a performance drop. We introduce PROTST, a novel transformer-based methodology for binary code embedding. PROTST employs a hierarchical training process based on a unique tree-like structure, where knowledge progressively flows from fundamental tasks at the root to more specialized tasks at the leaves. This progressive teacher-student paradigm allows the model to build upon previously learned knowledge, resulting in high-quality embeddings that can be effectively leveraged for diverse downstream binary analysis tasks. The effectiveness of PROTST is evaluated in seven binary analysis tasks, demonstrating an average of 14.8% improvement in F1 and MRR compared to traditional two-stage training, and a 16.6 % improvement when analyzing obfuscated code.
Hanxiao Lu, Hongyu Cai, Yiming Liang, Antonio Bianchi, Z. Berkay Celik
SANER3
2025 A lightweight and robust detection network for diverse glass surface defects via scale- and shape-aware feature extraction
Huan Yu 0002, Jin Wang 0015, Jingru Yang, Yiming Liang, Zhan Wang 0002, Haiyan He, Guodong Lu
Eng. Appl. Artif. Intell.4
2025 Toward Practical Autonomous Flight Simulation for Flapping Wing Biomimetic Robots With Experimental Validation
abstract
The utilization of a well-established flapping wing robot simulation holds significant importance in the advancement of flapping wing mechanisms and algorithms. This research paper introduces a pioneering application-oriented flapping wing robot simulation platform that exhibits high compatibility with diverse mechanical designs and adaptability to various robotic tasks. Initially, the paper presents the blade element theory and the quasi-steady model as computational approaches for determining the aerodynamics of flapping wings based on their kinematics. The computation incorporates translational lift, translational drag, rotational lift, added mass force, and clap-and-fling mechanism. The simulation platform is validated through flight control tasks, providing a comprehensive assessment of its performance. Moreover, this study addresses the challenges of attitude tracking and trajectory tracking control for a specially designed flapping wing robot. The proposed control strategies are evaluated through real flight experiments, offering practical insights into the robot flight capabilities. Note to Practitioners—One of the primary contributions of this research is the introduction of a novel and robust flapping wing flight controller that effectively addresses the challenges associated with attitude tracking control and positional trajectory tracking. The proposed controller demonstrates superior performance compared to existing algorithms, as evidenced by comparative simulations conducted on the developed simulation platform. Furthermore, real flight experiments are performed on a self-made flapping wing robot using the same control algorithm and parameters employed in the simulations. Another highlight lies in the transferability from simulation to reality. The utilization of a high-fidelity simulation platform enables in-depth exploration and scrutiny of complex behaviors manifested in diverse flight tasks. This, in turn, allows for meticulous analysis and examination to gain valuable insights into the practical implementation of flapping wing robot flight.
Yongchun Fang, Jifu Yan, Yiming Liang, Tiefeng Li
IEEE Trans Autom. Sci. Eng.5
2024 Uniform information density explains subject doubling in French
Yiming Liang, Pascal Amsili, Heather Burnett, Vera Demberg
CogSci1
2024 Random Frame: a Data Augmentation for Glass Detection
Yiming Liang
ICPR (33)1
2021 Inter-clausal Anaphora in Chinese Conditionals: a Multi-factorial Analysis
Shunting Chen, Pascal Amsili, Yiming Liang
PACLIC3
2018 LSTM Multiple Object Tracker Combining Multiple Cues
abstract
Traditional methods for multiple object tracking usually consider features at image level and reason about simple space and time constraints. However, in this paper we propose a multiple object tracker based on LSTM network to learn temporally correlated features. Our tracker learns features on velocity, position and appearance aspects of the objects to improve tracking accuracy. In order to deal with occlusion problems, velocity model is also used to predict the lost detections for the occluded identities. Besides, pre-association is performed before appearance model to improve the efficiency of calculation. Final association problem is solved by taking the best match. Experiments in MOT17 dataset show that our tracker achieves leading performance compared to other state-of-the-art algorithms.
Yiming Liang
ICIP1
2017 Multi-camera Tracking Exploiting Person Re-ID Technique
Yiming Liang, Yue Zhou 0005
ICONIP (3)1