Zhiyang Xu

dblp:267/2280 · DBLP profile ↗
← Back
16ranked-venue papers
3as first author
16since 2021 · last 2026
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 3 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Artificial intelligence - assisted pump-locked random lasers
Junhua Tong, Zhiyang Xu, Naeem Iqbal, Kun Ge, Tianrui Zhai
Sci. China Inf. Sci.3
2025 R2I-Bench: Benchmarking Reasoning-Driven Text-to-Image Generation
abstract
based on instance-specific, reasoning-oriented evaluation questions that assess three critical dimensions: text-image alignment, reasoning accuracy, and image quality.Extensive experiments with 17 representative T2I models, including a strong pipeline-based framework that decouples reasoning and generation using the state-of-the-art language and image generation models, demonstrate consistently limited reasoning performance, highlighting the need for more robust, reasoning-aware architectures in the next generation of T2I systems.
Kaijie Chen, Zihao Lin 0003, Zhiyang Xu, Ying Shen 0006, Yuguang Yao, Joy Rimchala, Jiaxin Zhang 0005, Lifu Huang
EMNLP3
2025 SPARTUN3D: Situated Spatial Understanding of 3D World in Large Language Model
abstract
Integrating the 3D world into large language models (3D-based LLMs) has been a promising research direction for 3D scene understanding. However, current 3D-based LLMs fall short in situated understanding due to two key limitations: 1) existing 3D datasets are constructed from a global perspective of the 3D scenes and lack situated context. 2) the architectures of the current 3D-based LLMs lack an explicit mechanism for aligning situated spatial information between 3D representations and natural language, limiting their performance in tasks requiring precise spatial reasoning. In this work, we address these issues by introducing a scalable situated 3D dataset, named Spartun3D, that incorporates various situated spatial information. In addition, we propose a situated spatial alignment module to enhance the learning between 3D visual representations and their corresponding textual descriptions. Our experimental results demonstrate that both our dataset and alignment module enhance situated spatial understanding ability.
Yue Zhang 0004, Zhiyang Xu, Ying Shen 0001, Parisa Kordjamshidi, Lifu Huang
ICLR2
2025 Modality-Specialized Synergizers for Interleaved Vision-Language Generalists
abstract
Recent advancements in Vision-Language Models (VLMs) have led to the emergence of Vision-Language Generalists (VLGs) capable of understanding and generating both text and images. However, seamlessly generating an arbitrary sequence of text and images remains a challenging task for the current VLGs. One primary limitation lies in applying a unified architecture and the same set of parameters to simultaneously model discrete text tokens and continuous image features. Recent works attempt to tackle this fundamental problem by introducing modality-aware expert models. However, they employ identical architectures to process both text and images, disregarding the intrinsic inductive biases in these two modalities. In this work, we introduce Modality-Specialized Synergizers (MoSS), a novel design that efficiently optimizes existing unified architectures of VLGs with modality-specialized adaptation layers, i.e., a Convolutional LoRA for modeling the local priors of image patches and a Linear LoRA for processing sequential text. This design enables more effective modeling of modality-specific features while maintaining the strong cross-modal integration gained from pretraining. In addition, to improve the instruction-following capability on interleaved text-and-image generation, we introduce LeafInstruct, the first open-sourced interleaved instruction tuning dataset comprising 184,982 high-quality instances on more than 10 diverse domains. Extensive experiments show that VLGs integrated with MoSS achieve state-of-the-art performance, significantly surpassing baseline VLGs in complex interleaved generation tasks. Furthermore, our method exhibits strong generalizability on different VLGs.
Zhiyang Xu, Minqian Liu, Ying Shen 0006, Joy Rimchala, Jiaxin Zhang 0005, Qifan Wang 0001, Lifu Huang
ICLR1
2025 UniHGKR: Unified Instruction-aware Heterogeneous Knowledge Retrievers
abstract
Dehai Min, Zhiyang Xu, Guilin Qi, Lifu Huang, Chenyu You. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Dehai Min, Zhiyang Xu, Guilin Qi, Lifu Huang, Chenyu You
NAACL (Long Papers)2
2025 AR-RAG: Autoregressive Retrieval Augmentation for Image Generation
abstract
We introduce Autoregressive Retrieval Augmentation (AR-RAG), a novel paradigm that enhances image generation by autoregressively incorporating k-nearest neighbor retrievals at the patch level. Unlike prior methods that perform a single, static retrieval before generation and condition the entire generation on fixed reference images, AR-RAG performs context-aware retrievals at each generation step, using prior-generated patches as queries to retrieve and incorporate the most relevant patch-level visual references, enabling the model to respond to evolving generation needs while avoiding limitations (e.g., over-copying, stylistic bias, etc.) prevalent in existing methods. To realize AR-RAG, we propose two parallel frameworks: (1) Distribution-Augmentation in Decoding (DAiD), a training-free plug-and-use decoding strategy that directly merges the distribution of model-predicted patches with the distribution of retrieved patches, and (2) Feature-Augmentation in Decoding (FAiD), a parameter-efficient fine-tuning method that progressively smooths the features of retrieved patches via multi-scale convolution operations and leverages them to augment the image generation process. We validate the effectiveness of AR-RAG on widely adopted benchmarks, including Midjourney-30K, GenEval and DPG-Bench, demonstrating significant performance gains over state-of-the-art image generation models.
Jingyuan Qi, Zhiyang Xu, Qifan Wang 0001, Lifu Huang
NeurIPS2
2024 MULTISCRIPT: Multimodal Script Learning for Supporting Open Domain Everyday Tasks
abstract
Automatically generating scripts (i.e. sequences of key steps described in text) from video demonstrations and reasoning about the subsequent steps are crucial to the modern AI virtual assistants to guide humans to complete everyday tasks, especially unfamiliar ones. However, current methods for generative script learning rely heavily on well-structured preceding steps described in text and/or images or are limited to a certain domain, resulting in a disparity with real-world user scenarios. To address these limitations, we present a new benchmark challenge – MULTISCRIPT, with two new tasks on task-oriented multimodal script learning: (1) multimodal script generation, and (2) subsequent step prediction. For both tasks, the input consists of a target task name and a video illustrating what has been done to complete the target task, and the expected output is (1) a sequence of structured step descriptions in text based on the demonstration video, and (2) a single text description for the subsequent step, respectively. Built from WikiHow, MULTISCRIPT covers multimodal scripts in videos and text descriptions for over 6,655 human everyday tasks across 19 diverse domains. To establish baseline performance on MULTISCRIPT, we propose two knowledge-guided multimodal generative frameworks that incorporate the task-related knowledge prompted from large language models such as Vicuna. Experimental results show that our proposed approaches significantly improve over the competitive baselines.
Jingyuan Qi, Minqian Liu, Ying Shen 0006, Zhiyang Xu, Lifu Huang
AAAI4
2024 Multimodal Instruction Tuning with Conditional Mixture of LoRA
abstract
Multimodal Large Language Models (MLLMs) have demonstrated remarkable proficiency in diverse tasks across different domains, with an increasing focus on improving their zeroshot generalization capabilities for unseen multimodal tasks.Multimodal instruction tuning has emerged as a successful strategy for achieving zero-shot generalization by fine-tuning pretrained models on diverse multimodal tasks through instructions.As MLLMs grow in complexity and size, the need for parameterefficient fine-tuning methods like Low-Rank Adaption (LoRA), which fine-tunes with a minimal set of parameters, becomes essential.However, applying LoRA in multimodal instruction tuning presents the challenge of task interference, which leads to performance degradation, especially when dealing with a broad array of multimodal tasks.To address this, this paper introduces a novel approach that integrates multimodal instruction tuning with Conditional Mixture-of-LoRA (MixLoRA).It innovates upon LoRA by dynamically constructing low-rank adaptation matrices tailored to the unique demands of each input instance, aiming to mitigate task interference.Experimental results on various multimodal evaluation datasets indicate that MixLoRA not only outperforms the conventional LoRA with the same or even higher ranks, demonstrating its efficacy and adaptability in diverse multimodal tasks 1 .
Ying Shen 0001, Zhiyang Xu, Qifan Wang 0001, Wenpeng Yin 0001, Lifu Huang
ACL (1)2
2024 Ameli: Enhancing Multimodal Entity Linking with Fine-Grained Attributes
abstract
Barry Yao, Sijia Wang, Yu Chen, Qifan Wang, Minqian Liu, Zhiyang Xu, Licheng Yu, Lifu Huang. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Barry Menglong Yao, Yu Chen 0022, Qifan Wang 0001, Minqian Liu, Zhiyang Xu, Licheng Yu, Lifu Huang
EACL (1)6
2024 Holistic Evaluation for Interleaved Text-and-Image Generation
abstract
Interleaved text-and-image generation has been an intriguing research direction, where the models are required to generate both images and text pieces in an arbitrary order.Despite the emerging advancements in interleaved generation, the progress in its evaluation still significantly lags behind.Existing evaluation benchmarks do not support arbitrarily interleaved images and text for both inputs and outputs, and they only cover a limited number of domains and use cases.Also, current works predominantly use similarity-based metrics which fall short in assessing the quality in open-ended scenarios.To this end, we introduce INTER-LEAVEDBENCH, the first benchmark carefully curated for the evaluation of interleaved textand-image generation.INTERLEAVEDBENCH features a rich array of tasks to cover diverse real-world use cases.In addition, we present INTERLEAVEDEVAL, a strong reference-free metric powered by GPT-4o to deliver accurate and explainable evaluation.We carefully define five essential evaluation aspects for IN-TERLEAVEDEVAL, including text quality, perceptual quality, image coherence, text-image coherence, and helpfulness, to ensure a comprehensive and fine-grained assessment.Through extensive experiments and rigorous human evaluation, we show that our benchmark and metric can effectively evaluate the existing models with a strong correlation with human judgments surpassing previous reference-based metrics.We also provide substantial findings and insights to foster future research in interleaved generation and its evaluation. 1
Minqian Liu, Zhiyang Xu, Zihao Lin 0003, Trevor Ashby, Joy Rimchala, Jiaxin Zhang 0005, Lifu Huang
EMNLP2
2024 X-Eval: Generalizable Multi-aspect Text Evaluation via Augmented Instruction Tuning with Auxiliary Evaluation Aspects
abstract
Minqian Liu, Ying Shen, Zhiyang Xu, Yixin Cao, Eunah Cho, Vaibhav Kumar, Reza Ghanadan, Lifu Huang. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Minqian Liu, Ying Shen 0006, Zhiyang Xu, Yixin Cao 0002, Eunah Cho, Vaibhav Kumar, Reza Ghanadan, Lifu Huang
NAACL-HLT3
2023 MultiInstruct: Improving Multi-Modal Zero-Shot Learning via Instruction Tuning
abstract
Instruction tuning, a new learning paradigm that fine-tunes pre-trained language models on tasks specified through instructions, has shown promising zero-shot performance on various natural language processing tasks.However, it has yet to be explored for vision and multimodal tasks.In this work, we introduce MUL-TIINSTRUCT, the first multimodal instruction tuning benchmark dataset that consists of 62 diverse multimodal tasks in a unified seq-toseq format covering 10 broad categories.The tasks are derived from 21 existing open-source datasets and each task is equipped with 5 expertwritten instructions.We take OFA (Wang et al., 2022a) as the base pre-trained model for multimodal instruction tuning, and to further improve its zero-shot performance, we explore multiple transfer learning strategies to leverage the large-scale NATURAL INSTRUCTIONS dataset (Mishra et al., 2022).Experimental results demonstrate strong zero-shot performance on various unseen multimodal tasks and the benefit of transfer learning from a text-only instruction dataset.We also design a new evaluation metric -Sensitivity, to evaluate how sensitive the model is to the variety of instructions.Our results indicate that fine-tuning the model on a diverse set of tasks and instructions leads to a reduced sensitivity to variations in instructions for each task 1 .
Zhiyang Xu, Ying Shen 0006, Lifu Huang
ACL (1)1
2023 The Art of SOCRATIC QUESTIONING: Recursive Thinking with Large Language Models
abstract
Chain-of-Thought (CoT) prompting enables large language models to solve complex reasoning problems by generating intermediate steps.However, confined by its inherent singlepass and sequential generation process, CoT heavily relies on the initial decisions, causing errors in early steps to accumulate and impact the final answers.In contrast, humans adopt recursive thinking when tackling complex reasoning problems, i.e., iteratively breaking the original problem into approachable subproblems and aggregating their answers to resolve the original one.Inspired by the human cognitive process, we propose SOCRATIC QUESTIONING, a divide-and-conquer style algorithm that mimics the recursive thinking process.Specifically, SOCRATIC QUESTIONING leverages large language models to raise and answer sub-questions until collecting enough information to tackle the original question.Unlike CoT, SOCRATIC QUESTIONING explicitly navigates the thinking space, stimulates effective recursive thinking, and is more robust towards errors in the thinking process.Extensive experiments on several complex reasoning tasks, including MMLU, MATH, LogiQA, and visual question-answering demonstrate significant performance improvements over the stateof-the-art prompting methods, such as CoT, and Tree-of-Thought.The qualitative analysis clearly shows that the intermediate reasoning steps elicited by SOCRATIC QUESTIONING are similar to humans' recursively thinking process of complex reasoning problems 12 .
Jingyuan Qi, Zhiyang Xu, Ying Shen 0006, Minqian Liu, Qifan Wang 0001, Lifu Huang
EMNLP2
2022 Structured Energy Network As a Loss
abstract
Belanger & McCallum (2016) and Gygli et al. (2017) have shown that an energy network can capture arbitrary dependencies amongst the output variables in structured prediction; however, their reliance on gradient-based inference (GBI) makes the inference slow and unstable. In this work, we propose Structured Energy As Loss (SEAL) to take advantage of the expressivity of energy networks without incurring the high inference cost. This is a novel learning framework that uses an energy network as a trainable loss function (loss-net) to train a separate neural network (task-net), which is then used to perform the inference through a forward pass. We establish SEAL as a general framework wherein various learning strategies like margin-based, regression, and noise-contrastive, could be employed to learn the parameters of loss-net. Through extensive evaluation on multi-label classification, semantic role labeling, and image segmentation, we demonstrate that SEAL provides various useful design choices, is faster at inference than GBI, and leads to significant performance gains over the baselines.
Jay-Yoon Lee, Dhruvesh Patel, Purujit Goyal, Wenlong Zhao 0001, Zhiyang Xu, Andrew McCallum
NeurIPS5
2022 RGB WGM lasing woven in fiber braiding cavity
Kun Ge, Zhiyang Xu, Jun Ruan, Libin Cui, Tianrui Zhai
Sci. China Inf. Sci.2
2021 Improved Latent Tree Induction with Distant Supervision via Span Constraints
abstract
Zhiyang Xu, Andrew Drozdov, Jay Yoon Lee, Tim O’Gorman, Subendhu Rongali, Dylan Finkbeiner, Shilpa Suresh, Mohit Iyyer, Andrew McCallum. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021.
Zhiyang Xu, Andrew Drozdov, Jay-Yoon Lee, Tim O'Gorman, Subendhu Rongali, Dylan Finkbeiner, Shilpa Suresh, Mohit Iyyer, Andrew McCallum
EMNLP (1)1