Zeju Qiu

dblp:276/4222 · DBLP profile ↗
← Back
7ranked-venue papers
6as first author
6since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 6 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Efficient and distributed learning · 62% Language models and text generation · 12% Generative modeling · 12%
Computer graphics and multimedia
1 paper
Geometric modeling and processing · 100%

Topics — the 10 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning › parameter-efficient fine-tuning
orthogonal finetuning
2.332025
Orthogonal Finetuning Made Scalable · EMNLP 2025
Parameter-Efficient Orthogonal Finetuning via Butterfly Factorization · ICLR 2024
Controlling Text-to-Image Diffusion by Orthogonal Finetuning · NeurIPS 2023
Machine learning › Efficient and distributed learning
parameter-efficient fine-tuning
2.332025
Orthogonal Finetuning Made Scalable · EMNLP 2025
Parameter-Efficient Orthogonal Finetuning via Butterfly Factorization · ICLR 2024
Controlling Text-to-Image Diffusion by Orthogonal Finetuning · NeurIPS 2023
Machine learning › Efficient and distributed learning
model compression
1.622025
Orthogonal Finetuning Made Scalable · EMNLP 2025
Parameter-Efficient Orthogonal Finetuning via Butterfly Factorization · ICLR 2024
Natural language and speech › Language models and text generation
large language model training
0.912025
Reparameterized LLM Training via Orthogonal Equivalence Transformation · NeurIPS 2025
Computer vision › Vision and language
multimodal reasoning
0.912025
Can Large Language Models Understand Symbolic Graphics Programs? · ICLR 2025
Machine learning › Efficient and distributed learning › model compression › quantization
quantized model fine-tuning
0.912025
Orthogonal Finetuning Made Scalable · EMNLP 2025
Machine learning › Generative modeling › diffusion model
controllable generation
0.712023
Controlling Text-to-Image Diffusion by Orthogonal Finetuning · NeurIPS 2023
Machine learning › Generative modeling › diffusion model › text-to-image generation
text-to-image diffusion model
0.712023
Controlling Text-to-Image Diffusion by Orthogonal Finetuning · NeurIPS 2023
Natural language and speech › Language models and text generation
instruction tuning
0.312025
Can Large Language Models Understand Symbolic Graphics Programs? · ICLR 2025
Natural language and speech › Language models and text generation › large language model › large language model adaptation › personalization
subject-driven generation
0.212023
Controlling Text-to-Image Diffusion by Orthogonal Finetuning · NeurIPS 2023

Methods — techniques the papers use, named apart from their topics

symbolic instruction tuning · 1.7benchmarking · 1.7truncated neumann series · 0.9spectral normalization · 0.9orthogonal matrix · 0.9matrix-free computation · 0.9cayley-neumann parameterization · 0.9orthogonal parameterization · 0.8fast fourier transform · 0.8butterfly structure · 0.8
YearPublicationVenuePosition
2025 Orthogonal Finetuning Made Scalable
abstract
Orthogonal finetuning (OFT) offers highly parameter-efficient adaptation while preventing catastrophic forgetting, but its high runtime and memory demands limit practical deployment. We identify the core computational bottleneck in OFT as its weight-centric implementation, which relies on costly matrix-matrix multiplications with cubic complexity. To overcome this, we propose OFTv2, an input-centric reformulation that instead uses matrix-vector multiplications (i.e., matrix-free computation), reducing the computational cost to quadratic. We further introduce the Cayley–Neumann parameterization, an efficient orthogonal parameterization that approximates the matrix inversion in the Cayley transform via a truncated Neumann series. These modifications allow OFTv2 to achieve up to 10x faster training and 3x lower GPU memory usage without compromising performance. In addition, we extend OFTv2 to support finetuning quantized foundation models and show that it outperforms the popular QLoRA in training stability, efficiency, and memory usage.
Zeju Qiu, Weiyang Liu, Adrian Weller, Bernhard Schölkopf
EMNLP1
2025 Can Large Language Models Understand Symbolic Graphics Programs?
abstract
Against the backdrop of enthusiasm for large language models (LLMs), there is a growing need to scientifically assess their capabilities and shortcomings. This is nontrivial in part because it is difficult to find tasks which the models have not encountered during training. Utilizing symbolic graphics programs, we propose a domain well-suited to test multiple spatial-semantic reasoning skills of LLMs. Popular in computer graphics, these programs procedurally generate visual data. While LLMs exhibit impressive skills in general program synthesis and analysis, symbolic graphics programs offer a new layer of evaluation: they allow us to test an LLM's ability to answer semantic questions about the images or 3D geometries without a vision encoder. To semantically understand the symbolic programs, LLMs would need to possess the ability to "imagine" and reason how the corresponding graphics content would look with only the symbolic description of the local curvatures and strokes. We use this task to evaluate LLMs by creating a large benchmark for the semantic visual understanding of symbolic graphics programs, built procedurally with minimal human effort. Particular emphasis is placed on transformations of images that leave the image level semantics invariant while introducing significant changes to the underlying program. We evaluate commercial and open-source LLMs on our benchmark to assess their ability to reason about visual output of programs, finding that LLMs considered stronger at reasoning generally perform better. Lastly, we introduce a novel method to improve this ability -- Symbolic Instruction Tuning (SIT), in which the LLM is finetuned with pre-collected instruction data on symbolic graphics programs. Interestingly, we find that SIT not only improves LLM's understanding on symbolic programs, but it also improves general reasoning ability on various other benchmarks.
Zeju Qiu, Weiyang Liu, Haiwen Feng, Zhen Liu 0019, Tim Z. Xiao, Katie Collins, Josh Tenenbaum, Adrian Weller, Michael J. Black, Bernhard Schölkopf
ICLR1
2025 Reparameterized LLM Training via Orthogonal Equivalence Transformation
abstract
While large language models (LLMs) are driving the rapid advancement of artificial intelligence, effectively and reliably training these large models remains one of the field's most significant challenges. To address this challenge, we propose POET, a novel reParameterized training algorithm that uses Orthogonal Equivalence Transformation to optimize neurons. Specifically, POET reparameterizes each neuron with two learnable orthogonal matrices and a fixed random weight matrix. Because of its provable preservation of spectral properties of weight matrices, POET can stably optimize the objective function with improved generalization. We further develop efficient approximations that make POET flexible and scalable for training large-scale neural networks. Extensive experiments validate the effectiveness and scalability of POET in training LLMs.
Zeju Qiu, Simon Buchholz, Tim Z. Xiao, Maximilian Dax, Bernhard Schölkopf, Weiyang Liu
NeurIPS1
2024 Parameter-Efficient Orthogonal Finetuning via Butterfly Factorization
abstract
Large foundation models are becoming ubiquitous, but training them from scratch is prohibitively expensive. Thus, efficiently adapting these powerful models to downstream tasks is increasingly important. In this paper, we study a principled finetuning paradigm -- Orthogonal Finetuning (OFT) -- for downstream task adaptation. Despite demonstrating good generalizability, OFT still uses a fairly large number of trainable parameters due to the high dimensionality of orthogonal matrices. To address this, we start by examining OFT from an information transmission perspective, and then identify a few key desiderata that enable better parameter-efficiency. Inspired by how the Cooley-Tukey fast Fourier transform algorithm enables efficient information transmission, we propose an efficient orthogonal parameterization using butterfly structures. We apply this parameterization to OFT, creating a novel parameter-efficient finetuning method, called Orthogonal Butterfly (BOFT). By subsuming OFT as a special case, BOFT introduces a generalized orthogonal finetuning framework. Finally, we conduct an extensive empirical study of adapting large vision transformers, large language models, and text-to-image diffusion models to various downstream tasks in computer vision and natural language. The results validate the effectiveness of BOFT as a generic finetuning method.
Weiyang Liu, Zeju Qiu, Yao Feng 0001, Yuliang Xiu, Yuxuan Xue 0001, Longhui Yu, Haiwen Feng, Zhen Liu 0019, Juyeon Heo, Songyou Peng, Yandong Wen, Michael J. Black, Adrian Weller, Bernhard Schölkopf
ICLR2
2023 Iterative Teaching by Data Hallucination
abstract
We consider the problem of iterative machine teaching, where a teacher sequentially provides examples based on the status of a learner under a discrete input space (i.e., a pool of finite samples), which greatly limits the teacher’s capability. To address this issue, we study iterative teaching under a continuous input space where the input example (i.e., image) can be either generated by solving an optimization problem or drawn directly from a continuous distribution. Specifically, we propose data hallucination teaching (DHT) where the teacher can generate input data intelligently based on labels, the learner’s status and the target concept. We study a number of challenging teaching setups (e.g., linear/neural learners in omniscient and black-box settings). Extensive empirical results verify the effectiveness of DHT.
Zeju Qiu, Weiyang Liu, Tim Z. Xiao, Zhen Liu 0019, Umang Bhatt, Yucen Luo, Adrian Weller, Bernhard Schölkopf
AISTATS1
2023 Controlling Text-to-Image Diffusion by Orthogonal Finetuning
abstract
Large text-to-image diffusion models have impressive capabilities in generating photorealistic images from text prompts. How to effectively guide or control these powerful models to perform different downstream tasks becomes an important open problem. To tackle this challenge, we introduce a principled finetuning method -- Orthogonal Finetuning (OFT), for adapting text-to-image diffusion models to downstream tasks. Unlike existing methods, OFT can provably preserve hyperspherical energy which characterizes the pairwise neuron relationship on the unit hypersphere. We find that this property is crucial for preserving the semantic generation ability of text-to-image diffusion models. To improve finetuning stability, we further propose Constrained Orthogonal Finetuning (COFT) which imposes an additional radius constraint to the hypersphere. Specifically, we consider two important finetuning text-to-image tasks: subject-driven generation where the goal is to generate subject-specific images given a few images of a subject and a text prompt, and controllable generation where the goal is to enable the model to take in additional control signals. We empirically show that our OFT framework outperforms existing methods in generation quality and convergence speed.
Zeju Qiu, Weiyang Liu, Haiwen Feng, Yuxuan Xue 0001, Yao Feng 0001, Zhen Liu 0019, Adrian Weller, Bernhard Schölkopf
NeurIPS1
2020 Hand Pose-based Task Learning from Visual Observations with Semantic Skill Extraction
abstract
Learning from Demonstrations is a promising technique to transfer task knowledge from a user to a robot. We propose a framework for task programming by observing the human hand pose and object locations solely with a depth camera. By extracting skills from the demonstrations, we are able to represent what the robot has learned, generalize to unseen object locations and optimize the robotic execution instead of replaying a non-optimal behavior. A two-staged segmentation algorithm that employs skill template matching via Hidden Markov Models has been developed to extract motion primitives from the demonstration and gives them semantic meanings. In this way, the transfer of task knowledge has been improved from a simple replay of the demonstration towards a semantically annotated, optimized and generalized execution. We evaluated the extraction of a set of skills in simulation and prove that the task execution can be optimized by such means.
Zeju Qiu, Thomas Eiband, Shile Li, Dongheui Lee
RO-MAN1