Jiazheng Xu

dblp:313/9484 · DBLP profile ↗
← Back
13ranked-venue papers
2as first author
13since 2021 · last 2026
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 2 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 VisionReward: Fine-Grained Multi-Dimensional Human Preference Learning for Image and Video Generation
abstract
Visual generative models have achieved remarkable progress in synthesizing photorealistic images and videos, yet aligning their outputs with human preferences across critical dimensions remains a persistent challenge. Though reinforcement learning from human feedback offers promise for preference alignment, existing reward models for visual generation face limitations, including black-box scoring without interpretability and potentially resultant unexpected biases. We present VisionReward, a general framework for learning human visual preferences in both image and video generation. Specifically, we employ a hierarchical visual assessment framework to capture fine-grained human preferences, and leverages linear weighting to enable interpretable preference learning. Furthermore, we propose a multi-dimensional consistent strategy when using VisionReward as a reward model during preference optimization for visual generation. Experiments show that VisionReward can significantly outperform existing image and video reward models on both machine metrics and human evaluation. Notably, VisionReward surpasses VideoScore by 17.2% in preference prediction accuracy, and text-to-video models with VisionReward achieve a 31.6% higher pairwise win rate compared to the same models using VideoScore.
Jiazheng Xu, Yuanming Yang, Wenbo Duan, Shen Yang 0001, Qunlin Jin, Shurun Li, Jiayan Teng, Zhuoyi Yang, Wendi Zheng, Xiao Liu 0036, Ming Ding 0004, Shiyu Huang 0001, Xiaotao Gu, Minlie Huang, Jie Tang 0001, Yuxiao Dong
AAAI1
2025 LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks
abstract
Yushi Bai, Shangqing Tu, Jiajie Zhang, Hao Peng, Xiaozhi Wang, Xin Lv, Shulin Cao, Jiazheng Xu, Lei Hou, Yuxiao Dong, Jie Tang, Juanzi Li. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Yushi Bai, Shangqing Tu, Hao Peng 0015, Xiaozhi Wang, Shulin Cao, Jiazheng Xu, Lei Hou 0001, Yuxiao Dong, Jie Tang 0001, Juan-Zi Li
ACL (1)8
2025 AlignMMBench: Evaluating Chinese Multimodal Alignment in Large Vision-Language Models
abstract
Yuhang Wu, Wenmeng Yu, Yean Cheng, Yan Wang, Xiaohan Zhang, Jiazheng Xu, Ming Ding, Yuxiao Dong. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Wenmeng Yu, Yean Cheng, Yan Wang 0120, Jiazheng Xu, Ming Ding 0004, Yuxiao Dong
ACL (1)6
2025 On the Out-Of-Distribution Generalization of Large Multimodal Models
abstract
We investigate the generalization boundaries of current Large Multimodal Models (LMMs) under out-of-distribution scenarios and domain-specific tasks. We evaluate their zero-shot generalization across synthetic images, real-world distributional shifts, and specialized datasets like medical and molecular imagery. Empirical results indicate that LMMs struggle with generalization beyond common training domains, limiting their direct application without adaptation. To understand the cause of unreliable performance, we analyze three hypotheses: semantic misinterpretation, visual feature extraction insufficiency, and mapping deficiency. Results identify mapping deficiency as the primary hurdle. To address this problem, we show that in-context learning (ICL) can significantly enhance LMMs’ generalization. We further explore the robustness of ICL under distribution shifts and show its vulnerability to domain shifts, label shifts, and spurious correlation shifts between in-context examples and test data, opening new avenues for overcoming generalization barriers.
Xingxuan Zhang, Jiansheng Li, Wenjing Chu, Junjia Hai, Renzhe Xu, Shikai Guan, Jiazheng Xu, Liping Jing, Peng Cui 0001
CVPR8
2025 VPO: Aligning Text-to-Video Generation Models with Prompt Optimization
abstract
Video generation models have achieved remarkable progress in text-to-video tasks. These models are typically trained on text-video pairs with highly detailed and carefully crafted descriptions, while real-world user inputs during inference are often concise, vague, or poorly structured. This gap makes prompt optimization crucial for generating high-quality videos. Current methods often rely on large language models (LLMs) to refine prompts through in-context learning, but suffer from several limitations: they may distort user intent, omit critical details, or introduce safety risks. Moreover, they optimize prompts without considering the impact on the final video quality, which can lead to suboptimal results. To address these issues, we introduce VPO, a principled framework that optimizes prompts based on three core principles: harmlessness, accuracy, and helpfulness. The generated prompts faithfully preserve user intents and, more importantly, enhance the safety and quality of generated videos. To achieve this, VPO employs a two-stage optimization approach. First, we construct and refine a supervised fine-tuning (SFT) dataset based on principles of safety and alignment. Second, we introduce both text-level and video-level feedback to further optimize the SFT model with preference learning. Our extensive experiments demonstrate that VPO significantly improves safety, alignment, and video quality compared to baseline methods. Moreover, VPO shows strong generalization across video generation models. Furthermore, we demonstrate that VPO could outperform and be combined with RLHF methods on video generation models, underscoring the effectiveness of VPO in aligning video generation models. Our code and data are publicly available at https://github.com/thu-coai/VPO.
Ruiliang Lyu, Xiaotao Gu, Xiao Liu 0036, Jiazheng Xu, Yida Lu, Jiayan Teng, Zhuoyi Yang, Yuxiao Dong, Jie Tang 0001, Hongning Wang, Minlie Huang
ICCV5
2025 CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer
abstract
We present CogVideoX, a large-scale text-to-video generation model based on diffusion transformer, which can generate 10-second continuous videos that align seamlessly with text prompts, with a frame rate of 16 fps and resolution of 768 x 1360 pixels. Previous video generation models often struggled with limited motion and short durations. It is especially difficult to generate videos with coherent narratives based on text. We propose several designs to address these issues. First, we introduce a 3D Variational Autoencoder (VAE) to compress videos across spatial and temporal dimensions, enhancing both the compression rate and video fidelity. Second, to improve text-video alignment, we propose an expert transformer with expert adaptive LayerNorm to facilitate the deep fusion between the two modalities. Third, by employing progressive training and multi-resolution frame packing, CogVideoX excels at generating coherent, long-duration videos with diverse shapes and dynamic movements. In addition, we develop an effective pipeline that includes various pre-processing strategies for text and video data. Our innovative video captioning model significantly improves generation quality and semantic alignment. Results show that CogVideoX achieves state-of-the-art performance in both automated benchmarks and human evaluation. We publish the code and model checkpoints of CogVideoX along with our VAE model and video captioning model at https://github.com/THUDM/CogVideo.
Zhuoyi Yang, Jiayan Teng, Wendi Zheng, Ming Ding 0004, Shiyu Huang 0001, Jiazheng Xu, Yuanming Yang, Wenyi Hong, Guanyu Feng, Da Yin, Yean Cheng, Bin Xu 0001, Xiaotao Gu, Yuxiao Dong, Jie Tang 0001
ICLR6
2025 DiCLIP: Integrating DINOv2 and CLIP for Zero-Shot and Few-Shot Anomaly Detection with Versatile Combination Prompts
abstract
Industrial Anomaly Detection (AD) encounters a significant cold-start challenge due to the requirement of a large number of labeled normal samples, which are often difficult to obtain in new production lines. Although Zero-Shot Anomaly Detection (ZSAD) and Few-Shot Anomaly Detection (FSAD) have been proposed as potential solutions, existing methods suffer from limitations in generalization ability and are prone to contamination by anomalies in few-shot scenarios. To address these issues, we propose DiCLIP, a unified framework that integrates ZSAD and FSAD through three key innovations. First, Versatile Combination Prompts Learning combines static, dynamic, and anomaly-sensitive prompts to leverage textual anomaly cues together with image features for accurate anomaly localization in images. Second, the Anomaly-Aware Memory Bank utilizes ZSAD priors to filter contaminated features, enabling anomaly detection based on a small number of anomaly samples. Third, Adaptive Threshold Optimization integrates semantic alignment from ZSAD with feature matching from FSAD to release the constraint of a uniform threshold for test images, thereby achieving higher-precision segmentation and localization performance. Extensive experiments on the standard MVTec and VisA benchmark datasets demonstrate the superior performance of DiCLIP, highlighting its effectiveness and practical value for industrial deployment.
Xinxu Cai, Lihang Sun, Zhenshen Qu, Jiazheng Xu
IJCNN4
2025 Real-time inner wall surface defect detection based on multi-morphological feature fusion network
abstract
In industrial manufacturing, defects on the inner wall surface are crucial for quality and safety assessment. However, existing detection methods are limited by low resolution and glare interference. This study presents a Multi-morphological Feature Fusion Network for Object Detection (MFFN-OD) for 360°detection of inner wall image defects. First, it cleverly integrates panoramic imaging with conventional features through a dual-branch backbone and annular features, ensuring rotation invariance and holistic feature preservation. Second, we develop an Adaptive Multi-morphological Feature Alignment Module (AMFAM) that combats centrally polarized defects by automatically adjusting feature alignment, reducing noise, and increasing accuracy, as well as a feature interaction module with a focus on strengthening multiscale feature fusion. Third, we introduce an Asymptotic Feature Pyramid Network with Auxiliary Features (AFPN-AF) to further refine fusion, close semantic gaps, and improve performance. Experimental results show that MFFN-OD achieves 96.1% mean Average Precision (mAP) and 94.3% Average Precision (AP) for demanding faults with fast detection of 17 milliseconds per frame, meeting industrial requirements for accuracy and real-time performance.
Zhenshen Qu, Xinxu Cai, Jiazheng Xu, Chuan Lin 0003
Eng. Appl. Artif. Intell.4
2024 CogAgent: A Visual Language Model for GUI Agents
abstract
People are spending an enormous amount of time on dig-ital devices through graphical user interfaces (GUIs), e.g., computer or smartphone screens. Large language models (LLMs) such as ChatGPT can assist people in tasks like writing emails, but struggle to understand and interact with GUIs, thus limiting their potential to increase automation levels. In this paper, we introduce CogAgent, an 18-billion-parameter visual language model (VLM) specializing in GUI understanding and navigation. By utilizing both low-resolution and high-resolution image encoders, CogA-gent supports input at a resolution of1120 × 1120, enabling it to recognize tiny page elements and text. As a general-ist visual language model, CogAgent achieves the state of the art on five text-rich and four general VQA benchmarks, including VQAv2, OK- VQA, Text- Vqa, St- Vqa, ChartQA, infoVQA, DocVQA, MM-Vet, and POPE. CogAgent, using only screenshots as input, outperforms LLM-based methods that consume extracted HTML text on both PC and Android GUI navigation tasks-Mind2Web and AITW, ad-vancing the state of the art. The model and codes are available at https://github.com/THUDM/CogVLM.
Wenyi Hong, Qingsong Lv, Jiazheng Xu, Wenmeng Yu, Junhui Ji, Yan Wang 0120, Yuxiao Dong, Ming Ding 0004, Jie Tang 0001
CVPR4
2024 CogVLM: Visual Expert for Pretrained Language Models
abstract
We introduce CogVLM, a powerful open-source visual language foundation model. Different from the popular \emph{shallow alignment} method which maps image features into the input space of language model, CogVLM bridges the gap between the frozen pretrained language model and image encoder by a trainable visual expert module in the attention and FFN layers. As a result, CogVLM enables a deep fusion of vision language features without sacrificing any performance on NLP tasks. CogVLM-17B achieves state-of-the-art performance on 17 classic cross-modal benchmarks, including 1) image captioning datasets: NoCaps, Flicker30k, 2) VQA datasets: OKVQA, TextVQA, OCRVQA, ScienceQA, 3) LVLM benchmarks: MM-Vet, MMBench, SEED-Bench, LLaVABench, POPE, MMMU, MathVista, 4) visual grounding datasets: RefCOCO, RefCOCO+, RefCOCOg, Visual7W. Codes and checkpoints are available at Github.
Qingsong Lv, Wenmeng Yu, Wenyi Hong, Ji Qi 0003, Yan Wang 0120, Junhui Ji, Zhuoyi Yang, Xixuan Song, Jiazheng Xu, Keqin Chen, Bin Xu 0001, Juan-Zi Li, Yuxiao Dong, Ming Ding 0004, Jie Tang 0001
NeurIPS11
2023 ImageReward: Learning and Evaluating Human Preferences for Text-to-Image Generation
abstract
We present a comprehensive solution to learn and improve text-to-image models from human preference feedback. To begin with, we build ImageReward---the first general-purpose text-to-image human preference reward model---to effectively encode human preferences. Its training is based on our systematic annotation pipeline including rating and ranking, which collects 137k expert comparisons to date. In human evaluation, ImageReward outperforms existing scoring models and metrics, making it a promising automatic metric for evaluating text-to-image synthesis. On top of it, we propose Reward Feedback Learning (ReFL), a direct tuning algorithm to optimize diffusion models against a scorer. Both automatic and human evaluation support ReFL's advantages over compared methods. All code and datasets are provided at \url{https://github.com/THUDM/ImageReward}.
Jiazheng Xu, Xiao Liu 0036, Yuxuan Tong, Qinkai Li, Ming Ding 0004, Jie Tang 0001, Yuxiao Dong
NeurIPS1
2023 User Behavior Simulation for Search Result Re-ranking
abstract
Result ranking is one of the major concerns for Web search technologies. Most existing methodologies rank search results in descending order of relevance. To model the interactions among search results, reinforcement learning (RL algorithms have been widely adopted for ranking tasks. However, the online training of RL methods is time and resource consuming at scale. As an alternative, learning ranking policies in the simulation environment is much more feasible and efficient. In this article, we propose two different simulation environments for the offline training of the RL ranking agent: the Context-aware Click Simulator (CCS) and the Fine-grained User Behavior Simulator with GAN (UserGAN). Based on the simulation environment, we also design a User Behavior Simulation for Reinforcement Learning (UBS4RL) re-ranking framework, which consists of three modules: a feature extractor for heterogeneous search results, a user simulator for collecting simulated user feedback, and a ranking agent for generation of optimized result lists. Extensive experiments on both simulated and practical Web search datasets show that (1) the proposed user simulators can capture and simulate fine-grained user behavior patterns by training on large-scale search logs, (2) the temporal information of user searching process is a strong signal for ranking evaluation, and (3) learning ranking policies from the simulation environment can effectively improve the search ranking performance.
Yiqun Liu 0001, Jiaxin Mao, Weizhi Ma, Jiazheng Xu, Shaoping Ma, Qi Tian 0001
ACM Trans. Inf. Syst.5
2022 Regulatory Instruments for Fair Personalized Pricing
abstract
Personalized pricing is a business strategy to charge different prices to individual consumers based on their characteristics and behaviors. It has become common practice in many industries nowadays due to the availability of a growing amount of high granular consumer data. The discriminatory nature of personalized pricing has triggered heated debates among policymakers and academics on how to design regulation policies to balance market efficiency and equity. In this paper, we propose two sound policy instruments, i.e., capping the range of the personalized prices or their ratios. We investigate the optimal pricing strategy of a profit-maximizing monopoly under both regulatory constraints and the impact of imposing them on consumer surplus, producer surplus, and social welfare. We theoretically prove that both proposed constraints can help balance consumer surplus and producer surplus at the expense of total surplus for common demand distributions, such as uniform, logistic, and exponential distributions. Experiments on both simulation and real-world datasets demonstrate the correctness of these theoretical results1. Our findings and insights shed light on regulatory policy design for the increasingly monopolized business in the digital era.
Renzhe Xu, Xingxuan Zhang, Peng Cui 0001, Bo Li 0064, Zheyan Shen, Jiazheng Xu
WWW6