VLDB 2026 Research / reviewers in the wild / expert
Yean Cheng
dblp:261/9981
· DBLP profile ↗
8ranked-venue papers
1as first author
7since 2021 · last 2025
0000-0002-2846-0450ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Generative modeling · 42% Vision and language · 24% 3D vision · 21% | |
| Computer graphics and multimedia
3 papers |
Image and video processing · 47% Computational photography and imaging · 41% Visual content generation and editing · 12% |
Topics — the 14 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling
diffusion model |
0.9 | 1 | 2025 | CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer · ICLR 2025 |
Machine learning › Generative modeling › video generation
long video generation |
0.9 | 1 | 2025 | CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer · ICLR 2025 |
Computer vision › Video understanding and tracking
long video understanding |
0.9 | 1 | 2025 | LVBench: An Extreme Long Video Understanding Benchmark · ICCV 2025 |
Computer vision › Vision and language › vision-language model
multimodal large language model |
0.9 | 1 | 2025 | AlignMMBench: Evaluating Chinese Multimodal Alignment in Large Vision-Language Models · ACL (1) 2025 |
Machine learning › Generative modeling › video generation
text-to-video generation |
0.9 | 1 | 2025 | CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer · ICLR 2025 |
Machine learning › Generative modeling
video generation |
0.9 | 1 | 2025 | CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer · ICLR 2025 |
Computer vision › Vision and language
video-language model |
0.9 | 1 | 2025 | MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models · CVPR 2025 |
Computer vision › 3D vision › inverse rendering
illumination estimation |
0.8 | 1 | 2024 | SPLiT: Single Portrait Lighting Estimation via a Tetrad of Face Intrinsics · IEEE Trans. Pattern Anal. Mach. Intell. 2024 |
Computer vision › 3D vision
neural radiance field |
0.8 | 1 | 2024 | Colorizing Monochromatic Radiance Fields · AAAI 2024 |
Computational photography and imaging
intrinsic image decomposition |
0.8 | 1 | 2024 | SPLiT: Single Portrait Lighting Estimation via a Tetrad of Face Intrinsics · IEEE Trans. Pattern Anal. Mach. Intell. 2024 |
Image and video processing › super-resolution
image super-resolution |
0.4 | 1 | 2020 | Structure-Preserving Super Resolution With Gradient Guidance · CVPR 2020 |
Image and video processing › super-resolution › image super-resolution
single image super-resolution |
0.4 | 1 | 2020 | Structure-Preserving Super Resolution With Gradient Guidance · CVPR 2020 |
Computer vision › Vision and language › cross-modal alignment › visual-semantic alignment
video-text alignment |
0.3 | 1 | 2025 | CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer · ICLR 2025 |
Visual content generation and editing
image colorization |
0.2 | 1 | 2024 | Colorizing Monochromatic Radiance Fields · AAAI 2024 |
Methods — techniques the papers use, named apart from their topics
image colorization module · 1.5cascaded network · 1.5vision-language model · 0.9progressive training · 0.9multimodal large language model · 0.9multi-resolution frame packing · 0.9frame rate scaling · 0.9expert transformer · 0.9benchmark evaluation · 0.93d variational autoencoder · 0.9two-stage lighting estimator · 0.8gradient loss · 0.4generative adversarial network · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | AlignMMBench: Evaluating Chinese Multimodal Alignment in Large Vision-Language ModelsabstractYuhang Wu, Wenmeng Yu, Yean Cheng, Yan Wang, Xiaohan Zhang, Jiazheng Xu, Ming Ding, Yuxiao Dong. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Wenmeng Yu, Yean Cheng, Yan Wang 0120, Jiazheng Xu, Ming Ding 0004, Yuxiao Dong |
ACL (1) | 3 |
| 2025 | MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language ModelsabstractIn recent years, vision language models (VLMs) have made significant advancements in video understanding. However, a crucial capability — fine-grained motion comprehension — remains under-explored in current benchmarks. To address this gap, we propose MotionBench, a comprehensive evaluation benchmark designed to assess the fine-grained motion comprehension of video understanding models. MotionBench evaluates models’ motion-level perception through six primary categories of motion-oriented question types and includes data collected from diverse sources, ensuring a broad representation of real-world video content. Experimental results reveal that existing VLMs perform poorly in understanding fine-grained motions. To enhance VLM’s ability to perceive fine-grained motion within a limited sequence length of LLM, we conduct extensive experiments reviewing VLM architectures optimized for video feature compression and propose a novel and efficient Through-Encoder (TE) Fusion method. Experiments show that higher frame rate inputs and TE Fusion yield improvements in motion understanding, yet there is still substantial room for enhancement. Our benchmark aims to guide and motivate the development of more capable video understanding models, emphasizing the importance of fine-grained motion comprehension. Project page: https://motion-bench.github.io. Wenyi Hong, Yean Cheng, Zhuoyi Yang, Lefan Wang, Xiaotao Gu, Shiyu Huang 0001, Yuxiao Dong, Jie Tang 0001 |
CVPR | 2 |
| 2025 | LVBench: An Extreme Long Video Understanding BenchmarkabstractRecent progress in multimodal large language models has markedly enhanced the understanding of short videos (typically under one minute), and several evaluation datasets have emerged accordingly. However, these advancements fall short of meeting the demands of real-world applications such as embodied intelligence for long-term decision-making, in-depth movie reviews and discussions, and live sports commentary, all of which require comprehension of long videos spanning several hours. To address this gap, we introduce LVBench, a benchmark specifically designed for long video understanding. Our dataset comprises publicly sourced videos and encompasses a diverse set of tasks aimed at long video comprehension and information extraction. LVBench is designed to challenge multimodal models to demonstrate long-term memory and extended comprehension capabilities. Our extensive evaluations reveal that current multimodal models still underperform on these demanding long video understanding tasks. Through LVBench, we aim to spur the development of more advanced models capable of tackling the complexities of long video comprehension. Our data and code are publicly available at: https://lvbench.github.io. Zehai He, Wenyi Hong, Yean Cheng, Ji Qi 0003, Ming Ding 0004, Xiaotao Gu, Shiyu Huang 0001, Bin Xu 0001, Yuxiao Dong, Jie Tang 0001 |
ICCV | 4 |
| 2025 | CogVideoX: Text-to-Video Diffusion Models with An Expert TransformerabstractWe present CogVideoX, a large-scale text-to-video generation model based on diffusion transformer, which can generate 10-second continuous videos that align seamlessly with text prompts, with a frame rate of 16 fps and resolution of 768 x 1360 pixels.
Previous video generation models often struggled with limited motion and short durations.
It is especially difficult to generate videos with coherent narratives based on text.
We propose several designs to address these issues.
First, we introduce a 3D Variational Autoencoder (VAE) to compress videos across spatial and temporal dimensions, enhancing both the compression rate and video fidelity.
Second, to improve text-video alignment, we propose an expert transformer with expert adaptive LayerNorm to facilitate the deep fusion between the two modalities.
Third, by employing progressive training and multi-resolution frame packing, CogVideoX excels at generating coherent, long-duration videos with diverse shapes and dynamic movements.
In addition, we develop an effective pipeline that includes various pre-processing strategies for text and video data.
Our innovative video captioning model significantly improves generation quality and semantic alignment.
Results show that CogVideoX achieves state-of-the-art performance in both automated benchmarks and human evaluation.
We publish the code and model checkpoints of CogVideoX along with our VAE model and video captioning model at https://github.com/THUDM/CogVideo. Zhuoyi Yang, Jiayan Teng, Wendi Zheng, Ming Ding 0004, Shiyu Huang 0001, Jiazheng Xu, Yuanming Yang, Wenyi Hong, Guanyu Feng, Da Yin, Yean Cheng, Bin Xu 0001, Xiaotao Gu, Yuxiao Dong, Jie Tang 0001 |
ICLR | 14 |
| 2024 | Colorizing Monochromatic Radiance FieldsabstractThough Neural Radiance Fields (NeRF) can produce colorful 3D representations of the world by using a set of 2D images, such ability becomes non-existent when only monochromatic images are provided. Since color is necessary in representing the world, reproducing color from monochromatic radiance fields becomes crucial. To achieve this goal, instead of manipulating the monochromatic radiance fields directly, we consider it as a representation-prediction task in the Lab color space. By first constructing the luminance and density representation using monochromatic images, our prediction stage can recreate color representation on the basis of an image colorization module. We then reproduce a colorful implicit model through the representation of luminance, density, and color. Extensive experiments have been conducted to validate the effectiveness of our approaches. Our project page: https://liquidammonia.github.io/color-nerf. Yean Cheng, Renjie Wan, Shuchen Weng, Chengxuan Zhu, Yakun Chang, Boxin Shi |
AAAI | 1 |
| 2024 | SPLiT: Single Portrait Lighting Estimation via a Tetrad of Face IntrinsicsabstractThis paper proposes a novel pipeline to estimate a non-parametric environment map with high dynamic range from a single human face image. Lighting-independent and -dependent intrinsic images of the face are first estimated separately in a cascaded network. The influence of face geometry on the two lighting-dependent intrinsics, diffuse shading and specular reflection, are further eliminated by distributing the intrinsics pixel-wise onto spherical representations using the surface normal as indices. This results in two representations simulating images of a diffuse sphere and a glossy sphere under the input scene lighting. Taking into account the distinctive nature of light sources and ambient terms, we further introduce a two-stage lighting estimator to predict both accurate and realistic lighting from these two representations. Our model is trained supervisedly on a large-scale and high-quality synthetic face image dataset. We demonstrate that our method allows accurate and detailed lighting estimation and intrinsic decomposition, outperforming state-of-the-art methods both qualitatively and quantitatively on real face images. Yean Cheng, Yongjie Zhu, Si Li 0001, Gang Pan 0001, Boxin Shi |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | Fault Diagnosis of Energy Networks Based on Improved Spatial-Temporal Graph Neural Network With Massive Missing DataabstractIn order to ensure the safe and reliable operation of the energy system, real-time fault diagnosis technology is indispensable. Energy systems are typically complex systems consisting of multiple subsystems that are coupled with each other. Before and after the occurrence of a fault, the system is generally in an abnormal or even harsh environment, which may cause a large number of randomly missing measurement data and make the application of fault diagnosis technology extremely difficult. In this paper, the graph attention network (GAT) is improved by a Gaussian mixture model (GMM) for incomplete-data representation. The iteratively updated expectation of the GMM serves as the characterization of missing data, which significantly improves the ability to fill in missing data. The GAT fuses multi-source data according to the topology structure so as to comprehensively exploit the spatial information. The gated recurrent units (GRU) extract dynamic fault information from embedded spatial features and classify the time series into various fault types. Moreover, we propose a loss function in the form of weighted focal loss so that the fault-class imbalance issue brought by the data deficiency can be solved. The proposed uniform spatial-temporal graph neural network classification framework together with the GMM (GM-STGNN) can effectively improve fault diagnosis performance and is applied on an experimental platform of an authentic industrial estate. Results of comparative experiments under different conditions of both sufficient and deficient data illustrate the efficiency and advancement of the proposed method.Note to Practitioners—This paper presents a fault diagnosis method for large-scale energy systems with massive missing data. The proposed GM-STGNN framework can be applied in complex energy networks consisting of coupling subsystems, such as power grids, heating networks, and gas networks. With an incomplete-data representation mechanism, the proposed method utilizes topology information to comprehensively exploit spatial features, it also recurrently transmits historical embedded features and extracts dynamic fault characteristics. Therefore, it can effectively improve energy-network fault identification accuracy when more than half of the sample exists vacant values randomly. In the training procedure, after pre-setting the model scale, data acquired by multi-source sensors is put into the model according to the real topology structure, and corresponding fault labels serve as the supervision. The statistical characteristics of missing data are learned with neural-network parameters until the loss converges. In practical application, the sampling data is divided by a time window of a few seconds. The missing data is mitigated by the estimated expectation of the GMM. Therefore, real-time fault classification results can be obtained with high accuracy. The effectiveness of the proposed method is illustrated by fault diagnosis of a typical distributed heating network under the noise influence. Benefiting from the ability to learn fault knowledge, the proposed method can be easily applied to new scenarios where the process data and topology structure of the system are known. Jingfei Zhang, Yean Cheng, Xiao He 0001 |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2020 | Structure-Preserving Super Resolution With Gradient GuidanceabstractStructures matter in single image super resolution (SISR). Recent studies benefiting from generative adversarial network (GAN) have promoted the development of SISR by recovering photo-realistic images. However, there are always undesired structural distortions in the recovered images. In this paper, we propose a structure-preserving super resolution method to alleviate the above issue while maintaining the merits of GAN-based methods to generate perceptual-pleasant details. Specifically, we exploit gradient maps of images to guide the recovery in two aspects. On the one hand, we restore high-resolution gradient maps by a gradient branch to provide additional structure priors for the SR process. On the other hand, we propose a gradient loss which imposes a second-order restriction on the super-resolved images. Along with the previous image-space loss functions, the gradient-space objectives help generative networks concentrate more on geometric structures. Moreover, our method is model-agnostic, which can be potentially used for off-the-shelf SR networks. Experimental results show that we achieve the best PI and LPIPS performance and meanwhile comparable PSNR and SSIM compared with state-of-the-art perceptual-driven SR methods. Visual results demonstrate our superiority in restoring structures while generating natural SR images. Yongming Rao, Yean Cheng, Ce Chen, Jiwen Lu, Jie Zhou 0001 |
CVPR | 3 |