Cheng Guan

dblp:61/3722 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
2since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 2 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
3D vision · 35% Video understanding and tracking · 18% Generative modeling · 18%
Computer graphics and multimedia
1 paper
Image and video processing · 100%

Topics — the 10 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
diffusion model
0.912025
Jasmine: Harnessing Diffusion Prior for Self-supervised Depth Estimation · NeurIPS 2025
Computer vision › 3D vision › depth estimation
monocular depth estimation
0.912025
Jasmine: Harnessing Diffusion Prior for Self-supervised Depth Estimation · NeurIPS 2025
Computer vision › 3D vision › depth estimation
self-supervised depth estimation
0.912025
Jasmine: Harnessing Diffusion Prior for Self-supervised Depth Estimation · NeurIPS 2025
Computer vision › Video understanding and tracking
action recognition
0.412020
Multi-Group Multi-Attention: Towards Discriminative Spatiotemporal Representation · ACM Multimedia 2020
Machine learning › Deep learning architectures and training
attention mechanism
0.412020
Multi-Group Multi-Attention: Towards Discriminative Spatiotemporal Representation · ACM Multimedia 2020
Machine learning › Deep learning architectures and training › attention mechanism › multi-dimensional attention
spatio-temporal attention
0.412020
Multi-Group Multi-Attention: Towards Discriminative Spatiotemporal Representation · ACM Multimedia 2020
Machine learning › Representation and self-supervised learning › representation learning
spatio-temporal representation learning
0.412020
Multi-Group Multi-Attention: Towards Discriminative Spatiotemporal Representation · ACM Multimedia 2020
Computer vision › Video understanding and tracking › dynamic scene analysis › video scene understanding
video visual relation detection
0.412020
CoTeRe-Net: Discovering Collaborative Ternary Relations in Videos · ECCV (6) 2020
Image and video processing
image reconstruction
0.312025
Jasmine: Harnessing Diffusion Prior for Self-supervised Depth Estimation · NeurIPS 2025
Knowledge, reasoning and agents › Knowledge representation and reasoning
relational reasoning
0.112020
CoTeRe-Net: Discovering Collaborative Ternary Relations in Videos · ECCV (6) 2020

Methods — techniques the papers use, named apart from their topics

stable diffusion · 1.7scale-shift GRU · 1.7mix-batch image reconstruction · 1.7multi-group attention · 0.43d convolutional neural network · 0.4
YearPublicationVenuePosition
2025 Jasmine: Harnessing Diffusion Prior for Self-supervised Depth Estimation
abstract
In this paper, we propose \textbf{Jasmine}, the first Stable Diffusion (SD)-based self-supervised framework for monocular depth estimation, which effectively harnesses SD’s visual priors to enhance the sharpness and generalization of unsupervised prediction. Previous SD-based methods are all supervised since adapting diffusion models for dense prediction requires high-precision supervision. In contrast, self-supervised reprojection suffers from inherent challenges (\textit{e.g.}, occlusions, texture-less regions, illumination variance), and the predictions exhibit blurs and artifacts that severely compromise SD's latent priors. To resolve this, we construct a novel surrogate task of mix-batch image reconstruction. Without any additional supervision, it preserves the detail priors of SD models by reconstructing the images themselves while preventing depth estimation from degradation. Furthermore, to address the inherent misalignment between SD's scale and shift invariant estimation and self-supervised scale-invariant depth estimation, we build the Scale-Shift GRU. It not only bridges this distribution gap but also isolates the fine-grained texture of SD output against the interference of reprojection loss. Extensive experiments demonstrate that Jasmine achieves SoTA performance on the KITTI benchmark and exhibits superior zero-shot generalization across multiple datasets.
Jiyuan Wang 0001, Chunyu Lin, Cheng Guan, Lang Nie, Kang Liao, Yao Zhao 0001
NeurIPS3
2024 Assessing illumination fatigue in tunnel workers through eye-tracking technology: A laboratory study
Jingzheng Zhu, Cheng Guan
Adv. Eng. Informatics3
2020 CoTeRe-Net: Discovering Collaborative Ternary Relations in Videos
Zhensheng Shi, Cheng Guan, Liangjie Cao, Ju Liang, Zhaorui Gu, Haiyong Zheng
ECCV (6)2
2020 Multi-Group Multi-Attention: Towards Discriminative Spatiotemporal Representation
abstract
Learning spatiotemporal features is very effective but challenging for video understanding especially action recognition. In this paper, we propose Multi-Group Multi-Attention, dubbed MGMA, paying more attention to "where and when" the action happens, for learning discriminative spatiotemporal representation in videos. The contribution of MGMA is three-fold: First, by devising a new spatiotemporal separable attention mechanism, it can learn temporal attention and spatial attention separately for fine-grained spatiotemporal representation. Second, through designing a novel multi-group structure, it can capture multi-attention rendered spatiotemporal features better. Finally, our MGMA module is lightweight and flexible yet effective, so that can be easily embedded into any 3D Convolutional Neural Network (3D-CNN) architecture. We embed multiple MGMA modules into 3D-CNN to train an end-to-end, RGB-only model and evaluate on four popular benchmarks: UCF101 and HMDB51, Something-Something V1 and V2. Ablation study and experimental comparison demonstrate the strength of our MGMA, which achieves superior performance compared to state-of-the-arts. Our code is available at https://github.com/zhenglab/mgma.
Zhensheng Shi, Liangjie Cao, Cheng Guan, Ju Liang, Zhaorui Gu, Haiyong Zheng
ACM Multimedia3
2020 Lip Image Segmentation Based on a Fuzzy Convolutional Neural Network
abstract
Research has shown that the human lip and its movements are a rich source of information related to speech content and speaker's identity. Lip image segmentation, as a fundamental step in many lip-reading and visual speaker authentication systems, is of vital importance. Because of variations in lip color, lighting conditions and especially the complex appearance of an open mouth, accurate lip region segmentation is still a challenging task. To address this problem, this article proposes a new fuzzy deep neural network having an architecture that integrates fuzzy units and traditional convolutional units. The convolutional units are used to extract discriminative features at different scales to provide comprehensive information for pixel-level lip segmentation. The fuzzy logic modules are employed to handle various kinds of uncertainties and to provide a more robust segmentation result. An end-to-end training scheme is then used to learn the optimal parameters for both the fuzzy and the convolutional units. A dataset containing more than 48 000 images of various speakers, under different lighting conditions, was used to evaluate lip segmentation performance. According to the experimental results, the proposed method achieves state-of-the-art performance when compared with other algorithms.
Cheng Guan, Shi-Lin Wang, Alan Wee-Chung Liew
IEEE Trans. Fuzzy Syst.1
2019 Lip Image Segmentation in Mobile Devices Based on Alternative Knowledge Distillation
abstract
Lip image segmentation, as the first step in many lip-related tasks (e.g. automatic lipreading), is of vital significance for the subsequent procedures. Nowadays, with the increasing computational power of the mobile devices, mobile applications become more and more popular. In this paper, a new approach is proposed, which is able to segment the lip region in natural scenes and is of acceptable computational complexity to be implemented in mobile devices. Two networks including a complex teacher network and a compact student network with the same structure are employed. With the proposed remedy loss and the alternative knowledge distillation scheme, the student network can learn useful knowledge from the teacher network effectively and efficiently, and even rectify some of its segmentation errors. A dataset containing 49 people captured under natural scenes by various cellphone cameras is adopted for evaluation and the experiment results have demonstrated that the proposed student network even outperforms the teacher network with much less computational cost.
Cheng Guan, Shi-Lin Wang, Gongshen Liu, Alan Wee-Chung Liew
ICIP1
2004 Adaptive time-varying sliding mode control for hydraulic servo system
abstract
This paper studies the position control of an electro-hydraulic servo system. Because the dynamics of the system are highly nonlinear and have large extent of model uncertainties including big changes in load and hydraulic parameters, which are unmatched, a time-varying sliding mode control approach combined with adaptive control is proposed based on Lyapunov analysis. The time-varying sliding mode control avoids the reaching phase of conventional sliding mode control, thus the control method proposed can be robust all the time. Adaptive control is used to identify the system parameters to overcome the influence of the uncertain parameters and disturbances. Simulation results indicate that the control approach has nice global robustness and improves position tracking accuracy considerably.
Cheng Guan, Shanan Zhu
ICARCV1