Zhenhuan Liu

dblp:262/3005 · DBLP profile ↗
← Back
8ranked-venue papers
5as first author
6since 2021 · last 2024
0000-0001-9932-9225ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Generative modeling · 95% Video understanding and tracking · 5%
Computer graphics and multimedia
2 papers
Visual content generation and editing · 100%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Bioinformatics and computational biology · 100%

Topics — the 14 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
generative domain adaptation
0.712023
Text-Driven Generative Domain Adaptation with Spectral Consistency Regularization · ICCV 2023
Machine learning › Generative modeling › generative adversarial network › GAN training
mode collapse
0.712023
Text-Driven Generative Domain Adaptation with Spectral Consistency Regularization · ICCV 2023
Machine learning › Generative modeling › diffusion model › image editing
instruction-based image editing
0.612022
LS-GAN: Iterative Language-based Image Manipulation via Long and Short Term Consistency Reasoning · ACM Multimedia 2022
Bioinformatics and computational biology › biological network › network biology › network inference
gene regulatory network inference
0.612022
Accurate inference of gene regulatory interactions from spatial gene expression with deep contrastive learning · Bioinform. 2022
Bioinformatics and computational biology › transcriptomics
spatial gene expression
0.612022
Accurate inference of gene regulatory interactions from spatial gene expression with deep contrastive learning · Bioinform. 2022
Visual content generation and editing › image editing › interactive image editing
iterative image editing
0.612022
LS-GAN: Iterative Language-based Image Manipulation via Long and Short Term Consistency Reasoning · ACM Multimedia 2022
Visual content generation and editing
style transfer
0.612022
Unsupervised Coherent Video Cartoonization with Perceptual Motion Consistency · AAAI 2022
Machine learning › Generative modeling › image generation
conditional image generation
0.412020
IR-GAN: Image Manipulation with Linguistic Instruction by Increment Reasoning · ACM Multimedia 2020
Machine learning › Generative modeling
multimodal generation
0.412020
IR-GAN: Image Manipulation with Linguistic Instruction by Increment Reasoning · ACM Multimedia 2020
Machine learning › Generative modeling › diffusion model
text-to-image generation
0.212023
Text-Driven Generative Domain Adaptation with Spectral Consistency Regularization · ICCV 2023
Machine learning › Generative modeling
generative adversarial network
0.212022
LS-GAN: Iterative Language-based Image Manipulation via Long and Short Term Consistency Reasoning · ACM Multimedia 2022
Computer vision › Video understanding and tracking › temporal modeling
temporal consistency
0.212022
Unsupervised Coherent Video Cartoonization with Perceptual Motion Consistency · AAAI 2022
Bioinformatics and computational biology › gene regulation
gene regulatory network
0.212022
Accurate inference of gene regulatory interactions from spatial gene expression with deep contrastive learning · Bioinform. 2022
Bioinformatics and computational biology
systems biology
0.212022
Accurate inference of gene regulatory interactions from spatial gene expression with deep contrastive learning · Bioinform. 2022

Methods — techniques the papers use, named apart from their topics

generative adversarial network · 1.6unsupervised learning · 1.1spatially-adaptive semantic alignment · 1.1phrase encoding · 1.1perceptual motion consistency · 1.1consistency reasoning · 1.1spectral consistency regularization · 0.7granularity adaptive regularization · 0.7siamese convolutional neural network · 0.6deep learning · 0.6contrastive learning · 0.6reasoning discriminator · 0.4increment reasoning · 0.4
YearPublicationVenuePosition
2024 T-Code: Simple Temporal Latent Code for Efficient Dynamic View Synthesis
Zhenhuan Liu
ICONIP (8)1
2023 FASTC: A Fast Attentional Framework for Semantic Traversability Classification Using Point Cloud
abstract
Producing traversability maps and understanding the surroundings are crucial prerequisites for autonomous navigation. In this paper, we address the problem of traversability assessment using point clouds. We propose a novel pillar feature extraction module that utilizes PointNet to capture features from point clouds organized in vertical volume and a 2D encoder-decoder structure to conduct traversability classification instead of the widely used 3D convolutions. This results in less computational cost while even better performance is achieved at the same time. We then propose a new spatio-temporal attention module to fuse multi-frame information, which can properly handle the varying density problem of LIDAR point clouds, and this makes our module able to assess distant areas more accurately. Comprehensive experimental results on augmented Semantic KITTI and RELLIS-3D datasets show that our method is able to achieve superior performance over existing approaches both quantitatively and quantitatively. Our code is publicly available at https://github.com/chenyirui/FASTC.
Yirui Chen, Pengjin Wei, Zhenhuan Liu, Bingchao Wang, Jie Yang 0002, Wei Liu 0044
ECAI3
2023 Text-Driven Generative Domain Adaptation with Spectral Consistency Regularization
abstract
Combined with the generative prior of pre-trained models and the flexibility of text, text-driven generative domain adaptation can generate images from a wide range of target domains. However, current methods still suffer from overfitting and the mode collapse problem. In this paper, we analyze the mode collapse from the geometric point of view and reveal its relationship to the Hessian matrix of generator. To alleviate it, we propose the spectral consistency regularization to preserve the diversity of source domain without restricting the semantic adaptation to target domain. We also design granularity adaptive regularization to flexibly control the balance between diversity and stylization for target model. We conduct experiments for broad target domains compared with state-of-the-art methods and extensive ablation studies. The experiments demonstrate the effectiveness of our method to preserve the diversity of source domain and generate high fidelity target images. Source code has been released in https://github.com/Victarry/Adaptation-SCR.
Zhenhuan Liu, Liang Li 0003, Jiayu Xiao, Zhengjun Zha, Qingming Huang
ICCV1
2022 Unsupervised Coherent Video Cartoonization with Perceptual Motion Consistency
abstract
In recent years, creative content generations like style transfer and neural photo editing have attracted more and more attention. Among these, cartoonization of real-world scenes has promising applications in entertainment and industry. Different from image translations focusing on improving the style effect of generated images, video cartoonization has additional requirements on the temporal consistency. In this paper, we propose a spatially-adaptive semantic alignment framework with perceptual motion consistency for coherent video cartoonization in an unsupervised manner. The semantic alignment module is designed to restore deformation of semantic structure caused by spatial information lost in the encoder-decoder architecture. Furthermore, we introduce the spatio-temporal correlative map as a style-independent, global-aware regularization on perceptual motion consistency. Deriving from similarity measurement of high-level features in photo and cartoon frames, it captures global semantic information beyond raw pixel-value of optical flow. Besides, the similarity measurement disentangles temporal relationship from domain-specific style properties, which helps regularize the temporal consistency without hurting style effects of cartoon images. Qualitative and quantitative experiments demonstrate our method is able to generate highly stylistic and temporal consistent cartoon videos.
Zhenhuan Liu, Liang Li 0003, Huajie Jiang, Xin Jin 0004, Dandan Tu, Shuhui Wang, Zhengjun Zha
AAAI1
2022 LS-GAN: Iterative Language-based Image Manipulation via Long and Short Term Consistency Reasoning
abstract
Iterative language-based image manipulation aims to edit images step by step according to user's linguistic instructions. The existing methods mostly focus on aligning the attributes and appearance of new-added visual elements with current instruction. However, they fail to maintain consistency between instructions and images as iterative rounds increase. To address this issue, we propose a novel Long and Short term consistency reasoning Generative Adversarial Network (LS-GAN), which enhances the awareness of previous objects with current instruction and better maintains the consistency with the user's intent under the continuous iterations. Specifically, we first design a Context-aware Phrase Encoder (CPE) to learn the user's intention by extracting different phrase-level information about the instruction. Further, we introduce a Long and Short term Consistency Reasoning (LSCR) mechanism. The long-term reasoning improves the model on semantic understanding and positional reasoning, while short-term reasoning ensures the ability to construct visual scenes based on linguistic instructions. Extensive results show that LS-GAN improves the generation quality in terms of both object identity and position, and achieves the state-of-the-art performance on two public datasets.
Gaoxiang Cong 0001, Liang Li 0003, Zhenhuan Liu, Yunbin Tu, Weijun Qin, Shenyuan Zhang, Chengang Yan, Bin Jiang 0011
ACM Multimedia3
2022 Accurate inference of gene regulatory interactions from spatial gene expression with deep contrastive learning
abstract
MOTIVATION: Reverse engineering of gene regulatory networks (GRNs) has long been an attractive research topic in system biology. Computational prediction of gene regulatory interactions has remained a challenging problem due to the complexity of gene expression and scarce information resources. The high-throughput spatial gene expression data, like in situ hybridization images that exhibit temporal and spatial expression patterns, has provided abundant and reliable information for the inference of GRNs. However, computational tools for analyzing the spatial gene expression data are highly underdeveloped. RESULTS: In this study, we develop a new method for identifying gene regulatory interactions from gene expression images, called ConGRI. The method is featured by a contrastive learning scheme and deep Siamese convolutional neural network architecture, which automatically learns high-level feature embeddings for the expression images and then feeds the embeddings to an artificial neural network to determine whether or not the interaction exists. We apply the method to a Drosophila embryogenesis dataset and identify GRNs of eye development and mesoderm development. Experimental results show that ConGRI outperforms previous traditional and deep learning methods by a large margin, which achieves accuracies of 76.7% and 68.7% for the GRNs of early eye development and mesoderm development, respectively. It also reveals some master regulators for Drosophila eye development. AVAILABILITYAND IMPLEMENTATION: https://github.com/lugimzheng/ConGRI. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Lujing Zheng, Zhenhuan Liu, Yang Yang 0030, Hong-Bin Shen
Bioinform.2
2020 Tripfly: Predicting Gene-gene Interaction of Drosophila Eye Development Using Triplet Loss
abstract
The reconstruction of gene regulatory network (GRN) is of significance in system biology. In recent years, benefiting from the advances of deep learning technologies, image-based gene expression data, which contains spatial expression patterns, has become a new resource in network inference. Most of the existing image-based GRN inference models are based on unsupervised models, due to the lack of labeled data. And a few methods employ supervised learning models, whose performance is limited by the scale of training data.In this study, in order to predict the gene regulatory network of the eye development of Drosophila embryos, we develop a weakly supervised learning method. We generate image triplets of genes according to their orientation and developing stage. Then we build a deep convolutional neural network, using triplet loss to train a siamese network and extract the relationship between genes. The new method achieves promising results in the prediction of gene regulatory relationship in the eye development of Drosophila with a total accuracy of over 72%.
Zhenhuan Liu, Jiafeng Chen, Yang Yang 0030
BIBM1
2020 IR-GAN: Image Manipulation with Linguistic Instruction by Increment Reasoning
abstract
Conditional image generation is an active research topic including text2image and image translation. Recently image manipulation with linguistic instruction brings new challenges of multimodal conditional generation. However, traditional conditional image generation models mainly focus on generating high-quality and visually realistic images, and lack resolving the partial consistency between image and instruction. To address this issue, we propose an Increment Reasoning Generative Adversarial Network (IR-GAN), which aims to reason the consistency between visual increment in images and semantic increment in instructions. First, we introduce the word-level and instruction-level instruction encoders to learn user's intention from history-correlated instructions as semantic increment. Second, we embed the representation of semantic increment into that of source image for generating target image, where source image plays the role of referring auxiliary. Finally, we propose a reasoning discriminator to measure the consistency between visual increment and semantic increment, which purifies user's intention and guarantees the good logic of generated target image. Extensive experiments and visualization conducted on two datasets show the effectiveness of IR-GAN.
Zhenhuan Liu, Jincan Deng, Liang Li 0003, Shaofei Cai, Qianqian Xu 0001, Shuhui Wang, Qingming Huang
ACM Multimedia1