Liheng Zhang

dblp:169/3159 · DBLP profile ↗
← Back
13ranked-venue papers
6as first author
3since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 first-author · 2 since 2021Artificial intelligence and machine learning · 8 · 5 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
10 papers
Representation and self-supervised learning · 26% Efficient and distributed learning · 23% 3D vision · 13%
Computer graphics and multimedia
2 papers
Visual content generation and editing · 82% Virtual and augmented reality · 18%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computational finance and economics · 100%

Topics — the 26 heaviest of 28, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision
3d reconstruction
1.012026
MGHead: Motion-Aware Animated Gaussian Head Avatars With Anchored Skeletal Structures · IEEE Trans. Multim. 2026
Computer vision › Vision and language › vision-language model
3d vision-language model
1.012026
HCC-3D: Hierarchical Compensatory Compression for 98% 3D Token Reduction in Vision-Language Models · AAAI 2026
Machine learning › Efficient and distributed learning
model compression
1.012026
HCC-3D: Hierarchical Compensatory Compression for 98% 3D Token Reduction in Vision-Language Models · AAAI 2026
Machine learning › Efficient and distributed learning › model compression
token compression
1.012026
HCC-3D: Hierarchical Compensatory Compression for 98% 3D Token Reduction in Vision-Language Models · AAAI 2026
Visual content generation and editing
3d content creation
1.012026
MGHead: Motion-Aware Animated Gaussian Head Avatars With Anchored Skeletal Structures · IEEE Trans. Multim. 2026
Visual content generation and editing › avatar generation
animatable head avatar
1.012026
MGHead: Motion-Aware Animated Gaussian Head Avatars With Anchored Skeletal Structures · IEEE Trans. Multim. 2026
Machine learning › Representation and self-supervised learning › equivariance
equivariant representation
1.022022
Learning Generalized Transformation Equivariant Representations Via AutoEncoding Transformations · IEEE Trans. Pattern Anal. Mach. Intell. 2022
AVT: Unsupervised Learning of Transformation Equivariant Representations by Autoencoding Variational Transformations · ICCV 2019
Machine learning › Representation and self-supervised learning › representation learning
unsupervised representation learning
0.932022
AVT: Unsupervised Learning of Transformation Equivariant Representations by Autoencoding Variational Transformations · ICCV 2019
AET vs. AED: Unsupervised Representation Learning by Auto-Encoding Transformations Rather Than Data · CVPR 2019
Learning Generalized Transformation Equivariant Representations Via AutoEncoding Transformations · IEEE Trans. Pattern Anal. Mach. Intell. 2022
Machine learning › Generative modeling
variational autoencoder
0.612022
Learning Generalized Transformation Equivariant Representations Via AutoEncoding Transformations · IEEE Trans. Pattern Anal. Mach. Intell. 2022
Machine learning › Trustworthy machine learning
robustness
0.412020
WCP: Worst-Case Perturbations for Semi-Supervised Deep Learning · CVPR 2020
Machine learning › Learning paradigms
semi-supervised learning
0.412020
WCP: Worst-Case Perturbations for Semi-Supervised Deep Learning · CVPR 2020
Machine learning › Deep learning architectures and training
capsule network
0.312018
CapProNet: Deep Feature Learning via Orthogonal Projections onto Capsule Subspaces · NeurIPS 2018
Machine learning › Generative modeling
generative adversarial network
0.312018
Global Versus Localized Generative Adversarial Nets · CVPR 2018
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction
manifold learning
0.312018
Global Versus Localized Generative Adversarial Nets · CVPR 2018
Machine learning › Efficient and distributed learning
orthogonal projection
0.312018
CapProNet: Deep Feature Learning via Orthogonal Projections onto Capsule Subspaces · NeurIPS 2018
Computer vision › 3D vision
point cloud analysis
0.312026
HCC-3D: Hierarchical Compensatory Compression for 98% 3D Token Reduction in Vision-Language Models · AAAI 2026
Machine learning › Deep learning architectures and training
recurrent neural network
0.312017
Stock Price Prediction via Discovering Multi-Frequency Trading Patterns · KDD 2017
Computational finance and economics › financial time series
financial time series prediction
0.312017
Stock Price Prediction via Discovering Multi-Frequency Trading Patterns · KDD 2017
Computational finance and economics › financial market prediction
stock price prediction
0.312017
Stock Price Prediction via Discovering Multi-Frequency Trading Patterns · KDD 2017
Computer vision › Video understanding and tracking
object tracking
0.212015
Vision-Inertial Hybrid Tracking for Robust and Efficient Augmented Reality on Smartphones · ACM Multimedia 2015
Virtual and augmented reality
augmented reality
0.212015
Vision-Inertial Hybrid Tracking for Robust and Efficient Augmented Reality on Smartphones · ACM Multimedia 2015
Virtual and augmented reality › tracking
pose tracking
0.212015
Vision-Inertial Hybrid Tracking for Robust and Efficient Augmented Reality on Smartphones · ACM Multimedia 2015
Machine learning › Learning theory
classification
0.112018
Global Versus Localized Generative Adversarial Nets · CVPR 2018
Machine learning › Trustworthy machine learning › robustness › robust learning
robust classification
0.112018
Global Versus Localized Generative Adversarial Nets · CVPR 2018
Machine learning › Deep learning architectures and training
sequence modeling
0.112017
Stock Price Prediction via Discovering Multi-Frequency Trading Patterns · KDD 2017
Wearable and physiological sensing
inertial sensing
0.112015
Vision-Inertial Hybrid Tracking for Robust and Efficient Augmented Reality on Smartphones · ACM Multimedia 2015

Methods — techniques the papers use, named apart from their topics

linear blend skinning · 2.03d morphable model · 2.03d gaussian splatting · 2.0global structure compression · 1.0adaptive detail mining · 1.0mutual information maximization · 1.0variational inference · 0.6autoencoding · 0.6kalman filter · 0.4featureless registration · 0.4feature-based tracking · 0.4regularization · 0.4dropconnect · 0.4recurrent network · 0.3inverse fourier transform · 0.3discrete fourier transform · 0.3
YearPublicationVenuePosition
2026 HCC-3D: Hierarchical Compensatory Compression for 98% 3D Token Reduction in Vision-Language Models
abstract
3D understanding has drawn significant attention recently, leveraging Vision-Language Models (VLMs) to enable multi-modal reasoning between point cloud and text data. Current 3D-VLMs directly embed the 3D point clouds into 3D tokens, following large 2D-VLMs with powerful reasoning capabilities. However, this framework has a great computational cost limiting its application, where we identify that the bottleneck lies in processing all 3D tokens in the Large Language Model (LLM) part. This raises the question: how can we reduce the computational overhead introduced by 3D tokens while preserving the integrity of their essential information? To address this question, we introduce Hierarchical Compensatory Compression (HCC-3D) to efficiently compress 3D tokens while maintaining critical detail retention. Specifically, we first propose a global structure compression (GSC), in which we design global queries to compress all 3D tokens into a few key tokens while keeping overall structural information. Then, to compensate for the information loss in GSC, we further propose an adaptive detail mining (ADM) module that selectively recompresses salient but under-attended features through complementary scoring. Extensive experiments demonstrate that HCC-3D not only achieves extreme compression ratios (approximately 98%) compared to previous 3D VLMs, but also achieves new state-of-the-art performance, showing the great improvements on both efficiency and performance.
Liheng Zhang, Bingfeng Zhang, Weifeng Liu 0001
AAAI1
2026 MGHead: Motion-Aware Animated Gaussian Head Avatars With Anchored Skeletal Structures
abstract
Creating photorealistic and animatable 3D head avatars across multiple views is still an ongoing challenge in AR/VR applications. Although previous studies adopt Neural Radiance Fields (NeRF) with 3D Morphable Model (3DMM) as a prior to produce impressive results in generating 3D heads, they incur considerable time costs and lack rendering quality. In this paper, we propose a novel approach called MGHead that synthesizes high-fidelity dynamic 3D head avatar with realistic appearance and complex deformation. It exploits anchor-based 3D Gaussians to model geometric shape and extends the fundamental deformation structure of a parametric morphable face model to each fixed point, providing greater robustness in responding to expressive variations. Notably, we introduce an expression and pose-dependent attribute adapter that injects driving signals into anchor features to refine neural Gaussian attributes. This adjustment compensates for the deficiencies of linear blend skinning in capturing high-frequency dynamics effectively, further improving the expressive realism and natural appearance of head avatar. Extensive experiments demonstrate the superiority of our model in visual detail quality and quantitative evaluations.
Haozhi Gu, Zubo Lu, Liheng Zhang, Weihao Yu 0002, Jin Huang 0007
IEEE Trans. Multim.3
2022 Learning Generalized Transformation Equivariant Representations Via AutoEncoding Transformations
abstract
Transformation equivariant representations (TERs) aim to capture the intrinsic visual structures that equivary to various transformations by expanding the notion of translation equivariance underlying the success of convolutional neural networks (CNNs). For this purpose, we present both deterministic AutoEncoding Transformations (AET) and probabilistic AutoEncoding Variational Transformations (AVT) models to learn visual representations from generic groups of transformations. While the AET is trained by directly decoding the transformations from the learned representations, the AVT is trained by maximizing the joint mutual information between the learned representation and transformations. This results in generalized TERs (GTERs) equivariant against transformations in a more general fashion by capturing complex patterns of visual structures beyond the conventional linear equivariance under a transformation group. The presented approach can be extended to (semi-)supervised models by jointly maximizing the mutual information of the learned representation with both labels and transformations. Experiments demonstrate the proposed models outperform the state-of-the-art models in both unsupervised and (semi-)supervised tasks. Moreover, we show that the unsupervised representation can even surpass the fully supervised representation pretrained on ImageNet when they are fine-tuned for the object detection task.
Guo-Jun Qi, Liheng Zhang, Feng Lin 0009, Xiao Wang 0013
IEEE Trans. Pattern Anal. Mach. Intell.2
2020 WCP: Worst-Case Perturbations for Semi-Supervised Deep Learning
abstract
In this paper, we present a novel regularization mechanism for training deep networks by minimizing the {\em Worse-Case Perturbation} (WCP). It is based on the idea that a robust model is least likely to be affected by small perturbations, such that its output decisions should be as stable as possible on both labeled and unlabeled examples. We will consider two forms of WCP regularizations -- additive and DropConnect perturbations, which impose additive noises on network weights, and make structural changes by dropping the network connections, respectively. We will show that the worse cases of both perturbations can be derived by solving respective optimization problems with spectral methods. The WCP can be minimized on both labeled and unlabeled data so that networks can be trained in a semi-supervised fashion. This leads to a novel paradigm of semi-supervised classifiers by stabilizing the predicted outputs in presence of the worse-case perturbations imposed on the network weights and structures.
Liheng Zhang, Guo-Jun Qi
CVPR1
2019 AET vs. AED: Unsupervised Representation Learning by Auto-Encoding Transformations Rather Than Data
abstract
The success of deep neural networks often relies on a large amount of labeled examples, which can be difficult to obtain in many real scenarios. To address this challenge, unsupervised methods are strongly preferred for training neural networks without using any labeled data. In this paper, we present a novel paradigm of unsupervised representation learning by Auto-Encoding Transformation (AET) in contrast to the conventional Auto-Encoding Data (AED) approach. Given a randomly sampled transformation, AET seeks to predict it merely from the encoded features as accurately as possible at the output end. The idea is the following: as long as the unsupervised features successfully encode the essential information about the visual structures of original and transformed images, the transformation can be well predicted. We will show that this AET paradigm allows us to instantiate a large variety of transformations, from parameterized, to non-parameterized and GAN-induced ones. Our experiments show that AET greatly improves over existing unsupervised approaches, setting new state-of-the-art performances being greatly closer to the upper bounds by their fully supervised counterparts on CIFAR-10, ImageNet and Places datasets.
Liheng Zhang, Guo-Jun Qi, Jiebo Luo 0001
CVPR1
2019 AVT: Unsupervised Learning of Transformation Equivariant Representations by Autoencoding Variational Transformations
abstract
The learning of Transformation-Equivariant Representations (TERs), which is introduced by Hinton et al. \cite{hinton2011transforming}, has been considered as a principle to reveal visual structures under various transformations. It contains the celebrated Convolutional Neural Networks (CNNs) as a special case that only equivary to the translations. In contrast, we seek to train TERs for a generic class of transformations and train them in an {\em unsupervised} fashion. To this end, we present a novel principled method by Autoencoding Variational Transformations (AVT), compared with the conventional approach to autoencoding data. Formally, given transformed images, the AVT seeks to train the networks by maximizing the mutual information between the transformations and representations. This ensures the resultant TERs of individual images contain the {\em intrinsic} information about their visual structures that would equivary {\em extricably} under various transformations in a generalized {\em nonlinear} case. Technically, we show that the resultant optimization problem can be efficiently solved by maximizing a variational lower-bound of the mutual information. This variational approach introduces a transformation decoder to approximate the intractable posterior of transformations, resulting in an autoencoding architecture with a pair of the representation encoder and the transformation decoder. Experiments demonstrate the proposed AVT model sets a new record for the performances on unsupervised tasks, greatly closing the performance gap to the supervised models.
Guo-Jun Qi, Liheng Zhang, Chang Wen Chen, Qi Tian 0001
ICCV2
2019 Gated Context Aggregation Network for Image Dehazing and Deraining
abstract
Image dehazing aims to recover the uncorrupted content from a hazy image. Instead of leveraging traditional low-level or handcrafted image priors as the restoration constraints, e.g., dark channels and increased contrast, we propose an end-to-end gated context aggregation network to directly restore the final haze-free image. In this network, we adopt the latest smoothed dilation technique to help remove the gridding artifacts caused by the widely-used dilated convolution with negligible extra parameters, and leverage a gated sub-network to fuse the features from different levels. Extensive experiments demonstrate that our method can surpass previous state-of-the-art methods by a large margin both quantitatively and qualitatively. In addition, to demonstrate the generality of the proposed method, we further apply it to the image deraining task, which also achieves the state-of-the-art performance.
Dongdong Chen 0001, Mingming He, Qingnan Fan, Jing Liao 0001, Liheng Zhang, Dongdong Hou, Lu Yuan 0001, Gang Hua 0001
WACV5
2019 Cascade Attention Machine for Occluded Landmark Detection in 2D X-Ray Angiography
abstract
In cardiac interventions, localization of guiding catheter tip in 2D fluoroscopic images is important to specify vessel branches and calibrate vessels with stenosis. While detection of guiding catheter tip is not trivial in contrast-free images due to low dose radiation as well as occlusion by other devices, it is even more challenging in contrast-filled images. As contrast-filled vessels become visible in X-ray imaging, the landmark of guiding catheter tip can often be completely occluded by the contrast medium. It is difficult even for human eyes to precisely localize the catheter tip from a single angiography image. Physicians have to rely on information before the inject of contrast medium to localize the guiding catheter tip occluded by contrast medium. Automatic landmark detection when occlusion happens is important and can significantly simplify the intervention workflow. To address this problem, we propose a novel Cascade Attention Machine (CAM) model. It borrows the idea of how human experts localize the catheter tip by first per-forming landmark detection when occlusion does not happen, then leveraging this information as prior knowledge to assist the occluded detection. Attention maps are computed from non-occluded detection to further refine the heatmaps for occluded detection to guide the inference focusing on related regions. Experiments on X-ray angiography demonstrate the promising performance compared with the state-of-the-art baselines. It shows that the CAM can capture the relation between situations with and without occlusion to achieve precise detection of occluded landmark.
Liheng Zhang, Vivek K. Singh 0002, Guo-Jun Qi, Terrence Chen
WACV1
2018 Global Versus Localized Generative Adversarial Nets
abstract
In this paper, we present a novel localized Generative Adversarial Net (GAN) to learn on the manifold of real data. Compared with the classic GAN that globally parameterizes a manifold, the Localized GAN (LGAN) uses local coordinate charts to parameterize distinct local geometry of how data points can transform at different locations on the manifold. Specifically, around each point there exists a local generator that can produce data following diverse patterns of transformations on the manifold. The locality nature of LGAN enables local generators to adapt to and directly access the local geometry without need to invert the generator in a global GAN. Furthermore, it can prevent the manifold from being locally collapsed to a dimensionally deficient tangent subspace by imposing an orthonormality prior between tangents. This provides a geometric approach to alleviating mode collapse at least locally on the manifold by imposing independence between data transformations in different tangent directions. We will also demonstrate the LGAN can be applied to train a robust classifier that prefers locally consistent classification decisions on the manifold, and the resultant regularizer is closely related with the Laplace-Beltrami operator. Our experiments show that the proposed LGANs can not only produce diverse image transformations, but also deliver superior classification performances.
Guo-Jun Qi, Liheng Zhang, Hao Hu 0010, Marzieh Edraki, Jingdong Wang 0001, Xian-Sheng Hua 0001
CVPR2
2018 Automated Pulmonary Nodule Detection: High Sensitivity with Few Candidates
Bin Wang 0065, Guo-Jun Qi, Sheng Tang, Liheng Zhang, Lixi Deng, Yongdong Zhang 0001
MICCAI (2)4
2018 CapProNet: Deep Feature Learning via Orthogonal Projections onto Capsule Subspaces
abstract
In this paper, we formalize the idea behind capsule nets of using a capsule vector rather than a neuron activation to predict the label of samples. To this end, we propose to learn a group of capsule subspaces onto which an input feature vector is projected. Then the lengths of resultant capsules are used to score the probability of belonging to different classes. We train such a Capsule Projection Network (CapProNet) by learning an orthogonal projection matrix for each capsule subspace, and show that each capsule subspace is updated until it contains input feature vectors corresponding to the associated class. With low dimensionality of capsule subspace as well as an iterative method to estimate the matrix inverse, only a small negligible computing overhead is incurred to train the network. Experiment results on image datasets show the presented network can greatly improve the performance of state-of-the-art Resnet backbones by $10-20\%$ with almost the same computing cost.
Liheng Zhang, Marzieh Edraki, Guo-Jun Qi
NeurIPS1
2017 Stock Price Prediction via Discovering Multi-Frequency Trading Patterns
abstract
Stock prices are formed based on short and/or long-term commercial and trading activities that reflect different frequencies of trading patterns. However, these patterns are often elusive as they are affected by many uncertain political-economic factors in the real world, such as corporate performances, government policies, and even breaking news circulated across markets. Moreover, time series of stock prices are non-stationary and non-linear, making the prediction of future price trends much challenging. To address them, we propose a novel State Frequency Memory (SFM) recurrent network to capture the multi-frequency trading patterns from past market data to make long and short term predictions over time. Inspired by Discrete Fourier Transform (DFT), the SFM decomposes the hidden states of memory cells into multiple frequency components, each of which models a particular frequency of latent trading pattern underlying the fluctuation of stock price. Then the future stock prices are predicted as a nonlinear mapping of the combination of these components in an Inverse Fourier Transform (IFT) fashion. Modeling multi-frequency trading patterns can enable more accurate predictions for various time ranges: while a short-term prediction usually depends on high frequency trading patterns, a long-term prediction should focus more on the low frequency trading patterns targeting at long-term return. Unfortunately, no existing model explicitly distinguishes between various frequencies of trading patterns to make dynamic predictions in literature. The experiments on the real market data also demonstrate more competitive performance by the SFM as compared with the state-of-the-art methods.
Liheng Zhang, Charu C. Aggarwal, Guo-Jun Qi
KDD1
2015 Vision-Inertial Hybrid Tracking for Robust and Efficient Augmented Reality on Smartphones
abstract
This paper aims at robust and efficient pose tracking for augmented reality on modern smartphones. Existing methods, relying on either vision analysis or motion sensing, are either too computationally expensive to achieve real-time performance on a smartphone, or too noisy to achieve sufficient robustness. This paper presents a hybrid tracking system which can achieve real-time performance with high robustness. Our system utilizes an efficient featureless method based on pixel-based registration to track the object pose on every frame. The featureless tracking result is revised from time to time by a feature-based method to reduce tracking errors. Both featureless and feature-based tracking results are sensitive to large motion blurs. To improve the robustness, an adaptive Kamlan filter is proposed to fuse the visual tracking results with the inertial tracking results computed form phone's built-in sensors. Our hybrid method is evaluated on a dataset consisting of 16 video clips with synchronized inertial sensing data. Experimental results demonstrated the superior performance of our method to state-of-the-art visual tracking methods [5, 12] on smartphones. The dataset will be made publicly available with the publication of this paper.
Xin Yang 0008, Xun Si, Tangli Xue, Liheng Zhang, Kwang-Ting Cheng
ACM Multimedia4