Ze Huang

dblp:117/2935 · DBLP profile ↗
← Back
10ranked-venue papers
4as first author
8since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
3D vision · 40% Generative modeling · 18% Knowledge representation and reasoning · 10%

Topics — the 14 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision
3d reconstruction
0.912025
MS-Road: Towards Spatiotemporal-Consistent Large-Scale Road Reconstruction · ACM Multimedia 2025
Computer vision › 3D vision
3d scene understanding
0.912025
From Flatland to Space: Teaching Vision-Language Models to Perceive and Reason in 3D · NeurIPS 2025
Computer vision › 3D vision › 3d scene reconstruction
road surface reconstruction
0.912025
MS-Road: Towards Spatiotemporal-Consistent Large-Scale Road Reconstruction · ACM Multimedia 2025
Knowledge, reasoning and agents › Knowledge representation and reasoning
spatial reasoning
0.912025
From Flatland to Space: Teaching Vision-Language Models to Perceive and Reason in 3D · NeurIPS 2025
Machine learning › Generative modeling › diffusion model › controllable generation
controllable scene generation
0.812024
WoVoGen: World Volume-Aware Diffusion for Controllable Multi-camera Driving Scene Generation · ECCV (80) 2024
Machine learning › Generative modeling
diffusion model
0.812024
WoVoGen: World Volume-Aware Diffusion for Controllable Multi-camera Driving Scene Generation · ECCV (80) 2024
Robotics › Autonomous driving › scenario generation
driving scene generation
0.812024
WoVoGen: World Volume-Aware Diffusion for Controllable Multi-camera Driving Scene Generation · ECCV (80) 2024
Computer vision › Image recognition and object detection › object detection
oriented object detection
0.812024
Orientation-Aware Multi-Modal Learning for Road Intersection Identification and Mapping · ICRA 2024
Computer vision › 3D vision
point cloud processing
0.812024
Orientation-Aware Multi-Modal Learning for Road Intersection Identification and Mapping · ICRA 2024
Machine learning › Kernel, tree and ensemble methods › ensemble learning
neural decision forests
0.412020
Dynamic Connected Neural Decision Classifier and Regressor with Dynamic Softing Pruning · ICDM 2020
Computer vision › Vision and language
spatial reasoning benchmark
0.312025
From Flatland to Space: Teaching Vision-Language Models to Perceive and Reason in 3D · NeurIPS 2025
Computer vision › Vision and language
vision-language model
0.312025
From Flatland to Space: Teaching Vision-Language Models to Perceive and Reason in 3D · NeurIPS 2025
Machine learning › Efficient and distributed learning
model compression
0.112020
Dynamic Connected Neural Decision Classifier and Regressor with Dynamic Softing Pruning · ICDM 2020
Machine learning › Efficient and distributed learning › model compression
pruning
0.112020
Dynamic Connected Neural Decision Classifier and Regressor with Dynamic Softing Pruning · ICDM 2020

Methods — techniques the papers use, named apart from their topics

spatio-temporal consistency · 0.9self-supervised learning · 0.9multi-view consistency · 0.9annotation pipeline · 0.92d spatial data generation · 0.9world volume awareness · 0.8trigonometry-induced regression · 0.8multimodal fusion · 0.8diffusion model · 0.8end-to-end training · 0.4
YearPublicationVenuePosition
2025 MS-Road: Towards Spatiotemporal-Consistent Large-Scale Road Reconstruction
abstract
Road surface reconstruction is crucial for autonomous driving, providing accurate and up-to-date road geometry for navigation, safety assessment, and infrastructure maintenance. Camera-based methods have become increasingly cost-effective and scalable for this task. However, achieving high-quality large-scale reconstruction remains challenging due to inconsistent observations of the same road surface points. These inconsistencies arise both within single sessions-caused by factors such as vehicle shadows and exposure shifts-and across multiple sessions, where changes in lighting conditions and viewpoints further exacerbate the problem. To address these challenges, we propose MS-Road, a camera-based approach for large-scale road surface reconstruction with strong geometric consistency. MS-Road leveraging self-supervised learning and multi-view consistent constrain to tackle two key issues: inconsistency in road appearance across different observations and inaccurate road height localization. By enforcing spatiotemporal consistency in both geometric and visual aspects, our method produces more reliable reconstructions within and across sessions. Experiments on two public datasets and a real-world dataset demonstrate that our approach achieves robust and high-fidelity reconstruction under diverse and challenging conditions.
Ze Huang, Zhongyang Xiao, Mingliang Song, Hongyuan Yuan, Kevin Li Sun, Li Zhang 0040
ACM Multimedia1
2025 From Flatland to Space: Teaching Vision-Language Models to Perceive and Reason in 3D
abstract
Recent advances in LVLMs have improved vision-language understanding, but they still struggle with spatial perception, limiting their ability to reason about complex 3D scenes. Unlike previous approaches that incorporate 3D representations into models to improve spatial understanding, we aim to unlock the potential of VLMs by leveraging spatially relevant image data. To this end, we introduce a novel 2D spatial data generation and annotation pipeline built upon scene data with 3D ground-truth. This pipeline enables the creation of a diverse set of spatial tasks, ranging from basic perception tasks to more complex reasoning tasks. Leveraging this pipeline, we construct SPAR-7M, a large-scale dataset generated from thousands of scenes across multiple public datasets. In addition, we introduce SPAR-Bench, a benchmark designed to offer a more comprehensive evaluation of spatial capabilities compared to existing spatial benchmarks, supporting both single-view and multi-view inputs. Training on both SPAR-7M and large-scale 2D datasets enables our models to achieve state-of-the-art performance on 2D spatial benchmarks. Further fine-tuning on 3D task-specific datasets yields competitive results, underscoring the effectiveness of our dataset in enhancing spatial reasoning.
Yurui Chen, Yueming Xu, Ze Huang, Jilin Mei, Junhui Chen, Yanpeng Zhou, Yu-Jie Yuan, Xinyue Cai, Xingyue Quan, Hang Xu 0004, Li Zhang 0040
NeurIPS4
2025 4D-VLA: Spatiotemporal Vision-Language-Action Pretraining with Cross-Scene Calibration
abstract
Leveraging diverse robotic data for pretraining remains a critical challenge. Existing methods typically model the dataset’s action distribution using simple observations as inputs. However, these inputs are often incomplete, resulting in a dispersed conditional action distribution—an issue we refer to as coordinate system chaos and state chaos. This inconsistency significantly hampers pretraining efficiency. To address this, we propose 4D-VLA, a novel approach that effectively integrates 4D information into the input to mitigate these sources of chaos. Our model introduces depth and temporal information into visual features with sequential RGB-D inputs, aligning the coordinate systems of the robot and the scene. This alignment endows the model with strong spatiotemporal reasoning capabilities while minimizing training overhead. Additionally, we introduce Memory bank sampling, a frame sampling strategy designed to extract informative frames from historical images, further improving effectiveness and efficiency. Experimental results demonstrate that our pretraining method and architectural components substantially enhance model performance. In both simulated and real-world experiments, our model achieves a significant increase in success rate over OpenVLA.To further assess spatial perception and generalization to novel views, we introduce MV-Bench, a multi-view simulation benchmark. Our model consistently outperforms existing methods, demonstrating stronger spatial understanding and adaptability.
Yurui Chen, Yueming Xu, Ze Huang, Yanpeng Zhou, Yu-Jie Yuan, Xinyue Cai, Xingyue Quan, Hang Xu 0004, Li Zhang 0040
NeurIPS4
2024 WoVoGen: World Volume-Aware Diffusion for Controllable Multi-camera Driving Scene Generation
Ze Huang, Zeyu Yang 0004, Li Zhang 0040
ECCV (80)2
2024 Orientation-Aware Multi-Modal Learning for Road Intersection Identification and Mapping
abstract
Accurate identification of road intersections is the pivotal task for automatic construction of high-definition maps, particularly in unstructured scenes. Existing methods predominantly rely on single-modal data and thus show an obvious unimodal limitation, i.e., lack of contextual information. Moreover, these approaches overlook the benefits of leveraging multi-modal data fusion and representation learning that is crucial for generalizability. To this end, we propose a novel orientation-aware multi-modal learning paradigm, which formulates intersection identification as an oriented object detection task. Specifically, heterogeneous fusion is introduced to harmonize disparate data modalities, i.e., vector maps, point clouds, and vehicle trajectories, into a unified feature space. Concurrently, we present trigonometry-induced adaptive regression to elevate orientation estimation, while mitigating issues related to scale imbalance and boundary confusion through dual-objective matching with spatial adaptation. To evaluate our methodology, we assemble the first-of-its-kind multi-modal benchmark tailored for complex low-speed environments, complete with fine-grained semantic annotations for intersections. Comprehensive empirical analyses, including ablation studies, affirm both the superior performance of our proposed framework and the efficacy of its constituent modules.
Qibin He 0001, Zhongyang Xiao, Ze Huang, Hongyuan Yuan, Li Sun 0005
ICRA3
2023 MLPST: MLP is All You Need for Spatio-Temporal Prediction
abstract
Traffic prediction is a typical spatio-temporal data mining task and has great significance to the public transportation system. Considering the demand for its grand application, we recognize key factors for an ideal spatio-temporal prediction method: efficient, lightweight, and effective. However, the current deep model-based spatio-temporal prediction solutions generally own intricate architectures with cumbersome optimization, which can hardly meet these expectations. To accomplish the above goals, we propose an intuitive and novel framework, MLPST, a pure multi-layer perceptron architecture for traffic prediction. Specifically, we first capture spatial relationships from both local and global receptive fields. Then, temporal dependencies in different intervals are comprehensively considered. Through compact and swift MLP processing, MLPST can well capture the spatial and temporal dependencies while requiring only linear computational complexity, as well as model parameters that are more than an order of magnitude lower than baselines. Extensive experiments validated the superior effectiveness and efficiency of MLPST against advanced baselines, and among models with optimal accuracy, MLPST achieves the best time and space efficiency.
Zijian Zhang 0009, Ze Huang, Zhiwei Hu, Xiangyu Zhao 0001, Zitao Liu 0001, Junbo Zhang 0004, S. Joe Qin
CIKM2
2023 EventPoint: Self-Supervised Interest Point Detection and Description for Event-based Camera
abstract
This paper proposes a self-supervised learned local detector and descriptor, called EventPoint, for event stream/camera tracking and registration. Event-based cameras have grown in popularity because of their biological inspiration and low power consumption. Despite this, applying local features directly to the event stream is difficult due to its peculiar data structure. We propose a new time-surface-like event stream representation method called Ten-code. The event stream data processed by Tencode can obtain the pixel-level positioning of interest points while also simultaneously extracting descriptors through a neural network. Instead of using costly and unreliable manual annotation, our network leverages the prior knowledge of local feature extraction on color images and conducts self-supervised learning via homographic and spatio-temporal adaptation. To the best of our knowledge, our proposed method is the first research on event-based local features learning using a deep neural network. We provide comprehensive experiments of feature point detection and matching, and three public datasets are used for evaluation (i.e. DSEC, N-Caltech101, and HVGA ATIS Corner Dataset). The experimental findings demonstrate that our method outperforms SOTA in terms of feature point detection and description.
Ze Huang, Li Sun 0005, Cheng Zhao 0002, Songzhi Su
WACV1
2022 VEFNet: an Event-RGB Cross Modality Fusion Network for Visual Place Recognition
abstract
Visual Place Recognition (VPR) on natural image is challenging due to the illumination variance and seasonal changes. In terms of long-term localization, the emerging event stream cameras are naturally resilient to appearance changes. In this paper, we propose a novel multi-modal network, e.g. VEFNet for VPR by learning location-specific cross RGB-event modality feature representations. Specifically, we firstly extract dense visual features via shared Convolutional Neural Network (CNN) backbone from RGB and event frames separately. Then, two branch features are fed to the cross-modality attention module to establish correspondences between the dual-modality. We also employ a self-attention module to enhance the contextual integration within densely encoded features. Finally, the learned global descriptor is used as the place representation of the dual-modality inputs for VPR. Experimental results demonstrate the state-of-the-art (SOTA) performance on the public datasets
Ze Huang, Li Sun 0005, Cheng Zhao 0002, Min Huang 0004, Songzhi Su
ICIP1
2020 Dynamic Connected Neural Decision Classifier and Regressor with Dynamic Softing Pruning
abstract
To deal with datasets of different complexity, this paper presents an efTo deal with various datasets over different complexity, this paper presents an self-adaptive learning model that combines the proposed Dynamic Connected Neural Decision Networks (DNDN) and a new pruning method–Dynamic Soft Pruning (DSP). DNDN is a combination of random forests and deep neural networks that enjoys both the advantages of strong classification capability of tree-like structure and representation learning capability of network structure. Based on Deep Neural Decision Forests (DNDF), this paper adopts an end-to-end training approach by representing the classification distribution with multiple randomly initialized softmax layers, which further allows an ensemble of multiple random forests attached to layers of neural network with different depth. We also propose a soft pruning method DSP to reduce the redundant connections of the network adaptively to avoid over-fitting simple dataset. The model demonstrates no performance loss compared with unpruned models and even higher robustness over different data and feature distribution. Extensive experiments on different datasets demonstrate the superiority of the proposed model over other popular algorithms in solving classification tasks.ficient learning model that combines the proposed Dynamic Connected Neural Decision Networks (DNDN) and a new pruning method–Dynamic Soft Pruning (DSP). DNDN is a combination of random forests and deep neural networks thereby it enjoys both the properties of powerful classification capability and representation learning functionality. Different from Deep Neural Decision Forests (DNDF), this paper adopts an end-to-end training approach by representing the classification distribution with multiple randomly initialized softmax layers, which enables the placement of the forest trees after each layer in the neural network and greatly improves the training speed and stability. Furthermore, DSP is proposed to reduce the redundant connections of the network in a soft fashion which has high flexibility but demonstrates no performance loss compared with previous approaches. Extensive experiments on different datasets demonstrate the superiority of the proposed model over other popular algorithms in solving classification tasks.
Jianfei Song, Ze Huang
ICDM7
2012 Mining Google Scholar Citations: An Exploratory Study
Ze Huang, Bo Yuan 0003
ICIC (1)1