EDBT 2026 Demo / reviewers in the wild / expert
Ze Huang
dblp:117/2935
· DBLP profile ↗
10ranked-venue papers
4as first author
8since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
3D vision · 40% Generative modeling · 18% Knowledge representation and reasoning · 10% |
Topics — the 14 heaviest of 14, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › 3D vision
3d reconstruction |
0.9 | 1 | 2025 | MS-Road: Towards Spatiotemporal-Consistent Large-Scale Road Reconstruction · ACM Multimedia 2025 |
Computer vision › 3D vision
3d scene understanding |
0.9 | 1 | 2025 | From Flatland to Space: Teaching Vision-Language Models to Perceive and Reason in 3D · NeurIPS 2025 |
Computer vision › 3D vision › 3d scene reconstruction
road surface reconstruction |
0.9 | 1 | 2025 | MS-Road: Towards Spatiotemporal-Consistent Large-Scale Road Reconstruction · ACM Multimedia 2025 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
spatial reasoning |
0.9 | 1 | 2025 | From Flatland to Space: Teaching Vision-Language Models to Perceive and Reason in 3D · NeurIPS 2025 |
Machine learning › Generative modeling › diffusion model › controllable generation
controllable scene generation |
0.8 | 1 | 2024 | WoVoGen: World Volume-Aware Diffusion for Controllable Multi-camera Driving Scene Generation · ECCV (80) 2024 |
Machine learning › Generative modeling
diffusion model |
0.8 | 1 | 2024 | WoVoGen: World Volume-Aware Diffusion for Controllable Multi-camera Driving Scene Generation · ECCV (80) 2024 |
Robotics › Autonomous driving › scenario generation
driving scene generation |
0.8 | 1 | 2024 | WoVoGen: World Volume-Aware Diffusion for Controllable Multi-camera Driving Scene Generation · ECCV (80) 2024 |
Computer vision › Image recognition and object detection › object detection
oriented object detection |
0.8 | 1 | 2024 | Orientation-Aware Multi-Modal Learning for Road Intersection Identification and Mapping · ICRA 2024 |
Computer vision › 3D vision
point cloud processing |
0.8 | 1 | 2024 | Orientation-Aware Multi-Modal Learning for Road Intersection Identification and Mapping · ICRA 2024 |
Machine learning › Kernel, tree and ensemble methods › ensemble learning
neural decision forests |
0.4 | 1 | 2020 | Dynamic Connected Neural Decision Classifier and Regressor with Dynamic Softing Pruning · ICDM 2020 |
Computer vision › Vision and language
spatial reasoning benchmark |
0.3 | 1 | 2025 | From Flatland to Space: Teaching Vision-Language Models to Perceive and Reason in 3D · NeurIPS 2025 |
Computer vision › Vision and language
vision-language model |
0.3 | 1 | 2025 | From Flatland to Space: Teaching Vision-Language Models to Perceive and Reason in 3D · NeurIPS 2025 |
Machine learning › Efficient and distributed learning
model compression |
0.1 | 1 | 2020 | Dynamic Connected Neural Decision Classifier and Regressor with Dynamic Softing Pruning · ICDM 2020 |
Machine learning › Efficient and distributed learning › model compression
pruning |
0.1 | 1 | 2020 | Dynamic Connected Neural Decision Classifier and Regressor with Dynamic Softing Pruning · ICDM 2020 |
Methods — techniques the papers use, named apart from their topics
spatio-temporal consistency · 0.9self-supervised learning · 0.9multi-view consistency · 0.9annotation pipeline · 0.92d spatial data generation · 0.9world volume awareness · 0.8trigonometry-induced regression · 0.8multimodal fusion · 0.8diffusion model · 0.8end-to-end training · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MS-Road: Towards Spatiotemporal-Consistent Large-Scale Road ReconstructionabstractRoad surface reconstruction is crucial for autonomous driving, providing accurate and up-to-date road geometry for navigation, safety assessment, and infrastructure maintenance. Camera-based methods have become increasingly cost-effective and scalable for this task. However, achieving high-quality large-scale reconstruction remains challenging due to inconsistent observations of the same road surface points. These inconsistencies arise both within single sessions-caused by factors such as vehicle shadows and exposure shifts-and across multiple sessions, where changes in lighting conditions and viewpoints further exacerbate the problem. To address these challenges, we propose MS-Road, a camera-based approach for large-scale road surface reconstruction with strong geometric consistency. MS-Road leveraging self-supervised learning and multi-view consistent constrain to tackle two key issues: inconsistency in road appearance across different observations and inaccurate road height localization. By enforcing spatiotemporal consistency in both geometric and visual aspects, our method produces more reliable reconstructions within and across sessions. Experiments on two public datasets and a real-world dataset demonstrate that our approach achieves robust and high-fidelity reconstruction under diverse and challenging conditions. Ze Huang, Zhongyang Xiao, Mingliang Song, Hongyuan Yuan, Kevin Li Sun, Li Zhang 0040 |
ACM Multimedia | 1 |
| 2025 | From Flatland to Space: Teaching Vision-Language Models to Perceive and Reason in 3DabstractRecent advances in LVLMs have improved vision-language understanding, but they still struggle with spatial perception, limiting their ability to reason about complex 3D scenes. Unlike previous approaches that incorporate 3D representations into models to improve spatial understanding, we aim to unlock the potential of VLMs by leveraging spatially relevant image data. To this end, we introduce a novel 2D spatial data generation and annotation pipeline built upon scene data with 3D ground-truth. This pipeline enables the creation of a diverse set of spatial tasks, ranging from basic perception tasks to more complex reasoning tasks. Leveraging this pipeline, we construct SPAR-7M, a large-scale dataset generated from thousands of scenes across multiple public datasets. In addition, we introduce SPAR-Bench, a benchmark designed to offer a more comprehensive evaluation of spatial capabilities compared to existing spatial benchmarks, supporting both single-view and multi-view inputs. Training on both SPAR-7M and large-scale 2D datasets enables our models to achieve state-of-the-art performance on 2D spatial benchmarks. Further fine-tuning on 3D task-specific datasets yields competitive results, underscoring the effectiveness of our dataset in enhancing spatial reasoning. Yurui Chen, Yueming Xu, Ze Huang, Jilin Mei, Junhui Chen, Yanpeng Zhou, Yu-Jie Yuan, Xinyue Cai, Xingyue Quan, Hang Xu 0004, Li Zhang 0040 |
NeurIPS | 4 |
| 2025 | 4D-VLA: Spatiotemporal Vision-Language-Action Pretraining with Cross-Scene CalibrationabstractLeveraging diverse robotic data for pretraining remains a critical challenge. Existing methods typically model the dataset’s action distribution using simple observations as inputs. However, these inputs are often incomplete, resulting in a dispersed conditional action distribution—an issue we refer to as coordinate system chaos and state chaos. This inconsistency significantly hampers pretraining efficiency. To address this, we propose 4D-VLA, a novel approach that effectively integrates 4D information into the input to mitigate these sources of chaos. Our model introduces depth and temporal information into visual features with sequential RGB-D inputs, aligning the coordinate systems of the robot and the scene. This alignment endows the model with strong spatiotemporal reasoning capabilities while minimizing training overhead. Additionally, we introduce Memory bank sampling, a frame sampling strategy designed to extract informative frames from historical images, further improving effectiveness and efficiency. Experimental results demonstrate that our pretraining method and architectural components substantially enhance model performance. In both simulated and real-world experiments, our model achieves a significant increase in success rate over OpenVLA.To further assess spatial perception and generalization to novel views, we introduce MV-Bench, a multi-view simulation benchmark. Our model consistently outperforms existing methods, demonstrating stronger spatial understanding and adaptability. Yurui Chen, Yueming Xu, Ze Huang, Yanpeng Zhou, Yu-Jie Yuan, Xinyue Cai, Xingyue Quan, Hang Xu 0004, Li Zhang 0040 |
NeurIPS | 4 |
| 2024 | WoVoGen: World Volume-Aware Diffusion for Controllable Multi-camera Driving Scene Generation
Ze Huang, Zeyu Yang 0004, Li Zhang 0040 |
ECCV (80) | 2 |
| 2024 | Orientation-Aware Multi-Modal Learning for Road Intersection Identification and MappingabstractAccurate identification of road intersections is the pivotal task for automatic construction of high-definition maps, particularly in unstructured scenes. Existing methods predominantly rely on single-modal data and thus show an obvious unimodal limitation, i.e., lack of contextual information. Moreover, these approaches overlook the benefits of leveraging multi-modal data fusion and representation learning that is crucial for generalizability. To this end, we propose a novel orientation-aware multi-modal learning paradigm, which formulates intersection identification as an oriented object detection task. Specifically, heterogeneous fusion is introduced to harmonize disparate data modalities, i.e., vector maps, point clouds, and vehicle trajectories, into a unified feature space. Concurrently, we present trigonometry-induced adaptive regression to elevate orientation estimation, while mitigating issues related to scale imbalance and boundary confusion through dual-objective matching with spatial adaptation. To evaluate our methodology, we assemble the first-of-its-kind multi-modal benchmark tailored for complex low-speed environments, complete with fine-grained semantic annotations for intersections. Comprehensive empirical analyses, including ablation studies, affirm both the superior performance of our proposed framework and the efficacy of its constituent modules. Qibin He 0001, Zhongyang Xiao, Ze Huang, Hongyuan Yuan, Li Sun 0005 |
ICRA | 3 |
| 2023 | MLPST: MLP is All You Need for Spatio-Temporal PredictionabstractTraffic prediction is a typical spatio-temporal data mining task and has great significance to the public transportation system. Considering the demand for its grand application, we recognize key factors for an ideal spatio-temporal prediction method: efficient, lightweight, and effective. However, the current deep model-based spatio-temporal prediction solutions generally own intricate architectures with cumbersome optimization, which can hardly meet these expectations. To accomplish the above goals, we propose an intuitive and novel framework, MLPST, a pure multi-layer perceptron architecture for traffic prediction. Specifically, we first capture spatial relationships from both local and global receptive fields. Then, temporal dependencies in different intervals are comprehensively considered. Through compact and swift MLP processing, MLPST can well capture the spatial and temporal dependencies while requiring only linear computational complexity, as well as model parameters that are more than an order of magnitude lower than baselines. Extensive experiments validated the superior effectiveness and efficiency of MLPST against advanced baselines, and among models with optimal accuracy, MLPST achieves the best time and space efficiency. Zijian Zhang 0009, Ze Huang, Zhiwei Hu, Xiangyu Zhao 0001, Zitao Liu 0001, Junbo Zhang 0004, S. Joe Qin |
CIKM | 2 |
| 2023 | EventPoint: Self-Supervised Interest Point Detection and Description for Event-based CameraabstractThis paper proposes a self-supervised learned local detector and descriptor, called EventPoint, for event stream/camera tracking and registration. Event-based cameras have grown in popularity because of their biological inspiration and low power consumption. Despite this, applying local features directly to the event stream is difficult due to its peculiar data structure. We propose a new time-surface-like event stream representation method called Ten-code. The event stream data processed by Tencode can obtain the pixel-level positioning of interest points while also simultaneously extracting descriptors through a neural network. Instead of using costly and unreliable manual annotation, our network leverages the prior knowledge of local feature extraction on color images and conducts self-supervised learning via homographic and spatio-temporal adaptation. To the best of our knowledge, our proposed method is the first research on event-based local features learning using a deep neural network. We provide comprehensive experiments of feature point detection and matching, and three public datasets are used for evaluation (i.e. DSEC, N-Caltech101, and HVGA ATIS Corner Dataset). The experimental findings demonstrate that our method outperforms SOTA in terms of feature point detection and description. Ze Huang, Li Sun 0005, Cheng Zhao 0002, Songzhi Su |
WACV | 1 |
| 2022 | VEFNet: an Event-RGB Cross Modality Fusion Network for Visual Place RecognitionabstractVisual Place Recognition (VPR) on natural image is challenging due to the illumination variance and seasonal changes. In terms of long-term localization, the emerging event stream cameras are naturally resilient to appearance changes. In this paper, we propose a novel multi-modal network, e.g. VEFNet for VPR by learning location-specific cross RGB-event modality feature representations. Specifically, we firstly extract dense visual features via shared Convolutional Neural Network (CNN) backbone from RGB and event frames separately. Then, two branch features are fed to the cross-modality attention module to establish correspondences between the dual-modality. We also employ a self-attention module to enhance the contextual integration within densely encoded features. Finally, the learned global descriptor is used as the place representation of the dual-modality inputs for VPR. Experimental results demonstrate the state-of-the-art (SOTA) performance on the public datasets Ze Huang, Li Sun 0005, Cheng Zhao 0002, Min Huang 0004, Songzhi Su |
ICIP | 1 |
| 2020 | Dynamic Connected Neural Decision Classifier and Regressor with Dynamic Softing PruningabstractTo deal with datasets of different complexity, this paper presents an efTo deal with various datasets over different complexity, this paper presents an self-adaptive learning model that combines the proposed Dynamic Connected Neural Decision Networks (DNDN) and a new pruning method–Dynamic Soft Pruning (DSP). DNDN is a combination of random forests and deep neural networks that enjoys both the advantages of strong classification capability of tree-like structure and representation learning capability of network structure. Based on Deep Neural Decision Forests (DNDF), this paper adopts an end-to-end training approach by representing the classification distribution with multiple randomly initialized softmax layers, which further allows an ensemble of multiple random forests attached to layers of neural network with different depth. We also propose a soft pruning method DSP to reduce the redundant connections of the network adaptively to avoid over-fitting simple dataset. The model demonstrates no performance loss compared with unpruned models and even higher robustness over different data and feature distribution. Extensive experiments on different datasets demonstrate the superiority of the proposed model over other popular algorithms in solving classification tasks.ficient learning model that combines the proposed Dynamic Connected Neural Decision Networks (DNDN) and a new pruning method–Dynamic Soft Pruning (DSP). DNDN is a combination of random forests and deep neural networks thereby it enjoys both the properties of powerful classification capability and representation learning functionality. Different from Deep Neural Decision Forests (DNDF), this paper adopts an end-to-end training approach by representing the classification distribution with multiple randomly initialized softmax layers, which enables the placement of the forest trees after each layer in the neural network and greatly improves the training speed and stability. Furthermore, DSP is proposed to reduce the redundant connections of the network in a soft fashion which has high flexibility but demonstrates no performance loss compared with previous approaches. Extensive experiments on different datasets demonstrate the superiority of the proposed model over other popular algorithms in solving classification tasks. Jianfei Song, Ze Huang |
ICDM | 7 |
| 2012 | Mining Google Scholar Citations: An Exploratory Study
Ze Huang, Bo Yuan 0003 |
ICIC (1) | 1 |