VLDB 2026 Research / reviewers in the wild / expert
Zehao Huang
dblp:197/1644
· DBLP profile ↗
20ranked-venue papers
2as first author
15since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 1 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MV2DFusion: Leveraging Modality-Specific Object Semantics for Multi-Modal 3D DetectionabstractThe rise of autonomous vehicles has significantly increased the demand for robust 3D object detection systems. While cameras and LiDAR sensors each offer unique advantages-cameras provide rich texture information and LiDAR offers precise 3D spatial data-relying on a single modality often leads to performance limitations. This paper introduces MV2DFusion, a multi-modal detection framework that integrates the strengths of both worlds through an advanced query-based fusion mechanism. By introducing an image query generator to align with image-specific attributes and a point cloud query generator, MV2DFusion effectively combines modality-specific object semantics without biasing toward one single modality. Then the sparse fusion process can be accomplished based on the valuable object semantics, ensuring efficient and accurate object detection across various scenarios. Our framework's flexibility allows it to integrate with any image and point cloud-based detectors, showcasing its adaptability and potential for future advancements. Extensive evaluations on the nuScenes and Argoverse2 datasets demonstrate that MV2DFusion achieves state-of-the-art performance, particularly excelling in long-range detection scenarios. Zitian Wang, Zehao Huang, Yulu Gao, Naiyan Wang, Si Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | Finding Associative Entities in Knowledge Graph by Incorporating User BehaviorsabstractThe task of finding associative entities in knowledge graph (KG) is to provide a ranking list of entities according to their association degrees. However, many entities are not only linked in KG but also associated in terms of user behaviors, which facilitates finding associative entities accurately. This manuscript incorporates KG with user-generated data to propose the Association Entity Graph Model (AEGM) to evaluate the association degrees. They first propose the joint weighting function to evaluate the entity associations and prove its submodularity theoretically as well as the greedy algorithm to select the candidates efficiently. They define the entity association information to score the entity association and give the hill climbing search based algorithm for AEGM construction. Following, they embed AEGM to calculate the association degrees and obtain the associative entities efficiently. Extensive experiments on three datasets show that the proposed method can achieve a better performance than some state-of-the-art competitors in accurately finding associative entities. Peizhong Yang, Kun Yue, Liang Duan, Zehao Huang |
J. Database Manag. | 5 |
| 2025 | Anchor3DLane++: 3D Lane Detection via Sample-Adaptive Sparse 3D Anchor RegressionabstractIn this paper, we focus on the challenging task of monocular 3D lane detection. Previous methods typically adopt inverse perspective mapping (IPM) to transform the Front-Viewed (FV) images or features into the Bird-Eye-Viewed (BEV) space for lane detection. However, IPM's dependence on flat ground assumption and context information loss in BEV representations lead to inaccurate 3D information estimation. Though efforts have been made to bypass BEV and directly predict 3D lanes from FV representations, their performances still fall behind BEV-based methods due to a lack of structured modeling of 3D lanes. In this paper, we propose a novel BEV-free method named Anchor3DLane++ which defines 3D lane anchors as structural representations and makes predictions directly from FV features. We also design a Prototype-based Adaptive Anchor Generation (PAAG) module to generate sample-adaptive sparse 3D anchors dynamically. In addition, an Equal-Width (EW) loss is developed to leverage the parallel property of lanes for regularization. Furthermore, camera-LiDAR fusion is also explored based on Anchor3DLane++ to leverage complementary information. Extensive experiments on three popular 3D lane detection benchmarks show that our Anchor3DLane++ outperforms previous state-of-the-art methods. Code is available at: https://github.com/tusen-ai/Anchor3DLane. Shaofei Huang 0001, Zhenwei Shen, Zehao Huang, Yue Liao, Jizhong Han, Naiyan Wang, Si Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | Neural Networks With Linear Adaptive Batch Normalization and Swarm Intelligence Calibration for Real-Time Gaze Estimation on SmartphonesabstractEye tracking has emerged as a valuable tool for both research and clinical applications. However, traditional eye‐tracking systems are often bulky and expensive, limiting their widespread adoption in various fields. Smartphone eye tracking has become feasible with advanced deep learning and edge computing technologies. However, the field still faces practical challenges related to large‐scale datasets, model inference speed, and gaze estimation accuracy. The present study created a new dataset that contains over 3.2 million face images collected with recent phone models and presents a comprehensive smartphone eye‐tracking pipeline comprising a deep neural network framework (MGazeNet), a personalized model calibration method, and a heuristic gaze signal filter. The MGazeNet model introduced a linear adaptive batch normalization module to efficiently combine eye and face features, achieving the state‐of‐the‐art gaze estimation accuracy of 1.59 cm on the GazeCapture dataset and 1.48 cm on our custom dataset. In addition, an algorithm that utilizes multiverse optimization to optimize the hyperparameters of support vector regression (MVO–SVR) was proposed to improve eye‐tracking calibration accuracy with 13 or fewer ground‐truth gaze points, further improving gaze estimation accuracy to 0.89 cm. This integrated approach allows for eye tracking with accuracy comparable to that of research‐grade eye trackers, offering new application possibilities for smartphone eye tracking. Gancheng Zhu, Yongkai Li, Shuai Zhang 0027, Xiaoting Duan, Zehao Huang, Zhaomin Yao, Zhiguo Wang 0003 |
Int. J. Intell. Syst. | 5 |
| 2024 | Learnable Graph Matching: A Practical Paradigm for Data AssociationabstractData association is at the core of many computer vision tasks, e.g., multiple object tracking, image matching, and point cloud registration. however, current data association solutions have some defects: they mostly ignore the intra-view context information; besides, they either train deep association models in an end-to-end way and hardly utilize the advantage of optimization-based assignment methods, or only use an off-the-shelf neural network to extract features. In this paper, we propose a general learnable graph matching method to address these issues. Especially, we model the intra-view relationships as an undirected graph. Then data association turns into a general graph matching problem between graphs. Furthermore, to make optimization end-to-end differentiable, we relax the original graph matching problem into continuous quadratic programming and then incorporate training into a deep graph neural network with KKT conditions and implicit function theorem. In MOT task, our method achieves state-of-the-art performance on several MOT datasets. For image matching, our method outperforms state-of-the-art methods on a popular indoor dataset, ScanNet. For point cloud registration, we also achieve competitive results. Jiawei He 0002, Zehao Huang, Naiyan Wang, Zhaoxiang Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | Fully Sparse Fusion for 3D Object DetectionabstractCurrently prevalent multi-modal 3D detection methods rely on dense detectors that usually use dense Bird's-Eye-View (BEV) feature maps. However, the cost of such BEV feature maps is quadratic to the detection range, making it not scalable for long-range detection. Recently, LiDAR-only fully sparse architecture has been gaining attention for its high efficiency in long-range perception. In this paper, we study how to develop a multi-modal fully sparse detector. Specifically, our proposed detector integrates the well-studied 2D instance segmentation into the LiDAR side, which is parallel to the 3D instance segmentation part in the LiDAR-only baseline. The proposed instance-based fusion framework maintains full sparsity while overcoming the constraints associated with the LiDAR-only fully sparse detector. Our framework showcases state-of-the-art performance on the widely used nuScenes dataset, Waymo Open Dataset, and the long-range Argoverse 2 dataset. Notably, the inference speed of our proposed method under the long-range perception setting is 2.7× faster than that of other state-of-the-art multimodal 3D detection methods. Yingyan Li, Lue Fan, Yang Liu 0347, Zehao Huang, Yuntao Chen, Naiyan Wang, Zhaoxiang Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Anchor3DLane: Learning to Regress 3D Anchors for Monocular 3D Lane DetectionabstractMonocular 3D lane detection is a challenging task due to its lack of depth information. A popular solution is to first transform the front-viewed (FV) images or features into the bird-eye-view (BEV) space with inverse perspective mapping (IPM) and detect lanes from BEV features. However, the reliance of IPM on flat ground assumption and loss of context information make it inaccurate to restore 3D information from BEV representations. An attempt has been made to get rid of BEV and predict 3D lanes from FV representations directly, while it still underperforms other BEV-based methods given its lack of structured representation for 3D lanes. In this paper, we define 3D lane anchors in the 3D space and propose a BEV-free method named Anchor3DLane to predict 3D lanes directly from FV representations. 3D lane anchors are projected to the FV features to extract their features which contain both good structural and context information to make accurate predictions. In addition, we also develop a global optimization method that makes use of the equal-width property between lanes to reduce the lateral error of predictions. Extensive experiments on three popular 3D lane detection benchmarks show that our Anchor3DLane outperforms previous BEV-based methods and achieves state-of-the-art performances. The code is available at: https://github.com/tusenai/Anchor3DLane. Shaofei Huang 0001, Zhenwei Shen, Zehao Huang, Jiao Dai, Jizhong Han, Naiyan Wang, Si Liu 0001 |
CVPR | 3 |
| 2023 | Object as Query: Lifting any 2D Object Detector to 3D Detectionabstract3D object detection from multi-view images has drawn much attention over the past few years. Existing methods mainly establish 3D representations from multi-view images and adopt a dense detection head for object detection, or employ object queries distributed in 3D space to localize objects. In this paper, we design Multi-View 2D Objects guided 3D Object Detector (MV2D), which can lift any 2D object detector to multi-view 3D object detection. Since 2D detections can provide valuable priors for object existence, MV2D exploits 2D detectors to generate object queries conditioned on the rich image semantics. These dynamically generated queries help MV2D to recall objects in the field of view and show a strong capability of localizing 3D objects. For the generated queries, we design a sparse cross attention module to force them to focus on the features of specific objects, which suppresses interference from noises. The evaluation results on the nuScenes dataset demonstrate the dynamic object queries and sparse feature aggregation can promote 3D detection capability. MV2D also exhibits a state-of-the-art performance among existing methods. We hope MV2D can serve as a new baseline for future research. Code is available at https://github.com/tusen-ai/MV2D. Zitian Wang, Zehao Huang, Jiahui Fu 0003, Naiyan Wang, Si Liu 0001 |
ICCV | 2 |
| 2023 | Social distance control for quadruped robots in a gated spike filter neural network framework
Shuai Zhang 0027, Yongkai Li, Zehao Huang, Zhiguo Wang 0003 |
Appl. Intell. | 3 |
| 2023 | CCS-Net: Cascade Detection Network With the Convolution Kernel Switch Block and Statistics Optimal Anchors Block in Hypopharyngeal Cancer MRIabstractMagnetic resonance imaging (MRI) is a common diagnostic method for hypopharyngeal cancer (HPC). It is a challenge to automatically detect HPC tumors and swollen lymph nodes (HPC risk areas) from MRI slices because of the small size and irregular shape of HPC risk areas. Herein, we propose a cascade detection network with Convolution Kernel Switch (CKS) Block and Statistics Optimal Anchors (SOA) Block in HPC MRI (CCS-Net). CKS Block can adaptively switch standard convolution to deformable convolution in some appropriate layers to detect irregular objects more efficiently without taking up too much computing resources. SOA Block can automatically generate the optimal anchors based on the size distribution of objects. Compared with other methods, our method achieves splendid detection performance and outperforms other methods on the HPC dataset (more than 1800 T2 MRI slices), achieving the highest AP50of 78.90%. Experiments show that the proposed network can be the basis of a computer aided diagnosis utility that helps achieve faster and more accurate diagnostic decisions for HPC. Yang Miao 0002, Zehao Huang, Ning Pei, Changming An |
IEEE J. Biomed. Health Informatics | 6 |
| 2022 | QueryDet: Cascaded Sparse Query for Accelerating High-Resolution Small Object DetectionabstractWhile general object detection with deep learning has achieved great success in the past few years, the performance and efficiency of detecting small objects are far from satisfactory. The most common and effective way to promote small object detection is to use high-resolution images or feature maps. However, both approaches induce costly computation since the computational cost grows squarely as the size of images and features increases. To get the best of two worlds, we propose QueryDet that uses a novel query mechanism to accelerate the inference speed of feature-pyramid based object detectors. The pipeline composes two steps: it first predicts the coarse locations of small objects on low-resolution features and then computes the accurate detection results using high-resolution features sparsely guided by those coarse positions. In this way, we can not only harvest the benefit of high-resolution feature maps but also avoid useless computation for the background area. On the popular COCO dataset, the proposed method improves the detection mAP by 1.0 and mAP-small by 2.0, and the high-resolution inference speed is improved to 3.0× on average. On VisDrone dataset, which contains more small objects, we create a new state-of-the-art while gaining a 2.3× high-resolution acceleration on average. Code is available at https://github.com/ChenhongyiYang/QueryDet=PyTorch. Chenhongyi Yang, Zehao Huang, Naiyan Wang |
CVPR | 2 |
| 2022 | Step Into My Mind Palace: Exploration of a Collaborative Paragogy Tool in VRabstractVirtual Reality (VR) can mediate remote collaborative learning and can support pedagogical processes like paragogy. Within education, methods such as spaced repetition and memory palaces exist to support the cognitive process of remembering. We identify an opportunity to enhance learner-led collaborative paragogy involving these methods through immersive VR experiences. We present CleVR, a VR-mediated collaboration-based system that supports the memory palace and spaced repetition techniques. As an exploratory study, we aim to identify the applicability, viability and user perception for such a system combining these two techniques in VR. CleVR is a novel implementation which provides a location-driven metaphor to populate and present multiple resources related to a topic for peer-led exploration. We discuss the design and provide a prototype implementation of CleVR. We conducted two studies, a targeted expert user review and a broader proof of concept survey. The results of the studies show interesting outcomes, with the system described as ‘engaging’, ‘useful’ and ‘fun’. Our findings provide insights to the potential of using Virtual Reality Learning Environments (VRLE) geared towards collaborative learner-led activities. Robert Sims, Barry Chang, Verity Bennett, Advaith Krishnan, Abdalslam Aboubakar, George Coman, Abdulrazak Bahrami, Zehao Huang, Christopher Clarke, Abhijit Karnik |
iLRN | 8 |
| 2021 | Learnable Graph Matching: Incorporating Graph Partitioning With Deep Feature Learning for Multiple Object TrackingabstractData association across frames is at the core of Multiple Object Tracking (MOT) task. This problem is usually solved by a traditional graph-based optimization or directly learned via deep learning. Despite their popularity, we find some points worth studying in current paradigm: 1) Existing methods mostly ignore the context information among tracklets and intra-frame detections, which makes the tracker hard to survive in challenging cases like severe occlusion. 2) The end-to-end association methods solely rely on the data fitting power of deep neural networks, while they hardly utilize the advantage of optimization-based assignment methods. 3) The graph-based optimization methods mostly utilize a separate neural network to extract features, which brings the inconsistency between training and inference. Therefore, in this paper we propose a novel learnable graph matching method to address these issues. Briefly speaking, we model the relationships between tracklets and the intra-frame detections as a general undirected graph. Then the association problem turns into a general graph matching between tracklet graph and detection graph. Furthermore, to make the optimization end-to-end differentiable, we relax the original graph matching into continuous quadratic programming and then incorporate the training of it into a deep graph network with the help of the implicit function theorem. Lastly, our method GMTracker, achieves state-of-the-art performance on several standard MOT datasets. Our code is available at https://github.com/jiaweihe1996/GMTracker. Jiawei He 0002, Zehao Huang, Naiyan Wang, Zhaoxiang Zhang 0001 |
CVPR | 2 |
| 2021 | Direct Differentiable Augmentation SearchabstractData augmentation has been an indispensable tool to improve the performance of deep neural networks, however the augmentation can hardly transfer among different tasks and datasets. Consequently, a recent trend is to adopt AutoML technique to learn proper augmentation policy without extensive hand-crafted tuning. In this paper, we propose an efficient differentiable search algorithm called Direct Differentiable Augmentation Search (DDAS). It exploits meta-learning with one-step gradient update and continuous relaxation to the expected training loss for efficient search. Our DDAS can achieve efficient augmentation search without relying on approximations such as Gumbel-Softmax or second order gradient approximation. To further reduce the adverse effect of improper augmentations, we organize the search space into a two level hierarchy, in which we first decide whether to apply augmentation, and then determine the specific augmentation policy. On standard image classification benchmarks, our DDAS achieves state-of-the-art performance and efficiency tradeoff while reducing the search cost dramatically, e.g. 0.15 GPU hours for CIFAR-10. In addition, we also use DDAS to search augmentation for object detection task and achieve comparable performance with AutoAugment [8], while being 1000× faster. Code will be released in https://github.com/zxcvfd13502/DDAS_code Aoming Liu, Zehao Huang, Zhiwu Huang, Naiyan Wang |
ICCV | 2 |
| 2021 | You Only Search Once: Single Shot Neural Architecture Search via Direct Sparse OptimizationabstractRecently neural architecture search (NAS) has raised great interest in both academia and industry. However, it remains challenging because of its huge and non-continuous search space. Instead of applying evolutionary algorithm or reinforcement learning as previous works, this paper proposes a direct sparse optimization NAS (DSO-NAS) method. The motivation behind DSO-NAS is to address the task in the view of model pruning. To achieve this goal, we start from a completely connected block, and then introduce scaling factors to scale the information flow between operations. Next, sparse regularizations are imposed to prune useless connections in the architecture. Lastly, an efficient and theoretically sound optimization method is derived to solve it. Our method enjoys both advantages of differentiability and efficiency, therefore it can be directly applied to large datasets like ImageNet and tasks beyond classification. Particularly, on the CIFAR-10 dataset, DSO-NAS achieves an average test error 2.74 percent, while on the ImageNet dataset DSO-NAS achieves 25.4 percent test error under 600M FLOPs with 8 GPUs in 18 hours. As for semantic segmentation task, DSO-NAS also achieve competitive result compared with manually designed architectures on the PASCAL VOC dataset. Code is available at https://github.com/XinbangZhang/DSO-NAS. Xinbang Zhang, Zehao Huang, Naiyan Wang, Shiming Xiang, Chunhong Pan |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2019 | SimpleDet: A Simple and Versatile Distributed Framework for Object Detection and Instance RecognitionabstractObject detection and instance recognition play a central role in many AI applications like autonomous driving, video surveillance and medical image analysis. However, training object detection models on large scale datasets remains computationally expensive and time consuming. This paper presents an efficient and open source object detection framework called SimpleDet which enables the training of state-of-the-art detection models on consumer grade hardware at large scale. SimpleDet covers a wide range of models including both high-performance and high-speed ones. SimpleDet is well-optimized for both low precision training and distributed training and achieves 70% higher throughput for the Mask R-CNN detector compared with existing frameworks. Codes, examples and documents of SimpleDet can be found at https://github.com/tusimple/simpledet. Yuntao Chen, Chenxia Han, Yanghao Li, Zehao Huang, Naiyan Wang, Zhaoxiang Zhang 0001 |
J. Mach. Learn. Res. | 4 |
| 2018 | Data-Driven Sparse Structure Selection for Deep Neural Networks
Zehao Huang, Naiyan Wang |
ECCV (16) | 1 |
| 2017 | GPS-Simulated Trajectory Detection
Han Su 0001, Wei Chen 0070, Min Nie, Bolong Zheng, Zehao Huang, Defu Lian |
DASFAA (2) | 6 |
| 2017 | Image super-resolution via deep dilated convolutional networksabstractDeep learning techniques have been successfully applied in single image super-resolution (SR). Recently, researches have shown that increasing the depth of network can significantly improve SR performance. Very deep networks for SR achieved a large improvement than former methods. However, simply increasing depths basically introduce more parameters and this lead to cumbersome computational cost. In this paper, we present a general and effective method to accelerate very deep networks for single image SR. Our method is based on dilated convolution operation, which support exponential expansion of the receptive field without increasing filter size. With the help of dilated convolution, shallow networks can achieve large receptive field and exploit contextual information in an efficient way. Based on a very deep network, we propose a 12 layers dilated convolutional network for SR (DCNSR). While accelerating 2x speed, our shallow network achieves better performance than original deep networks and shows state-of-the-art reconstructed results. Zehao Huang, Lingfeng Wang 0002, Gaofeng Meng, Chunhong Pan |
ICIP | 1 |
| 2017 | Ensemble based deep networks for image super-resolution
Lingfeng Wang 0002, Zehao Huang, Yongchao Gong, Chunhong Pan |
Pattern Recognit. | 2 |