VLDB 2026 Research / reviewers in the wild / expert
Hu Zhang 0005
dblp:69/5169-5
· DBLP profile ↗
13ranked-venue papers
4as first author
10since 2021 · last 2025
0009-0009-9892-9515ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 4 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Blind Bitstream-corrupted Video Recovery via Metadata-guided Diffusion ModelabstractBitstream-corrupted video recovery aims to fill in realistic video content due to bitstream corruption during video storage or transmission. Most existing methods typically assume that the predefined masks of the corrupted regions are known in advance. However, manually annotating these masks is laborious and time-consuming, limiting the applicability of existing methods in real-world scenarios. Therefore, we expect to relax this assumption by defining a new blind video recovery setting where the recovery of corrupted regions does not rely on predefined masks. There are two significant challenges in this setting: (i) without predefined masks, how accurately can a model identify the regions requiring recovery? (ii) how to recover contents from extensive and irregular regions, especially when large portions of frames are severely degraded? To address these challenges, we introduce a Metadata-Guided Diffusion Model, dubbed M-GDM. To enable a diffusion model focusing on the corrupted regions, we leverage intrinsic video metadata as a corruption indicator and design a dual-stream metadata encoder. This encoder first embeds the motion vectors and frame types of a video separately and then merges them into a unified metadata representation. The metadata representation will interact with the corrupted latent feature through cross-attention mechanisms at each diffusion step. Meanwhile, to preserve the intact regions, we propose a prior-driven mask predictor that generates pseudo masks for the corrupted regions by leveraging the metadata prior and diffusion prior. These pseudo masks enable the separation and recombination of intact and recovered regions through hard masking. However, imperfections in pseudo mask predictions and hard masking processes often result in boundary artifacts. Thus, we introduce a post-refinement module that refines the hard-masked outputs, enhancing the consistency between intact and recovered regions. Extensive experiment results validate the effectiveness of our method and demonstrate its superiority in the blind video recovery task. Hu Zhang 0005, Dadong Wang, Xin Yu 0002 |
CVPR | 2 |
| 2025 | Harnessing Uncertainty-Aware Bounding Boxes for Unsupervised 3D Object DetectionabstractUnsupervised 3D object detection aims to identify objects of interest from unlabeled raw data, such as LiDAR points. Recent approaches usually adopt pseudo 3D bounding boxes (3D bboxes) from clustering algorithm to initialize the model training. However, pseudo bboxes inevitably contain noise, and such inaccuracies accumulate to the final model, compromising the performance. Therefore, in an attempt to mitigate the negative impact of inaccurate pseudo bboxes, we introduce a new uncertainty-aware framework for unsupervised 3D object detection, dubbed UA3D. In particular, our method consists of two phases: uncertainty estimation and uncertainty regularization. (1) In the uncertainty estimation phase, we incorporate an extra auxiliary detection branch alongside the original primary detector. The prediction disparity between the primary and auxiliary detectors could reflect fine-grained uncertainty at the box coordinate level. (2) Based on the assessed uncertainty, we adaptively adjust the weight of every 3D bbox coordinate via uncertainty regularization, refining the training process on pseudo bboxes. For pseudo bbox coordinate with high uncertainty, we assign a relatively low loss weight. Extensive experiments verify that the proposed method is robust against the noisy pseudo bboxes, yielding substantial improvements on nuScenes and Lyft compared to existing approaches, with increases of +6.9% AP$_{BEV}$ and +2.5% AP$_{3D}$ on nuScenes, and +4.1% AP$_{BEV}$ and +2.0% AP$_{3D}$ on Lyft. Ruiyang Zhang, Hu Zhang 0005, Zhedong Zheng |
ICCV | 2 |
| 2025 | Queryable 3D Scene Representation: A Multi-Modal Framework for Semantic Reasoning and Robotic Task PlanningabstractTo enable robots to comprehend high-level human instructions and perform complex tasks, a key challenge lies in achieving comprehensive scene understanding: interpreting and interacting with the 3D environment in a meaningful way. This requires a smart map that fuses accurate geometric structure with rich, human-understandable semantics. To address this, we introduce the 3D Queryable Scene Representation (3D QSR), a novel framework built on multimedia data that unifies three complementary 3D representations: (1) 3D-consistent novel view rendering and segmentation from panoptic reconstruction, (2) precise geometry from 3D point clouds, and (3) structured, scalable organization via 3D scene graphs. Built on an object-centric design, the framework integrates with large vision-language models to enable semantic queryability by linking multimodal object embeddings, and supporting object-level retrieval of geometric, visual, and semantic information. The retrieved data are then loaded into a robotic task planner for downstream execution. Xun Li 0004, Rodrigo Santa Cruz, Mingze Xi, Hu Zhang 0005, Madhawa Perera, Ziwei Wang 0003, Ahalya Ravendran, Brandon J. Matthews, Matt Adcock, Dadong Wang, Jiajun Liu 0004 |
ACM Multimedia | 4 |
| 2024 | OpenSight: A Simple Open-Vocabulary Framework for LiDAR-Based Object Detection
Hu Zhang 0005, Xin Yu 0002, Zi Huang, Kaicheng Yu |
ECCV (84) | 1 |
| 2024 | Approaching Outside: Scaling Unsupervised 3D Object Detection from 2D Scene
Ruiyang Zhang, Hu Zhang 0005, Hang Yu 0006, Zhedong Zheng |
ECCV (11) | 2 |
| 2024 | M3A: A multimodal misinformation dataset for media authenticity analysisabstractWith the development of various generative models, misinformation in news media becomes more deceptive and easier to create, posing a significant problem. However, existing datasets for misinformation study often have limited modalities, constrained sources, and a narrow range of topics. These limitations make it difficult to train models that can effectively combat real-world misinformation. To address this, we propose a comprehensive, large-scale Multimodal Misinformation dataset for Media Authenticity Analysis ( M 3 A ), featuring broad sources and fine-grained annotations for topics and sentiments. To curate M 3 A , we collect genuine news content from 60 renowned news outlets worldwide and generate fake samples using multiple techniques. These include altering named entities in texts, swapping modalities between samples, creating new modalities, and misrepresenting movie content as news. M 3 A contains 708K genuine news samples and over 6M fake news samples, spanning text, images, audio, and video. M 3 A provides detailed multi-class labels, crucial for various misinformation detection tasks, including out-of-context detection and deepfake detection. For each task, we offer extensive benchmarks using state-of-the-art models, aiming to enhance the development of robust misinformation detection systems. • We present M 3 A , a large-scale multimodal misinformation dataset with diverse news samples. • M 3 A includes texts, images, audio, and videos from multiple reputable news outlets. • M 3 A addresses limitations in existing datasets in misinformation generation methods and scale. • We provide multi-class annotations in M 3 A for various key tasks in misinformation detection. • We propose benchmarks for M 3 A using state-of-the-art models and out-of-distribution testing. Qingzheng Xu, Huiqiang Chen, Heming Du, Hu Zhang 0005, Szymon Lukasik, Tianqing Zhu, Xin Yu 0002 |
Comput. Vis. Image Underst. | 4 |
| 2024 | BAVS: Bootstrapping Audio-Visual Segmentation by Integrating Foundation KnowledgeabstractGiven an audio-visual pair, audio-visual segmentation (AVS) aims to locate sounding sources by predicting pixel-wise maps. Previous methods assume that each sound component in an audio signal always has a visual counterpart in the image. However, this assumption overlooks that off-screen sounds and background noise often contaminate the audio recordings in real-world scenarios. They impose significant challenges on building a consistent semantic mapping between audio and visual signals for AVS models and thus impede precise sound localization. In this work, we propose a two-stage bootstrapping audio-visual segmentation framework by incorporating multi-modal foundation knowledge$^{1}$In a nutshell, our BAVS is designed to eliminate the interference of background noise or off-screen sounds in segmentation by establishing the audio-visual correspondences in an explicit manner. In the first stage, we employ a segmentation model to localize potential sounding objects from visual data without being affected by contaminated audio signals. Meanwhile, we also utilize a foundation audio classification model to discern audio semantics. Considering the audio tags provided by the audio foundation model are noisy, associating object masks with audio tags is not trivial. Thus, in the second stage, we develop an audio-visual semantic integration strategy (AVIS) to localize the authentic-sounding objects. Here, we construct an audio-visual tree based on the hierarchical correspondence between sounds and object categories. We then examine the label concurrency between the localized objects and classified audio tags by tracing the audio-visual tree. With AVIS, we can effectively segment real-sounding objects. Extensive experiments demonstrate the superiority of our method on AVS datasets, particularly in scenarios involving background noise. Our project website ishttps://yenanliu.github.io/AVSS.github.io/. Chen Liu 0028, Peike Li, Hu Zhang 0005, Lincheng Li, Zi Huang, Dadong Wang, Xin Yu 0002 |
IEEE Trans. Multim. | 3 |
| 2024 | Divide and Retain: A Dual-Phase Modeling for Long-Tailed Visual RecognitionabstractThis work explores visual recognition models on real-world datasets exhibiting a long-tailed distribution. Most of previous works are based on a holistic perspective that the overall gradient for training model is directly obtained by considering all classes jointly. However, due to the extreme data imbalance in long-tailed datasets, joint consideration of different classes tends to induce the gradient distortion problem; i.e., the overall gradient tends to suffer from shifted direction toward data-rich classes and enlarged variances caused by data-poor classes. The gradient distortion problem impairs the training of our models. To avoid such drawbacks, we propose to disentangle the overall gradient and aim to consider the gradient on data-rich classes and that on data-poor classes separately. We tackle the long-tailed visual recognition problem via a dual-phase-based method. In the first phase, only data-rich classes are concerned to update model parameters, where only separated gradient on data-rich classes is used. In the second phase, the rest data-poor classes are involved to learn a complete classifier for all classes. More importantly, to ensure the smooth transition from phase I to phase II, we propose an exemplar bank and a memory-retentive loss. In general, the exemplar bank reserves a few representative examples from data-rich classes. It is used to maintain the information of data-rich classes when transiting. The memory-retentive loss constrains the change of model parameters from phase I to phase II based on the exemplar bank and data-poor classes. The extensive experimental results on four commonly used long-tailed benchmarks, including CIFAR100-LT, Places-LT, ImageNet-LT, and iNaturalist 2018, highlight the excellent performance of our proposed method. Hu Zhang 0005, Linchao Zhu, Yi Yang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | Audio-Visual Segmentation by Exploring Cross-Modal Mutual SemanticsabstractThe audio-visual segmentation (AVS) task aims to segment sounding objects from a given video. Existing works mainly focus on fusing audio and visual features of a given video to achieve sounding object masks. However, we observed that prior arts are prone to segment a certain salient object in a video regardless of the audio information. This is because sounding objects are often the most salient ones in the AVS dataset. Thus, current AVS methods might fail to localize genuine sounding objects due to the dataset bias. In this work, we present an audio-visual instance-aware segmentation approach to overcome the dataset bias. In a nutshell, our method first localizes potential sounding objects in a video by an object segmentation network, and then associates the sounding object candidates with the given audio. We notice that an object could be a sounding object in one video but a silent one in another video. This would bring ambiguity in training our object segmentation network as only sounding objects have corresponding segmentation masks. We thus propose a silent object-aware segmentation objective to alleviate the ambiguity. Moreover, since the category information of audio is unknown, especially for multiple sounding sources, we propose to explore the audio-visual semantic correlation and then associate audio with potential objects. Specifically, we attend predicted audio category scores to potential instance masks and these scores will highlight corresponding sounding instances while suppressing inaudible ones. When we enforce the attended instance masks to resemble the ground-truth mask, we are able to establish audio-visual semantics correlation. Experimental results on the AVS benchmarks demonstrate that our method can effectively segment sounding objects without being biased to salient objects and also achieves state-of-the-art performance in both the single-source and multi-source scenarios. Chen Liu 0028, Peike Li, Xingqun Qi, Hu Zhang 0005, Lincheng Li, Dadong Wang, Xin Yu 0002 |
ACM Multimedia | 4 |
| 2023 | RVD: A Handheld Device-Based Fundus Video Dataset for Retinal Vessel SegmentationabstractRetinal vessel segmentation is generally grounded in image-based datasets collected with bench-top devices. The static images naturally lose the dynamic characteristics of retina fluctuation, resulting in diminished dataset richness, and the usage of bench-top devices further restricts dataset scalability due to its limited accessibility. Considering these limitations, we introduce the first video-based retinal dataset by employing handheld devices for data acquisition. The dataset comprises 635 smartphone-based fundus videos collected from four different clinics, involving 415 patients from 50 to 75 years old. It delivers comprehensive and precise annotations of retinal structures in both spatial and temporal dimensions, aiming to advance the landscape of vasculature segmentation. Specifically, the dataset provides three levels of spatial annotations: binary vessel masks for overall retinal structure delineation, general vein-artery masks for distinguishing the vein and artery, and fine-grained vein-artery masks for further characterizing the granularities of each artery and vein. In addition, the dataset offers temporal annotations that capture the vessel pulsation characteristics, assisting in detecting ocular diseases that require fine-grained recognition of hemodynamic fluctuation. In application, our dataset exhibits a significant domain shift with respect to data captured by bench-top devices, thus posing great challenges to existing methods. Thanks to rich annotations and data scales, our dataset potentially paves the path for more advanced retinal analysis and accurate disease diagnosis. In the experiments, we provide evaluation metrics and benchmark results on our dataset, reflecting both the potential and challenges it offers for vessel segmentation tasks. We hope this challenging dataset would significantly contribute to the development of eye disease diagnosis and early prevention. Wahiduzzaman Khan, Hongwei Sheng, Hu Zhang 0005, Heming Du, Sen Wang 0001, Minas Theodore Coroneo, Farshid Hajati, Sahar Shariflou, Michael Kalloniatis, Jack Phu, Ashish Agar, Zi Huang, S. Mojtaba Golzan, Xin Yu 0002 |
NeurIPS | 3 |
| 2020 | Motion-Excited Sampler: Video Adversarial Attack with Sparked Prior
Hu Zhang 0005, Linchao Zhu, Yi Zhu 0004, Yi Yang 0001 |
ECCV (20) | 1 |
| 2020 | Query-efficient Meta Attack to Deep Neural Networks
Jiawei Du 0002, Hu Zhang 0005, Joey Tianyi Zhou, Yi Yang 0001, Jiashi Feng |
ICLR | 2 |
| 2019 | Generalized Majorization-Minimization for Non-Convex OptimizationabstractMajorization-Minimization (MM) algorithms optimize an objective function by iteratively minimizing its majorizing surrogate and offer attractively fast convergence rate for convex problems. However, their convergence behaviors for non-convex problems remain unclear. In this paper, we propose a novel MM surrogate function from strictly upper bounding the objective to bounding the objective in expectation. With this generalized surrogate conception, we develop a new optimization algorithm, termed SPI-MM, that leverages the recent proposed SPIDER for more efficient non-convex optimization. We prove that for finite-sum problems, the SPI-MM algorithm converges to an stationary point within deterministic and lower stochastic gradient complexity. To our best knowledge, this work gives the first non-asymptotic convergence analysis for MM-alike algorithms in general non-convex optimization. Extensive empirical studies on non-convex logistic regression and sparse PCA demonstrate the advantageous efficiency of the proposed algorithm and validate our theoretical results. Hu Zhang 0005, Pan Zhou 0002, Yi Yang 0001, Jiashi Feng |
IJCAI | 1 |