EDBT 2026 Demo / reviewers in the wild / expert
Yuzhi Huang
dblp:95/8495
· DBLP profile ↗
8ranked-venue papers
2as first author
8since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Systems, architecture and hardware · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Segmentation and scene understanding · 38% 3D vision · 32% Generative modeling · 14% |
Topics — the 15 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › 3D vision
camera pose estimation |
0.9 | 1 | 2025 | DynamicVerse: A Physically-Aware Multimodal Framework for 4D World Modeling · NeurIPS 2025 |
Computer vision › 3D vision
depth estimation |
0.9 | 1 | 2025 | DynamicVerse: A Physically-Aware Multimodal Framework for 4D World Modeling · NeurIPS 2025 |
Computer vision › 3D vision › 3d scene reconstruction
dynamic scene reconstruction |
0.9 | 1 | 2025 | DynamicVerse: A Physically-Aware Multimodal Framework for 4D World Modeling · NeurIPS 2025 |
Computer vision › Video understanding and tracking
video anomaly detection |
0.9 | 1 | 2025 | Track Any Anomalous Object: A Granular Video Anomaly Detection Pipeline · CVPR 2025 |
Computer vision › 3D vision › depth estimation
video depth estimation |
0.9 | 1 | 2025 | DynamicVerse: A Physically-Aware Multimodal Framework for 4D World Modeling · NeurIPS 2025 |
Computer vision › Segmentation and scene understanding › medical image segmentation
ambiguous medical image segmentation |
0.8 | 1 | 2024 | P2SAM: Probabilistically Prompted SAMs Are Efficient Segmentator for Ambiguous Medical Images · ACM Multimedia 2024 |
Machine learning › Generative modeling › variational autoencoder
conditional variational autoencoder |
0.8 | 1 | 2024 | Flaws can be Applause: Unleashing Potential of Segmenting Ambiguous Objects in SAM · NeurIPS 2024 |
Computer vision › Segmentation and scene understanding
image segmentation |
0.8 | 1 | 2024 | Flaws can be Applause: Unleashing Potential of Segmenting Ambiguous Objects in SAM · NeurIPS 2024 |
Computer vision › Segmentation and scene understanding
medical image segmentation |
0.8 | 1 | 2024 | P2SAM: Probabilistically Prompted SAMs Are Efficient Segmentator for Ambiguous Medical Images · ACM Multimedia 2024 |
Computer vision › Segmentation and scene understanding › image segmentation
probabilistic segmentation |
0.8 | 1 | 2024 | P2SAM: Probabilistically Prompted SAMs Are Efficient Segmentator for Ambiguous Medical Images · ACM Multimedia 2024 |
Machine learning › Generative modeling
variational autoencoder |
0.8 | 1 | 2024 | Flaws can be Applause: Unleashing Potential of Segmenting Ambiguous Objects in SAM · NeurIPS 2024 |
Computer vision › Vision and language
multimodal understanding |
0.3 | 1 | 2025 | DynamicVerse: A Physically-Aware Multimodal Framework for 4D World Modeling · NeurIPS 2025 |
Computer vision › Video understanding and tracking
multi-object tracking |
0.3 | 1 | 2025 | Track Any Anomalous Object: A Granular Video Anomaly Detection Pipeline · CVPR 2025 |
Computer vision › Segmentation and scene understanding
foundation model segmentation |
0.2 | 1 | 2024 | Flaws can be Applause: Unleashing Potential of Segmenting Ambiguous Objects in SAM · NeurIPS 2024 |
Computer vision › Vision and language › vision-language model › prompt learning
prompt-based adaptation |
0.2 | 1 | 2024 | P2SAM: Probabilistically Prompted SAMs Are Efficient Segmentator for Ambiguous Medical Images · ACM Multimedia 2024 |
Methods — techniques the papers use, named apart from their topics
structure from motion · 0.9pixel-level tracking · 0.9multimodal large models · 0.9bundle adjustment · 0.9anomaly scoring · 0.9segment anything model · 0.8probabilistic prompt space · 0.8latent probability distribution · 0.8conditional variational autoencoder · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Track Any Anomalous Object: A Granular Video Anomaly Detection PipelineabstractVideo anomaly detection (VAD) is crucial in scenarios such as surveillance and autonomous driving, where timely detection of unexpected activities is essential. Albeit existing methods have primarily focused on detecting anomalous objects in videos—either by identifying anomalous frames or objects—they often neglect finer-grained analysis, such as anomalous pixels, which limits their ability to capture a broader range of anomalies. To address this challenge, we propose an innovative VAD framework called Track Any Anomalous Object (TAO), which introduces a Granular Video Anomaly Detection Framework that, for the first time, integrates the detection of multiple fine-grained anomalous objects into a unified framework. Unlike methods that assign anomaly scores to every pixel at each moment, our approach transforms the problem into pixel-level tracking of anomalous objects. By linking anomaly scores to subsequent tasks such as image segmentation and video tracking, our method eliminates the need for threshold selection and achieves more precise anomaly localization, even in long and challenging video sequences. Experiments on extensive datasets demonstrate that TAO achieves state-of-the-art performance, setting a new progress for VAD by providing a practical, granular, and holistic solution. For more information, visit the project page at: https://tao-25.github.io/ Yuzhi Huang, Chenxin Li, Zixu Lin, Yunlong Lin, Hengyu Liu 0007, Wuyang Li, Xinyu Liu 0001, Jiechao Gao, Yue Huang 0001, Xinghao Ding, Yixuan Yuan |
CVPR | 1 |
| 2025 | An EfficientNetV2 Deep Learning Framework for Adolescent Bone-age Prediction Using Hand RadiographsabstractGiven the increasing demand for assessment of adolescent growth and development, prediction of adolescent hand-bone age has become important in the field of medical image analysis. Recent convolutional neural networks(CNNs), especially the EfficientNetV2 model, have significantly improved the accuracy of bone-age prediction. In this study, we trained a CNN model based on EfficientNetV2 to predict bone age based on 10,000 X-ray images of adolescents hand bones. The model extracted deep features from X-ray images, and after training, predicted bone age with remarkable accuracy. Moreover, when tested on real dataset, the model reduced the mean absolute error (MAE) of predicted bone age to 0.752 years. Our study confirms that deep learning methods aid in medical image analysis. Our bone-age prediction model is both objective and quantitative, and will find applications in clinical practice and when inferring the developmental cycle of minors. TakMan Lo, Lixin Deng, Yizhu Tang, Kaip Tse, Jiatao Wu, Yuzhi Huang, Renzhi Lu |
INDIN | 6 |
| 2025 | Bone Age Assessment Using EfficientNet V2 and Multi-Model FusionabstractBone age assessment is a vital indicator for evaluating the growth and development of children and adolescents. Traditional manual evaluation methods are inefficient and highly subjective, failing to meet modern clinical demands. Recent advancements in deep learning, especially in image recognition and medical image analysis, have provided new opportunities for automated and precise bone age assessment. This paper presents a comprehensive study on deep learning-based bone age prediction, covering data preprocessing, model architecture, training strategies, and evaluation methods. The proposed method, leveraging data augmentation, U-Net segmentation, and the EfficientNet V2 architecture, achieves high accuracy and robustness in bone age prediction, demonstrating significant potential for clinical applications and providing valuable insights for future research. Lixin Deng, Yizhu Tang, Kaip Tse, Jiatao Wu, Yuzhi Huang, TakMan Lo, Renzhi Lu |
INDIN | 6 |
| 2025 | Bone Age Prediction using a Convolutional Neural Network-based Regression Algorithm employing Attention-Directing and ClusterabstractBone age assessment (BAA) is a critical research topic in pediatric radiology, with growing interest in developing automated BAA methods. This study proposes a bone age prediction model integrating cluster analysis and convolutional neural network (CNN) regression, further enhanced by a multi-scale attention mechanism to construct a "divide-and-focus" dual-driven deep learning framework. Targeting age-sensitive regional features in hand radiographs, we innovatively design an adaptive spatial attention module that achieves hierarchical anatomical feature enhancement through saliency detection of attention-guided regions of interest (ROI). The algorithm first uses multiconstrained clustering of K methods to generate age-specific subsets, followed by parallel execution on each subset: 1) attention-guided ROI segmentation and feature enhancement; 2) validation of the base CNN regression networks (including ResNet, DenseNet and EfficientNetV2); 3) set of cross-subset models with Bayesian-optimized weighting strategies for final prediction. By synergistically integrating the data distribution priors with attention-driven anatomical priors, the method delivers interpretable solutions when performing medical image regression tasks. The modular design ensures compatibility with mainstream CNN architectures. The method will aid pediatric growth monitoring and the diagnosis of endocrine disorders. Tinghong Ye, Lixin Deng, Yizhu Tang, TakMan Lo, Kaip Tse, Jiatao Wu, Yuzhi Huang, Renzhi Lu |
INDIN | 8 |
| 2025 | DynamicVerse: A Physically-Aware Multimodal Framework for 4D World ModelingabstractUnderstanding the dynamic physical world, characterized by its evolving 3D structure, real-world motion, and semantic content with textual descriptions, is crucial for human-agent interaction and enables embodied agents to perceive and act within real environments with human‑like capabilities. However, existing datasets are often derived from limited simulators or utilize traditional Structure-from-Motion for up-to-scale annotation and offer limited descriptive captioning, which restricts the capacity of foundation models to accurately interpret real-world dynamics from monocular videos, commonly sourced from the internet. To bridge these gaps, we introduce **DynamicVerse**, a physical‑scale, multimodal 4D world modeling framework for dynamic real-world video. We employ large vision, geometric, and multimodal models to interpret metric-scale static geometry, real-world dynamic motion, instance-level masks, and holistic descriptive captions. By integrating window-based Bundle Adjustment with global optimization, our method converts long real-world video sequences into a comprehensive 4D multimodal format. DynamicVerse delivers a large-scale dataset consists of 100K+ videos with 800K+ annotated masks and 10M+ frames from internet videos. Experimental evaluations on three benchmark tasks, namely video depth estimation, camera pose estimation, and camera intrinsics estimation, demonstrate that our 4D modeling achieves superior performance in capturing physical-scale measurements with greater global accuracy than existing methods. Kairun Wen, Yuzhi Huang, Runyu Chen, Hui Zheng 0003, Yunlong Lin, Panwang Pan, Chenxin Li, Wenyan Cong, Junbin Lu, Chenguo Lin, Dilin Wang, Zhicheng Yan 0001, Hongyu Xu, Justin Theiss, Yue Huang 0001, Xinghao Ding, Zhiwen Fan |
NeurIPS | 2 |
| 2024 | P2SAM: Probabilistically Prompted SAMs Are Efficient Segmentator for Ambiguous Medical ImagesabstractGenerating diverse plausible outputs from a single input is crucial for addressing visual ambiguities, exemplified in medical imaging where experts may provide varying semantic segmentation annotations for the same image.Existing methods handles ambiguous segmentation relying on probabilistic modeling and extensive multi-output annotated data while often struggles with limited ambiguously labeled datasets common in real-world applications.To surmount the challenge, we propose P²SAM, a novel framework that leverages the Segment Anything Model (SAM)'s prior knowledge for ambiguous object segmentation. By transforming SAM's sensitivity to prompts into an advantage, we introduce a prior probabilistic space for prompts.Experimental results show that P²SAM significantly enhances medical segmentation precision and diversity using minimal ambiguously annotated samples. Benchmarking against state-of-the-art methods demonstrates superior performance with just 5.5% of the training data (+12% Dmax). This approach marks a significant advancement towards deploying probabilistic models in data-limited real-world scenarios. Yuzhi Huang, Chenxin Li, Zixu Lin, Hengyu Liu 0007, Haote Xu, Yifan Liu 0010, Yue Huang 0001, Xinghao Ding, Xiaotong Tu, Yixuan Yuan |
ACM Multimedia | 1 |
| 2024 | Flaws can be Applause: Unleashing Potential of Segmenting Ambiguous Objects in SAMabstractAs the vision foundation models like the Segment Anything Model (SAM) demonstrate potent universality, they also present challenges in giving ambiguous and uncertain predictions. Significant variations in the model output and granularity can occur with simply subtle changes in the prompt, contradicting the consensus requirement for the robustness of a model. While some established works have been dedicated to stabilizing and fortifying the prediction of SAM, this paper takes a unique path to explore how this flaw can be inverted into an advantage when modeling inherently ambiguous data distributions. We introduce an optimization framework based on a conditional variational autoencoder, which jointly models the prompt and the granularity of the object with a latent probability distribution. This approach enables the model to adaptively perceive and represent the real ambiguous label distribution, taming SAM to produce a series of diverse, convincing, and reasonable segmentation outputs controllably. Extensive experiments on several practical deployment scenarios involving ambiguity demonstrates the exceptional performance of our framework. Project page: \url{https://a-sa-m.github.io/}. Chenxin Li, Yuzhi Huang, Wuyang Li, Hengyu Liu 0007, Xinyu Liu 0001, Qing Xu 0014, Zhen Chen 0013, Yue Huang 0001, Yixuan Yuan |
NeurIPS | 2 |
| 2022 | Graph-based Weakly Supervised Framework for Semantic Relevance Learning in E-commerceabstractProduct searching is fundamental in online e-commerce systems, it needs to quickly and accurately find the products that users required. Relevance is essential for e-commerce search, which role is avoiding displaying products that do not match search intent and optimizing user experience. Measuring semantic relevance is necessary because distributional biases between search queries and product titles may lead to large lexical differences between relevant textual expressions. Several problems limit the performance of semantic relevance learning, including extremely long-tail product distribution and low-quality labeled data. Recent works attempt to conduct relevance learning through user behaviors. However, noisy user behavior can easily cause inadequately semantic modeling. Therefore, it is valuable but challenging to utilize user behavior in relevance learning. In this paper, we first propose a weakly supervised contrastive learning framework that focuses on how to provide effective semantic supervision and generate reasonable representation. We utilize topology structure information contained in a user behavior heterogeneous graph to design a semantically aware data construction strategy. Besides, we propose a contrastive learning framework suitable for e-commerce scenarios with targeted improvements in data augmentation and training objectives. For relevance calculation, we propose a novel hybrid method that combines fine-tuning and transfer learning. It eliminates the negative impacts caused by distributional bias and guarantees semantic matching capabilities. Extensive experiments and analyses show the promising performance of proposed methods in relevance learning. Yuzhi Huang, Tianshu Wu, Hongbo Deng, Jian Xu 0015, Bo Zheng 0007 |
CIKM | 2 |