Haiwen Huang

dblp:220/3988 · DBLP profile ↗
← Back
14ranked-venue papers
5as first author
13since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 4 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Sel3DCraft: Interactive Visual Prompts for User-Friendly Text-to-3D Generation
abstract
Text-to-3D (T23D) generation has transformed digital content creation, yet remains bottlenecked by blind trial-and-error prompting processes that yield unpredictable results. While visual prompt engineering has advanced in text-to-image domains, its application to 3D generation presents unique challenges requiring multi-view consistency evaluation and spatial understanding. We present Sel3DCraft, a visual prompt engineering system for T23D that transforms unstructured exploration into a guided visual process. Our approach introduces three key innovations: a dual-branch structure combining retrieval and generation for diverse candidate exploration; a multi-view hybrid scoring approach that leverages MLLMs with innovative high-level metrics to assess 3D models with human-expert consistency; and a prompt-driven visual analytics suite that enables intuitive defect identification and refinement. Extensive testing and a user study demonstrate that Sel3DCraft surpasses other T23D systems in supporting creativity for designers.
Tianyi Liang 0002, Haiwen Huang, Shiqi Jiang 0001, Yifei Huang 0006, Liangyu Chen 0001, Changbo Wang, Chenhui Li 0001
IEEE Trans. Vis. Comput. Graph.3
2025 LoftUp: Learning a Coordinate-Based Feature Upsampler for Vision Foundation Models
Haiwen Huang, Anpei Chen, Volodymyr Havrylov, Andreas Geiger 0001
ICCV1
2025 Scientists' First Exam: Probing Cognitive Abilities of MLLM via Perception, Understanding, and Reasoning
abstract
Scientific discoveries increasingly rely on complex multimodal reasoning based on information-intensive scientific data and domain-specific expertise. Empowered by expert-level scientific benchmarks, scientific Multimodal Large Language Models (MLLMs) hold the potential to significantly enhance this discovery process in realistic workflows. However, current scientific benchmarks mostly focus on evaluating the knowledge understanding capabilities of MLLMs, leading to an inadequate assessment of their perception and reasoning abilities. To address this gap, we present the Scientists’ First Exam (SFE) benchmark, designed to evaluate the scientific cognitive capacities of MLLMs through three interconnected levels: scientific signal perception, scientific attribute understanding, scientific comparative reasoning. Specifically, SFE comprises 830 expert-verified VQA pairs across three question types, spanning 66 multimodal tasks across five high-value disciplines. Extensive experiments reveal that current state-of-the-art GPT-o3 and InternVL-3 achieve only 34.08% and 26.52% on SFE, highlighting significant room for MLLMs to improve in scientific realms. We hope the insights obtained in SFE will facilitate further developments in AI-enhanced scientific discoveries.
Yuhao Zhou 0005, Ruoyao Xiao, Qiantai Feng, Zijie Guo, Yuejin Yang, Wenxuan Huang 0001, Dan Si, Xiuqi Yao, Jia Bu, Haiwen Huang, Tianfan Fu, Shixiang Tang, Ben Fei, Dongzhan Zhou, Fenghua Ling, Yan Lu 0001, Chenhui Li 0001, Guanjie Zheng, Lei Bai 0001
NeurIPS15
2025 Prompt2Color: A prompt-based framework for image-derived color generation and visualization optimization
Jiayun Hu, Shiqi Jiang 0001, Haiwen Huang, Changbo Wang, Chenhui Li 0001
Comput. Graph.3
2025 SUPQA: LLM-based Geo-Visualization for Subjective Urban Performance Question-Answering
abstract
Abstract As urbanization accelerates, urban performance has become a growing concern, impacting every aspect of residents' lives. However, urban performance exploration is a tedious and highly subjective process for users. Users need to manually collect and integrate various information, or spend a large amount of time and effort due to the steep learning curves of existing specialized tools. To address these challenges, we introduce SUPQA, a novel approach for urban performance exploration using natural language as input and interactive geographic visualizations as output. Our approach leverages Large Language Models (LLMs) to effectively interpret user intents and quantify various urban performance measures. We integrate progressive navigation and multi‐geographic scale analysis in our visualization system, explaining the reasoning process and streamlining users' decision‐making workflow. Two usage scenarios and evaluations demonstrate the effectiveness of SUPQA in helping residents and planners acquire desired information more efficiently and enhancing the quality of decision‐making.
Haiwen Huang, Juntong Chen, Changbo Wang, Chenhui Li 0001
Comput. Graph. Forum1
2024 SalienTime: User-driven Selection of Salient Time Steps for Large-Scale Geospatial Data Visualization
abstract
The voluminous nature of geospatial temporal data from physical monitors and simulation models poses challenges to efficient data access, often resulting in cumbersome temporal selection experiences in web-based data portals. Thus, selecting a subset of time steps for prioritized visualization and pre-loading is highly desirable. Addressing this issue, this paper establishes a multifaceted definition of salient time steps via extensive need-finding studies with domain experts to understand their workflows. Building on this, we propose a novel approach that leverages autoencoders and dynamic programming to facilitate user-driven temporal selections. Structural features, statistical variations, and distance penalties are incorporated to make more flexible selections. User-specified priorities, spatial regions, and aggregations are used to combine different perspectives. We design and implement a web-based interface to enable efficient and context-aware selection of time steps and evaluate its efficacy and usability through case studies, quantitative evaluations, and expert interviews.
Juntong Chen, Haiwen Huang, Huayuan Ye, Zhong Peng, Chenhui Li 0001, Changbo Wang
CHI2
2024 Renovating Names in Open-Vocabulary Segmentation Benchmarks
abstract
Names are essential to both human cognition and vision-language models. Open-vocabulary models utilize class names as text prompts to generalize to categories unseen during training. However, the precision of these names is often overlooked in existing datasets. In this paper, we address this underexplored problem by presenting a framework for "renovating" names in open-vocabulary segmentation benchmarks (RENOVATE). Our framework features a renaming model that enhances the quality of names for each visual segment. Through experiments, we demonstrate that our renovated names help train stronger open-vocabulary models with up to 15% relative improvement and significantly enhance training efficiency with improved data quality. We also show that our renovated names improve evaluation by better measuring misclassification and enabling fine-grained model analysis. We provide our code and relabelings for several popular segmentation datasets to the research community on our project page: https://andrehuang.github.io/renovate.
Haiwen Huang, Songyou Peng, Andreas Geiger 0001
NeurIPS1
2024 A class-aware multi-stage UDA framework for prostate zonal segmentation
Zibo Ma, Yue Mi, Bo Zhang 0032, Zheng Zhang 0038, Yu Bai 0020, Jingyun Wu, Haiwen Huang, Wendong Wang 0003
Multim. Tools Appl.7
2023 GOOD: Exploring geometric cues for detecting objects in an open world
Haiwen Huang, Andreas Geiger 0001
ICLR1
2022 ACL-Net: Adaptive and Collaborative Learning Network for Multi-Site Prostate MRI Segmentation
abstract
High-performance deep learning models require large amounts of data with high quality annotations for model training, while the labeling work usually takes a lot of time for the experts. Meanwhile, the inter-observer variability al-ways exist between annotations from different experts and the distribution shift between the data acquired from different medical institutions. To address these challenges, we propose an end-to-end domain adaptive collaborative learning network for multi-institutional prostate MRI segmentation. Specifically, we introduce an unpaired image translation module to match the image domains between different institutions, which can alleviate the heterogeneity between 1.5T and 3T prostate MR images during model training. Moreover, we design a self-taught strategy to transfer domain-aware knowledge to jointly learn generic and unique representations. Furthermore, we evaluate our approach in scenarios with limited or without annotations, experimental results show that our approach has better adaptation performance than traditional supervised learning approaches, and has the potential to extend to unsupervised domain adaptation scenario. We also evaluate our approach with prostate MRI segmentation benchmark datasets, experimental results show that our approach outperforms several state-of-the-art methods.
Zibo Ma, Bo Zhang 0032, Zheng Zhang 0038, Wendong Wang 0003, Yue Mi, Haiwen Huang, Jingyun Wu
IEEE Big Data6
2022 LSRML: A latent space regularization based meta-learning framework for MR image segmentation
Bo Zhang 0032, Yunpeng Tan, Zheng Zhang 0038, Xiuzhuang Zhou, Jingyun Wu, Yue Mi, Haiwen Huang, Wendong Wang 0003
Pattern Recognit.8
2022 A Scalable Graph-Based Framework for Multi-Organ Histology Image Classification
abstract
Graph-based approaches are successful for histology image classification tasks but still face many challenges, such as: 1) the lack of nuclei-level labels and the significant variations between histology images make it extremely difficult to extract discriminative high-level nuclei features like nuclei type, texture and micro-environment; 2) graph-based approaches cannot handle large-scale cell graph nodes typically contained in histology images; and 3) graph neural networks (GNNs) struggle to learn the long-range dependency of cell graphs. To address the above challenges, we propose a scalable graph-based framework for multi-organ histology image classification. We develop a two-step masked nuclei patches supervised training approach to extract discriminative high-level nuclei features for histology images without nuclei-level labels. Additionally, we introduce a nuclei sampling strategy to make our graph-based framework scalable for large-scale cell graphs. Furthermore, we proposeHierArchicalTransformer Graph NeuralNetwork (HAT-Net+) for cell graph classi- fications. HAT-Net+ adopts Transformer to model the long-range dependency of cell graphs and a parameter-free approach to adaptively fuse different hierarchical graph representations of each layer. We achieved the state-of-the-art results on four public histology image classification datasets: CRC dataset (100%), Extended CRC dataset (98%), UZH dataset (96.9%) and BACH dataset (88%). Unlike other methods, our approach can be used in various histology image classification tasks, even for images without nuclei-level labels, indicating its potential in cancer diagnosis. The code is available athttps://github.com/suyouooooo/HAT-Net.
Yu Bai 0020, Yue Mi, Yihan Su, Bo Zhang 0032, Zheng Zhang 0038, Jingyun Wu, Haiwen Huang, Yongping Xiong, Xiangyang Gong, Wendong Wang 0003
IEEE J. Biomed. Health Informatics7
2021 MFSL-Net: A Modality Fusion and Shape Learning based Cascaded Network for Prostate Tumor Segmentation
abstract
Contouring prostate tumor in magnetic resonance images is a prerequisite for diagnosis. Automatically segmenting blurred lesion regions is challenging and requires fully leveraging multi-parameter MR images. This paper proposes MFSL-Net, an end-to-end network that cascades two novel sub-networks: 1) a modality fusion network that selectively fuses information of two MRI modalities by expanding a dual-stream CNN with spatial and channel attention modules; 2) a shape learning network that integrates shape learning and context learning to recognize the shape and edge information while preserving high-resolution semantic information. We justify MFSL-Net’s design by ablation experiments and compare its performance with the state-of-the-art approaches. Experimental results show a 3.6% improvement in Dice Similarity Coefficient, which confirms the effectiveness of MFSL-Net.
Bo Zhang 0032, Zheng Zhang 0038, Yue Mi, Jingyun Wu, Haiwen Huang, Xirong Que, Wendong Wang 0003
IEEE BigData6
2019 Nostalgic Adam: Weighting More of the Past Gradients When Designing the Adaptive Learning Rate
abstract
First-order optimization algorithms have been proven prominent in deep learning. In particu- lar, algorithms such as RMSProp and Adam are extremely popular. However, recent works have pointed out the lack of “long-term memory” in Adam-like algorithms, which could hamper their performance and lead to divergence. In our study, we observe that there are benefits of weighting more of the past gradients when designing the adaptive learning rate. We therefore propose an algorithm called the Nostalgic Adam (NosAdam) with theoretically guaranteed convergence at the best known convergence rate. NosAdam can be regarded as a fix to the non-convergence issue of Adam in alternative to the recent work of [Reddi et al., 2018]. Our preliminary numerical experiments show that NosAdam is a promising alternative al- gorithm to Adam. The proofs, code and other supplementary materials are already released.
Haiwen Huang
IJCAI1