Nadine Chang

dblp:227/2758 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
6since 2021 · last 2026
0000-0003-4765-8478ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Language models and text generation · 25% Vision and language · 22% Image recognition and object detection · 13%
Databases, data mining, and information retrieval
1 paper
Machine learning and data management · 50% Data integration and cleaning · 50%

Topics — the 14 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language
multimodal reasoning
0.912025
OmniDrive: A Holistic Vision-Language Dataset for Autonomous Driving with Counterfactual Reasoning · CVPR 2025
Natural language and speech › Language models and text generation › prompting
prompt robustness
0.912025
PARC: A Quantitative Framework Uncovering the Symmetries within Vision Language Models · CVPR 2025
Natural language and speech › Language models and text generation › prompting
prompt sensitivity
0.912025
PARC: A Quantitative Framework Uncovering the Symmetries within Vision Language Models · CVPR 2025
Machine learning › Trustworthy machine learning
robustness
0.912025
PARC: A Quantitative Framework Uncovering the Symmetries within Vision Language Models · CVPR 2025
Computer vision › Vision and language
vision-language model
0.912025
PARC: A Quantitative Framework Uncovering the Symmetries within Vision Language Models · CVPR 2025
Robotics › Autonomous driving
vision-language model for driving
0.912025
OmniDrive: A Holistic Vision-Language Dataset for Autonomous Driving with Counterfactual Reasoning · CVPR 2025
Data integration and cleaning › data transformation
data enrichment
0.912025
SSE: Multimodal Semantic Data Selection and Enrichment for Industrial-scale Data Assimilation · KDD (1) 2025
Machine learning and data management
data selection
0.912025
SSE: Multimodal Semantic Data Selection and Enrichment for Industrial-scale Data Assimilation · KDD (1) 2025
Computer vision › Segmentation and scene understanding
instance segmentation
0.512021
Image-Level or Object-Level? A Tale of Two Resampling Strategies for Long-Tailed Detection · ICML 2021
Computer vision › Segmentation and scene understanding › instance segmentation
long-tailed instance segmentation
0.512021
Image-Level or Object-Level? A Tale of Two Resampling Strategies for Long-Tailed Detection · ICML 2021
Computer vision › Image recognition and object detection › object detection › robust object detection
long-tailed object detection
0.512021
Image-Level or Object-Level? A Tale of Two Resampling Strategies for Long-Tailed Detection · ICML 2021
Computer vision › Image recognition and object detection
object detection
0.512021
Image-Level or Object-Level? A Tale of Two Resampling Strategies for Long-Tailed Detection · ICML 2021
Machine learning › Probabilistic and Bayesian machine learning
data assimilation
0.312025
SSE: Multimodal Semantic Data Selection and Enrichment for Industrial-scale Data Assimilation · KDD (1) 2025
Natural language and speech › Language models and text generation › prompting
prompt engineering
0.312025
PARC: A Quantitative Framework Uncovering the Symmetries within Vision Language Models · CVPR 2025

Methods — techniques the papers use, named apart from their topics

multimodal learning · 1.7reliability scoring · 0.9large language model · 0.9counterfactual reasoning · 0.9calibration · 0.9episodic memory bank · 0.5dynamic resampling · 0.5
YearPublicationVenuePosition
2026 GHOST: Getting to the Bottom of Hallucinations with A Multi-round Consistency Benchmark
Vibashan VS, Nadine Chang, Jenny Schmalfuss, Vishal M. Patel, Zhiding Yu, José M. Álvarez 0004
WACV2
2025 PARC: A Quantitative Framework Uncovering the Symmetries within Vision Language Models
abstract
Vision language models (VLMs) respond to user-crafted text prompts and visual inputs, and are applied to numerous real-world problems. VLMs integrate visual modalities with large language models (LLMs), which are well known to be prompt-sensitive. Hence, it is crucial to determine whether VLMs inherit this instability to varying prompts. We therefore investigate which prompt variations VLMs are most sensitive to and which VLMs are most agnostic to prompt variations. To this end, we introduce PARC (Prompt Analysis via Reliability and Calibration), a VLM prompt sensitivity analysis framework built on three pillars: (1) plausible prompt variations in both the language and vision domain, (2) a novel model reliability score with built-in guarantees, and (3) a calibration step that enables dataset-and prompt-spanning prompt variation analysis. Regarding prompt variations, PARC’s evaluation shows that VLMs mirror LLM language prompt sensitivity in the vision domain, and most destructive variations change the expected answer. Regarding models, outstandingly robust VLMs among 22 evaluated models come from the InternVL2 family. We further find indications that prompt sensitivity is linked to training data. https://github.com/NVlabs/PARC
Jenny Schmalfuss, Nadine Chang, Vibashan VS, Maying Shen, Andrés Bruhn, José M. Álvarez 0004
CVPR2
2025 OmniDrive: A Holistic Vision-Language Dataset for Autonomous Driving with Counterfactual Reasoning
abstract
The advances in vision-language models (VLMs) have led to a growing interest in autonomous driving to leverage their strong reasoning capabilities. However, extending these capabilities from 2D to full 3D understanding is crucial for real-world applications. To address this challenge, we propose OmniDrive, a holistic vision-language dataset that aligns agent models with 3D driving tasks through counter-factual reasoning. This approach enhances decision-making by evaluating potential scenarios and their outcomes, similar to human drivers considering alternative actions. Our counterfactual-based synthetic data annotation process generates large-scale, high-quality datasets, providing denser supervision signals that bridge planning trajectories and language-based reasoning. Futher, we explore two advanced OmniDrive-Agent frameworks, namely Omni-L and Omni-Q, to assess the importance of vision-language alignment versus 3D perception, revealing critical insights into designing effective LLM-agents. Significant improvements on the DriveLM Q&A benchmark and nuScenes open-loop planning demonstrate the effectiveness of our dataset and methods.
Zhiding Yu, Xiaohui Jiang, Shiyi Lan, Nadine Chang, Jan Kautz, José M. Álvarez 0004
CVPR6
2025 Enhancing Autonomous Driving Safety with Collision Scenario Integration
abstract
Autonomous vehicle safety is crucial for the successful deployment of self-driving cars. However, most existing planning methods rely heavily on imitation learning, which limits their ability to leverage collision data effectively. Moreover, collecting collision or near-collision data is inherently challenging, as it involves risks and raises ethical and practical concerns. In this paper, we propose SafeFusion, a training framework to learn from collision data. Instead of over-relying on imitation learning, SafeFusion integrates safety-oriented metrics during training to enable collision avoidance learning. In addition, to address the scarcity of collision data, we propose CollisionGen, a scalable data generation pipeline to generate diverse, high-quality scenarios using natural language prompts, generative models, and rule-based filtering. Experimental results show that our approach improves planning performance in collision-prone scenarios by 56% over previous state-of-the-art planners while maintaining effectiveness in regular driving situations. Our work provides a scalable and effective solution for advancing the safety of autonomous driving systems.
Shiyi Lan, Xinglong Sun, Nadine Chang, Zhenxin Li, Zhiding Yu, José M. Álvarez 0004
IROS4
2025 SSE: Multimodal Semantic Data Selection and Enrichment for Industrial-scale Data Assimilation
Maying Shen, Nadine Chang, Sifei Liu, José M. Álvarez 0004
KDD (1)2
2021 Image-Level or Object-Level? A Tale of Two Resampling Strategies for Long-Tailed Detection
abstract
Training on datasets with long-tailed distributions has been challenging for major recognition tasks such as classification and detection. To deal with this challenge, image resampling is typically introduced as a simple but effective approach. However, we observe that long-tailed detection differs from classification since multiple classes may be present in one image. As a result, image resampling alone is not enough to yield a sufficiently balanced distribution at the object-level. We address object-level resampling by introducing an object-centric sampling strategy based on a dynamic, episodic memory bank. Our proposed strategy has two benefits: 1) convenient object-level resampling without significant extra computation, and 2) implicit feature-level augmentation from model updates. We show that image-level and object-level resamplings are both important, and thus unify them with a joint resampling strategy. Our method achieves state-of-the-art performance on the rare categories of LVIS, with 1.89% and 3.13% relative improvements over Forest R-CNN on detection and instance segmentation.
Nadine Chang, Zhiding Yu, Yu-Xiong Wang, Anima Anandkumar, Sanja Fidler, José M. Álvarez 0004
ICML1