Alexander Xiong

dblp:289/8719 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
3since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 3 since 2021Computer networks · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Trustworthy machine learning · 100%

Topics — the 6 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning › robustness › model robustness evaluation
adversarial robustness evaluation
1.122025
MMDT: Decoding the Trustworthiness and Safety of Multimodal Foundation Models · ICLR 2025
VMDT: Decoding the Trustworthiness of Video Foundation Models · NeurIPS 2025
Machine learning › Trustworthy machine learning › privacy
privacy evaluation
1.122025
VMDT: Decoding the Trustworthiness of Video Foundation Models · NeurIPS 2025
MMDT: Decoding the Trustworthiness and Safety of Multimodal Foundation Models · ICLR 2025
Machine learning › Trustworthy machine learning › fairness
fairness and bias evaluation
0.912025
MMDT: Decoding the Trustworthiness and Safety of Multimodal Foundation Models · ICLR 2025
Machine learning › Trustworthy machine learning › fairness
fairness evaluation
0.912025
VMDT: Decoding the Trustworthiness of Video Foundation Models · NeurIPS 2025
Machine learning › Trustworthy machine learning › hallucination
hallucination evaluation
0.912025
VMDT: Decoding the Trustworthiness of Video Foundation Models · NeurIPS 2025
Machine learning › Trustworthy machine learning
safety evaluation
0.912025
VMDT: Decoding the Trustworthiness of Video Foundation Models · NeurIPS 2025

Methods — techniques the papers use, named apart from their topics

benchmark construction · 1.7red teaming · 0.9large language model evaluation · 0.9
YearPublicationVenuePosition
2025 MMDT: Decoding the Trustworthiness and Safety of Multimodal Foundation Models
abstract
Multimodal foundation models (MMFMs) play a crucial role in various applications, including autonomous driving, healthcare, and virtual assistants. However, several studies have revealed vulnerabilities in these models, such as generating unsafe content by text-to-image models. Existing benchmarks on multimodal models either predominantly assess the helpfulness of these models, or only focus on limited perspectives such as fairness and privacy. In this paper, we present the first unified platform, MMDT (Multimodal DecodingTrust), designed to provide a comprehensive safety and trustworthiness evaluation for MMFMs. Our platform assesses models from multiple perspectives, including safety, hallucination, fairness/bias, privacy, adversarial robustness, and out-of-distribution (OOD) generalization. We have designed various evaluation scenarios and red teaming algorithms under different tasks for each perspective to generate challenging data, forming a high-quality benchmark. We evaluate a range of multimodal models using MMDT, and our findings reveal a series of vulnerabilities and areas for improvement across these perspectives. This work introduces the first comprehensive and unique safety and trustworthiness evaluation platform for MMFMs, paving the way for developing safer and more reliable MMFMs and systems. Our platform and benchmark are available at https://mmdecodingtrust.github.io/.
Chejian Xu, Jiawei Zhang 0013, Zhaorun Chen, Chulin Xie, Mintong Kang, Yujin Potter, Zhun Wang, Zhuowen Yuan, Alexander Xiong, Zidi Xiong, Lingzhi Yuan, Yi Zeng 0005, Peiyang Xu, Chengquan Guo, Andy Zhou, Jeffrey Ziwei Tan, Xuandong Zhao, Francesco Pinto, Zhen Xiang
ICLR9
2025 VMDT: Decoding the Trustworthiness of Video Foundation Models
abstract
As foundation models become more sophisticated, ensuring their trustworthiness becomes increasingly critical; yet, unlike text and image, the video modality still lacks comprehensive trustworthiness benchmarks. We introduce VMDT (Video-Modal DecodingTrust), the first unified platform for evaluating text-to-video (T2V) and video-to-text (V2T) models across five key trustworthiness dimensions: safety, hallucination, fairness, privacy, and adversarial robustness. Through our extensive evaluation of 7 T2V models and 19 V2T models using VMDT, we uncover several significant insights. For instance, all open-source T2V models evaluated fail to recognize harmful queries and often generate harmful videos, while exhibiting higher levels of unfairness compared to image modality models. In V2T models, unfairness and privacy risks rise with scale, whereas hallucination and adversarial robustness improve---though overall performance remains low. Uniquely, safety shows no correlation with model size, implying that factors other than scale govern current safety levels. Our findings highlight the urgent need for developing more robust and trustworthy video foundation models, and VMDT provides a systematic framework for measuring and tracking progress toward this goal. The code is available at https://sunblaze-ucb.github.io/VMDT-page/.
Yujin Potter, Zhun Wang, Nicholas Crispino, Kyle Montgomery, Alexander Xiong, Ethan Y. Chang, Francesco Pinto, Rahul Gupta 0001, Morteza Ziyadi, Christos Christodoulopoulos 0001, Bo Li 0026, Chenguang Wang 0001, Dawn Song
NeurIPS5
2022 Deep object detection for waterbird monitoring using aerial imagery
abstract
Monitoring of colonial waterbird nesting islands is essential to tracking waterbird population trends, which are used for evaluating ecosystem health and informing conservation management decisions. Recently, unmanned aerial vehicles, or drones, have emerged as a viable technology to precisely monitor waterbird colonies. However, manually counting waterbirds from hundreds, or potentially thousands, of aerial images is both difficult and time-consuming. In this work, we present a deep learning pipeline that can be used to precisely detect, count, and monitor waterbirds using aerial imagery collected by a commercial drone. By utilizing convolutional neural network-based object detectors, we show that we can detect 16 classes of waterbird species that are commonly found in colonial nesting islands along the Texas coast. Our experiments using Faster R-CNN and RetinaNet object detectors give mean interpolated average precision scores of 67.9% and 63.1% respectively.
Krish Kabra, Alexander Xiong, Minxuan Luo, William Lu, Tianjiao Yu, Dhananjay Singh 0003, Raul Garcia, Maojie Tang, Hank Arnold, Anna Vallery, Richard Gibbons, Arko Barman
ICMLA2
2020 Privacy Preserving Inference with Convolutional Neural Network Ensemble
abstract
Machine Learning as a Service on cloud not only provides a solution to scale demanding workloads, but also allows broader accessibility for the utilization of trained deep neural networks. For example, in the medical field, cloud-based deep-learning assisted diagnoses can be life-saving, especially in developing areas where experienced doctors and domain expertise are lacking. However, preserving end-users' data privacy while using cloud service for deep learning is a challenge. Some recent works based on fully homomorphic encryption have enabled neural-network predictions on encrypted input data. In this paper, we further extend the capability of privacy preserving deep neural network inference, through a joint decision made by multiple deep neural network models on encrypted data, to address bias caused by unbalanced local training datasets. In particular, we design and implement a privacy preserving prediction method through an ensemble of convolutional neural networks. The extensive experiment results show that our method can achieve higher accuracy compared to individual models, and preserve the user data privacy at the same level. We also verify the time efficiency of our implementation.
Alexander Xiong, Michael Nguyen, Andrew So, Tingting Chen 0001
IPCCC1