Noel Codella

dblp:172/1174 · also Noel C. F. Codella, Noel Christopher Codella · DBLP profile ↗
← Back
22ranked-venue papers
4as first author
10since 2021 · last 2026
0000-0001-6735-9067ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 15 · 4 first-author · 7 since 2021Artificial intelligence and machine learning · 11 · 8 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-authorHuman-computer interaction and ubiquitous computing · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Exploring the Future of AI in Clinical Collaboration: A Study on Tumor Board Case Preparation
abstract
Multidisciplinary tumor boards (MTBs) bring specialists together to identify therapies for complex cancer cases, but preparing for them is time-intensive. Clinicians must extract key details from extensive records and evaluate treatment options. While large language models (LLMs) show promise in medicine for basic tasks like summarizing notes, little is known about their role in high-stakes tasks like MTB preparation. We conducted a mixed-methods study with 16 oncologists using two AI systems to prepare patient cases for MTB: an off-the-shelf assistant (Copilot) and a task-specific multi-agent system (Healthcare Agent Orchestrator, HAO). We analyzed oncologist prompts, AI responses, and oncologists’ perception of AI. Participants showed greater willingness to adopt HAO but were often overconfident in AI summaries and skeptical of AI-recommended therapies. Trust calibration strategies, such as source links and agent-trajectories, failed to align trust with system capabilities. We conclude with how AI systems should be built to support clinicians in high-stakes tasks.
Amanda K. Hall, Ruican Rachel Zhong, Selin S. Everett, Alyssa Unell, Matthias Blondeel, Jonathan Carlson, Katie Claveau, Thulasee Jose, Tristan Naumann, David C. Rhew, Naiteek Sangani, Frank Tuan, James Weinstein, Varun Mishra 0001, Elizabeth D. Mynatt, T. Scott Saponas, Leonardo Schettini, J. Samuel Preston, Yu Gu 0017, Naoto Usuyama, Zelalem Gero, Cliff Wong, Noel Codella, Hoifung Poon, Shrey Jain, Matthew P. Lungren, Eric Horvitz
CHI26
2026 Generative Enhancement for 3D Medical Images
abstract
Abstract The limited availability of 3D medical image datasets, due to privacy concerns and high collection or annotation costs, poses significant challenges in the field of medical imaging. There are few solutions for realistic 3D medical image synthesis due to difficulties in backbone design and fewer 3D training samples compared to 2D counterparts. In this paper, we propose GEM-3D , a novel generative approach to the synthesis of 3D medical images and the enhancement of existing datasets using conditional diffusion models. Our method begins with a 2D slice, noted as the informed slice to serve the patient prior, and propagates the generation process using a 3D segmentation mask. By decomposing the 3D medical images into editable masks and patient prior information, GEM-3D offers a flexible yet effective solution for generating versatile 3D images from existing datasets. Moreover, as the informed slice contains patient-wise information, GEM-3D can also facilitate counterfactual image synthesis and dataset-level de-enhancement with desired control. Experiments on brain MRI and abdomen CT images demonstrate that GEM-3D is capable of synthesizing high-quality 3D medical images with volumetric consistency, offering a straightforward solution for dataset enhancement during inference. The code is available at https://github.com/HKU-MedAI/GEM-3D .
Lingting Zhu, Noel Codella, Dongdong Chen 0001, Zhenchao Jin, Lu Yuan 0001, Lequan Yu
Int. J. Comput. Vis.2
2025 ImageCLEF 2025: Multimedia Retrieval in Medical, Social Media and Content Recommendation Applications
Bogdan Ionescu, Henning Müller, Dan-Cristian Stanciu, Ahmad Idrissi-Yaghir, Ahmedkhan Radzhabov, Alba Garcia Seco de Herrera, Alexandra-Georgiana Andrei, Andrea M. Storås, Asma Ben Abacha, Benjamin Bracke, Benjamin Lecouteux, Benno Stein 0001, Cécile Macaire, Christoph M. Friedrich, Cynthia Sabrina Schmidt, Diandra Fabre, Didier Schwab, Dimitar Dimitrov 0003, Emmanuelle Esperança-Rodier, Mihai Gabriel Constantin, Helmut Becker, Hendrik Damm, Henning Schäfer, Ivan Rodkin, Ivan Koychev, Johannes Kiesel, Johannes Rückert, Josep Malvehy, Liviu-Daniel Stefan, Louise Bloch, Martin Potthast, Maximilian Heinrich, Michael Riegler 0001, Mihai Dogariu, Noel Codella, Pål Halvorsen, Preslav Nakov, Raphael Brüngel, Roberto A. Novoa, Rocktim Jyoti Das, Steven Alexander Hicks, Sushant Gautam, Tabea Margareta Grace Pakull, Vajira Thambawita, Vassili Kovalev, Wen-Wai Yim, Zhuohan Xie
ECIR (5)35
2024 Fully Authentic Visual Question Answering Dataset from Online Communities
Chongyan Chen, Mengchen Liu, Noel Codella, Yunsheng Li, Lu Yuan 0001, Danna Gurari
ECCV (48)3
2023 i-Code: An Integrative and Composable Multimodal Learning Framework
abstract
Human intelligence is multimodal; we integrate visual, linguistic, and acoustic signals to maintain a holistic worldview. Most current pretraining methods, however, are limited to one or two modalities. We present i-Code, a self-supervised pretraining framework where users may flexibly combine the modalities of vision, speech, and language into unified and general-purpose vector representations. In this framework, data from each modality are first given to pretrained single-modality encoders. The encoder outputs are then integrated with a multimodal fusion network, which uses novel merge- and co-attention mechanisms to effectively combine information from the different modalities. The entire system is pretrained end-to-end with new objectives including masked modality unit modeling and cross-modality contrastive learning. Unlike previous research using only video for pretraining, the i-Code framework can dynamically process single, dual, and triple-modality data during training and inference, flexibly projecting different combinations of modalities into a single representation space. Experimental results demonstrate how i-Code can outperform state-of-the-art techniques on five multimodal understanding tasks and single-modality benchmarks, improving by as much as 11% and demonstrating the power of integrative multimodal pretraining.
Ziyi Yang 0011, Yuwei Fang, Chenguang Zhu 0001, Reid Pryzant, Dongdong Chen 0001, Yu Shi 0001, Yichong Xu, Yao Qian, Mei Gao, Liyang Lu, Yujia Xie, Robert Gmyr, Noel Codella, Naoyuki Kanda, Bin Xiao 0004, Lu Yuan 0001, Takuya Yoshioka, Michael Zeng 0001, Xuedong Huang 0001
AAAI14
2023 Streaming Video Model
abstract
Video understanding tasks have traditionally been modeled by two separate architectures, specially tailored for two distinct tasks. Sequence-based video tasks, such as action recognition, use a video backbone to directly extract spatiotemporal features, while frame-based video tasks, such as multiple object tracking (MOT), rely on single fixed-image backbone to extract spatial features. In contrast, we propose to unify video understanding tasks into one novel streaming video architecture, referred to as Streaming Vision Transformer (S-ViT). S-ViT first produces frame-level features with a memory-enabled temporally-aware spatial encoder to serve the frame-based video tasks. Then the frame features are input into a task-related temporal decoder to obtain spatiotemporal features for sequence-based tasks. The efficiency and efficacy of S-ViT is demonstrated by the state-of-the-art accuracy in the sequence-based action recognition task and the competitive advantage over conventional architecture in the frame-based MOT task. We believe that the concept of streaming video model and the implementation of S-ViT are solid steps towards a unified deep learning architecture for video understanding. Code will be available at https://github.com/yuzhms/Streaming-Video-Model.
Chong Luo 0001, Chuanxin Tang, Dongdong Chen 0001, Noel Codella, Zhengjun Zha
CVPR5
2022 RegionCLIP: Region-based Language-Image Pretraining
abstract
Contrastive language-image pretraining (CLIP) using image-text pairs has achieved impressive results on image classification in both zero-shot and transfer learning set-tings. However, we show that directly applying such mod-els to recognize image regions for object detection leads to unsatisfactory performance due to a major domain shift: CLIP was trained to match an image as a whole to a text de-scription, without capturing the fine-grained alignment be-tween image regions and text spans. To mitigate this issue, we propose a new method called RegionCLIP that signifi-cantly extends CLIP to learn region-level visual representations, thus enabling fine-grained alignment between image regions and textual concepts. Our method leverages a CLIP model to match image regions with template captions, and then pretrains our model to align these region-text pairs in the feature space. When transferring our pretrained model to the open-vocabulary object detection task, our method outperforms the state of the art by 3.8 AP50 and 2.2 AP for novel categories on COCO and LVIS datasets, respectively. Further, the learned region representations support zero-shot inference for object detection, showing promising results on both COCO and LVIS datasets. Our code is available at https://github.com/microsoft/RegionCLIP.
Yiwu Zhong, Pengchuan Zhang, Chunyuan Li, Noel Codella, Liunian Harold Li, Luowei Zhou, Xiyang Dai, Lu Yuan 0001, Yin Li 0003, Jianfeng Gao 0001
CVPR5
2022 DaViT: Dual Attention Vision Transformers
Mingyu Ding, Bin Xiao 0004, Noel Codella, Ping Luo 0002, Jingdong Wang 0001, Lu Yuan 0001
ECCV (24)3
2022 Learning Visual Representation from Modality-Shared Contrastive Language-Image Pre-training
Haoxuan You, Luowei Zhou, Bin Xiao 0004, Noel Codella, Yu Cheng 0001, Ruochen Xu, Shih-Fu Chang, Lu Yuan 0001
ECCV (27)4
2021 CvT: Introducing Convolutions to Vision Transformers
abstract
We present in this paper a new architecture, named Convolutional vision Transformer (CvT), that improves Vision Transformer (ViT) in performance and efficiency by introducing convolutions into ViT to yield the best of both de-signs. This is accomplished through two primary modifications: a hierarchy of Transformers containing a new convolutional token embedding, and a convolutional Transformer block leveraging a convolutional projection. These changes introduce desirable properties of convolutional neural networks (CNNs) to the ViT architecture (i.e. shift, scale, and distortion invariance) while maintaining the merits of Transformers (i.e. dynamic attention, global context, and better generalization). We validate CvT by conducting extensive experiments, showing that this approach achieves state-of-the-art performance over other Vision Transformers and ResNets on ImageNet-1k, with fewer parameters and lower FLOPs. In addition, performance gains are maintained when pretrained on larger datasets (e.g. ImageNet-22k) and fine-tuned to downstream tasks. Pretrained on ImageNet-22k, our CvT-W24 obtains a top-1 accuracy of 87.7% on the ImageNet-1k val set. Finally, our results show that the positional encoding, a crucial component in existing Vision Transformers, can be safely re-moved in our model, simplifying the design for higher resolution vision tasks. Code will be released at https://github.com/microsoft/CvT.
Haiping Wu, Bin Xiao 0004, Noel Codella, Mengchen Liu, Xiyang Dai, Lu Yuan 0001, Lei Zhang 0001
ICCV3
2020 A Broader Study of Cross-Domain Few-Shot Learning
Yunhui Guo, Noel Codella, Leonid Karlinsky, James V. Codella, John R. Smith, Kate Saenko, Tajana Rosing, Rogério Feris
ECCV (27)2
2020 Fairness of Classifiers Across Skin Tones in Dermatology
Newton M. Kinyanjui, Timothy Odonga, Celia Cintas, Noel Codella, Rameswar Panda, Prasanna Sattigeri, Kush R. Varshney
MICCAI (6)4
2019 TED: Teaching AI to Explain its Decisions
abstract
Artificial intelligence systems are being increasingly deployed due to their potential to increase the efficiency, scale, consistency, fairness, and accuracy of decisions. However, as many of these systems are opaque in their operation, there is a growing demand for such systems to provide explanations for their decisions. Conventional approaches to this problem attempt to expose or discover the inner workings of a machine learning model with the hope that the resulting explanations will be meaningful to the consumer. In contrast, this paper suggests a new approach to this problem. It introduces a simple, practical framework, called Teaching Explanations for Decisions (TED), that provides meaningful explanations that match the mental model of the consumer. We illustrate the generality and effectiveness of this approach with two different examples, resulting in highly accurate explanations with no loss of prediction accuracy for these two examples.
Michael Hind, Dennis Wei, Murray Campbell, Noel Codella, Amit Dhurandhar, Aleksandra Mojsilovic, Karthikeyan Natesan Ramamurthy, Kush R. Varshney
AIES4
2019 Dermoscopy Image Analysis: Overview and Future Directions
abstract
Dermoscopy is a non-invasive skin imaging technique that permits visualization of features of pigmented melanocytic neoplasms that are not discernable by examination with the naked eye. While studies on the automated analysis of dermoscopy images date back to the late 1990s, because of various factors (lack of publicly available datasets, open-source software, computational power, etc.), the field progressed rather slowly in its first two decades. With the release of a large public dataset by the International Skin Imaging Collaboration in 2016, development of open-source software for convolutional neural networks, and the availability of inexpensive graphics processing units, dermoscopy image analysis has recently become a very active research field. In this paper, we present a brief overview of this exciting subfield of medical image analysis, primarily focusing on three aspects of it, namely, segmentation, feature extraction, and classification. We then provide future directions for researchers.
M. Emre Celebi 0001, Noel Codella, Alan Halpern
IEEE J. Biomed. Health Informatics2
2019 Guest Editorial Skin Lesion Image Analysis for Melanoma Detection
abstract
The papers in this special section focus on the use of image analysis to detect Melanoma. Melanoma is deadliest form of skin cancer, with roughly 91,000 new cases reported every year in the US and more than 9,000 deaths. Unlike many other cancer types, the incidence rate of melanoma has been steadily increasing in the past several decades. Early diagnosis is crucial since melanoma can be cured with a simple excision, if detected early. The goals of this special issue are to summarize the state-ofthe- art in the automated analysis of skin lesion images and to provide future directions for this exciting subfield of medical image analysis. The intended audience includes researchers and practicing clinicians, who are increasingly using digital analytic tools.
M. Emre Celebi 0001, Noel Codella, Alan Halpern, Dinggang Shen
IEEE J. Biomed. Health Informatics2
2017 Leveraging multiple cues for recognizing family photos
Xiaolong Wang 0006, Guodong Guo, Michele Merler, Noel Codella, M. V. Rohith, John R. Smith, Chandra Kambhamettu
Image Vis. Comput.4
2014 Automated Medical Image Modality Recognition by Fusion of Visual and Text Information
Noel Codella, Jonathan H. Connell, Sharath Pankanti, Michele Merler, John R. Smith
MICCAI (2)1
2014 Modeling Attributes from Category-Attribute Proportions
abstract
Attribute-based representation has been widely used in visual recognition and retrieval due to its interpretability and cross-category generalization properties. However, classic attribute learning requires manually labeling attributes on the images, which is very expensive, and not scalable. In this paper, we propose to model attributes from category-attribute proportions. The proposed framework can model attributes without attribute labels on the images. Specifically, given a multi-class image datasets with N categories, we model an attribute, based on an N-dimensional category-attribute proportion vector, where each element of the vector characterizes the proportion of images in the corresponding category having the attribute. The attribute learning can be formulated as a learning from label proportion (LLP) problem. Our method is based on a newly proposed machine learning algorithm called $\propto$SVM. Finding the category-attribute proportions is much easier than manually labeling images, but it is still not a trivial task. We further propose to estimate the proportions from multiple modalities such as human commonsense knowledge, NLP tools, and other domain knowledge. The value of the proposed approach is demonstrated by various applications including modeling animal attributes, visual sentiment attributes, and scene attributes.
Felix X. Yu, Liangliang Cao, Michele Merler, Noel Codella, Tao Chen 0015, John R. Smith, Shih-Fu Chang
ACM Multimedia4
2013 Large-scale video event classification using dynamic temporal pyramid matching of visual semantics
abstract
Video event classification and retrieval has recently emerged as a challenging research topic. In addition to the variation in appearance of visual content and the large scale of the collections to be analyzed, this domain presents new and unique challenges in the modeling of the explicit temporal structure and implicit temporal trends of content within the video events. In this study, we present a technique for video event classification that captures temporal information over semantics using a scalable and efficient modeling scheme. An architecture for partitioning videos into a linear temporal pyramid, using segments of equal length and segments determined by the patterns of the underlying data, is applied over a rich underlying semantic description at the frame level using a taxonomy of nearly 1000 concepts containing 500,000 training images. Forward model selection with data bagging is used to prune the space of temporal features and data for efficiency. The system is implemented in the Hadoop Map-Reduce environment for arbitrary scalability. Our method is applied to the TRECVID Multimedia Event Detection 2012 task. Results demonstrate a significant boost in performance of over 50%, in terms of mean average precision, compared to common max or average pooling, and 17.7% compared to more complex pooling strategies that ignore temporal content.
Noel Codella, Gang Hua 0001, Liangliang Cao, Michele Merler, Leiguang Gong, Matthew L. Hill, John R. Smith
ICIP1
2013 Learning by focusing: A new framework for concept recognition and feature selection
abstract
In this paper, we develop a new method for feature selection and category learning. We first introduce two observations from our experiments: (1) It is easier to distinguish two concepts than to learn an isolated concept. (2) To distinguish different concept pairs we can find different selections of optimal features. These two observations may partly explain the success of human vision learning, especially why an infant can simultaneously capture distinguished visual features when learning new concepts. Based on these two observations, we developed a new learning-by-focusing method which first constructs focalized concept discriminators for pairs of concepts, and then builds nonlinear classifiers using the discrimination scores. We build datasets for four concept structure: vehicle, human affliction, sports, and animals, and experiments on all the four datasets verify the success of our new approach.
Liangliang Cao, Leiguang Gong, John R. Kender, Noel Codella, John R. Smith
ICME4
2012 Cardiac anatomy as a biometric
abstract
In this study, we propose a novel biometric signature for human identification based on anatomically unique structures of the left ventricle of the heart. An algorithm is developed that analyzes the 3 primary anatomical structures of the left ventricle: the endocardium, myocardium, and papillary muscles. Comparisons of these analyses between probe and gallery images produces a similarity score that is used as the basis of the biometric. The performance of the algorithm is tested on a cohort of 10 de-identified subjects imaged by Cardiac MRI. Perfect matching between individuals is obtained with good separation between the genuine and impostor classes. In summary, this study demonstrates using anatomy of the left ventricle of the human heart for the purposes of a biometric signature.
Noel Codella, Jonathan H. Connell, Nalini K. Ratha, Jonathan W. Weinsaft
ICIP1
2012 Video Event Detection Using Temporal Pyramids of Visual Semantics with Kernel Optimization and Model Subspace Boosting
abstract
In this study, we present a system for video event classification that generates a temporal pyramid of static visual semantics using minimum-value, maximum-value, and average-value aggregation techniques. Kernel optimization and model subspace boosting are then applied to customize the pyramid for each event. SVM models are independently trained for each level in the pyramid using kernel selection according to 3-fold cross-validation. Kernels that both enforce static temporal order and permit temporal alignment are evaluated. Model subspace boosting is used to select the best combination of pyramid levels and aggregation techniques for each event. The NIST TRECVID Multimedia Event Detection (MED) 2011 dataset was used for evaluation. Results demonstrate that kernel optimizations using both temporally static and dynamic kernels together achieves better performance than any one particular method alone. In addition, model sub-space boosting reduces the size of the model by 80%, while maintaining 96% of the performance gain.
Noel Codella, Apostol Natsev, Gang Hua 0001, Matthew L. Hill, Liangliang Cao, Leiguang Gong, John R. Smith
ICME1