Ashish Shah

dblp:01/2068 · DBLP profile ↗
← Back
9ranked-venue papers
0as first author
5since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Software engineering, systems software and programming languages · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Vision and language · 32% Video understanding and tracking · 16% Segmentation and scene understanding · 14%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Cloud and datacenter computing · 50% Distributed systems · 50%

Topics — the 20 heaviest of 20, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Video understanding and tracking
long video understanding
0.812024
MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding · CVPR 2024
Computer vision › Vision and language › vision-language model
multimodal large language model
0.812024
MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding · CVPR 2024
Computer vision › Video understanding and tracking
video question answering
0.812024
MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding · CVPR 2024
Machine learning › Transfer learning and domain adaptation › zero-shot learning
zero-shot retrieval
0.812024
Spherical Linear Interpolation and Text-Anchoring for Zero-Shot Composed Image Retrieval · ECCV (19) 2024
Information retrieval › image retrieval
composed image retrieval
0.812024
Spherical Linear Interpolation and Text-Anchoring for Zero-Shot Composed Image Retrieval · ECCV (19) 2024
Machine learning › Representation and self-supervised learning
contrastive learning
0.712023
Open Vocabulary Semantic Segmentation with Patch Aligned Contrastive Learning · CVPR 2023
Computer vision › Segmentation and scene understanding › semantic segmentation
open-vocabulary segmentation
0.712023
Open Vocabulary Semantic Segmentation with Patch Aligned Contrastive Learning · CVPR 2023
Computer vision › Segmentation and scene understanding
semantic segmentation
0.712023
Open Vocabulary Semantic Segmentation with Patch Aligned Contrastive Learning · CVPR 2023
Computer vision › Vision and language
vision-language pretraining
0.712023
Open Vocabulary Semantic Segmentation with Patch Aligned Contrastive Learning · CVPR 2023
Machine learning › Time series and sequential data
anomaly detection
0.612022
Few-Shot Fast-Adaptive Anomaly Detection · NeurIPS 2022
Machine learning › Generative modeling
energy-based model
0.612022
Few-Shot Fast-Adaptive Anomaly Detection · NeurIPS 2022
Machine learning › Time series and sequential data › anomaly detection
few-shot anomaly detection
0.612022
Few-Shot Fast-Adaptive Anomaly Detection · NeurIPS 2022
Computer vision › Vision and language
image captioning
0.612022
Object-Centric Unsupervised Image Captioning · ECCV (36) 2022
Computer vision › Vision and language › image captioning › grounded image captioning
object captioning
0.612022
Object-Centric Unsupervised Image Captioning · ECCV (36) 2022
Computer vision › Vision and language › image captioning
unpaired image captioning
0.612022
Object-Centric Unsupervised Image Captioning · ECCV (36) 2022
Cloud and datacenter computing › datacenter operations
datacenter reliability
0.312018
Maelstrom: Mitigating Datacenter-level Disasters by Draining Interdependent Traffic Safely and Efficiently · OSDI 2018
Distributed systems
fault tolerance
0.312018
Maelstrom: Mitigating Datacenter-level Disasters by Draining Interdependent Traffic Safely and Efficiently · OSDI 2018
Machine learning › Transfer learning and domain adaptation › zero-shot learning
zero-shot classification
0.212023
Open Vocabulary Semantic Segmentation with Patch Aligned Contrastive Learning · CVPR 2023
Machine learning › Transfer learning and domain adaptation › few-shot learning
few-shot adaptation
0.212022
Few-Shot Fast-Adaptive Anomaly Detection · NeurIPS 2022
Machine learning › Transfer learning and domain adaptation
meta-learning
0.212022
Few-Shot Fast-Adaptive Anomaly Detection · NeurIPS 2022

Methods — techniques the papers use, named apart from their topics

text anchoring · 1.5spherical linear interpolation · 1.5memory bank · 0.8large language model · 0.8contrastive learning · 0.7CLIP · 0.7unsupervised learning · 0.6object-centric representation learning · 0.6langevin dynamics · 0.6energy-based model · 0.6
YearPublicationVenuePosition
2024 MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding
abstract
With the success of large language models (LLMs), integrating the vision model into LLMs to build vision-language foundation models has gained much more interest recently. However, existing LLM-based large multimodal models (e.g., Video-LLaMA, VideoChat) can only take in a limited number of frames for short video understanding. In this study, we mainly focus on designing an efficient and effective model for long-term video understanding. Instead of trying to process more frames simultaneously like most existing work, we propose to process videos in an online manner and store past video information in a memory bank. This allows our model to reference historical video content for long-term analysis without exceeding LLMs' context length constraints or GPU memory limits. Our memory bank can be seamlessly integrated into current multimodal LLMs in an off-the-shelf manner. We conduct extensive experiments on various video understanding tasks, such as long-video understanding, video question answering, and video captioning, and our model can achieve state-of-the-art performances across multiple datasets.
Bo He 0004, Hengduo Li, Young Kyun Jang, Menglin Jia, Xuefei Cao, Ashish Shah, Abhinav Shrivastava, Ser-Nam Lim
CVPR6
2024 Spherical Linear Interpolation and Text-Anchoring for Zero-Shot Composed Image Retrieval
Young Kyun Jang, Dat Huynh, Ashish Shah, Wen-Kai Chen, Ser-Nam Lim
ECCV (19)3
2023 Open Vocabulary Semantic Segmentation with Patch Aligned Contrastive Learning
abstract
We introduce Patch Aligned Contrastive Learning (PACL), a modified compatibility function for CLIP's contrastive loss, intending to train an alignment between the patch tokens of the vision encoder and the CLS token of the text encoder. With such an alignment, a model can identify regions of an image corresponding to a given text input, and therefore transfer seamlessly to the task of open vocabulary semantic segmentation without requiring any segmentation annotations during training. Using pre-trained CLIP encoders with PACL, we are able to set the state-of-the-art on the task of open vocabulary zero-shot segmentation on 4 different segmentation benchmarks: Pascal VOC, Pascal Context, COCO Stuff and ADE20K. Furthermore, we show that PACL is also applicable to image-level predictions and when used with a CLIP backbone, provides a general improvement in zero-shot classification accuracy compared to CLIP, across a suite of 12 image classification datasets.
Jishnu Mukhoti, Tsung-Yu Lin, Omid Poursaeed, Rui Wang 0043, Ashish Shah, Philip Torr 0001, Ser-Nam Lim
CVPR5
2022 Object-Centric Unsupervised Image Captioning
Zihang Meng, Xuefei Cao, Ashish Shah, Ser-Nam Lim
ECCV (36)4
2022 Few-Shot Fast-Adaptive Anomaly Detection
abstract
The ability to detect anomaly has long been recognized as an inherent human ability, yet to date, practical AI solutions to mimic such capability have been lacking. This lack of progress can be attributed to several factors. To begin with, the distribution of ``abnormalities'' is intractable. Anything outside of a given normal population is by definition an anomaly. This explains why a large volume of work in this area has been dedicated to modeling the normal distribution of a given task followed by detecting deviations from it. This direction is however unsatisfying as it would require modeling the normal distribution of every task that comes along, which includes tedious data collection. In this paper, we report our work aiming to handle these issues. To deal with the intractability of abnormal distribution, we leverage Energy Based Model (EBM). EBMs learn to associates low energies to correct values and higher energies to incorrect values. At its core, the EBM employs Langevin Dynamics (LD) in generating these incorrect samples based on an iterative optimization procedure, alleviating the intractable problem of modeling the world of anomalies. Then, in order to avoid training an anomaly detector for every task, we utilize an adaptive sparse coding layer. Our intention is to design a plug and play feature that can be used to quickly update what is normal during inference time. Lastly, to avoid tedious data collection, this mentioned update of the sparse coding layer needs to be achievable with just a few shots. Here, we employ a meta learning scheme that simulates such a few shot setting during training. We support our findings with strong empirical evidence.
Yipin Zhou, Rui Wang 0043, Tsung-Yu Lin, Ashish Shah, Ser-Nam Lim
NeurIPS5
2018 Maelstrom: Mitigating Datacenter-level Disasters by Draining Interdependent Traffic Safely and Efficiently
Kaushik Veeraraghavan, Justin Meza, Scott Michelson, Sankaralingam Panneerselvam, Alex Gyori, Sonia Margulis, Daniel Obenshain, Shruti Padmanabha, Ashish Shah, Yee Jiun Song, Tianyin Xu
OSDI10
1997 A single chip radio transceiver for DECT
abstract
A single chip radio transceiver for the DECT (Digital Enhanced Cordless Telephone) system has been implemented using Si 0.6 /spl mu/m BiCMOS technology. The device implements Rx down-converting mixers, a limiting IF strip, second down-converting image rejecting mixer, PLL demodulator with S-field sample and hold, bit slicer, synthesiser, UHF VCO, frequency doubler and transmit buffer. Synthesiser programming and the chip power management is achieved via the programming of an integrated 3 wire serial port. The chip also includes two separate supply voltage regulators, one of which is dedicated for use by the single UHF VCO.
Simon Atkinson, Ashish Shah, Jon Strange
PIMRC2
1996 The generalised maximum SINR array processor for personal communication systems in a multipath environment
abstract
This paper presents the generalised maximum signal-to-interference-plus-noise ratio (MSINR) array processor, which is the extension of MSINR from the single to the multiple desired-signal case. The problem of optimal multiple desired signal processing is of particular importance in a multipath fading environment. Therefore, the performance of the MSINR processor is evaluated for personal communication systems (PCS) in a Rayleigh fading channel with interference. Strong interference is expected in the PCS channel, due to the overlay in the spectrum allocated to PCS. It is shown that the generalised MSINR processor outperforms other methods the minimum variance distortionless response (MVDR) and the minimum mean square error (MMSE) interference cancellers, yielding the best bit error rate (BER).
Dignus-Jan Moelker, Ashish Shah, Yeheskel Bar-Ness
PIMRC2
1995 Chirico-a framework for computerization of medical practice guidelines
abstract
Methodologies based on an elaboration of the knowledge acquisition and design structuring (KADS) philosophy were developed and a suite of tools based on this methodology was implemented. The tools implement object oriented support at the domain-layer, Bayesian reasoning combined with a Bayesian compatible version of "fuzzy sets" to support the inference-layer, and mechanisms for knowledge level task control to support both the task-level and strategic-level. These tools were built for the task of computerizing medical practice guidelines. The tools were successfully applied to two practice guidelines, one selected by symptom (febrile neutropenia) and the other by drug class (CSF).
Clifton Davis, Christoph F. Eick, Balasubramanian Krishnamurthy, Ashish Shah, Lee Wanke
ICTAI4