Ali Ansari 0001

dblp:200/9876-1 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
3since 2021 · last 2024
0000-0002-9798-6966ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 since 2021Systems, architecture and hardware · 2 · 2 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Processor architecture and microarchitecture · 75% Memory systems · 25%
Artificial intelligence
2 papers
Time series and sequential data · 44% Trustworthy machine learning · 44% Generative modeling · 13%
Network and information security
1 paper
Security and privacy of machine learning · 100%

Topics — the 10 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Processor architecture and microarchitecture › instruction fetch
instruction prefetching
1.122023
MANA: Microarchitecting a Temporal Instruction Prefetcher · IEEE Trans. Computers 2023
Divide and Conquer Frontend Bottleneck · ISCA 2020
Machine learning › Time series and sequential data
anomaly detection
0.812024
RODEO: Robust Outlier Detection via Exposing Adaptive Out-of-Distribution Samples · ICML 2024
Machine learning › Trustworthy machine learning › robustness
out-of-distribution detection
0.812024
Scanning Trojaned Models Using Out-of-Distribution Samples · NeurIPS 2024
Security and privacy of machine learning › adversarial attack › backdoor attack › backdoor defense
backdoor detection
0.812024
Scanning Trojaned Models Using Out-of-Distribution Samples · NeurIPS 2024
Memory systems
cache
0.712023
MANA: Microarchitecting a Temporal Instruction Prefetcher · IEEE Trans. Computers 2023
Processor architecture and microarchitecture
branch prediction
0.412020
Divide and Conquer Frontend Bottleneck · ISCA 2020
Processor architecture and microarchitecture › branch prediction
branch target buffer
0.412020
Divide and Conquer Frontend Bottleneck · ISCA 2020
Processor architecture and microarchitecture
front-end
0.412020
Divide and Conquer Frontend Bottleneck · ISCA 2020
Machine learning › Generative modeling › diffusion model
text-to-image generation
0.212024
RODEO: Robust Outlier Detection via Exposing Adaptive Out-of-Distribution Samples · ICML 2024
Memory systems › cache
cache miss reduction
0.212023
MANA: Microarchitecting a Temporal Instruction Prefetcher · IEEE Trans. Computers 2023

Methods — techniques the papers use, named apart from their topics

out-of-distribution sampling · 1.5adversarial perturbation · 1.5outlier exposure · 0.8adversarial training · 0.8storage cost reduction · 0.7metadata record design · 0.7
YearPublicationVenuePosition
2024 RODEO: Robust Outlier Detection via Exposing Adaptive Out-of-Distribution Samples
abstract
In recent years, there have been significant improvements in various forms of image outlier detection. However, outlier detection performance under adversarial settings lags far behind that in standard settings. This is due to the lack of effective exposure to adversarial scenarios during training, especially on unseen outliers, leading detection models failing to learn robust features. To bridge this gap, we introduce RODEO, a data-centric approach that generates effective outliers for robust outlier detection. More specifically, we show that incorporating outlier exposure (OE) and adversarial training could be an effective strategy for this purpose, as long as the exposed training outliers meet certain characteristics, including diversity, and both conceptual differentiability and analogy to the inlier samples. We leverage a text-to-image model to achieve this goal. We demonstrate both quantitatively and qualitatively that our adaptive OE method effectively generates ”diverse” and ”near-distribution” outliers, leveraging information from both text and image domains. Moreover, our experimental results show that utilizing our synthesized outliers significantly enhances the performance of the outlier detector, particularly in adversarial settings.
Hossein Mirzaei, Hamid Reza Dehbashi, Ali Ansari 0001, Sepehr Ghobadi, Masoud Hadi, Arshia Soltani Moakhar, Mohammad Azizmalayeri, Mahdieh Soleymani Baghshah, Mohammad H. Rohban
ICML4
2024 Scanning Trojaned Models Using Out-of-Distribution Samples
abstract
Scanning for trojan (backdoor) in deep neural networks is crucial due to their significant real-world applications. There has been an increasing focus on developing effective general trojan scanning methods across various trojan attacks. Despite advancements, there remains a shortage of methods that perform effectively without preconceived assumptions about the backdoor attack method. Additionally, we have observed that current methods struggle to identify classifiers trojaned using adversarial training. Motivated by these challenges, our study introduces a novel scanning method named TRODO (TROjan scanning by Detection of adversarial shifts in Out-of-distribution samples). TRODO leverages the concept of "blind spots"—regions where trojaned classifiers erroneously identify out-of-distribution (OOD) samples as in-distribution (ID). We scan for these blind spots by adversarially shifting OOD samples towards in-distribution. The increased likelihood of perturbed OOD samples being classified as ID serves as a signature for trojan detection. TRODO is both trojan and label mapping agnostic, effective even against adversarially trained trojaned classifiers. It is applicable even in scenarios where training data is absent, demonstrating high accuracy and adaptability across various scenarios and datasets, highlighting its potential as a robust trojan scanning strategy.
Hossein Mirzaei, Ali Ansari 0001, Bahar Dibaei Nia, Mojtaba Nafez, Moein Madadi, Sepehr Rezaee, Zeinab Taghavi 0001, Arad Maleki, Kian Shamsaie, Mahdi Hajialilue, Jafar Habibi, Mohammad Sabokrou, Mohammad H. Rohban
NeurIPS2
2023 MANA: Microarchitecting a Temporal Instruction Prefetcher
abstract
L1 instruction(L1-l) cache misses are a source of performance bottleneck. While many instruction prefetchers have been proposed, most of them leave a considerable potential uncovered. In 2011, Proactive Instruction Fetch (PIF) showed that a hardware prefetcher could effectively eliminate all instruction-cache misses. However, its enormous storage cost makes it impractical. Consequently, reducing the storage cost was the main research focus in instruction prefetching in the past decade. Several instruction prefetchers, including RDIP and Shotgun, were proposed to offer PIF-level performance with significantly lower storage overhead. However, our findings show that there is a considerable performance gap between these proposals and PIF. While these proposals use different mechanisms for prefetching, the performance gap is mainly not because of the mechanism, and instead, is due to not having sufficient storage. We make the case that the key to designing a powerful and cost-effective instruction prefetcher is choosing a metadata record and microarchitecting the prefetcher to minimize the storage. Our proposal, MANA, offers PIF-level performance with 15.7x lower storage cost. MANA outperforms RDIP and Shotgun by 12.5 and 29%, respectively. We also evaluate a version of MANA with no storage overhead and show that it offers 98% of the peak performance benefits.
Ali Ansari 0001, Fatemeh Golshan, Rahil Barati, Pejman Lotfi-Kamran, Hamid Sarbazi-Azad
IEEE Trans. Computers1
2020 Divide and Conquer Frontend Bottleneck
abstract
The frontend stalls caused by instruction and BTB misses are a significant source of performance degradation in server processors. Prefetchers are commonly employed to mitigate frontend bottleneck. However, next-line prefetchers, which are available in server processors, are incapable of eliminating a considerable number of L1 instruction misses. Temporal instruction prefetchers, on the other hand, effectively remove most of the instruction and BTB misses but impose significant area overhead.Recently, an old idea of using BTB-directed instruction prefetching is revived to address the limitations of temporal instruction prefetchers. While this approach leads to prefetchers with low area overhead, it requires significant changes to the frontend of a processor. Moreover, as this approach relies on the BTB content for prefetching, BTB misses stall the prefetcher, and likely lead to costly instruction misses. Especially as instruction misses are usually more expensive than BTB misses, the dependence of instruction prefetching to the BTB content is harmful to workloads with very large instruction footprints. Moreover, BTB-directed instruction prefetchers, as proposed in prior work, cannot be applied to variable-length ISAs.In this work, we showcase the harmful effects of making instruction prefetchers depend on the BTB content. Moreover, we divide the frontend bottleneck into three categories and use a divide-and-conquer approach to propose simple and effective solutions for each one. Sequential misses can be covered by an accurate and timely sequential prefetcher named SN4L, a lightweight discontinuity prefetcher named Dis eliminates discontinuity misses, and the BTB misses are reduced by pre-decoding the prefetched blocks. We also discuss how our proposal can be used for variable-length ISAs with low storage overhead. Our proposal, SN4L+ Dis+BTB, imposes the same area overhead as the state-of-the-art BTB-directed prefetcher, and at the same time, outperforms it by 5% on average and up to 16%.
Ali Ansari 0001, Pejman Lotfi-Kamran, Hamid Sarbazi-Azad
ISCA1