Robert Martin

dblp:71/3286 · DBLP profile ↗
← Back
10ranked-venue papers
1as first author
2since 2021 · last 2023
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 1 since 2021Databases, data management, data science and information retrieval · 3Systems, architecture and hardware · 2 · 1 first-authorComputer networks · 2Software engineering, systems software and programming languages · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
3 papers
Bioinformatics and computational biology · 58% Medical and health informatics · 21% Computational science and engineering · 21%
Network and information security
2 papers
Hardware security and side channels · 80% Privacy and data protection · 20%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
Performance modeling and evaluation · 62% Memory systems · 38%

Topics — the 12 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology › biomedical text mining
biomedical entity linking
0.712023
BELB: a biomedical entity linking benchmark · Bioinform. 2023
Bioinformatics and computational biology
biomedical text mining
0.712023
BELB: a biomedical entity linking benchmark · Bioinform. 2023
Medical and health informatics
biomedical natural language processing
0.612022
BigBio: A Framework for Data-Centric Biomedical Natural Language Processing · NeurIPS 2022
Computational science and engineering
multi-task learning
0.612022
BigBio: A Framework for Data-Centric Biomedical Natural Language Processing · NeurIPS 2022
Hardware security and side channels
microarchitectural side channel
0.322012
TimeWarp: Rethinking timekeeping and performance monitoring mechanisms to mitigate side-channel attacks · ISCA 2012
Side-channel vulnerability factor: A metric for measuring information leakage · ISCA 2012
Privacy and data protection › information leakage
information leakage metric
0.112012
Side-channel vulnerability factor: A metric for measuring information leakage · ISCA 2012
Hardware security and side channels
side-channel attack
0.112012
Side-channel vulnerability factor: A metric for measuring information leakage · ISCA 2012
Hardware security and side channels
side-channel countermeasures
0.112012
TimeWarp: Rethinking timekeeping and performance monitoring mechanisms to mitigate side-channel attacks · ISCA 2012
Performance modeling and evaluation
performance monitoring
0.112012
TimeWarp: Rethinking timekeeping and performance monitoring mechanisms to mitigate side-channel attacks · ISCA 2012
Bioinformatics and computational biology
computational neuroscience
0.112008
Dependence of Orientation Tuning on Recurrent Excitation and Inhibition in a Network Model of V1 · NIPS 2008
Memory systems
cache side channel
0.012012
Side-channel vulnerability factor: A metric for measuring information leakage · ISCA 2012
Memory systems
on-chip memory
0.012012
Side-channel vulnerability factor: A metric for measuring information leakage · ISCA 2012

Methods — techniques the papers use, named apart from their topics

rule-based entity linking · 0.7pre-trained language model · 0.7language prompting · 0.6instruction tuning · 0.6timekeeping limitation · 0.3side-channel vulnerability factor · 0.3performance counter modification · 0.3parameter fitting · 0.1hodgkin-huxley neuron model · 0.1
YearPublicationVenuePosition
2023 BELB: a biomedical entity linking benchmark
abstract
MOTIVATION: Biomedical entity linking (BEL) is the task of grounding entity mentions to a knowledge base (KB). It plays a vital role in information extraction pipelines for the life sciences literature. We review recent work in the field and find that, as the task is absent from existing benchmarks for biomedical text mining, different studies adopt different experimental setups making comparisons based on published numbers problematic. Furthermore, neural systems are tested primarily on instances linked to the broad coverage KB UMLS, leaving their performance to more specialized ones, e.g. genes or variants, understudied. RESULTS: We therefore developed BELB, a biomedical entity linking benchmark, providing access in a unified format to 11 corpora linked to 7 KBs and spanning six entity types: gene, disease, chemical, species, cell line, and variant. BELB greatly reduces preprocessing overhead in testing BEL systems on multiple corpora offering a standardized testbed for reproducible experiments. Using BELB, we perform an extensive evaluation of six rule-based entity-specific systems and three recent neural approaches leveraging pre-trained language models. Our results reveal a mixed picture showing that neural approaches fail to perform consistently across entity types, highlighting the need of further studies towards entity-agnostic models. AVAILABILITY AND IMPLEMENTATION: The source code of BELB is available at: https://github.com/sg-wbi/belb. The code to reproduce our experiments can be found at: https://github.com/sg-wbi/belb-exp.
Samuele Garda, Leon Weber-Genzel, Robert Martin, Ulf Leser
Bioinform.3
2022 BigBio: A Framework for Data-Centric Biomedical Natural Language Processing
abstract
Training and evaluating language models increasingly requires the construction of meta-datasets -- diverse collections of curated data with clear provenance. Natural language prompting has recently lead to improved zero-shot generalization by transforming existing, supervised datasets into a variety of novel instruction tuning tasks, highlighting the benefits of meta-dataset curation. While successful in general-domain text, translating these data-centric approaches to biomedical language modeling remains challenging, as labeled biomedical datasets are significantly underrepresented in popular data hubs. To address this challenge, we introduce BigBio a community library of 126+ biomedical NLP datasets, currently covering 13 task categories and 10+ languages. BigBio facilitates reproducible meta-dataset curation via programmatic access to datasets and their metadata, and is compatible with current platforms for prompt engineering and end-to-end few/zero shot language model evaluation. We discuss our process for task schema harmonization, data auditing, contribution guidelines, and outline two illustrative use cases: zero-shot evaluation of biomedical prompts and large-scale, multi-task learning. BigBio is an ongoing community effort and is available at https://github.com/bigscience-workshop/biomedical
Jason Alan Fries, Leon Weber-Genzel, Natasha Seelam, Gabriel Altay, Debajyoti Datta, Samuele Garda, Sunny Kang, Rosaline Su, Wojciech Kusa, Samuel Cahyawijaya, Fabio Barth, Simon Ott, Matthias Samwald, Stephen H. Bach, Stella Biderman, Mario Sänger, Bo Wang 0044, Alison Callahan, Daniel León Periñán, Théo Gigant, Patrick Haller 0002, Jenny Chim, José D. Posada, John M. Giorgi, Karthik Rangasai Sivaraman, Marc Pàmies, Marianna Nezhurina, Robert Martin, Michael Cullan, Moritz Freidank, Nathan Dahlberg, Shubhanshu Mishra, Shamik Bose, Nicholas Broad, Yanis Labrak, Shlok Deshmukh, Sid Kiblawi, Ayush Singh, Minh Chien Vu, Trishala Neeraj, Jonas Golde, Albert Villanova del Moral, Benjamin Beilharz
NeurIPS28
2017 Reuters tracer: Toward automated news production using large scale social media data
abstract
To deal with the sheer volume of information and gain competitive advantage, the news industry has started to explore and invest in news automation. In this paper, we present Reuters Tracer, a system that automates end-to-end news production using Twitter data. It is capable of detecting, classifying, annotating, and disseminating news in real time for Reuters journalists without manual intervention. In contrast to other similar systems, Tracer is topic and domain agnostic. It has a bottom-up approach to news detection, and does not rely on a predefined set of sources or subjects. Instead, it identifies emerging conversations from 12+ million tweets per day and selects those that are news-like. Then, it contextualizes each story by adding a summary and a topic to it, estimating its newsworthiness, veracity, novelty, and scope, and geotags it. Designing algorithms to generate news that meets the standards of Reuters journalists in accuracy and timeliness is quite challenging. But Tracer is able to achieve competitive precision, recall, timeliness, and veracity on news detection and delivery. In this paper, we reveal our key algorithm designs and evaluations that helped us achieve this goal, and lessons learned along the way.
Xiaomo Liu, Armineh Nourbakhsh, Quanzhi Li, Sameena Shah, Robert Martin, John Duprey
IEEE BigData5
2016 Reuters Tracer: A Large Scale System of Detecting & Verifying Real-Time News Events from Twitter
abstract
News professionals are facing the challenge of discovering news from more diverse and unreliable information in the age of social media. More and more news events break on social media first and are picked up by news media subsequently. The recent Brussels attack is such an example. At Reuters, a global news agency, we have observed the necessity of providing a more effective tool that can help our journalists to quickly discover news on social media, verify them and then inform the public.
Xiaomo Liu, Quanzhi Li, Armineh Nourbakhsh, Merine Thomas, Kajsa Anderson, Russ Kociuba, Mark Vedder, Steven Pomerville, Ramdev Wudali, Robert Martin, John Duprey, Arun Vachher, William Keenan, Sameena Shah
CIKM11
2015 TR Discover: A Natural Language Interface for Querying and Analyzing Interlinked Datasets
Dezhao Song, Frank Schilder, Charese Smiley, Chris Brew, Tom Zielund, Hiroko Bretz, Robert Martin, Chris Dale, John Duprey, Johanna Harrison
ISWC (2)7
2012 Side-channel vulnerability factor: A metric for measuring information leakage
abstract
There have been many attacks that exploit side-effects of program execution to expose secret information and many proposed countermeasures to protect against these attacks. However there is currently no systematic, holistic methodology for understanding information leakage. As a result, it is not well known how design decisions affect information leakage or the vulnerability of systems to side-channel attacks. In this paper, we propose a metric for measuring information leakage called the Side-channel Vulnerability Factor (SVF). SVF is based on our observation that all side-channel attacks ranging from physical to microarchitectural to software rely on recognizing leaked execution patterns. SVF quantifies patterns in attackers' observations and measures their correlation to the victim's actual execution patterns and in doing so captures systems' vulnerability to side-channel attacks. In a detailed case study of on-chip memory systems, SVF measurements help expose unexpected vulnerabilities in whole-system designs and shows how designers can make performance-security trade-offs. Thus, SVF provides a quantitative approach to secure computer architecture.
John Demme, Robert Martin, Adam Waksman, Simha Sethumadhavan
ISCA2
2012 TimeWarp: Rethinking timekeeping and performance monitoring mechanisms to mitigate side-channel attacks
abstract
Over the past two decades, several microarchitectural side channels have been exploited to create sophisticated security attacks. Solutions to this problem have mainly focused on fixing the source of leaks either by limiting the flow of information through the side channel by modifying hardware, or by refactoring vulnerable software to protect sensitive data from leaking. These solutions are reactive and not preventative: while the modifications may protect against a single attack, they do nothing to prevent future side channel attacks that exploit other microarchitectural side channels or exploit the same side channel in a novel way. In this paper we present a general mitigation strategy that focuses on the infrastructure used to measure side channel leaks rather than the source of leaks, and thus applies to all known and unknown microarchitectural side channel leaks. Our approach is to limit the fidelity of fine grain timekeeping and performance counters, making it difficult for an attacker to distinguish between different microarchitectural events, thus thwarting attacks. We demonstrate the strength of our proposed security modifications, and validate that our changes do not break existing software. Our proposed changes require minor - or in some cases, no - hardware modifications and do not result in any substantial performance degradation, yet offer the most comprehensive protection against microarchitectural side channels to date.
Robert Martin, John Demme, Simha Sethumadhavan
ISCA1
2009 Modeling of a Public Safety Communication System for Emergency Response
abstract
This paper describes simulation work in modeling a city-wide trunked radio system and evaluating its performance under stress. Specifically, traffic models are developed to simulate both routine background traffic and the traffic load associated with an emergency scenario of a high-rise apartment fire. An analysis of a one-month radio data log provided insights on the traffic distribution among different talk-groups, as well as building individual traffic loading profiles for the talk-groups directly involved in the emergency response. The performance of the radio system is evaluated under different traffic loading intensity levels. Initial simulation results show that this emergency response communication system provides adequate capacity even under intense traffic loading. However, one area of concern is the build-up of waiting calls within some heavily loaded talk-group, which results in long waiting time to access the channel.
Chao Chen 0001, Carlos A. Pomalaza-Raez, M. Colone, Robert Martin, Jim Isaacs
ICC4
2008 Dependence of Orientation Tuning on Recurrent Excitation and Inhibition in a Network Model of V1
abstract
One major role of primary visual cortex (V1) in vision is the encoding of the orientation of lines and contours. The role of the local recurrent network in these computations is, however, still a matter of debate. To address this issue, we analyze intracellular recording data of cat V1, which combine measuring the tuning of a range of neuronal properties with a precise localization of the recording sites in the orientation preference map. For the analysis, we consider a network model of Hodgkin-Huxley type neurons arranged according to a biologically plausible two-dimensional topographic orientation preference map. We then systematically vary the strength of the recurrent excitation and inhibition relative to the strength of the afferent input. Each parametrization gives rise to a different model instance for which the tuning of model neurons at different locations of the orientation map is compared to the experimentally measured orientation tuning of membrane potential, spike output, excitatory, and inhibitory conductances. A quantitative analysis shows that the data provides strong evidence for a network model in which the afferent input is dominated by strong, balanced contributions of recurrent excitation and inhibition. This recurrent regime is close to a regime of 'instability', where strong, self-sustained activity of the network occurs. The firing rate of neurons in the best-fitting network is particularly sensitive to small modulations of model parameters, which could be one of the functional benefits of a network operating in this particular regime.
Klaus Wimmer 0002, Marcel Stimberg, Robert Martin, Lars Schwabe, Jorge Mariño, James Schummers, David C. Lyon, Mriganka Sur, Klaus Obermayer
NIPS3
2006 Robot teleoperation featuring commercially available wireless network cards
Elizabeth A. Thompson, Eric Harmison, Robert Carper, Robert Martin, Jim Isaacs
J. Netw. Comput. Appl.4