VLDB 2026 Research / reviewers in the wild / expert
Daniel Citron
dblp:02/4975
· DBLP profile ↗
9ranked-venue papers
4as first author
0since 2021 · last 2013
0000-0002-1282-4176ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 3 first-authorSoftware engineering, systems software and programming languages · 4 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Performance modeling and evaluation · 62% Processor architecture and microarchitecture · 27% Memory systems · 9% | |
| Computer graphics and multimedia
1 paper |
Multimedia systems and quality of experience · 100% |
Topics — the 5 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Performance modeling and evaluation
benchmarking |
0.0 | 1 | 2003 | MisSPECulation: Partial and Misleading Use of SPEC CPU2000 in Computer Architecture Conferences · ISCA 2003 |
Performance modeling and evaluation
workload characterization |
0.0 | 1 | 2003 | MisSPECulation: Partial and Misleading Use of SPEC CPU2000 in Computer Architecture Conferences · ISCA 2003 |
Processor architecture and microarchitecture › microprocessor design › processor core design
functional units |
0.0 | 1 | 1998 | Accelerating Multi-Media Processing by Implementing Memoing in Multiplication and Division Units · ASPLOS 1998 |
Performance modeling and evaluation › system-level analysis › architecture evaluation
processor performance evaluation |
0.0 | 1 | 2003 | MisSPECulation: Partial and Misleading Use of SPEC CPU2000 in Computer Architecture Conferences · ISCA 2003 |
Interconnection networks and networks-on-chip › bus-based interconnection
bus-based communication |
0.0 | 1 | 1995 | Creating a Wider Bus Using Caching Techniques · HPCA 1995 |
Methods — techniques the papers use, named apart from their topics
memoing · 0.0lookup table · 0.0statistical projection · 0.0amdahl's law analysis · 0.0simulation · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2013 | Evaluating the FITTEST Automated Testing Tools: An Industrial Case StudyabstractThis paper aims at evaluating a set of automated tools of the FITTEST EU project within an industrial case study. The case study was conducted at the IBM Research lab in Haifa, by a team responsible for building the testing environment for future development versions of an IBM system management product. The main function of that product is resource management in a networked environment. This case study has investigated whether current IBM Research testing practices could be improved or complemented by using some of the automated testing tools that were developed within the FITTEST EU project. Although the existing Test Suite from IBM Research (TSibm) that was selected for comparison is substantially smaller than the Test Suite generated by FITTEST (TSfittest), the effectiveness of TSfittest, measured by the injected faults coverage is significantly higher (50% vs 70%). With respect to efficiency, by normalizing the execution times, we found the TSfittest runs faster (9.18 vs. 6.99). This is due to the fact that the TSfittest includes shorter tests. Within IBM Research and for the testing of the target product in the simulated environment: the FITTEST tools can increase the effectiveness of the current practice and the test cases automatically generated by the FITTEST tools can help in more efficient identification of the source of the identified faults. Moreover, the FITTEST tools have shown the ability to automate testing within a real industry case. Duy Cu Nguyen, Bilha Mendelson, Daniel Citron, Onn Shehory, Tanja E. J. Vos, Nelly Condori-Fernández |
ESEM | 3 |
| 2008 | Aggressive Function Inlining: Preventing Loop Blockings in the Instruction Cache
Yosi Ben-Asher, Omer Boehm, Daniel Citron, Gadi Haber, Moshe Klausner, Roy Levin, Yousef Shajrawi |
HiPEAC | 3 |
| 2005 | Instrumenting annotated programsabstractInstrumentation is commonly used to track application behavior: to collect program profiles; to monitor component health and performance; to aid in component testing; and more. Program annotation enables developers and tools to pass extra information to later stages of software development and execution. For example, the .NET runtime relies on annotations for a significant chunk of the services it provides. Both mechanisms are evolving into important parts of software development %, in the context of modern platforms such as Java and .NET.Instrumentation tools are generally not aware of the semantics of information passed via the annotation mechanism. This is especially true for post-compiler, e.g., run-time, instrumentation. The problem is that instrumentation may affect the correctness of annotations, rendering them invalid or misleading, and producing unforeseen side-effects during program execution. This problem has not been addressed so far.In this paper, we show the subtle interaction that takes place between annotations and instrumentation using several real-life examples. Many annotations are intended to provide information for the runtime; the virtual environment is a prominent annotation consumer, and must be aware of this conflict. It may also be required to provide runtime support to other annotation consumers. We propose an annotation taxonomy and show how instrumentation affects various annotations that were used in research and in industrial applications. We show how the annotations can expose enough information about themselves to prevent the instrumentation from accidentally corrupting the annotations. We demonstrate this approach on our annotations benchmark. Marina Biberstein, Vugranam C. Sreedhar, Bilha Mendelson, Daniel Citron, Alberto Giammaria |
VEE | 4 |
| 2004 | Reducing program image size by extracting frozen code and dataabstractConstraints on the memory size of embedded systems require reducing the image size of executing programs. Common techniques include code compression and reduced instruction sets. We propose a novel technique that eliminates large portions of the executable image without compromising execution time (due to decompression) or code generation (due to reduced instruction sets). Frozen code and data portions are identified using profiling techniques and removed from the loadable image. They are replaced with branches to code stubs that load them in the unlikely case that they are accessed. The executable is sustained in a runnable mode.Analysis of the frozen portions reveals that most are error and uncommon input handlers. Only a minority of the code (less than 1%) that was identified as frozen during a training run, is also accessed with production datasets.The technique was applied on three benchmark suites (SPEC CINT2000, SPEC CFP2000, and MediaBench) and results in image size reductions of up to 73%, 92%, and 85% per suite, The average reductions are 59%, 79%, and 78% per suite. Daniel Citron, Gadi Haber, Roy Levin |
EMSOFT | 1 |
| 2004 | Overlapping Memory Operations with Circuit Evaluation in Reconfigurable ComputingabstractSummary form only given. We consider the problem of compiling programs, written in a general high-level programming language, into hardware circuits executed by an FPGA (field programmable gate array) unit. In particular, we consider the problem of synthesizing nested loops that frequently access array elements stored in an external memory (outside the FPGA). We propose an aggressive compilation scheme, based on loop unrolling and code flattening techniques, where array references from/to the external memory are overlapped with uninterrupted hardware evaluation of the synthesized loop's circuit. We implement a restricted programming language called DOL based on the proposed compilation scheme and our experimental results provide preliminary evidence that aggressive compilation can be used to compile large code segments into circuits, including overlapping of hardware operations and memory references. Yosi Ben-Asher, Daniel Citron, Gadi Haber |
IPDPS | 2 |
| 2004 | The future of simulation: A field of dreamsabstractQuantitative evaluation of next-generation computer architectures and processor enhancements is possible only by running simulations. However, since the insights that are gained through simulation are predicated on the accuracy of the simulation results, and since the design decisions for future processor architectures -- which cost billions of dollars to design and implement -- are based on those insights, periodic examination of the simulation process becomes a necessity, rather than a luxury. Accordingly, this panel discusses the deficiencies of existing simulators, benchmarks, and simulation methodologies and techniques, and, in addition, what future directions are available for each. Brad Calder, Daniel Citron, Yale N. Patt, James E. Smith 0001 |
ISPASS | 2 |
| 2003 | MisSPECulation: Partial and Misleading Use of SPEC CPU2000 in Computer Architecture ConferencesabstractA majority of the papers published in leading computer architecture conferences use SPEC CPU2000, or its predecessor SPEC CPU95, which has become the de facto standard for measuring processor and/or memory-hierarchy performance. However, in most cases a subset of the suite’s benchmarks are simulated. For example: 27 papers were published in ISCA 2002, 16 used SPEC CINT2000, 4 used the whole suite, and only 3 papers explained their omissions. This paper quantifies the extent of this phenomenon in the ISCA, Micro, and HPCA conferences: 173 papers were surveyed, 115 used benchmarks from SPEC CINT, but only 23 used the whole suite. If this current trend continues, by the year 2005 80 % of the papers will use the full CINT2000 suite, a year after CPU2004 shall be announced. We claim that results based upon a subset of a benchmark suite are speculative and conflict with Amdahl’s Law. The law implies that we must present the speedup of using the proposed technique on the whole suite. Projecting the law (by statistically supplying values for the missing benchmarks) to several published papers reduces promising results to average ones. Speedups are reduced from 1.42 to 1.16 in one case, from 1.43 to 1.13 in another, and from 1.76 to 1.15 in a third. Finally, we have found that the disregard for CFP2000 is unwarranted in papers that explore the data cache domain, the suite displays a higher data cache miss rate than CINT2000, which is used more frequently. 1. Daniel Citron |
ISCA | 1 |
| 1998 | Accelerating Multi-Media Processing by Implementing Memoing in Multiplication and Division UnitsabstractThis paper proposes a technique that enables performing multi-cycle (multiplication, division, square-root …) computations in a single cycle. The technique is based on the notion of memoing: saving the input and output of previous calculations and using the output if the input is encountered again. This technique is especially suitable for Multi-Media (MM) processing. In MM applications the local entropy of the data tends to be low which results in repeated operations on the same datum.The inputs and outputs of assembly level operations are stored in cache-like lookup tables and accessed in parallel to the conventional computation. A successful lookup gives the result of a multi-cycle computation in a single cycle, and a failed lookup doesn't necessitate a penalty in computation time.Results of simulations have shown that on the average, for a modestly sized memo-table, about 40% of the floating point multiplications and 50% of the floating point divisions, in Multi-Media applications, can be avoided by using the values within the memo-table, leading to an average computational speedup of more than 20%. Daniel Citron, Dror G. Feitelson, Larry Rudolph |
ASPLOS | 1 |
| 1995 | Creating a Wider Bus Using Caching TechniquesabstractThe effective bandwidth of a bus and external communication ports can be increased by using a variant of data compression techniques that compacts words instead of data streams. The compaction is performed by caching the high order bits into a table and sending the index into the table along with the low order bits. A coherent table at the receiving end expands the word into it original form. Compaction/expansion units can be placed between processor and memory, between processor and local bus, and between devices that access the system bus. Simulations have shown that over 90% of all informative transferred can be sent in a single cycle when using a 32 bit processor connected by a 16 bit wide bus to a 32 bit memory module. This is for all forms of data, address, data, and instructions, and when a cache-based processor is used.> Daniel Citron, Larry Rudolph |
HPCA | 1 |