VLDB 2026 Research / reviewers in the wild / expert
Bogong Su
dblp:64/3897
· DBLP profile ↗
16ranked-venue papers
6as first author
1since 2021 · last 2021
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 4 first-authorGraphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-authorSoftware engineering, systems software and programming languages · 3Applied, interdisciplinary, general and emerging computing · 2Artificial intelligence and machine learning · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Software engineering, system software, and programming languages
6 papers |
Compilers and program optimization · 92% Program analysis · 8% | |
| Computer architecture, parallel and distributed computing, and storage systems
4 papers |
Processor architecture and microarchitecture · 100% |
Topics — the 10 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Compilers and program optimization › instruction scheduling
software pipelining |
0.0 | 4 | 1994 | A study of pointer aliasing for software pipelining using run-time disambiguation · MICRO 1994 GPMB - software pipelining branch-intensive loops · MICRO 1993 A VLIW architecture for optimal execution of branch-intensive loops · MICRO 1992 |
Compilers and program optimization
instruction scheduling |
0.0 | 4 | 1994 | A study of pointer aliasing for software pipelining using run-time disambiguation · MICRO 1994 Foresighted Instruction Scheduling Under Timing Constraints · IEEE Trans. Computers 1992 A VLIW architecture for optimal execution of branch-intensive loops · MICRO 1992 |
Processor architecture and microarchitecture
instruction-level parallelism |
0.0 | 4 | 1994 | A VLIW architecture for optimal execution of branch-intensive loops · MICRO 1992 A study of pointer aliasing for software pipelining using run-time disambiguation · MICRO 1994 Foresighted Instruction Scheduling Under Timing Constraints · IEEE Trans. Computers 1992 |
Compilers and program optimization
parallelization |
0.0 | 1 | 1994 | A study of pointer aliasing for software pipelining using run-time disambiguation · MICRO 1994 |
Compilers and program optimization › instruction scheduling
instruction-level parallelism |
0.0 | 1 | 1993 | GPMB - software pipelining branch-intensive loops · MICRO 1993 |
Compilers and program optimization › dynamic optimization
profile-guided optimization |
0.0 | 1 | 1993 | GPMB - software pipelining branch-intensive loops · MICRO 1993 |
Program analysis
static analysis |
0.0 | 1 | 1993 | GPMB - software pipelining branch-intensive loops · MICRO 1993 |
Processor architecture and microarchitecture › instruction-level parallelism
VLIW |
0.0 | 1 | 1992 | A VLIW architecture for optimal execution of branch-intensive loops · MICRO 1992 |
Compilers and program optimization › instruction scheduling
microcode compaction |
0.0 | 1 | 1983 | A Preliminary Evaluatin of Trace Scheduling for Global Microcode Compaction · IEEE Trans. Computers 1983 |
Compilers and program optimization › instruction scheduling
trace scheduling |
0.0 | 1 | 1983 | A Preliminary Evaluatin of Trace Scheduling for Global Microcode Compaction · IEEE Trans. Computers 1983 |
Methods — techniques the papers use, named apart from their topics
run-time alias disambiguation · 0.0lookahead scheduling · 0.0data dependency graph · 0.0global scheduling · 0.0static program analysis · 0.0branch prediction · 0.0empirical evaluation · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | Using machine learning techniques for DSP software performance prediction at source code levelabstractEfficient performance prediction at the source code level is essential in reducing the turnaround time of software development. In this paper, we introduce a new prediction model, which combines several machine learning algorithms, such as KNN, clustering, similarity, sample and attribute weighting with multiple linear regression techniques, to predict the execution time of Digital Signal Processing (DSP) software at the source code level. Prediction at source code level tends to both under-predict the performance for certain testing samples and over-predict for some other samples. Therefore, we propose a new algorithm called MAX/MIN algorithm to select the best-predicted execution time. To validate the new model, we measure experimentally the execution time of a set of functions selected from PHY DSP Benchmark and run them on TIC64 DSP processor. It is observed that the average absolute relative prediction error is less than 10% between the computed performance from the new model and the actual measured execution time. Erh-Wen Hu, Bogong Su, Jian Wang 0046 |
Connect. Sci. | 3 |
| 2017 | Software performance prediction at source levelabstractPerformance prediction is critical in embedded system design for reducing the turnaround time of software. Using simulation to measure the performance of the whole source code is often too slow, particularly after the modification of the source code due to changes in problem specification. In this paper we present a comprehensive method that combines analytical modeling and statistical approach to predicting the performance of application software at source code level. We take samples from EEMBC and SMV benchmarks and gather the static attributes from the source code of those samples as our learning set. To determine the effectiveness of our new approach, we select several functions from PHY Benchmark as our testing set. We then apply multiple linear regression technique enhanced with the inclusion of new approaches by using the popular statistical tool SPSS23 to predict the performance of these functions. Comparing our predicted results with the actual measured values, the outcome is promising as the average relative error is within 20%. Erh-Wen Hu, Bogong Su, Jian Wang 0046 |
SERA | 2 |
| 2003 | De-pipeline a software-pipelined loopabstractSoftware pipelining is a loop optimization technique that has been widely implemented in modem optimizing compilers. In order to utilize fully the instruction level parallelism of the recent VLIW DSP processors, DSP programs have to be optimized by software pipelining. However, because of the transformation of the original sequential code, a software-pipelined loop is often difficult to understand, test, and debug. It is also very difficult to reuse and port a software-pipelined loop to other processors, especially when the original sequential code is unavailable. We propose a de-pipelining technique, which converts the optimized assembly code of a software-pipelined loop back to a semantically equivalent sequential counterpart. Preliminary experiments on 20 assembly programs verifies the validity of the proposed de-pipelining algorithm. Bogong Su, Jian Wang 0046, Erh-Wen Hu, Joseph B. Manzano |
ICASSP (2) | 1 |
| 2000 | A scalable loop optimization approach for scalable DSP processorsabstractThis paper proposes the possibility of reuse of the existing optimized DSP code on a scalable high-performance VLIW DSP processor. Since loops are the critical paths in most DSP applications, we focus on issues related to loop optimization. In our approach, we first perform a loop alignment transformation on the source level; we then reuse the existing optimized loop code on the assembly level. The approach is highly portable because it is independent of DSP hardware details. It can be used directly by a DSP programmer on the source level and/or by a DSP compiler designer to implement independent optimization modules. Jian Wang 0046, Bogong Su, Erh-Wen Hu |
ICASSP | 2 |
| 1999 | Source-level loop optimization for DSP code generationabstractThe performance of current C compilers for DSP is almost unacceptable. One of the most important reasons is the lack of implementing software pipelining. This paper presents a remedy called source-level loop optimization. DSP programmers can use source-level loop optimization first then input its result to the DSP compiler to obtain better assembly code. The implementation of source-level loop optimization is easier than that of software pipelining. The preliminary result with the DSP compiler-challenge C code shows that source-level loop optimization is a portable and efficient approach. The detailed method and working examples are presented. Bogong Su, Jian Wang 0046, Andrew Esguerra |
ICASSP | 1 |
| 1998 | Software pipelining of nested loops for real-time DSP applicationsabstractModern DSP processors have been integrated with instruction-level parallelism (lLP), which presents a challenge to exploit ILP within DSP applications. Software pipelining is an efficient technique used to expose ILP for loop programs and has been widely used for current microprocessors. It has been also used in DSP compilers, but only for the innermost loops. This paper proposes a new approach which extends software pipelining from innermost loops to whole nested loops in DSP applications. Given a perfect loop, we apply an existing software pipelining approach for the innermost loops, then use the so-called pipelining-dovetailing transformation to extend software pipelining to the outer loops. We also present a transformation to convert a non-perfect nested loop into a perfect one. We have verified the above transformations with some nested loops selected from DSP compiler-challenge C code. The preliminary results are further presented in this paper. Jian Wang 0046, Bogong Su |
ICASSP | 2 |
| 1998 | Building a Retargetable Local Instruction SchedulerabstractWhile high-performance architectures have included some Instruction-Level Parallelism (ILP) for at least 25 years, recent computer designs have exploited ILP to a significant degree. Although a local scheduler is not sufficient for generation of excellent ILP code, it is necessary as many global scheduling and software pipelining techniques rely on a local scheduler. Global scheduling techniques are well-documented, yet practical discussions of local schedulers are notable in their absence. This paper strives to remedy that disparity by describing a list scheduling framework and several important practical details that, taken together, allow implementation of an efficient local instruction scheduler that is easily retargetable for ILP architectures. The foundation of our machine-independent instruction scheduler is a timing model that allows easy retargetability to a wide range of architectures. In addition to describing how a general list-scheduler can be implemented within the framework of our timing model, experimental results indicate that lookahead scheduling can profoundly improve a scheduler's ability to produce a legal schedule. Further experimental data shows that deciding to schedule a data dependence DAG (DDD) in forward or reverse order depends significantly upon that target architecture, suggesting the possibility of scheduling in each direction and using the best of the two schedules. In contrast, experiments demonstrate little difference in code quality for schedules generated by either instruction-driven or operation-driven schedulers. Thus, the inherent flexibility of operation-driven methods suggests including that approach in a retargetable instruction scheduler. List scheduling is, of course, a heuristic scheduling method. A variety of scheduling heuristics are presented. In addition, the paper describes a method, using a genetic algorithm search, to ‘fine-tune’ the weights of twenty-four individual heuristics to form a DDD-node heuristic tuned to a specific architecture. © 1998 John Wiley & Sons, Ltd. Vicki H. Allan, Steven J. Beaty, Bogong Su, Philip H. Sweany |
Softw. Pract. Exp. | 3 |
| 1994 | A study of pointer aliasing for software pipelining using run-time disambiguationabstractRun-time alias disambiguation (RTD) has been proposed as a technique for pointer aliasing. This paper suggests several RTD approaches which may be used for DOACROSS scheduling to exploit coarse-grained parallelism. We analyze the rerollability problem in the transformation of those RTD approaches to software pipelining in order to exploit the instruction level parallelism available in loops. Finally, we give some suggestion as to how to address the rerollability problem. Bogong Su, Stanley Habib, Jian Wang 0046, Youfeng Wu |
MICRO | 1 |
| 1994 | Using timed Petri net to model instruction-level loop scheduling with resource constraints
Jian Wang 0046, Christine Eisenbeis, Bogong Su |
J. Comput. Sci. Technol. | 3 |
| 1993 | GPMB - software pipelining branch-intensive loopsabstractCompile-time code transformations which expose instruction-level parallelism (ILP) typically take into account the constraints imposed by all execution scenarios in the program. However, there are additional opportunities to increase ILP along some execution sequences if the constraints from alternative execution sequences can be ignored. Traditionally, profile information has been used to identify important execution sequences for aggressive compiler optimization and scheduling. The paper presents a set of static program analysis heuristics used in the IMPACT compiler to identify execution sequences for aggressive optimization. The authors show that the static program analysis heuristics identify execution sequences without hazardous conditions that tend to prohibit compiler optimizations. As a result, the static program analysis approach often achieves optimization results comparable to profile information in spite of its inferior branch prediction accuracies. This observation makes a strong case for using static program analysis with or without profile information to facilitate aggressive compiler optimization and scheduling.> Zhizhong Tang, Chihong Zhang, Bogong Su, Stanley Habib |
MICRO | 5 |
| 1993 | URPR-1: A single-chip VLIW architecture
Bogong Su, Jian Wang 0046, Zhizhong Tang, Chihong Zhang |
Microprocess. Microprogramming | 1 |
| 1992 | A VLIW architecture for optimal execution of branch-intensive loops
Bogong Su, Zhizhong Tang, Stanley Habib |
MICRO | 1 |
| 1992 | Foresighted Instruction Scheduling Under Timing ConstraintsabstractWhen data dependency graph arcs representing data dependency information are annotated with minimum and maximum timing information, new algorithms are required. Foresighted compaction is a list scheduling technique in which look ahead is used in making decisions. Foresighted compaction is very effective in reducing, failure inherent in greedy compaction algorithms.> Vicki H. Allan, Bogong Su, Pantung Wijaya, Jian Wang 0046 |
IEEE Trans. Computers | 2 |
| 1991 | Decomposition and allocation of flat-structured problemsabstractA classification is made of the applications of distributed problem solving (DPS) into hierarchically structured problems (HP) and flatly structured problems (FP). A formal description is given of an FP and its solving system. An improved problem decomposition (IPD) algorithm is presented by which a problem is decomposed in size into several tasks of identical property and each task is allocated to a proximate agent. Heuristic state space search is used to balance load. Experiments on a distributed transport dispatching system, DTD-1, indicate that this algorithm is effective and superior. > Bogong Su, Chunyi Shi |
COMPSAC | 2 |
| 1991 | GURPR*: A New Global Software Pipelining AlgorithmabstractArticle Free Access Share on GURPR*: a new global software pipelining algorithm Authors: Bogong Su The College of Staten Island of The City University of New York and Tsinghua University, Beijing, China The College of Staten Island of The City University of New York and Tsinghua University, Beijing, ChinaView Profile , Jian Wang Dept. of Computer Science and Technology, Tsinghua University, Beijing, China Dept. of Computer Science and Technology, Tsinghua University, Beijing, ChinaView Profile Authors Info & Claims MICRO 24: Proceedings of the 24th annual international symposium on MicroarchitectureSeptember 1991 Pages 212–216https://doi.org/10.1145/123465.123509Published:01 September 1991Publication History 38citation233DownloadsMetricsTotal Citations38Total Downloads233Last 12 Months7Last 6 weeks2 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF Bogong Su, Jian Wang 0046 |
MICRO | 1 |
| 1983 | A Preliminary Evaluatin of Trace Scheduling for Global Microcode CompactionabstractFisher has recently described a new procedure for global microcode compaction which he calls "trace scheduling." We have implemented this procedure and tested it on several microcode sequences. We report in this correspondence on the relative effectiveness of local compaction, manual compaction, and trace scheduling on these sequences. Ralph Grishman, Bogong Su |
IEEE Trans. Computers | 2 |