Dongming Jiang

dblp:97/4373 · DBLP profile ↗
← Back
17ranked-venue papers
8as first author
3since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 5 first-authorSoftware engineering, systems software and programming languages · 4 · 2 first-authorDatabases, data management, data science and information retrieval · 3 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Language models and text generation · 50% Knowledge representation and reasoning · 50%
Computer architecture, parallel and distributed computing, and storage systems
8 papers
Memory systems · 33% Performance modeling and evaluation · 20% Parallel and multicore computing · 17%
Computer networks
1 paper
Network optimization and economics · 100%
Theoretical computer science
1 paper
Mathematical optimization · 100%

Topics — the 17 heaviest of 23, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation › LLM agents
agent memory
1.012026
MAGMA: A Multi-Graph based Agentic Memory Architecture for AI Agents · ACL (1) 2026
Network optimization and economics › network economics
revenue sharing
0.112011
Efficient online ad serving in a display advertising exchange · WSDM 2011
Memory systems › virtual memory management
shared virtual memory
0.132001
Accelerating shared virtual memory via general-purpose network interface support · ACM Trans. Comput. Syst. 2001
Limits to the Performance of Software Shared Memory: A Layered Approach · HPCA 1999
Application Restructuring and Performance Portability on Shared Virtual Memory and Hardware-Coherent Multiprocessors · PPoPP 1997
Memory systems
cache coherence
0.021999
Scaling Application Performance on a Cache-Coherent Multiprocessors · ISCA 1999
A Methodology and an Evaluation of the SGI Origin2000 · SIGMETRICS 1998
Interconnection networks and networks-on-chip
cluster interconnect
0.011999
Limits to the Performance of Software Shared Memory: A Layered Approach · HPCA 1999
Memory systems › shared memory › distributed shared memory
software distributed shared memory
0.011999
Limits to the Performance of Software Shared Memory: A Layered Approach · HPCA 1999
Parallel and multicore computing
synchronization
0.011999
Evaluating Synchronization on Shared Address Space Multiprocessors: Methodology and Performance · SIGMETRICS 1999
Parallel and multicore computing › synchronization
synchronization mechanisms
0.011999
Evaluating Synchronization on Shared Address Space Multiprocessors: Methodology and Performance · SIGMETRICS 1999
High-performance computing › scientific visualization
parallel volume rendering
0.011997
Parallel Shear-Warp Volume Rendering on Shared Address Space Multiprocessors · PPoPP 1997
High-performance computing › performance engineering
performance portability
0.011997
Application Restructuring and Performance Portability on Shared Virtual Memory and Hardware-Coherent Multiprocessors · PPoPP 1997
Memory systems › cache coherence
cache-coherent shared memory
0.012001
Accelerating shared virtual memory via general-purpose network interface support · ACM Trans. Comput. Syst. 2001
High-performance computing
cluster computing
0.012001
Accelerating shared virtual memory via general-purpose network interface support · ACM Trans. Comput. Syst. 2001
High-performance computing › cluster computing
SMP cluster
0.012001
Accelerating shared virtual memory via general-purpose network interface support · ACM Trans. Comput. Syst. 2001
Performance modeling and evaluation
benchmarking
0.011999
Evaluating Synchronization on Shared Address Space Multiprocessors: Methodology and Performance · SIGMETRICS 1999
Processor architecture and microarchitecture › multiprocessor architecture
cache-coherent multiprocessor
0.011999
Evaluating Synchronization on Shared Address Space Multiprocessors: Methodology and Performance · SIGMETRICS 1999
Performance modeling and evaluation › benchmarking
microbenchmarking
0.011999
Evaluating Synchronization on Shared Address Space Multiprocessors: Methodology and Performance · SIGMETRICS 1999
Performance modeling and evaluation
parallel performance evaluation
0.011997
Parallel Shear-Warp Volume Rendering on Shared Address Space Multiprocessors · PPoPP 1997

Methods — techniques the papers use, named apart from their topics

policy-guided traversal · 1.0graph-based retrieval · 1.0precomputation · 0.4polynomial-time algorithm · 0.4candidate ordering · 0.4microbenchmarking · 0.0programmable network interface · 0.0asynchronous home-based LRC protocol · 0.0workload characterization · 0.0performance analysis · 0.0layered performance study · 0.0application restructuring · 0.0
YearPublicationVenuePosition
2026 MAGMA: A Multi-Graph based Agentic Memory Architecture for AI Agents
abstract
Memory-Augmented Generation (MAG) extends Large Language Models with external memory to support long-context reasoning, but existing approaches largely rely on semantic similarity over monolithic memory stores, entangling temporal, causal, and entity information.This design limits interpretability and alignment between query intent and retrieved evidence, leading to suboptimal reasoning accuracy.In this paper, we propose MAGMA, a multi-graph agentic memory architecture that represents each memory item across orthogonal semantic, temporal, causal, and entity graphs.MAGMA formulates retrieval as policy-guided traversal over these relational views, enabling query-adaptive selection and structured context construction.By decoupling memory representation from retrieval logic, MAGMA provides transparent reasoning paths and fine-grained control over retrieval.Experiments on LoCoMo and LongMemEval demonstrate that MAGMA consistently outperforms state-of-the-art agentic memory systems in long-horizon reasoning tasks.
Dongming Jiang, Guanpeng Li, Bingzhe Li
ACL (1)1
2026 Learning Where to Embed: Noise-Aware Positional Embedding for Query Retrieval in Small-Object Detection
abstract
Transformer-based detectors have advanced small-object detection, but they often remain inefficient and vulnerable to background-induced query noise, which motivates deep decoders to refine low-quality queries. We present HELP (Heatmap-guided Embedding Learning Paradigm), a noise-aware positional-semantic fusion framework that studies where to embed positional information by selectively preserving positional encodings in foreground-salient regions while suppressing background clutter. Within HELP, we introduce Heatmap-guided Positional Embedding (HPE) as the core embedding mechanism and visualize it with a heatbar for interpretable diagnosis and fine-tuning. HPE is integrated into both the encoder and decoder: it guides noise-suppressed feature encoding by injecting heatmap-aware positional encoding, and it enables high-quality query retrieval by filtering background-dominant embeddings via a gradient-based mask filter before decoding. To address feature sparsity in complex small targets, we integrate Linear-Snake Convolution to enrich retrieval-relevant representations. The gradient-based heatmap supervision is used during training only, incurring no additional gradient computation at inference. As a result, our design reduces decoder layers from eight to three and achieves a 59.4% parameter reduction (66.3M vs. 163M) while maintaining consistent accuracy gains under a reduced compute budget across benchmarks. Code Repository: https://github.com/yidimopozhibai/Noise-Suppressed-Query-Retrieval.
Yangchen Zeng, Zhenyu Yu, Dongming Jiang, Yifan Hong 0001, Zhanhua Hu, Jiao Luo, Kangning Cui
ICMR3
2026 CEMG: Collaborative-Enhanced Multimodal Generative Recommendation
Yuzhen Lin, Xuanjing Chen, Ivonne Xu, Dongming Jiang
MMM (1)6
2018 An iteration-based interactive analysis method to design dynamic service-oriented systems
abstract
Summary Service‐oriented paradigm presents numerous new software development patterns and idioms. Software systems are implemented by composing existing third‐party services deployed in the open environment, which is significantly different from traditional software development methodologies in which systems are built through developing modules after system design in a closed environment. Therefore, it is urgent to raise a new design method to adapt to this new circumstance. We concentrate on reusing as many deployed services as possible then introduce a new life cycle model named Taiji model to illustrate this development process. The iteration‐based interactive analysis method following the model is proposed to design service‐oriented systems based on the view of extracting non‐creative activities from a creative activity through defining new notations or applying new rules. The method includes the interactive analysis process that analyzes requirements with deployed services in a local point of view and the iterative analysis process that redesigns system with new knowledge in a global perspective. Meanwhile, the reusable service threshold value is defined to build the uncertain candidate service set (UCSS) of each module in analysis process. The reliability and flexibility of systems can be improved through the quantitative static structure analysis on the basis of the UCSS of systems. Meanwhile, a practical dynamic service binding method that selects services according to actual states of invoking them is presented on the basis of the UCSS containing them. Finally, we also give a case study to illustrate the feasibility of this method. Copyright © 2017 John Wiley & Sons, Ltd.
Wuping Xie, Jinyun Xue, Dongming Jiang, Lan Song
Softw. Pract. Exp.3
2013 A Novel Service Selection Based on Resource-Directive Decomposition
Dongming Jiang, Wuping Xie, Zhen You
WAIM1
2012 A reputation model based on hierarchical bayesian estimation for Web services
abstract
The motivation of Web service comes from its interoperational ability so a large number of Web services can interact with others and constitute an open network, Web service network. The success of Web services selection rely on, not only its Qos capability advertised, but the trustworthy of QoS to large degree. How to evaluate the trustworthy of services QoS information, however, is a challenge in Web service network. Reputation system, a mechanism which assesses the future QoS performance by the past behavior of service, is one of promising approaches to facilitate users make optimal decision. In this paper, we present a hybrid framework of reputation model for Web service. Based on this hybrid architecture, clients build their specific social communities, by which they obtain service's prior reputation. At the same time, the central reputation system fuses the rating data from clients by bayesian estimation. The result of experiments illustrated our approach is more efficiency and accuracy in several aspects, especially when dealing with strategics services.
Dongming Jiang, Jinyun Xue, Wuping Xie
CSCWD1
2011 Efficient online ad serving in a display advertising exchange
abstract
We introduce and formalize a novel constrained path optimization problem that is the heart of the real-time ad serving task in the Yahoo! (formerly RightMedia) Display Advertising Exchange. In the Exchange, the ad server's task for each display opportunity is to compute, with low latency, an optimal valid path through a directed graph representing the business arrangements between the hundreds of thousands of business entities that are participating in the Exchange. These entities include not only publishers and advertisers, but also intermediate entities called "ad networks" which have delegated their ad serving responsibilities to the Exchange. Path optimality is determined by the payment to the publisher, and is affected by an advertiser's bid and also by the revenue-sharing agreements between the entities in the chosen path leading back to the publisher. Path validity is determined by constraints which focus on the following three issues: 1) suitability of the opportunity's web page and its publisher 2)suitability of the user who is currently viewing that web page, and 3) suitability of a candidate ad and its advertiser. Because the Exchange's constrained path optimization task is novel, there are no published algorithms for it. This paper describes two different algorithms that have both been successfully used in the actual Yahoo! ad server. The first algorithm has the advantage of being extremely simple, while the second is more robust thanks to its polynomial worst-case running time. In both cases, meeting latency caps has required that the basic algorithms be improved by optimizations; we will describe a candidate ordering scheme and a pre-computation scheme that have both been effective in reducing latency in the real ad serving system that serves over ten billion ad calls per day.
Kevin J. Lang, Joaquin Delgado, Dongming Jiang, Bhaskar Ghosh, Shirshanka Das, Amita Gajewar, Swaroop Jagadish, Arathi Seshan, Chavdar Botev, Michael Ortega-Binderberger, Sunil Nagaraj, Raymie Stata
WSDM3
2003 Shared virtual memory clusters: bridging the cost-performance gap between SMPs and hardware DSM systems
Angelos Bilas, Dongming Jiang, Jaswinder Pal Singh
J. Parallel Distributed Comput.2
2001 Accelerating shared virtual memory via general-purpose network interface support
abstract
Clusters of symmetric multiprocessors (SMPs) are important platforms for high-performance computing. With the success of hardware cache-coherent distributed shared memory (DSM), a lot of effort has also been made to support the coherent shared-address-space programming model in software on clusters. Much research has been done in fast communication on clusters and in protocols for supporting software shared memory across them. However, the performance of software virtual memory (SVM) is still far from that achieved on hardware DSM systems. The goal of this paper is to improve the performance of SVM on system area network clusters by considering communication and protocol layer interactions. We first examine what are the important communication system bottlenecks that stand in the way of improving parallel performance of SVM clusters; in particular, which parameters of the communication architecture are most important to improve further relative to processor speed, which ones are already adequate on modern systems for most applications, and how will this change with technology in the future. We find that the most important communication subsystem cost to improve is the overhead of generating and delivery interrupts for asynchronous protocol processing. Then we proceed to show, that by providing simple and general support for asynchronous message handling in a commodity network interface (NI) and by altering SVM protocols appropriately, protocol activity can be decoupled from asynchronous message handling, and the need for interrupts or polling can be eliminated. The NI mechanisms needed are generic, not SVM-dependent. We prototype the mechanisms and such asynchronous home-basedLRCprotocol, calledGeNIMA(GEneral-purpose Network Interface support for shared Memory Abstractions), on a cluster of SMPs with a programmable NI. We find that the performance improvements are substantial, bringing performance on a small-scale SMP cluster much closer to that of hardware-coherent shared memory for many applications, and we show the value of each of the mechanisms in different applications.
Angelos Bilas, Dongming Jiang, Jaswinder Pal Singh
ACM Trans. Comput. Syst.2
1999 Limits to the Performance of Software Shared Memory: A Layered Approach
abstract
Much research has been done in fast communication on clusters and in protocols for supporting software shared memory across them. However, the end performance of applications that were written for the more proven hardware-coherent shared memory is still not very good on these systems. Three major layers of software (and hardware) stand between the end user and parallel performance, each with its own functionality and performance characteristics. They include the communication layer, the software protocol layer that supports the programming model, and the application layer. These layers provide a useful framework to identify the key remaining limitations and bottlenecks in software shared memory systems, as well as the areas where optimization efforts might yield the greatest performance improvements. This paper performs such an integrated study, using this layered framework, for two types of software distributed shared memory systems: page-based shared virtual memory (SVM) and fine-grained software systems (FG). For the two system layers (communication and protocol), we focus on the performance costs of basic operations in the layers rather than on their functionalities. This is possible because their functionalities are now fairly mature. The less mature applications layer is treated through application restructuring. We examine the layers individually and in combination, understanding their implications for the two types of protocols and exposing the synergies among layers.
Angelos Bilas, Dongming Jiang, Yuanyuan Zhou 0001, Jaswinder Pal Singh
HPCA2
1999 Application scaling under shared virtual memory on a cluster of SMPs
abstract
In this paper we examine how application performance scales on a state-of-the-art shared virtual memory (SVM) system on a cluster with 64 processors, comprising 4-way SMPs connected with a fast system area network.The protocol we use is home-based and takes advantage of general-purpose data movement and mutual exclusion support provided by a programmable network interface.We find that while the level of application restructuring needed is quite high compared to applications that perform well on a hardware-coherent system of this scale, and larger problem sizes are needed for good performance, SVM, surprisingly, performs quite well at the 64-processor scale for a fairly wide range of applications, achieving at least half the parallel efficiency of a high-end hardware-coherent system and often much more.We explore further application restructurings than those developed earlier for smaller-scale SVM systems, examine the main remaining system and application bottlenecks, and point out directions for future research.
Dongming Jiang, Brian O'Kelley, Angelos Bilas, Jaswinder Pal Singh
International Conference on Supercomputing1
1999 Scaling Application Performance on a Cache-Coherent Multiprocessors
abstract
Hardware-coherent, distributed shared address space systems are increasingly successful at moderate scale. However, it is unclear whether, or with how much difficulty, the performance of a load-store shared address space programming model scales to large processor counts on real applications. We examine this question using an aggressive case-study machine, the SGI Origin2000, up to 128 processors. We show for the first time that scalable performance can indeed be achieved in this programming model on a wide range of applications, including challenging kernels like FFT. However, this does not come easily, even for applications considered to be already highly optimized, and is very often not simply a matter of increasing problem size. Rather, substantial further application restructuring is often needed, which is usually quite algorithmic in nature. We examine how the restructurings compare with those needed for performance portability to shared virtual memory on clusters, and we comment on common programming guidelines for performance portability and scalability as well as on how the programming difficulty compares with that of explicit message passing. We also examine where applications spend their time on this large machine, the impact of special hardware features that the machine provides, and the impact of mapping to the network topology.
Dongming Jiang, Jaswinder Pal Singh
ISCA1
1999 Evaluating Synchronization on Shared Address Space Multiprocessors: Methodology and Performance
abstract
Synchronization is an area that exhibits rich hardware-software interactions in multiprocessors.It was studied extensively using microbenchmarks a decade ago.However, its performance implications are not well understood on modern systems or on real applications.We study the impact of synchronization primitives and algorithms on a modern, 64processor, hardware-coherent shared address space multiprocessor: the SGI Origin 2000.In addition to the actual results on a modern system, we examine the key methodological issues in studying synchronization, for both microbenchmarks and applications.We find that although the efficient hardware support (Fetch&Op) for synchronization provided on our machine usually helps lock and barrier microbenchmarks, it does not help in improving application performance when compared to good software algorithms that use the processor-provided LL-SC instructions.This is true even in applications that spend a significant amount of time in synchronization operations.More elaborate hardware support is unlikely to have a significant benefit either.From the applications' perspective, it is usually the waiting time due to load imbalance or serialization that dominates synchronization time, not the overhead of the synchronization operations themselves, even in apparently balanced cases where the overhead may be expected to be substantial.
Dongming Jiang, Rohit Chandra, Jaswinder Pal Singh
SIGMETRICS2
1998 Monitoring Shared Virtual Memory Performance on a Myrinet-based PC Cluster
abstract
Network-connected clusters of PCs or workstations are becoming a widespread parallel computing platform. Performance methodologies that use either simulation or high-level software instrumentation cannot adequately measure the detailed behavior of such systems. The availability of new network technologies based on programmable network interfaces opens a new avenue of research in analyzing and improving the performance of software shared memory protocols. We have developed monitoring firmware embedded in the programmable network interfaces of a Myrinet-based PC cluster. Timestamps on network packets facilitate the collection of low-level statistics on, e.g., network latencies, interrupt handler times and inter-node synchronization. This paper describes our use of the low-level software performance monitor to measure and understand the performance of a Shared Virtual Memory (SVM) system implemented on a Myrinetbased cluster, running the SPLASH-2 benchmarks. We measured time spent in vari...
Dongming Jiang, Liviu Iftode, Margaret Martonosi, Douglas W. Clark
International Conference on Supercomputing2
1998 A Methodology and an Evaluation of the SGI Origin2000
abstract
As hardware-coherent, distributed shared memory (DSM) multiprocessing becomes popular commercially, it is important to evaluate modern realizations to understand how they perform and scale for a range of interesting applications and to identify the nature of the key bottlenecks. This paper evaluates the SGI Origin2000---the machine that perhaps has the most aggressive communication architecture of the recent cache-coherent offerings---and, in doing so, articulates a sound methodology for evaluating real systems. We examine data access and synchronization microbenchmarks; speedups for different application classes, problem sizes and scaling models; detailed interactions and time breakdowns using performance tools; and the impact of special hardware support. We find that overall the Origin appears to deliver on the promise of cache-coherent shared address space multiprocessing, at least at the 32-processor scale we examine. The machine is quite easy to program for performance and has fewer organizational problems than previous systems we have examined. However, some important trouble spots are also identified, especially related to contention that is apparently caused by engineering decisions to share resources among processors.
Dongming Jiang, Jaswinder Pal Singh
SIGMETRICS1
1997 Parallel Shear-Warp Volume Rendering on Shared Address Space Multiprocessors
abstract
This paper presents a new parallel volume rendering algorithm and implementation, based on shear warp factorization, for shared address space multiprocessors. Starting from an existing parallel shear-warp renderer, we use increasingly detailed performance measurements on real machines and simulators to understand performance bottlenecks. This leads us to a new parallel implementation that substantially outperforms and out-scales the old one on a range of shared address space platforms, from bus-based centralized memory machine to hardware-coherent distributed memory machines to networks of computers connected by page-based shared virtual memory. The results demonstrate that real time volume rendering is promising on general purpose multiprocessors, and illustrate the utility of tool hierarchies in conjunction with algorithmic and application knowledge to understand memory system interactions and improve parallel algorithms.
Dongming Jiang, Jaswinder Pal Singh
PPoPP1
1997 Application Restructuring and Performance Portability on Shared Virtual Memory and Hardware-Coherent Multiprocessors
abstract
The performance portability of parallel programs across a wide range of emerging coherent shared address space systems is not well understood. Programs that run well on efficient, hardware cache-coherent systems often do not perform well on less optimal or more commodity-based communication architectures. This paper studies this issue of performance portability, with the commodity communication architecture of interest being page-grained shared virtual memory. We begin with applications that perform well on moderat scale hardware cache-coherent systems, and find that they do not do so well on SVM systems. Then, we examine whether and how the applications can be improved for SVM systems --- through data structuring or algorithmic enhancements---and the nature and difficulty of the optimization. Finally, we examine the impact of the successful optimizations on hardware-coherent platforms themselves, to see whether they are helpful, harmful or neutral on those platforms. We develop a systematic methodology to explore optimizations in different structured classes. The results, and the difficulty of the optimizations, lead insight not only into performance portability but also into the viability of SVM as a platform for these types of applications.
Dongming Jiang, Hongzhang Shan, Jaswinder Pal Singh
PPoPP1