VLDB 2026 Research / reviewers in the wild / expert
Ohyoung Jang
dblp:27/10978
· DBLP profile ↗
5ranked-venue papers
0as first author
0since 2021 · last 2018
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3Artificial intelligence and machine learning · 1Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
4 papers |
Memory systems · 69% GPUs and heterogeneous computing · 10% Interconnection networks and networks-on-chip · 9% | |
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 91% Knowledge graphs · 9% | |
| Software engineering, system software, and programming languages
3 papers |
Compilers and program optimization · 100% |
Topics — the 21 heaviest of 22, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Information retrieval › interactive information retrieval › conversational information seeking
conversational search |
0.3 | 1 | 2018 | Conversational Semantic Search: Looking Beyond Web Search, Q&A and Dialog Systems · WSDM 2018 |
Information retrieval
question answering and dialogue systems |
0.3 | 1 | 2018 | Conversational Semantic Search: Looking Beyond Web Search, Q&A and Dialog Systems · WSDM 2018 |
Information retrieval › search engines
semantic search |
0.3 | 1 | 2018 | Conversational Semantic Search: Looking Beyond Web Search, Q&A and Dialog Systems · WSDM 2018 |
Memory systems
data layout optimization |
0.3 | 2 | 2013 | Data layout optimization for GPGPU architectures · PPoPP 2013 A data layout optimization framework for NUCA-based multicores · MICRO 2011 |
Memory systems
cache |
0.2 | 1 | 2013 | Data layout optimization for GPGPU architectures · PPoPP 2013 |
GPUs and heterogeneous computing
GPU programming |
0.2 | 1 | 2013 | Data layout optimization for GPGPU architectures · PPoPP 2013 |
Memory systems › memory hierarchy
memory hierarchy optimization |
0.2 | 1 | 2013 | Data layout optimization for GPGPU architectures · PPoPP 2013 |
Compilers and program optimization › vectorization
superword level parallelism |
0.1 | 1 | 2012 | A compiler framework for extracting superword level parallelism · PLDI 2012 |
Memory systems
cache coherence |
0.1 | 1 | 2012 | A hybrid NoC design for cache coherence optimization for chip multiprocessors · DAC 2012 |
Interconnection networks and networks-on-chip
network-on-chip design |
0.1 | 1 | 2012 | A hybrid NoC design for cache coherence optimization for chip multiprocessors · DAC 2012 |
Memory systems › cache coherence › cache coherence protocol
snoopy coherence |
0.1 | 1 | 2012 | A hybrid NoC design for cache coherence optimization for chip multiprocessors · DAC 2012 |
Memory systems
cache design |
0.1 | 1 | 2011 | A data layout optimization framework for NUCA-based multicores · MICRO 2011 |
Parallel and multicore computing › task allocation
computation-to-core mapping |
0.1 | 1 | 2011 | A data layout optimization framework for NUCA-based multicores · MICRO 2011 |
Memory systems › memory hierarchy › cache hierarchy
non-uniform cache access |
0.1 | 1 | 2011 | A data layout optimization framework for NUCA-based multicores · MICRO 2011 |
Knowledge graphs › knowledge graph querying
knowledge graph question answering |
0.1 | 1 | 2018 | Conversational Semantic Search: Looking Beyond Web Search, Q&A and Dialog Systems · WSDM 2018 |
Compilers and program optimization › loop optimization
loop nest optimization |
0.0 | 1 | 2013 | Data layout optimization for GPGPU architectures · PPoPP 2013 |
Compilers and program optimization › memory optimization
data layout optimization |
0.0 | 1 | 2012 | A compiler framework for extracting superword level parallelism · PLDI 2012 |
Processor architecture and microarchitecture
chip multiprocessor |
0.0 | 1 | 2012 | A hybrid NoC design for cache coherence optimization for chip multiprocessors · DAC 2012 |
Processor architecture and microarchitecture › SIMD
SIMD instructions |
0.0 | 1 | 2012 | A compiler framework for extracting superword level parallelism · PLDI 2012 |
Compilers and program optimization
compiler optimization |
0.0 | 1 | 2011 | A data layout optimization framework for NUCA-based multicores · MICRO 2011 |
Compilers and program optimization › memory optimization
data locality optimization |
0.0 | 1 | 2011 | A data layout optimization framework for NUCA-based multicores · MICRO 2011 |
Methods — techniques the papers use, named apart from their topics
semantic functional unit composition · 0.3data localization · 0.3affine loop nest analysis · 0.3statement scheduling · 0.3statement grouping · 0.3data layout optimization · 0.3layout customization · 0.2full-system simulation · 0.2array tiling · 0.2hybrid noc · 0.1dynamic link reconfiguration · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2018 | Conversational Semantic Search: Looking Beyond Web Search, Q&A and Dialog SystemsabstractUser expectations of web search are changing. They are expecting search engines to answer questions, to be more conversational, and to offer means to complete tasks on their behalf. At the same time, to increase the breadth of tasks that personal digital assistants (PDAs), such as Microsoft»s Cortana or Amazon»s Alexa, are capable of, PDAs need to better utilize information about the world, a significant amount of which is available in the knowledge bases and answers built for search engines. It thus seems likely that the underlying systems that power web search and PDAs will converge. This demonstration presents a system that merges elements of traditional multi-turn dialog systems with web based question answering. This demo focuses on the automatic composition of semantic functional units, Botlets, to generate responses to user»s natural language (NL) queries. We show that such a system can be trained to combine information from search engine answers with PDA tasks to enable new user experiences. Paul A. Crook, Alex Marin, Vipul Agarwal, Samantha Anderson, Ohyoung Jang, Aliasgar Lanewala, Karthik Tangirala, Imed Zitouni |
WSDM | 5 |
| 2013 | Data layout optimization for GPGPU architecturesabstractGPUs are being widely used in accelerating general-purpose applications, leading to the emergence of GPGPU architectures. New programming models, e.g., Compute Unified Device Architecture (CUDA), have been proposed to facilitate programming general-purpose computations in GPGPUs. However, writing high-performance CUDA codes manually is still tedious and difficult. In particular, the organization of the data in the memory space can greatly affect the performance due to the unique features of a custom GPGPU memory hierarchy. In this work, we propose an automatic data layout transformation framework to solve the key issues associated with a GPGPU memory hierarchy (i.e., channel skewing, data coalescing, and bank conflicts). Our approach employs a widely applicable strategy based on a novel concept called data localization. Specifically, we try to optimize the layout of the arrays accessed in affine loop nests, for both the device memory and shared memory, at both coarse grain and fine grain parallelization levels. We performed an experimental evaluation of our data layout optimization strategy using 15 benchmarks on an NVIDIA CUDA GPU device. The results show that the proposed data transformation approach brings around 4.3X speedup on average. Jun Liu 0008, Wei Ding 0008, Ohyoung Jang, Mahmut T. Kandemir |
PPoPP | 3 |
| 2012 | A hybrid NoC design for cache coherence optimization for chip multiprocessorsabstractOn chip many-core systems, evolving from prior multi-processor systems, are considered as a promising solution to the performance scalability and power consumption problems. The long communication distance between the traditional multi-processors makes directory-based cache coherence protocols better solutions compared to bus-based snooping protocols even with the overheads from indirections. However, much smaller distances between the CMP cores enhance the reachability of buses, revitalizing the applicability of snooping protocols for cache-to-cache transfers. In this work, we propose a hybrid NoC design to provide optimized support for cache coherency. In our design, on-chip links can be dynamically configured as either point-to-point links between NoC nodes or short buses to facilitate localized snooping. By taking advantage of the best of both worlds, bus-based snooping coherency and NoC-based directory coherency, our approach brings both power and performance benefits. Hui Zhao 0013, Ohyoung Jang, Wei Ding 0008, Mahmut T. Kandemir, Mary Jane Irwin |
DAC | 2 |
| 2012 | A compiler framework for extracting superword level parallelismabstractSIMD (single-instruction multiple-data) instruction set extensions are quite common today in both high performance and embedded microprocessors, and enable the exploitation of a specific type of data parallelism called SLP (Superword Level Parallelism). While prior research shows that significant performance savings are possible when SLP is exploited, placing SIMD instructions in an application code manually can be very difficult and error prone. In this paper, we propose a novel automated compiler framework for improving superword level parallelism exploitation. The key part of our framework consists of two stages: superword statement generation and data layout optimization. The first stage is our main contribution and has two phases, statement grouping and statement scheduling, of which the primary goals are to increase SIMD parallelism and, more importantly, capture more superword reuses among the superword statements through global data access and reuse pattern analysis. Further, as a complementary optimization, our data layout optimization organizes data in memory space such that the price of memory operations for SLP is minimized. The results from our compiler implementation and tests on two systems indicate performance improvements as high as 15.2% over a state-of-the-art SLP optimization algorithm. Jun Liu 0008, Ohyoung Jang, Wei Ding 0008, Mahmut T. Kandemir |
PLDI | 3 |
| 2011 | A data layout optimization framework for NUCA-based multicoresabstractFuture multicore architectures are likely to include a large number of cores connected using an on-chip network with Non-uniform Cache Access (NUCA). In such architectures, whether a data request is satisfied from a local cache or a remote cache can make an important difference. To exploit this NUCA property, prior research explored both architectural enhancements as well as compiler-based code optimization strategies. In this work, we take an alternate view, and explore data layout optimizations to improve locality of data accesses in a NUCA-based system. Our proposed approach includes three steps: array tiling, computation-to-core mapping, and layout customization. The first of these tries to identify the affinity between data and computation taking into account parallelization information, with the goal of minimizing remote accesses. The second step maps computations (and their associated data) to cores with the goal of minimizing average distance-to-data, and the last step further customizes the memory layout taking into account the data placement policy adopted by the underlying architecture. We evaluated the success of this three-step approach in enhancing on-chip cache behavior using all application programs from the SPECOMP suite on a full-system simulator. Our results show that the proposed approach improves on average data access latency and execution time by 24.7% and 18.4%, respectively, in the case of static NUCA, and 18.1% and 12.7%, respectively, in the case of dynamic NUCA. Wei Ding 0008, Mahmut T. Kandemir, Jun Liu 0008, Ohyoung Jang |
MICRO | 5 |