Chien-Hao Lee

dblp:20/1906 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
0since 2021 · last 2019
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Memory systems · 100%

Topics — the 3 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems
memory hierarchy
0.012004
Tolerating memory latency through push prefetching for pointer-intensive applications · ACM Trans. Archit. Code Optim. 2004
Memory systems › memory access optimization
pointer-chasing prefetching
0.012004
Tolerating memory latency through push prefetching for pointer-intensive applications · ACM Trans. Archit. Code Optim. 2004
Memory systems › cache
prefetching
0.012004
Tolerating memory latency through push prefetching for pointer-intensive applications · ACM Trans. Archit. Code Optim. 2004

Methods — techniques the papers use, named apart from their topics

hardware/software cooperative prefetching · 0.0data movement model · 0.0
YearPublicationVenuePosition
2019 Adaptive Resource Allocation for ICIC in Downlink NOMA Systems
abstract
Inter-cell interference coordination (ICIC) has been widely studied for mitigating the effects of severe inter-cell interference (ICI) in cell- edge users. However, based on the scarcity of frequency resources in orthogonal multiple access systems, the ICIC methods proposed in the previous papers have difficulty in maintaining the overall performance and fairness of a system. Non-orthogonal multiple access (NOMA) is a promising radio access technology that can serve multiple users simultaneously with the same frequency resources. However, most previous work has not considered the ICI problem in NOMA systems. We propose a centralized adaptive ICIC framework for downlink NOMA systems, including a distributed clustering algorithm, a distributed power allocation algorithm, and a centralized frequency allocation algorithm. Simulation results demonstrate that the proposed framework outperforms all benchmark frameworks and can improve both the overall performance of a system and fairness among users.
Chien-Hao Lee, Makoto Kobayashi, Hung-Yu Wei 0001, Shunsuke Saruwatari, Takashi Watanabe 0001
VTC Fall1
2019 Deep Q-Network Based Adaptive Resource Allocation with User Grouping on ICIC
abstract
In cellular networks, inter-cell interference is the main factor in the reduction of service quality for users, so intercell interference coordination (ICIC) has been widely studied to mitigate severe interference. However, in some previous work, cell- edge users are sacrificed to improve the performance of the overall system. Apart from this, most previous methods change the ICIC configuration frequently to achieve the optimal results, but in practice, the frequent ICIC reconfiguration results in large overhead for small cells. Thus, a centralized dynamic ICIC scheme is proposed in this work, including Q-learning assisted deep neural network based ICIC framework and Type-Balanced User Grouping algorithm. The simulation results show that the proposed ICIC scheme outperforms the benchmarks in both sparse and dense user distribution.
Chien-Hao Lee, Kuang-Hsun Lin, Hung-Yu Wei 0001
VTC Spring1
2004 HotSpot cache: joint temporal and spatial locality exploitation for i-cache energy reduction
abstract
Power consumption is an important design issue of current embedded systems. It has been shown that the instruction cache accounts for a significant portion of the power dissipation of the whole chip. Several studies propose to add a cache (L0 cache) that is very small relative to the conventional L1 cache on chip for power optimization since a smaller cache has lower load capacitance. However, energy savings often come at the cost of performance degradation. In this paper, we propose a novel instruction cache architecture, the HotSpot cache, that achieves energy savings without sacrificing performance. The HotSpot cache identifies frequently accessed instructions dynamically and stores them in the L0 cache. Other instructions are placed only in the L1 cache. A steering mechanism is employed to direct an instruction to its allocated cache in the instruction fetch stage. The simulation results show that the HotSpot cache can achieve 52% instruction cache energy reduction on the average for a set of multimedia applications without performance degradation.
Chia-Lin Yang, Chien-Hao Lee
ISLPED2
2004 Tolerating memory latency through push prefetching for pointer-intensive applications
abstract
Prefetching is often used to overlap memory latency with computation for array-based applications. However, prefetching for pointer-intensive applications remains a challenge because of the irregular memory access pattern and pointer-chasing problem. In this paper, we proposed a cooperative hardware/software prefetching framework, the push architecture, which is designed specifically for linked data structures. The push architecture exploits program structure for future address generation instead of relying on past address history. It identifies the load instructions that traverse a LDS and uses a prefetch engine to execute them ahead of the CPU execution. This allows the prefetch engine to successfully generate future addresses. To overcome the serial nature of LDS address generation, the push architecture employs a novel data movement model. It attaches the prefetch engine to each level of the memory hierarchy and pushes , rather than pulls , data to the CPU. This push model decouples the pointer dereference from the transfer of the current node up to the processor. Thus a series of pointer dereferences becomes a pipelined process rather than a serial process. Simulation results show that the push architecture can reduce up to 100% of memory stall time on a suite of pointer-intensive applications, reducing overall execution time by an average 15%.
Chia-Lin Yang, Alvin R. Lebeck, Hung-Wei Tseng 0001, Chien-Hao Lee
ACM Trans. Archit. Code Optim.4