Quang Dinh

dblp:64/3078 · DBLP profile ↗
← Back
6ranked-venue papers
5as first author
0since 2021 · last 2015
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 5 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Parallel and multicore computing · 34% High-performance computing · 23% Processor architecture and microarchitecture · 16%

Topics — the 13 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Parallel and multicore computing › parallel algorithms › parallel geometric algorithms
parallel mesh processing
0.212015
Scalable and efficient implementation of 3d unstructured meshes computation: a case study on matrix assembly · PPoPP 2015
Parallel and multicore computing
parallel programming models
0.212015
Scalable and efficient implementation of 3d unstructured meshes computation: a case study on matrix assembly · PPoPP 2015
High-performance computing
unstructured mesh computation
0.212015
Scalable and efficient implementation of 3d unstructured meshes computation: a case study on matrix assembly · PPoPP 2015
Energy-efficient computing
dynamic power reduction
0.112010
A Routing Approach to Reduce Glitches in Low Power FPGAs · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2010
Electronic design automation › physical design › routing
FPGA routing
0.112010
A Routing Approach to Reduce Glitches in Low Power FPGAs · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2010
Energy-efficient computing › dynamic power reduction
glitch reduction
0.112010
A Routing Approach to Reduce Glitches in Low Power FPGAs · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2010
Processor architecture and microarchitecture › instruction set architecture
application-specific instruction-set processor
0.112008
Efficient ASIP design for configurable processors with fine-grained resource sharing · FPGA 2008
Processor architecture and microarchitecture › instruction set architecture › instruction set extension
custom instruction generation
0.112008
Efficient ASIP design for configurable processors with fine-grained resource sharing · FPGA 2008
Electronic design automation
high-level synthesis
0.112008
Efficient ASIP design for configurable processors with fine-grained resource sharing · FPGA 2008
High-performance computing
domain decomposition
0.112015
Scalable and efficient implementation of 3d unstructured meshes computation: a case study on matrix assembly · PPoPP 2015
Parallel and multicore computing
load balancing
0.112015
Scalable and efficient implementation of 3d unstructured meshes computation: a case study on matrix assembly · PPoPP 2015
Processor architecture and microarchitecture
many-core architecture
0.112015
Scalable and efficient implementation of 3d unstructured meshes computation: a case study on matrix assembly · PPoPP 2015
High-performance computing › code optimization
vectorization
0.112015
Scalable and efficient implementation of 3d unstructured meshes computation: a case study on matrix assembly · PPoPP 2015

Methods — techniques the papers use, named apart from their topics

mesh coloring · 0.2domain decomposition · 0.2divide-and-conquer · 0.2target-delay routing · 0.1arrival time alignment · 0.1resource sharing heuristics · 0.1
YearPublicationVenuePosition
2015 Scalable and efficient implementation of 3d unstructured meshes computation: a case study on matrix assembly
abstract
Exposing massive parallelism on 3D unstructured meshes computation with efficient load balancing and minimal synchronizations is challenging. Current approaches relying on domain decomposition and mesh coloring struggle to scale with the increasing number of cores per nodes, especially with new many-core processors. In this paper, we propose an hybrid approach using domain decomposition to exploit distributed memory parallelism, Divide-and-Conquer, D&C, to exploit shared memory parallelism and improve locality, and mesh coloring at core level to exploit vectors. It illustrates a new trade-off for many-cores between structuredness, memory locality, and vectorization. We evaluate our approach on the finite element matrix assembly of an industrial fluid dynamic code developed by Dassault Aviation. We compare our D&C approach to domain decomposition and to mesh coloring. D&C achieves a high parallel efficiency, a good data locality as well as an improved bandwidth usage. It competes on current nodes with the optimized pure MPI version with a minimum 10% speed-up. D&C shows an impressive 319x strong scaling on 512 cores (32 nodes) with only 2000 vertices per core. Finally, the Intel Xeon Phi version has a performance similar to 10 Intel E5-2665 Xeon Sandy Bridge cores and 95% parallel efficiency on the 60 physical cores. Running on 4 Xeon Phi (240 cores), D&C has 92% efficiency on the physical cores and performance similar to 33 Intel E5-2665 Xeon Sandy Bridge cores.
Loïc Thébault, Quang Dinh
PPoPP3
2010 Dynamic power estimation for deep submicron circuits with process variation
abstract
Dynamic power consumption in CMOS circuits is usually estimated based on the number of signal transitions. However, when considering glitches, this is not accurate because narrow glitches consume less power than wide glitches. Glitch width and transition density modeling is further complicated by the effect of process variation. This paper presents a fast and accurate dynamic power estimation method that considers the detailed effect of process variation. First, we extend the probabilistic modeling approach to handle timing variations. Then the power consumption of a logic gate is computed based on the transition waveforms of its inputs. Both mean values and standard deviations of the dynamic power are estimated with high confidence based on accurate device characterization data. Compared with SPICE-based Monte Carlo simulations for small circuits, our power estimator reports power results within 3% error for the mean and 5% error for the standard deviation with six orders of magnitude speedup. For medium and large benchmarks, it is impossible to run Monte Carlo simulations with enough samples due to very long runtime, while our estimator can finish within minutes.
Quang Dinh, Deming Chen, Martin D. F. Wong
ASP-DAC1
2010 BDD-based circuit restructuring for reducing dynamic power
abstract
As advances in process technology continue to scale down transistors, low power design is becoming more critical. Clock gating is a dynamic power saving technique that can freeze some flip-flops and prevent portion of the circuit from unneeded switching. In this paper, we consider fine-grained clock gating through pipelining, in which control signals from one pipeline stage are used to freeze some logic in the next pipeline stage. We present a novel BDD-based decomposition algorithm to restructure the circuit and expose possible control signals that would maximize power saving. We then use ILP formulation to select the optimal set of control signals for the circuit. We show that the constraint matrix is totally unimodular, and solve this selection problem optimally using linear programming. Comparing to a previous work, we get similar and 9% better dynamic power saving for small and medium circuits, respectively. For the largest MCNC circuits, which the previous technique cannot handle, we get an average of 19% dynamic power saving with 9.3% area overhead comparing to the original, non-restructured circuits.
Quang Dinh, Deming Chen, Martin D. F. Wong
ICCD1
2010 A Routing Approach to Reduce Glitches in Low Power FPGAs
abstract
This paper presents a novel approach to reduce dynamic power in field-programmable gate arrays (FPGAs) by reducing glitches during routing. It finds alternative routes for early-arriving signals so that signal arrival times at look-up tables are aligned. We developed an efficient algorithm to find routes with target delays and then built a glitch-aware router aiming at reducing dynamic power. To the best of our knowledge, this is the first glitch-aware routing algorithm for FPGAs. Experiments show that an average of 27% reduction in glitch power is achieved, which translates into an 11% reduction in dynamic power, compared to the glitch-unaware versatile place and route's router.
Quang Dinh, Deming Chen, Martin D. F. Wong
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2009 A routing approach to reduce glitches in low power FPGAs
abstract
Glitches (spurious transitions) are common in electronic circuits. In this paper we present a novel approach to reduce dynamic power in FPGAs by reducing glitches during the routing step. This approach involves finding alternative routes for early-arriving signals, so that signal arrival times at LUTs are aligned and no glitches are generated. This approach does not require additional circuitry to balance signals as done in previous work, but uses the available programmable routing resources instead. We develop an efficient algorithm to find routes with target delays. Based on this algorithm, we then build a glitch-aware router, named GlitchReroute, aiming at reducing dynamic power. To the best of our knowledge, this is the first glitch-aware routing algorithm for FPGAs. Experiments show that an average of 23% reduction in glitch power is achieved, which translates into a 9.8% reduction in dynamic power, compared to the glitch-unaware VPR router.
Quang Dinh, Deming Chen, Martin D. F. Wong
ISPD1
2008 Efficient ASIP design for configurable processors with fine-grained resource sharing
abstract
Application-Specific Instruction-set Processors (ASIP) can improve execution speed by using custom instructions. Several ASIP design automation flows have been proposed recently. In this paper, we investigate two techniques to improve these flows, so that ASIP can be efficiently applied to simple computer architectures in embedded applications. Firstly, we efficiently generate custom instructions with multi-cycle IO (which allows multi-outputs), thus removing the constraint imposed by the ports of the register file. Secondly, we allow identical portions of different custom instructions to be shared, thus allowing more custom instructions under the same area constraint. To handle the greatly increased exploration space, we propose several heuristics to keep the problem tractable. Experimental results show that we can achieve 3x speedup in some cases
Quang Dinh, Deming Chen, Martin D. F. Wong
FPGA1