Junjie Gu

dblp:61/3002 · DBLP profile ↗
← Back
13ranked-venue papers
8as first author
4since 2021 · last 2026
0000-0002-8593-4288ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 4 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 2 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
YearPublicationVenuePosition
2026 A Compact 0.2-19GHz Zero-IF Reconfigurable Quadrature Transmitter with Dual-Band Selection in 28nm CMOS Process
An Sun, Junjie Gu, Haoqi Qin, Weitao He, Yaxin Zeng, Hao Xu 0005, Na Yan 0004
ISCAS2
2025 Design and Analysis of a 26-32-GHz 6-bit Passive Vector Modulation Phase Shifter for CMOS Bidirectional Transceiver
abstract
This article presents a 26–32-GHz 6-bit bidirectional passive vector modulation phase shifter (PVM-PS) in 40-nm CMOS for phased array systems. The passive phase shifter comprises a center-tap transformer-based quadrature generator/combiner, two 6-bit X-type attenuators, and a differential Wilkinson power combiner/divider. The symmetric design enables bidirectional signal propagation and offers flexible system configuration. Passive switches are sized to optimize the tradeoff among gain variation, insertion loss, and linearity. The phase shifter implemented in 40 nm covers a range of 360° with 5.625° resolution and the rms phase error is between 0.4° and 1.3°. It exhibits <1-dB magnitude imbalance and <1.2° phase imbalance between forward and reverse propagation modes. Its OP1dB is above −1 dBm across the operation frequency.
Yechen Tian, Junjie Gu, Hao Xu 0005, Weitian Liu, Zongming Duan, Hao Gao 0001, Na Yan 0004
IEEE Trans. Very Large Scale Integr. Syst.3
2024 An innovative unsupervised gait recognition based tracking system for safeguarding large-scale nature reserves in complex terrain
Chichun Zhou 0001, Xiaolin Guan, Zhuohang Yu, Zhen-Yu Zhang, Junjie Gu
Expert Syst. Appl.6
2021 A Two-Way Power-Combining 60GHz CMOS Power Amplifier with 22.0% PAE and 19.4dBm Psat in 65nm Bulk CMOS
abstract
In this paper, the analysis and design of a 57-64 CMOS Power Amplifier is discussed. The power-combining technique and the capacitor neutralization technique are applied to boost the performance of the purposed PA. The PA is designed in 65nm bulk CMOS process to achieve a saturated output power of 19.4dBm and a peak power-added efficiency of 22%. The power amplifier consumes 300mW from a 1.2V power supply at the output-refereed 1dB compression point of 16.1dBm, and the corresponding power-added efficiency is 13.0%. The passive devices, such as the transformers, the power-combiner and the pads, are designed by Electromagnetic Field Simulation on ADS momentum.
Junjie Gu, Guixiang Jin, Hongtao Xu, Hao Min, Na Yan 0004
ISCAS1
2019 IGC: The Open Source Intel Graphics Compiler
abstract
With increasing general purpose programming capability, GPUs have become the mainstay for a wide variety of compute intensive tasks from cloud to edge computing. Because of its availability on nearly every desktop and mobile processor that Intel ships, Intel integrated GPU offers a plethora of opportunities for researchers and application developers to make significant real-world impact. In this paper we present the Intel Graphics Compiler (IGC), the LLVM-based production compiler for Intel HD and Iris graphics. IGC supports all major graphics and compute APIs, and its OpenCL compute stack including compute runtime, compiler frontend and backend, and architecture specification is fully open-source, giving a unique opportunity for developers to optimize the entire stack. We highlight several custom optimizations that address the challenges for GPU compilation. Examples include SIMD size selection, divergence analysis, instruction scheduling, addressing mode selection, and redundant copy elimination. These optimizations take advantage of features in the Intel GPU architecture such as a larger register file, indirect register addressing, and multiple memory addressing modes. Experimental results show that our optimizations deliver significant speedup on a number of OpenCL benchmarks; compared to the baseline, we see a geometric mean of 12% speed up across benchmarks with a peak gain of 45%.
Anupama Chandrasekhar, Wei-Yu Chen, Junjie Gu, Shruthi Hebbur Prasanna Kumar, Guei-Yuan Lueh, Pankaj Mistry, Thomas Raoux, Konrad Trifunovic
CGO5
2017 Fast and Automated Electromigration Analysis for CMOS RF PA Design
Junjie Gu, Haipeng Fu, Weicong Na, Qijun Zhang, Jianguo Ma
J. Electron. Test.1
2016 Design and Temperature Reliability Testing for A 0.6-2.14GHz Broadband Power Amplifier
Qian-Fu Cheng, Junjie Gu, Haipeng Fu
J. Electron. Test.3
2014 Investigate Performance of Expected Maximization on the Knowledge Tracing Model
Junjie Gu, Hang Cai, Joseph E. Beck
Intelligent Tutoring Systems1
2014 Personalizing Knowledge Tracing: Should We Individualize Slip, Guess, Prior or Learn Rate?
Junjie Gu, Neil T. Heffernan
Intelligent Tutoring Systems1
2008 Adaptive inferential feed-forward control algorithm and application with reduced model
abstract
In the process of industrial production, the system may be affected by many external factors, which can be equivalent to many measurable or immeasurable disturbances. The adaptive inferential feed-forward control algorithm adopts reduced model to design controller which resolves complexity of the adaptive algorithm when the order of model is unknown or too high. Meanwhile, the regularization technique is used to translate the unknown dynamic process into bounded disturbance and the relative dead zone technique is involved to identify parameters of the system, which guarantees bounded stability of the self-tuning control system. Through combining adaptive control, inferential control feed-forward control and adaptive prediction, the algorithm effectively eliminates the influence of measurable and immeasurable disturbances on the system. Finally, the validity and practicability of this algorithm is substantiates by the simulation result of superheated steam temperature control system.
Junjie Gu, Luanying Zhang, Zhiming Qin
ICARCV1
2000 Efficient Interprocedural Array Data-Flow Analysis for Automatic Program Parallelization
abstract
Since sequential languages such as Fortran and C are more machine-independent than current parallel languages, it is highly desirable to develop powerful parallelization tools which can generate parallel codes, automatically or semi-automatically, targeting different parallel architectures. Array data-flow analysis is known to be crucial to the success of automatic parallelization. Such an analysis should be performed interprocedurally and symbolically and it often needs to handle the predicates represented by IF conditions. Unfortunately, such a powerful program analysis can be extremely time-consuming if it is not carefully designed. How to enhance the efficiency of this analysis to a practical level remains an issue largely untouched to date. This paper presents techniques for efficient interprocedural array data-flow analysis and documents experimental results of its implementation in a research parallelizing compiler. Our techniques are based on guarded array regions and the resulting tool runs faster, by one or two orders of magnitude, than other similarly powerful tools.
Junjie Gu, Zhiyuan Li 0001
IEEE Trans. Software Eng.1
1997 Experience with Efficient Array Data-Flow Analysis for Array Privatization
abstract
Array data flow analysis is known to be crucial to the success of array privatization, one of the most important techniques for program parallelization. It is clear that array data flow analysis should be performed interprocedurally and symbolically, and that it often needs to handle the predicates represented by IF conditions. Unfortunately, such a powerful program analysis can be extremely time-consuming if not carefully designed. How to enhance the efficiency of thk analysis to a practical level remains an issue largely untouched to date. This paper documents our experience with building a highly efficient array data flow analyzer which is based on guarded array regions and which runs faster, by one or two orders of magnitude, than other similarly powerful tools.
Junjie Gu, Zhiyuan Li 0001, Gyungho Lee
PPoPP1
1995 Symbolic Array Dataflow Analysis for Array Privatization and Program Parallelization
abstract
Array dataflow information plays an important role for successful automatic parallelization of Fortran programs. This paper proposes a powerful symbolic array dataflow analysis to support array privatization and loop parallelization for programs with arbitrary control flow graphs and acyclic call graphs. Our scheme summarizes array access information using guarded array regions and propagates such regions over a Hierarchical Supergraph (HSG). The use of guards allows us to use the information in IF conditions to sharpen the array dataflow analysis and thereby to handle difficult cases which elude other existing techniques. The guarded array regions retain the simplicity of set operations for regular array regions in common cases, and they enhance regular array regions in complicated cases by using guards to handle complex symbolic expressions and array shapes. Scalar values that appear in array subscripts and loop limits are substituted on the fly during the array information propagation, which disambiguates the symbolic values precisely for set operations. We present efficient algorithms that implement our scheme. Initial experiments of applying our analysis to Perfect Benchmarks show promising results of improved array privatization.
Junjie Gu, Zhiyuan Li 0001, Gyungho Lee
SC1