Guangjun Xie

dblp:54/1672 · DBLP profile ↗
← Back
24ranked-venue papers
2as first author
13since 2021 · last 2026
0000-0002-7801-2875ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 13 · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Security and privacy · 3 · 1 first-authorArtificial intelligence and machine learning · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 High accurate approximate adders using hybrid gates
Yongqiang Zhang 0006, Jiao Qin, Xin Cheng 0001, Guangjun Xie
Integr.4
2026 iFCN: An Automated RTL-to-Device Framework for Molecular Field-Coupled Nanocomputing Circuits
abstract
Molecular Field-Coupled Nanocomputing (MolFCN) offers a promising post-CMOS alternative, characterized by ultra-low power consumption and high integration density. However, existing MolFCN design flows face critical challenges, including rigid clock-phase constraints, inefficient placement and routing, and the absence of accurate gate-to-device mapping. This paper introducesiFCN, an automated RTL-to-device-level design framework specifically optimized for MolFCN circuits. Building upon prior heuristic methods,iFCNincorporates inverter pruning at the RTL level to simplify circuit structure, utilizes Morton-coded linear quadtrees for efficient spatial indexing, and enhances A* routing to maximize path reuse for multi-fanout nets. To address limitations of fixed-phase clocking, we propose a hierarchical placement approach guided by Graph Convolutional Networks (GCN). The GCN learns connectivity-aware node embeddings that guide recursive partitioning and intra-layer ordering, significantly reducing wire crossings and improving placement quality. Subsequently, a lightweight adaptive method heuristically assigns clock phases to each layout layer, with careful consideration of timing and topological constraints. Additionally, we present an accurate gate-to-cell mapping algorithm to facilitate direct physical simulation and energy analyses. Benchmark results show a 30% reduction in runtime and a 10% improvement in layout area compared to our heuristic algorithm. Furthermore, the proposed method achieves comparable runtime performance to the state-of-the-artfictiontool, completing the layout of circuits with over 150 nodes in under one second. In addition, Comparative analyses against 12nm CMOS designs confirm that MolFCN circuits generated byiFCNexhibit superior area utilization and reduced power consumption, highlighting the practical potential of MolFCN for future ultra-low-power computing. All source code and data are available athttps://github.com/li-yangshuai/iFCN
Yangshuai Li, Xiansheng Tong, Rongjie Zhu, Qian Han, Guangjun Xie
IEEE Trans. Circuits Syst. I Regul. Pap.6
2025 Reconfigurable 10T SRAM for Energy-Efficient CAM Operation and In-Memory Computing
abstract
The limitations of the von Neumann architecture in terms of power consumption and throughput are increasingly evident. In-memory computing is a promising computing paradigm to alleviate this limitation. This article proposes a high-speed and low-power 10T compute-static random-access memory (CSRAM) capable of conducting rowwise search operations and executing in-memory logic functions efficiently. A self-suppressed discharge scheme is implemented to curtail the power consumption of the search operation by reducing the discharge swing of the match lines (MLs). The rowwise search scheme avoids vertical data storage, enhancing the compatibility between different operation modes. The proposed 10T SRAM architecture addresses the issue of sneak currents effectively when multiple lines are activated. Additionally, decoupled read ports eliminate compute access disturbance. To validate the design, a 4Kb array is designed with a 40-nm CMOS technology. At a supply voltage (VDD) of 1.1 V, the in-memory logic operations are capable of operating at a frequency of 752 MHz, consuming 29.2 fJ/bit. In binary content-addressable memory (BCAM) search mode, the minimum energy consumption of 0.51 fJ/bit occurs at 0.8 V and 120 MHz.
Zhang Zhang 0004, Zhihao Chen 0008, Jiedong Wang, Guangjun Xie
IEEE Trans. Very Large Scale Integr. Syst.4
2024 A 10T SRAM with Two Read and Write Modes across Row and Column for CAM Operation and Computing In-Memory
abstract
With SRAM-based computing in-memory (CIM), parallel searching is implemented through the multi-row activation scheme, which necessitates words to be stored in the column-wise fashion. However, existing column-wise write schemes usually require multi-cycle and cause write performance degradation for the SRAM with only row access transistors. In this study, we propose a novel 10T SRAM with both row and column access transistors, supporting data writing across row and column without additional data moving, overcoming the above problem. Furthermore, the proposed SRAM features horizontal and vertical read ports to enable two-direction logic operations, search operation, and matrix transposition, significantly enhancing computational flexibility. Besides, the array can be used to perform arithmetic operations. The 10T SRAM design is validated in a 4 Kb array with a 40-nm CMOS technology. It achieves a frequency of 917 MHz at 1.1V for logic operations. For binary content-addressable memory (BCAM) search operations, the energy consumption is 0.82 fJ/search/bit at 0.7 V in the worst case, and the frequency is up to 807 MHz at 1.1V.
Zhang Zhang 0004, Zhihao Chen 0008, Sikai Chen, Guangjun Xie, Jianmin Zeng
ISCAS4
2024 Cascaded refinement residual attention network for image outpainting
Yizhong Yang, Zhang Zhang 0004, Guangjun Xie
Multim. Syst.5
2024 Design of a Stochastic Computing Architecture for the Phansalkar Algorithm
abstract
Binarization plays a key role in image processing. Its performance directly affects the success of subsequent character segmentation and recognition. The Phansalkar algorithm performs excellent in processing heavily degraded or poor-quality images. However, this algorithm incurs significant hardware costs. In this article, efficient stochastic computing (SC) functions and an architecture are proposed for the Phansalkar algorithm. Highly accurate stochastic elements are designed for this architecture, including a stochastic mean circuit (SMC), a stochastic unipolar subtractor (USUB), a stochastic square root circuit (SQRT), and a stochastic exponential circuit (SEXP). Simulation results show that the SC architecture using 64-bit streams for the Phansalkar algorithm provides sufficient accuracy. Physical implementation indicates the effectiveness of the proposed architecture in lowering hardware costs for this algorithm compared with the binary counterpart.
Yongqiang Zhang 0006, Jiao Qin, Jie Han 0001, Guangjun Xie
IEEE Trans. Very Large Scale Integr. Syst.4
2023 CFS: Scaling Metadata Service for Distributed File System via Pruned Scope of Critical Sections
abstract
There is a fundamental tension between metadata scalability and POSIX semantics within distributed file systems. The bottleneck lies in the coordination, mainly locking, used for ensuring strong metadata consistency, namely, atomicity and isolation. CFS is a scalable, fully POSIX-compliant distributed file system that eliminates the metadata management bottleneck via pruning the scope of critical sections for reduced locking overhead. First, CFS adopts a tiered metadata organization to scale file attributes and the remaining namespace hierarchies independently with appropriate partitioning and indexing methods, eliminating cross-shard distributed coordination. Second, it further scales up the single metadata shard performance by single-shard atomic primitives, shortening the metadata requests' lifespan and removing spurious conflicts. Third, CFS drops the metadata proxy layer but employs the light-weight, scalable client-side metadata resolving. CFS has been running in the production environment of Baidu AI Cloud for three years. Our evaluation with a 50-node cluster and microbenchmarks shows that CFS simultaneously improves the throughput of baselines like HopsFS and InfiniFS by 1.76--75.82× and 1.22--4.10×, and reduces their average latency by up to 91.71% and 54.54%, respectively. Under cases with higher contention and larger directories, CFS' throughput benefits expand by one order of magnitude. For three real-world workloads with data accesses, CFS introduces 1.62--2.55× end-to-end throughput speedups and 35.06--62.47% tail latency reductions over InfiniFS.
Yiduo Wang 0002, Yufei Wu 0011, Cheng Li 0001, Biao Cao, Yinlong Xu 0001, Guangjun Xie
EuroSys10
2023 Cascaded deep residual learning network for single image dehazing
Yizhong Yang, Ce Hou, Haixia Huang, Zhang Zhang 0004, Guangjun Xie
Multim. Syst.5
2023 A multi-scale feature fusion spatial-channel attention model for background subtraction
Yizhong Yang, Tingting Xia, Dajin Li, Zhang Zhang 0004, Guangjun Xie
Multim. Syst.5
2023 An Energy-Efficient Binary-Interfaced Stochastic Multiplier Using Parallel Datapaths
abstract
Stochastic computing (SC) typically requires a low design complexity compared with weighted binary computing, so it has been successfully applied in neural networks (NNs). Usually, SC utilizes random bitstreams as its medium, which makes it suffer from a long delay that offsets its advantages. This drawback can be alleviated by utilizing parallel datapaths, which, however, will significantly increase the hardware cost due to the requirement of multiple parallel computing units. In this article, a hybrid bit-splitting generator (HBSG) is proposed to efficiently produce parallel bitstreams in a single clock cycle to reduce delay. The HBSG uniformly splits binary numbers into R segments, each of which is encoded in parallel by using hardwired connections according to the weight of each bit. A binary-interfaced parallel stochastic multiplier (BipSMul) using the HBSG is then proposed to accelerate the multiplication in SC. Experimental results show that the BipSMul is more energy efficient than the state-of-the-art parallel and serial stochastic designs, as well as their binary and Booth counterparts, in delay, power-delay product (PDP), and area-delay product (ADP).
Yongqiang Zhang 0006, Siting Liu 0001, Jie Han 0001, Zhendong Lin, Xin Cheng 0001, Guangjun Xie
IEEE Trans. Very Large Scale Integr. Syst.7
2022 Field-Coupled Nanocomputing Placement and Routing With Genetic and A* Algorithms
abstract
Field-Coupled Nanocomputing technologies have great potential to surpass CMOS technology because of their lower power consumption and higher device concentration. To ease the burden of placement and routing (P&R) problems for FCN circuits, many delicate two-dimensional clocking schemes have been proposed, upon which algorithms can solve the P&R problems more strategically. In this paper, we propose a two-level optimization strategy by using a genetic algorithm (GA) combined with an enhanced A* algorithm. Some circuit design requirements, such as clock synchronization, layout area, etc., are cleverly designed in the fitness value function of the GA. Numerical results demonstrate the effectiveness of the hybrid algorithm. In particular, compared to current tools, such as fiction and Ropper, the proposed algorithm can achieve an optimal solution with a higher success rate and a sizeable applicable circuit scale. In addition, the concept of design rule checking (DRC) was proposed in FCN and integrated into the algorithm, making the P&R results mapping from gate-level to cell-level more smoothly. Besides, the number of cross wires is significantly reduced, and the distribution of IO ports can be more effectively controlled.
Yangshuai Li, Guangjun Xie, Qian Han, Xiaoshuai Li, Gaisheng Li, Bing Zhang 0016
IEEE Trans. Circuits Syst. I Regul. Pap.2
2022 STPNet: A Spatial-Temporal Propagation Network for Background Subtraction
abstract
In background subtraction tasks, spatial and temporal contexts are beneficial in detecting moving objects. The methods based on Deep Neural Networks in this task has explored different topologies, which are composed of the conventional operations of convolutional neural networks, such as Convolutional Long-short Term Memory layer (ConvLSTM), 2D convolutional layer, or 3D convolutional layer, to capture these contexts. In this work, we propose a new background subtraction algorithm named spatial–temporal propagation network. An end-to-end network with novel layers, whose process of operation is equivalent to that the feature maps multiply with affinity matrices, is proposed to capture the spatial–temporal correlation in video sequences and aggregate the deep features from the consecutive frames. Experimental results on CDnet-2014 and LASIESTA datasets show that this novel layer provides an alternative way for our network to aggregate multiscale spatial–temporal features. Meanwhile, the proposed network achieves state-of-the-art performance and is generalizable to unseen videos.
Yizhong Yang, Jiahao Ruan, Yongqiang Zhang 0006, Xin Cheng 0001, Zhang Zhang 0004, Guangjun Xie
IEEE Trans. Circuits Syst. Video Technol.6
2022 MSE-Net: generative image inpainting with multi-scale encoder
Yizhong Yang, Zhihang Cheng, Haotian Yu, Yongqiang Zhang 0006, Xin Cheng 0001, Zhang Zhang 0004, Guangjun Xie
Vis. Comput.7
2020 A matrix representation method for decoders using majority gate characteristics in quantum-dot cellular automata
Feifei Deng, Guangjun Xie, Renjun Zhu, Yongqiang Zhang 0006
J. Supercomput.2
2018 A Qi compatible wireless power receiver with integrated full-wave synchronous rectifier
Chubin Wu, Zhang Zhang 0004, Jianmin Zeng, Xin Cheng 0001, Guangjun Xie
Sci. China Inf. Sci.5
2018 The Fundamental Primitives with Fault-Tolerance in Quantum-Dot Cellular Automata
Mengbo Sun, Hongjun Lv, Yongqiang Zhang 0006, Guangjun Xie
J. Electron. Test.4
2018 Novel designs of full adder in quantum-dot cellular automata technology
Lei Wang 0141, Guangjun Xie
J. Supercomput.2
2017 Detecting TCP-Based DDoS Attacks in Baidu Cloud Computing Data Centers
abstract
Cloud computing data centers have become one of the most important infrastructures in the big-data era. When considering the security of data centers, distributed denial of service (DDoS) attacks are one of the most serious problems. Here we consider DDoS attacks leveraging TCP traffic, which are increasingly rampant but are difficult to detect. To detect DDoS attacks, we identify two attack modes: fixed source IP attacks (FSIA) and random source IP attacks (RSIA), based on the source IP address used by attackers. We also propose a real-time TCP-based DDoS detection approach, which extracts effective features of TCP traffic and distinguishes malicious traffic from normal traffic by two decision tree classifiers. We evaluate the proposed approach using a simulated dataset and real datasets, including the ISCX IDS dataset, the CAIDA DDoS Attack 2007 dataset, and a Baidu Cloud Computing Platform dataset. Experimental results show that the proposed approach can achieve attack detection rate higher than 99% with a false alarm rate less than 1%. This approach will be deployed to the victim-end DDoS defense system in Baidu cloud computing data center.
Jiahui Jiao, Benjun Ye, Rebecca J. Stones, Gang Wang 0001, Xiaoguang Liu 0001, Shaoyan Wang, Guangjun Xie
SRDS8
2011 hUBI: An Optimized Hybrid Mapping Scheme for NAND Flash-Based SSDs
abstract
NAND flash-based SSDs have become attractive alternatives to hard disk drivers due to their high random read performances and low power consumptions. However, the poor random write performances highly limit their popularization in commercial applications. In this paper, we propose a novel mapping scheme called hybrid mapping unsorted block images (hUBI). hUBI aims for (1) optimized random write performances, (2) low write laten- cies, and (3) low space consumptions. It optimizes traditional hybrid mapping schemes to achieve (1). In hUBI, the merge operation involves only a single block. So, (2) is guaranteed. To avoid introducing a high space cost maintaining the metadata, hUBI puts its metadata into out-of-band (OOB) areas of SSD pages, which obtains (3). Our experimental results show that hUBI provides a considerable random write performance and holds low write latencies at the same time.
Guangjun Xie, Guangzhi Xu, Gang Wang 0001, Jing Liu 0010
TrustCom1
2008 Some rewrite optimizations of DB2 XQuery navigation
abstract
IBM® DB2® 9 is a truly hybrid commercial database system that combines XML and relational data. It provides native support for XML storage and indexing, and query evaluation support for XQuery. By building a hybrid system, the designers of DB2 9 were able to use the existing SQL query evaluation and optimization techniques to develop similar methods for XQuery. However, SQL and XQuery are sufficiently different that new optimization techniques can and are being developed in the new XQuery domain. This paper describes a few such techniques, all based on static rewrites of XQuery expressions.
Guangjun Xie, Jarek Gryz, Calisto Zuzarte
CIKM1
2008 Constructing Double-Erasure HoVer Codes Using Latin Squares
abstract
Storage applications are in urgent need of multi-erasure codes. But there is no consensus on the best coding technique. Hafner has presented a class of multi-erasure codes named HoVer codes [1]. This kind of codes has a unique data/parity layout which provides a range of implementation options that cover a large portion of the performance/efficiency trade-off space. Thus it can be applied to many scenarios by simple tuning. In this paper, we give a combinatorial representation of a family of double-erasure HoVer codes - create a mapping between this family of codes and Latin squares. We also present two families of double-erasure HoVer codes respectively based on the column-Hamiltonian Latin squares (of odd order) and a family of Latin squares of even order. Compared with the double-erasure HoVer codes presented in [1], the new codes enable greater flexibility in performance and efficiency trade-off.
Gang Wang 0001, Xiaoguang Liu 0001, Sheng Lin 0002, Guangjun Xie, Jing Liu 0010
ICPADS4
2008 Generalizing RDP Codes Using the Combinatorial Method
abstract
In this paper, we present PDH Latin - a new class of 2-erasure horizontal codes with dependent parity symbols based on column-Hamiltonian Latin squares (CHLS). We prove that PDH Latin codes are MDS codes. We also present a new class of 2-erasure parity independent mixed codes based on CHLS - PIMLatin. We show that the performance of the new codes is comparable to or better than other codes of this kind. They have perfect parameter flexibility and structure variety that benefit performance. We also discuss code shortening technologies that can improve parameter flexibility, structure variety and reliability. Borrowing ideas from vertical shortening, we develop a 2-erasure array code construction method using non-Hamiltonian Latin squares.
Gang Wang 0001, Xiaoguang Liu 0001, Sheng Lin 0002, Guangjun Xie, Jing Liu 0010
NCA4
2007 Constructing double- and triple-erasure-correcting codes with high availability using mirroring and parity approaches
abstract
With the rapid progress of the capacity and slow pace of the speed/MTTF of hard disks, and increasing size of storage systems, the reliability and availability of storage systems become more and more serious. This paper discusses the method of constructing double- and triple-erasure-correcting codes via combining mirroring and parity approaches in details, and presents a double-erasure code MPDC and a triple-erasure code MPPDC based on one-factorizations of complete graphs. The two codes are simple, easy to implement, and have no disk number limitation. They achieve perfect fault-free load balance and approximately optimal reconstruction load balance. The simulation results show that, compared with other double- and triple-erasure codes, MPDC and MPPDC have comparative light-load and moderate-load performance and better heavy-load performance in fault-free mode. Because parity declustering is used, the two codes are far superior to the other double- and triple-erasure codes in degraded- and reconstruction-mode performance.
Gang Wang 0001, Xiaoguang Liu 0001, Sheng Lin 0002, Guangjun Xie, Jing Liu 0010
ICPADS4
2007 Combinatorial Constructions of Multi-erasure-Correcting Codes with Independent Parity Symbols for Storage Systems
abstract
In this paper, we present a new class of t-erasure horizontal codes with independent parity symbols based on Column-Hamiltonian Latin squares (CHLS). We call the codes PIHLatin (parity independent horizontal Latin) codes. We prove the necessary and sufficient condition of the existence of PIHLatin codes for t=2. For tges3, we prove some necessary conditions of the existence of PIHLatin codes. We also prove the bijection between 2-erasure PIHLatin-like codes and CHLSs and prove the mapping from t-erasure PIHLatin-like codes to t-1 mutually orthogonal CHLSs for t>2. The performance analysis shows that PIHLatin codes are superior to other multi-erasure array codes in flexibility and variety. Moreover, PIHLatin codes are suitable for both traditional disk arrays and distributed storage systems.
Gang Wang 0001, Sheng Lin 0002, Xiaoguang Liu 0001, Guangjun Xie, Jing Liu 0010
PRDC4