Chenglong Xiao

dblp:17/9547 · DBLP profile ↗
← Back
16ranked-venue papers
6as first author
11since 2021 · last 2026
0000-0001-7013-4985ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 11 · 5 first-author · 6 since 2021Theory of computation · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 MPM-LLM4DSE: Reaching the Pareto Frontier in HLS with Multimodal Learning and LLM-Driven Exploration
abstract
High-Level Synthesis (HLS) design space exploration (DSE) seeks Pareto-optimal designs within expansive pragma configuration spaces. To accelerate HLS DSE, graph neural networks (GNNs) are commonly employed as surrogates for HLS tools to predict quality of results (QoR) metrics, while multi-objective optimization algorithms expedite the exploration. However, GNN-based prediction methods may not fully capture the rich semantic features inherent in behavioral descriptions, and conventional multi-objective optimization algorithms often do not explicitly account for the domain-specific knowledge regarding how pragma directives influence QoR. To address these limitations, this paper proposes the MPM-LLM4DSE framework, which incorporates a multimodal prediction model (MPM) that simultaneously fuses features from behavioral descriptions and control and data flow graphs. Furthermore, the framework employs a large language model (LLM) as an optimizer, accompanied by a tailored prompt engineering methodology. This methodology incorporates pragma impact analysis on QoR to guide the LLM in generating high-quality configurations (LLM4DSE). Experimental results demonstrate that our multimodal predictive model signifi-cantly outperforms state-of-the-art work ProgSG by up to 10.25×. Furthermore, in DSE tasks, the proposed LLM4DSE achieves an average performance gain of 39.90% over prior methods, validating the effectiveness of our prompting methodology. Code and models are available at https://github.com/wslcccc/MPM-LLM4DSE.
Chenglong Xiao
DATE3
2026 Diffusion Model Based Multi-Objective Evolutionary Algorithm for High-Level Synthesis Design Space Exploration
Runxin Lin, Chenglong Xiao
ISCAS3
2026 SGFormer-RGCN: Accurate Performance Prediction for High-Level Synthesis Design Space Exploration
Ruiqi Tang, Chenglong Xiao
ISCAS3
2026 An algorithm with a delay of O(kΔ) for enumerating connected induced subgraphs of size k
Chenglong Xiao, Chengyong Mao
Inf. Comput.1
2025 Relation Logical Reasoning and Relation-aware Entity Encoding for Temporal Knowledge Graph Reasoning
abstract
Temporal Knowledge Graph Reasoning (TKGR) aims to predict future facts based on historical data. Current mainstream models primarily use embedding techniques, which predict missing facts by representing entities and relations as low-dimensional vectors. However, these models often consider only the structural information of individual entities and relations, overlooking the broader structure of the entire TKG. To address these limitations, we propose a novel model called Relation Logical Reasoning and Relation-aware Entity Encoding (RLEE), drawing inspiration from attention mechanisms and logical rule-based techniques. RLEE introduces a two-layer representation of the TKG: an entity layer and a relation layer. At the relation layer, we extract relation paths to mine potential logical correlations between different relations, learning relation embeddings through a process of relation logical reasoning. At the entity layer, we use the relation-aware attention mechanism to learn the entity embeddings specific to the predicted query relations. These learned relation and entity embeddings are then used to predict facts at future timestamps. When evaluated on five commonly used public datasets, RLEE consistently outperforms state-of-the-art baselines.
Longzhou Liu, Chenglong Xiao, Tingwen Liu
COLING2
2025 A Reinforcement Learning Based Multi-objective Evolutionary Algorithm for High-level Synthesis Design Space Exploration
abstract
By applying various optimization directives (e.g., loop unrolling, function inlining), high-level synthesis (HLS) enables designers to generate diverse hardware implementations with varying performance and resource trade-offs. The interplay of these directives significantly influences the performance and cost of the synthesized designs, making Design Space Exploration (DSE) crucial for identifying configurations that yield Pareto-optimal solutions. To address this complex multi-objective optimization challenge, we propose an approach that integrates the evolutionary algorithm with reinforcement learning. Specifically, Q-learning is employed to adaptively optimize the exploration strategy throughout the search process, while the evolutionary algorithm effectively balances global and local searches. Compared to state-of-the-art DSE approaches, experimental results demonstrate that our proposed approach achieves better performance in synthesis runs, synthesis time, and final ADRS, with improvements of 21.00%-34.00%, 25.22%-38.22%, and 80.80%97.00%, respectively, and achieves 115X-837X speedup.
Yuda Qian, Chenglong Xiao
ISCAS3
2025 Decomposition based estimation of distribution algorithm for high-level synthesis design space exploration
Huiliang Hong, Chenglong Xiao
Integr.4
2025 Novel Algorithms for Efficient Mining of Connected Induced Subgraphs of a Given Cardinality
Chenglong Xiao
J. Comput. Sci. Technol.2
2024 Fine-Grained Features Alignment and Fusion for Text-Video Cross-Modal Retrieval
abstract
Text-video cross-modal retrieval is an increasingly prominent and challenging task that has garnered significant attention. Traditional models typically embed videos and texts into global vectors, aiming to capture the global features of these modalities. While the models often fall short in capturing fine-grained semantic details. Relying solely on global features proves insufficient to address this challenge. Hence, there is a pressing need to bridge the gap between different modalities by incorporating fine-grained features. In light of this, we propose a highly efficient model designed to capture the fine-grained features of videos and texts including question answer semantic alignment, object alignment and text-video feature fusion. For texts, our model includes the incorporation of entity information and part-of-speech information including adjectives, nouns and verbs information, while for videos, the identification of objects plays a crucial role in facilitating text-video retrieval. Our model undergoes extensive training on the WebVid and CC3M datasets, yielding unequivocal evidence of its superior performance over baseline models. It excels particularly in zero-shot text-video cross-modal retrieval tasks, offering substantial reductions in required computational resources.
Shuili Zhang, Hongzhang Mu, Quangang Li, Chenglong Xiao, Tingwen Liu
ICASSP4
2024 Rethinking High-Level Synthesis Design Space Exploration from a Contrastive Perspective
abstract
In High-Level Synthesis (HLS), the design space of RTL implementations growing exponentially with the number of directives, coupled with the expensive nature of the syn-thesis process to obtain the characteristics of a single RTL implementation, pose challenges in exhaustively exploring the design space to identify Pareto-optimal designs. To find Pareto-optimal designs, most existing studies adopt regression methods or call HLS tools to provide the performance and cost of each design. Considering the essence of finding Pareto-optimal designs, we observed that predicting the relative dominance relation-ship between designs is sufficient while predicting the absolute performance and cost of each design is not a must. We propose an alternative approach to identify Pareto-optimal designs by establishing a Contrastive Learning (CL) framework that determines the dominance relationship between designs based on classification methods. Furthermore, we integrate the CL framework with three state-of-the-art Design Space Exploration (DSE) approaches, namely CL-DSE, with the goal of effectively reducing the number of syntheses. Experimental comparisons of 39 regression methods and 17 CL-assisted classification methods reveal that the classification methods generally overwhelm the regression methods for predicting the dominance relationship between designs. Experimental results also reveal that the CL-DSE approaches can significantly decrease the overall running time by reducing the number of syntheses, while retaining the quality of results in comparison to the DSE approaches. The proposed CL-DSE approaches are publicly available at https://github.com/hong64/CL-DSE.
Huiliang Hong, Chenglong Xiao
ICCD2
2024 Algorithms with improved delay for enumerating connected induced subgraphs of a large cardinality
Chenglong Xiao, Emmanuel Casseau
Inf. Process. Lett.2
2017 Parallel custom instruction identification for extensible processors
Chenglong Xiao, Wanjun Liu, Emmanuel Casseau
J. Syst. Archit.1
2014 Improving high-level synthesis effectiveness through custom operator identification
abstract
It is increasingly common to see custom operators appear in various fields of circuit design. Custom operators that can be implemented in special hardware units make it possible to improve performance and reduce area. In this paper, we propose a design flow for identifying custom operators for high-level synthesis. Experimental results show that our approach achieves on average 19%, and up to 37% area reduction, compared to a traditional high-level synthesis. Meanwhile, the latency is reduced on average by 22%, and up to 59%. In addition, on average 74% and up to 81% code size reduction can be achieved, so synthesis runtime can be reduced.
Chenglong Xiao, Emmanuel Casseau
ISCAS1
2012 Exact custom instruction enumeration for extensible processors
Chenglong Xiao, Emmanuel Casseau
Integr.1
2011 Efficient custom instruction enumeration for extensible processors
abstract
In recent years, the use of extensible processors has been increased. Extensible processors extend the base instruction set of a general-purpose processor with a set of custom instructions. Custom instructions that can be implemented in special hardware units make it possible to improve performance and decrease power consumption in extensible processors. The key issue involved is to generate and select automatically custom instructions from a high-level application code. In this paper, we propose an efficient and flexible algorithm for the exact enumeration of custom instructions. The algorithm can be tuned to generate all possible patterns or only connected patterns. Compared to a previously proposed well-known algorithm, our algorithm can achieve orders of magnitude speedup.
Chenglong Xiao, Emmanuel Casseau
ASAP1
2011 An efficient algorithm for custom instruction enumeration
abstract
In order to meet growing market demands in flexibility and performance, the use of extensible processors has been increased. Extensible processors extend the base instruction set of a general-purpose processor with a set of custom instructions. Custom instruction that can be implemented in special hardware unit is a vital component for improving performance in extensible processors. The key issue involved is to generate and select automatically custom instructions from high-level application code. In this paper, we propose a new efficient algorithm for automatic generation of all candidate instructions (or patterns). Our pattern generation algorithm identifies all feasible connected and disjoint patterns under different constraints. Compared to a previously proposed well-known algorithm, our algorithm solves the problem more efficiently by taking advantage of topological property of data flow graph (DFG) as well as overcoming the drawbacks of the previously proposed algorithm.
Chenglong Xiao, Emmanuel Casseau
ACM Great Lakes Symposium on VLSI1