Hongjun Dai

dblp:40/1029 · DBLP profile ↗
← Back
30ranked-venue papers
6as first author
13since 2021 · last 2026
0000-0002-1075-8750ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 12 · 5 first-author · 3 since 2021Artificial intelligence and machine learning · 6 · 6 since 2021Databases, data management, data science and information retrieval · 6 · 4 since 2021Computer networks · 4 · 2 since 2021Security and privacy · 2 · 1 first-authorSoftware engineering, systems software and programming languages · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Region Embedding With Adaptive Correlation Discovery for Predicting Urban Socioeconomic Indicators
abstract
A recent trend in urban computing involves utilizing multi-modal data for urban region embedding, which can be further expanded in a variety of downstream urban sensing tasks. Many previous studies rely on multi-graph embedding techniques and follow a two-stage paradigm: first building a k-nearest neighbor graph based on fixed region correlations for each view, and then blending multi-view information in a posterior stage to learn region representations. However, multi-graph construction and multi-graph representation learning are not associated in most existing two-stage studies, and the relationship between them is not leveraged, which can provide complementary information to each other. In this paper, we unify these two stages into one by constructing learnable weighted complete graphs of regions and propose a new one-stage Region Embedding method with Adaptive region correlation Discovery (READ). Specifically, READ comprises three modules, including a disentangled region feature learning module utilizing a city-context Transformer to encode regions' semantic and mobility features, and an adaptive weighted multi-graph construction module that builds multiple complete graphs with learnable weights based on disentangled features of regions. In addition, we propose a multi-graph representation learning module to yield effective region representations that integrate information from multiple graphs. We conduct thorough experiments on three downstream tasks to assess READ. Experimental results demonstrate that READ considerably outperforms state-of-the-art baseline methods in urban region embedding.
Meng Chen 0003, Hongwei Jia, Zechen Li 0003, Weiming Huang 0001, Kai Zhao 0011, Yongshun Gong, Hongjun Dai
IEEE Trans. Knowl. Data Eng.8
2025 Black-Box Test-Time Prompt Tuning for Vision-Language Models
abstract
Test-time prompt tuning (TPT) aims to adjust the vision-language models (e.g., CLIP) with learnable prompts during the inference phase. However, previous works overlooked that pre-trained models as a service (MaaS) have become a noticeable trend due to their commercial usage and potential risk of misuse. In the context of MaaS, users can only design prompts in inputs and query the black-box vision-language models through inference APIs, rendering the previous paradigm of utilizing gradient for prompt tuning is infeasible. In this paper, we propose black-box test-time prompt tuning (B²TPT), a novel framework that addresses the challenge of optimizing prompts without gradients in an unsupervised manner. Specifically, B²TPT designs a consistent or confident (CoC) pseudo-labeling strategy to generate high-quality pseudo-labels from the outputs. Subsequently, we propose to optimize low-dimensional intrinsic prompts using a derivative-free evolution algorithm and to project them onto the original text and vision prompts. This strategy addresses the gradient-free challenge while reducing complexity. Extensive experiments across 15 datasets demonstrate the superiority of B²TPT. The results show that B²TPT not only outperforms CLIP's zero-shot inference at test time, but also surpasses other gradient-based TPT methods.
Fan'an Meng, Chaoran Cui, Hongjun Dai, Shuai Gong
AAAI3
2025 DebateNav: Structured Multi-VLM Expert Debate for Robust Zero-Shot Object Navigation
abstract
Zero-shot object navigation presents a highly challenging task in embodied AI, requiring an agent to interpret natural language instructions, perceive complex visual environments, and plan actions without any task-specific training. While recent approaches have introduced large language models (LLMs) as high-level planners, they often rely on static, one-shot inference and struggle with ambiguous or partially observable scenes. This paper proposes DebateNav, a novel multi-agent decision framework that integrates multiple vision-language model (VLM) experts under the supervision of a central LLM controller. Each VLM is assigned a unique expert role (e.g., object detection, risk assessment, spatial reasoning), and together they engage in structured multi-round debates when perception conflicts arise. The LLM controller performs task decomposition, memory-guided exploration, and final arbitration based on expert arguments. To enhance the perception and decision process, DebateNav incorporates a multimodal image fusion module combining RGB, depth, and segmentation inputs, as well as a map memory and trajectory tracking system that helps avoid redundant exploration and supports long-horizon planning. The system is evaluated on a subset of the HM3D dataset with approximately 5,000 tasks, achieving a$\mathbf{5 2. 3 \%}$success rate and$\mathbf{1 5. 5}$SPL under strict zero-shot conditions. Extensive ablation studies confirm the effectiveness of the expert debate mechanism, multi-modal fusion, memory system, and LLM-based arbitration. The results demonstrate that DebateNav outperforms several recent baselines and establishes a new perspective on collaborative, interpretable planning for zero-shot embodied navigation.
Henghui Sun, Weixing Tan, Lei Liu 0003, Zhongmin Yan, Xudong Lu 0001, Hongjun Dai
HPCC6
2025 A Hybrid Pipeline and Large Language Model System for Task-Oriented Dialogue
abstract
In task-oriented dialogue systems, intent recognition and entity extraction are key for driving system understanding and state updates. However, traditional structured systems often show limited robustness and generalization when facing complex expressions. To improve understanding in these scenarios, this paper proposes a hybrid dialogue system framework. The framework fuses a traditional pipeline architecture with a large language model (LLM). It is built on a conventional pipeline. It evaluates the uncertainty of pipeline predictions online. When confidence falls below a preset threshold, it automatically invokes the LLM for semantic enhancement and reanalysis. This dynamic step compensates for the shortcomings of the structured process. To ensure controllability and structural alignment of generative outputs, we design a dual-control mechanism. The mechanism integrates a domain-adaptive prompt selection strategy with an output normalization protocol. This design significantly improves LLM accuracy and consistency in entity extraction. Experiments on the MultiWOZ 2.2 dataset show that the hybrid framework outperforms traditional structured methods. It handles complex inputs and colloquial queries more effectively.
Yihan Zheng, Weixing Tan, Lei Liu 0003, Zhongmin Yan, Hongjun Dai
HPCC6
2025 Cross-City Latent Space Alignment for Consistency Region Embedding
abstract
Learning urban region embeddings has substantially advanced urban analysis, but their typical focus on individual cities leads to disparate embedding spaces, hindering cross-city knowledge transfer and the reuse of downstream task predictors. To tackle this issue, we present Consistency Region Embedding (CoRE), a unified framework integrating region embedding learning with cross-city latent space alignment. CoRE first embeds regions from two cities into separate latent spaces, followed by the alignment of latent space manifolds and fine-grained individual regions from both cities. This ensures compatible and comparable embeddings within aligned latent spaces, enabling predictions of various socioeconomic indicators without ground truth labels by migrating knowledge from label-rich cities. Extensive experiments show CoRE outperforms competitive baselines, confirming its effectiveness for cross-city knowledge transfer via aligned latent spaces.
Meng Chen 0003, Hongwei Jia, Zechen Li 0003, Wenzhen Jia, Kai Zhao 0011, Hongjun Dai, Weiming Huang 0001
ICML6
2025 A High-Performance AI Processor Architecture: Integrating Multi-controller with Hybrid DDR Memory
Zixuan Ding, Hongjun Dai
WASA (1)5
2025 vGPU Performance Testing Framework for Large Model Inference
Youli Zhang, Hongjun Dai, Huifeng Liu
WASA (2)3
2025 CATScaler: A Convolution-Augmented Transformer Scaling Framework for Cloud-Native Applications
abstract
Efficient container scaling is crucial for enhancing the availability and scalability of cloud-native applications through adaptive resource management. In cloud computing, the default autoscaling feature of Kubernetes scales pods only when the cluster or application exceeds a predefined threshold. However, this reactive approach often leads to significant resource waste during demand fluctuations because it cannot predict future workload changes and adjust resources in advance. This paper presents CATScaler, a novel Convolution-Augmented Transformer Scaler designed to proactively optimize resource allocation in serverless environments. CATScaler is a proactive approach composed of two modules: workload prediction and elastic auto-scaling. In the prediction module, we develop a convolution-augmented transformer to accurately predict workload changes at both local and global levels. Additionally, we incorporate reversible instance normalization to mitigate the shift caused by the difference between workload data and training data. In the auto-scaling module, we implement an instance-counting method to handle the nonlinear relationships between variables. Experiments using two real datasets from Alibaba Cloud and Huawei Cloud demonstrate the effectiveness of CATScaler. The tests conducted on a cluster of 4 servers demonstrated that CATScaler reduced response time latency by 1.1× compared to Kubernetes' default scaler and decreased service violation rates by 3.2×.
Fan'an Meng, Hongjun Dai, Guoqing Cong, Hailiang Zhao
IEEE Trans. Serv. Comput.2
2024 sEMG-Based Multi-view Feature-Constrained Representation Learning
Hongjun Dai
KSEM (1)2
2024 Energy Consumption Prediction Method for Refrigeration Systems Based on Adversarial Networks and Transformer Networks
Huifeng Liu, Youli Zhang, Hongjun Dai, Minghao Shao, Hongyu Xu
KSEM (5)5
2024 TryonCM2: Try-on-Enhanced Fashion Compatibility Modeling Framework
abstract
Recently, fashion compatibility modeling, which can score the matching degree of several complementary fashion items, has gained increasing research attention. Previous studies have primarily learned the features of fashion items and utilize their interaction as the fashion compatibility. However, the try-on looking of an outfit help us to learn the fashion compatibility in a combined manner, where items are spatially distributed and partially covered by other items. Inspired by this, we design a try-on-enhanced fashion compatibility modeling framework, named TryonCM2, which incorporates the try-on appearance with the item interaction to enhance the fashion compatibility modeling. Specifically, we treat each outfit as a sequence of items and adopt the bidirectional long short-term memory (LSTM) network to capture the latent interaction of fashion items. Meanwhile, we synthesize a try-on template image to depict the try-on appearance of an outfit. And then, we regard the outfit as a sequence of multiple image stripes, i.e., local content, of the try-on template, and adopt the bidirectional LSTM network to capture the contextual structure in the try-on appearance. Ultimately, we combine the fashion compatibility lying in the item interaction and try-on appearance as the final compatibility of the outfit. Both the objective and subjective experiments on the existing FOTOS dataset demonstrate the superiority of our framework over the state-of-the-art methods.
Xuemeng Song, Jianlong Wu, Hongjun Dai, Liqiang Nie
IEEE Trans. Neural Networks Learn. Syst.5
2023 A hardware-independent time estimation method for inference process of convolutional layers on GPU
Chengzhen Meng, Hongjun Dai
Perform. Evaluation2
2021 Optimization of Remote Desktop with CNN-based Image Compression Model
Hejun Wang, Hongjun Dai, Meikang Qiu, Meiqin Liu 0001
KSEM2
2020 A Heuristic Services Binding Algorithm to Improve Fault-Tolerance in Microservice based Edge Computing Architecture
abstract
In microservice architecture, Mobile Edge Computing (MEC) has been widely developed to improve Quality of Service (QoS). This distributed technology deploys different microservices to intricate network environment. A fault-tolerant deployment and scheduling scheme is significant for overall architecture to enhance robustness. Especially in some time-sensitive service scenarios, any network fluctuation caused services invocation failures may lead to irreparable loss. This work introduces a heuristic based services binding algorithm to improve fault-tolerant MEC in microservice architecture with the help of Cache-enabled Edge Nodes. This challenge is modeled as a constrained optimization problem on services composition and network nodes graphs. In addition, this work proposes a graph based state and action value functions to heuristically generate solutions. In overview system, the suitable combination will balance the performance and robustness.
Hongjun Dai
SERVICES2
2020 Fashion Compatibility Modeling through a Multi-modal Try-on-guided Scheme
abstract
Recent years have witnessed a growing trend of fashion compatibility modeling, which scores the matching degree of the given outfit and then provides people with some dressing advice. Existing methods have primarily solved this problem by analyzing the discrete interaction among multiple complementary items. However, the fashion items would present certain occlusion and deformation when they are worn on the body. Therefore, the discrete item interaction cannot capture the fashion compatibility in a combined manner due to the neglect of a crucial factor: the overall try-on appearance. In light of this, we propose a multi-modal try-on-guided compatibility modeling scheme to jointly characterize the discrete interaction and try-on appearance of the outfit. In particular, we first propose a multi-modal try-on template generator to automatically generate a try-on template from the visual and textual information of the outfit, depicting the overall look of its composing fashion items. Then, we introduce a new compatibility modeling scheme which integrates the outfit try-on appearance into the traditional discrete item interaction modeling. To fulfill the proposal, we construct a large-scale real-world dataset from SSENSE, named FOTOS, consisting of 11,000 well-matched outfits and their corresponding realistic try-on images. Extensive experiments have demonstrated its superiority to state-of-the-arts.
Jianlong Wu, Xuemeng Song, Hongjun Dai, Liqiang Nie
SIGIR4
2020 A service recommendation algorithm with the transfer learning based matrix factorization to improve cloud security
Hongjun Dai, Zhilou Yu, Rui Li 0090
Inf. Sci.2
2020 A Trust Verification Architecture with Hardware Root for Secure Clouds
abstract
Cloud security has become a vital issue within thousands of inter-connected servers in clouds, as malicious attacks or discovered vulnerabilities may spread more rapidly than ever. Based on the opinion that hardware is more secure and trustworthy, a trust platform module (TPM) is used as an external chip to ensure the trust verification, while it's unsuitable as virtual machine (VM) migration, hybrid servers, distributed storage with a low performance. So, we design a novel cloud architecture with a special physical server named as the trust verification server (TVS) to provide trust services according to the TPM specification, then the servers in the cloud can use TVS remotely as a high-performance TPM chip. In this paper, we design the TVS with accelerator hardware, upgrade the cloud architecture with an additional certificate authority (CA) server, and use TVS with a non-interference trust measurement model. The experiments show that the TVS can work efficiently with huge performance improvements at more than 100 times compared with the use of TPM in the cloud. This can be used to solve the complex cloud security problems such as VM sprawl and VM escape.
Zhilou Yu, Hongjun Dai, Xiaoming Xi, Meikang Qiu
IEEE Trans. Sustain. Comput.2
2019 A scheduling algorithm for autonomous driving tasks on mobile edge computing servers
Hongjun Dai, Zhilou Yu
J. Syst. Archit.1
2018 Set variation-aware shared LLC management for CPU-GPU heterogeneous architecture
abstract
Heterogeneous CPU-GPU multiprocessor systems-on-chip (HMPSoC) becomes a popular architecture choice for high performance embedded systems, where shared last-level cache (LLC) management becomes a critical design consideration. We observe that within a sampling period, CPU and GPU may have distinct access behaviors over various LLC sets. In this work, we propose a light-weighted and fined-grained cache management policy to cope with the CPU-GPU access behavior variation among cache sets. In particular, CPU and GPU requests are prioritized disparately in each LLC set during cache block insertion and promotion, based on the per-core utility behaviors and a per-set CPU-GPU miss counter. Experimental results show that our LLC management scheme outperforms the two state-of-the-art schemes TAP-RRIP and LSP by 12.6% and 10.01%, respectively.
Zhaoying Li 0004, Lei Ju 0001, Hongjun Dai, Mengying Zhao, Zhiping Jia
DATE3
2018 A distributed multi-level model with dynamic replacement for the storage of smart edge computing
Jiarong Xing, Hongjun Dai, Zhilou Yu
J. Syst. Archit.2
2017 A chaos-oriented prediction and suppression model to enhance the security for cyber physical power systems
Hongjun Dai, Shulin Zhao 0002
J. Parallel Distributed Comput.1
2017 Shared write buffer to boost applications on SpMT architecture
John M. Ye, Tianzhou Chen, Hongjun Dai
J. Supercomput.4
2016 Explore prediction for instruction level redundant execution in fault tolerant microprocessors
Hongjun Dai
J. Syst. Archit.1
2015 A big data inspired chaotic solution for fuzzy feedback linearization model in cyber-physical systems
Lei Liu 0003, Shulin Zhao 0002, Zhilou Yu, Hongjun Dai
Ad Hoc Networks4
2015 Dynamic malicious node detection with semi-supervised multivariate classification in cognitive wireless sensor networks
abstract
Summary Usually, wireless sensor networks are distributed massively with a number of nodes in an open large‐scale environment, and they are vulnerable to malicious attacks because the communications change dynamically and unpredictably. In this paper, we present a detection method based on multivariate classification to find out the malicious sensor nodes. It learns the features of a few type‐known node, classifies them with dynamical multivariate classification, and then establishes the sample space of all sensor nodes in the network activities to deduce the malicious nodes. The experiment results show that as long as the value of sensor node preferences and the number of active sensor nodes is stable, the false detection rate is stabilized below 0.5%. This proves that the algorithm can be used to the cognitive wireless sensor networks widely. Copyright © 2014 John Wiley & Sons, Ltd.
Hongjun Dai, Huabo Liu, Zhiping Jia
Concurr. Comput. Pract. Exp.1
2015 Security enhancement of cloud servers with a redundancy-based fault-tolerant cache structure
Hongjun Dai, Shulin Zhao 0002, Jiutian Zhang, Meikang Qiu, Lixin Tao
Future Gener. Comput. Syst.1
2014 A piecewise geometry method for optimizing the motion planning of data mule in tele-health wireless sensor networks
Hongjun Dai, Zhiping Jia, Meikang Qiu, Bin Wang 0002
Wirel. Networks2
2012 A Multivariate Classification Algorithm for Malicious Node Detection in Large-Scale WSNs
abstract
WSN is a distributed network exposed to an open environment, which is vulnerable to malicious nodes. To find out malicious nodes among a WSN with mass sensor nodes, this paper presents a malicious detection method based on multi-variate classification. Given the types of a few sensor nodes, it extracts sensor nodes' preferences related with the known types of malicious node, establishes the sample space of all sensor nodes that participate in network activities. Then, according to the study on the type-known sensor nodes' samples based on the multivariate classification algorithm, a classifier is generated, and all of the unknown-type sensor nodes are classified. The experiment results show that as long as the value of sensor nodes preferences and the number of active sensor nodes is stable, the false detection rate is stabilized under 0.5%.
Hongjun Dai, Huabo Liu, Zhiping Jia, Tianzhou Chen
TrustCom1
2011 Verification-Based Multi-backup Firmware Architecture, an Assurance of Trusted Boot Process for the Embedded Systems
abstract
NAND flash has been widely used as the only non-volatile storage device in the embedded systems. However, it has high rates of bad block, which may lead the stored programs damaged. Especially for the firmware including bootloader and OS, this will lead the system crash immediately. This paper proposes a novel verification-based multi-backup firmware architecture (VMFA) to improve the reliability with the multiple copies of firmware in NAND flash. According to the theory of chain of trust, during the boot process, the integrity of one program should be checked before it gets the right to execute, and the program can be executed only on condition that its integrity is valid. Meanwhile, the system can automatically load and measure the backup copies and verify the integrity when the original program is damaged. Some experiments are taken on a real development platform and the VMFA is measured with time module to analyze the boot time. The results show that the system can work well with VMFA and the boot process can be ensured with the suitable verifications.
Hongfei Yin, Hongjun Dai, Zhiping Jia
TrustCom2
2010 A Fault-tolerant Architecture with Error Correcting Code for the Instruction-level Temporal Redundancy
abstract
Soft error has become an increasingly significant problem in modern computing systems. To overcome soft errors, it has reported that the instruction-level temporal redundancy in out-of-order cores suffers a performance penalty up to 45%. In this work, we propose the fault-tolerant double execution architecture with the fast error correcting code (such as two-dimensional error code) in the instruction reuse buffer. Experimental results show that it gains back IPC loss between 9.14% and 10.15%, with an average around 9.22% compared with the conventional double execution approach.
Hongjun Dai, Tianzhou Chen, Meikang Qiu
EUC2