Zhaoxiang Liu

dblp:93/2654 · DBLP profile ↗
← Back
34ranked-venue papers
5as first author
33since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 19 · 1 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 9 since 2021Systems, architecture and hardware · 6 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 5 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 HiMo-CLIP: Modeling Semantic Hierarchy and Monotonicity in Vision-Language Alignment
abstract
Contrastive vision-language models like CLIP have achieved impressive results in image-text retrieval by aligning image and text representations in a shared embedding space. However, these models often treat text as flat sequences, limiting their ability to handle complex, compositional, and long-form descriptions. In particular, they fail to capture two essential properties of language: semantic hierarchy, which reflects the multi-level compositional structure of text, and semantic monotonicity, where richer descriptions should result in stronger alignment with visual content. To address these limitations, we propose HiMo-CLIP, a representation-level framework that enhances CLIP-style models without modifying the encoder architecture. HiMo-CLIP introduces two key components: a hierarchical decomposition (HiDe) module that extracts latent semantic components from long-form text via in-batch PCA, enabling flexible, batch-aware alignment across different semantic granularities, and a monotonicity-aware contrastive loss (MoLo) that jointly aligns global and component-level representations, encouraging the model to internalize semantic ordering and alignment strength as a function of textual completeness. These components work together to produce structured, cognitively aligned cross-modal representations. Experiments on multiple image-text retrieval benchmarks show that HiMo-CLIP consistently outperforms strong baselines, particularly under long or compositional descriptions.
Ruijia Wu, Fei Shen 0004, Shaoan Zhao, Qiang Hui, Huanlin Gao, Zhaoxiang Liu, Kai Wang 0012, Shiguo Lian
AAAI8
2026 HumorReject: Decoupling LLM Safety from Refusal Prefix via a Little Humor
abstract
Large Language Models (LLMs) commonly rely on explicit refusal prefixes for safety, making them vulnerable to prefix injection attacks. We introduce HumorReject, a novel data-driven approach that reimagines LLM safety by decoupling it from refusal prefixes through humor as an indirect refusal strategy. Rather than explicitly rejecting harmful instructions, HumorReject responds with contextually appropriate humor that naturally defuses potentially dangerous requests. Our approach effectively addresses common "over-defense" issues while demonstrating superior robustness against various attack vectors. Our findings suggest that improvements in training data design can be as important as the alignment algorithm itself in achieving effective LLM safety.
Zihui Wu, Haichang Gao, Jiacheng Luo, Zhaoxiang Liu
AAAI4
2026 GlitchMiner: Mining Glitch Tokens in Large Language Models via Gradient-based Discrete Optimization
abstract
Glitch tokens—inputs that trigger unpredictable or anomalous behavior in Large Language Models (LLMs)—pose significant challenges to model reliability and safety. Existing detection methods primarily rely on heuristic embedding patterns or statistical anomalies within internal representations, limiting their generalizability across different model architectures and potentially missing anomalies that deviate from observed patterns. We introduce GlitchMiner, an behavior-driven framework designed to identify glitch tokens by maximizing predictive entropy. Leveraging a gradient-guided local search strategy, GlitchMiner efficiently explores the discrete token space without relying on model-specific heuristics or large-batch sampling. Extensive experiments across ten LLMs from five major model families demonstrate that GlitchMiner consistently outperforms existing approaches in detection accuracy and query efficiency, providing a generalizable and scalable solution for effective glitch token discovery.
Zihui Wu, Haichang Gao, Ping Wang 0003, Shudong Zhang, Zhaoxiang Liu, Shiguo Lian
AAAI5
2026 Enhanced data techniques and optimization in conversational gesture generation
Xiang Wang 0018, Yifeng Peng, Zhaoxiang Liu, Kai Wang 0012, Shiguo Lian
CCF Trans. Pervasive Comput. Interact.3
2026 ADI-SAM: Adapting segment anything model for degraded images
Yang Zhao 0028, Zhaoxiang Liu, Yibing Nan, Ke Lu 0002, Shiguo Lian
Neurocomputing2
2026 KAConvNet: Kolmogorov-Arnold convolutional networks for vision recognition
Zhaoxiang Liu, Zhicheng Ma, Kaikai Zhao, Kai Wang 0012, Shiguo Lian
Image Vis. Comput.1
2026 RelPose-TTA: Energy-based relative pose correction for test-time adaptation of category-level object pose estimation
Yue Zhan, Xin Wang 0135, Zhaoxiang Liu, Shiguo Lian, Tangwen Yang
Image Vis. Comput.3
2025 DMFF-EPI: A Dual-Modality Feature Fusion Network with Contrastive Learning for Enhancer-Promoter Interaction Prediction
abstract
Enhancer-Promoter Interactions (EPIs) play a pivotal role in transcriptional gene regulation, and their precise identification is critical for understanding disease mechanisms. Despite significant advances in computational models, achieving robust cross-cell-line generalization and comprehensive feature representation remains an open challenge. To tackle this, we propose DMFF-EPI, a dual-modality feature fusion network enhanced with contrastive learning for enhancer-promoter interaction prediction. Specifically, DMFF-EPI constructs parallel feature extraction modules for enhancer and promoter sequences. One branch combines a convolutional neural network (CNN) with a bidirectional gated recurrent unit (BiGRU) to capture both local motifs and long-range sequential dependencies. The parallel branch employs a multilayer perceptron (MLP) to encode k-mer frequency embeddings as prior structural knowledge. The core innovation lies in a contrastive learning-based regularizer that optimizes the feature space by aligning two complementary representations derived from the same sample-dynamic sequence features and static frequency features. This alignment enforces intra-sample consistency and improves model robustness across cell lines. A joint loss function combining binary cross-entropy for EPI classification and contrastive loss for feature alignment enables more stable training and stronger generalization. We conduct extensive performance evaluations on six publicly available cell lines (GM12878, K562, HeLa-S3, IMR90, HUVEC, and NHEK). The results suggest that DMFF-EPI achieves superior performance compared to existing baseline models in terms of both AUROC and AUPR. Cross-cell-line prediction experiments further evaluate the model's generalization to unseen cellular contexts. The results demonstrate that DMFF-EPI maintains competitive performance even under challenging cross-cell-line scenarios, highlighting its robustness and broad applicability across diverse biological contexts. The source code is publicly available at https://github.com/Wujiahui95/DMFF-EPI.
Amu Erermeri, Zhaoxiang Liu, Haitao Fu
BIBM3
2025 Optimizing for the Shortest Path in Denoising Diffusion Model
abstract
In this research, we propose a novel denoising diffusion model based on shortest-path modeling that optimizes residual propagation to enhance both denoising efficiency and quality. Drawing on Denoising Diffusion Implicit Models (DDIM) and insights from graph theory, our model, termed the Shortest Path Diffusion Model (ShortDF), treats the denoising process as a shortest-path problem aimed at minimizing reconstruction error. By optimizing the initial residuals, we improve the efficiency of the reverse diffusion process and the quality of the generated samples. Extensive experiments on multiple standard benchmarks demonstrate that ShortDF significantly reduces diffusion time (or steps) while enhancing the visual fidelity of generated samples compared to prior arts. This work, we suppose, paves the way for interactive diffusion-based applications and establishes a foundation for rapid data generation. Code is available at https://github.com/UnicomAI/ShortDF.
Xingpeng Zhang, Zhaoxiang Liu, Kai Wang 0012, Min Wang 0031, Yanlin Qian, Shiguo Lian
CVPR3
2025 ILearnRobot: An Interactive Learning-Based Multi-modal Robot with Continuous Improvement
Kohou Wang, Zhaoxiang Liu, Kai Wang 0012, Shiguo Lian
ICIC (14)2
2025 Art3D-Fusion: A Hybrid Framework for Visual Synthesis with Artistic Control
Kohou Wang, Zhaoxiang Liu, Zezhou Chen, Xin Wang 0135, Kai Wang 0012, Shiguo Lian
ICIG (1)3
2025 Data Leakage Detection in Large Vision-Language Models via Multimodal Perturbation
Xin Wang 0135, Zhaoxiang Liu, Yue Zhan, Kaikai Zhao, Kai Wang 0012, Shiguo Lian
ICIG (1)2
2025 CP3: Customizable 3D Pop-Out Effect Creation for Immersive Content Using Multimodal Models
abstract
In this paper, a multi-modal model based 3D pop-out video generation framework (CP3) is proposed to solve the shortcomings of the existing video generation technology for accurate control of 3D pop-out effects. 3D pop-out effects create an immersive visual experience by changing the disparity of a particular object so that it appears beyond the screen. However, although software has made some progress in this area, there is currently no effective way to accurately control 3D pop-out effects and generate high-quality video. In addition, the lack of high-quality 3D pop-out effect data sets is also one of the bottlenecks in the field. Therefore, the CP3 framework proposed in this paper utilizes multi-modal models to help 3D video creators make 3D pop-out effects, enhance the audience's sense of immersion and visual comfort, and thus promote the development of 3D effect generation technology. To support the training and evaluation of this framework, a new dataset containing 37000 frames of pop-out effects is constructed, such as text guidance, segmentation results, depth maps, optical flow, and the trajectory of the pop-out target. Through the 3D UNet model based on the potential de-noising diffusion mechanism, combined with the 3D-try module in the CP3 framework and Mask Encoder, this paper has achieved remarkable results in the generation of 3D pop-out effect videos. The results of the experiment show that the CP3 framework demonstrates its advantages in generating immersive 3D pop-out effects in comparison to existing technologies.
Zezhou Chen, Zhaoxiang Liu, Kai Wang 0012, Shiguo Lian
ACM Multimedia6
2025 LeMiCa: Lexicographic Minimax Path Caching for Efficient Diffusion-Based Video Generation
abstract
We present LeMiCa, a training-free and efficient acceleration framework for diffusion-based video generation. While existing caching strategies primarily focus on reducing local heuristic errors, they often overlook the accumulation of global errors, leading to noticeable content degradation between accelerated and original videos. To address this issue, we formulate cache scheduling as a directed graph with error-weighted edges and introduce a Lexicographic Minimax Path Optimization strategy that explicitly bounds the worst-case path error. This approach substantially improves the consistency of global content and style across generated frames. Extensive experiments on multiple text-to-video benchmarks demonstrate that LeMiCa delivers dual improvements in both inference speed and generation quality. Notably, our method achieves a 2.9× speedup on the Latte model and reaches an LPIPS score of 0.05 on Open-Sora, outperforming prior caching techniques. Importantly, these gains come with minimal perceptual quality degradation, making LeMiCa a robust and generalizable paradigm for accelerating diffusion-based video generation. We believe this approach can serve as a strong foundation for future research on efficient and reliable video synthesis.
Huanlin Gao, Fuyuan Shi, Zhaoxiang Liu, Kai Wang 0012, Shiguo Lian
NeurIPS5
2025 DistRMI: a deep distance-aware neural network for explainable RNA loop motif-small molecule interaction prediction
abstract
RNA participates in the occurrence and development of various diseases by regulating gene expression. Owing to its potential to circumvent the limitations of traditional "undruggable" protein targets, it is regarded as a core direction for next-generation precision therapy. Against this backdrop, the accurate and interpretable prediction of RNA-small molecule interactions has become a key link in accelerating the discovery of RNA-targeted drugs. However, existing methods suffer from insufficient prediction accuracy and interpretability, failing to effectively guide lead compound screening or elucidate the mechanism of action. This study presents DistRMI, which integrates Transformers and graph neural networks to capture, respectively, the sequence information of RNA loop motifs and the chemical topological features of small molecules while introducing distance priors between them and leveraging a distance-aware attention mechanism to capture their interaction information. The results show that DistRMI outperforms baseline models, and its performance remains robust even when confronted with unknown RNA loop motifs and small molecules. Visualization of attention weights reveals that bases near the paired bases of RNA loop motifs contribute significantly. Furthermore, retrospective case studies validate the model's reliability. Predicting the binding preferences between RNA loop motifs and small molecules while providing interpretability facilitates an in-depth understanding of RNA-small molecule interactions, promotes in-depth research on RNA and related drugs, and opens up new avenues for disease treatment.
Zhaoxiang Liu, Qiqi Zhu, Qingyan Tian, Yingxiang Deng, Dengguo Wei, Haitao Fu
Briefings Bioinform.1
2025 PSTF-AttControl: Per-subject-tuning-free personalized image generation with controllable face attributes
Zhaoxiang Liu, Zezhou Chen, Kai Wang 0012, Shiguo Lian
Image Vis. Comput.2
2025 Hybrid Attention Transformers with fast Fourier convolution for light field image super-resolution
Zhicheng Ma, Yuduo Guo, Zhaoxiang Liu, Shiguo Lian, Sen Wan
Image Vis. Comput.3
2025 MITS: A large-scale multimodal benchmark dataset for Intelligent Traffic Surveillance
Kaikai Zhao, Zhaoxiang Liu, Xin Wang 0135, Zhicheng Ma, Yajun Xu, Wenjing Zhang 0006, Yibing Nan, Kai Wang 0012, Shiguo Lian
Image Vis. Comput.2
2024 Microscope: Causality Inference Crossing the Hardware and Software Boundary from Hardware Perspective
abstract
The increasing complexity of System-on-Chip (SoC) designs and the rise of third-party vendors in the semiconductor industry have led to unprecedented security concerns. Traditional formal methods struggle to address software-exploited hardware bugs, and existing solutions for hardware-software co-verification often fall short. This paper presents Microscope, a novel framework for inferring software instruction patterns that can trigger hardware vulnerabilities in SoC designs. Microscope enhances the Structural Causal Model (SCM) with hardware features, creating a scalable Hardware Structural Causal Model (HW-SCM). A domain-specific language (DSL) in SMT-LIB represents the HW-SCM and predefined security properties, with incremental SMT solving deducing possible instructions. Microscope identifies causality to determine whether a hardware threat could result from any software events, providing a valuable resource for patching hardware bugs and generating test input. Extensive experimentation demonstrates Microscope’s capability to infer the causality of a wide range of vulnerabilities and bugs located in SoC-level benchmarks.
Zhaoxiang Liu, Kejun Chen, Dean Sullivan, Orlando Arias, Raj Gautam Dutta, Yier Jin, Xiaolong Guo 0001
ASPDAC1
2024 Poster: BlindMarket: A Trustworthy Chip Designs Marketplace for IP Vendors and Users
abstract
Due to the globalization of the semiconductor supply chain, chip fabrication now involves multiple parties, including intellectual property (IP) vendors and Electronic Design Automation (EDA) tool vendors. Involving multiple entities and valuable IP naturally raises security and privacy concerns. Various frameworks and tools, such as the IEEE 1735 standard for IP protection, have been developed to mitigate the risk of theft. However, existing solutions fail to address all the threats envisioned by the zero-trust model. We propose a novel zero-trust formal verification framework that requires only two essential parties: IP users and IP vendors. This framework leverages secure multiparty computation to ensure the security and privacy of the hardware verification process. Our proposed solution allows IP users and IP vendors to independently convert the hardware design and assertions into conjunctive normal form (CNF), and then apply privacy-preserving SAT solving to verify the conformance of the design to the specification. This paper introduces a domain-specific secure decision procedure, hw-ppSAT, designed to overcome the scalability challenges of using SAT solving in hardware design verification. Our approach also leverages property-based hardware optimizations and domain-specific heuristics to enhance the verification process. We showcase the framework's effectiveness through its application to several open-source benchmarks.
Zhaoxiang Liu, Ning Luo 0002, Samuel Judson, Raj Gautam Dutta, Xiaolong Guo 0001, Mark Santolucito
CCS1
2024 A Multimodal Benchmark Dataset and Model for Crop Disease Diagnosis
Zhaoxiang Liu, Zezhou Chen, Kohou Wang, Kai Wang 0012, Shiguo Lian
ECCV (86)2
2024 DTjRTL: A Configurable Framework for Automated Hardware Trojan Insertion at RTL
abstract
Shifts in the IC supply chain have necessitated outsourcing design or fabrication to third-party vendors, introducing various hardware security issues, notably Hardware Trojans (HTs) as a prominent risk. The research in detecting and preventing HTs faces challenges due to the lack of standardized benchmarks and measurements. This paper introduces a framework to automatically generate dynamic functional HTs in a configurable and systematical manner at Register Transfer Level (RTL). The objective is not to produce HTs that are difficult to activate but to systematically create a diverse set of HT designs. This approach serves dual purposes: it aids the research community in testing their detection frameworks and facilitates buggy design benchmark creation for competitive exercises between blue and red teams. Our framework accepts RTL designs and configuration parameters, automating the generation of HT-inserted designs at RTL. We present an evaluation of the generated HT designs focusing on hardware cost overhead and post-synthesis survivability by verifying HT presence at both RT and gate levels. Results indicate that HTs employing only combinational logic are easier to optimize away but result in lower overhead compared to HTs that incorporate additional sequential logic.
Ruochen Dai, Zhaoxiang Liu, Orlando Arias, Xiaolong Guo 0001, Tuba Yavuz
ACM Great Lakes Symposium on VLSI2
2024 Spatial-Temporal Transformer Network for Continuous Action Recognition in Industrial Assembly
Shanghua Tang, Shaoan Zhao, Yimin Lin, Kai Wang 0012, Zhaoxiang Liu, Shiguo Lian
ICIC (10)9
2024 Self-supervised Visual Anomaly Detection with Image Patch Generation and Comparison Networks
Kaikai Zhao, Yimin Lin, Zhaoxiang Liu, Kai Wang 0012, Shiguo Lian
ICIC (10)5
2024 TP3M: Transformer-based Pseudo 3D Image Matching with Reference Image
abstract
Image matching is still challenging in such scenes with large viewpoints or illumination changes or with low textures. In this paper, we propose a Transformer-based pseudo 3D image matching method. It upgrades the 2D features extracted from the source image to 3D features with the help of a reference image and matches to the 2D features extracted from the destination image by the coarse-to-fine 3D matching. Our key discovery is that by introducing the reference image, the source image’s fine points are screened and furtherly their feature descriptors are enriched from 2D to 3D, which improves the match performance with the destination image. Experimental results on multiple datasets show that the proposed method achieves the state-of-the-art on the tasks of homography estimation, pose estimation and visual localization especially in challenging scenes.
Liming Han, Zhaoxiang Liu, Shiguo Lian
ICRA2
2024 A Large Vision-Language Model based Environment Perception System for Visually Impaired People
abstract
It is a challenging task for visually impaired people to perceive their surrounding environment due to the complexity of the natural scenes. Their personal and social activities are thus highly limited. This paper introduces a Large Vision-Language Model(LVLM) based environment perception system which helps them to better understand the surrounding environment, by capturing the current scene they face with a wearable device, and then letting them retrieve the analysis results through the device. The visually impaired people could acquire a global description of the scene by long pressing the screen to activate the LVLM output, retrieve the categories of the objects in the scene resulting from a segmentation model by tapping or swiping the screen, and get a detailed description of the objects they are interested in by double-tapping the screen. To help visually impaired people more accurately perceive the world, this paper proposes incorporating the segmentation result of the RGB image as external knowledge into the input of LVLM to reduce the LVLM’s hallucination. Technical experiments on POPE, MME and LLaVA-QA90 show that the system could provide a more accurate description of the scene compared to Qwen-VL-Chat, exploratory experiments show that the system helps visually impaired people to perceive the surrounding environment effectively.
Zezhou Chen, Zhaoxiang Liu, Kai Wang 0012, Kohou Wang, Shiguo Lian
IROS2
2024 Reparameterization-Based Parameter-Efficient Fine-Tuning Methods for Large Language Models: A Systematic Survey
Zezhou Chen, Zhaoxiang Liu, Kai Wang 0012, Shiguo Lian
NLPCC (3)2
2024 What is the Best Model? Application-Driven Evaluation for Large Language Models
Shiguo Lian, Kaikai Zhao, Xuejiao Lei, Bikun Yang, Wenjing Zhang 0006, Kai Wang 0012, Zhaoxiang Liu
NLPCC (3)8
2024 Optimized Conversational Gesture Generation with Enhanced Motion Feature Extraction and Cascaded Generator
Xiang Wang 0018, Yifeng Peng, Zhaoxiang Liu, Shijie Dong, Ruitao Liu, Kai Wang 0012, Shiguo Lian
NLPCC (3)3
2024 Hybrid attention transformer with re-parameterized large kernel convolution for image super-resolution
Zhicheng Ma, Zhaoxiang Liu, Kai Wang 0012, Shiguo Lian
Image Vis. Comput.2
2023 Patch-Wise Auto-Encoder for Visual Anomaly Detection
abstract
Anomaly detection without priors of the anomalies is challenging. In the field of unsupervised anomaly detection, traditional auto-encoder (AE) tends to fail based on the assumption that by training only on normal images, the model will not be able to reconstruct abnormal images correctly. On the contrary, we propose a novel patch-wise auto-encoder (Patch AE) framework, which aims at enhancing the reconstruction ability of AE to anomalies instead of weakening it. Each patch of image is reconstructed by corresponding spatially distributed feature vector of the learned feature representation, i.e., patch-wise reconstruction, which ensures anomaly- sensitivity of AE. Our method is simple and efficient. It advances the state-of-the-art performances on Mvtec AD benchmark, which proves the effectiveness of our model. It shows great potential in practical industrial application scenarios.
Yajie Cui, Zhaoxiang Liu, Shiguo Lian
ICIP2
2022 RTSEC: Automated RTL Code Augmentation for Hardware Security Enhancement
abstract
Current hardware designs have increased in complexity, resulting in a reduced ability to perform security checks on them. Further, the addition of any security features to these designs is still largely manual which further complicates the design and integration process. In this paper, we address these shortcomings by introducing Rtsec as a framework which is capable of performing security analysis on designs as well as integrating security features directly into the HDL code, a feature that commercial EDA tools do not provide. Rtsec first breaks down HDL code into an Abstract Syntax Tree which is then used to infer the logic of the design. We demonstrate how Rtsec can be utilized to automatically include security mechanisms in RTL designs: watermarking and logic locking. We also compare the efficacy of our analysis algorithms with state of the art tools, demonstrating that Rtsec has capabilities equal or superior to those of state of the art tools while also providing the means of enhancing security features to the design.
Orlando Arias, Zhaoxiang Liu, Xiaolong Guo 0001, Yier Jin, Shuo Wang 0003
DATE2
2022 Inter-IP Malicious Modification Detection through Static Information Flow Tracking
abstract
To help expand the usage of formal methods in the hardware security domain. We propose a static register-transfer level (RTL) security analysis framework and an electronic design automation (EDA) tool named If-Tracker to support the proposed framework. Through this framework, a data-flow model will be automatically extracted from the RTL description of the SoC. Information flow security properties will then be generated. The tool checks all possible inter-IP paths to verify whether any property violations exist. The effectiveness of the proposed framework is demonstrated on customized SoC designs using AMBA bus where malicious modifications are inserted across multiple IPs. Existing IP level security analysis tools cannot detect such Trojans. Compared to commercial formal tools such as Cadence JasperGold and Synopsys VC-Formal, our framework provides a much simpler user interface and can identify more types of malicious modifications.
Zhaoxiang Liu, Orlando Arias, Weimin Fu, Yier Jin, Xiaolong Guo 0001
DATE1
2019 Deep Global-Relative Networks for End-to-End 6-DoF Visual Localization and Odometry
Yimin Lin, Zhaoxiang Liu, Chaopeng Wang, Guoguang Du 0001, Jinqiang Bai, Shiguo Lian
PRICAI (2)2