Zhijie Xu

dblp:58/4078 · DBLP profile ↗
← Back
58ranked-venue papers
9as first author
40since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 25 · 5 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 21 · 1 first-author · 12 since 2021Human-computer interaction and ubiquitous computing · 7 · 1 first-author · 1 since 2021Systems, architecture and hardware · 5 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 4 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
YearPublicationVenuePosition
2026 SEVADE: Self-Evolving Multi-Agent Analysis with Decoupled Evaluation for Hallucination-Resistant Sarcasm Detection
abstract
Sarcasm detection is a crucial yet challenging Natural Language Processing task. Existing Large Language Model methods are often limited by single-perspective analysis, static reasoning pathways, and a susceptibility to hallucination when processing complex ironic rhetoric, which impacts their accuracy and reliability. To address these challenges, we propose SEVADE, a novel Self-Evolving multi-agent Analysis framework with Decoupled Evaluation for hallucination-resistant sarcasm detection. The core of our framework is a Dynamic Agentive Reasoning Engine (DARE), which utilizes a team of specialized agents grounded in linguistic theory to perform a multifaceted deconstruction of the text and generate a structured reasoning chain. Subsequently, a separate lightweight rationale adjudicator (RA) performs the final classification based solely on this reasoning chain. This decoupled architecture is designed to mitigate the risk of hallucination by separating complex reasoning from the final judgment. Extensive experiments on four benchmark datasets demonstrate that our framework achieves state-of-the-art performance, with average improvements of 7.01% in Accuracy and 6.55% in Macro-F1 score.
Mingxuan Hu, Yushan Pan, Zhijie Xu, Yangbin Chen
AAAI6
2026 Placing Any Object at Any 3D Position
Ming Kong 0001, Zhanbin Hu, Zhijie Xu
AAAI5
2026 Algorithm-hardware co-design of binary neural network for efficient super resolution on FPGA
Yuanxin Su, Yushan Pan, Zhijie Xu, Xinfei Guo
Integr.5
2026 Unified compositional controller: A training-free framework for highly controllable text-to-image generation
Ruiyue Liu, Zhijie Xu, Jianqin Zhang
Inf. Sci.2
2026 Spatio-temporal feature discrimination for self-supervised skeleton action representation learning
Hongwei Chen 0002, Zhijie Xu, Xinlu Zong, Changyong Lin
Multim. Syst.2
2026 SFDiff: shadow removal via semantic-guided prototypes and frequency-aware modulation
Xinqing Huang, Zhijie Xu, Jianqin Zhang
Mach. Vis. Appl.2
2026 A text-guided cross-hierarchical fusion and multi-task learning framework for multimodal sentiment analysis
Yushan Pan, Zuhe Li, Di Wu 0035, Zhiyang Zhao, Yuanping Xu, Zhijie Xu
Neural Networks8
2026 Learning Symmetric Invariants from Symmetric Samples
abstract
Invariant synthesis is a fundamental problem in program verification, yet existing learning-based approaches rarely exploit the inherent symmetry present in many programs, particularly parameterized and concurrent systems. Such symmetry induces a symmetric reachable state space, naturally yielding symmetric samples and admitting symmetric invariants, motivating the task of learning symmetric invariants from symmetric samples. To this end, we introduce symmetric decision trees (SDTs), a novel hypothesis class that enforces symmetry structurally, guaranteeing symmetric invariants by construction. Furthermore, we develop a learning algorithm to construct SDTs and integrate it as the learner within the Horn-ICE framework, yielding our approach, Horn-SDT. Empirical evaluation on parameterized programs demonstrates that Horn-SDT achieves faster convergence and constructs more compact trees compared to non-symmetric baselines.
Zhijie Xu, Fei He 0001
Proc. ACM Program. Lang.1
2026 Unsupervised Large-Scale Point Cloud Registration via Spherical Projection Consistency
abstract
In real-world applications, large-scale outdoor Li- DAR point cloud registration faces significant challenges, including data sparsity, occlusions, and reliance on expensive ground-truth pose labels. To address these issues, this paper proposes an unsupervised registration framework based on spherical projection consistency. Specifically, both the input point cloud and its spatially transformed counterpart are projected into range images, and their spatial consistency is exploited as a supervision signal. Furthermore, a multi-scale patch-topatch feature fusion module is introduced to effectively integrate features from point clouds and range images, thereby enhancing feature discriminability. In addition, a dynamic masking strategy is applied in the range image domain to mitigate the impact of sparsity variations and further strengthen the spatial consistency of the supervision signal. Extensive experiments on the KITTI and NuScenes datasets demonstrate that the proposed method outperforms state-of-the-art unsupervised approaches, validating its effectiveness and robustness in large-scale outdoor scenarios.
Yejun Shou, Shuai Liu 0009, Zhijie Xu, Yanlong Cao
IEEE Signal Process. Lett.4
2026 AL-HCL: Active Learning and Hierarchical Contrastive Learning for Multimodal Sentiment Analysis With Fusion Guidance
Xiaojiang He, Yushan Pan, Zhijie Xu, Zuhe Li, Xinfei Guo, Chenguang Yang 0001
IEEE Trans. Affect. Comput.3
2026 Fast Photometric Stereo by Time- and Spectral-Multiplexing With Crosstalk Handling
abstract
Photometric stereo is widely used to recover detailed surface normals. However, previous methods fail to balance the accuracy and efficiency. Conventional photometric stereo achieves high accuracy but suffers from low efficiency due to spectral-multiplexing and inefficient algorithms. In contrast, multispectral photometric stereo captures images efficiently with spectral-multiplexing, but its accuracy is harmed by crosstalk. In this paper, we aim to resolve the crosstalk issue to achieve fast photometric stereo (FPS) at low cost. First, we analyze the formulation and impact of crosstalk, showing that it significantly affects normal estimation, with external factors being primary contributors to crosstalk and internal factors being the secondary. Subsequently, we propose the FPS framework with a fast data capture scheme that combines time- and spectral-multiplexing to introduce constraints on crosstalk regarding both internal and external factors, along with a lightweight network, FPS-Net, to remove crosstalk caused by those factors based on constraints under such scheme. Finally, we build a real-world crosstalk-affected FPS dataset to evaluate the performance in handling crosstalk for normal estimation. Experimental results show the superior accuracy and efficiency of our method. The code and dataset are available at https://github.com/wxy-zju/FPS-Net.
Xiaoyao Wei, Lingfeng Shen, Zhijie Xu, Yanlong Cao
IEEE Trans. Image Process.4
2026 Clue and Context Fusion for Sarcasm Detection with Large Multimodal Models
abstract
Detecting sarcasm in social media is fundamentally different from general VLM benchmarks: it is a pragmatic contradiction problem in which the literal signal in one modality is intentionally misaligned with the intended meaning, while dominant pre-training (e.g., CLIP-style contrastive agreement) biases models toward modality alignment rather than incongruity detection. We present SCARF, a contradiction-aware framework that equips large multimodal models with explicit sarcasm cues and context-sensitive retrieval. SCARF constructs coarse scene cues and fine localized evidence via tag-constrained QA, then distills them with visual tokens into a [FUSION] control vector for the LLM; a label-contrastive retriever supplies type- and context-matched exemplars, and a local multi-view encoder surfaces micro-cues. With the same backbone and training data, SCARF attains 87.92% Acc/86.67% F1 on MMSD2.0 and 77.14% Acc/76.44% F1 zero-shot on XDMSD, outperforming a comparably fine-tuned LLaVA-1.5. Ablations show sarcasm clue fusion is the main driver of gains, and tag-constrained QA improves rationale grounding and reduces hallucinations.
Yushan Pan, Ding Wang 0006, Wei Wang 0042, Xiaowei Huang 0001, Zhijie Xu
ACM Trans. Intell. Syst. Technol.6
2026 SFD-DETR: an end-to-end spatial-frequency dual-guided multi-scale network for tiny object detection in UAV aerial images
Weiquan Li, Zhijie Xu, Jianqin Zhang
J. Supercomput.2
2026 Modality balancing network for pedestrian detection based on cross-modal compensation fusion and multimodal feature alignment
Zuhe Li, Ruochong Fu, Zhiyang Zhao, Penghao Ouyang, Zhijie Xu, Yushan Pan
Vis. Comput.8
2025 NAT-3D: A Non-Rigid Approach for Accurate 3D Tracking of Basic Surgical Instruments
abstract
3D tracking technology serves as an essential enabler in numerous domains, most notably within digital healthcare, where it supports critical procedures such as surgical navigation, trajectory planning, and high-fidelity surgical simulation, etc. However, due to the slender, textureless, and reflective characteristics of basic surgical instruments, coupled with the complexity of their diverse motion patterns, existing systems and methods face significant limitations in achieving accurate 3D tracking of these instruments. We propose NAT-3D, an effective and efficient non-rigid 3D tracking approach based on multi-modal region estimation, integrating kinematic structural constraints and nonrigid constraints. This method enables accurate tracking and dynamic mapping of a variety of basic surgical instruments, including scalpels, scissors, clamps, and forceps, without the need for additional markers. The tracking covers various movement modes, including rigid motion, local mechanical motion, and non-rigid composite motion. Extensive experiments demonstrate that our method outperforms previous algorithms in terms of robustness, accuracy, applicability, and real-time performance. In addition, we introduce a novel dataset for Basic Surgical Instrument Tracking (BSIT), which can serve as a benchmark for future related research.
Zihan Deng, Jier Zhang, Yushan Pan, JunJun Pan, Zhijie Xu
BIBM6
2025 CAF-I: A Collaborative Multi-agent Framework for Enhanced Irony Detection with Large Language Models
Mingxuan Hu, Yangbin Chen, Zhijie Xu
ICONIP (1)5
2025 Causal-GATNet: Explainable Evidence-Aware Fake News Detection by Causal Inference
abstract
The proliferation of AI-generated disinformation has amplified the demand for evidence-based fake news detection. Although existing techniques leverage relevant evidence to verify the authenticity of news, they often rely on spurious correlations between superficial patterns and labels instead of genuine evidence-claim reasoning. This tendency degrades the model generalization when applied to real-world scenarios with distribution shifts. To address this challenge, we propose Causal-GATNet1, a unified causal inference architecture that integrates causal intervention with adaptive evidence gating. Our approach: 1) Employs a dual-path causal debiasing framework that compares the standard prediction using all available evidence with a counterfactual prediction generated by blocking bias-inducing evidence, thereby yielding an unbiased output through prediction differencing. 2) Incorporates the Gated Affine Transformation (GAT) mechanism to enhance evidence reliability by combining the source credibility of evidence with their semantics. Experiments on multiple datasets demonstrate that integrating Causal-GATNet with established baseline models such as BERT and MAC significantly improves prediction accuracy and bias mitigation compared to the original implementations.
Jianwei Cai, Fuyu Xing, Haiyang Zhang 0004, Yangbin Chen, Zhijie Xu
INDIN6
2025 Investigation into the Coexistence of AI and Digital Media Workforce
abstract
This study explores the influence of artificial intelligence on digital media by analyzing user-generated content from Reddit, Twitter, and Tumblr. Natural language processing (NLP) and machine learning techniques are applied to classify sentiment and examine its relationship with user engagement and temporal trends. BERTweet, a transformer-based model, is employed for sentiment analysis, while machine learning models are used for further classification. The findings reveal a positive correlation between sentiment intensity and user engagement, while overall sentiment trends remain stable. This study contributes to the understanding of AI-driven discourse in digital media, demonstrating the effectiveness of advanced NLP techniques in social media analysis.
Ruixuan Chen, Haiyang Zhang 0004, Zhijie Xu, Zihan Deng, Yunchu Peng
INDIN3
2025 Enhanced multi-scale feature adaptive fusion sparse convolutional network for large-scale scenes semantic segmentation
Lingfeng Shen, Yanlong Cao, Yejun Shou, Zhijie Xu
Comput. Graph.7
2025 Towards understanding human actions through long-short-term semantic motion encoding
Chaolong Zhang 0002, Yuanping Xu, Zhijie Xu, Chao Kong, Benjun Guo, Jin Jin 0004
Eng. Appl. Artif. Intell.3
2025 A self-supervised facial expression restoration and recognition method based on improved masked auto-encoder
Chaolong Zhang 0002, Yuanping Xu, Zhijie Xu, Rongqiang Gou, Jin Jin 0004, Jian Huang 0017
Neurocomputing3
2025 Empowering couriers: balancing algorithmic control and autonomy for enhanced well-being
abstract
Abstract This paper explores the complex socio-technical dynamics that shape the work experience of food delivery workers in China through a qualitative study of 28 couriers. Our findings reveal a serious disconnect between algorithmic systems and workers’ lived experiences. We identify five key tensions in this algorithmic labour environment: (1) complementarity and conflict between algorithmic logic and workers’ experiential knowledge; (2) business models that shape technological systems without adequately accounting for task complexity; (3) rating systems that simultaneously quantify performance but creates ”service inflation” and psychological stress; (4) privacy concerns create a tension between service efficiency and information protection; and (5) complaint handling mechanisms systematically shift risks to workers through disproportionate penalties and defense costs. These findings contribute to the human–computer interaction literature by showing how algorithmic management embodies power structures rather than being a neutral tool, how workers’ adaptability constitutes valuable but invisible labor, how the trade-off between privacy and efficiency manifests differently in the Chinese cultural context, and how procedural justice issues extend beyond technical design. Finally, we suggest the implications of designing more equitable sociotechnical systems, namely better integrating workers’ tacit knowledge, matching business models to technical capabilities, developing culturally appropriate privacy solutions, and establishing more balanced evaluation and grievance resolution processes.
Haosu Xu, Zhijie Xu, Yushan Pan
Interact. Comput.4
2025 Graph-Aware Hybrid Encoding for Hyperspectral Image Classification
abstract
Hyperspectral image (HSI) classification faces critical challenges in effectively modeling the intricate spectral-spatial structures and non-Euclidean relationships. Traditional methods often struggle to simultaneously capture local details, global contextual dependencies, and graph-structured correlations, leading to limited classification accuracy. To address the above issues, this letter proposes a Graph-Aware Hybrid Encoding (GAHE) framework. To fully exploit the spectral-spatial characteristics and graph structural dependencies inherent in HSI, the proposed method is structured into three key components: a multi-scale selective graph-aware attention module, a hybrid projection encoding module, and a graph sensitive aggregation module. The three modules work in a complementary manner to progressively refine and enhance feature representations across multiple scales and modalities. Comparing with advanced classification methods,the experimental results demonstrate that the proposed GAHE method shows better classification performance.
Yuquan Gan, Zhijie Xu, Yushan Pan
IEEE Geosci. Remote. Sens. Lett.4
2025 A Unified Framework for Event-Based Frame Interpolation With Ad-Hoc Deblurring in the Wild
abstract
Effective video frame interpolation hinges on the adept handling of motion in the input scene. Prior work acknowledges asynchronous event information for this, but often overlooks whether motion induces blur in the video, limiting its scope to sharp frame interpolation. We instead propose a unified framework for event-based frame interpolation that performs deblurring ad-hoc and thus works both on sharp and blurry input videos. Our model consists in a bidirectional recurrent network that incorporates the temporal dimension of interpolation and fuses information from the input frames and the events adaptively based on their temporal proximity. To enhance the generalization from synthetic data to real event cameras, we integrate self-supervised framework with the proposed model to enhance the generalization on real-world datasets in the wild. At the dataset level, we introduce a novel real-world high-resolution dataset with events and color videos named HighREV, which provides a challenging evaluation setting for the examined task. Extensive experiments show that our network consistently outperforms previous state-of-the-art methods on frame interpolation, single image deblurring, and the joint task of both. Experiments on domain transfer reveal that self-supervised training effectively mitigates the performance degradation observed when transitioning from synthetic data to real-world data. Code and datasets are available at https://github.com/AHupuJR/REFID.
Lei Sun 0009, Daniel Gehrig, Christos Sakaridis, Mathias Gehrig, Jingyun Liang, Zhijie Xu, Kaiwei Wang, Luc Van Gool, Davide Scaramuzza 0001
IEEE Trans. Pattern Anal. Mach. Intell.7
2025 Verifying chip designs at RTL level
Nan Zhang 0001, Zhijie Xu, Cong Tian 0001, Chaofeng Yu
Sci. Comput. Program.2
2025 Scale-Selectable Global Information and Discrepancy Learning Network for Multimodal Sentiment Analysis
abstract
Multimodal sentiment analysis and depression detection are pivotal for advancing human-computer interaction, yet significant challenges remain. First, the limited extraction of global contextual information within individual modalities risks the loss of modal-specific features. Second, existing methods often prioritize unaligned textual interactions, neglecting critical inter-modal discrepancies. To address these issues, we propose the Scale-Selectable Global and Discrepancy Learning Network (SSGDL), an innovative framework that integrates two core modules: the Cross-Shaped Dynamic Scale Attention Module (CSDSA) and the Primary-Secondary modal Discrepancy Learning Module (PS-MDL). The CS-DSA dynamically selects scales and employs cross-shaped attention to capture comprehensive global context and intricate internal correlations, effectively producing a fused modal representation. Meanwhile, the PS-MDL designates the fused modal as primary and utilizes cross-attention mechanisms to learn discrepancy representations between it and other modalities (textual, acoustic, and visual). By leveraging intermodal discrepancies, SSGDL achieves a more nuanced and holistic understanding of emotional content. Extensive experiments on three benchmark multimodal sentiment analysis datasets (MOSI, MOSEI, SIMS) and a depression detection dataset (AVEC2019) demonstrate that SSGDL consistently outperforms state-of-theart approaches, setting a new benchmark for multimodal affective computing.
Xiaojiang He, Yushan Pan, Xinfei Guo, Zhijie Xu, Chenguang Yang 0001
IEEE Trans. Affect. Comput.4
2025 Learnable semantic completion and distance-induced graph matching for domain adaptive object detection
Zhijie Xu, Yuanyuan Qiu, Jianqin Zhang
J. Supercomput.1
2024 iS-MAP: Neural Implicit Mapping and Positioning for Structural Environments
Yanlong Cao, Yejun Shou, Lingfeng Shen, Xiaoyao Wei, Zhijie Xu
ACCV (9)6
2024 Multi-Object Tracking for Unmanned Aerial Vehicles Based on Multi-Frame Feature Fusion
abstract
To address the issues of tracking trajectory loss caused by small object size, frequent view angle changes and object occlusion in the multi-object tracking task of Unmanned Aerial Vehicle (UAV), in this paper, we propose a multi-object tracker for UAV based on multi-frame feature fusion. First, in order to more fully extract and utilize the interframe information, we design an attention-based adaptive multi-frame fusion module, which introduces Efficient Channel Attention (ECA) to trade-off the importance of the information in the history frames and the current frame. Second, we use a high-resolution feature extraction network as backbone network to extract features. The proposed method is evaluated on the UAV multi-object tracking datasets of Visdrone2019 and UAVDT. Compared with other mainstream multi-object tracking algorithms, our method achieves higher accuracy and fewer identity switches, which effectively improves multi-object tracking performance.
Jiayin Wen, Dianwei Wang, Zhijie Xu
ICASSP5
2024 Structerf-SLAM: Neural implicit representation SLAM for structural environments
Yanlong Cao, Xiaoyao Wei, Yejun Shou, Lingfeng Shen, Zhijie Xu
Comput. Graph.6
2024 Image recognition based on lightweight convolutional neural network: Recent advances
abstract
Image recognition is an important task in computer vision with broad applications. In recent years, with the advent of deep learning, lightweight convolutional neural network (CNN) has brought new opportunities for image recognition, which allows high-performance recognition algorithms to run on resource-constrained devices with strong representation and generalization capabilities. This paper first presents an overview of several classical lightweight CNN models. Then, a comprehensive review is provided on recent image recognition techniques using lightweight CNN. According to the strategies applied to optimize image recognition performance, existing methods are classified into three categories: (1) model compression, (2) optimization of lightweight network, and (3) combining Transformer with lightweight network. In addition, some representative methods are tested on three commonly used datasets for performance comparison. Finally, technical challenges and future research trends in this field are discussed.
Ying Liu 0026, Jiahao Xue, Daxiang Li 0002, Weidong Zhang 0005, Tuan Kiang Chiew, Zhijie Xu
Image Vis. Comput.6
2024 Facial emotion recognition with a reduced feature set for video game and metaverse avatars
abstract
Abstract This paper presents a novel real‐time facial feature extraction algorithm, producing a small feature set, suitable for implementing emotion recognition with online game and metaverse avatars. The algorithm aims to reduce data transmission and storage requirements, hurdles in the adoption of emotion recognition in these mediums. The early results presented show a facial emotion recognition accuracy of up to 92% on one benchmark dataset, with an overall accuracy of 77.2% across a wide range of datasets, demonstrating the early promise of the research.
Darren Bellenger, Minsi Chen, Zhijie Xu
Comput. Animat. Virtual Worlds3
2024 psoResNet: An improved PSO-based residual network search algorithm
abstract
Neural Architecture Search (NAS) methods are widely employed to address the time-consuming and costly challenges associated with manual operation and design of deep convolutional neural networks (DCNNs). Nonetheless, prevailing methods still encounter several pressing obstacles, including limited network architecture design, excessively lengthy search periods, and insufficient utilization of the search space. In light of these concerns, this study proposes an optimization strategy for residual networks that leverages an enhanced Particle swarm optimization algorithm. Primarily, low-complexity residual architecture block is employed as the foundational unit for architecture exploration, facilitating a more diverse investigation into network architectures while minimizing parameters. Additionally, we employ a depth initialization strategy to confine the search space within a reasonable range, thereby mitigating unnecessary particle exploration. Lastly, we present a novel approach for computing particle differences and updating velocity mechanisms to enhance the exploration of updated trajectories. This method significantly contributes to the improved utilization of the search space and the augmentation of particle diversity. Moreover, we constructed a crime-dataset comprising 13 classes to assess the effectiveness of the proposed algorithm. Experimental results demonstrate that our algorithm can design lightweight networks with superior classification performance on both benchmark datasets and the crime-dataset.
Dianwei Wang, Leilei Zhai, Yuanqing Li 0003, Zhijie Xu
Neural Networks5
2023 An End-to-End Multi-stage Network for Ultrasound Video Object Segmentation
abstract
Real-time tracking and segmentation of ultrasound video sequence are prerequisite for identifying and analyzing lesions. While significant progress has been made in natural video object segmentation, developing a model for ultrasound video is still challenging due to problems such as low distinguishability and low visual saliency of the target objects, large variation between adjacent frames. These challenges are inherently complex and cannot be effectively tackled through a single process. This paper develops an end-to-end multi-stage network (EMNet) for ultrasound video object segmentation. EMNet consists of two stages. The inital mask generation stage comprises a contrast-enhanced layer to enhance visual contrast between targets and backgrounds. In this stage, a module that adopts the encoder-attention-decoder structure is designed for mask induction. After obtaining the initial segmentation mask, the mask refinement stage is followed to further improve initial segmentation. To prevent the propagation of errors, a gating mechanism is designed to control the fusion of segmentation probability maps in the initial and refinement stages. By transforming certain fixed parameters in different stage into trainable parameters and establishing an end-to-end learning process, we optimized the performance of our approach. We evaluate EMNet on real-world lymphoma ultrasound video dataset. Compared with the best results among seven competing baselines, EMNet achieves the best performance in terms of ℐ&ℱ and ℱ measures, the second-best performance with Param and FPS measures, which demonstrates the competitive performance in terms of both speed and accuracy.
Yijie Dong, Zhijie Xu, Qiao Pan, Dehua Chen, Jianwen Su
BIBM4
2023 An Improved Method for CFNet Identifying Glioma Cells
Lin Yuan 0001, Jinling Lai, Zhen Shen 0003, Wendong Yu, Hongwei Wei, Zhijie Xu
ICIC (3)7
2023 Verifying Chips Design at RTL Level
Nan Zhang 0001, Cong Tian 0001, Zhijie Xu, Chaofeng Yu
TASE5
2022 Latent Space Simulation for Carbon Capture Design Optimization
abstract
The CO2 capture efficiency in solvent-based carbon capture systems (CCSs) critically depends on the gas-solvent interfacial area (IA), making maximization of IA a foundational challenge in CCS design. While the IA associated with a particular CCS design can be estimated via a computational fluid dynamics (CFD) simulation, using CFD to derive the IAs associated with numerous CCS designs is prohibitively costly. Fortunately, previous works such as Deep Fluids (DF) (Kim et al., 2019) show that large simulation speedups are achievable by replacing CFD simulators with neural network (NN) surrogates that faithfully mimic the CFD simulation process. This raises the possibility of a fast, accurate replacement for a CFD simulator and therefore efficient approximation of the IAs required by CCS design optimization. Thus, here, we build on the DF approach to develop surrogates that can successfully be applied to our complex carbon-capture CFD simulations. Our optimized DF-style surrogates produce large speedups (4000x) while obtaining IA relative errors as low as 4% on unseen CCS configurations that lie within the range of training configurations. This hints at the promise of NN surrogates for our CCS design optimization problem. Nonetheless, DF has inherent limitations with respect to CCS design (e.g., limited transferability of trained models to new CCS packings). We conclude with ideas to address these challenges.
Brian R. Bartoldson, Yucheng Fu, David P. Widemann, Sam Nguyen, Zhijie Xu, Brenda Ng
AAAI7
2022 Para-CFlows: $C^k$-universal diffeomorphism approximators as superior neural surrogates
abstract
Invertible neural networks based on Coupling Flows (CFlows) have various applications such as image synthesis and data compression. The approximation universality for CFlows is of paramount importance to ensure the model expressiveness. In this paper, we prove that CFlows}can approximate any diffeomorphism in $C^k$-norm if its layers can approximate certain single-coordinate transforms. Specifically, we derive that a composition of affine coupling layers and invertible linear transforms achieves this universality. Furthermore, in parametric cases where the diffeomorphism depends on some extra parameters, we prove the corresponding approximation theorems for parametric coupling flows named Para-CFlows. In practice, we apply Para-CFlows as a neural surrogate model in contextual Bayesian optimization tasks, to demonstrate its superiority over other neural surrogate models in terms of optimization performance and gradient approximations.
Junlong Lyu, Zhitang Chen, Chang Feng, Wenjing Cun, Shengyu Zhu 0001, Yanhui Geng, Zhijie Xu, Chen Yongwei
NeurIPS7
2022 Hybrid handcrafted and learned feature framework for human action recognition
Chaolong Zhang 0002, Yuanping Xu, Zhijie Xu, Jian Huang 0017
Appl. Intell.3
2021 Hierarchical visual localization for visually impaired people using multimodal images
Ruiqi Cheng, Weijian Hu, Yicheng Fang, Kaiwei Wang, Zhijie Xu
Expert Syst. Appl.6
2020 An Augmented Treble Stream Deep Neural Network for Video Analysis
abstract
Video analysis for human action recognition is one of the most important research areas in pattern recognition and computer vision due to its wide applications. Deep learning-based approaches have been proven more effective than conventional feature engineering-based models. However, the performance is still unreliable when facing real-world application scenarios. Inspired by the Convolutional Neural Network (CNN) and Recurrent Long-Short Term Model (LSTM), this paper presents an augmented treble-stream deep neural network architecture that supports direct extraction of spatial-temporal features from video streams and their corresponding dense optical flows. This innovative approach assists effective detection of complex video event features that are annotated by rich event "appearance" and motion features. Substantially improved recognition accuracy is recorded during the experiments that are carried and benchmarked over public video event datasets, for example, UCF 101 and HMDB 51. Analytical evaluation approves the validity and effectiveness of the treble-stream neural network design.
Chaolong Zhang 0002, Yuanping Xu, Zhijie Xu, Mei Gong, Benjun Guo, Dengguo Yao
IV3
2019 Dual-channel CNN for efficient abnormal behavior identification through crowd feature engineering
Yuanping Xu, Zhijie Xu, Jia He 0003, Jiliu Zhou, Chaolong Zhang 0002
Mach. Vis. Appl.3
2019 A generic parallel computational framework of lifting wavelet transform for online engineering surface filtration
Yuanping Xu, Chaolong Zhang 0002, Zhijie Xu, Jiliu Zhou, Kaiwei Wang, Jian Huang 0017
Signal Process.3
2018 A Graphical Simulator for Modeling Complex Crowd Behaviors
abstract
Abnormal crowd behaviors of varied real-world settings could represent or pose serious threat to public safety. The video data required for relevant analysis are often difficult to acquire due to security, privacy and data protection issues. Without large amounts of realistic crowd data, it is difficult to develop and verify crowd behavioral models, event detection techniques, and corresponding test and evaluations. This paper presented a synthetic method for generating crowd movements and tendency based on existing social and behavioral studies. Graph and tree searching algorithms as well as game engine-enabled techniques have been adopted in the study. The main outcomes of this research include a categorization model for entity-based behaviors following a linear aggregation approach; and the construction of an innovative agent-based pipeline for the synthesis of A-Star path-finding algorithm and an enhanced Social Force Model. A Spatial-Temporal Texture (STT) technique has been adopted for the evaluation of the model's effectiveness. Tests have highlighted the visual similarities between STTs extracted from the simulations and their counterparts - video recordings - from the real-world.
Yu Hao 0002, Zhijie Xu, Ying Liu 0026, Jing Wang 0033, JiuLun Fan 0001
IV2
2016 The Visualisation of Cognitive Structures in Forensic Statements
abstract
Forensic statements are often unstructured, intricate, and thus difficult to interpret and assess. This is due to the varied nature and format of how interviews with victims, witnesses, or suspects are conducted. It is even more difficult for police investigators, lawyers or other legal practitioners to grasp intuitively and accurately the key information pertaining within the varied statements. This research investigates the opportunities in the convergence of linguistic approaches to extracting and reconstructing the cognitive structure, i.e. "Text-Worlds", in a statement, and the computerised operational settings for enabling effective and hopefully more accurate interpretation of forensic discourse through visualisation.
Jing Wang 0033, Yufang Ho, Zhijie Xu, Dan McIntyre, Jane Lugea
IV3
2016 Spatio-temporal texture modelling for real-time crowd anomaly detection
Jing Wang 0033, Zhijie Xu
Comput. Vis. Image Underst.2
2014 Learning with positive and unlabeled examples using biased twin support vector machine
Zhijie Xu, Zhiquan Qi, Jianqin Zhang
Neural Comput. Appl.1
2013 STV-based video feature processing for action recognition
Jing Wang 0033, Zhijie Xu
Signal Process.2
2012 Multi-source Transfer Learning with Multi-view Adaboost
Zhijie Xu, Shiliang Sun
ICONIP (3)1
2011 Multi-view Transfer Learning with Adaboost
abstract
Transfer learning, serving as one of the most important research directions in machine learning, has been studied in various fields in recent years. In this paper, we integrate the theory of multi-view learning into transfer learning and propose a new algorithm named Multi-View Transfer Learning with Adaboost (MV-TL Adaboost). Different from many previous works on transfer learning, we not only focus on using the labeled data from one task to help to learn another task, but also consider how to transfer them in different views synchronously. We regard both the source and target task as a collection of several constituent views and each of these two tasks can be learned from every views at the same time. Moreover, this kind of multi-view transfer learning is implemented with adaboost algorithm. Furthermore, we analyze the effectiveness and feasibility of MV-TL Adaboost. Experimental results also validate the effectiveness of our proposed approach.
Zhijie Xu, Shiliang Sun
ICTAI1
2011 Part-Based Transfer Learning
Zhijie Xu, Shiliang Sun
ISNN (3)1
2011 Developing a knowledge-based system for complex geometrical product specification (GPS) data manipulation
Yuanping Xu, Zhijie Xu, Xiangqian Jiang, Paul J. Scott
Knowl. Based Syst.2
2010 An Algorithm on Multi-View Adaboost
Zhijie Xu, Shiliang Sun
ICONIP (1)1
2010 Parallel implementation of wavelet-based image denoising on programmable PC-grade graphics hardware
Zhijie Xu
Signal Process.2
2008 GPGPU-based Gaussian Filtering for Surface Metrological Data Processing
abstract
Engineering surfaces are characterized by the form, waviness and roughness features that are comprised of a range of spatial wavelengths. Filtering techniques are commonly adopted to separate these different wavelength components into well-defined bandwidths for further processing. The Gaussian filtered surface in which a 2D Gaussian filter is employed for surface assessments has been recommended by the ISO 11562-1996 and ASME B46-1995 standards to establish a reference surface. For Gaussian filtering, computational efficiency is a key problem when it is issued on a large set of surface metrology data. In the past this problem was tackled through reducing computation amount by the design and adoption of some fast algorithms. In this paper, a general purpose computing on GPU (GPGPU) framework is discussed to accelerate 2D Gaussian filtering for surface characterization. This framework takes advantage of the GPUpsilas parallel computing ability and has achieved better data efficiency without reducing the computational amount while maintaining the filtering quality. Filtering results and their accuracy from this model have been compared with the results obtained from the MATLAB simulation kits and the satisfied outcomes were observed.
Zhijie Xu, Xiangqian Jiang
IV2
2005 A modified clustering algorithm for data mining
abstract
Clustering is a widely used technique of finding interesting patterns residing in the dataset that were not obviously known. It is a division of data into groups of similar objects. The clustering of large data sets has received a lot of attention in recent years, however, clustering is still a challenging task since many cluster algorithms fail to do well in scaling with the size of the data set and the number of dimensions that describe the points, or in finding arbitrary shapes of clusters, or dealing effectively with the presence of noise. This paper describes a clustering method for unsupervised classification of objects in large data sets. The new methodology combines the simulating annealing algorithm with CLARANS (clustering large application based upon randomized search) in order to cluster large data sets efficiently. At last, the method is experimented on the generated data set. The result shows that the approach is quick than CLARANS and can produce a similar division of data as CLARANS.
Zhijie Xu, Laisheng Wang, Jiancheng Luo, Jianqin Zhang
IGARSS1
2003 Virtual reality for Neuropsychological diagnosis and rehabilitation: A Survey
abstract
Recently the considerable potential of virtual reality has been recognized for the scientific study, diagnosis and treatment in the field of mental healthcare. It provides a unique opportunity to provide a natural interface and to mimic actual challenges faced by impaired subjects in their lives. In spite of certain limitations, some work has emerged which provides guidelines for future research efforts. A shortlist of such applications includes assessment and treatment of phobias; obsessive compulsive disorders, posttraumatic stress disorders and hyperactive attention deficit syndromes. We highlight these recent and ongoing research projects and compare this unique paradigm with conventional nonVR techniques by discussion of potential benefits and possible shortcomings. Due to the nature of research, ethical and human factors cannot be over emphasized. We also describe the limitations of VR in the field of practical psychology as media hype has over sold its potential. With recent advances in technology, it is not unrealistic to hope that VR will soon be a part of a clinician's armoury.
Yasir Khan, Zhijie Xu, Mark Stigant
IV2
2003 Using Motion Platform as a Haptic Display for Virtual Inertia Simulation
abstract
Haptic displays serve great importance in creating the immersiveness sense that block virtual reality (VR) users from the contradictory impressions from the real world. We introduce the general structure of virtual reality systems. It reports common approaches in constructing different virtual scenarios and the way they are utilised. We integrate a motion platform with an application program interface (API) based virtual environment authoriser to provide force feedback for simulating inertia related applications such as virtual gliding, driving, and snowboarding.
Zhijie Xu
IV1