Xujie Zhang

dblp:126/4384 · DBLP profile ↗
← Back
22ranked-venue papers
4as first author
19since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 1 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 6 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 CSNF-TimeGAN: Class-prior spatiotemporal nonlinear feature-based time series generative adversarial network for blast furnace fault diagnosis
Yuelin Yang, Chunjie Yang 0001, Siwei Lou, Dali Gao, Xujie Zhang
Adv. Eng. Informatics5
2026 Dynamical Component Extraction-Based Fault Detection for Industrial IoT With Application to Ironmaking Process
abstract
The Industrial Internet of Things (IIoT) has become a crucial infrastructure in the process industry, particularly in the era of Industry 4.0. Ensuring operational safety in industrial processes necessitates fault detection techniques, which play a pivotal role in IIoT systems. These systems continuously collect high-dimensional process data, which often exhibit dynamic behavior due to the inherent complexity of industrial operations. Consequently, the dynamic characteristics of such data pose significant challenges for fault detection. As a powerful dimensionality reduction technique, Dynamical Component Analysis (DyCA) decomposes multivariate measurements of a dynamical system into a deterministic component which can be described by a system of differential equations and independent noise components. DyCA incorporates the covariance matrices of both the signals, and their derivative, as well as their cross-correlation. By doing so, it identifies a low-dimensional subspace that minimizes the error in the underlying ordinary differential equations. The DyCA components are estimated to capture low-dimensional trajectories that characterize the process dynamics. This study proposes a novel data-driven fault detection method based on dynamical component analysis for dynamic processes. Leveraging these DyCA components that represent the low-dimensional trajectories to describe the process dynamics, Hotelling’sT2and Square Prediction Error (SPE) statistics are utilized as monitoring metrics for fault detection. Case studies on the widely utilized Tennessee Eastman process benchmark and a real-world blast furnace ironmaking process are conducted to demonstrate the effectiveness and capability of the proposed DyCA based fault detection method, comparing its performance with other relevant methods.
Ping Wu 0001, Yicheng Yu, Xujie Zhang, Siwei Lou, Jinfeng Gao 0002, Qian Zhang 0002, Chunjie Yang 0001
IEEE Internet Things J.3
2026 DyCVDA: Dynamical and Canonical Variate Dissimilarity Analysis-Based Fault Detection for Blast Furnace Ironmaking Process
abstract
The blast furnace ironmaking process (BFIP) serves as the core unit within the iron and steel industry. However, its harsh operating environment often leads to process abnormalities, resulting in unplanned shutdowns and product quality degradation. Data driven fault detection is crucial for ensuring operational safety and maintaining product quality in this context. Nevertheless, BFIP data are inherently characterized by strong dynamic and highly correlated properties, which present significant challenges for effective fault detection. To address these challenges, this work proposes a novel data driven fault detection method based on Dynamic and Canonical Variate Dissimilarity Analysis (DyCVDA). Within the developed DyCVDA framework, high-dimensional process data are first decomposed into deterministic components and noise components by optimally satisfying a set of coupled ordinary differential equations, where deterministic components consist of time-dependent amplitudes and corresponding multivariate modes. Thereby, the process’s dynamic characteristic can be captured. Subsequently, low-dimensional subspaces of deterministic components are extracted to represent the underlying dynamic behavior. Canonical subspaces are further derived by maximizing the correlation between past and future observations of these low-dimensional projections to exploit the correlation structure within the process data. Finally, a canonical variate dissimilarity index is employed as the monitoring statistic for fault detection. Experimental results on two case studies, including the popular Tennessee Eastman Process industrial benchmark and a real world blast furnace ironmaking process, demonstrate the effectiveness of the proposed DyCVDA method, with favorable performance comparisons against other relevant approaches.
Ping Wu 0001, Yicheng Yu, Jinfeng Gao 0002, Xujie Zhang, Siwei Lou, Chunjie Yang 0001
IEEE Trans Autom. Sci. Eng.5
2025 DreamFit: Garment-Centric Human Generation via a Lightweight Anything-Dressing Encoder
abstract
Diffusion models for garment-centric human generation from text or image prompts have garnered emerging attention for their great application potential. However, existing methods often face a dilemma: lightweight approaches, such as adapters, are prone to generate inconsistent textures; while finetune-based methods involve high training costs and struggle to maintain the generalization capabilities of pretrained diffusion models, limiting their performance across diverse scenarios. To address these challenges, we propose DreamFit, which incorporates a lightweight Anything-Dressing Encoder specifically tailored for the garment-centric human generation. DreamFit has three key advantages: (1) Lightweight training: with the proposed adaptive attention and LoRA modules, DreamFit significantly minimizes the model complexity to 83.4M trainable parameters. (2) Anything-Dressing: Our model generalizes surprisingly well to a wide range of (non-)garments, creative styles, and prompt instructions, consistently delivering high-quality results across diverse scenarios. (3) Plug-and-play: DreamFit is engineered for smooth integration with any community control plugins for diffusion models, ensuring easy compatibility and minimizing adoption barriers. To further enhance generation quality, DreamFit leverages pretrained large multi-modal models (LMMs) to enrich the prompt with fine-grained garment descriptions, thereby reducing the prompt gap between training and inference. We conduct comprehensive experiments on both 768 x 512 high-resolution benchmarks and in-the-wild images. DreamFit surpasses all existing methods, highlighting its state-of-the-art capabilities of garment-centric human generation.
Ente Lin, Xujie Zhang, Fuwei Zhao, Yuxuan Luo 0002, Long Zeng 0001, Xiaodan Liang
AAAI2
2025 CatVTON: Concatenation Is All You Need for Virtual Try-On with Diffusion Models
abstract
Virtual try-on methods based on diffusion models achieve realistic effects but often require additional encoding modules, a large number of training parameters, and complex preprocessing, which increases the burden on training and inference. In this work, we re-evaluate the necessity of additional modules and analyze how to improve training efficiency and reduce redundant steps in the inference process. Based on these insights, we propose CatVTON, a simple and efficient virtual try-on diffusion model that transfers in-shop or worn garments of arbitrary categories to target individuals by concatenating them along spatial dimensions as inputs of the diffusion model. The efficiency of CatVTON is reflected in three aspects: (1) Lightweight network. CatVTON consists only of a VAE and a simplified denoising UNet, removing redundant image and text encoders as well as cross-attentions, and includes just 899.06M parameters. (2) Parameter-efficient training. Through experimental analysis, we identify self-attention modules as crucial for adapting pre-trained diffusion models to the virtual try-on task, enabling high-quality results with only 49.57M training parameters. (3) Simplified inference. CatVTON eliminates unnecessary preprocessing, such as pose estimation, human parsing, and captioning, requiring only a person image and garment reference to guide the virtual try-on process, reducing over 49% memory usage compared to other diffusion-based methods. Extensive experiments demonstrate that CatVTON achieves superior qualitative and quantitative results compared to baseline methods and demonstrates strong generalization performance in in-the-wild scenarios, despite being trained solely on public datasets with 73K samples.
Zheng Chong, Xujie Zhang, Dongmei Jiang, Xiaodan Liang
ICLR7
2025 FA-SconvAE-LSTM: Feature-Aligned Stacked Convolutional Autoencoder with Long Short-Term Memory Network for Soft Sensor Modeling
Ping Wu 0001, Zengdi Miao, Jinfeng Gao 0002, Xujie Zhang, Siwei Lou, Chunjie Yang 0001
Eng. Appl. Artif. Intell.5
2024 GarmentAligner: Text-to-Garment Generation via Retrieval-Augmented Multi-level Corrections
Zheng Chong, Xujie Zhang, Yuhao Cheng, Yiqiang Yan, Xiaodan Liang
ECCV (25)3
2024 Design and comprehensive analysis of improved Proportional-Integral-Retarded protocol for second-order multi-agent systems
Xujie Zhang, Qingbin Gao, Jiazhi Cai, Wenfu Xu
Inf. Sci.1
2024 Data-Driven Joint Fault Diagnosis Based on RMK-ASSA and DBSKNet for Blast Furnace Iron-Making Process
abstract
Blast furnace iron-making process (BFIP) is one of the most critical procedures in the iron and steel industry where timely detection and accurate classification of faults have always been of core focus. However, the coupling effects of system’s nonlinear and nonstationary characteristics often cause process consistent underlying information to be buried, allowing accurate extraction to be a significant challenge. This also complicates the development of BFIP fault diagnosis model. Therefore, we propose a novel data-driven joint fault diagnosis strategy that employs regularized mutual kernel analytic stationary subspace analysis (RMK-ASSA) and deep broad stationary kernel network (DBSKNet) to eliminate this interference. To develop this method, we first construct an RMK-ASSA approach to address the poor modeling accuracy caused by standard analytic stationary subspace analysis (ASSA)’s inability to handle complex process nonlinearity. Global and local kernels are utilized to account for multiple nonlinearities in BFIP data. The weight of different nonlinear data is calculated by regularized principal component analysis, and the main information is imported into ASSA to obtain more robust and accurate modeling results by eliminating the interference of redundant noise. Subsequently, we design a DBSKNet-based classifier to implement the fault diagnosis task. This network further considers the nonlinearity by boosting kernel structure in depth and width while distinguishing the respective contributions of different kernels to fault diagnosis results. Finally, a double-layer loop parameter optimization algorithm is used for optimizing. Simulated cases and practical BFIP tests validate that RMK-ASSA eliminates the negative impact caused by nonstationary data and that the proposed joint fault diagnosis strategy outperforms other methods.Note to Practitioners—BFIP’s nonlinear and nonstationary coupling properties pose unique challenges in eliminating distractions, constructing fault classifiers and accurately detecting process anomalies. To tackle these challenges, this paper proposes a joint fault diagnosis strategy based on RMK-ASSA and DBSKNet. RMK-ASSA effectively estimates nonlinear consistent features, while DBSKNet mines rich deep nonlinear information, accurately distinguishing variations in BFIP data under different working conditions. Experimental results demonstrate that this data-driven strategy can perform high-quality fault diagnosis, enabling field engineers to execute operations efficiently.
Siwei Lou, Chunjie Yang 0001, Ping Wu 0001, Yuelin Yang, Liyuan Kong, Xujie Zhang
IEEE Trans Autom. Sci. Eng.6
2024 Blast Furnace Ironmaking Process Monitoring With Time-Constrained Global and Local Nonlinear Analytic Stationary Subspace Analysis
abstract
In this article, a novel time-constrained global and local nonlinear analytic stationary subspace analysis (Tc-GLNASSA) is proposed to enhance blast furnace ironmaking process (BFIP) monitoring. Although the existing analytic stationary subspace analysis method has been available for deriving process consistent relationships. However, the presence of complex nonlinear, periodic nonstationary, and time-varying smelting conditions renders the satisfactory estimation of stationary projections unattainable. To this end, we leverage multiple kernel functions and manifold learning methods to establish a global and local nonlinear structure with time constraints, which will identify the unique nonlinearities excited by periodic nonstationarity. Meanwhile, a singular value decomposition-based modeling efficiency promotion strategy is constructed to reduce the proposed Tc-GLNASSA's computational complexity significantly. The orthogonality of model update scheme is analyzed theoretically, and an overall BFIP monitoring framework is given. Ultimately, practical BFIP case studies fully demonstrate the effectiveness of our proposal.
Siwei Lou, Chunjie Yang 0001, Xujie Zhang, Hanwen Zhang 0002, Ping Wu 0001
IEEE Trans. Ind. Informatics3
2024 From Complexity to Clarity: M2KCSVA's Nonlinear Temporal Correlation Analysis and Stationary Estimation Pave the Way for Fault Diagnosis in Ironmaking Processes
abstract
As the infrastructure industry continues to evolve toward digitalization, ongoing development of spatial and temporal data-based intelligent sensing guarantees safe operation. However, blast furnace ironmaking processes (BFIP) encounter a tricky dilemma in this revolution. Data-driven multivariate statistical analysis always fails for expected diagnosis performance due to complex dynamic, nonlinear, and nonstationary characteristics. To address this issue, we propose a novel method named modified mixed kernel-aided canonical stationary variate analysis (M2KCSVA). To start with, the past and future matrices and multiview nonlinear mapping of mixed kernel are properly considered to explore canonical stationary variables (CSVs) for both temporal correlation and weak stationarity. Especially, efficiency improvement procedures based on singular value decomposition and iterative modeling flow are deployed to reduce the computational cost and estimate accurate CSVs. In addition, we retain the smooth information without autocorrelation in the residuals for further analysis using stationary subspace analysis to generate static stationary variables. The corresponding two statistics and exponential difference contributions are computed for simultaneous fault detection and identification with an intuitive interpretation of dynamic and static stationary information. Experiments through an actual BFIP demonstrate that M2KCSVA surpasses comparison methods in terms of efficiency and accuracy.
Siwei Lou, Chunjie Yang 0001, Xujie Zhang, Hanwen Zhang 0002, Ping Wu 0001
IEEE Trans. Ind. Informatics3
2024 SFENOSA: A Novel KPI-Related Process Monitoring Method by Slow Feature Extraction and Elastic Net Orthonormal Subspace Analysis
abstract
Key performance indicators (KPIs), such as product quality variables or critical parameters in major units, play a crucial role in ensuring the desired performances in industrial processes. Nonetheless, focusing solely on monitoring process variables may result in the generation of nuisance alarms in response to disturbances that do not have a significant or meaningful impact on KPI variables. In this article, a novel KPI-related process monitoring method based on slow feature extraction and elastic net orthonormal subspace analysis (SFENOSA) is proposed. Traditional orthonormal subspace analysis (OSA) divides process data and KPI data subspaces into three orthonormal subspaces using least squares. To deal with the overfitting problem in high-dimensional space and enhance the robustness caused by correlated variables, the elastic net orthonormal subspace analysis (ENOSA) is developed by employing elastic net regularization in the OSA. Furthermore, to address the dynamic characteristics inherent in industrial processes, the slow feature analysis is naturally integrated into the framework of ENOSA for KPI-related process monitoring. Specifically, using the slow features extracted from process variables as the input and the KPI variables as the output, an ENOSA model is built. Based on the developed SFENOSA model, several monitoring statistics are established for KPI-related process monitoring. Experimental results on a numerical example, the well-known Tennessee Eastman process, and a real blast furnace ironmaking process demonstrate the superior performance of the proposed SFENOSA compared to the related methods.
Ping Wu 0001, Xujie Zhang, Siwei Lou, Jinfeng Gao 0002, Chunjie Yang 0001
IEEE Trans. Ind. Informatics3
2023 DiffCloth: Diffusion Based Garment Synthesis and Manipulation via Structural Cross-modal Semantic Alignment
abstract
Cross-modal garment synthesis and manipulation will significantly benefit the way fashion designers generate garments and modify their designs via flexible linguistic interfaces. However, despite the significant progress that has been made in generic image synthesis using diffusion models, producing garment images with garment part level semantics that are well aligned with input text prompts and then flexibly manipulating the generated results still remains a problem. Current approaches follow the general text-to-image paradigm and mine cross-modal relations via simple cross-attention modules, neglecting the structural correspondence between visual and textual representations in the fashion design domain. In this work, we instead introduce DiffCloth, a diffusion-based pipeline for cross-modal garment synthesis and manipulation, which empowers diffusion models with flexible compositionality in the fashion domain by structurally aligning the cross-modal semantics. Specifically, we formulate the part-level cross-modal alignment as a bipartite matching problem between the linguistic Attribute-Phrases (AP) and the visual garment parts which are obtained via constituency parsing and semantic segmentation, respectively. To mitigate the issue of attribute confusion, we further propose a semantic-bundled cross-attention to preserve the spatial structure similarities between the attention maps of attribute adjectives and part nouns in each AP. Moreover, DiffCloth allows for manipulation of the generated results by simply replacing APs in the text prompts. The manipulation-irrelevant regions are recognized by blended masks obtained from the bundled attention maps of the APs and kept unchanged. Extensive experiments on the CM-Fashion benchmark demonstrate that DiffCloth both yields state-of-the-art garment synthesis results by leveraging the inherent structural information and supports flexible manipulation with region consistency.
Xujie Zhang, Michael Kampffmeyer, Guansong Lu, Liang Lin 0004, Hang Xu 0004, Xiaodan Liang
ICCV1
2023 SRLI: Handling Irregular Time Series with a Novel Self-supervised Model Based on Contrastive Learning
Xujie Zhang, Qilong Han, Dan Lu 0004
ICONIP (13)2
2022 ARMANI: Part-level Garment-Text Alignment for Unified Cross-Modal Fashion Design
abstract
Cross-modal fashion image synthesis has emerged as one of the most promising directions in the generation domain due to the vast untapped potential of incorporating multiple modalities and the wide range of fashion image applications. To facilitate accurate generation, cross-modal synthesis methods typically rely on Contrastive Language-Image Pre-training (CLIP) to align textual and garment information. In this work, we argue that simply aligning texture and garment information is not sufficient to capture the semantics of the visual information and therefore propose MaskCLIP. MaskCLIP decomposes the garments into semantic parts, ensuring fine-grained and semantically accurate alignment between the visual and text information. Building on MaskCLIP, we propose ARMANI, a unified cross-modal fashion designer with part-level garment-text alignment. ARMANI discretizes an image into uniform tokens based on a learned cross-modal codebook in its first stage and uses a Transformer to model the distribution of image tokens for a real image given the tokens of the control signals in its second stage. Contrary to prior approaches that also rely on two-stage paradigms, ARMANI introduces textual tokens into the codebook, making it possible for the model to utilize fine-grain semantic information to generate more realistic images. Further, by introducing a cross-modal Transformer, ARMANI is versatile and can accomplish image synthesis from various control signals, such as pure text, sketch images, and partial images. Extensive experiments conducted on our newly collected cross-modal fashion dataset demonstrate that ARMANI generates photo-realistic images in diverse synthesis tasks and outperforms existing state-of-the-art cross-modal image synthesis approaches. Our code is available at https://github.com/Harvey594/ARMANI.
Xujie Zhang, Yu Sha, Michael Kampffmeyer, Zhenyu Xie, Zequn Jie, Chengwen Huang, Jianqing Peng, Xiaodan Liang
ACM Multimedia1
2021 UltraPose: Synthesizing Dense Pose with 1 Billion Points by Human-body Decoupling 3D Model
abstract
Recovering dense human poses from images plays a critical role in establishing an image-to-surface correspondence between RGB images and the 3D surface of the human body, serving the foundation of rich real-world applications, such as virtual humans, monocular-to-3d reconstruction. However, the popular DensePose-COCO dataset relies on a sophisticated manual annotation system, leading to severe limitations in acquiring the denser and more accurate annotated pose resources. In this work, we introduce a new 3D human-body model with a series of decoupled parameters that could freely control the generation of the body. Furthermore, we build a data generation system based on this decoupling 3D model, and construct an ultra dense synthetic benchmark UltraPose, containing around 1.3 billion corresponding points. Compared to the existing manually annotated DensePose-COCO dataset, the synthetic UltraPose has ultra dense image-to-surface correspondences without annotation cost and error. Our proposed UltraPose provides the largest benchmark and data resources for lifting the model capability in predicting more accurate dense poses. To promote future researches in this field, we also propose a transformer-based method to model the dense correspondence between 2D and 3D worlds. The proposed model trained on synthetic UltraPose can be applied to real-world scenarios, indicating the effectiveness of our benchmark and model.1
Haonan Yan, Xujie Zhang, Shengkai Zhang, Nianhong Jiao, Xiaodan Liang, Tianxiang Zheng 0001
ICCV3
2021 WAS-VTON: Warping Architecture Search for Virtual Try-on Network
abstract
Despite recent progress on image-based virtual try-on, current methods are constraint by shared warping networks and thus fail to synthesize natural try-on results when faced with clothing categories that require different warping operations. In this paper, we address this problem by finding clothing category-specific warping networks for the virtual try-on task via Neural Architecture Search (NAS). We introduce a NAS-Warping Module and elaborately design a bilevel hierarchical search space to identify the optimal network-level and operation-level flow estimation architecture. Given the network-level search space, containing different numbers of warping blocks, and the operation-level search space with different convolution operations, we jointly learn a combination of repeatable warping cells and convolution operations specifically for the clothing-person alignment. Moreover, a NAS-Fusion Module is proposed to synthesize more natural final try-on results, which is realized by leveraging particular skip connections to produce better-fused features that are required for seamlessly fusing the warped clothing and the unchanged person part. We adopt an efficient and stable one-shot searching strategy to search the above two modules. Extensive experiments demonstrate that our WAS-VTON significantly outperforms the previous fixed-architecture try-on methods with more natural warping results and virtual try-on results.
Zhenyu Xie, Xujie Zhang, Fuwei Zhao, Haoye Dong, Michael Kampffmeyer, Haonan Yan, Xiaodan Liang
ACM Multimedia2
2021 Data-Driven Fault Diagnosis Using Deep Canonical Variate Analysis and Fisher Discriminant Analysis
abstract
In this article, a novel data-driven fault diagnosis method by combining deep canonical variate analysis and Fisher discriminant analysis (DCVA-FDA) is proposed for complex industrial processes. Inspired by the recently developed deep canonical correlation analysis, a new nonlinear canonical variate analysis (CVA) called DCVA is first developed by incorporating deep neural networks into CVA. Based on DCVA, a residual generator is designed for the fault diagnosis process. FDA is applied in the feature space spanned by residual vectors. Then, a Bayesian inference classifier is performed in the reduced dimensional space of FDA to label the class of process data. A continuous stirred-tank reactor and an industrial benchmark of the Tennessee Eastman process are carried out to test the performance of DCVA-FDA fault diagnosis. The experimental results demonstrate that the proposed DCVA-FDA fault diagnosis is able to significantly improve the fault diagnosis performance when compared to other methods also examined in this article.
Ping Wu 0001, Siwei Lou, Xujie Zhang, Jiajun He 0002, Jinfeng Gao 0002
IEEE Trans. Ind. Informatics3
2021 Weakly Supervised Person Re-ID: Differentiable Graphical Learning and a New Benchmark
abstract
Person reidentification (Re-ID) benefits greatly from the accurate annotations of existing data sets (e.g., CUHK03 and Market-1501), which are quite expensive because each image in these data sets has to be assigned with a proper label. In this work, we ease the annotation of Re-ID by replacing the accurate annotation with inaccurate annotation, i.e., we group the images into bags in terms of time and assign a bag-level label for each bag. This greatly reduces the annotation effort and leads to the creation of a large-scale Re-ID benchmark called SYSU- 30k . The new benchmark contains 30k individuals, which is about 20 times larger than CUHK03 (1.3k individuals) and Market-1501 (1.5k individuals), and 30 times larger than ImageNet (1k categories). It sums up to 29606918 images. Learning a Re-ID model with bag-level annotation is called the weakly supervised Re-ID problem. To solve this problem, we introduce a differentiable graphical model to capture the dependencies from all images in a bag and generate a reliable pseudolabel for each person's image. The pseudolabel is further used to supervise the learning of the Re-ID model. Compared with the fully supervised Re-ID models, our method achieves state-of-the-art performance on SYSU- 30k and other data sets. The code, data set, and pretrained model will be available at https://github.com/wanggrun/SYSU-30k.
Guangrun Wang, Guangcong Wang, Xujie Zhang, Jian-Huang Lai, Zhengtao Yu 0001, Liang Lin 0004
IEEE Trans. Neural Networks Learn. Syst.3
2020 Fashion Editing With Adversarial Parsing Learning
abstract
Interactive fashion image manipulation, which enables users to edit images with sketches and color strokes, is an interesting research problem with great application value. Existing works often treat it as a general inpainting task and do not fully leverage the semantic structural information in fashion images. Moreover, they directly utilize conventional convolution and normalization layers to restore the incomplete image, which tends to wash away the sketch and color information. In this paper, we propose a novel Fashion Editing Generative Adversarial Network (FE-GAN), which is capable of manipulating fashion images by free-form sketches and sparse color strokes. FE-GAN consists of two modules: 1) a free-form parsing network that learns to control the human parsing generation by manipulating sketch and color; 2) a parsing-aware inpainting network that renders detailed textures with semantic guidance from the human parsing map. A new attention normalization layer is further applied at multiple scales in the decoder of the inpainting network to enhance the quality of the synthesized image. Extensive experiments on high-resolution fashion image datasets demonstrate that the proposed FE-GAN significantly outperforms the state-of-the-art methods on fashion image manipulation.
Haoye Dong, Xiaodan Liang, Xujie Zhang, Xiaohui Shen, Zhenyu Xie, Jian Yin 0001
CVPR4
2020 Formation Control of Time-varying Multi-agent System Based on BP Neural Network
abstract
In this paper, the formation control problem of time-varying multi-agent system (MAS) based on back propagation neural network (BPNN) is investigated. The system consists of two robots, which effectively improve the system operation efficiency through the coordination of robots. Moreover, leader-follower based strategy is used to formation. Among them, the leader robot can be controlled wirelessly through blue-tooth to share monitoring data in real time. The front image can be captured by the on-board camera mounted on the follower robot and optimized by the least square method (LSM). Then machine vision is used to locate the leader robot and the follower robot's position is adjusted in real time through proportion integral differential (PID) cnotroller based on BPNN to realize the formation. Compared with the traditional PID controller, BPNN PID cnotroller takes the weight coefficient of NN as an adaptive parameter, only few parameters need to be updated online, which can greatly lessen the computational burden. Finally, simulation experiments are given to illustrate the results.
Jinfeng Gao 0002, Xujie Zhang, Jiajun He 0002
ICARCV3
2012 Multilevel halftone screen design: Keeping texture or keeping smoothness?
abstract
Multilevel halftoning algorithms are becoming increasingly important as the capabilities of image output devices improve. The traditional approach to multilevel halftoning is hampered by the appearance of contouring in the vicinity of native tones of the output device. To overcome this limitation, we propose a novel framework based on maintaining a consistent periodic, clustered-dot halftone texture across the tone scale. We develop metrics for granularity and structure dissimilarity, and show how these can be used to guide the manner in which the halftone texture evolves from native tone to native tone, across the tone scale. Experimental results confirm the benefits of our new approach.
Xujie Zhang, Alex Veis, Robert Ulichney, Jan P. Allebach
ICIP1