Xiangzhong Fang

dblp:20/5598 · DBLP profile ↗
← Back
49ranked-venue papers
0as first author
23since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 33 · 14 since 2021Artificial intelligence and machine learning · 18 · 14 since 2021Applied, interdisciplinary, general and emerging computing · 2Systems, architecture and hardware · 1
YearPublicationVenuePosition
2026 FastFLUX: Pruning FLUX with Block-wise Replacement and Sandwich Training
abstract
Recent advancements in text-to-image (T2I) generation have led to the emergence of highly expressive models such as diffusion transformers (DiTs), exemplified by FLUX. However, their massive parameter sizes lead to slow inference, high memory usage, and poor deployability. Existing acceleration methods (e.g., single-step distillation and attention pruning) often suffer from significant performance degradation and incur substantial training costs. To address these limitations, we propose FastFLUX, an architecture-level pruning framework designed to enhance the inference efficiency of FLUX. At its core is the Block-wise Replacement with Linear Layers (BRLL) method, which replaces structurally complex residual branches in ResBlocks with lightweight linear layers while preserving the original shortcut connections for stability. Furthermore, we introduce Sandwich Training (ST), a localized fine-tuning strategy that leverages LoRA to supervise neighboring blocks, mitigating performance drops caused by structural replacement. Experiments show that our FastFLUX maintains high image quality under both qualitative and quantitative evaluations, while significantly improving inference speed, even with 20% of the hierarchy pruned.
Fuhan Cai, Jie Li 0022, Wenbo Li 0002, Jian Chen 0011, Xiangzhong Fang
AAAI6
2026 Reasoning with OmniThought: A Large CoT Dataset with Verbosity and Cognitive Difficulty Annotations
abstract
Tasks such as mathematical problem solving and coding require models to leverage chain-ofthought (CoT) processes, enabling human-like reasoning strategies.However, the advancement of large reasoning models (LRMs) is hindered by the lack of comprehensive CoT datasets.Existing resources often fail to provide extensive reasoning problems with coherent CoT processes distilled from multiple teacher models, and do not account for multifaceted properties describing the internal characteristics of CoTs.To address these challenges, we introduce OmniThought, a largescale dataset featuring 2 million CoT processes generated and validated by multiple powerful LRMs.Each CoT process in OmniThought is annotated with novel Reasoning Verbosity (RV) and Cognitive Difficulty (CD) scores, which characterize the appropriateness of CoT verbosity and the cognitive difficulty level for models to comprehend these reasoning processes.We further establish a self-reliant pipeline to curate this dataset.Extensive experiments using Qwen2.5 and Qwen3 of various sizes demonstrate the positive impact of our RV and CD scores on LRM training effectiveness.Based on the OmniThought dataset, we train and release a series of high-performing LRMs with enhanced reasoning abilities and optimized CoT output length.Our contributions advance the development of LRMs across different scales for solving complex reasoning tasks. 1
Wenrui Cai 0001, Chengyu Wang 0001, Jun Huang 0007, Xiangzhong Fang
ACL (1)5
2025 Neural Block Compression: Variable Bitrates Feature Blocks for Texture Representation
abstract
The imperative for compression of material textures emerges from the critical demand for high-quality rendering, which necessitates sophisticated textures that, in turn, require substantial storage and memory resources. Thus, low-bitrate compression is crucial, especially in modern games demanding higher texture resolutions. Concurrent methodologies in texture compression predominantly employ a block-based paradigm based on color space, which inevitably leads to representational redundancies and a limited compression scope, particularly at lower bitrates. In the context of mobile devices, bandwidth during texture loading and runtime memory are major bottlenecks, making existing compression algorithms inadequate for high-resolution textures. To mitigate these limitations, we propose a novel multi-resolution texture compression scheme, Neural Block Compression (NBC), developed within the neural feature domain. Our encoding scheme is constructed on a hierarchy of multi-resolution neural feature blocks, and the key ingredient is the variable bitrates quantization scheme. It allocates higher bitrates to higher feature mip-levels and lower bitrates to lower feature mip-levels, thereby extending the concept of block compression from color domain into neural feature domain. Extensive experiments demonstrate the superior texture compression quality achieved by the proposed scheme, especially at low bitrates.
Yishun Dou, Xiangzhong Fang, Wenjun Zhang 0001, Bingbing Ni
AAAI4
2025 Task-Adaptive Channel Attention Graph Network for Few-Shot 3D Point Cloud Classification
Zhongqiang Zhang 0004, Jingren Xie, Xiangzhong Fang
CGI (1)4
2025 Enhancing Reasoning Abilities of Small LLMs with Cognitive Alignment
abstract
The reasoning capabilities of large reasoning models (LRMs), such as OpenAI's o1 and DeepSeek-R1, have seen substantial advancements through deep thinking.However, these enhancements come with significant resource demands, underscoring the need for training effective small reasoning models.A critical challenge is that small models possess different reasoning capacities and cognitive trajectories compared with their larger counterparts.Hence, directly distilling chain-of-thought (CoT) rationales from large LRMs to smaller ones can sometimes be ineffective and often requires a substantial amount of annotated data.In this paper, we first introduce a novel Critique-Rethink-Verify (CRV) system, designed for training smaller yet powerful LRMs.Our CRV system consists of multiple LLM agents, each specializing in unique tasks: (i) critiquing the CoT rationales according to the cognitive capabilities of smaller models, (ii) rethinking and refining these CoTs based on the critiques, and (iii) verifying the correctness of the refined results.Building on the CRV system, we further propose the Cognitive Preference Optimization (CogPO) algorithm to continuously enhance the reasoning abilities of smaller models by aligning their reasoning processes with their cognitive capacities.Comprehensive evaluations on challenging reasoning benchmarks demonstrate the efficacy of our CRV+CogPO framework, which outperforms other methods by a large margin. 1
Wenrui Cai 0001, Chengyu Wang 0001, Jun Huang 0007, Xiangzhong Fang
EMNLP5
2025 Linear Multistep Solver Distillation for Fast Sampling of Diffusion Models
abstract
Sampling from diffusion models can be seen as solving the corresponding probability flow ordinary differential equation (ODE). The solving process requires a significant number of function evaluations (NFE), making it time-consuming. Recently, several solver search frameworks have attempted to find better-performing model-specific solvers. However, predicting the impact of intermediate solving strategies on final sample quality remains challenging, rendering the search process inefficient. In this paper, we propose a novel method for designing solving strategies. We first introduce a unified prediction formula for linear multistep solvers. Subsequently, we present a solver distillation framework, which enables a student solver to mimic the sampling trajectory generated by a teacher solver with more steps. We utilize the mean Euclidean distance between the student and teacher sampling trajectories as a metric, facilitating rapid adjustment and optimization of intermediate solving strategies. The design space of our framework encompasses multiple aspects, including prediction coefficients, time step schedules, and time scaling factors. Our framework has the ability to complete a solver search for Stable-Diffusion in under 12 total GPU hours. Compared to previous reinforcement learning-based search frameworks, our approach achieves over a 10$\times$ increase in search efficiency. With just 5 NFE, we achieve FID scores of 3.23 on CIFAR10, 7.16 on ImageNet-64, 5.44 on LSUN-Bedroom, and 12.52 on MS-COCO, resulting in a 2$\times$ sampling acceleration ratio compared to handcrafted solvers.
Xiangzhong Fang, Hanting Chen, Yunhe Wang 0001
ICLR2
2025 Generalized Category Discovery via Reciprocal Learning and Class-Wise Distribution Regularization
abstract
Generalized Category Discovery (GCD) aims to identify unlabeled samples by leveraging the base knowledge from labeled ones, where the unlabeled set consists of both base and novel classes. Since clustering methods are time-consuming at inference, parametric-based approaches have become more popular. However, recent parametric-based methods suffer from inferior base discrimination due to unreliable self-supervision. To address this issue, we propose a Reciprocal Learning Framework (RLF) that introduces an auxiliary branch devoted to base classification. During training, the main branch filters the pseudo-base samples to the auxiliary branch. In response, the auxiliary branch provides more reliable soft labels for the main branch, leading to a virtuous cycle. Furthermore, we introduce Class-wise Distribution Regularization (CDR) to mitigate the learning bias towards base classes. CDR essentially increases the prediction confidence of the unlabeled data and boosts the novel class performance. Combined with both components, our proposed method, RLCD, achieves superior performance in all classes with negligible extra computation. Comprehensive experiments across seven GCD datasets validate its superiority. Our codes are available at https://github.com/APORduo/RLCD.
Zhiquan Tan, Linglan Zhao, Xiangzhong Fang, Weiran Huang 0001
ICML5
2025 Few-Shot Class-Incremental Learning via Asymmetric Supervised Contrastive Learning
abstract
Few-Shot Class-Incremental Learning (FSCIL) is to continuously learn novel classes from a few samples without forgetting previous knowledge. Adapting directly to limited novel data typically results in significant forgetting of base class knowledge. Consequently, prevailing FSCIL methods are devoted to training a strong initial model that can be frozen in incremental sessions. However, these works face a dilemma in poor generalization: they benefit mainly from base class performance, yet underperform in novel classes. To alleviate this issue, we design a two-stage training framework to simultaneously enhance generalization for novel classes and maintain base class discrimination. In the first stage, an asymmetric supervised contrastive learning (AsyCon) algorithm is proposed. AsyCon introduces a predicted feature to achieve an asymmetric alignment of positive pairs. It alleviates over-similarity within positive features, allowing the model to better transfer to new classes in incremental sessions. In the second stage, the model is finetuned for promoting its performance on base classes. To maintain the generalization obtained in the first stage, we employ an L2 normalized regularization (LR) to keep the feature consistent with the model in the first stage. The finetuned model, termed AsyCLR, effectively balances generalization and discrimination, significantly outperforming existing FSCIL works especially in novel class accuracy. Experiments on CUB200, CIFAR100, and mini-ImageNet verify the effectiveness of our method. Additionally, our method also performs well in the standard few-shot recognition scenario due to its strong generalization ability. Our codes are available at https://github.com/APORduo/AsyCLR.
Duo Liu 0001, Linglan Zhao, Fan Lyu, Xiangzhong Fang, Liang Wang 0001
IEEE Trans. Circuits Syst. Video Technol.5
2024 NER-guided Comprehensive Hierarchy-aware Prompt Tuning for Hierarchical Text Classification
abstract
Hierarchical text classification (HTC) is a significant but challenging task in natural language processing (NLP) due to its complex taxonomic label hierarchy. Recently, there have been a number of approaches that applied prompt learning to HTC problems, demonstrating impressive efficacy. The majority of prompt-based studies emphasize global hierarchical features by employing graph networks to represent the hierarchical structure as a whole, with limited research on maintaining path consistency within the internal hierarchy of the structure. In this paper, we formulate prompt-based HTC as a named entity recognition (NER) task and introduce conditional random fields (CRF) and Global Pointer to establish hierarchical dependencies. Specifically, we approach single- and multi-path HTC as flat and nested entity recognition tasks and model them using span- and token-based methods. By narrowing the gap between HTC and NER, we maintain the consistency of internal paths within the hierarchical structure through a simple and effective way. Extensive experiments on three public datasets show that our method achieves state-of-the-art (SoTA) performance.
Fuhan Cai, Xiaozhe Yang, Xiangzhong Fang
LREC/COLING6
2024 Learning Quantized Adaptive Conditions for Diffusion Models
Yuchan Tian, Huaao Tang, Jie Hu 0021, Xiangzhong Fang, Hanting Chen
ECCV (81)6
2024 COPHTC: Contrastive Learning with Prompt Tuning for Hierarchical Text Classification
abstract
Hierarchical Text Classification (HTC) is an essential yet challenging task in natural language processing (NLP) due to its complex label structure. Recently, a number of approaches have employed prompt learning in HTC, achieving noteworthy outcomes. However, prompt-based HTC does not further optimize the representation of samples based on label relationships, nor does it dynamically adjust the relative positions of sample pairs from the perspective of the embedding space. In this work, we propose a COntrastive-enhanced Promptbased model for HTC tasks (COPHTC). More specifically, we integrate contrastive learning into prompt tuning while employing momentum updates and a dynamic queue to offer a greater variety of positive and negative samples for the input texts. Extensive experiments on three public datasets verify the effectiveness of COPHTC.
Fuhan Cai, Xiangzhong Fang
ICASSP4
2024 Distillation Excluding Positives for Few-Shot Class-Incremental Learning
abstract
Few-Shot Class-Incremental Learning (FSCIL) defines a challenging task to continually recognize novel classes with few training data without forgetting old classes. Considering the catastrophic forgetting and overfitting issues, mainstream FSCIL methods resort to obtaining a strong model in the base session and freezing it in incremental sessions. Although prevailing methods perform well in the base classes, they often struggle with poor novel class generalization. To strengthen the representation of these models, this paper focuses on Knowledge Distillation (KD). Since existing KD methods are incompetent for FSCIL and introduce limited improvement, we propose the Distillation Excluding Positives (DEP) method for boosting performance on FSCIL tasks. Specifically, DEP consists of negative relationship distillation and asymmetric self-feature distillation. It can mitigate the over-similarity of intra-class features, leading to a more generalized model. Extensive experiments on three FSCIL benchmarks validate the superiority of DEP over current SOTAs.
Linglan Zhao, Fuhan Cai, Xiangzhong Fang
ICME5
2024 Mix background and foreground separately: Transformer-based Augmentation Strategies for Domain Generalization
abstract
Domain generalization (DG) aims to alleviate the severe performance degradation of deep neural networks when a domain shift exists between training and testing data. The different distributions of the backgrounds of samples represent a vital factor contributing to the domain gap. Inspired by the augmentation-based method, we aim to mix up the backgrounds of samples to generate new images for training. The Mixup method, a classical approach, randomly blends two samples at the image level. Therefore, it cannot resolve the causal relationship between backgrounds and semantic labels. To solve this problem, we introduce a new method that separates the foreground and background of samples at the patch level using a Vision Transformer network (ViT). Concretely, we calculate attention scores of each patch based on self-attention modules in ViT. Then, we identify the background or foreground by the rank of attention scores. Next, we present a Background-Mix method to blend the common background of samples from different domains, regularizing the model to ignore causal relations between background and semantic label. Moreover, the different appearances of objects across distinct source domains also contribute to the performance drop on the target domain. Therefore, we design a Foreground-Mix method to mix only the object part, excluding the background. This can enable the model to classify objects in multiple patterns of representation. The entire framework is collectively referred to as BFMix. Extensive experiments demonstrate that our method achieves state-of-the-art performance.
Fuhan Cai, Xiangzhong Fang
ICME5
2023 Few-Shot Class-Incremental Learning via Class-Aware Bilateral Distillation
abstract
Few-Shot Class-Incremental Learning (FSCIL) aims to continually learn novel classes based on only few training samples, which poses a more challenging task than the well-studied Class-Incremental Learning (CIL) due to data scarcity. While knowledge distillation, a prevailing technique in CIL, can alleviate the catastrophic forgetting of older classes by regularizing outputs between current and previous model, it fails to consider the overfitting risk of novel classes in FSCIL. To adapt the powerful distillation technique for FSCIL, we propose a novel distillation structure, by taking the unique challenge of overfitting into account. Concretely, we draw knowledge from two complementary teachers. One is the model trained on abundant data from base classes that carries rich general knowledge, which can be leveraged for easing the overfitting of current novel classes. The other is the updated model from last incremental session that contains the adapted knowledge of previous novel classes, which is used for alleviating their forgetting. To combine the guidances, an adaptive strategy conditioned on the class-wise semantic similarities is introduced. Besides, for better preserving base class knowledge when accommodating novel concepts, we adopt a two-branch network with an attention-based aggregation module to dynamically merge predictions from two complementary branches. Extensive experiments on 3 popular FSCIL datasets: mini-ImageNet, CIFAR100 and CUB200 validate the effectiveness of our method by surpassing existing works by a significant margin. Code is available at https://github.com/LinglanZhao/BiDistFSCIL.
Linglan Zhao, Jing Lu 0004, Yunlu Xu, Zhanzhan Cheng, Dashan Guo, Xiangzhong Fang
CVPR7
2023 Rethinking Self-Supervision for Few-Shot Class-Incremental Learning
abstract
Few-Shot Class-Incremental Learning (FSCIL) focuses on progressively absorbing new concepts given only limited training data. For tackling this challenge, several recent FS-CIL works resort to pre-training models with Self-Supervised Learning (SSL) to obtain features that can generalize well to new classes. However, to avoid overfitting and catastrophic forgetting, previous works only leverage SSL in the base session and keep all or most parameters fixed in incremental sessions, resulting in inadequate adaptation to novel classes. Thus, in this paper, we explore the setting where more parameters can be updated for adapting to novel concepts, and discover that the model pre-trained with SSL leads to degraded performance even compared to that without SSL. It can be attributed to the severer forgetting of base class knowledge. To address this issue, we propose an imprinting-based distillation module for effectively regularizing the adaption process, and a mathematically provable routing strategy for further improved results. The effectiveness of our approach is verified on 3 popular FSCIL benchmarks by significantly outperforming previous methods.
Linglan Zhao, Jing Lu 0004, Zhanzhan Cheng, Xiangzhong Fang
ICME5
2023 Efficient Attention for Domain Generalization
Fuhan Cai, Xiangzhong Fang
ICONIP (9)5
2023 Boosting domain generalization by domain-aware knowledge distillation
Zhongqiang Zhang 0004, Fuhan Cai, Xiangzhong Fang
Knowl. Based Syst.5
2022 Im2Oil: Stroke-Based Oil Painting Rendering with Linearly Controllable Fineness Via Adaptive Sampling
abstract
This paper proposes a novel stroke-based rendering (SBR) method that translates images into vivid oil paintings. Previous SBR techniques usually formulate the oil painting problem as pixel-wise approximation. Different from this technique route, we treat oil painting creation as an adaptive sampling problem. Firstly, we compute a probability density map based on the texture complexity of the input image. Then we use the Voronoi algorithm to sample a set of pixels as the stroke anchors. Next, we search and generate an individual oil stroke at each anchor. Finally, we place all the strokes on the canvas to obtain the oil painting. By adjusting the hyper-parameter maximum sampling probability, we can control the oil painting fineness in a linear manner. Comparison with existing state-of-the-art oil painting techniques shows that our results have higher fidelity and more realistic textures. A user opinion test demonstrates that people behave more preference toward our oil paintings than the results of other methods. More interesting results and the code are in https://github.com/TZYSJTU/Im2Oil.
Zhengyan Tong, Xiaohang Wang 0004, Shengchao Yuan, Xuanhong Chen, Xiangzhong Fang
ACM Multimedia6
2022 Boosting Few-shot visual recognition via saliency-guided complementary attention
Linglan Zhao, Dashan Guo, Wei Li 0084, Xiangzhong Fang
Neurocomputing5
2021 A Strong Baseline for Semi-Supervised Incremental Few-Shot Learning
Linglan Zhao, Dashan Guo, Yunlu Xu, Liang Qiao 0001, Zhanzhan Cheng, Shiliang Pu, Xiangzhong Fang
BMVC8
2021 Saliency-Guided Complementary Attention for Improved Few-Shot Learning
abstract
Despite significant progress in recent deep neural networks, most deep learning algorithms rely heavily on abundant training samples. To address this problem, we propose an effective and interpretable few-shot classification model using Saliency-Guided Complementary Attention (SGCA), which aims to learn transferable representations and to build a robust classification module simultaneously. Concretely, we propose to train our feature extractor using an auxiliary task to separate object regions from background clutter guided by saliency detection signals. In addition, to make the separation beneficial to the downstream tasks, we introduce a complementary attention mechanism to force the classification module to focus on various informative parts of the image. Extensive experiments on few-shot learning tasks demonstrate the effectiveness of our proposed method, e.g., we achieve 68.81% and 84.60% for 5-way 1-shot and 5-shot settings on mini-ImageNet, respectively.
Linglan Zhao, Dashan Guo, Wei Li 0084, Xiangzhong Fang
ICME5
2021 Class-wise Metric Scaling for Improved Few-Shot Classification
abstract
Few-shot classification aims to generalize basic knowledge to recognize novel categories from a few samples. Recent centroid-based methods achieve promising classification performance with the nearest neighbor rule. However, we consider that those methods intrinsically ignore per-class distribution, as the decision boundaries are biased due to the diversity of intra-class variances. Hence, we propose a class-wise metric scaling (CMS) mechanism, which can be applied to both training and testing stages. Concretely, metric scalars are set as learnable parameters in the training stage, helping to learn a more discriminative and transferable feature representation. As for testing, we construct a convex optimization problem to generate an optimal scalar vector for refining the nearest neighbor decisions. Besides, we also involve a low-ranking bilinear pooling layer for improved representation capacity, which further provides significant performance gains. Extensive experiments are conducted on a series of feature extractor backbones, datasets, and testing modes, which have shown consistent improvements compared to prior SOTA methods, e.g., we achieve accuracies of 66.64 % and 83.63 % for 5-way 1-shot and 5-shot settings on the mini-ImageNet, respectively. Under the semi-supervised inductive mode, results are further up to 78.34 % and 87.53 %, respectively.
Linglan Zhao, Wei Li 0084, Dashan Guo, Xiangzhong Fang
WACV5
2021 PDA: Proxy-based domain adaptation for few-shot image recognition
Linglan Zhao, Xiangzhong Fang
Image Vis. Comput.3
2020 Moflowgan: Video Generation With Flow Guidance
abstract
In recent years, video generation has attracted a lot of attention in the computer vision community. Unlike image generation which only focuses on appearance, video generation requires modeling both content information and motion dynamics. In this work, we propose MoFlowGAN, which explicitly models motion dynamics by a content-motion decomposition architecture with an additional flow generator. The decomposition architecture models content and motion separately and is instantiated by a compact variant of BigGAN [1]. Besides, the flow generator generates optical flow directly based on high-level feature maps of adjacent frames as a strong supervision, hence the searching space of motion patterns is highly reduced. Our proposed MoFlowGAN achieves the state-of-the-art results on both MUG facial expression and UCF-101 datasets.
Wei Li 0084, Zehuan Yuan, Xiangzhong Fang, Changhu Wang
ICME3
2020 Visual question answering with attention transfer and a cross-modal gating mechanism
Wei Li 0084, Jianhui Sun, Linglan Zhao, Xiangzhong Fang
Pattern Recognit. Lett.5
2019 Refining Proposals with Neighboring Contexts for Temporal Action Detection
abstract
While many methods have been proposed for generating temporal proposals, the performance of existing temporal detection pipelines is still limited by the quality of proposals. In this paper, we introduce a new refining model for temporal action detection, which incorporates the evaluation of Intersection-over-Union (IoU) value into the action classification, and then regresses a better segment using the neighboring contexts of one candidate proposal. To refine one candidate proposal, we augment it with two neighboring proposals of equal length, which capture the contextual information from the past and future segments. After extracting regional features for the augmented proposals, we utilize the dilated convolutions for contextual modeling to regress the offsets between the candidate proposal and the target segment within this augmented area. Extensive experiments on THUMOS14 demonstrate that our method successfully refines the generated proposals and achieves superior detection performance over other methods.
Dashan Guo, Wei Li 0084, Jianhui Sun, Xiangzhong Fang
ICME5
2018 Multimodal architecture for video captioning with memory networks and an attention mechanism
Wei Li 0084, Dashan Guo, Xiangzhong Fang
Pattern Recognit. Lett.3
2018 Fully Convolutional Network for Multiscale Temporal Action Proposals
abstract
Similar to the function of object proposals in localizing objects within images, temporal action proposals can facilitate the extraction of semantic segments and simplify the computations required for temporal action localization in untrimmed videos. In this paper, we propose a fully convolutional network to identify multiscale temporal action proposals (FCN-TAP) that utilizes only the temporal convolutions to retrieve accurate action proposals for video sequences. Using gated linear units, our network enables simple but powerful inferences, and by parallelizing the computations, it significantly improves performances compared with previous recurrent models. To capture more temporal contexts with fewer parameters, we apply dilated convolutions to expand the receptive fields of our network. Moreover, we divide the receptive fields into multiple scale ranges and then refine the corresponding temporal boundaries using duration regression at each scale. To generate suitable segments with arbitrary durations for training, we design a new strategy to select sampled candidates within the corresponding scale range. The power of our method is demonstrated through experiments on the THUMOS'14 and ActivityNet datasets, where FCN-TAP performs better and achieves a remarkable speedup compared to other state-of-the-art methods. Additional experiments show that our method generates high-quality proposals and improves the localization stage of existing action detection pipelines.
Dashan Guo, Wei Li 0084, Xiangzhong Fang
IEEE Trans. Multim.3
2018 Exact Confidence Limits for the Acceleration Factor Under Constant-Stress Partially Accelerated Life Tests With Type-I Censoring
abstract
In modern lifetime assessment and reliability analysis, accelerated life test has frequently been used to yield information quickly so that the life distribution of products can be estimated. This paper considers the estimate of the acceleration factor for the exponentially distributed lifetimes under the constant-stress partially accelerated life test with Type-I censored data. In many applications with small sample sizes, the approximate confidence limits for the parameters based on large-sample asymptotic distributions or the bootstrap method are usually not accurate enough. This study defines two ordering relations on the sample space by the generalized maximum likelihood estimator of the acceleration factor, and proposes one new approach of constructing the exact lower and upper confidence limits for the acceleration factor. An efficient procedure of computing the exact lower and upper confidence limits for the acceleration factor is presented via the EM algorithm. The approximate confidence limits for the acceleration factor using the asymptotic and bootstrap methods are also derived in this study. The proposed exact approach is compared with the two approximate methods by carrying out extensive simulation studies, and it is shown that the exact method performs well and is robust in small sample settings. Finally, we present a real example of the accelerated life test to illustrate all methods studied in this paper.
Deqiang Zheng, Xiangzhong Fang
IEEE Trans. Reliab.2
2017 Lossless image compression algorithm and hardware architecture for bandwidth reduction of external memory
abstract
In high definition (HD) video coders, huge memory access bandwidth is the major throughput bottleneck. Lossless embedded compression is an efficient solution to alleviate the bandwidth burden, in which image are compressed before writing into local memory and decompressed after retrieving from local memory. This study proposes a hardware‐oriented lossless image compression algorithm, supporting block and line random access flexibly for adapting diverse hardware video codec architectures. The major contributions are characterised as follows. First, block or pixel‐level adaptive prediction is proposed to fully utilise the image spatial correlation by employing adaptive mode decision. Second, multiple‐range semi‐fixed (SF) variable length coding (VLC) is employed to describe the prediction residue, and adaptive block size selection is employed for SF VLC to fully utilise the statistical redundancy. In addition, Huffman VLC is further employed to represent the control syntax elements. Third, four‐stage pipeline hardware architecture is proposed to implement the proposed algorithm. Simulation results show that the proposed algorithm achieves competitive rate compression performance compared with reference algorithms. The proposed hardware architecture is verified supporting real‐time processing for quad‐HD videos at the frequency of 166 MHz. The proposed work achieves reducing memory access bandwidth by ∼55.2%, which is useful for hardwired video coding.
Shizhong Li, Hai Bing Yin, Xiangzhong Fang, Huijuan Lu
IET Image Process.3
2017 Capturing Temporal Structures for Video Captioning by Spatio-temporal Contexts and Channel Attention Mechanism
Dashan Guo, Wei Li 0084, Xiangzhong Fang
Neural Process. Lett.3
2013 A Video Text Detection and Tracking System
abstract
Faced with the increasing large scale video databases, retrieving videos quickly and efficiently has become a crucial problem. Video text, which carries high level semantic information, is a type of important source that is useful for this task. In this paper, we introduce a video text detecting and tracking approach. By these methods we can obtain clear binary text images, and these text images can be processed by OCR (Optical Character Recognition) software directly. Our approach including two parts, one is stroke-model based video text detection and localization method, the other is SURF (Speeded Up Robust Features) based text region tracking method. In our detection and localization approach, we use stroke model and morphological operation to roughly identify candidate text regions. Combine stroke-map and edge response to localize text lines in each candidate text regions. Several heuristics and SVM (Support Vector Machine) used to verifying text blocks. The core part of our text tracking method is fast approximate nearest-neighbour search algorithm for extracted SURF features. Text-ending frame is determined based on SURF feature point numbers, while, text motion estimation is based on correct matches in adjacent frames. Experimental result on large number of different video clips shows that our approach can effectively detect and track both static texts and scrolling texts.
Tuoerhongjiang Yusufu, Xiangzhong Fang
ISM3
2012 Frame Rate Up-Conversion for Depth-Based 3D Video
abstract
A novel frame rate up-conversion (FRUC) scheme for depth-based 3D video is proposed. Differing from the existing conventional FRUC methods which are designed for two dimensional (2D) video only, the proposed method is designed for depth-based 3D video, which increases the frame rate of both color sequences and its associated depth maps by jointly considering the image intensity, depth information, and spatial-temporal correlations. The proposed method contains motion estimation, block irregular segmentation (BIS), depth-constrained motion vector post-processing, forward and backward motion compensation, and edge-preserved combination. Experimental results show that the intermediate color and depth images interpolated by our proposed method provide a good image quality both objectively and subjectively. Moreover, compared with conversional FRUC methods, the color and depth map sequences up-converted by the proposed method are more suitable for further virtual view synthesis.
Qingchun Lu, Xiangzhong Fang, Yongzhe Wang
ICME2
2011 Integrating Visual Saliency and Consistency for Re-Ranking Image Search Results
abstract
In this paper, we propose a new algorithm for image re-ranking in web image search applications. The proposed method focuses on investigating the following two mechanisms: 1) Visual consistency. In most web image search cases, the images that closely related to the search query are visually similar. These visually consistent images which occur most frequently in the first few web pages will be given higher ranks. 2) Visual saliency. From visual aspect, it is obvious that salient images would be easier to catch users' eyes, and it is observed that these visually salient images in the front pages are often relevant to the user's query. By integrating the above two mechanisms, our method can efficiently re-rank the images from search engines and obtain a more satisfactory search result. Experimental results on a real-world web image dataset demonstrate that our approach can effectively improve the performance of image retrieval.
Xiaokang Yang 0001, Xiangzhong Fang, Weisi Lin, Rui Zhang 0052
IEEE Trans. Multim.3
2010 Integrating visual saliency and consistency for re-ranking image search results
abstract
The paper investigates two mechanisms, visual consistency and visual saliency, in web image search: (1) In most current web image search engines, such as Google Image Search and Yahoo Image Search, the images that closely related to the search query are typically visually similar. These visually consistent images which occur most frequently in the first few web pages will be given higher ranks. (2) From visual aspect, it is obvious that salient images would be easier to catch users' eyes and more likely to be clicked than the cluttered ones in low-level vision. In addition, we also observe the fact that the visually salient images in the front pages are often relevant to the user's query. The principal novelty of this paper is in combining visual saliency and consistency to re-rank the results from search engines to make the re-ranked images more satisfying in both vision and content. The experimental results on a real world web image dataset demonstrate that our approach can effectively improve the performance of image retrieval.
Xiaokang Yang 0001, Rui Zhang 0052, Fuxiang Lu, Xiangzhong Fang
ICIP5
2010 Adaptively Adjusted Gaussian Mixture Models for Surveillance Applications
Tianci Huang, Xiangzhong Fang, Jingbang Qiu, Takeshi Ikenaga
MMM2
2010 Content-adaptive spatial scalability for scalable video coding
abstract
This paper presents an enhancement of the SVC extension of the H.264/AVC standard by content-adaptive spatial scalability (CASS). CASS introduces a novel functionality which is important for high quality content distribution. The video streams (spatial layers), which are used as input to the encoder, are created by content-adaptive and art-directable retargeting of existing high resolution video. Video is retargeted to resolutions and aspect ratios which are mainly dictated by target display devices. Thereby no content is cut off, but visually important content is preserved at the expense of a non-linear distortion of visually unimportant areas. The non-linear dependencies between such video streams are efficiently exploited by CASS for scalable coding. This is achieved by integrating warping-based non-linear texture prediction and warp coding into the SVC framework. The results indicate high prediction accuracy of non-linear predictors and high compression efficiency with limited increase in bit rate and complexity compared to the standard SVC for the case of INTRA only coding.
Yongzhe Wang, Nikolce Stefanoski, Xiangzhong Fang, Aljoscha Smolic
PCS3
2009 Parallel HD Encoding on CELL
abstract
The Cell Broadband Engine Architecture (CBEA) is an excellent architecture for high performance distributed computing and multimedia processing. While the Cell/BE processor is capable of high definition H.264 encoding, there are still no such implementations available. In this paper, we present a parallel implementation of a HD H.264 encoder on this heterogeneous nine cores processor. First we implement a real time SD encoder on a single SPU by optimizing Motion Estimation algorithm, DMA transfers etc. Then we propose a pipelined parallel encoding algorithm for multicore processors, and use this algorithm to get a real time HD H.264 encoder (1920×1080@31fps) by using eight SPEs (58fps on 16 SPEs).
Xiangzhong Fang, Ci Wang, Satoshi Goto
ISCAS2
2008 Multilevel Framework to Detect and Handle Vehicle Occlusion
abstract
This paper presents a multilevel framework to detect and handle vehicle occlusion. The proposed framework consists of the intraframe, interframe, and tracking levels. On the intraframe level, occlusion is detected by evaluating thecompactness ratioandinterior distance ratioof vehicles, and the detected occlusion is handled by removing a “cutting region” of the occluded vehicles. On the interframe level, occlusion is detected by performing subtractive clustering on the motion vectors of vehicles, and the occluded vehicles are separated according to the binary classification of motion vectors. On the tracking level, occlusion layer images are adaptively constructed and maintained, and the detected vehicles are tracked in both the captured images and the occlusion layer images by performing a bidirectional occlusion reasoning algorithm. The proposed intraframe, interframe, and tracking levels are sequentially implemented in our framework. Experiments on various typical scenes exhibit the effectiveness of the proposed framework. Quantitative evaluation and comparison demonstrate that the proposed method outperforms state-of-the-art methods.
Wei Zhang 0025, Q. M. Jonathan Wu, Xiaokang Yang 0001, Xiangzhong Fang
IEEE Trans. Intell. Transp. Syst.4
2007 Multiframe Super-Resolution Reconstruction Based on Cycle-Spinning
abstract
A multiframe super-resolution (SR) reconstruction algorithm based on cycle-spinning (CS) is proposed. We utilize the relative motion information of sequential images to construct a CS-based framework for the resolution enhancement. The unique feature of the proposed algorithm is that it is effective for low-resolution (LR) images with various point spread function (PSF) and noise characteristics, even if the degradation models are unknown for the imaging system. Moreover, the computational complexity is inexpensive. Experiments demonstrate the effectiveness of the proposed method and show the superiority to previous methods in objective and subjective qualities.
Xiangzhong Fang, Songyu Yu
ICASSP (1)2
2007 Spread-Spectrum Audio Watermark Robust Against Pitch-Scale Modification
abstract
A new digital audio watermarking for copyright protection is described in this paper. The watermark is embedded in the time domain by taking use of spread-spectrum technique and exploiting the hearing thresholds in the psychoacoustic model. The scheme of detection is based on high-pass filtering and correlation. The decision threshold is adaptively determined according to the statistical distribution of correlation values. Besides, the postprocessing technique is introduced for resisting time-scale modification. Experimental results show that the proposed watermarking scheme is high robust to almost any attacks, including cropping, low-pass filtering, noise addition, MP3 compression, sampling rate changing, requantization, time-scale modification, and pitch-scale modification.
Jianling Hu, Xiangzhong Fang
ICME3
2007 Moving Cast Shadows Detection Using Ratio Edge
abstract
Moving objects segmentation plays a very important role in real-time image analysis. However, as one of the common parts in the natural scenes, shadows severely interfere with the accuracy of moving objects detection in video surveillance. In this paper, we present a novel method for moving cast shadows detection. Based on the analysis of the physical model of moving shadows, we prove that the ratio edge is illumination invariant. The distribution of the ratio edge is discussed and a significance test is performed to classify each moving pixel into foreground object or moving shadow. Intensity constraint and geometric heuristics are imposed to further improve the performance. Experiments on various typical scenes exhibit the robustness of the proposed method. Extensively quantitative evaluation and comparison demonstrate that the proposed method significantly outperforms state-of-the-art methods.
Wei Zhang 0025, Xiangzhong Fang, Xiaokang Yang 0001, Q. M. Jonathan Wu
IEEE Trans. Multim.2
2006 Motion Vector Smoothing for True Motion Estimation
abstract
This paper proposes a new motion vector (MV) smoothing algorithm to track the real motion in image sequences for MPEG video encoders. First, a pre-checking algorithm is employed to eliminate wrong motion vectors and preserve all possible motion vectors. For each block considered, the motion similarity between the neighboring blocks and the number of candidate motion vectors are jointly exploited to adaptively grow the filtering support, which is supposed to have homogeneous motion and sufficient spatial gradient. Then, all candidate motion vectors are checked within the filtering support using a new motion smoothness-constrained matching criteria. The simulation results show that the proposed algorithm can efficiently track the real motion resulting in smooth motion vector field (MVF).
Hai Bing Yin, Xiangzhong Fang, Hua Yang 0001, Songyu Yu, Xiaokang Yang 0001
ICASSP (2)2
2006 Multi-Rate, Dynamic and Compliant Region of Interest Coding for JPEG2000
abstract
A method is proposed to encode multiple regions of interest (ROI) in JPEG2000 image. It rearranges truncation point for every codeblock in each layer. It assigns higher bitrate to ROI and lower bitrate to non-ROI and combines them to codestream. The proposed strategy produces a fully compliant JPEG2000 codestream. It allows transmission of different ROIs with different priorities and supports dynamic delineation and prioritization of them. Experimental results demonstrating the validity of the proposed approach are presented
Xiangzhong Fang, Haibin Yin, Songyu Yu
ICME2
2006 Moving vehicles segmentation based on Bayesian framework for Gaussian motion model
Wei Zhang 0025, Xiangzhong Fang, Xiaokang Yang 0001
Pattern Recognit. Lett.2
2005 Rate Control for Motion JPEG2000 Using Correlation Prediction
abstract
Motion JPEG2000 has been widely used for its fine scalability with bit stream and its efficiency. A single scalable bitstream can provide precise rate control for variable bitrate (VBR) traffic. This paper presents an algorithm that makes use of correlation among the video frames to get consistent quality. By using one frame's distortion-rate to predict the next several frames' distortion-rate values, according to correlation among frames, we use only a few buffers to achieve highly consistent quality which used to use many buffers to achieve. Experimental results using motion JPEG2000 demonstrate substantial benefits.
Xiangzhong Fang, Rong Hou
ICASSP (2)2
2005 Joint Rate Control for Multiple Sequences coding based on H.264 standard
abstract
The objective of joint rate control is to dynamically distribute the channel capacity among video sequences according to their respective complexities, thus a more uniform picture quality and a more efficient utilization of channel capacity are achieved. Most existing approaches are based on MPEG2 coding platform. This paper presents a novel joint rate control scheme for multiple video sequences coding based on H.264 standard. A novel complexity measure that adapts to the characteristics of H.264 video coding is proposed. Experimental results show that the proposed scheme maintains a good balance in picture quality among the sequences as well as within a sequence
Xiangzhong Fang, Hongkai Xiong
ICME2
2005 A new rate control algorithm for macroblock-level coders
abstract
Rate control plays an important role in video coding. It regulates the coded bits to satisfy the channel rate while keeping good video quality. In Tsai's paper, a new algorithm was proposed which rearranging the macroblocks' coding order according to their significance in each frame. More complex macroblocks will be coded with more priority in coding order. But it is required to rearrange the coding order and output the encoded bits in the original scan order. In this paper, a new rate control algorithm is proposed which recalculates the mquant of each macroblock inside a frame according to its significance. Furthermore, it can be applied to all macroblock-level coders. Simulation results show that this new algorithm can achieve obvious improvement in video quality like Tsai's algorithm while has lower complexity
Yutao Dong, Xiangzhong Fang, Hao Liu 0010
MMSP2
2005 Unequal Forced Intra-Refresh for Real-time Multicast Video
abstract
In motion-compensated video coding, the errors caused by packet loss not only impair the reconstruction quality of current frame, but also lead to error propagation to subsequent frames. Based on the error-propagation analysis in a group of pictures (GOP), we propose an unequal forced intra-refresh scheme to increase error resilience of multicast video. According to a GOP-level error-propagation model, the proposed scheme can distribute the unequal number of forced intra-mode MBs to different P-frames of a GOP. Experimental results show that the proposed scheme can effectively mitigate the error-propagation effect and achieve about 0.1~1.1 dB gains over the traditional average scheme in H.264/AVC
Hao Liu 0010, Wenjun Zhang 0001, Yutao Dong, Xiangzhong Fang
MMSP4