VLDB 2026 Research / reviewers in the wild / expert
Xiangzhong Fang
dblp:20/5598
· DBLP profile ↗
49ranked-venue papers
0as first author
23since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 33 · 14 since 2021Artificial intelligence and machine learning · 18 · 14 since 2021Applied, interdisciplinary, general and emerging computing · 2Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FastFLUX: Pruning FLUX with Block-wise Replacement and Sandwich TrainingabstractRecent advancements in text-to-image (T2I) generation have led to the emergence of highly expressive models such as diffusion transformers (DiTs), exemplified by FLUX. However, their massive parameter sizes lead to slow inference, high memory usage, and poor deployability. Existing acceleration methods (e.g., single-step distillation and attention pruning) often suffer from significant performance degradation and incur substantial training costs. To address these limitations, we propose FastFLUX, an architecture-level pruning framework designed to enhance the inference efficiency of FLUX. At its core is the Block-wise Replacement with Linear Layers (BRLL) method, which replaces structurally complex residual branches in ResBlocks with lightweight linear layers while preserving the original shortcut connections for stability. Furthermore, we introduce Sandwich Training (ST), a localized fine-tuning strategy that leverages LoRA to supervise neighboring blocks, mitigating performance drops caused by structural replacement. Experiments show that our FastFLUX maintains high image quality under both qualitative and quantitative evaluations, while significantly improving inference speed, even with 20% of the hierarchy pruned. Fuhan Cai, Jie Li 0022, Wenbo Li 0002, Jian Chen 0011, Xiangzhong Fang |
AAAI | 6 |
| 2026 | Reasoning with OmniThought: A Large CoT Dataset with Verbosity and Cognitive Difficulty AnnotationsabstractTasks such as mathematical problem solving and coding require models to leverage chain-ofthought (CoT) processes, enabling human-like reasoning strategies.However, the advancement of large reasoning models (LRMs) is hindered by the lack of comprehensive CoT datasets.Existing resources often fail to provide extensive reasoning problems with coherent CoT processes distilled from multiple teacher models, and do not account for multifaceted properties describing the internal characteristics of CoTs.To address these challenges, we introduce OmniThought, a largescale dataset featuring 2 million CoT processes generated and validated by multiple powerful LRMs.Each CoT process in OmniThought is annotated with novel Reasoning Verbosity (RV) and Cognitive Difficulty (CD) scores, which characterize the appropriateness of CoT verbosity and the cognitive difficulty level for models to comprehend these reasoning processes.We further establish a self-reliant pipeline to curate this dataset.Extensive experiments using Qwen2.5 and Qwen3 of various sizes demonstrate the positive impact of our RV and CD scores on LRM training effectiveness.Based on the OmniThought dataset, we train and release a series of high-performing LRMs with enhanced reasoning abilities and optimized CoT output length.Our contributions advance the development of LRMs across different scales for solving complex reasoning tasks. 1 Wenrui Cai 0001, Chengyu Wang 0001, Jun Huang 0007, Xiangzhong Fang |
ACL (1) | 5 |
| 2025 | Neural Block Compression: Variable Bitrates Feature Blocks for Texture RepresentationabstractThe imperative for compression of material textures emerges from the critical demand for high-quality rendering, which necessitates sophisticated textures that, in turn, require substantial storage and memory resources. Thus, low-bitrate compression is crucial, especially in modern games demanding higher texture resolutions. Concurrent methodologies in texture compression predominantly employ a block-based paradigm based on color space, which inevitably leads to representational redundancies and a limited compression scope, particularly at lower bitrates. In the context of mobile devices, bandwidth during texture loading and runtime memory are major bottlenecks, making existing compression algorithms inadequate for high-resolution textures. To mitigate these limitations, we propose a novel multi-resolution texture compression scheme, Neural Block Compression (NBC), developed within the neural feature domain. Our encoding scheme is constructed on a hierarchy of multi-resolution neural feature blocks, and the key ingredient is the variable bitrates quantization scheme. It allocates higher bitrates to higher feature mip-levels and lower bitrates to lower feature mip-levels, thereby extending the concept of block compression from color domain into neural feature domain. Extensive experiments demonstrate the superior texture compression quality achieved by the proposed scheme, especially at low bitrates. Yishun Dou, Xiangzhong Fang, Wenjun Zhang 0001, Bingbing Ni |
AAAI | 4 |
| 2025 | Task-Adaptive Channel Attention Graph Network for Few-Shot 3D Point Cloud Classification
Zhongqiang Zhang 0004, Jingren Xie, Xiangzhong Fang |
CGI (1) | 4 |
| 2025 | Enhancing Reasoning Abilities of Small LLMs with Cognitive AlignmentabstractThe reasoning capabilities of large reasoning models (LRMs), such as OpenAI's o1 and DeepSeek-R1, have seen substantial advancements through deep thinking.However, these enhancements come with significant resource demands, underscoring the need for training effective small reasoning models.A critical challenge is that small models possess different reasoning capacities and cognitive trajectories compared with their larger counterparts.Hence, directly distilling chain-of-thought (CoT) rationales from large LRMs to smaller ones can sometimes be ineffective and often requires a substantial amount of annotated data.In this paper, we first introduce a novel Critique-Rethink-Verify (CRV) system, designed for training smaller yet powerful LRMs.Our CRV system consists of multiple LLM agents, each specializing in unique tasks: (i) critiquing the CoT rationales according to the cognitive capabilities of smaller models, (ii) rethinking and refining these CoTs based on the critiques, and (iii) verifying the correctness of the refined results.Building on the CRV system, we further propose the Cognitive Preference Optimization (CogPO) algorithm to continuously enhance the reasoning abilities of smaller models by aligning their reasoning processes with their cognitive capacities.Comprehensive evaluations on challenging reasoning benchmarks demonstrate the efficacy of our CRV+CogPO framework, which outperforms other methods by a large margin. 1 Wenrui Cai 0001, Chengyu Wang 0001, Jun Huang 0007, Xiangzhong Fang |
EMNLP | 5 |
| 2025 | Linear Multistep Solver Distillation for Fast Sampling of Diffusion ModelsabstractSampling from diffusion models can be seen as solving the corresponding
probability flow ordinary differential equation (ODE).
The solving process requires a significant number of function
evaluations (NFE), making it time-consuming.
Recently, several solver search frameworks have attempted to find
better-performing model-specific solvers. However, predicting the impact of
intermediate solving strategies on final sample quality remains challenging,
rendering the search process inefficient.
In this paper, we propose a novel method for designing
solving strategies. We first introduce a unified prediction formula
for linear multistep solvers. Subsequently, we present a solver distillation
framework, which enables a student solver to mimic the sampling trajectory
generated by a teacher solver with more steps. We utilize the mean Euclidean
distance between the student and teacher sampling trajectories as a metric,
facilitating rapid adjustment and optimization of intermediate solving strategies.
The design space of our framework encompasses multiple aspects,
including prediction coefficients, time step schedules, and time scaling
factors.
Our framework has the ability to complete a solver search
for Stable-Diffusion in under 12 total GPU hours.
Compared to previous reinforcement learning-based
search frameworks,
our approach achieves over a 10$\times$ increase in search efficiency.
With just 5 NFE, we achieve FID scores of 3.23 on CIFAR10, 7.16 on ImageNet-64,
5.44 on LSUN-Bedroom, and 12.52 on MS-COCO, resulting in a 2$\times$ sampling acceleration ratio
compared to handcrafted solvers. Xiangzhong Fang, Hanting Chen, Yunhe Wang 0001 |
ICLR | 2 |
| 2025 | Generalized Category Discovery via Reciprocal Learning and Class-Wise Distribution RegularizationabstractGeneralized Category Discovery (GCD) aims to identify unlabeled samples by leveraging the base knowledge from labeled ones, where the unlabeled set consists of both base and novel classes.
Since clustering methods are time-consuming at inference, parametric-based approaches have become more popular.
However, recent parametric-based methods suffer from inferior base discrimination due to unreliable self-supervision.
To address this issue, we propose a Reciprocal Learning Framework (RLF) that introduces an auxiliary branch devoted to base classification.
During training, the main branch filters the pseudo-base samples to the auxiliary branch.
In response, the auxiliary branch provides more reliable soft labels for the main branch, leading to a virtuous cycle.
Furthermore, we introduce Class-wise Distribution Regularization (CDR) to mitigate the learning bias towards base classes.
CDR essentially increases the prediction confidence of the unlabeled data and boosts the novel class performance.
Combined with both components, our proposed method, RLCD, achieves superior performance in all classes with negligible extra computation.
Comprehensive experiments across seven GCD datasets validate its superiority.
Our codes are available at https://github.com/APORduo/RLCD. Zhiquan Tan, Linglan Zhao, Xiangzhong Fang, Weiran Huang 0001 |
ICML | 5 |
| 2025 | Few-Shot Class-Incremental Learning via Asymmetric Supervised Contrastive LearningabstractFew-Shot Class-Incremental Learning (FSCIL) is to continuously learn novel classes from a few samples without forgetting previous knowledge. Adapting directly to limited novel data typically results in significant forgetting of base class knowledge. Consequently, prevailing FSCIL methods are devoted to training a strong initial model that can be frozen in incremental sessions. However, these works face a dilemma in poor generalization: they benefit mainly from base class performance, yet underperform in novel classes. To alleviate this issue, we design a two-stage training framework to simultaneously enhance generalization for novel classes and maintain base class discrimination. In the first stage, an asymmetric supervised contrastive learning (AsyCon) algorithm is proposed. AsyCon introduces a predicted feature to achieve an asymmetric alignment of positive pairs. It alleviates over-similarity within positive features, allowing the model to better transfer to new classes in incremental sessions. In the second stage, the model is finetuned for promoting its performance on base classes. To maintain the generalization obtained in the first stage, we employ an L2 normalized regularization (LR) to keep the feature consistent with the model in the first stage. The finetuned model, termed AsyCLR, effectively balances generalization and discrimination, significantly outperforming existing FSCIL works especially in novel class accuracy. Experiments on CUB200, CIFAR100, and mini-ImageNet verify the effectiveness of our method. Additionally, our method also performs well in the standard few-shot recognition scenario due to its strong generalization ability. Our codes are available at https://github.com/APORduo/AsyCLR. Duo Liu 0001, Linglan Zhao, Fan Lyu, Xiangzhong Fang, Liang Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | NER-guided Comprehensive Hierarchy-aware Prompt Tuning for Hierarchical Text ClassificationabstractHierarchical text classification (HTC) is a significant but challenging task in natural language processing (NLP) due to its complex taxonomic label hierarchy. Recently, there have been a number of approaches that applied prompt learning to HTC problems, demonstrating impressive efficacy. The majority of prompt-based studies emphasize global hierarchical features by employing graph networks to represent the hierarchical structure as a whole, with limited research on maintaining path consistency within the internal hierarchy of the structure. In this paper, we formulate prompt-based HTC as a named entity recognition (NER) task and introduce conditional random fields (CRF) and Global Pointer to establish hierarchical dependencies. Specifically, we approach single- and multi-path HTC as flat and nested entity recognition tasks and model them using span- and token-based methods. By narrowing the gap between HTC and NER, we maintain the consistency of internal paths within the hierarchical structure through a simple and effective way. Extensive experiments on three public datasets show that our method achieves state-of-the-art (SoTA) performance. Fuhan Cai, Xiaozhe Yang, Xiangzhong Fang |
LREC/COLING | 6 |
| 2024 | Learning Quantized Adaptive Conditions for Diffusion Models
Yuchan Tian, Huaao Tang, Jie Hu 0021, Xiangzhong Fang, Hanting Chen |
ECCV (81) | 6 |
| 2024 | COPHTC: Contrastive Learning with Prompt Tuning for Hierarchical Text ClassificationabstractHierarchical Text Classification (HTC) is an essential yet challenging task in natural language processing (NLP) due to its complex label structure. Recently, a number of approaches have employed prompt learning in HTC, achieving noteworthy outcomes. However, prompt-based HTC does not further optimize the representation of samples based on label relationships, nor does it dynamically adjust the relative positions of sample pairs from the perspective of the embedding space. In this work, we propose a COntrastive-enhanced Promptbased model for HTC tasks (COPHTC). More specifically, we integrate contrastive learning into prompt tuning while employing momentum updates and a dynamic queue to offer a greater variety of positive and negative samples for the input texts. Extensive experiments on three public datasets verify the effectiveness of COPHTC. Fuhan Cai, Xiangzhong Fang |
ICASSP | 4 |
| 2024 | Distillation Excluding Positives for Few-Shot Class-Incremental LearningabstractFew-Shot Class-Incremental Learning (FSCIL) defines a challenging task to continually recognize novel classes with few training data without forgetting old classes. Considering the catastrophic forgetting and overfitting issues, mainstream FSCIL methods resort to obtaining a strong model in the base session and freezing it in incremental sessions. Although prevailing methods perform well in the base classes, they often struggle with poor novel class generalization. To strengthen the representation of these models, this paper focuses on Knowledge Distillation (KD). Since existing KD methods are incompetent for FSCIL and introduce limited improvement, we propose the Distillation Excluding Positives (DEP) method for boosting performance on FSCIL tasks. Specifically, DEP consists of negative relationship distillation and asymmetric self-feature distillation. It can mitigate the over-similarity of intra-class features, leading to a more generalized model. Extensive experiments on three FSCIL benchmarks validate the superiority of DEP over current SOTAs. Linglan Zhao, Fuhan Cai, Xiangzhong Fang |
ICME | 5 |
| 2024 | Mix background and foreground separately: Transformer-based Augmentation Strategies for Domain GeneralizationabstractDomain generalization (DG) aims to alleviate the severe performance degradation of deep neural networks when a domain shift exists between training and testing data. The different distributions of the backgrounds of samples represent a vital factor contributing to the domain gap. Inspired by the augmentation-based method, we aim to mix up the backgrounds of samples to generate new images for training. The Mixup method, a classical approach, randomly blends two samples at the image level. Therefore, it cannot resolve the causal relationship between backgrounds and semantic labels. To solve this problem, we introduce a new method that separates the foreground and background of samples at the patch level using a Vision Transformer network (ViT). Concretely, we calculate attention scores of each patch based on self-attention modules in ViT. Then, we identify the background or foreground by the rank of attention scores. Next, we present a Background-Mix method to blend the common background of samples from different domains, regularizing the model to ignore causal relations between background and semantic label. Moreover, the different appearances of objects across distinct source domains also contribute to the performance drop on the target domain. Therefore, we design a Foreground-Mix method to mix only the object part, excluding the background. This can enable the model to classify objects in multiple patterns of representation. The entire framework is collectively referred to as BFMix. Extensive experiments demonstrate that our method achieves state-of-the-art performance. Fuhan Cai, Xiangzhong Fang |
ICME | 5 |
| 2023 | Few-Shot Class-Incremental Learning via Class-Aware Bilateral DistillationabstractFew-Shot Class-Incremental Learning (FSCIL) aims to continually learn novel classes based on only few training samples, which poses a more challenging task than the well-studied Class-Incremental Learning (CIL) due to data scarcity. While knowledge distillation, a prevailing technique in CIL, can alleviate the catastrophic forgetting of older classes by regularizing outputs between current and previous model, it fails to consider the overfitting risk of novel classes in FSCIL. To adapt the powerful distillation technique for FSCIL, we propose a novel distillation structure, by taking the unique challenge of overfitting into account. Concretely, we draw knowledge from two complementary teachers. One is the model trained on abundant data from base classes that carries rich general knowledge, which can be leveraged for easing the overfitting of current novel classes. The other is the updated model from last incremental session that contains the adapted knowledge of previous novel classes, which is used for alleviating their forgetting. To combine the guidances, an adaptive strategy conditioned on the class-wise semantic similarities is introduced. Besides, for better preserving base class knowledge when accommodating novel concepts, we adopt a two-branch network with an attention-based aggregation module to dynamically merge predictions from two complementary branches. Extensive experiments on 3 popular FSCIL datasets: mini-ImageNet, CIFAR100 and CUB200 validate the effectiveness of our method by surpassing existing works by a significant margin. Code is available at https://github.com/LinglanZhao/BiDistFSCIL. Linglan Zhao, Jing Lu 0004, Yunlu Xu, Zhanzhan Cheng, Dashan Guo, Xiangzhong Fang |
CVPR | 7 |
| 2023 | Rethinking Self-Supervision for Few-Shot Class-Incremental LearningabstractFew-Shot Class-Incremental Learning (FSCIL) focuses on progressively absorbing new concepts given only limited training data. For tackling this challenge, several recent FS-CIL works resort to pre-training models with Self-Supervised Learning (SSL) to obtain features that can generalize well to new classes. However, to avoid overfitting and catastrophic forgetting, previous works only leverage SSL in the base session and keep all or most parameters fixed in incremental sessions, resulting in inadequate adaptation to novel classes. Thus, in this paper, we explore the setting where more parameters can be updated for adapting to novel concepts, and discover that the model pre-trained with SSL leads to degraded performance even compared to that without SSL. It can be attributed to the severer forgetting of base class knowledge. To address this issue, we propose an imprinting-based distillation module for effectively regularizing the adaption process, and a mathematically provable routing strategy for further improved results. The effectiveness of our approach is verified on 3 popular FSCIL benchmarks by significantly outperforming previous methods. Linglan Zhao, Jing Lu 0004, Zhanzhan Cheng, Xiangzhong Fang |
ICME | 5 |
| 2023 | Efficient Attention for Domain Generalization
Fuhan Cai, Xiangzhong Fang |
ICONIP (9) | 5 |
| 2023 | Boosting domain generalization by domain-aware knowledge distillation
Zhongqiang Zhang 0004, Fuhan Cai, Xiangzhong Fang |
Knowl. Based Syst. | 5 |
| 2022 | Im2Oil: Stroke-Based Oil Painting Rendering with Linearly Controllable Fineness Via Adaptive SamplingabstractThis paper proposes a novel stroke-based rendering (SBR) method that translates images into vivid oil paintings. Previous SBR techniques usually formulate the oil painting problem as pixel-wise approximation. Different from this technique route, we treat oil painting creation as an adaptive sampling problem. Firstly, we compute a probability density map based on the texture complexity of the input image. Then we use the Voronoi algorithm to sample a set of pixels as the stroke anchors. Next, we search and generate an individual oil stroke at each anchor. Finally, we place all the strokes on the canvas to obtain the oil painting. By adjusting the hyper-parameter maximum sampling probability, we can control the oil painting fineness in a linear manner. Comparison with existing state-of-the-art oil painting techniques shows that our results have higher fidelity and more realistic textures. A user opinion test demonstrates that people behave more preference toward our oil paintings than the results of other methods. More interesting results and the code are in https://github.com/TZYSJTU/Im2Oil. Zhengyan Tong, Xiaohang Wang 0004, Shengchao Yuan, Xuanhong Chen, Xiangzhong Fang |
ACM Multimedia | 6 |
| 2022 | Boosting Few-shot visual recognition via saliency-guided complementary attention
Linglan Zhao, Dashan Guo, Wei Li 0084, Xiangzhong Fang |
Neurocomputing | 5 |
| 2021 | A Strong Baseline for Semi-Supervised Incremental Few-Shot Learning
Linglan Zhao, Dashan Guo, Yunlu Xu, Liang Qiao 0001, Zhanzhan Cheng, Shiliang Pu, Xiangzhong Fang |
BMVC | 8 |
| 2021 | Saliency-Guided Complementary Attention for Improved Few-Shot LearningabstractDespite significant progress in recent deep neural networks, most deep learning algorithms rely heavily on abundant training samples. To address this problem, we propose an effective and interpretable few-shot classification model using Saliency-Guided Complementary Attention (SGCA), which aims to learn transferable representations and to build a robust classification module simultaneously. Concretely, we propose to train our feature extractor using an auxiliary task to separate object regions from background clutter guided by saliency detection signals. In addition, to make the separation beneficial to the downstream tasks, we introduce a complementary attention mechanism to force the classification module to focus on various informative parts of the image. Extensive experiments on few-shot learning tasks demonstrate the effectiveness of our proposed method, e.g., we achieve 68.81% and 84.60% for 5-way 1-shot and 5-shot settings on mini-ImageNet, respectively. Linglan Zhao, Dashan Guo, Wei Li 0084, Xiangzhong Fang |
ICME | 5 |
| 2021 | Class-wise Metric Scaling for Improved Few-Shot ClassificationabstractFew-shot classification aims to generalize basic knowledge to recognize novel categories from a few samples. Recent centroid-based methods achieve promising classification performance with the nearest neighbor rule. However, we consider that those methods intrinsically ignore per-class distribution, as the decision boundaries are biased due to the diversity of intra-class variances. Hence, we propose a class-wise metric scaling (CMS) mechanism, which can be applied to both training and testing stages. Concretely, metric scalars are set as learnable parameters in the training stage, helping to learn a more discriminative and transferable feature representation. As for testing, we construct a convex optimization problem to generate an optimal scalar vector for refining the nearest neighbor decisions. Besides, we also involve a low-ranking bilinear pooling layer for improved representation capacity, which further provides significant performance gains. Extensive experiments are conducted on a series of feature extractor backbones, datasets, and testing modes, which have shown consistent improvements compared to prior SOTA methods, e.g., we achieve accuracies of 66.64 % and 83.63 % for 5-way 1-shot and 5-shot settings on the mini-ImageNet, respectively. Under the semi-supervised inductive mode, results are further up to 78.34 % and 87.53 %, respectively. Linglan Zhao, Wei Li 0084, Dashan Guo, Xiangzhong Fang |
WACV | 5 |
| 2021 | PDA: Proxy-based domain adaptation for few-shot image recognition
Linglan Zhao, Xiangzhong Fang |
Image Vis. Comput. | 3 |
| 2020 | Moflowgan: Video Generation With Flow GuidanceabstractIn recent years, video generation has attracted a lot of attention in the computer vision community. Unlike image generation which only focuses on appearance, video generation requires modeling both content information and motion dynamics. In this work, we propose MoFlowGAN, which explicitly models motion dynamics by a content-motion decomposition architecture with an additional flow generator. The decomposition architecture models content and motion separately and is instantiated by a compact variant of BigGAN [1]. Besides, the flow generator generates optical flow directly based on high-level feature maps of adjacent frames as a strong supervision, hence the searching space of motion patterns is highly reduced. Our proposed MoFlowGAN achieves the state-of-the-art results on both MUG facial expression and UCF-101 datasets. Wei Li 0084, Zehuan Yuan, Xiangzhong Fang, Changhu Wang |
ICME | 3 |
| 2020 | Visual question answering with attention transfer and a cross-modal gating mechanism
Wei Li 0084, Jianhui Sun, Linglan Zhao, Xiangzhong Fang |
Pattern Recognit. Lett. | 5 |
| 2019 | Refining Proposals with Neighboring Contexts for Temporal Action DetectionabstractWhile many methods have been proposed for generating temporal proposals, the performance of existing temporal detection pipelines is still limited by the quality of proposals. In this paper, we introduce a new refining model for temporal action detection, which incorporates the evaluation of Intersection-over-Union (IoU) value into the action classification, and then regresses a better segment using the neighboring contexts of one candidate proposal. To refine one candidate proposal, we augment it with two neighboring proposals of equal length, which capture the contextual information from the past and future segments. After extracting regional features for the augmented proposals, we utilize the dilated convolutions for contextual modeling to regress the offsets between the candidate proposal and the target segment within this augmented area. Extensive experiments on THUMOS14 demonstrate that our method successfully refines the generated proposals and achieves superior detection performance over other methods. Dashan Guo, Wei Li 0084, Jianhui Sun, Xiangzhong Fang |
ICME | 5 |
| 2018 | Multimodal architecture for video captioning with memory networks and an attention mechanism
Wei Li 0084, Dashan Guo, Xiangzhong Fang |
Pattern Recognit. Lett. | 3 |
| 2018 | Fully Convolutional Network for Multiscale Temporal Action ProposalsabstractSimilar to the function of object proposals in localizing objects within images, temporal action proposals can facilitate the extraction of semantic segments and simplify the computations required for temporal action localization in untrimmed videos. In this paper, we propose a fully convolutional network to identify multiscale temporal action proposals (FCN-TAP) that utilizes only the temporal convolutions to retrieve accurate action proposals for video sequences. Using gated linear units, our network enables simple but powerful inferences, and by parallelizing the computations, it significantly improves performances compared with previous recurrent models. To capture more temporal contexts with fewer parameters, we apply dilated convolutions to expand the receptive fields of our network. Moreover, we divide the receptive fields into multiple scale ranges and then refine the corresponding temporal boundaries using duration regression at each scale. To generate suitable segments with arbitrary durations for training, we design a new strategy to select sampled candidates within the corresponding scale range. The power of our method is demonstrated through experiments on the THUMOS'14 and ActivityNet datasets, where FCN-TAP performs better and achieves a remarkable speedup compared to other state-of-the-art methods. Additional experiments show that our method generates high-quality proposals and improves the localization stage of existing action detection pipelines. Dashan Guo, Wei Li 0084, Xiangzhong Fang |
IEEE Trans. Multim. | 3 |
| 2018 | Exact Confidence Limits for the Acceleration Factor Under Constant-Stress Partially Accelerated Life Tests With Type-I CensoringabstractIn modern lifetime assessment and reliability analysis, accelerated life test has frequently been used to yield information quickly so that the life distribution of products can be estimated. This paper considers the estimate of the acceleration factor for the exponentially distributed lifetimes under the constant-stress partially accelerated life test with Type-I censored data. In many applications with small sample sizes, the approximate confidence limits for the parameters based on large-sample asymptotic distributions or the bootstrap method are usually not accurate enough. This study defines two ordering relations on the sample space by the generalized maximum likelihood estimator of the acceleration factor, and proposes one new approach of constructing the exact lower and upper confidence limits for the acceleration factor. An efficient procedure of computing the exact lower and upper confidence limits for the acceleration factor is presented via the EM algorithm. The approximate confidence limits for the acceleration factor using the asymptotic and bootstrap methods are also derived in this study. The proposed exact approach is compared with the two approximate methods by carrying out extensive simulation studies, and it is shown that the exact method performs well and is robust in small sample settings. Finally, we present a real example of the accelerated life test to illustrate all methods studied in this paper. Deqiang Zheng, Xiangzhong Fang |
IEEE Trans. Reliab. | 2 |
| 2017 | Lossless image compression algorithm and hardware architecture for bandwidth reduction of external memoryabstractIn high definition (HD) video coders, huge memory access bandwidth is the major throughput bottleneck. Lossless embedded compression is an efficient solution to alleviate the bandwidth burden, in which image are compressed before writing into local memory and decompressed after retrieving from local memory. This study proposes a hardware‐oriented lossless image compression algorithm, supporting block and line random access flexibly for adapting diverse hardware video codec architectures. The major contributions are characterised as follows. First, block or pixel‐level adaptive prediction is proposed to fully utilise the image spatial correlation by employing adaptive mode decision. Second, multiple‐range semi‐fixed (SF) variable length coding (VLC) is employed to describe the prediction residue, and adaptive block size selection is employed for SF VLC to fully utilise the statistical redundancy. In addition, Huffman VLC is further employed to represent the control syntax elements. Third, four‐stage pipeline hardware architecture is proposed to implement the proposed algorithm. Simulation results show that the proposed algorithm achieves competitive rate compression performance compared with reference algorithms. The proposed hardware architecture is verified supporting real‐time processing for quad‐HD videos at the frequency of 166 MHz. The proposed work achieves reducing memory access bandwidth by ∼55.2%, which is useful for hardwired video coding. Shizhong Li, Hai Bing Yin, Xiangzhong Fang, Huijuan Lu |
IET Image Process. | 3 |
| 2017 | Capturing Temporal Structures for Video Captioning by Spatio-temporal Contexts and Channel Attention Mechanism
Dashan Guo, Wei Li 0084, Xiangzhong Fang |
Neural Process. Lett. | 3 |
| 2013 | A Video Text Detection and Tracking SystemabstractFaced with the increasing large scale video databases, retrieving videos quickly and efficiently has become a crucial problem. Video text, which carries high level semantic information, is a type of important source that is useful for this task. In this paper, we introduce a video text detecting and tracking approach. By these methods we can obtain clear binary text images, and these text images can be processed by OCR (Optical Character Recognition) software directly. Our approach including two parts, one is stroke-model based video text detection and localization method, the other is SURF (Speeded Up Robust Features) based text region tracking method. In our detection and localization approach, we use stroke model and morphological operation to roughly identify candidate text regions. Combine stroke-map and edge response to localize text lines in each candidate text regions. Several heuristics and SVM (Support Vector Machine) used to verifying text blocks. The core part of our text tracking method is fast approximate nearest-neighbour search algorithm for extracted SURF features. Text-ending frame is determined based on SURF feature point numbers, while, text motion estimation is based on correct matches in adjacent frames. Experimental result on large number of different video clips shows that our approach can effectively detect and track both static texts and scrolling texts. Tuoerhongjiang Yusufu, Xiangzhong Fang |
ISM | 3 |
| 2012 | Frame Rate Up-Conversion for Depth-Based 3D VideoabstractA novel frame rate up-conversion (FRUC) scheme for depth-based 3D video is proposed. Differing from the existing conventional FRUC methods which are designed for two dimensional (2D) video only, the proposed method is designed for depth-based 3D video, which increases the frame rate of both color sequences and its associated depth maps by jointly considering the image intensity, depth information, and spatial-temporal correlations. The proposed method contains motion estimation, block irregular segmentation (BIS), depth-constrained motion vector post-processing, forward and backward motion compensation, and edge-preserved combination. Experimental results show that the intermediate color and depth images interpolated by our proposed method provide a good image quality both objectively and subjectively. Moreover, compared with conversional FRUC methods, the color and depth map sequences up-converted by the proposed method are more suitable for further virtual view synthesis. Qingchun Lu, Xiangzhong Fang, Yongzhe Wang |
ICME | 2 |
| 2011 | Integrating Visual Saliency and Consistency for Re-Ranking Image Search ResultsabstractIn this paper, we propose a new algorithm for image re-ranking in web image search applications. The proposed method focuses on investigating the following two mechanisms: 1) Visual consistency. In most web image search cases, the images that closely related to the search query are visually similar. These visually consistent images which occur most frequently in the first few web pages will be given higher ranks. 2) Visual saliency. From visual aspect, it is obvious that salient images would be easier to catch users' eyes, and it is observed that these visually salient images in the front pages are often relevant to the user's query. By integrating the above two mechanisms, our method can efficiently re-rank the images from search engines and obtain a more satisfactory search result. Experimental results on a real-world web image dataset demonstrate that our approach can effectively improve the performance of image retrieval. Xiaokang Yang 0001, Xiangzhong Fang, Weisi Lin, Rui Zhang 0052 |
IEEE Trans. Multim. | 3 |
| 2010 | Integrating visual saliency and consistency for re-ranking image search resultsabstractThe paper investigates two mechanisms, visual consistency and visual saliency, in web image search: (1) In most current web image search engines, such as Google Image Search and Yahoo Image Search, the images that closely related to the search query are typically visually similar. These visually consistent images which occur most frequently in the first few web pages will be given higher ranks. (2) From visual aspect, it is obvious that salient images would be easier to catch users' eyes and more likely to be clicked than the cluttered ones in low-level vision. In addition, we also observe the fact that the visually salient images in the front pages are often relevant to the user's query. The principal novelty of this paper is in combining visual saliency and consistency to re-rank the results from search engines to make the re-ranked images more satisfying in both vision and content. The experimental results on a real world web image dataset demonstrate that our approach can effectively improve the performance of image retrieval. Xiaokang Yang 0001, Rui Zhang 0052, Fuxiang Lu, Xiangzhong Fang |
ICIP | 5 |
| 2010 | Adaptively Adjusted Gaussian Mixture Models for Surveillance Applications
Tianci Huang, Xiangzhong Fang, Jingbang Qiu, Takeshi Ikenaga |
MMM | 2 |
| 2010 | Content-adaptive spatial scalability for scalable video codingabstractThis paper presents an enhancement of the SVC extension of the H.264/AVC standard by content-adaptive spatial scalability (CASS). CASS introduces a novel functionality which is important for high quality content distribution. The video streams (spatial layers), which are used as input to the encoder, are created by content-adaptive and art-directable retargeting of existing high resolution video. Video is retargeted to resolutions and aspect ratios which are mainly dictated by target display devices. Thereby no content is cut off, but visually important content is preserved at the expense of a non-linear distortion of visually unimportant areas. The non-linear dependencies between such video streams are efficiently exploited by CASS for scalable coding. This is achieved by integrating warping-based non-linear texture prediction and warp coding into the SVC framework. The results indicate high prediction accuracy of non-linear predictors and high compression efficiency with limited increase in bit rate and complexity compared to the standard SVC for the case of INTRA only coding. Yongzhe Wang, Nikolce Stefanoski, Xiangzhong Fang, Aljoscha Smolic |
PCS | 3 |
| 2009 | Parallel HD Encoding on CELLabstractThe Cell Broadband Engine Architecture (CBEA) is an excellent architecture for high performance distributed computing and multimedia processing. While the Cell/BE processor is capable of high definition H.264 encoding, there are still no such implementations available. In this paper, we present a parallel implementation of a HD H.264 encoder on this heterogeneous nine cores processor. First we implement a real time SD encoder on a single SPU by optimizing Motion Estimation algorithm, DMA transfers etc. Then we propose a pipelined parallel encoding algorithm for multicore processors, and use this algorithm to get a real time HD H.264 encoder (1920×1080@31fps) by using eight SPEs (58fps on 16 SPEs). Xiangzhong Fang, Ci Wang, Satoshi Goto |
ISCAS | 2 |
| 2008 | Multilevel Framework to Detect and Handle Vehicle OcclusionabstractThis paper presents a multilevel framework to detect and handle vehicle occlusion. The proposed framework consists of the intraframe, interframe, and tracking levels. On the intraframe level, occlusion is detected by evaluating thecompactness ratioandinterior distance ratioof vehicles, and the detected occlusion is handled by removing a “cutting region” of the occluded vehicles. On the interframe level, occlusion is detected by performing subtractive clustering on the motion vectors of vehicles, and the occluded vehicles are separated according to the binary classification of motion vectors. On the tracking level, occlusion layer images are adaptively constructed and maintained, and the detected vehicles are tracked in both the captured images and the occlusion layer images by performing a bidirectional occlusion reasoning algorithm. The proposed intraframe, interframe, and tracking levels are sequentially implemented in our framework. Experiments on various typical scenes exhibit the effectiveness of the proposed framework. Quantitative evaluation and comparison demonstrate that the proposed method outperforms state-of-the-art methods. Wei Zhang 0025, Q. M. Jonathan Wu, Xiaokang Yang 0001, Xiangzhong Fang |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2007 | Multiframe Super-Resolution Reconstruction Based on Cycle-SpinningabstractA multiframe super-resolution (SR) reconstruction algorithm based on cycle-spinning (CS) is proposed. We utilize the relative motion information of sequential images to construct a CS-based framework for the resolution enhancement. The unique feature of the proposed algorithm is that it is effective for low-resolution (LR) images with various point spread function (PSF) and noise characteristics, even if the degradation models are unknown for the imaging system. Moreover, the computational complexity is inexpensive. Experiments demonstrate the effectiveness of the proposed method and show the superiority to previous methods in objective and subjective qualities. Xiangzhong Fang, Songyu Yu |
ICASSP (1) | 2 |
| 2007 | Spread-Spectrum Audio Watermark Robust Against Pitch-Scale ModificationabstractA new digital audio watermarking for copyright protection is described in this paper. The watermark is embedded in the time domain by taking use of spread-spectrum technique and exploiting the hearing thresholds in the psychoacoustic model. The scheme of detection is based on high-pass filtering and correlation. The decision threshold is adaptively determined according to the statistical distribution of correlation values. Besides, the postprocessing technique is introduced for resisting time-scale modification. Experimental results show that the proposed watermarking scheme is high robust to almost any attacks, including cropping, low-pass filtering, noise addition, MP3 compression, sampling rate changing, requantization, time-scale modification, and pitch-scale modification. Jianling Hu, Xiangzhong Fang |
ICME | 3 |
| 2007 | Moving Cast Shadows Detection Using Ratio EdgeabstractMoving objects segmentation plays a very important role in real-time image analysis. However, as one of the common parts in the natural scenes, shadows severely interfere with the accuracy of moving objects detection in video surveillance. In this paper, we present a novel method for moving cast shadows detection. Based on the analysis of the physical model of moving shadows, we prove that the ratio edge is illumination invariant. The distribution of the ratio edge is discussed and a significance test is performed to classify each moving pixel into foreground object or moving shadow. Intensity constraint and geometric heuristics are imposed to further improve the performance. Experiments on various typical scenes exhibit the robustness of the proposed method. Extensively quantitative evaluation and comparison demonstrate that the proposed method significantly outperforms state-of-the-art methods. Wei Zhang 0025, Xiangzhong Fang, Xiaokang Yang 0001, Q. M. Jonathan Wu |
IEEE Trans. Multim. | 2 |
| 2006 | Motion Vector Smoothing for True Motion EstimationabstractThis paper proposes a new motion vector (MV) smoothing algorithm to track the real motion in image sequences for MPEG video encoders. First, a pre-checking algorithm is employed to eliminate wrong motion vectors and preserve all possible motion vectors. For each block considered, the motion similarity between the neighboring blocks and the number of candidate motion vectors are jointly exploited to adaptively grow the filtering support, which is supposed to have homogeneous motion and sufficient spatial gradient. Then, all candidate motion vectors are checked within the filtering support using a new motion smoothness-constrained matching criteria. The simulation results show that the proposed algorithm can efficiently track the real motion resulting in smooth motion vector field (MVF). Hai Bing Yin, Xiangzhong Fang, Hua Yang 0001, Songyu Yu, Xiaokang Yang 0001 |
ICASSP (2) | 2 |
| 2006 | Multi-Rate, Dynamic and Compliant Region of Interest Coding for JPEG2000abstractA method is proposed to encode multiple regions of interest (ROI) in JPEG2000 image. It rearranges truncation point for every codeblock in each layer. It assigns higher bitrate to ROI and lower bitrate to non-ROI and combines them to codestream. The proposed strategy produces a fully compliant JPEG2000 codestream. It allows transmission of different ROIs with different priorities and supports dynamic delineation and prioritization of them. Experimental results demonstrating the validity of the proposed approach are presented Xiangzhong Fang, Haibin Yin, Songyu Yu |
ICME | 2 |
| 2006 | Moving vehicles segmentation based on Bayesian framework for Gaussian motion model
Wei Zhang 0025, Xiangzhong Fang, Xiaokang Yang 0001 |
Pattern Recognit. Lett. | 2 |
| 2005 | Rate Control for Motion JPEG2000 Using Correlation PredictionabstractMotion JPEG2000 has been widely used for its fine scalability with bit stream and its efficiency. A single scalable bitstream can provide precise rate control for variable bitrate (VBR) traffic. This paper presents an algorithm that makes use of correlation among the video frames to get consistent quality. By using one frame's distortion-rate to predict the next several frames' distortion-rate values, according to correlation among frames, we use only a few buffers to achieve highly consistent quality which used to use many buffers to achieve. Experimental results using motion JPEG2000 demonstrate substantial benefits. Xiangzhong Fang, Rong Hou |
ICASSP (2) | 2 |
| 2005 | Joint Rate Control for Multiple Sequences coding based on H.264 standardabstractThe objective of joint rate control is to dynamically distribute the channel capacity among video sequences according to their respective complexities, thus a more uniform picture quality and a more efficient utilization of channel capacity are achieved. Most existing approaches are based on MPEG2 coding platform. This paper presents a novel joint rate control scheme for multiple video sequences coding based on H.264 standard. A novel complexity measure that adapts to the characteristics of H.264 video coding is proposed. Experimental results show that the proposed scheme maintains a good balance in picture quality among the sequences as well as within a sequence Xiangzhong Fang, Hongkai Xiong |
ICME | 2 |
| 2005 | A new rate control algorithm for macroblock-level codersabstractRate control plays an important role in video coding. It regulates the coded bits to satisfy the channel rate while keeping good video quality. In Tsai's paper, a new algorithm was proposed which rearranging the macroblocks' coding order according to their significance in each frame. More complex macroblocks will be coded with more priority in coding order. But it is required to rearrange the coding order and output the encoded bits in the original scan order. In this paper, a new rate control algorithm is proposed which recalculates the mquant of each macroblock inside a frame according to its significance. Furthermore, it can be applied to all macroblock-level coders. Simulation results show that this new algorithm can achieve obvious improvement in video quality like Tsai's algorithm while has lower complexity Yutao Dong, Xiangzhong Fang, Hao Liu 0010 |
MMSP | 2 |
| 2005 | Unequal Forced Intra-Refresh for Real-time Multicast VideoabstractIn motion-compensated video coding, the errors caused by packet loss not only impair the reconstruction quality of current frame, but also lead to error propagation to subsequent frames. Based on the error-propagation analysis in a group of pictures (GOP), we propose an unequal forced intra-refresh scheme to increase error resilience of multicast video. According to a GOP-level error-propagation model, the proposed scheme can distribute the unequal number of forced intra-mode MBs to different P-frames of a GOP. Experimental results show that the proposed scheme can effectively mitigate the error-propagation effect and achieve about 0.1~1.1 dB gains over the traditional average scheme in H.264/AVC Hao Liu 0010, Wenjun Zhang 0001, Yutao Dong, Xiangzhong Fang |
MMSP | 4 |