EDBT 2026 Demo / reviewers in the wild / expert
Qingyi Gu
dblp:86/8369
· DBLP profile ↗
50ranked-venue papers
6as first author
31since 2021 · last 2026
0000-0001-8332-5350ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 34 · 2 first-author · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 23 · 2 first-author · 17 since 2021Systems, architecture and hardware · 14 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MambaSeg: Harnessing Mamba for Accurate and Efficient Image-Event Semantic SegmentationabstractSemantic segmentation is a fundamental task in computer vision with wide-ranging applications, including autonomous driving and robotics. While RGB-based methods have achieved strong performance with CNNs and Transformers, their effectiveness degrades under fast motion, low-light, or high dynamic range conditions due to limitations of frame cameras. Event cameras offer complementary advantages such as high temporal resolution and low latency, yet lack color and texture, making them insufficient on their own. To address this, recent research has explored multimodal fusion of RGB and event data; however, many existing approaches are computationally expensive and focus primarily on spatial fusion, neglecting the temporal dynamics inherent in event streams. In this work, we propose MambaSeg, a novel dual-branch semantic segmentation framework that employs parallel Mamba encoders to efficiently model RGB images and event streams. To reduce cross-modal ambiguity, we introduce the Dual-Dimensional Interaction Module (DDIM), comprising a Cross-Spatial Interaction Module (CSIM) and a Cross-Temporal Interaction Module (CTIM), which jointly perform fine-grained fusion along both spatial and temporal dimensions. This design improves cross-modal alignment, reduces ambiguity, and leverages the complementary properties of each modality. Extensive experiments on the DDD17 and DSEC datasets demonstrate that MambaSeg achieves state-of-the-art segmentation performance while significantly reducing computational cost, showcasing its promise for efficient, scalable, and robust multimodal perception. Fuqiang Gu, Yuanke Li, Xianlei Long, Kangping Ji, Chao Chen 0004, Qingyi Gu, Zhen-Liang Ni |
AAAI | 6 |
| 2026 | SAQ-SAM: Semantically-Aligned Quantization for Segment Anything ModelabstractSegment Anything Model (SAM) exhibits remarkable zero-shot segmentation capability; however, its prohibitive computational costs make edge deployment challenging. Although post-training quantization (PTQ) offers a promising compression solution, existing methods yield unsatisfactory results when applied to SAM, owing to its specialized model components and promptable workflow: (i) The mask decoder's attention exhibits extreme activation outliers, and we find that aggressive clipping (even 100x), without smoothing or isolation, is effective in suppressing outliers while maintaining performance. Unfortunately, traditional distribution-based metrics (e.g., MSE) fail to provide such large-scale clipping. (ii) Existing quantization reconstruction methods neglect semantic interactivity of SAM, leading to misalignment between image feature and prompt intention. To address the above issues, we propose SAQ-SAM in this paper, which boosts PTQ for SAM from the perspective of semantic alignment. Specifically, we propose Perceptual-Consistency Clipping, which exploits attention focus overlap to promote aggressive clipping while preserving semantic capabilities. Furthermore, we propose Prompt-Aware Reconstruction, which incorporates image-prompt interactions by leveraging cross-attention in mask decoder, thus facilitating alignment in both distribution and semantic. Moreover, to ensure the interaction efficiency, we design a layer-skipping strategy for image tokens in encoder. Extensive experiments are conducted on various SAM sizes and tasks, including instance segmentation, oriented object detection, and semantic segmentation, and the results show that our method consistently exhibits advantages. For example, when quantizing SAM-B to 4-bit, SAQ-SAM achieves 11.7% higher mAP than the baseline in instance segmentation task. Zhikai Li, Chengzhi Hu, Qingyi Gu |
AAAI | 5 |
| 2026 | DapQ-DiT: Distribution-Aware Post-Training Quantization for Efficient Generative Tasks in Diffusion TransformersabstractDiffusion Transformers (DiTs) have demonstrated remarkable performance in image and video generation tasks. However, their high computational and memory overheads severely restrict their practical deployment on resource-constrained devices. Post-training quantization (PTQ), an efficient and practical model compression technique, serves as a solution to alleviate this issue. Nevertheless, existing PTQ methods tailored for DiTs suffer from significant performance degradation when conducting low-bit weight–activation quantization. In this work, we identify two key factors responsible for such performance degradation. First, the weights of DiTs exhibit Gaussian-like distributions, which makes uniform quantization poorly matched to the actual weight density and introduces large quantization errors. Second, activation outliers with extremely large magnitudes, especially in specific linear layers, significantly widen the value range and severely reduce the suppression effectiveness of fixed Hadamard rotation. To address the above degradation issues, we propose DapQ-DiT, a novel distribution-aware post-training quantization framework tailored for DiTs. First, we introduce an arctan quantizer that explicitly adapts to Gaussian-like weight distributions, concentrating more quantization intervals in the high-density central region while preserving representation accuracy for critical weights. Second, we enhance the fixed Hadamard rotation by leveraging principal components derived from activation covariance, which allows the transformation to better align with real activation distributions and more effectively suppress diverse and extreme activation outliers. Extensive experiments conducted on text-to-image generation with PixArt and text-to-video generation with OpenSORA demonstrate that DapQ-DiT consistently outperforms existing PTQ methods across various prompt sets and diverse bit-width configurations. Lianwei Yang, Haokun Lin, Zhenan Sun, Qingyi Gu |
ICMR | 5 |
| 2026 | Privacy-Preserving SAM Quantization for Efficient Edge Intelligence in Healthcare
Zhikai Li, Qingyi Gu |
Int. J. Comput. Vis. | 3 |
| 2026 | Reshape and rotate: Adaptive weight reshaping and fine-grained rotation for ultra-low-bit diffusion transformers quantization
Lianwei Yang, Haokun Lin, Caifeng Shan, Zhenan Sun, Qingyi Gu |
Neurocomputing | 6 |
| 2026 | EDA-DM: Enhanced Distribution Alignment for Post-Training Quantization of Diffusion ModelsabstractDiffusion models have achieved great success in image generation tasks. However, the lengthy denoising process and complex neural networks hinder their low-latency applications in real-world scenarios. Quantization can effectively reduce model complexity, and post-training quantization (PTQ), which does not require fine-tuning, is highly promising for compressing and accelerating diffusion models. Unfortunately, we find that due to the highly dynamic activations, existing PTQ methods suffer from distribution mismatch issues at both calibration sample level and reconstruction output level, which makes the performance far from satisfactory. In this paper, we propose EDA-DM, a standardized PTQ method that efficiently addresses the above issues. Specifically, at the calibration sample level, we extract information from the density and diversity of latent space feature maps, which guides the selection of calibration samples to align with the overall sample distribution; and at the reconstruction output level, we theoretically analyze the reasons for previous reconstruction failures and, based on this insight, optimize block reconstruction using the Hessian loss of layers, aligning the outputs of quantized model and full-precision model at different network granularity. Extensive experiments demonstrate that EDA-DM significantly outperforms the existing PTQ methods across various models and datasets. Our method achieves a $1.83\times $ speedup and $4\times $ compression for the popular Stable-Diffusion on MS-COCO, with only a 0.05 loss in CLIP score. Code is available at http://github.com/BienLuky/EDA-DM. Zhikai Li, Junrui Xiao, Mengjuan Chen, Qingyi Gu |
IEEE Trans. Image Process. | 6 |
| 2026 | CoLeQ: Improving Data-Free Quantization via Contrastive LearningabstractModel quantization is an effective approach to reduce the complexity of neural networks, enabling them to be deployed on resource-constrained edge devices. Recently, data-free quantization has been widely investigated, since it does not access the original datasets and can address the widely-held data privacy and security concerns. Its idea is to generate fake data depending on the prior information in the full-precision (FP) model, and then fine-tune the quantized model with them under the supervision of the FP model. The quantization performance relies heavily on the validity of the generated data, however, existing methods suffer from two severe issues: mode collapse and (catastrophic) example forgetting, leading to non-trivial accuracy degradation. In this work, we proposeContrastiveLearningQuantization (CoLeQ), which achieves data diversity enhancement and old knowledge restoration via contrastive learning to address the above issues. Specifically, we introduce the MoCo paradigm that maintains a dynamic momentum queue of the encoded features to data-free quantization. The contrastive learning objective is used to improve data diversity by facilitating the separation of generated samples from the already generated ones in previous mini-batches, thus mitigating the mode collapse problem. Moreover, we design a tied-weight decoder to restore the previous samples from the encoded features in the queue without additional parameters and training, hence cost-effectively preventing the example forgetting problem. Extensive experiments are conducted to evaluate the effectiveness of CoLeQ, and the results demonstrate a consistent superiority compared to state-of-the-art methods. Zhikai Li, Mengjuan Chen, Junrui Xiao, Qingyi Gu |
IEEE Trans. Multim. | 4 |
| 2025 | K-Sort Arena: Efficient and Reliable Benchmarking for Generative Models via K-wise Human PreferencesabstractThe rapid advancement of visual generative models necessitates efficient and reliable evaluation methods. Arena platform, which gathers user votes on model comparisons, can rank models with human preferences. However, traditional Arena methods, while established, require an excessive number of comparisons for ranking to converge and are vulnerable to preference noise in voting, suggesting the need for better approaches tailored to contemporary evaluation challenges. In this paper, we introduce K-Sort Arena, an efficient and reliable platform based on a key insight: images and videos possess higher perceptual intuitiveness than texts, enabling rapid evaluation of multiple samples simultaneously. Consequently, K-Sort Arena employs K-wise comparisons, allowing K models to engage in free-forall competitions, which yield much richer information than pairwise comparisons. To enhance the robustness of the system, we leverage probabilistic modeling and Bayesian updating techniques. We propose an exploration-exploitation-based matchmaking strategy to facilitate more informative comparisons. In our experiments, K-Sort Arena exhibits 16.3× faster convergence compared to the widely used Elo algorithm. To further validate the superiority and obtain a comprehensive leaderboard, we collect human feedback via crowdsourced evaluations of numerous cutting-edge text-to-image and text-to-video models. Thanks to its high efficiency, K-Sort Arena can continuously incorporate emerging models and update the leaderboard with minimal votes. Our project has undergone several months of internal testing and is now available online. Zhikai Li, Dongrong Fu, Qingyi Gu, Kurt Keutzer, Zhen Dong 0003 |
CVPR | 5 |
| 2025 | CacheQuant: Comprehensively Accelerated Diffusion ModelsabstractDiffusion models have gradually gained prominence in the field of image synthesis, showcasing remarkable generative capabilities. Nevertheless, the slow inference and complex networks, resulting from redundancy at both temporal and structural levels, hinder their low-latency applications in real-world scenarios. Current acceleration methods for diffusion models focus separately on temporal and structural levels. However, independent optimization at each level to further push the acceleration limits results in significant performance degradation. On the other hand, integrating optimizations at both levels can compound the acceleration effects. Unfortunately, we find that the optimizations at these two levels are not entirely orthogonal. Performing separate optimizations and then simply integrating them results in unsatisfactory performance. To tackle this issue, we propose CacheQuant, a novel training-free paradigm that comprehensively accelerates diffusion models by jointly optimizing model caching and quantization techniques. Specifically, we employ a dynamic programming approach to determine the optimal cache schedule, in which the properties of caching and quantization are carefully considered to minimize errors. Additionally, we propose decoupled error correction to further mitigate the coupled and accumulated errors step by step. Experimental results show that CacheQuant achieves a 5.18× speedup and 4× compression for Stable Diffusion on MS-COCO, with only a 0.02 loss in CLIP score. Our code are open-sourced. Zhikai Li, Qingyi Gu |
CVPR | 3 |
| 2025 | SLTNet: Efficient Event-based Semantic Segmentation with Spike-driven Lightweight Transformer-based NetworksabstractEvent-based semantic segmentation has great potential in autonomous driving and robotics due to the advantages of event cameras, such as high dynamic range, low latency, and low power cost. Unfortunately, current artificial neural network (ANN)-based segmentation methods suffer from high computational demands, the requirements for image frames, and massive energy consumption, limiting their efficiency and application on resource-constrained edge/mobile platforms. To address these problems, we introduce SLTNet, a Spike-driven Lightweight Transformer-based Network designed for event-based semantic segmentation. Specifically, SLTNet is built on efficient spike-driven convolution blocks (SCBs) to extract rich semantic features while reducing the model’s parameters. Then, to enhance the long-range contextual feature interaction, we propose novel spike-driven transformer blocks (STBs) with binary mask operations. Based on these basic blocks, SLTNet employs a high-efficiency single-branch architecture while maintaining the low energy consumption of the Spiking Neural Network (SNN). Finally, extensive experiments on DDD17 and DSEC-Semantic datasets demonstrate that SLTNet outperforms state-of-the-art (SOTA) SNN-based methods by at most 9.06% and 9.39% mIoU, respectively, with extremely 4.58× lower energy consumption and 114 FPS inference speed. Our code is open-sourced and available at https://github.com/longxianlei/SLTNet-v1.0. Xianlei Long, Xiaxin Zhu, Fangming Guo, Wanyi Zhang, Qingyi Gu, Chao Chen 0004, Fuqiang Gu |
IROS | 5 |
| 2025 | DilateQuant: Accurate and Efficient Quantization-Aware Training for Diffusion Models via Weight DilationabstractModel quantization is a promising method for accelerating and compressing diffusion models. Nevertheless, since post-training quantization (PTQ) fails catastrophically at low-bit cases, quantization-aware training (QAT) is essential. Unfortunately, the wide range and time-varying activations in diffusion models sharply increase the complexity of quantization, making existing QAT methods inefficient. Equivalent scaling can effectively reduce activation range, but previous methods remain the overall quantization error unchanged. More critically, these methods significantly disrupt the original weight distribution, resulting in poor weight initialization and challenging convergence during QAT training. In this paper, we propose a novel QAT framework for diffusion models, called DilateQuant. Specifically, we propose Weight Dilation (WD) that maximally dilates the unsaturated in-channel weights to a constrained range through equivalent scaling. WD decreases the activation range while preserving the original weight range, which steadily reduces the quantization error and ensures model convergence. To further enhance accuracy and efficiency, we design a Temporal Parallel Quantizer (TPQ) to address the time-varying activations and introduce a Block-wise Knowledge Distillation (BKD) to reduce resource consumption in training. Extensive experiments demonstrate that DilateQuant significantly outperforms existing methods in terms of accuracy and efficiency. Zhikai Li, Mengjuan Chen, Qingyi Gu |
ACM Multimedia | 6 |
| 2025 | BinaryViT: Toward Efficient and Accurate Binary Vision TransformersabstractVision Transformers (ViTs) have emerged as the new fundamental architecture for most computer vision fields. However, the considerable memory and computation costs also hinder their application on resource-limited devices. Currently, binarization has demonstrated remarkable potential as a model compression technique in traditional Convolutional Neural Networks (CNNs), albeit with some accuracy loss. In this paper, we focus on binarization of ViTs, which is still under-studied and suffering a significant performance drop. We start with constructing a strong baseline of binary ViTs, integrating some of the best practices from binary CNNs, which forms the foundation of our exploration. Subsequently, we identify that the severe performance degradation of the baseline is mainly caused by the weight oscillation around the quantization boundary and the information distortion in the activation of ViTs. To address these challenges, we introduce BinaryViT, a precise full binarization framework tailored for Vision Transformers (ViTs), effectively pushing the binarization of ViTs to its limit. Specifically, we propose a novel gradient regularization scheme (GRS), which mitigates oscillations by fostering a smooth moving of latent weights to be away from the quantization boundary during the training process. Additionally, we have devised an Activation Shift Module (ASM) that dynamically adjusts the activation distribution prior to the sign function, thereby minimizing the information distortion stemming from the significant inter-channel variations. Extensive experiments on ImageNet dataset show that our BinaryViT consistently surpasses the strong baseline by 2.05% and improves the accuracy of fully binarized ViTs to a usable level. Furthermore, our method achieves impressive savings of$16.2\times $and$17.7\times $in model size and OPs compared to the full-precision DeiT-S. Junrui Xiao, Zhikai Li, Lianwei Yang, Qingyi Gu |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | MGRQ: Post-Training Quantization For Vision Transformer With Mixed Granularity ReconstructionabstractPost-training quantization (PTQ) efficiently compresses vision models, but unfortunately, it accompanies a certain degree of accuracy degradation. Reconstruction methods aim to enhance model performance by narrowing the gap between the quantized model and the full-precision model, often yielding promising results. However, efforts to significantly improve the performance of PTQ through reconstruction in the Vision Transformer (ViT) have shown limited efficacy. In this paper, we conduct a thorough analysis of the reasons for this limited effectiveness and propose MGRQ (Mixed Granularity Reconstruction Quantization) as a solution to address this issue. Unlike previous reconstruction schemes, MGRQ introduces a mixed granularity reconstruction approach. Specifically, MGRQ enhances the performance of PTQ by introducing Extra-Block Global Supervision and Intra-Block Local Supervision, building upon Optimized Block-wise Reconstruction. Extra-Block Global Supervision considers the relationship between block outputs and the model’s output, aiding block-wise reconstruction through global supervision. Meanwhile, Intra-Block Local Supervision reduces generalization errors by aligning the distribution of outputs at each layer within a block. Subsequently, MGRQ is further optimized for reconstruction through Mixed Granularity Loss Fusion. Extensive experiments conducted on various ViT models illustrate the effectiveness of MGRQ. Notably, MGRQ demonstrates robust performance in low-bit quantization, thereby enhancing the practicality of the quantized model. Lianwei Yang, Zhikai Li, Junrui Xiao, Haisong Gong, Qingyi Gu |
ICIP | 5 |
| 2024 | A Novel Wide-Area Multiobject Detection System with High-Probability Region SearchingabstractIn recent years, wide-area visual surveillance systems have been widely applied in various industrial and transportation scenarios. These systems, however, face significant challenges when implementing multi-object detection due to conflicts arising from the need for high-resolution imaging, efficient object searching, and accurate localization. To address these challenges, this paper presents a hybrid system that incorporates a wide-angle camera, a high-speed search camera, and a galvano-mirror. In this system, the wide-angle camera offers panoramic images as prior information, which helps the search camera capture detailed images of the targeted objects. This integrated approach enhances the overall efficiency and effectiveness of wide-area visual detection systems. Specifically, in this study, we introduce a wide-angle camera-based method to generate a panoramic probability map (PPM) for estimating high-probability regions of target object presence. Then, we propose a probability searching module that uses the PPM-generated prior information to dynamically adjust the sampling range and refine target coordinates based on uncertainty variance computed by the object detector. Finally, the integration of PPM and the probability searching module yields an efficient hybrid vision system capable of achieving 120 fps multi-object search and detection. Extensive experiments are conducted to verify the system’s effectiveness and robustness. Xianlei Long, Chao Chen 0004, Fuqiang Gu, Qingyi Gu |
ICRA | 5 |
| 2024 | HTQ: Exploring the High-Dimensional Trade-Off of mixed-precision quantization
Zhikai Li, Xianlei Long, Junrui Xiao, Qingyi Gu |
Pattern Recognit. | 4 |
| 2024 | PSAQ-ViT V2: Toward Accurate and General Data-Free Quantization for Vision TransformersabstractData-free quantization can potentially address data privacy and security concerns in model compression and thus has been widely investigated. Recently, patch similarity aware data-free quantization for vision transformers (PSAQ-ViT) designs a relative value metric, patch similarity, to generate data from pretrained vision transformers (ViTs), achieving the first attempt at data-free quantization for ViTs. In this article, we propose PSAQ-ViT V2, a more accurate and general data-free quantization framework for ViTs, built on top of PSAQ-ViT. More specifically, following the patch similarity metric in PSAQ-ViT, we introduce an adaptive teacher-student strategy, which facilitates the constant cyclic evolution of the generated samples and the quantized model (student) in a competitive and interactive fashion under the supervision of the full-precision (FP) model (teacher), thus significantly improving the accuracy of the quantized model. Moreover, without the auxiliary category guidance, we employ the task- and model-independent prior information, making the general-purpose scheme compatible with a broad range of vision tasks and models. Extensive experiments are conducted on various models on image classification, object detection, and semantic segmentation tasks, and PSAQ-ViT V2, with the naive quantization strategy and without access to real-world data, consistently achieves competitive results, showing potential as a powerful baseline on data-free quantization for ViTs. For instance, with Swin-S as the (backbone) model, 8-bit quantization reaches 82.13 top-1 accuracy on ImageNet, 50.9 box AP and 44.1 mask AP on COCO, and 47.2 mean Intersection over Union (mIoU) on ADE20K. We hope that accurate and general PSAQ-ViT V2 can serve as a potential and practice solution in real-world applications involving sensitive data. Code is released and merged at: https://github.com/zkkli/PSAQ-ViT. Zhikai Li, Mengjuan Chen, Junrui Xiao, Qingyi Gu |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2023 | I-ViT: Integer-only Quantization for Efficient Vision Transformer InferenceabstractVision Transformers (ViTs) have achieved state-of-the-art performance on various computer vision applications. However, these models have considerable storage and computational overheads, making their deployment and efficient inference on edge devices challenging. Quantization is a promising approach to reducing model complexity, and the dyadic arithmetic pipeline can allow the quantized models to perform efficient integer-only inference. Unfortunately, dyadic arithmetic is based on the homogeneity condition in convolutional neural networks, which is not applicable to the non-linear components in ViTs, making integer-only inference of ViTs an open issue. In this paper, we propose I-ViT, an integer-only quantization scheme for ViTs, to enable ViTs to perform the entire computational graph of inference with integer arithmetic and bit-shifting, and without any floating-point arithmetic. In I-ViT, linear operations (e.g., MatMul and Dense) follow the integer-only pipeline with dyadic arithmetic, and non-linear operations (e.g., Softmax, GELU, and LayerNorm) are approximated by the proposed light-weight integer-only arithmetic methods. More specifically, I-ViT applies the proposed Shiftmax and ShiftGELU, which are designed to use integer bit-shifting to approximate the corresponding floating-point operations. We evaluate I-ViT on various benchmark models and the results show that integer-only INT8 quantization achieves comparable (or even slightly higher) accuracy to the full-precision (FP) baseline. Furthermore, we utilize TVM for practical hardware deployment on the GPU’s integer arithmetic units, achieving 3.72 ~ 4.11 inference speedup compared to the FP model. Code of both Pytorch and TVM is released at https://github.com/zkkli/I-ViT. Zhikai Li, Qingyi Gu |
ICCV | 2 |
| 2023 | RepQ-ViT: Scale Reparameterization for Post-Training Quantization of Vision TransformersabstractPost-training quantization (PTQ), which only requires a tiny dataset for calibration without end-to-end retraining, is a light and practical model compression technique. Recently, several PTQ schemes for vision transformers (ViTs) have been presented; unfortunately, they typically suffer from non-trivial accuracy degradation, especially in low-bit cases. In this paper, we propose RepQ-ViT, a novel PTQ framework for ViTs based on quantization scale reparameterization, to address the above issues. RepQ-ViT decouples the quantization and inference processes, where the former employs complex quantizers and the latter employs scale-reparameterized simplified quantizers. This ensures both accurate quantization and efficient inference, which distinguishes it from existing approaches that sacrifice quantization performance to meet the target hardware. More specifically, we focus on two components with extreme distributions: post-LayerNorm activations with severe inter-channel variation and post-Softmax activations with power-law features, and initially apply channel-wise quantization and log$\sqrt 2 $ quantization, respectively. Then, we reparameterize the scales to hardware-friendly layer-wise quantization and log2 quantization for inference, with only slight accuracy or computational costs. Extensive experiments are conducted on multiple vision tasks with different model variants, proving that RepQ-ViT, without hyperparameters and expensive reconstruction procedures, can outperform existing strong baselines and encouragingly improve the accuracy of 4-bit PTQ of ViTs to a usable level. Code is available at https://github.com/zkkli/RepQ-ViT. Zhikai Li, Junrui Xiao, Lianwei Yang, Qingyi Gu |
ICCV | 4 |
| 2023 | Patch-wise Mixed-Precision Quantization of Vision TransformerabstractAs emerging hardware begins to support mixed bit-width arithmetic computation, mixed-precision quantization is widely used to reduce the complexity of neural networks. However, Vision Transformers (ViT ViTs) require complex self-attention computation to guarantee the learning of powerful feature representations, which makes mixed-precision quantization of ViTs still challenging. In this paper, we propose a novel patch-wise mixed-precision quantization (PMQ) for efficient inference of ViT$s$. Specifically, we design a lightweight global metric, which is faster than existing methods, to measure the sensitivity of each component in ViT$s$to quantization errors. Moreover, we also introduce a pareto frontier approach to automatically allocate the optimal bit-precision according to the sensitivity. To further reduce the computational complexity of self-attention in inference stage, we propose a patch-wise module to reallocate bit-width of patches in each layer. Extensive experiments on the ImageNet dataset shows that our method greatly reduces the search cost and facilitates the application of mixed-precision quantization to ViTs. Junrui Xiao, Zhikai Li, Lianwei Yang, Qingyi Gu |
IJCNN | 4 |
| 2023 | DCIFPN: Deformable cross-scale interaction feature pyramid network for object detectionabstractAbstract Exploiting multi‐scale features is one of the most effective methods to recognize objects of different scales in object detection. Since image pyramid is time‐consuming, Feature Pyramid Network (FPN) becomes the most popular component used for obtaining pyramidal features. Despite its effectiveness, there still exist some intrinsic defects. In this work, it is attributed to insufficient information flow and a Deformable Cross‐scale Interaction Feature Pyramid Network (DCIFPN) is proposed, which aims to promote the information transfer process with content‐aware sampling and dynamic aggregation weights. More specifically, Deformable Semantic Enhancement Module (DSEM) is designed that can construct accurate information flow with dynamic aggregation weights. In addition, Deformable Spatial Refinement Module (DSRM) is proposed to enhance high‐level features with low‐level location details. When DCIFPN is deployed on RetinaNet and FCOS with ResNet‐50, the performance is improved by 1.6 AP and 1.1 AP, respectively, on the challenging MS COCO benchmark. Apart from one‐stage detectors, DCIFPN is also applicable to two‐stage methods such as Faster R‐CNN and Mask R‐CNN. Further experiments on Pascal VOC and CrowdHuman datasets can verify the effectiveness and generalization of the method. Junrui Xiao, Zhikai Li, Qingyi Gu |
IET Image Process. | 4 |
| 2023 | Temporal action detection with dynamic weights based on curriculum learning
Yunze Chen, Junrui Xiao, Qingyi Gu |
Neurocomputing | 5 |
| 2023 | PWSNAS: Powering Weight Sharing NAS With General Search Space Shrinking FrameworkabstractNeural architecture search (NAS) depends heavily on an efficient and accurate performance estimator. To speed up the evaluation process, recent advances, like differentiable architecture search (DARTS) and One-Shot approaches, instead of training every model from scratch, train a weight-sharing super-network to reuse parameters among different candidates, in which all child models can be efficiently evaluated. Though these methods significantly boost search efficiency, they inherently suffer from inaccurate and unstable performance estimation. To this end, we propose a general and effective framework for powering weight-sharing NAS, namely, PWSNAS, by shrinking search space automatically, i.e., candidate operators will be discarded if they are less important. With the strategy, our approach can provide a promising search space of a smaller size by progressively simplifying the original search space, which can reduce difficulties for existing NAS methods to find superior architectures. In particular, we present two strategies to guide the shrinking process: detect redundant operators with a new angle-based metric and decrease the degree of weight sharing of a super-network by increasing parameters, which differentiates PWSNAS from existing shrinking methods. Comprehensive analysis experiments on NASBench-201 verify the superiority of our proposed metric over existing accuracy-based and magnitude-based metrics. PWSNAS can easily apply to the state-of-the-art NAS methods, e.g., single path one-shot neural architecture search (SPOS), FairNAS, ProxylessNAS, DARTS, and progressive DARTS (PDARTS). We evaluate PWSNAS and demonstrate consistent performance gains over baseline methods. Yiming Hu, Xingang Wang 0003, Qingyi Gu |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2022 | Patch Similarity Aware Data-Free Quantization for Vision Transformers
Zhikai Li, Liping Ma, Mengjuan Chen, Junrui Xiao, Qingyi Gu |
ECCV (11) | 5 |
| 2022 | G-Head: Gating Head for Multi-Task Learning in One-Stage Object DetectionabstractObject detection is commonly formulated as a multi-task learning problem in deep learning methods. Due to the di-vergence between classification and regression tasks, modern one-stage detectors typically utilize two parallel branches as the detection head, which might be sub-optimal. In this paper, we propose a new Gating Head (G-Head) to enhance the in-teraction between different tasks and promote the multi-task learning process. By introducing Multi-Scale Aggregation (MSA), Multi-Aspect Learning (MAL), and Gating Selec-tor (GS), our method can significantly boost the performance of existing one-stage frameworks with fewer parameters and computational costs. To validate the efficiency, effectiveness, and generalization of our G- Head, extensive experiments are conducted on the challenging MS COCO dataset. Without bells and whistles, we achieve a new state-of-the-art 48.7 AP under single-model and single-scale test. Qingyi Gu |
ICME | 2 |
| 2022 | A Flexible Calibration Algorithm for High-speed Bionic Vision System based on GalvanometerabstractTraditional gimbal-based bionic eye systems usually use a multi-degree-of-freedom mechanical platform to move the camera freely, which makes the structure complex and bulky. The galvanometer-based reflective bionic eye system uses a galvanometer to replace the traditional mechanical rotation structure, which separates the camera from the gimbal system, greatly simplifying the structure. However, there are currently few methods for calibrating such systems, mostly for object detection and tracking. In this paper, a flexible method for high-precision calibration of a galvanometer-based reflective bionic eye system is proposed. In this method, a planar target is used for the calibration of the bionic eye system. The effectiveness and accuracy of the method are evaluated by the reprojection error of the control voltage and the spatial localization of the binocular system. Experiments show that the error of the control voltage after calibration is less than 0.2%. At an indoor distance of about 7 m, the RMSE of spatial visual localization is less than 0.3 cm. Qing Li 0046, Mengjuan Chen, Qingyi Gu, Idaku Ishii |
IROS | 3 |
| 2022 | Class-wise boundary regression by uncertainty in temporal action detectionabstractAbstract Temporal action detection is a crucial aspect of video understanding. It aims to classify the action as well as locate the start and end boundaries of the action in the untrimmed videos. As deep learning is frequently utilized, the accuracy of annotation is crucial to boundary localization. However, it is observed that some annotation instances are ambiguous and the ambiguity varies between categories. To solve the problem above, a Gaussian model is built to estimate the boundary uncertainty for each instance. Based on instance uncertainty, category uncertainty is applied to describe the uncertainty of each category. By combining instance and category uncertainty, the boundaries of the selected proposals are refined and the ranking of candidate proposals is adjusted. Furthermore, overcorrection is avoided for categories with a high level of uncertainty. With the uncertainty approach, state‐of‐the‐art performance is achieved: 57.5% on THUMOS14 ([email protected]) and 35.4% on ActivityNet (mAP@Avg). Yunze Chen, Mengjuan Chen, Qingyi Gu |
IET Image Process. | 3 |
| 2022 | Dual-discriminator adversarial framework for data-free quantization
Zhikai Li, Liping Ma, Xianlei Long, Junrui Xiao, Qingyi Gu |
Neurocomputing | 5 |
| 2022 | Rethinking prediction alignment in one-stage object detection
Junrui Xiao, Zhikai Li, Qingyi Gu |
Neurocomputing | 4 |
| 2022 | Boosting semi-supervised face recognition with raw faces
Yunze Chen, Junjie Huang 0005, Xianlei Long, Qingyi Gu |
Image Vis. Comput. | 5 |
| 2021 | Improving One-Shot NAS with Shrinking-and-Expanding Supernet
Yiming Hu, Xingang Wang 0003, Lujun Li 0001, Qingyi Gu |
Pattern Recognit. | 4 |
| 2021 | Efficient Center Voting for Object Detection and 6D Pose Estimation in 3D Point CloudabstractWe present a novel and efficient approach to estimate 6D object poses of known objects in complex scenes represented by point clouds. Our approach is based on the well-known point pair feature (PPF) matching, which utilizes self-similar point pairs to compute potential matches and thereby cast votes for the object pose by a voting scheme. The main contribution of this paper is to present an improved PPF-based recognition framework, especially a new center voting strategy based on the relative geometric relationship between the object center and point pair features. Using this geometric relationship, we first generate votes to object centers resulting in vote clusters near real object centers. Then we group and aggregate these votes to generate a set of pose hypotheses. Finally, a pose verification operator is performed to filter out false positives and predict appropriate 6D poses of the target object. Our approach is also suitable to solve the multi-instance and multi-object detection tasks. Extensive experiments on a variety of challenging benchmark datasets demonstrate that the proposed algorithm is discriminative and robust towards similar-looking distractors, sensor noise, and geometrically simple shapes. The advantage of our work is further verified by comparing to the state-of-the-art approaches. Jianwei Guo 0003, Xuejun Xing, Weize Quan, Dong-Ming Yan 0001, Qingyi Gu, Yang Liu 0014, Xiaopeng Zhang 0001 |
IEEE Trans. Image Process. | 5 |
| 2020 | Refinement of Boundary Regression Using Uncertainty in Temporal Action Localization
Yunze Chen, Mengjuan Chen, Jiagang Zhu, Qingyi Gu |
BMVC | 6 |
| 2020 | Angle-Based Search Space Shrinking for Neural Architecture Search
Yiming Hu, Yuding Liang, Zichao Guo, Ruosi Wan, Xiangyu Zhang 0005, Qingyi Gu, Jian Sun 0001 |
ECCV (19) | 7 |
| 2020 | Natural Scene Facial Expression Recognition with Dimension Reduction NetworkabstractAs an external manifestation of human emotions, expression recognition plays an important role in human-computer interaction. Although existing expression recognition methods performs perfectly on constrained frontal faces, there are still many challenges in expression recognition in natural scenes due to different unrestricted conditions. Expression classification belongs to a pattern recognition problem where intra-class distance is greater than the inter-class distance, which leads to severe over-fitting when using neural networks for expression recognition. This paper proposes a novel net-work structure called Dimension Reduction Network which can effectively reduce generalization error. By adding a data dimension reduction module before the general classification network, a lot of redundant information is filtered, and only useful information is left. This can reduce the interference by irrelevant information when performing classification tasks and reduce generalization error. The proposed method does not require any modification to the classification network, only a small dimension reduction module needs to be added in front of the classification network. However, it can effectively reduce generalization error. We designed big and tiny versions of Dimension Reduction Network, both exceeds our baseline on AffectNet data set. The big version of our proposed method surpassed the state-of-the-art methods by more than 1.2% on AffectNet data set. Our code will open source3when the paper is accepted. Shenhua Hu, Yiming Hu, Xianlei Long, Mengjuan Chen, Qingyi Gu |
ICRA | 6 |
| 2019 | Cluster Regularized Quantization for Deep Networks CompressionabstractDeep neural networks (DNNs) have achieved great success in a wide range of computer vision areas, but the applications to mobile devices is limited due to their high storage and computational cost. Much efforts have been devoted to compress DNNs. In this paper, we propose a simple yet effective method for deep networks compression, named Cluster Regularized Quantization (CRQ), which can reduce the presentation precision of a full-precision model to ternary values without significant accuracy drop. In particular, the proposed method aims at reducing the quantization error by introducing a cluster regularization term, which is imposed on the full-precision weights to enable them naturally concentrate around the target values. Through explicitly regularizing the weights during the re-training stage, the full-precision model can achieve the smooth transition to the low-bit one. Comprehensive experiments on benchmark datasets demonstrate the effectiveness of the proposed method. Yiming Hu, Xianlei Long, Shenhua Hu, Jiagang Zhu, Xingang Wang 0003, Qingyi Gu |
ICIP | 7 |
| 2019 | Multi-Loss-Aware Channel Pruning of Deep NetworksabstractChannel pruning, which seeks to reduce the model size by removing redundant channels, is a popular solution for deep networks compression. Existing channel pruning methods usually conduct layer-wise channel selection by directly minimizing the reconstruction error of feature maps between the baseline model and the pruned one. However, they ignore the feature and semantic distributions within feature maps and real contribution of channels to the overall performance. In this paper, we propose a new channel pruning method by explicitly using both intermediate outputs of the baseline model and the classification loss of the pruned model to supervise layer-wise channel selection. Particularly, we introduce an additional loss to encode the differences in the feature and semantic distributions within feature maps between the baseline model and the pruned one. By considering the reconstruction error, the additional loss and the classification loss at the same time, our approach can significantly improve the performance of the pruned model. Comprehensive experiments on benchmark datasets demonstrate the effectiveness of the proposed method. Yiming Hu, Siyang Sun, Jiagang Zhu, Xingang Wang 0003, Qingyi Gu |
ICIP | 6 |
| 2017 | 12, 000-fps Multi-object detection using HOG descriptor and SVM classifierabstractThis paper describes a high-frame-rate (HFR) vision system that can detect multiple objects in an image of 512 × 512 pixels at 12,000 frames per seconds (fps). An optimized algorithm is proposed based on conventional Histograms of Oriented Gradient (HOG) descriptor and Support Vector Machine (SVM) classifier algorithms for hardware implementation. By implementing the proposed algorithm on a field-programmable gate array (FPGA) of a high-speed vision platform, multi-object in an image can be detected at 12,000 fps under complex background. In hardware implementation, 64 pixels were processed in parallel with 80 MHz camera clock. Source image and detection results can be transferred to personal computer (PC) in real-time for recording or post-processing. Our developed HFR multi-object detection system was verified by performing several evaluations. Yingjie Yin, Xilong Liu, De Xu, Qingyi Gu |
IROS | 5 |
| 2016 | Control scheme of nongrasping manipulation based on virtual connecting constraintabstractThe research field of nongrasping manipulation is a maturing area in robotic motion control. However, the common principles of motion planning for nongrasping manipulation systems have not yet been established. This paper proposes the concept of virtual connecting manipulation as a generalized motion planning framework for nongrasping manipulation systems. As a preliminary result, we had previously realized a flower-stick juggling task called “propeller motion” using an actual experimental system. In this paper, we apply the virtual connecting manipulation concept to a flower-stick juggling task and analyze the generated motion from the view point of analytical methodology. We conduct a stability analysis of the generated cyclic motion of the flower stick by using a Poincaré map, and the analytical results show that the generated cyclic motion is asymptotically stable. Tadayoshi Aoyama, Takeshi Takaki, Qingyi Gu, Idaku Ishii |
ICRA | 3 |
| 2015 | A scheme for manipulating a passive object using an active plateabstractWe propose a novel scheme for manipulating a passive object using an active plate. The objective of this study is to control an object's orientation with respect to the gravitational force direction by using an active plate for realizing hitherto unrealized object motion. In this context, motions of the object and active plate are designed to be cyclic. A state vector composed of the object's angle and angular velocity is defined, and the cyclic motion is expressed as a nonlinear discrete system. Fixed points of the state vector are searched for in the designed cyclic motion. A stability analysis around the fixed points is conducted using a Poincaré map. As a result, the fixed points are shown to be asymptotically stable. Finally, experimental results are used to verify that the object's angle can be manipulated with the designed cyclic motion using the plate. Tadayoshi Aoyama, Yuji Harada, Qingyi Gu, Takeshi Takaki, Idaku Ishii |
ICRA | 3 |
| 2015 | Realization of flower stick rotation using robotic armabstractFlower stick juggling is a dexterous task done by skillful jugglers. We aim to realize dexterous tasks done by humans using robotic systems. This work focuses on flower stick juggling and proposes a feedback control strategy for a flower stick juggling task called “propeller” as one of the robotic dexterous manipulations. The propeller motion is modeled by considering combined flower stick and a robotic manipulator. We developed a control strategy that allows stable cyclic rotation of the flower stick in the air. The control parametars in the control strategy are searched through numerical simulation. Finally, the flower stick propeller motion is realized using an actual robotic system. Tadayoshi Aoyama, Takeshi Takaki, Takumi Miura, Qingyi Gu, Idaku Ishii |
IROS | 4 |
| 2015 | Simultaneous Vision-Based Shape and Motion Analysis of Cells Fast-Flowing in a MicrochannelabstractThis paper proposes a novel concept for simultaneous cell shape and motion analysis in fast microchannel flows by implementing a multiobject feature extraction algorithm on a frame-straddling high-speed vision platform. The system can synchronize two camera inputs with the same view with only a tiny time delay on the sub-microsecond timescale. Real-time video processing is performed in hardware logic by extracting the moment features of multiple cells in 512 × 256 images at 4000 fps for the two camera inputs and their frame-straddling time can be adjusted from 0 to 0.25 ms in 9.9 ns steps. By setting the frame-straddling time in a certain range to avoid large image displacements between the two camera inputs, our frame-straddling high-speed vision platform can perform simultaneous shape and motion analysis of cells in fast microchannel flows of 1 m/s or greater. The results of real-time experiments conducted to analyze the deformabilities and velocities of sea urchin egg cells fast-flowing in microchannels verify the efficacy of our vision-based cell analysis system. Qingyi Gu, Tadayoshi Aoyama, Takeshi Takaki, Idaku Ishii |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2015 | LOC-Based High-Throughput Cell Morphology Analysis SystemabstractWe present a high-speed vision-based morphological analysis system for fast-flowing cells in a microchannel. Real-time video processing is performed in hardware logic by extracting the moment features and bounding boxes of multiple cells in 512 × 256-pixel images at 2000 fps. The extracted cell regions are pushed into a first-in-first-out (FIFO) buffer for real-time image-based morphological analysis after being shrunk proportionally to a certain size. By extracting the bounding boxes of the cell regions using hardware logic and shrinking the cell region to a certain size to reduce processing time, our high-speed vision system can perform fast morphological analysis of cells at 500 cells/s in fast microchannel flows. Moreover, snapshots of all passed cells under the microscope are stored for offline verification. The results of real-time experiments conducted to analyze the size, eccentricity, and transparency of sea urchin embryos flowing fast in microchannels verify the efficacy of our vision-based cell analysis system. Qingyi Gu, Tomohiro Kawahara, Tadayoshi Aoyama, Takeshi Takaki, Idaku Ishii, Ayumi Takemoto, Naoaki Sakamoto |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2014 | Rapid vision-based shape and motion analysis system for fast-flowing cells in a microchannelabstractThis paper proposes a novel method for simultaneous cell shape and motion analysis in rapid microchannel flows based on a multi-object feature extraction algorithm with a frame-straddling high-speed vision platform. This system can synchronize two camera inputs that share the same view with only a very small sub-microsecond time delay. Real-time video processing is performed using the hardware logic by extracting the moment features of multiple cells at 2000 fps or more, which are obtained from the two camera inputs, and their frame-straddling time can be adjusted from 0 to 0.5 ms in 9.9 ns steps. After setting the frame-straddling time within a certain range to avoid large image displacements between the two camera inputs, the frame-straddling high-speed vision platform can perform simultaneous shape and motion analysis of cells in fast microchannel flows of 1 m/s or greater. The results of real-time experiments conducted to analyze the deformabilities, velocities, and shapes of fast-flowing sea urchin egg cells in straight and L-type microchannels verified the efficacy of our vision-based cell analysis system. Qingyi Gu, Tadayoshi Aoyama, Takeshi Takaki, Idaku Ishii |
ICRA | 1 |
| 2014 | Real-time LOC-based morphological cell analysis system using high-speed visionabstractIn this paper, a high-speed vision-based morphological analysis system for fast-flowing cells in a microchannel implementing a multi-object feature extraction algorithm on a high-speed vision platform is proposed. Real-time video processing is performed in hardware logic by extracting the moment features and bounding boxes of multiple cells in 512×256-pixel images at 2000 fps. The extracted cell regions are pushed into a first-in-first-out (FIFO) buffer for real-time image-based morphological analysis after being shrunk proportionally to a certain size. By extracting the bounding boxes of the cell regions using hardware logic and shrinking the cell region to a certain size to reduce processing time, our high-speed vision system can perform fast morphological analysis of cells at 2 ms/cell in fast microchannel flows. The results of real-time experiments conducted to analyze the size, eccentricity, and transparency of fertilized sea urchin eggs fast flowing in microchannels verify the efficacy of our vision-based cell analysis system. Qingyi Gu, Tadayoshi Aoyama, Takeshi Takaki, Idaku Ishii, Ayumi Takemoto, Naoaki Sakamoto |
IROS | 1 |
| 2013 | Fast 3-D shape measurement using blink-dot projectionabstractWe propose a novel dot-pattern-projection three-dimensional (3-D) shape measurement method that can measure 3-D displacements of blink dots projected onto a measured object accurately even when it moves rapidly or is observed from a camera as moving rapidly. In our method, blinking dot patterns, in which each dot changes its size at different timings corresponding to its identification (ID) number, are projected from a projector at a high frame rate. 3-D shapes can be obtained without any miscorrespondence of the projected dots between frames by simultaneous tracking and identification of multiple dots projected onto a measured 3-D object in a camera view. Our method is implemented on a field-programmable gate array (FPGA)-based high-frame-rate (HFR) vision platform that can track and recognize as much as 15×15 blink-dot pattern in a 512×512 image in real time at 1000 fps, synchronized with an HFR projector. We demonstrate the performance of our system by showing real-time 3-D measurement results when our system is mounted on a parallel link manipulator as a sensing head. Jun Chen 0017, Qingyi Gu, Tadayoshi Aoyama, Takeshi Takaki, Idaku Ishii |
IROS | 2 |
| 2013 | Real-time feature-based video mosaicing at 500 fpsabstractWe conducted high-frame-rate (HFR) video mosaicing for real-time synthesis of a panoramic image by implementing an improved feature-based video mosaicing algorithm on a field-programmable gate array (FPGA)-based high-speed vision platform. In the implementation of the mosaicing algorithm, feature point extraction was accelerated by implementing a parallel processing circuit module for Harris corner detection in the FPGA on the high-speed vision platform. Feature point correspondence matching can be executed for hundreds of selected feature points in the current frame by searching those in the previous frame in their neighbor ranges, assuming that frame-to-frame image displacement becomes considerably smaller in HFR vision. The system we developed can mosaic 512×512 images at 500 fps as a single synthesized image in real time by stitching the images based on their estimated frame-to-frame changes in displacement and orientation. The results of an experiment conducted, in which an outdoor scene was captured using a hand-held camera-head that was quickly moved by hand, verify the performance of our system. Ken-ichi Okumura, Sushil Raut, Qingyi Gu, Tadayoshi Aoyama, Takeshi Takaki, Idaku Ishii |
IROS | 3 |
| 2013 | A fast multi-camera tracking system with heterogeneous lensesabstractWe have developed a fast target tracking system that utilizes four cameras with lenses of different focal lengths to track an object without blurring images, even when the object moves in the depth direction away from the cameras in three-dimensional (3-D) space. This system can maintain a well-focused camera view by switching among the four input images, instead of using lens motor control. The multi-camera system was mounted on a two-axis mechanical active vision platform. The active vision control and camera-view switching are executed by processing color 512 × 512 images from the four camera inputs at 500 fps in real time on a high-speed vision platform. The performance of our system was verified by its tracking results for objects moving rapidly in 3-D space. Xiaorong Zhao, Qingyi Gu, Tadayoshi Aoyama, Takeshi Takaki, Idaku Ishii |
IROS | 2 |
| 2013 | Fast FPGA-Based Multiobject Feature ExtractionabstractThis paper describes a high-frame-rate (HFR) vision system that can extract locations and features of multiple objects in an image at 2000 f/s for 512 × 512 images by implementing a cell-based multiobject feature extraction algorithm as hardware logic on a field-programmable gate array-based high-speed vision platform. In the hardware implementation of the algorithm, 25 higher-order local autocorrelation features of 1024 objects in an image can be simultaneously extracted for multiobject recognition by dividing the image into 8 × 8 cells concurrently with calculation of the zeroth and first-order moments to obtain the sizes and locations of multiple objects. Our developed HFR multiobject extraction system was verified by performing several experiments: tracking for multiple objects rotating at 16 r/s, recognition for multiple patterns projected at 1000 f/s, and recognition for human gestures with quick finger motion. Qingyi Gu, Takeshi Takaki, Idaku Ishii |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2010 | 2000-fps multi-object extraction based on cell-based labelingabstractReal-time multi-object extraction at 2000 fps was realized by designing a cell-based labeling algorithm. The algorithm can label the divided cells in an image by scanning the image only once to obtain their moment features, and the computational complexity required for labeling can be remarkably reduced. The cell-based labeling algorithm for 8 × 8 pixel cells was implemented on a high-speed vision platform, and multiple objects in an image of 512 × 512 pixels could be extracted at 2000 fps. An experiment was performed using a quickly rotating object to verify the performance of our multi-object extraction system. Qingyi Gu, Takeshi Takaki, Idaku Ishii |
ICIP | 1 |
| 2010 | 2000 fps real-time vision system with high-frame-rate video recordingabstractThis paper introduces a high-speed vision system called IDP Express, which can execute real-time image processing and high frame rate video recording simultaneously. In IDP Express, a dedicated FPGA (Field Programmable Gate Array) board processes 512 × 512 pixel images from two camera heads by implementing image processing algorithms as hardware logic; the input images and processed results are transferred to standard PC memory at a rate of 2000 fps or more. Owing to the simultaneous high-frame-rate video processing and recording, IDP Express can be used as an intelligent video logger for long-term high-speed phenomenon analysis even when the measured objects move quickly in a wide area. We applied IDP Express to a mechanical target tracking system to record a high-frame-rate video at high resolution for a crucial moment, which is magnified by tracking the measured objects with visual feedback control. Several experiments on moving objects that undergo sudden shape deformation were performed. The results of the experiments involving the explosion of a rotating balloon and the crash of falling custard pudding have been provided to verify the effectiveness of IDP Express. Idaku Ishii, Tetsuro Tatebe, Qingyi Gu, Yuta Moriue, Takeshi Takaki, Kenji Tajima |
ICRA | 3 |