VLDB 2026 Research / reviewers in the wild / expert
Hongwei Qin
dblp:161/1819
· DBLP profile ↗
33ranked-venue papers
6as first author
23since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 24 · 2 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 21 · 2 first-author · 15 since 2021Systems, architecture and hardware · 4 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Task-Aware Encoder Control for Deep Video CompressionabstractPrior research on deep video compression (DVC) for machine tasks typically necessitates training a unique codec for each specific task, mandating a dedicated decoder per task. In contrast, traditional video codecs employ a flexible encoder controller, enabling the adaptation of a single codec to different tasks through mechanisms like mode prediction. Drawing inspiration from this, we introduce an innovative encoder controller for deep video compression for machines. This controller features a mode prediction and a Group of Pictures (GoP) selection module. Our approach centralizes control at the encoding stage, allowing for adaptable encoder adjustments across different tasks, such as detection and tracking, while maintaining compatibility with a standard pre-trained DvC decoder. Empirical evidence demonstrates that our method is applica-ble across multiple tasks with various existing pre-trained Dv'Cs. Moreover, extensive experiments demonstrate that our method outperforms previous DVC by about 25% bi-trate for different tasks, with only one pre-trained decoder. Xingtong Ge, Jixiang Luo, Tongda Xu, Guo Lu, Dailan He, Yan Wang 0105, Jun Zhang 0004, Hongwei Qin |
CVPR | 10 |
| 2024 | Boosting Neural Representations for Videos with a Conditional DecoderabstractImplicit neural representations (INRs) have emerged as a promising approach for video storage and processing, showing remarkable versatility across various video tasks. However, existing methods often fail to fully leverage their representation capabilities, primarily due to inadequate alignment of intermediate features during target frame decoding. This paper introduces a universal boosting framework for current implicit video representation approaches. Specifically, we utilize a conditional decoder with a temporal-aware affine transform module, which uses the frame index as a prior condition to effectively align intermediate features with target frames. Besides, we introduce a sinusoidal NeRV-like block to generate diverse intermediate features and achieve a more balanced parameter distribution, thereby enhancing the model's capacity. With a high-frequency information-preserving reconstruction loss, our approach successfully boosts multiple baseline INRs in the reconstruction quality and convergence speed for video regression, and exhibits superior inpainting and interpolation results. Further, we integrate a consistent entropy minimization technique and develop video codecs based on these boosted INRs. Experiments on the UVG dataset confirm that our enhanced codecs significantly outperform baseline INRs and offer competitive rate-distortion performance compared to traditional and learning-based codecs. Code is available at htt ps://github.com/Xin j ieQ/Boosting-NeRV. Dailan He, Xingtong Ge, Tongda Xu, Yan Wang 0105, Hongwei Qin, Jun Zhang 0004 |
CVPR | 7 |
| 2024 | GaussianImage: 1000 FPS Image Representation and Compression by 2D Gaussian Splatting
Xingtong Ge, Tongda Xu, Dailan He, Yan Wang 0080, Hongwei Qin, Guo Lu, Jun Zhang 0004 |
ECCV (9) | 6 |
| 2024 | Idempotence and Perceptual Image CompressionabstractIdempotence is the stability of image codec to re-compression. At the first glance, it is unrelated to perceptual image compression. However, we find that theoretically: 1) Conditional generative model-based perceptual codec satisfies idempotence; 2) Unconditional generative model with idempotence constraint is equivalent to conditional generative codec. Based on this newfound equivalence, we propose a new paradigm of perceptual image codec by inverting unconditional generative model with idempotence constraints. Our codec is theoretically equivalent to conditional generative codec, and it does not require training new models. Instead, it only requires a pre-trained mean-square-error codec and unconditional generative model. Empirically, we show that our proposed approach outperforms state-of-the-art methods such as HiFiC and ILLM, in terms of Fréchet Inception Distance (FID). The source code is provided in https://github.com/tongdaxu/Idempotence-and-Perceptual-Image-Compression. Tongda Xu, Ziran Zhu, Dailan He, Yanghao Li, Zhe Wang 0070, Hongwei Qin, Yan Wang 0105, Ya-Qin Zhang |
ICLR | 8 |
| 2023 | A Simple Baseline for Video Restoration with Grouped Spatial-Temporal ShiftabstractVideo restoration, which aims to restore clear frames from degraded videos, has numerous important applications. The key to video restoration depends on utilizing inter-frame information. However, existing deep learning methods often rely on complicated network architectures, such as optical flow estimation, deformable convo-lution, and cross-frame self-attention layers, resulting in high computational costs. In this study, we propose a sim-ple yet effective framework for video restoration. Our approach is based on grouped spatial-temporal shift, which is a lightweight and straightforward technique that can implicitly capture inter-frame correspondences for multi-frame aggregation. By introducing grouped spatial shift, we attain expansive effective receptive fields. Combined with basic 2D convolution, this simple framework can effectively aggregate inter-frame information. Extensive experiments demonstrate that our framework outperforms the previous state-of-the-art method, while using less than a quarter of its computational cost, on both video deblurring and video denoising tasks. These results indicate the potential for our approach to significantly reduce computational overhead while maintaining high-quality results. Code is avaliable at https://github.com/dasonglil/Shift-Net. Dasong Li, Xiaoyu Shi 0002, Yi Zhang 0108, Ka Chun Cheung, Simon See, Xiaogang Wang 0001, Hongwei Qin, Hongsheng Li 0001 |
CVPR | 7 |
| 2023 | FlowFormer++: Masked Cost Volume Autoencoding for Pretraining Optical Flow EstimationabstractFlowFormer [24] introduces a transformer architecture into optical flow estimation and achieves state-of-the-art performance. The core component of FlowFormer is the transformer-based cost-volume encoder. Inspired by the recent success of masked autoencoding (MAE) pretraining in unleashing transformers' capacity of encoding visual representation, we propose Masked Cost Volume Autoencoding (MCVA) to enhance FlowFormer by pretraining the cost-volume encoder with a novel MAE scheme. Firstly, we introduce a block-sharing masking strategy to prevent masked information leakage, as the cost maps of neighboring source pixels are highly correlated. Secondly, we propose a novel pre-text reconstruction task, which encourages the cost-volume encoder to aggregate long-range information and ensures pretraining-finetuning consistency. We also show how to modify the FlowFormer architecture to accommodate masks during pretraining. Pretrained with MCVA, FlowFormer++ ranks 1st among published methods on both Sintel and KITTI-2015 benchmarks. Specifically, FlowFormer++ achieves 1.07 and 1.94 average end-point error (AEPE) on the clean and final pass of Sintel benchmark, leading to 7.76% and 7.18% error reductions from FlowFormer. FlowFormer++ obtains 4.52 F1-all on the KITTI-2015 test set, improving FlowFormer by 0.16. Xiaoyu Shi 0002, Dasong Li, Manyuan Zhang, Ka Chun Cheung, Simon See, Hongwei Qin, Jifeng Dai, Hongsheng Li 0001 |
CVPR | 7 |
| 2023 | VideoFlow: Exploiting Temporal Cues for Multi-frame Optical Flow EstimationabstractWe introduce VideoFlow, a novel optical flow estimation framework for videos. In contrast to previous methods that learn to estimate optical flow from two frames, VideoFlow concurrently estimates bi-directional optical flows for multiple frames that are available in videos by sufficiently exploiting temporal cues.We first propose a TRi-frame Optical Flow (TROF) module that estimates bi-directional optical flows for the center frame in a three-frame manner. The information of the frame triplet is iteratively fused onto the center frame. To extend TROF for handling more frames, we further propose a MOtion Propagation (MOP) module that bridges multiple TROFs and propagates motion features between adjacent TROFs. With the iterative flow estimation refinement, the information fused in individual TROFs can be propagated into the whole sequence via MOP. By effectively exploiting video information, VideoFlow presents extraordinary performance, ranking 1st on all public benchmarks. On the Sintel benchmark, VideoFlow achieves 1.649 and 0.991 average end-point-error (AEPE) on the final and clean passes, a 15.1% and 7.6% error reduction from the best published results (1.943 and 1.073 from FlowFormer++). On the KITTI-2015 benchmark, VideoFlow achieves an F1-all error of 3.65%, a 19.2% error reduction from the best published result (4.52% from FlowFormer++). Code is released at https://github.com/XiaoyuShi97/VideoFlow. Xiaoyu Shi 0002, Weikang Bian, Dasong Li, Manyuan Zhang, Ka Chun Cheung, Simon See, Hongwei Qin, Jifeng Dai, Hongsheng Li 0001 |
ICCV | 8 |
| 2023 | Unified Learning-Based Lossy and Lossless Jpeg RecompressionabstractJPEG is still the most widely used image compression algorithm. Most image compression algorithms only consider uncompressed original image, while ignoring a large number of already existing JPEG images. Recently, JPEG recompression approaches have been proposed to further reduce the size of JPEG files. However, those methods only consider JPEG lossless recompression, which is just a special case of the rate-distortion theorem. In this paper, we propose a unified lossly and lossless JPEG recompression framework, which consists of learned quantization table and Markovian hierarchical variational autoencoders. Experiments show that our method can achieve arbitrarily low distortion when the bitrate is close to the upper bound, namely the bitrate of the lossless compression model. To the best of our knowledge, this is the first learned method that bridges the gap between lossy and lossless recompression of JPEG images. Jianghui Zhang, Jixiang Luo, Tongda Xu, Yan Wang 0105, Hongwei Qin |
ICIP | 8 |
| 2023 | Bit Allocation using OptimizationabstractIn this paper, we consider the problem of bit allocation in Neural Video Compression (NVC). First, we reveal a fundamental relationship between bit allocation in NVC and Semi-Amortized Variational Inference (SAVI). Specifically, we show that SAVI with GoP (Group-of-Picture)-level likelihood is equivalent to pixel-level bit allocation with precise rate & quality dependency model. Based on this equivalence, we establish a new paradigm of bit allocation using SAVI. Different from previous bit allocation methods, our approach requires no empirical model and is thus optimal. Moreover, as the original SAVI using gradient ascent only applies to single-level latent, we extend the SAVI to multi-level such as NVC by recursively applying back-propagating through gradient ascent. Finally, we propose a tractable approximation for practical implementation. Our method can be applied to scenarios where performance outweights encoding speed, and serves as an empirical bound on the R-D performance of bit allocation. Experimental results show that current state-of-the-art bit allocation algorithms still have a room of $\approx 0.5$ dB PSNR to improve compared with ours. Code is available at https://github.com/tongdaxu/Bit-Allocation-Using-Optimization. Tongda Xu, Han Gao 0012, Chenjian Gao, Dailan He, Jinyong Pi, Jixiang Luo, Mao Ye 0001, Hongwei Qin, Yan Wang 0080, Ya-Qin Zhang |
ICML | 10 |
| 2022 | Practical Learned Lossless JPEG Recompression with Multi-Level Cross-Channel Entropy Model in the DCT DomainabstractJPEG is a popular image compression method widely used by individuals, data center, cloud storage and network filesystems. However, most recent progress on image compression mainly focuses on uncompressed images while ignoring trillions of already-existing JPEG images. To compress these JPEG images adequately and restore them back to JPEG format losslessly when needed, we propose a deep learning based JPEG recompression method that operates on DCT domain and propose a Multi-Level Cross-Channel Entropy Model to compress the most informative Y component. Experiments show that our method achieves state-of-the-art performance compared with traditional JPEG recompression methods including Lepton, JPEG XL and CMIX. To the best of our knowledge, this is the first learned compression method that losslessly transcodes JPEG images to more storage-saving bitstreams. Xinjie Shi, Dailan He, Hongwei Qin, Yan Wang 0080 |
CVPR | 6 |
| 2022 | ELIC: Efficient Learned Image Compression with Unevenly Grouped Space-Channel Contextual Adaptive CodingabstractRecently, learned image compression techniques have achieved remarkable performance, even surpassing the best manually designed lossy image coders. They are promising to be large-scale adopted. For the sake of practicality, a thorough investigation of the architecture design of learned image compression, regarding both compression performance and running speed, is essential. In this paper, we first propose uneven channel-conditional adaptive coding, motivated by the observation of energy compaction in learned image compression. Combining the proposed uneven grouping model with existing context models, we obtain a spatial-channel contextual adaptive model to improve the coding performance without damage to running speed. Then we study the structure of the main transform and propose an efficient model, ELIC, to achieve state-of-the-art speed and compression ability. With superior performance, the proposed model also supports extremely fast preview decoding and progressive decoding, which makes the coming application of learning-based image compression more promising. Dailan He, Zimin (Max) Yang, Weikun Peng, Hongwei Qin, Yan Wang 0080 |
CVPR | 5 |
| 2022 | IDR: Self-Supervised Image Denoising via Iterative Data RefinementabstractThe lack of large-scale noisy-clean image pairs restricts supervised denoising methods' deployment in actual applications. While existing unsupervised methods are able to learn image denoising without ground-truth clean images, they either show poor performance or work under impractical settings (e.g., paired noisy images). In this paper, we present a practical unsupervised image denoising method to achieve state-of-the-art denoising performance. Our method only requires single noisy images and a noise model, which is easily accessible in practical raw image denoising. It performs two steps iteratively: (1) Constructing a noisier-noisy dataset with random noise from the noise model; (2) training a model on the noisier-noisy dataset and using the trained model to refine noisy images to obtain the targets used in the next round. We further approximate our full iterative method with a fast algorithm for more efficient training while keeping its original high performance. Experiments on real-world, synthetic, and correlated noise show that our proposed unsupervised denoising approach has superior performances over existing unsupervised methods and competitive performance with supervised methods. In addition, we argue that existing denoising datasets are of low quality and contain only a small number of scenes. To evaluate raw image denoising performance in real-world applications, we build a high-quality raw image dataset SenseNoise-500 that contains 500 real-life scenes. The dataset can serve as a strong benchmark for better evaluating raw image denoising. Code and dataset will be released11https://github.com/zhangyi-3/IDR Yi Zhang 0108, Dasong Li, Ka Lung Law, Xiaogang Wang 0001, Hongwei Qin, Hongsheng Li 0001 |
CVPR | 5 |
| 2022 | FlowFormer: A Transformer Architecture for Optical Flow
Xiaoyu Shi 0002, Qiang Wang 0023, Ka Chun Cheung, Hongwei Qin, Jifeng Dai, Hongsheng Li 0001 |
ECCV (17) | 6 |
| 2022 | Learning Degradation Representations for Image Deblurring
Dasong Li, Yi Zhang 0108, Ka Chun Cheung, Xiaogang Wang 0001, Hongwei Qin, Hongsheng Li 0001 |
ECCV (18) | 5 |
| 2022 | Spatial Moment Pooling Improves Neural Image AssessmentabstractIn recent years, there has been widespread attention drawn to convolutional neural network (CNN) based blind image quality assessment (IQA). A large number of works start by extracting deep features from CNN. Then, those features are processed through spatial average pooling (SAP) and fully connected layers to predict quality. Inspired by full reference IQA and texture features, in this paper, we extend SAP (1stmoment) into spatial moment pooling (SMP) by incorporating higher order moments (such as variance, skewness). Moreover, we provide learning friendly normalization to circumvent numerical issue when computing gradients of higher moments. Experimental results suggest that simply upgrading SAP to SMP significantly enhances CNN-based blind IQA methods and achieves state of the art performance. Tongda Xu, Yifan Shao, Yan Wang 0080, Hongwei Qin |
ICIP | 4 |
| 2022 | Flexible Neural Image Compression via Code EditingabstractNeural image compression (NIC) has outperformed traditional image codecs in rate-distortion (R-D) performance. However, it usually requires a dedicated encoder-decoder pair for each point on R-D curve, which greatly hinders its practical deployment. While some recent works have enabled bitrate control via conditional coding, they impose strong prior during training and provide limited flexibility. In this paper we propose Code Editing, a highly flexible coding method for NIC based on semi-amortized inference and adaptive quantization. Our work is a new paradigm for variable bitrate NIC, and experimental results show that our method surpasses existing variable-rate methods. Furthermore, our approach is so flexible that it can also achieves ROI coding and multi-distortion trade-off with a single decoder. Our approach is compatible to all NIC methods with differentiable decoder NIC, and it can be even directly adopted on existing pre-trained models. Chenjian Gao, Tongda Xu, Dailan He, Yan Wang 0080, Hongwei Qin |
NeurIPS | 5 |
| 2022 | Multi-Sample Training for Neural Image CompressionabstractThis paper considers the problem of lossy neural image compression (NIC). Current state-of-the-art (SOTA) methods adopt uniform posterior to approximate quantization noise, and single-sample pathwise estimator to approximate the gradient of evidence lower bound (ELBO). In this paper, we propose to train NIC with multiple-sample importance weighted autoencoder (IWAE) target, which is tighter than ELBO and converges to log likelihood as sample size increases. First, we identify that the uniform posterior of NIC has special properties, which affect the variance and bias of pathwise and score function estimators of the IWAE target. Moreover, we provide insights on a commonly adopted trick in NIC from gradient variance perspective. Based on those analysis, we further propose multiple-sample NIC (MS-NIC), an enhanced IWAE target for NIC. Experimental results demonstrate that it improves SOTA NIC methods. Our MS-NIC is plug-and-play, and can be easily extended to neural video compression. Tongda Xu, Yan Wang 0080, Dailan He, Chenjian Gao, Han Gao 0012, Kunzan Liu, Hongwei Qin |
NeurIPS | 7 |
| 2022 | Efficient Burst Raw Denoising with Variance Stabilization and Multi-frequency Denoising Network
Dasong Li, Yi Zhang 0108, Ka Lung Law, Xiaogang Wang 0001, Hongwei Qin, Hongsheng Li 0001 |
Int. J. Comput. Vis. | 5 |
| 2022 | Improving F2FS Performance in Mobile Devices With Adaptive Reserved Space Based on TracebackabstractAs we all know, file and free space fragmentation negatively affect file system performance. F2FS is a file system designed for flash memory. However, it suffers from severe fragmentation due to out-of-place updates and the highly synchronous, multithreaded writing behaviors of applications. The adaptive reserved space (ARS) scheme chooses some files instead of all files to update in the reserved space which collects file features associated with fragmentation to construct datasets and uses decision trees to pick reserved files. However, ARS processes file features according to writefds. On the one hand,fdsand files are mapped many-to-many, so there is an issue in mapping reserved files according to reservedfdsthat a normal file is selected as a reserved file. On the other hand, the mapping of files tofdsfor big data traces is confusing. We propose ARST that traces the file of writefdto optimize ARS. Moreover, by selecting time-independent file features, ARST predicts whether files with little historical information are reserved to maximize the performance improvement brought by reserving space. Besides, adjustable reserved space and dynamic reservation strategies are adopted. We implement ARST on a HiKey960 development platform and a commercial smartphone with slight space and file creation time overheads. Experimental results show that ARST reduces file and free space fragmentation dramatically and improves file system performance. ARST reduces the running time of realistic workloads by up to 94.28% than F2FS with in-place updates and outperforms ARS by up to 49.06% for Wechat. Fang Wang 0001, Dan Feng 0001, Hongwei Qin, Shiyun Tu, Jiaxing Qian |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2021 | Checkerboard Context Model for Efficient Learned Image CompressionabstractFor learned image compression, the autoregressive context model is proved effective in improving the rate-distortion (RD) performance. Because it helps remove spatial redundancies among latent representations. However, the decoding process must be done in a strict scan order, which breaks the parallelization. We propose a parallelizable checkerboard context model (CCM) to solve the problem. Our two-pass checkerboard context calculation eliminates such limitations on spatial locations by re-organizing the decoding order. Speeding up the decoding process more than 40 times in our experiments, it achieves significantly improved computational efficiency with almost the same rate-distortion performance. To the best of our knowledge, this is the first exploration on parallelization-friendly spatial context model for learned image compression. Dailan He, Yaoyan Zheng, Baocheng Sun 0002, Yan Wang 0080, Hongwei Qin |
CVPR | 5 |
| 2021 | Rethinking Noise Synthesis and Modeling in Raw DenoisingabstractThe lack of large-scale real raw image denoising dataset gives rise to challenges on synthesizing realistic raw image noise for training denoising models. However, the real raw image noise is contributed by many noise sources and varies greatly among different sensors. Existing methods are unable to model all noise sources accurately, and building a noise model for each sensor is also laborious. In this paper, we introduce a new perspective to synthesize noise by directly sampling from the sensor’s real noise. It inherently generates accurate raw image noise for different camera sensors. Two efficient and generic techniques: pattern-aligned patch sampling and high-bit reconstruction help accurate synthesis of spatial-correlated noise and high-bit noise respectively. We conduct systematic experiments on SIDD and ELD datasets. The results show that (1) our method outperforms existing methods and demonstrates wide generalization on different sensors and lighting conditions. (2) Recent conclusions derived from DNN-based noise modeling methods are actually based on inaccurate noise parameters. The DNN-based methods still cannot outperform physics-based statistical methods. The code will be available at https://github.com/zhangyi-3/. Yi Zhang 0108, Hongwei Qin, Xiaogang Wang 0001, Hongsheng Li 0001 |
ICCV | 2 |
| 2021 | Better atomic writes by exposing the flash out-of-band area to file systemsabstractFile systems for mobile devices usually preserve data consistency by ordered I/Os. However, maintaining I/O ordering prevents applications from fully exploiting device parallelism and thus degrades the storage performance. In this paper, we propose NBStack to eliminate ordered I/Os without compromising data consistency. First, we augment the existing block interface to expose the Flash out-of-band area to file systems. Second, we build an enhanced block device prototype that supports the new interface. Third, we develop NBFS, a Linux file system, that leverages the new block interface to achieve atomic writes without enforcing I/O orderings. Experimental results show that NBStack doubles the performance of F2FS while providing strong consistency and durability guarantees. If applications are willing to trade-off durability, NBStack can further aggressively improve performance. Hongwei Qin, Dan Feng 0001, Wei Tong 0001, Sheng Qiu |
LCTES | 1 |
| 2021 | QBLKe: Host-side flash translation layer management for Open-Channel SSDs
Hongwei Qin, Dan Feng 0001, Wei Tong 0001, Mengye Peng, Jingning Liu |
J. Syst. Archit. | 1 |
| 2019 | Fully Quantized Network for Object DetectionabstractEfficient neural network inference is important in a number of practical domains, such as deployment in mobile settings. An effective method for increasing inference efficiency is to use low bitwidth arithmetic, which can subsequently be accelerated using dedicated hardware. However, designing effective quantization schemes while maintaining network accuracy is challenging. In particular, current techniques face difficulty in performing fully end-to-end quantization, making use of aggressively low bitwidth regimes such as 4-bit, and applying quantized networks to complex tasks such as object detection. In this paper, we demonstrate that many of these difficulties arise because of instability during the fine-tuning stage of the quantization process, and propose several novel techniques to overcome these instabilities. We apply our techniques to produce fully quantized 4-bit detectors based on RetinaNet and Faster R-CNN, and show that these achieve state-of-the-art performance for quantized detectors. The mAP loss due to quantization using our methods is more than 3.8x less than the loss from existing methods. Yan Wang 0080, Hongwei Qin |
CVPR | 4 |
| 2019 | QBLK: Towards Fully Exploiting the Parallelism of Open-Channel SSDsabstractBy exposing physical channels to host software, Open-Channel SSD shows great potential in future high performance storage systems. However, the existing scheme fails to achieve acceptable performance under heavy workloads. The main reasons reside not only in its single-buffer architecture, more importantly, but also in its line-based physical address management. Besides, the lock of address mapping table is also a performance burden under heavy workloads. We propose QBLK, an open source driver which tries to better exploit the parallelism of Open-Channel SSDs. Particularly, QBLK adopts four key techniques, namely (1) Multi-queue based buffering, (2) Per-channel based address management, (3) Lock-free address mapping, and (4) Fine-grained draining. Experimental results show that QBLK achieves up to 97.4% bandwidth improvement compared with the state-of-the-art PBLK scheme. Hongwei Qin, Dan Feng 0001, Wei Tong 0001, Jingning Liu |
DATE | 1 |
| 2019 | CeSR: A Cell State Remapping Strategy to Reduce Raw Bit Error Rate of MLC NAND FlashabstractThe following topics are dealt with: storage management; flash memories; parallel processing; cache storage; cloud computing; meta data; data compression; data handling; optimisation; learning (artificial intelligence). Wei Tong 0001, Jingning Liu, Dan Feng 0001, Hongwei Qin |
MSST | 5 |
| 2018 | Quantization Mimic: Towards Very Tiny CNN for Object Detection
Hongwei Qin, Wanli Ouyang |
ECCV (8) | 3 |
| 2017 | Scale-Aware Face DetectionabstractConvolutional neural network (CNN) based face detectors are inefficient in handling faces of diverse scales. They rely on either fitting a large single model to faces across a large scale range or multi-scale testing. Both are computationally expensive. We propose Scale-aware Face Detection (SAFD) to handle scale explicitly using CNN, and achieve better performance with less computation cost. Prior to detection, an efficient CNN predicts the scale distribution histogram of the faces. Then the scale histogram guides the zoom-in and zoom-out of the image. Since the faces will be approximately in uniform scale after zoom, they can be detected accurately even with much smaller CNN. Actually, more than 99% of the faces in AFW can be covered with less than two zooms per image. Extensive experiments on FDDB, MALF and AFW show advantages of SAFD. Zekun Hao, Yu Liu 0015, Hongwei Qin, Xiu Li 0001, Xiaolin Hu 0001 |
CVPR | 3 |
| 2017 | Depth Estimation by Parameter Transfer With a Lightweight Model for Single Still ImagesabstractIn this paper, we propose a novel method for automatic depth estimation from color images using parameter transfer. By modeling the correlation between color images and their depth maps with a set of parameters, we get a database of parameter sets. Given an input image, we extract the high-level features to find the best matched image sets from the database. Then the set of parameters corresponding to the best match are used to estimate the depth of the input image. Compared with the past learning-based methods, our trained model consists only of trained features and parameter sets, which occupy little space. We evaluate our depth estimation method on several benchmark RGB-D (RGB + depth) data sets. The experimental results are comparable to the state-of-the-art results, while the model size is very small and very suitable for mobile devices, demonstrating the promising performance of our proposed method. Hongwei Qin, Xiu Li 0001, Yangang Wang 0001, Yongbing Zhang 0002, Qionghai Dai |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2016 | Joint Training of Cascaded CNN for Face DetectionabstractCascade has been widely used in face detection, where classifier with low computation cost can be firstly used to shrink most of the background while keeping the recall. The cascade in detection is popularized by seminal Viola-Jones framework and then widely used in other pipelines, such as DPM and CNN. However, to our best knowledge, most of the previous detection methods use cascade in a greedy manner, where previous stages in cascade are fixed when training a new stage. So optimizations of different CNNs are isolated. In this paper, we propose joint training to achieve end-to-end optimization for CNN cascade. We show that the back propagation algorithm used in training CNN can be naturally used in training CNN cascade. We present how jointly training can be conducted on naive CNN cascade and more sophisticated region proposal network (RPN) and fast R-CNN. Experiments on face detection benchmarks verify the advantages of the joint training. Hongwei Qin, Xiu Li 0001, Xiaolin Hu 0001 |
CVPR | 1 |
| 2016 | DeepFish: Accurate underwater live fish recognition with a deep architecture
Hongwei Qin, Xiu Li 0001, Jian Liang 0002, YiGang Peng, Changshui Zhang |
Neurocomputing | 1 |
| 2016 | A modified simulated annealing algorithm and an excessive area model for floorplanning using fixed-outline constraintsabstractOutline-free floorplanning focuses on area and wirelength reductions, which are usually meaningless, since they can hardly satisfy modern design requirements. We concentrate on a more difficult and useful issue, fixed-outline floorplanning. This issue imposes fixed-outline constraints on the outline-free floorplanning, making the physical design more interesting and challenging. The contributions of this paper are primarily twofold. First, a modified simulated annealing (MSA) algorithm is proposed. In the beginning of the evolutionary process, a new attenuation formula is used to decrease the temperature slowly, to enhance MSA’s global searching capacity. After a period of time, the traditional attenuation formula is employed to decrease the temperature rapidly, to maintain MSA’s local searching capacity. Second, an excessive area model is designed to guide MSA to find feasible solutions readily. This can save much time for refining feasible solutions. Additionally, B*-tree representation is known as a very useful method for characterizing floorplanning. Therefore, it is employed to perform a perturbing operation for MSA. Finally, six groups of benchmark instances with different dead spaces and aspect ratios—circuits n10, n30, n50, n100, n200, and n300—are chosen to demonstrate the efficiency of our proposed method on fixed-outline floorplanning. Compared to several existing methods, the proposed method is more efficient in obtaining desirable objective function values associated with the chip area, wirelength, and fixed-outline constraints. De-Xuan Zou, Gaige Wang, Gai Pan, Hongwei Qin |
Frontiers Inf. Technol. Electron. Eng. | 4 |
| 2014 | DEPT: Depth Estimation by Parameter Transfer for Single Still Images
Xiu Li 0001, Hongwei Qin, Yangang Wang 0001, Yongbing Zhang 0002, Qionghai Dai |
ACCV (2) | 2 |