EDBT 2026 Demo / reviewers in the wild / expert
Qiang Wang 0023
dblp:64/5630-23
· DBLP profile ↗
41ranked-venue papers
8as first author
19since 2021 · last 2026
0000-0001-5632-4408ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 31 · 5 first-author · 11 since 2021Artificial intelligence and machine learning · 23 · 3 first-author · 14 since 2021Systems, architecture and hardware · 5 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-authorHuman-computer interaction and ubiquitous computing · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | EC-MVSNet: Enhanced Cascaded Multi-View Stereo with Cross-Scale Relevance IntegrationabstractCascade-based multi-scale architectures are currently the mainstream in Multi-view Stereo (MVS), achieving a balance between computational efficiency and reconstruction accuracy. However, existing cascade MVS methods suffer from significant limitations in cross-scale information utilization, where depth estimation processes operate independently across scales without fully exploiting the rich relevance between adjacent scales. To address this fundamental limitation, we propose an Enhanced Cascade Multi-View Stereo framework (EC-MVSNet), which introduces a novel cross-scale relevance integration strategy. Specifically, we introduce a Cross-Scale Feature-based Joint Construction (CFC) module to synergistically combine features from adjacent scales to build more reliable cost volumes. Additionally, a Cross-Scale Probability-guided Enhancement (CPE) module is proposed to propagate depth probability distributions across scales to guide cost volume enhancement. Furthermore, we propose a Monocular Feature-based Refinement (MFR) module to further enhance depth prediction accuracy by leveraging monocular priors. Extensive experiments demonstrate that EC-MVSNet achieves state-of-the-art performance on multiple benchmarks, validating the effectiveness of the cross-scale integration in improving MVS reconstruction quality. Shaoqian Wang, Jiadai Sun, Bin Fan 0002, Qiang Wang 0023, Yuchao Dai |
AAAI | 4 |
| 2024 | Efficient Learning on Successive Test Time AugmentationabstractTest time augmentation (TTA) has been a promising tool for improving the robustness against out-of-distribution data at inference time. Recent TTA methods try to learn predictive transformations which are supposed to provide the best performance gain on each test sample. However, existing methods are either restricted to predicting one single transformation for each sample or require multiple forward passes of the transformation predictor, leading to a sub-optimal solution regarding efficiency. In this paper, we propose a novel method to predict successive test time augmentations. For the first time, it only requires a single forward pass of the transformation predictor, while can output multiple desired transformations iteratively. The experimental results show that our method provides a significant and consistent improvement in model robustness against various corruptions while significantly surpassing state-of-the-arts in runtime. Siyang Pan, Jiaqian Yu, Qiang Wang 0023, ByungIn Yoo |
ICASSP | 6 |
| 2024 | Gradtrans: Transformer-Based Gradient Guidance for Image GenerationabstractImage generation has been attracting widespread attention in recent years along with the development of generative models. Existing works mostly focus on pursuing high-quality generated samples as a priority. In this work, we introduce a lightweight transformer-based module, called GradTrans, that provides a novel balance on the speed-performance trade-off with generative adversarial networks for image generation. GradTrans effectively leverages the instructive information in the discriminator network to guide the generator network for a higher generation quality at the inference stage without overburdening the cost. Extensive experiments are conducted for unconditional image generation task and style transfer task on diverse datasets, including CIFAR10, STL10 and Horse2Zebra, demonstrating that our proposed GradTrans can surpass different related methods with significantly superior performance, as well as being generalizable with large compatibility to different base models. Jiaqian Yu, Siyang Pan, Sangil Jung, Wu Bi, Seung In Park, Qiang Wang 0023, ByungIn Yoo |
ICIP | 7 |
| 2024 | DVI-SLAM: A Dual Visual Inertial SLAM NetworkabstractRecent deep learning based visual simultaneous localization and mapping (SLAM) methods have made significant progress. However, how to make full use of visual information as well as better integrate with inertial measurement unit (IMU) in visual SLAM has potential research value. This paper proposes a novel deep SLAM network with dual visual factors. The basic idea is to integrate both photometric factor and re-projection factor into the end-to-end differentiable structure through multi-factor data association module. We show that the proposed network dynamically learns and adjusts the confidence maps of both visual factors and it can be further extended to include the IMU factors as well. Extensive experiments validate that our proposed method significantly outperforms the state-of-the-art methods on several public datasets, including TartanAir, EuRoC and ETH3D-SLAM. Specifically, when dynamically fusing the three factors together, the absolute trajectory error for both monocular and stereo configurations on EuRoC dataset has reduced by 45.3% and 36.2% respectively. Xiongfeng Peng, SoonYong Cho, Qiang Wang 0023 |
ICRA | 6 |
| 2023 | TrajectoryFormer: 3D Object Tracking Transformer with Predictive Trajectory Hypothesesabstract3D multi-object tracking (MOT) is vital for many applications including autonomous driving vehicles and service robots. With the commonly used tracking-by-detection paradigm, 3D MOT has made important progress in recent years. However, these methods only use the detection boxes of the current frame to obtain trajectory-box association results, which makes it impossible for the tracker to recover objects missed by the detector. In this paper, we present TrajectoryFormer, a novel point-cloud-based 3D MOT framework. To recover the missed object by detector, we generates multiple trajectory hypotheses with hybrid candidate boxes, including temporally predicted boxes and current-frame detection boxes, for trajectory-box association. The predicted boxes can propagate object’s history trajectory information to the current frame and thus the network can tolerate short-term miss detection of the tracked objects. We combine long-term object motion feature and short-term object appearance feature to create per-hypothesis feature embedding, which reduces the computational overhead for spatial-temporal encoding. Additionally, we introduce a Global-Local Interaction Module to conduct information interaction among all hypotheses and models their spatial relations, leading to accurate estimation of hypotheses. Our TrajectoryFormer achieves state-of-the-art performance on the Waymo 3D MOT benchmarks. Code is available at https://github.com/poodarchu/EFG. Xuesong Chen 0001, Shaoshuai Shi, Benjin Zhu, Qiang Wang 0023, Ka Chun Cheung, Simon See, Hongsheng Li 0001 |
ICCV | 5 |
| 2023 | BadTrack: A Poison-Only Backdoor Attack on Visual Object TrackingabstractVisual object tracking (VOT) is one of the most fundamental tasks in computer vision community. State-of-the-art VOT trackers extract positive and negative examples that are used to guide the tracker to distinguish the object from the background. In this paper, we show that this characteristic can be exploited to introduce new threats and hence propose a simple yet effective poison-only backdoor attack. To be specific, we poison a small part of the training data by attaching a predefined trigger pattern to the background region of each video frame, so that the trigger appears almost exclusively in the extracted negative examples. To the best of our knowledge, this is the first work that reveals the threat of poison-only backdoor attack on VOT trackers. We experimentally show that our backdoor attack can significantly degrade the performance of both two-stream Siamese and one-stream Transformer trackers on the poisoned data while gaining comparable performance with the benign trackers on the clean data. Jiaqian Yu, Siyang Pan, Qiang Wang 0023 |
NeurIPS | 5 |
| 2023 | Prototype-guided Instance matching for multiple pedestrian tracking
Qiang Wang 0023, Wankou Yang, Chunyan Xu, Zhen Cui 0001 |
Neurocomputing | 1 |
| 2023 | LiDAR-only 3D object detection based on spatial context
Qiang Wang 0023, Dejun Zhu, Wankou Yang |
J. Vis. Commun. Image Represent. | 1 |
| 2023 | Fast Monocular Depth Estimation via Side Prediction Aggregation with Continuous Spatial RefinementabstractRecent works have validated the benefit of integrating spatial information into deep networks to improve pixel-level prediction tasks such as monocular depth estimation. However, how to efficiently and robustly integrate spatial cues retains as an open problem. In this paper, we introduce the Side Prediction Aggregation (termed SPA) method to enhance the embedding of scene structural information from low-level to high-level layers. To improve the estimation accuracy, the proposed method is further equipped with continuous Spatial Refinement Loss (termed SRL) at multiple resolutions with negligible extra computation. Besides, the proposed sequential network can further perform adversarial learning at multiple resolutions. Such an adversarial refinement strategy greatly improves the accuracy of estimated depth with a little extra computation. Without using any pre-trained models, our network achieves the the-state-of-art accuracy on KITTI, NYUD V2, and Cityscapes datasets, which has achieved real-time depth estimation online. Jipeng Wu, Rongrong Ji, Qiang Wang 0023, Shengchuan Zhang, Xiaoshuai Sun, Yan Wang 0059, Mingliang Xu 0001, Feiyue Huang |
IEEE Trans. Multim. | 3 |
| 2022 | Task Generalizable Spatial and Texture Aware Image Downsizing Network
Lin Ma 0002, Hongsheng Li 0001, Qiang Wang 0023 |
BMVC | 4 |
| 2022 | Learning a Structured Latent Space for Unsupervised Point Cloud CompletionabstractUnsupervised point cloud completion aims at estimating the corresponding complete point cloud of a partial point cloud in an unpaired manner. It is a crucial but challenging problem since there is no paired partial-complete supervision that can be exploited directly. In this work, we pro-pose a novel framework, which learns a unified and structured latent space that encoding both partial and complete point clouds. Specifically, we map a series of related par-tial point clouds into multiple complete shape and occlusion code pairs and fuse the codes to obtain their repre-sentations in the unified latent space. To enforce the learning of such a structured latent space, the proposed method adopts a series of constraints including structured ranking regularization, latent code swapping constraint, and distribution supervision on the related partial point clouds. By establishing such a unified and structured latent space, better partial-complete geometry consistency and shape completion accuracy can be achieved. Extensive experi-ments show that our proposed method consistently outper-forms state-of-the-art unsupervised methods on both syn-thetic ShapeNet and real-world KITTI, ScanNet, and Mat- terport3D datasets. Yingjie Cai, Kwan-Yee Lin, Qiang Wang 0023, Xiaogang Wang 0001, Hongsheng Li 0001 |
CVPR | 4 |
| 2022 | FlowFormer: A Transformer Architecture for Optical Flow
Xiaoyu Shi 0002, Qiang Wang 0023, Ka Chun Cheung, Hongwei Qin, Jifeng Dai, Hongsheng Li 0001 |
ECCV (17) | 4 |
| 2022 | DH-LC: Hierarchical Matching and Hybrid Bundle Adjustment Towards Accurate and Robust Loop ClosureabstractA loop closure module plays an important role in visual SLAM systems, which can reduce the accumulat-ed drift. This task faces the challenges of large viewpoint changes and expensive computational costs when optimizing the global map. This paper proposes DH-LC, a novel accurate and robust loop closure method that consists of hierarchical spatial feature matching (HSFM) and hybrid bundle adjustment (HBA). HSFM estimates a reliable relative pose between the query image and the retrieval image in a coarse-to-fine way. Specifically, 3D points are firstly triangulated and then clus-tered according to the spatial distribution. The cluster centers estimate coarse cube-level matching pairs in a larger perception field which can tolerate large viewpoint changes. HBA optimizes the global map efficiently by adaptively selecting incremental bundle adjustment or full bundle adjustment according to the accumulated drift and relative pose verification in the temporal window. Experimental results demonstrate that our proposed method easily detects loops in large viewpoint changes and efficiently optimizes the global map. When compared with the state-of-the-art methods, our method increases loop closure recall and improves SLAM localization accuracy with reducing the accumulated drift. Xiongfeng Peng, Qiang Wang 0023, Yun-Tae Kim |
IROS | 3 |
| 2022 | Attention-guided RGB-D Fusion Network for Category-level 6D Object Pose EstimationabstractThis work focuses on estimating 6D poses and sizes of category-level objects from a single RGB-D image. How to exploit the complementary RGB and depth features plays an important role in this task yet remains an open question. Due to the large intra-category texture and shape variations, an object instance in test may have different RGB and depth features from those of the object instances in training, which poses challenges to previous RGB-D fusion methods. To deal with such problem, an Attention-guided RGB-D Fusion Network (ARF-Net) is proposed in this work. Our key design is an ARF module that learns to adaptively fuse RGB and depth features with guidance from both structure-aware attention and relation-aware attention. Specifically, the structure-aware attention captures spatial relationship among object parts and the relation-aware attention captures the RGB-to-depth correlations between the appearance and geometric features. Our ARF -Net directly establishes canonical correspondences with a compact decoder based on the multi-modal features from our ARF module. Extensive experiments show that our method can effectively fuse RGB features to various popular point cloud encoders and provide consistent performance improvement. In particular, without reconstructing instance 3D models, our method with its relatively compact architecture outperforms all state-of-the-art models on CAMERA25 and REAL275 benchmarks by a large margin. Hao Wang 0144, Jiyeon Kim, Qiang Wang 0023 |
IROS | 4 |
| 2021 | UASNet: Uncertainty Adaptive Sampling Network for Deep Stereo MatchingabstractRecent studies have shown that cascade cost volume can play a vital role in deep stereo matching to achieve high resolution depth map with efficient hardware usage. However, how to construct good cascade volume as well as effective sampling for them are still under in-depth study. Previous cascade-based methods usually perform uniform sampling in a predicted disparity range based on variance, which easily misses the ground truth disparity and decreases disparity map accuracy. In this paper, we propose an uncertainty adaptive sampling network (UASNet) featuring two modules: an uncertainty distribution-guided range prediction (URP) model and an uncertainty-based disparity sampler (UDS) module. The URP explores the more discriminative uncertainty distribution to handle the complex matching ambiguities and to improve disparity range prediction. The UDS adaptively adjusts sampling interval to localize disparity with improved accuracy. With the proposed modules, our UASNet learns to construct cascade cost volume and predict full-resolution disparity map directly. Extensive experiments show that the proposed method achieves the highest ground truth covering ratio compared with other cascade cost volume based stereo matching methods. Our method also achieves top performance on both SceneFlow dataset and KITTI benchmark. Yamin Mao, Yuchao Dai, Qiang Wang 0023, Yun-Tae Kim, Hong-Seok Lee |
ICCV | 5 |
| 2021 | Learning Generalized Intersection Over Union for Dense Pixelwise PredictionabstractIntersection over union (IoU) score, also named Jaccard Index, is one of the most fundamental evaluation methods in machine learning. The original IoU computation cannot provide non-zero gradients and thus cannot be directly optimized by nowadays deep learning methods. Several recent works generalized IoU for bounding box regression, but they are not straightforward to adapt for pixelwise prediction. In particular, the original IoU fails to provide effective gradients for the non-overlapping and location-deviation cases, which results in performance plateau. In this paper, we propose PixIoU, a generalized IoU for pixelwise prediction that is sensitive to the distance for non-overlapping cases and the locations in prediction. We provide proofs that PixIoU holds many nice properties as the original IoU. To optimize the PixIoU, we also propose a loss function that is proved to be submodular, hence we can apply the Lovász functions, the efficient surrogates for submodular functions for learning this loss. Experimental results show consistent performance improvements by learning PixIoU over the original IoU for several different pixelwise prediction tasks on Pascal VOC, VOT-2020 and Cityscapes. Jiaqian Yu, Jingtao Xu, Qiang Wang 0023, ByungIn Yoo, Jae-Joon Han |
ICML | 5 |
| 2021 | Accurate Visual-Inertial SLAM by Feature Re-identificationabstractMost of the state-of-the-art visual inertial SLAM methods pay less attention to 2D-2D and 3D-2D matching with more reliable features in a long time span, which easily results in continuous estimation drift. In this paper, we propose an efficient drift-free visual-inertial SLAM method by a pose guided feature matching method to re-identify existing features from a spatial-temporal sensitive sub-global map. The re-identified features serve as augmented visual measurements to anchor the current frame and gradually decrease the accumulated error in the long run. When incorporating the measurements into the optimization module, it benefits to build a drift-free global map in the system. Extensive experiments show that our feature re-identification method is both effective and efficient. Specifically, when combining the feature re-identification with the state-of-the-art SLAM method [1], our method achieves 67.3% and 87.5% absolute trajectory error reduction with only a small additional computational cost on two public SLAM benchmark DBs: EuRoC and TUM-VI respectively. Xiongfeng Peng, Qiang Wang 0023, Yun-Tae Kim, Myungjae Jeon, Hong-Seok Lee |
IROS | 3 |
| 2021 | Accurate Visual-Inertial SLAM by Manhattan Frame Re-identificationabstractMost of the state-of-the-art visual-inertial SLAM methods pay less attention to the scene structure of man-made environments. In this paper, based on the assumption of multiple local Manhattan worlds (MWs), we propose a Manhattan frame (MF) re-identification method to build relative rotation constraints between MF matching pairs and tightly couple these constraints into global bundle adjust module. Specifically, a coarse-to-fine vanishing point (VP) estimation method and pose guided MF temporal consistency verification method are firstly proposed to improve the accuracy and robustness of MF estimation. Then unreliable MF matching pairs are filtered out by a spatial temporal consistency check. Finally, the relative rotation constraints of the remaining MF matching pairs are combined into global bundle adjustment energy function for further optimization. We have validated our proposed method on both synthetic and real-world datasets. When comparing with the baseline method [1], the real-time absolute trajectory error (ATE) of our proposed method has decreased by 29.1%, 19.8% on TartanAir hospital and EuRoC datasets respectively. Our method also exceeds existing state-of-the-art algorithms on both synthetic and real-world datasets. Xiongfeng Peng, Qiang Wang 0023, Yun-Tae Kim, Hong-Seok Lee |
IROS | 3 |
| 2021 | An Interconnected Feature Pyramid Networks for object detection
Qiang Wang 0023, Lukuan Zhou, Yuncong Yao, Yong Wang 0032, Jun Li 0033, Wankou Yang |
J. Vis. Commun. Image Represent. | 1 |
| 2020 | Synthetic Depth Transfer for Monocular 3D Object Pose Estimation in the WildabstractMonocular object pose estimation is an important yet challenging computer vision problem. Depth features can provide useful information for pose estimation. However, existing methods rely on real depth images to extract depth features, leading to its difficulty on various applications. In this paper, we aim at extracting RGB and depth features from a single RGB image with the help of synthetic RGB-depth image pairs for object pose estimation. Specifically, a deep convolutional neural network is proposed with an RGB-to-Depth Embedding module and a Synthetic-Real Adaptation module. The embedding module is trained with synthetic pair data to learn a depth-oriented embedding space between RGB and depth images optimized for object pose estimation. The adaptation module is to further align distributions from synthetic to real data. Compared to existing methods, our method does not need any real depth images and can be trained easily with large-scale synthetic data. Extensive experiments and comparisons show that our method achieves best performance on a challenging public PASCAL 3D+ dataset in all the metrics, which substantiates the superiority of our method and the above modules. Yueying Kao, Qiang Wang 0023, Zhouchen Lin, Wooshik Kim, Sunghoon Hong |
AAAI | 3 |
| 2019 | Referring Expression Comprehension with Semantic Visual Relationship and Word MappingabstractReferring expression comprehension, which locates the object instance described by a natural language expression, gains increasing interests in recent years. This paper aims at improving the task from two aspects: visual feature extraction and language features extraction. For visual feature extraction, we observe that most of the previous methods utilize only relative spatial information to model the visual relationship between object pairs while discarding rich semantic relationship between objects. This makes the visual-language matching difficult when the language expression contains semantic relationship to discriminate the referred object from other objects in the image. In this work, we propose a Semantic Visual Relationship Module (SVRM) to exploit this important information. For language feature extraction, a major problem comes from the long-tail distribution of words in the expressions. Since more than half of the words appear less than 20 times in the public datasets, deep models such as LSTM tend to fail to learn accurate representations for these words. To solve this problem, we propose a word2vec based word mapping method that maps these low frequency words to high frequency words with similar meaning. Experiments show that the proposed method outperforms existing state-of-the-art methods on three referring expression comprehension datasets. Wanli Ouyang, Qiang Wang 0023, Woo-Shik Kim, Sunghoon Hong |
ACM Multimedia | 4 |
| 2018 | An Appearance-and-Structure Fusion Network for Object Viewpoint EstimationabstractAutomatic object viewpoint estimation from a single image is an important but challenging problem in machine intelligence community. Although impressive performance has been achieved, current state-of-the-art methods still have difficulty to deal with the visual ambiguity and structure ambiguity in real world images. To tackle these problems, a novel Appearance-and-Structure Fusion network, which we call it ASFnet that estimates viewpoint by fusing both appearance and structure information, is proposed in this paper. The structure information is encoded by precise semantic keypoints and can help address the visual ambiguity. Meanwhile, distinguishable appearance features contribute to overcoming the structure ambiguity. Our ASFnet integrates an appearance path and a structure path to an end-to-end network and allows deep features effectively share supervision from both the two complementary aspects. A convolutional layer is learned to fuse the two path results adaptively. To balance the influence from the two supervision sources, a piecewise loss weight strategy is employed during training. Experimentally, our proposed network outperforms state-of-the-art methods on a public PASCAL 3D+ dataset, which verifies the effectiveness of our method and further corroborates the above proposition. Yueying Kao, Zairan Wang, Dongqing Zou, Qiang Wang 0023, Minsu Ahn, Sunghoon Hong |
IJCAI | 6 |
| 2018 | HCR-Net: A Hybrid of Classification and Regression Network for Object Pose EstimationabstractObject pose estimation from a single image is a fundamental and challenging problem in computer vision and robotics. Generally, current methods treat pose estimation as a classification or a regression problem. However, regression based methods usually suffer from the issue of imbalanced training data, while classification methods are difficult to discriminate nearby poses. In this paper, a hybrid CNN model, which we call it HCR-Net that integrates both a classification network and a regression network, is proposed to deal with these issues. Our model is inspired by that regression methods can get better accuracy on homogeneously distributed datasets while classification methods are more effective for coarse quantization of the poses even if the dataset is not well balanced. The classification methods and the regression methods essentially complement each other. Thus we integrate both them into a neural network in a hybrid fashion and train it end-to-end with two novel loss functions. As a result, our method surpass the state-of-the-art methods, even with imbalanced training data and much less data augmentation. The experimental results on the challenging Pascal3D+ database demonstrate that our method outperforms the state-of-the-arts significantly, achieving improvements on ACC and AVP metrics up to 4% and 6%, respectively. Zairan Wang, Yueying Kao, Dongqing Zou, Qiang Wang 0023, Minsu Ahn, Sunghoon Hong |
IJCAI | 5 |
| 2017 | Adaptive Temporal Pooling for Object Detection using Dynamic Vision Sensor
Wei-Heng Liu, Dongqing Zou, Qiang Wang 0023, Paul K. J. Park, Hyunsurk Ryu |
BMVC | 5 |
| 2017 | Robust Dense Depth Maps Generations from Sparse DVS Stereos
Dongqing Zou, Wei-Heng Liu, Qiang Wang 0023, Paul K. J. Park, Hyunsurk Ryu |
BMVC | 5 |
| 2016 | Performance improvement of deep learning based gesture recognition using spatiotemporal demosaicing techniqueabstractWe propose a novel method for the demosaicing of event-based images that offers substantial performance improvement of far-distance gesture recognition based on deep Convolutional Neural Network. Unlike the conventional demosaicing technique using the spatial color interpolation of Bayer patterns, our new approach utilizes spatiotemporal correlation between pixel arrays, whereby timestamps of high-resolution pixels are efficiently generated in real-time from the event data. In this paper, we describe this new method and evaluate its performance with a hand motion recognition task. Paul K. J. Park, Baek Hwan Cho, Jin Man Park, Kyoobin Lee, Ha Young Kim, Hyo Ah Kang, Hyun Goo Lee, Jooyeon Woo, Yohan Roh, Won Jo Lee, Chang-Woo Shin, Qiang Wang 0023, Hyunsurk Ryu |
ICIP | 12 |
| 2016 | Context-aware event-driven stereo matchingabstractSimilarity measuring plays as an import role in stereo matching, whether for visual data from standard cameras or for those from novel sensors such as Dynamic Vision Sensors (DVS). Generally speaking, robust feature descriptors contribute to designing a powerful similarity measurement, as demonstrated by classic stereo matching methods. However, the kind and representative ability of feature descriptors for DVS data are so limited that achieving accurate stereo matching on DVS data becomes very challenging. In this paper, a novel feature descriptor is proposed to improve the accuracy for DVS stereo matching. Our feature descriptor can describe the local context or distribution of the DVS data, contributing to constructing an effective similarity measurement for DVS data matching, yielding an accurate stereo matching result. Our method is evaluated by testing our method on groundtruth data and comparing with various standard stereo methods. Experiments demonstrate the efficiency and effectiveness of our method. Dongqing Zou, Qiang Wang 0023, Xiaotao Wang, Guangqi Shao, Paul K. J. Park |
ICIP | 3 |
| 2015 | Real-time human body parts localization from dynamic vision sensorabstractDynamic vision sensor (DVS) as a novel type of visual sensors can detect a moving object in a fast and cost effective way by outputting events on edges of the object. This paper proposes a body part localization method using structured output Deep Belief Network (s-DBN) to label the body parts in block of pixels in very fast fashion. Experiments show that our proposed algorithm achieves pixel accuracy 90.13% on body parts localization compared to Deep Belief Network (87.01%) and Random Forests (84.15%) under the same computational cost. For head/hand detection s-DBN has significant better accuracy of 99.3%/87.8% compared to DBN 98.7%/81.7% and RF 97.1%/47.1% under recall rate 99%/90%. Specifically, the process time on a 240×180 sized image is less than 1ms on Intel Core2 2.83GHZ CPU. Wentao Mao, Qiang Wang 0023, Xiaotao Wang, Shandong Wang, Guangqi Shao, Kyoobin Lee, Paul K. J. Park |
ICIP | 2 |
| 2013 | Learning a Structured Graphical Model with Boosted Top-Down Features for Ultrasound Image Segmentation
Zhihui Hao, Qiang Wang 0023, Xiaotao Wang, Jung-Bae Kim, Youngkyoo Hwang, Baek Hwan Cho, Won Ki Lee |
MICCAI (1) | 2 |
| 2012 | Multiscale superpixel classification for tumor segmentation in breast ultrasound imagesabstractTumor localization and segmentation in breast ultrasound (BUS) images is an important as well as intractable problem for computer-aided diagnosis (CAD) due to the high variation in shape and appearance. We propose a novel algorithm in this paper without making any assumption on tumor, compared to most previous works. Heterogeneous features are collected via a hierarchical over-segmentation framework, which we have shown has the multiscale property. The superpixels are then classified with their confidences nested into the bottom layer. The ultimate segmentation is made by using an efficient conditional random field model. Experiments on challenging data set show that our algorithm is able to handle almost all kinds of benign and malignant tumors, and also confirm the superiority of our work through a comparison with other two different approaches. Zhihui Hao, Qiang Wang 0023, Haibing Ren, Kuanhong Xu, Yeong Kyeong Seong, Ji-yeun Kim |
ICIP | 2 |
| 2012 | Combining CRF and Multi-hypothesis Detection for Accurate Lesion Segmentation in Breast Sonograms
Zhihui Hao, Qiang Wang 0023, Yeong Kyeong Seong, Jong-Ha Lee 0001, Haibing Ren, Ji-yeun Kim |
MICCAI (1) | 2 |
| 2009 | Performance driven face animation via non-rigid 3d trackingabstractIn this demo, a performance driven 3D face animation system is proposed. The proposed system consists of two key components: a robust non-rigid 3D tracking module and a MPEG4 compliant facial animation module. Firstly, the facial motion is tracked from source videos which contain both the rigid 3D head motion (6 DOF) and the non-rigid expression variations. Afterward, the tracked facial motion is parameterized via estimating a set of MPEG4 facial animation parameters(FAP). As the final step, these FAP values are transferred to the MPEG4-compliant face model for the animation purpose. The proposed tracking and animation system has a strong generalization ability and can be used in the indoor environment with no additional assumptions. Wayne Zhang 0001, Qiang Wang 0023, Xiaoou Tang |
ACM Multimedia | 2 |
| 2008 | Real Time Feature Based 3-D Deformable Face Tracking
Wayne Zhang 0001, Qiang Wang 0023, Xiaoou Tang |
ECCV (2) | 2 |
| 2006 | Real-Time Bayesian 3-D Pose TrackingabstractIn this paper, we propose a novel approach for real-time 3-D tracking of object pose from a single camera. We formulate the 3-D pose tracking task in a Bayesian framework which fuses feature correspondence information from both previous frame and some selected key-frames into the posterior distribution of pose. We also developed an inter-frame motion inference algorithm which can get reliable inter-frame feature correspondences and relative pose. Finally, the maximum a posteriori estimation of pose is obtained via stochastic sampling to achieve stable and drift-free tracking. Experiments show significant improvement of our algorithm over existing algorithms especially in the cases of tracking agile motion, severe occlusion, drastic illumination change, and large object scale change Qiang Wang 0023, Xiaoou Tang, Harry Shum |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2005 | Patch Based Blind Image Super ResolutionabstractIn this paper, a novel method for learning based image super resolution (SR) is presented. The basic idea is to bridge the gap between a set of low resolution (LR) images and the corresponding high resolution (HR) image using both the SR reconstruction constraint and a patch based image synthesis constraint in a general probabilistic framework. We show that in this framework, the estimation of the LR image formation parameters is straightforward. The whole framework is implemented via an annealed Gibbs sampling method. Experiments on SR on both single image and image sequence input show that the proposed method provides an automatic and stable way to compute super-resolution and the achieved result is encouraging for both synthetic and real LR images. Qiang Wang 0023, Xiaoou Tang, Harry Shum |
ICCV | 1 |
| 2005 | Automatic 3D Face Modeling from VideoabstractIn this paper, we develop an efficient technique for fully automatic recovery of accurate 3D face shape from videos captured by a low cost camera. The method is designed to work with a short video containing a face rotating from frontal view to profile view. The whole approach consists of three components. First, automatic initialization is performed in the first frame with approximately frontal face. Then, to handle the case of low quality image captured by low cost camera, the 2D feature matching, head poses and underlying 3D face shape are estimated and refined iteratively in an efficient way based on image sequence segmentation. Finally, to take advantage of the sparse structure of the proposed algorithm, sparse bundle adjustment technique is further employed to speed up the computation. We demonstrate the accuracy and robustness of the algorithm using a set of experiments Le Xin, Qiang Wang 0023, Jianhua Tao 0001, Xiaoou Tang, Tieniu Tan, Harry Shum |
ICCV | 2 |
| 2004 | Learning-Based Tracking of Complex Non-Rigid Motion
Qiang Wang 0023, Haizhou Ai, Guangyou Xu |
J. Comput. Sci. Technol. | 1 |
| 2003 | Learning Object Intrinsic Structure for Robust Visual TrackingabstractIn this paper, a novel method to learn the intrinsic object structure for robust visual tracking is proposed. The basic assumption is that the parameterized object state lies on a low dimensional manifold and can be learned from training data. Based on this assumption, firstly we derived the dimensionality reduction and density estimation algorithm for unsupervised learning of object intrinsic representation, the obtained non-rigid part of object state reduces even to 2 dimensions. Secondly the dynamical model is derived and trained based on this intrinsic representation. Thirdly the learned intrinsic object structure is integrated into a particle-filter style tracker. We will show that this intrinsic object representation has some interesting properties and based on which the newly derived dynamical model makes particle-filter style tracker more robust and reliable. Experiments show that the learned tracker performs much better than existing trackers on the tracking of complex non-rigid motions such as fish twisting with self-occlusion and large inter-frame lip motion. The proposed method also has the potential to solve other type of tracking problems. Qiang Wang 0023, Guangyou Xu, Haizhou Ai |
CVPR (2) | 1 |
| 2002 | Robust pose estimation for 3D face modeling from stereo sequencesabstractProposes a robust pose estimation algorithm from 2D correspondences, which is a key issue of a 3D face modeling system from calibrated stereo sequences. The estimated rigid motion parameters are utilized to obtain the perspective projection of a generic face model, which is then matched with 2D clues extracted from the image under corresponding pose to decide the shape of a specified face. The main merits of our method are: (1) In order to obtain robust and accurate results under the situation of dramatic pose variation, we first evaluate the reliability of 2D tracker. Then after eliminating erroneous 2D correspondences, we refine the rigid motion parameters estimated between successive poses by performing a non-linear, batch estimator to compute the parameters of all poses in a clip of stereo sequences simultaneously. (2) Full automaticity is achieved by detecting and matching new features when there are not enough reliable 2D tracking results. Experiments show that this algorithm is accurate and robust, and help our system reach a satisfactory face modeling result. Guangyou Xu, Qiang Wang 0023 |
ICIP (3) | 4 |
| 2002 | A Probabilistic Dynamic Contour Model for Accurate and Robust Lip TrackingabstractIn this paper a new condensation style contour tracking method called probabilistic dynamic contour (PDC) is proposed for lip tracking: a novel mixture dynamic model is designed to represent shape more compactly and to tolerate larger motions between frames, a measurement model is designed to include multiple visual cues. The proposed PDC tracker has the advantage that it is conceptually general but effectively suitable for lip tracking with the designed dynamic and measurement model. The new tracker improves the traditional condensation style tracker in three aspects: Firstly, the dynamic model is partially derived from the image sequence, so the tracker does not need to learn the dynamics in advance. Secondly, the measurement model is easy to be updated during tracking, which avoids modeling the foreground object in prior. Thirdly, to improve the tracker's speed, a compact representation of shape and a noise model are proposed to reduce the samples required to represent the posterior distribution. An experiment on lip contour tracking shows that the proposed method tracks contours robustly as well as accurately compared to the existing tracking method. Qiang Wang 0023, Haizhou Ai, Guangyou Xu |
ICMI | 1 |
| 1999 | Automating the Construction of Dynamic and Multi-Resolution 360° Panorama for Natural Scenes with Moving ObjectsabstractA new approach is presented to automatically build a dynamic and multi-resolution 360/spl deg/ panorama (DMP) from image sequences taken by a hand-held camera. A multi-resolution representation is built for the more interesting areas by means of camera zooming. The dynamic objects in the scene can be detected and represented separately. The DMP construction method is fast, robust and automatic, achieving 1 Hz in a 266 MHz PC. Zhigang Zhu 0001, Guangyou Xu, Qiang Wang 0023 |
VR | 4 |