EDBT 2026 Demo / reviewers in the wild / expert
Feng Wang 0015
dblp:90/4225-15
· DBLP profile ↗
28ranked-venue papers
4as first author
15since 2021 · last 2025
0000-0002-7669-2547ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 19 · 3 first-author · 9 since 2021Artificial intelligence and machine learning · 16 · 1 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 2Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Animation Anycolor: Enhancing Line Drawing Colorization with Keypoint MatchingabstractColorization is a crucial but labor-intensive and time-consuming process of animation production. The automation of animation line-drawing colorization has become a prominent research topic. Recently, methods based on pre-trained text-to-image models have been explored for the task of line-drawing colorization. However, these approaches may result in colorization errors when dealing with complex situations such as positional changes and extensive motion commonly encountered in animated scenes. These issues are primarily attributed to the inadequate semantic correspondence between the reference images and line-drawings. To tackle this problem, we introduce the Animation Anycolor framework. This approach leverages spatial attention mechanisms to integrate the appearance features of reference images, ensuring consistent feature transmission. Furthermore, we present a novel technique that employs keypoint matching to explicitly direct the network to recognize the feature correspondence areas between reference and target images, thus effectively mitigating color confusion. Our method preserves the accuracy and naturalness of color results in scenes characterized by positional shifts and character movement. Comparative evaluations indicate that our method outperforms the baseline by an average of 11.3% on the FID metric. Notably, this method improves the efficiency of line-drawing colorization and reduces production costs. It also introduces new insights by combining in-context correspondences with knowledge from the pre-trained model. This approach has broad application prospects in the animation industry. Liyao Wang, Zuzeng Lin, Danni Wu, Suzhe Zhang, Zixian Wu, Feng Wang 0015 |
ICASSP | 7 |
| 2025 | CycleVAR: Repurposing Autoregressive Model for Unsupervised One-Step Image TranslationabstractThe current conditional autoregressive image generation methods have shown promising results, yet their potential remains largely unexplored in the practical unsupervised image translation domain, which operates without explicit cross-domain correspondences. A critical limitation stems from the discrete quantization inherent in traditional Vector Quantization-based frameworks, which disrupts gradient flow between the Variational Autoencoder decoder and causal Transformer, impeding end-to-end optimization during adversarial training in image space. To tackle this issue, we propose using Softmax Relaxed Quantization, a novel approach that reformulates codebook selection as a continuous probability mixing process via Softmax, thereby preserving gradient propagation. Building upon this differentiable foundation, we introduce CycleVAR, which reformulates image-to-image translation as image-conditional visual autoregressive generation by injecting multi-scale source image tokens as contextual prompts, analogous to prefix-based conditioning in language models. CycleVAR exploits two modes to generate the target image tokens, including (1) serial multi-step generation, enabling iterative refinement across scales, and (2) parallel one-step generation synthesizing all resolution outputs in a single forward pass. Experimental findings indicate that the parallel one-step generation mode attains superior translation quality with quicker inference speed than the serial multi-step mode in unsupervised scenarios. Furthermore, both quantitative and qualitative results indicate that CycleVAR surpasses previous state-of-the-art unsupervised image translation models, \textit{e}.\textit{g}., CycleGAN-Turbo. Shengqian Li, Zuzeng Lin, Feng Wang 0015 |
ICCV | 4 |
| 2025 | LayerAnimate: Layer-Level Control for Animation
Yuxue Yang, Lue Fan, Zuzeng Lin, Feng Wang 0015, Zhaoxiang Zhang 0001 |
ICCV | 4 |
| 2025 | AnimeColor: Reference-based Animation Colorization with Diffusion TransformersabstractAnimation colorization plays a vital role in animation production, yet existing methods struggle to achieve color accuracy and temporal consistency. To address these challenges, we propose AnimeColor, a novel reference-based animation colorization framework leveraging Diffusion Transformers (DiT). Our approach integrates sketch sequences into a DiT-based video diffusion model, enabling sketch-controlled animation generation. We introduce two key components: a High-level Color Extractor (HCE) to capture semantic color information and a Low-level Color Guider (LCG) to extract fine-grained color details from reference images. These components work synergistically to guide the video diffusion process. Additionally, we employ a multi-stage training strategy to maximize the utilization of reference image color information. Extensive experiments demonstrate that AnimeColor outperforms existing methods in color accuracy, sketch alignment, temporal consistency, and visual quality. Our framework not only advances the state of the art in animation colorization but also provides a practical solution for industrial applications. The code will be made publicly available at https://github.com/IamCreateAI/AnimeColor. Liyao Wang, Danni Wu, Zuzeng Lin, Feng Wang 0015, Li Song 0001 |
ACM Multimedia | 6 |
| 2025 | FSD V2: Improving Fully Sparse 3D Object Detection With Virtual VoxelsabstractLiDAR-based fully sparse architecture has gained increasing attention. FSDv1 stands out as a representative work, achieving impressive efficacy and efficiency, albeit with intricate structures and handcrafted designs. In this paper, we present FSDv2, an evolution that aims to simplify the previous FSDv1 and eliminate the ad-hoc heuristics in its handcrafted instance-level representation, thus promoting better universality. To this end, we introduce virtual voxels, taking over the clustering-based instance segmentation in FSDv1. Virtual voxels not only address the notorious issue of the Center Feature Missing in fully sparse detectors but also endow the framework with a more elegant and streamlined approach. Besides, we develop a suite of components to complement the virtual voxel mechanism, including a virtual voxel encoder, a virtual voxel mixer, and a virtual voxel assignment strategy. We conduct experiments on three large-scale datasets: Waymo Open Dataset, Argoverse 2 dataset, and nuScenes dataset. Our results showcase state-of-the-art performance on all three datasets, highlighting the superiority of FSDv2 in long-range scenarios and its universality in achieving competitive performance across diverse scenarios. Moreover, we provide comprehensive experimental analysis to understand the workings of FSDv2. To facilitate further research, we have open-sourced the full code at https://github.com/tusen-ai/SST. Lue Fan, Feng Wang 0015, Naiyan Wang, Zhaoxiang Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | Efficient Deformable ConvNets: Rethinking Dynamic and Sparse Operator for Vision ApplicationsabstractWe introduce Deformable Convolution v4 (DCNv4), a highly efficient and effective operator designed for a broad spectrum of vision applications. DCNv4 addresses the limitations of its predecessor, DCNv3, with two key enhancements: 1. removing softmax normalization in spatial aggregation to enhance its dynamic property and expressive power and 2. optimizing memory access to minimize redundant operations for speedup. These improvements result in a significantly faster convergence compared to DCNv3 and a substantial increase in processing speed, with DCNv4 achieving more than three times the forward speed. DCNv4 demonstrates exceptional performance across various tasks, including image classification, instance and semantic segmentation, and notably, image generation. When integrated into generative models like U-Net in the latent diffusion model, DCNv4 outperforms its baseline, underscoring its possibility to enhance generative models. In practical applications, replacing DCNv3 with DCNv4 in the InternImage model to create FlashInternImage results in up to 80% speed increase and further performance improvement without further modifications. The advancements in speed and efficiency of DCNv4, combined with its robust performance across diverse vision tasks, show its potential as a foundational building block for future vision models. Yuwen Xiong, Yuntao Chen, Feng Wang 0015, Xizhou Zhu, Jiapeng Luo, Wenhai Wang, Tong Lu 0002, Hongsheng Li 0001, Yu Qiao 0001, Lewei Lu, Jie Zhou 0001, Jifeng Dai |
CVPR | 4 |
| 2024 | Frame Fusion with Vehicle Motion Prediction for 3D Object DetectionabstractIn LiDAR-based 3D detection, history point clouds contain rich temporal information helpful for future prediction. In the same way, history detections should contribute to future detections. In this paper, we propose a detection enhancement method, namely FrameFusion, which improves 3D object detection results by fusing history detection frames. In FrameFusion, we "forward" history frames to the current frame and apply weighted Non-Maximum-Suppression on dense bounding boxes to obtain a fused frame with merged boxes. To "forward" frames, we use vehicle motion models to estimate the future pose of the bounding boxes. Our method is flexible in motion model selection. We explore three motion models in our work and show how the unicycle model and the bicycle model improve turning cases. On Waymo Open Dataset, our FrameFusion method consistently improves the performance of various 3D detectors by about 2.0 vehicle LEVEL 2 APH with negligible latency and slightly enhances the performance of the temporal fusion method MPPNet. We also conduct extensive experiments on motion model selection. Feng Wang 0015, Naiyan Wang, Chao Ma 0001 |
ICRA | 2 |
| 2023 | Once Detected, Never Lost: Surpassing Human Performance in Offline LiDAR based 3D Object DetectionabstractThis paper aims for high-performance offline LiDAR-based 3D object detection. We first observe that experienced human annotators annotate objects from a track-centric perspective. They first label objects in a track with clear shapes, and then leverage the temporal coherence to infer the annotations of obscure objects. Drawing inspiration from this, we propose a high-performance offline detector in a track-centric perspective instead of the conventional object-centric perspective. Our method features a bidirectional tracking module and a track-centric learning module. Such design allows our detector to infer and refine a complete track once the object is detected at a certain moment. We refer this characteristic to "onCe detecTed, neveR Lost" and name the proposed system CTRL. Extensive experiments demonstrate the remarkable performance of our method, surpassing the human-level annotating accuracy and previous state-of-the-art methods in the highly competitive Waymo Open Dataset leaderboard without model ensemble. The code is available at https://github.com/tusen-ai/SST. Lue Fan, Yuxue Yang, Yiming Mao 0008, Feng Wang 0015, Yuntao Chen, Naiyan Wang, Zhaoxiang Zhang 0001 |
ICCV | 4 |
| 2023 | Echoes Beyond Points: Unleashing the Power of Raw Radar Data in Multi-modality FusionabstractRadar is ubiquitous in autonomous driving systems due to its low cost and good adaptability to bad weather. Nevertheless, the radar detection performance is usually inferior because its point cloud is sparse and not accurate due to the poor azimuth and elevation resolution. Moreover, point cloud generation algorithms already drop weak signals to reduce the false targets which may be suboptimal for the use of deep fusion. In this paper, we propose a novel method named EchoFusion to skip the existing radar signal processing pipeline and then incorporate the radar raw data with other sensors. Specifically, we first generate the Bird's Eye View (BEV) queries and then take corresponding spectrum features from radar to fuse with other sensors. By this approach, our method could utilize both rich and lossless distance and speed clues from radar echoes and rich semantic clues from images, making our method surpass all existing methods on the RADIal dataset, and approach the performance of LiDAR. The code will be released on https://github.com/tusen-ai/EchoFusion. Yang Liu 0347, Feng Wang 0015, Naiyan Wang, Zhaoxiang Zhang 0001 |
NeurIPS | 2 |
| 2023 | Super Sparse 3D Object DetectionabstractAs the perception range of LiDAR expands, LiDAR-based 3D object detection contributes ever-increasingly to the long-range perception in autonomous driving. Mainstream 3D object detectors often build dense feature maps, where the cost is quadratic to the perception range, making them hardly scale up to the long-range settings. To enable efficient long-range detection, we first propose a fully sparse object detector termed FSD. FSD is built upon the general sparse voxel encoder and a novel sparse instance recognition (SIR) module. SIR groups the points into instances and applies highly-efficient instance-wise feature extraction. The instance-wise grouping sidesteps the issue of the center feature missing, which hinders the design of the fully sparse architecture. To further enjoy the benefit of fully sparse characteristic, we leverage temporal information to remove data redundancy and propose a super sparse detector named FSD++. FSD++ first generates residual points, which indicate the point changes between consecutive frames. The residual points, along with a few previous foreground points, form the super sparse input data, greatly reducing data redundancy and computational overhead. We comprehensively analyze our method on the large-scale Waymo Open Dataset, and state-of-the-art performance is reported. To showcase the superiority of our method in long-range detection, we also conduct experiments on Argoverse 2 Dataset, where the perception range ([Formula: see text] m) is much larger than Waymo Open Dataset ([Formula: see text] m). Lue Fan, Yuxue Yang, Feng Wang 0015, Naiyan Wang, Zhaoxiang Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | Embracing Single Stride 3D Object Detector with Sparse TransformerabstractIn LiDAR-based 3D object detection for autonomous driving, the ratio of the object size to input scene size is significantly smaller compared to 2D detection cases. Over-looking this difference, many 3D detectors directly follow the common practice of 2D detectors, which downsample the feature maps even after quantizing the point clouds. In this paper, we start by rethinking how such multi-stride stereotype affects the LiDAR-based 3D object detectors. Our experiments point out that the downsampling operations bring few advantages, and lead to inevitable information loss. To remedy this issue, we propose Single-stride Sparse Transformer (SST) to maintain the original resolution from the beginning to the end of the network. Armed with transformers, our method addresses the problem of insufficient receptive field in single-stride architectures. It also cooperates well with the sparsity of point clouds and naturally avoids expensive computation. Eventually, our SST achieves state-of-the-art results on the large-scale Waymo Open Dataset. It is worth mentioning that our method can achieve exciting performance (83.8 LEVEL_1 AP on validation split) on small object (pedestrian) detection due to the characteristic of single stride. Our codes will be public soon. Lue Fan, Ziqi Pang, Tianyuan Zhang 0002, Yu-Xiong Wang, Hang Zhao 0021, Feng Wang 0015, Naiyan Wang, Zhaoxiang Zhang 0001 |
CVPR | 6 |
| 2022 | Fully Sparse 3D Object DetectionabstractAs the perception range of LiDAR increases, LiDAR-based 3D object detection becomes a dominant task in the long-range perception task of autonomous driving. The mainstream 3D object detectors usually build dense feature maps in the network backbone and prediction head. However, the computational and spatial costs on the dense feature map are quadratic to the perception range, which makes them hardly scale up to the long-range setting. To enable efficient long-range LiDAR-based object detection, we build a fully sparse 3D object detector (FSD). The computational and spatial cost of FSD is roughly linear to the number of points and independent of the perception range. FSD is built upon the general sparse voxel encoder and a novel sparse instance recognition (SIR) module. SIR first groups the points into instances and then applies instance-wise feature extraction and prediction. In this way, SIR resolves the issue of center feature missing, which hinders the design of the fully sparse architecture for all center-based or anchor-based detectors. Moreover, SIR avoids the time-consuming neighbor queries in previous point-based methods by grouping points into instances. We conduct extensive experiments on the large-scale Waymo Open Dataset to reveal the working mechanism of FSD, and state-of-the-art performance is reported. To demonstrate the superiority of FSD in long-range detection, we also conduct experiments on Argoverse 2 Dataset, which has a much larger perception range ($200m$) than Waymo Open Dataset ($75m$). On such a large perception range, FSD achieves state-of-the-art performance and is 2.4$\times$ faster than the dense counterpart. Codes will be released. Lue Fan, Feng Wang 0015, Naiyan Wang, Zhaoxiang Zhang 0001 |
NeurIPS | 2 |
| 2022 | Imbalanced regression for intensity series of pain expression from videos by regularizing spatio-temporal face nets
Xiang Xiang 0001, Feng Wang 0015, Yuwen Tan, Alan L. Yuille |
Pattern Recognit. Lett. | 2 |
| 2021 | LiDAR R-CNN: An Efficient and Universal 3D Object DetectorabstractLiDAR-based 3D detection in point cloud is essential in the perception system of autonomous driving. In this paper, we present LiDAR R-CNN, a second stage detector that can generally improve any existing 3D detector. To fulfil-l the real-time and high precision requirement in practice, we resort to point-based approach other than the popular voxel-based approach. However, we find an overlooked issue in previous work: Naively applying point-based methods like PointNet could make the learned features ignore the size of proposals. To this end, we analyze this problem in detail and propose several methods to remedy it, which bring significant performance improvement. Comprehensive experimental results on real-world datasets like Waymo Open Dataset (WOD) and KITTI dataset with various popular detectors demonstrate the universality and superiority of our LiDAR R-CNN. In particular, based on one variant of PointPillars, our method could achieve new state-of-the-art results with minor cost. Codes will be released at https://github.com/tusimple/LiDAR_RCNN. Feng Wang 0015, Naiyan Wang |
CVPR | 2 |
| 2021 | RangeDet: In Defense of Range View for LiDAR-based 3D Object DetectionabstractIn this paper, we propose an anchor-free single-stage LiDAR-based 3D object detector – RangeDet. The most notable difference with previous works is that our method is purely based on the range view representation. Compared with the commonly used voxelized or Bird’s Eye View (BEV) representations, the range view representation is more compact and without quantization error. Although there are works adopting it for semantic segmentation, its performance in object detection is largely behind voxelized or BEV counterparts. We first analyze the existing range-view-based methods and find two issues overlooked by previous works: 1) the scale variation between nearby and far away objects; 2) the inconsistency between the 2D range image coordinates used in feature extraction and the 3D Cartesian coordinates used in output. Then we deliberately design three components to address these issues in our RangeDet. We test our RangeDet in the large-scale Waymo Open Dataset (WOD). Our best model achieves 72.9/75.9/65.8 3D AP on vehicle/pedestrian/cyclist. These results outperform other range-view-based methods by a large margin, and are overall comparable with the state-of-the-art multi-view-based methods. Codes will be released at https://github.com/TuSimple/RangeDet. Lue Fan, Xuan Xiong, Feng Wang 0015, Naiyan Wang, Zhaoxiang Zhang 0001 |
ICCV | 3 |
| 2019 | Semantic Segmentation of High Resolution Remote Sensing Image Based on Batch-Attention MechanismabstractDeep convolution neural network has been widely used in recent works for semantic segmentation of High Resolution Remote Sensing(HRRS) images. Because of the limitation of GPU memory, HRRS images are usually split into several sub-images for training convolutional neural networks. For each sub-image, the segmentation model may not have enough information to predict the segmentation map very well. In order to alleviate this problem, we propose to apply a batch-attention module to capture the discriminative information from similar objects, which come from other sub-images in a mini-batch. We also utilize global attention upsample module as the decoder to provide global context and fuse high and low level information better. We evaluate our model on the Potsdam dataset and achieve 88.30% pixAcc and 73.78% mIoU. Yanzhou Su, Feng Wang 0015, Jian Cheng 0003 |
IGARSS | 4 |
| 2019 | A Discriminatively Learned CNN Embedding For Remote Sensing Image Scene ClassificationabstractIn this work, a discriminatively learned CNN embedding is proposed for remote sensing image scene classification. Our proposed siamese network simultaneously computes the classification loss function and the metric learning loss function of the two input images. Specifically, for the classification loss, we use the standard cross-entropy loss function to predict the classes of the images. For the metric learning loss, our siamese network learns to map the intra-class and inter-class input pairs to a feature space where intra-class inputs are close and inter-class inputs are separated by a margin. Concretely, for remote sensing image scene classification, we would like to map images from the same scene to feature vectors that are close, and map images from different scenes to feature vectors that are widely separated. Experiments are conducted on three different remote sensing image datasets to evaluate the effectiveness of our proposed approach. The results demonstrate that the proposed method achieves an excellent classification performance. Wen Wang 0012, Lijun Du, Yinxing Gao, Yanzhou Su, Feng Wang 0015, Jian Cheng 0003 |
IGARSS | 5 |
| 2018 | Visualizing deep neural network by alternately image blurring and deblurring
Feng Wang 0015, Haijun Liu 0001, Jian Cheng 0003 |
Neural Networks | 1 |
| 2018 | Additive Margin Softmax for Face VerificationabstractIn this letter, we propose a conceptually simple and intuitive learning objective function, i.e., additive margin softmax, for face verification. In general, face verification tasks can be viewed as metric learning problems, even though lots of face verification models are trained in classification schemes. It is possible when a large-margin strategy is introduced into the classification model to encourage intraclass variance minimization. As one alternative, angular softmax has been proposed to incorporate the margin. In this letter, we introduce another kind of margin to the softmax loss function, which is more intuitive and interpretable. Experiments on LFW and MegaFace show that our algorithm performs better when the evaluation criteria are designed for very low false alarm rate. Feng Wang 0015, Jian Cheng 0003, Weiyang Liu, Haijun Liu 0001 |
IEEE Signal Process. Lett. | 1 |
| 2018 | Sequential Subspace Clustering via Temporal Smoothness for Sequential Data SegmentationabstractThis paper develops a novel sequential subspace clustering method for sequential data. Inspired by the state-of-the-art methods, ordered subspace clustering, and temporal subspace clustering, we design a novel local temporal regularization term based on the concept of temporal predictability. Through minimizing the short-term variance on historical data, it can recover the temporal smoothness relationships in sequential data. Moreover, we claim that the local temporal regularization is more important than the global structural regularization for a specific task, such as sequential subspace clustering, which leads to a concise minimization objective function. To solve the bi-convex objective function, a simple and efficient optimization algorithm based on the alternate convex search method is devised to jointly learn the coding matrix and the dictionary. Furthermore, five baseline methods are also devised for comparison with our proposed method from different aspects. Extensive experimental results and comparisons with the state-of-the-art methods on three data sets demonstrate the effectiveness of the proposed temporal smoothness sequential subspace clustering method for sequential data. Haijun Liu 0001, Jian Cheng 0003, Feng Wang 0015 |
IEEE Trans. Image Process. | 3 |
| 2018 | Pedestrian Detection via Body Part Semantic and Contextual Information With DNNabstractPedestrian detection has achieved great improve-ments in recent years, while complex occlusion handling and high-accurate localization are still the most important problems. To take advantage of the body part semantic information and the contextual information for pedestrian detection, we propose the part and context network (PCN) in this paper. A PCN is composed of three branches: the basic branch; the part branch; and the context branch. It specially utilizes two branches to detect the pedestrians through the body part semantic information and the contextual information, respectively. In the part branch, the semantic information of body parts can communicate with each other via long short-term memory (LSTM). In the context branch, we adopt a local competition mechanism (maxout) for adaptive context scale selection. By combining the outputs of all branches, we develop a strong complementary pedestrian detector with a lower miss rate and higher localization accuracy, especially for the occlusion pedestrian. The combination of the body part semantic information and the contextual information in pedestrian detection is fully explored in this paper. Comprehensive evaluations on three challenging pedestrian detection datasets (i.e., Caltech, INRIA and KITTI) well demonstrate the effectiveness of our proposed PCN. Code for PCN is publicly available on GitHub https://github.com/sunnyxiaohu/pcn_pedestrian. Shiguang Wang, Jian Cheng 0003, Haijun Liu 0001, Feng Wang 0015, Hui Zhou 0005 |
IEEE Trans. Multim. | 4 |
| 2017 | Sequential Subspace Clustering via Temporal SmoothnessabstractThis paper develops a novel sequential subspace clustering method for sequential data. Inspired by state-of-the-art methods ordered subspace clustering (OSC) and temporal subspace clustering (TSC), we design a novel local temporal regularization term based on the concept of temporal predictability, which is measured by short-term variance against long-term variance, to recover the temporal smoothness relationships in sequential data. To solve the bi-convex objective function, a simple and efficient optimization algorithm based on the alternate convex search (ACS) method is devised to jointly learn the codings matrix and dictionary. Extensive experimental results and comparisons with state-of-the-art methods on gesture and face datasets demonstrate the effectiveness of the proposed temporal smoothness sequential subspace clustering method for sequential data. Haijun Liu 0001, Jian Cheng 0003, Feng Wang 0015 |
FG | 3 |
| 2017 | Kinship verification based on status-aware projection learningabstractKinship verification for parent-child is considered to be an asymmetric metric process, in which parents and children are associated with different status where the parents are priorly known to be significantly older than the children. To address the asymmetric metric learning, a status-aware projection learning (SaPL) method is proposed for facial image-based kinship verification, especially for the parent-child kinship. SaPL learns two status-specific projections to capture the significant appearance commonality between parents and children, respectively. Each status-specific projection consists of two components: a common component shared by the two status projections and a status-specific component. SaPL generally outperforms the one Mahalanobis distance metric. Extensive experimental results and comparisons with state-of-the-art approaches demonstrate the effectiveness of the proposed SaPL for kinship verification. Haijun Liu 0001, Jian Cheng 0003, Feng Wang 0015 |
ICIP | 3 |
| 2017 | Regularizing face verification nets for pain intensity regressionabstractLimited labeled data are available for the research of estimating facial expression intensities. For instance, the ability to train deep networks for automated pain assessment is limited by small datasets with labels of patient-reported pain intensities. Fortunately, fine-tuning from a data-extensive pre-trained domain, such as face verification, can alleviate this problem. In this paper, we propose a network that fine-tunes a state-of-the-art face verification network using a regularized regression loss and additional data with expression labels. In this way, the expression intensity regression task can benefit from the rich feature representations trained on a huge amount of data for face verification. The proposed regularized deep regressor is applied to estimate the pain expression intensity and verified on the widely-used UNBC-McMaster Shoulder-Pain dataset, achieving the state-of-the-art performance. A weighted evaluation metric is also proposed to address the imbalance issue of different pain intensities. Feng Wang 0015, Xiang Xiang 0001, Trac D. Tran, Austin Reiter, Gregory D. Hager, Harry Quon, Jian Cheng 0003, Alan L. Yuille |
ICIP | 1 |
| 2017 | Supervised hashing with jointly learning embedding and quantizationabstractCompared with unsupervised hashing, supervised hashing commonly illustrates better accuracy in many real applications by leveraging semantic (label) information. However, it is tough to solve the supervised hashing problem directly because it is essentially a discrete optimization problem. Some other works try to solve the discrete optimization problem directly using binary quadratic programming, but they are typically too complicated and time-consuming while some supervised hashing methods have to solve a relaxed continuous optimization problem by dropping the discrete constraints. However, these methods typically suffer from poor performance due to the errors caused by the relaxation manner. In this paper based on the general two-step framework: learning binary embedded codes and learning hash functions, we propose a new method to solve the problem introduced by relaxing the cost function. Inspired by the property of rotation invariance of learning embedding features, our method tries to jointly learn similarity-preserving representation and rotation transformation for better quantization alternatively. In experiments, our method shows significant improvement. Compared with the methods based on discrete optimization our methods obtains the competitive performance and even achieves the state-of-the-art performance in some image retrieval applications. Feng Wang 0015, Xiang Xiang 0001, Trac D. Tran |
ICIP | 2 |
| 2017 | MAT: A Multimodal Attentive Translator for Image CaptioningabstractIn this work we formulate the problem of image captioning as a multimodal translation task. Analogous to machine translation, we present a sequence-to-sequence recurrent neural networks (RNN) model for image caption generation. Different from most existing work where the whole image is represented by convolutional neural network (CNN) feature, we propose to represent the input image as a sequence of detected objects which feeds as the source sequence of the RNN model. In this way, the sequential representation of an image can be naturally translated to a sequence of words, as the target sequence of the RNN model. To represent the image in a sequential way, we extract the objects features in the image and arrange them in a order using convolutional neural networks. To further leverage the visual information from the encoded objects, a sequential attention layer is introduced to selectively attend to the objects that are related to generate corresponding words in the sentences. Extensive experiments are conducted to validate the proposed approach on popular benchmark dataset, i.e., MS COCO, and the proposed model surpasses the state-of-the-art methods in all metrics following the dataset splits of previous work. The proposed approach is also evaluated by the evaluation server of MS COCO captioning challenge, and achieves very competitive results, e.g., a CIDEr of 1.029 (c5) and 1.064 (c40). Fuchun Sun 0001, Changhu Wang, Feng Wang 0015, Alan L. Yuille |
IJCAI | 4 |
| 2017 | NormFace: L2 Hypersphere Embedding for Face VerificationabstractThanks to the recent developments of Convolutional Neural Networks, the performance of face verification methods has increased rapidly. In a typical face verification method, feature normalization is a critical step for boosting performance. This motivates us to introduce and study the effect of normalization during training. But we find this is non-trivial, despite normalization being differentiable. We identify and study four issues related to normalization through mathematical analysis, which yields understanding and helps with parameter settings. Based on this analysis we propose two strategies for training using normalized features. The first is a modification of softmax loss, which optimizes cosine similarity instead of inner-product. The second is a reformulation of metric learning by introducing an agent vector for each class. We show that both strategies, and small variants, consistently improve performance by between 0.2% to 0.4% on the LFW dataset based on two models. This is significant because the performance of the two models on LFW dataset is close to saturation at over 98%. Feng Wang 0015, Xiang Xiang 0001, Jian Cheng 0003, Alan L. Yuille |
ACM Multimedia | 1 |
| 2015 | Silhouette Analysis for Human Action Recognition Based on Supervised Temporal t-SNE and Incremental LearningabstractThis paper develops a human action recognition method for human silhouette sequences based on supervised temporal t-stochastic neighbor embedding (ST-tSNE) and incremental learning. Inspired by the SNE and its variants, ST-tSNE is proposed to learn the underlying relationship between action frames in a manifold, where the class label information and temporal information are introduced to well represent those frames from the same action class. As to the incremental learning, an important step for action recognition, we introduce three methods to perform the low-dimensional embedding of new data. Two of them are motivated by local methods, locally linear embedding and locality preserving projection. Those two techniques are proposed to learn explicit linear representations following the local neighbor relationship, and their effectiveness is investigated for preserving the intrinsic action structure. The rest one is based on manifold-oriented stochastic neighbor projection to find a linear projection from high-dimensional to low-dimensional space capturing the underlying pattern manifold. Extensive experimental results and comparisons with the state-of-the-art methods demonstrate the effectiveness and robustness of the proposed ST-tSNE and incremental learning methods in the human action silhouette analysis. Jian Cheng 0003, Haijun Liu 0001, Feng Wang 0015, Hongsheng Li 0001, Ce Zhu |
IEEE Trans. Image Process. | 3 |