EDBT 2026 Demo / reviewers in the wild / expert
Yan Wang 0059
dblp:59/2227-59
· DBLP profile ↗
43ranked-venue papers
8as first author
17since 2021 · last 2024
0000-0003-4309-3166ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 30 · 4 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 26 · 6 first-author · 11 since 2021Systems, architecture and hardware · 2Human-computer interaction and ubiquitous computing · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Emotion Recognition in Conversation via Dynamic PersonalityabstractEmotion recognition in conversation (ERC) is a field that aims to classify the emotion of each utterance within conversational contexts. This presents significant challenges, particularly in handling emotional ambiguity across various speakers and contextual factors. Existing ERC approaches have primarily focused on modeling conversational contexts while incorporating only superficial speaker attributes such as names, memories, and interactions. Recent works introduce personality as an essential deep speaker factor for emotion recognition, but relies on static personality, overlooking dynamic variability during conversations. Advances in personality psychology conceptualize personality as dynamic, proposing that personality states can change across situations. In this paper, we introduce ERC-DP, a novel model considering the dynamic personality of speakers during conversations. ERC-DP accounts for past utterances from the same speaker as situation impacting dynamic personality. It combines personality modeling with prompt design and fine-grained classification modules. Through a series of comprehensive experiments, ERC-DP demonstrates superior performance on three benchmark conversational datasets. Yan Wang 0059, Bo Wang 0011, Yachao Zhao, Xiaojia Jin, Jijun Zhang, Ruifang He, Yuexian Hou |
LREC/COLING | 1 |
| 2024 | A Comparative Study of Explicit and Implicit Gender Biases in Large Language Models via Self-evaluationabstractWhile extensive work has examined the explicit and implicit biases in large language models (LLMs), little research explores the relation between these two types of biases. This paper presents a comparative study of the explicit and implicit biases in LLMs grounded in social psychology. Social psychology distinguishes between explicit and implicit biases by whether the bias can be self-recognized by individuals. Aligning with this conceptualization, we propose a self-evaluation-based two-stage measurement of explicit and implicit biases within LLMs. First, the LLM is prompted to automatically fill templates with social targets to measure implicit bias toward these targets, where the bias is less likely to be self-recognized by the LLM. Then, the LLM is prompted to self-evaluate the templates filled by itself to measure explicit bias toward the same targets, where the bias is more likely to be self-recognized by the LLM. Experiments conducted on state-of-the-art LLMs reveal human-like inconsistency between explicit and implicit occupational gender biases. This work bridges a critical gap where prior studies concentrate solely on either explicit or implicit bias. We advocate that future work highlight the relation between explicit and implicit biases in LLMs. Yachao Zhao, Bo Wang 0011, Yan Wang 0059, Xiaojia Jin, Jijun Zhang, Ruifang He, Yuexian Hou |
LREC/COLING | 3 |
| 2024 | RepAn: Enhanced Annealing through Re-parameterizationabstractThe simulated annealing algorithm aims to improve model convergence through multiple restarts of training. However, existing annealing algorithms overlook the cor-relation between different cycles, neglecting the potential for incremental learning. We contend that a fixed network structure prevents the model from recognizing distinct features at different training stages. To this end, we propose RepAn, redesigning the irreversible re-parameterization (Rep) method and integrating it with annealing to enhance training. Specifically, the network goes through Rep, ex-pansion, restoration, and backpropagation operations during training, and iterating through these processes in each annealing round. Such a method exhibits good generalization and is easy to apply, and we provide theoretical expla-nations for its effectiveness. Experiments demonstrate that our method improves baseline performance by 6.38% on the CIFAR-100 dataset and 2.80% on ImageNet, achieving state-of-the-art performance in the Rep field. The code is available at https://github.com/xfey/RepAn. Xiawu Zheng, Yan Wang 0059, Fei Chao 0001, Chenglin Wu 0001, Liujuan Cao |
CVPR | 3 |
| 2024 | Asynchronous Large Language Model Enhanced Planner for Autonomous Driving
Ziqin Wang, Yan Wang 0059, Si Liu 0001 |
ECCV (36) | 4 |
| 2024 | LaMI-DETR: Open-Vocabulary Detection with Language Model Instruction
Penghui Du, Yifan Sun 0003, Luting Wang 0001, Yue Liao, Errui Ding, Yan Wang 0059, Jingdong Wang 0001, Si Liu 0001 |
ECCV (23) | 8 |
| 2023 | OMPQ: Orthogonal Mixed Precision QuantizationabstractTo bridge the ever-increasing gap between deep neural networks' complexity and hardware capability, network quantization has attracted more and more research attention. The latest trend of mixed precision quantization takes advantage of hardware's multiple bit-width arithmetic operations to unleash the full potential of network quantization. However, existing approaches rely heavily on an extremely time-consuming search process and various relaxations when seeking the optimal bit configuration. To address this issue, we propose to optimize a proxy metric of network orthogonality that can be efficiently solved with linear programming, which proves to be highly correlated with quantized model accuracy and bit-width. Our approach significantly reduces the search time and the required data amount by orders of magnitude, but without a compromise on quantization accuracy. Specifically, we achieve 72.08% Top-1 accuracy on ResNet-18 with 6.7Mb parameters, which does not require any searching iterations. Given the high efficiency and low data dependency of our algorithm, we use it for the post-training quantization, which achieves 71.27% Top-1 accuracy on MobileNetV2 with only 1.5Mb parameters. Yuexiao Ma, Taisong Jin, Xiawu Zheng, Yan Wang 0059, Huixia Li, Yongjian Wu 0001, Guannan Jiang, Wei Zhang 0217, Rongrong Ji |
AAAI | 4 |
| 2023 | Meta Architecture for Point Cloud AnalysisabstractRecent advances in 3D point cloud analysis bring a diverse set of network architectures to the field. However, the lack of a unified framework to interpret those networks makes any systematic comparison, contrast, or analysis challenging, and practically limits healthy development of the field. In this paper, we take the initiative to explore and propose a unified framework called PointMeta, to which the popular 3D point cloud analysis approaches could fit. This brings three benefits. First, it allows us to compare different approaches in a fair manner, and use quick experiments to verify any empirical observations or assumptions summarized from the comparison. Second, the big picture brought by PointMeta enables us to think across different components, and revisit common beliefs and key design decisions made by the popular approaches. Third, based on the learnings from the previous two analyses, by doing simple tweaks on the existing approaches, we are able to derive a basic building block, termed PointMetaBase. It shows very strong performance in efficiency and effectiveness through extensive experiments on challenging benchmarks, and thus verifies the necessity and benefits of high-level interpretation, contrast, and comparison like PointMeta. In particular, PointMetaBase surpasses the previous state-of-the-art method by 0.7%/1.4/%2.1% mIoU with only 2%/11%/13% of the computation cost on the S3DIS datasets. The code and models are available at https://github.com/linhaojia13/PointMetaBase. Haojia Lin, Xiawu Zheng, Lijiang Li, Fei Chao 0001, Shanshan Wang 0002, Yan Wang 0059, Yonghong Tian 0001, Rongrong Ji |
CVPR | 6 |
| 2023 | Automatic Network Pruning via Hilbert-Schmidt Independence Criterion Lasso under Information Bottleneck PrincipleabstractMost existing neural network pruning methods hand-crafted their importance criteria and structures to prune. This constructs heavy and unintended dependencies on heuristics and expert experience for both the objective and the parameters of the pruning approach. In this paper, we try to solve this problem by introducing a principled and unified framework based on Information Bottleneck (IB) theory, which further guides us to an automatic pruning approach. Specifically, we first formulate the channel pruning problem from an IB perspective, and then implement the IB principle by solving a Hilbert-Schmidt Independence Criterion (HSIC) Lasso problem under certain conditions. Based on the theoretical guidance, we then provide an automatic pruning scheme by searching for global penalty coefficients. Verified by extensive experiments, our method yields state-of-the-art performance on various benchmark networks and datasets. For example, with VGG-16, we achieve a 60%-FLOPs reduction by removing 76% of the parameters, with an improvement of 0.40% in top-1 accuracy on CIFAR-10. With ResNet-50, we achieve a 56%-FLOPs reduction by removing 50% of the parameters, with a small loss of 0.08% in the top-1 accuracy on ImageNet. The code is available at https://github.com/sunggo/APIB. Song Guo 0001, Lei Zhang 0001, Xiawu Zheng, Yan Wang 0059, Fei Chao 0001, Chenglin Wu 0001, Shengchuan Zhang, Rongrong Ji |
ICCV | 4 |
| 2023 | DDPNAS: Efficient Neural Architecture Search via Dynamic Distribution Pruning
Xiawu Zheng, Chenyi Yang 0002, Yan Wang 0059, Baochang Zhang 0001, Yongjian Wu 0001, Yunsheng Wu, Ling Shao 0001, Rongrong Ji |
Int. J. Comput. Vis. | 4 |
| 2023 | Learning Efficient GANs for Image Translation via Differentiable Masks and Co-Attention DistillationabstractGenerative Adversarial Networks (GANs) have been widely-used in image translation, but their high computational and storage costs impede the deployment on mobile devices. Prevalent methods for CNN compression cannot be directly applied to GANs due to the specificity of GAN tasks and the unstable adversarial training. To solve these, in this paper, we introduce a novel GAN compression method, termed DMAD, by proposing a Differentiable Mask and a co-Attention Distillation. The former searches for a light-weight generator architecture in a training-adaptive manner. To overcome channel inconsistency when pruning the residual connections, an adaptive cross-block group sparsity is further incorporated. The latter simultaneously distills informative attention maps from both the generator and discriminator of a pre-trained model to the searched generator, effectively stabilizing the adversarial training of our light-weight model. Experiments show that DMAD can reduce the Multiply Accumulate Operations (MACs) of CycleGAN by 13x and that of Pix2Pix by 4x while retaining a comparable performance against the full model. Our code can be available at https://github.com/SJLeo/DMAD. Mingbao Lin, Yan Wang 0059, Fei Chao 0001, Ling Shao 0001, Rongrong Ji |
IEEE Trans. Multim. | 3 |
| 2023 | Fast Monocular Depth Estimation via Side Prediction Aggregation with Continuous Spatial RefinementabstractRecent works have validated the benefit of integrating spatial information into deep networks to improve pixel-level prediction tasks such as monocular depth estimation. However, how to efficiently and robustly integrate spatial cues retains as an open problem. In this paper, we introduce the Side Prediction Aggregation (termed SPA) method to enhance the embedding of scene structural information from low-level to high-level layers. To improve the estimation accuracy, the proposed method is further equipped with continuous Spatial Refinement Loss (termed SRL) at multiple resolutions with negligible extra computation. Besides, the proposed sequential network can further perform adversarial learning at multiple resolutions. Such an adversarial refinement strategy greatly improves the accuracy of estimated depth with a little extra computation. Without using any pre-trained models, our network achieves the the-state-of-art accuracy on KITTI, NYUD V2, and Cityscapes datasets, which has achieved real-time depth estimation online. Jipeng Wu, Rongrong Ji, Qiang Wang 0023, Shengchuan Zhang, Xiaoshuai Sun, Yan Wang 0059, Mingliang Xu 0001, Feiyue Huang |
IEEE Trans. Multim. | 6 |
| 2023 | Distilling a Powerful Student Model via Online Knowledge DistillationabstractExisting online knowledge distillation approaches either adopt the student with the best performance or construct an ensemble model for better holistic performance. However, the former strategy ignores other students' information, while the latter increases the computational complexity during deployment. In this article, we propose a novel method for online knowledge distillation, termed feature fusion and self-distillation (FFSD), which comprises two key components: FFSD, toward solving the above problems in a unified framework. Different from previous works, where all students are treated equally, the proposed FFSD splits them into a leader student set and a common student set. Then, the feature fusion module converts the concatenation of feature maps from all common students into a fused feature map. The fused representation is used to assist the learning of the leader student. To enable the leader student to absorb more diverse information, we design an enhancement strategy to increase the diversity among students. Besides, a self-distillation module is adopted to convert the feature map of deeper layers into a shallower one. Then, the shallower layers are encouraged to mimic the transformed feature maps of the deeper layers, which helps the students to generalize better. After training, we simply adopt the leader student, which achieves superior performance, over the common students, without increasing the storage or inference cost. Extensive experiments on CIFAR-100 and ImageNet demonstrate the superiority of our FFSD over existing works. The code is available at https://github.com/SJLeo/FFSD. Mingbao Lin, Yan Wang 0059, Yongjian Wu 0001, Yonghong Tian 0001, Ling Shao 0001, Rongrong Ji |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2022 | Black-Box Dissector: Towards Erasing-Based Hard-Label Model Stealing Attack
Yixu Wang, Jie Li 0052, Hong Liu 0009, Yan Wang 0059, Yongjian Wu 0001, Feiyue Huang, Rongrong Ji |
ECCV (5) | 4 |
| 2022 | Towards Open-Ended Text-to-Face Generation, Combination and ManipulationabstractText-to-face (T2F) generation is an emerging research hot spot in multimedia, and its main challenge lies in the high fidelity requirement of generated portraits. Many existing works resort to exploring the latent space in a pre-trained generator, e.g., StyleGAN, which has obvious shortcomings in efficiency and generalization ability. In this paper, we propose a generative network for open-ended text-to-face generation, which is termed OpenFaceGAN. Differing from existing StyleGAN-based methods, OpenFaceGAN constructs an effective multi-modal latent space that directly converts the natural language description into a face. This mapping paradigm can fit the real data distribution well and make the model capable of open-ended and even zero-shot T2F generation. Our method improves the inference speed by an order of magnitude, e.g., 294 times than TediGAN. Based on OpenFaceGAN, we further explore text-guided face manipulation (editing). In particular, we propose a parameterized module, OpenEditor, to automatically disentangle the target latent code and update the original style information. OpenEditor also makes OpenFaceGAN directly applicable for most manipulation instructions without example-dependent searches or optimizations, greatly improving the efficiency of face manipulation. We conduct extensive experiments on two benchmark datasets namely Multi-Modal CelebA-HQ and Face2Text-v1.0. The experimental results not only show the superior performance of OpenFaceGAN to the existing T2F methods in both image quality and image-text matching degree but also greatly confirm its outstanding ability in the zero-shot generation. Codes will be released at: \textcolormagenta \urlhttps://github.com/pengjunn/OpenFace Jun Peng 0007, Han Pan, Yiyi Zhou, Xiaoshuai Sun, Yan Wang 0059, Yongjian Wu 0001, Rongrong Ji |
ACM Multimedia | 6 |
| 2022 | Towards Lightweight Transformer Via Group-Wise Transformation for Vision-and-Language TasksabstractDespite the exciting performance, Transformer is criticized for its excessive parameters and computation cost. However, compressing Transformer remains as an open problem due to its internal complexity of the layer designs, i.e., Multi-Head Attention (MHA) and Feed-Forward Network (FFN). To address this issue, we introduce Group-wise Transformation towards a universal yet lightweight Transformer for vision-and-language tasks, termed as LW-Transformer. LW-Transformer applies Group-wise Transformation to reduce both the parameters and computations of Transformer, while also preserving its two main properties, i.e., the efficient attention modeling on diverse subspaces of MHA, and the expanding-scaling feature transformation of FFN. We apply LW-Transformer to a set of Transformer-based networks, and quantitatively measure them on three vision-and-language tasks and six benchmark datasets. Experimental results show that while saving a large number of parameters and computations, LW-Transformer achieves very competitive performance against the original Transformer networks for vision-and-language tasks. To examine the generalization ability, we apply LW-Transformer to the task of image classification, and build its network based on a recently proposed image Transformer called Swin-Transformer, where the effectiveness can be also confirmed. Gen Luo, Yiyi Zhou, Xiaoshuai Sun, Yan Wang 0059, Liujuan Cao, Yongjian Wu 0001, Feiyue Huang, Rongrong Ji |
IEEE Trans. Image Process. | 4 |
| 2022 | Network Pruning Using Adaptive Exemplar FiltersabstractPopular network pruning algorithms reduce redundant information by optimizing hand-crafted models, and may cause suboptimal performance and long time in selecting filters. We innovatively introduce adaptive exemplar filters to simplify the algorithm design, resulting in an automatic and efficient pruning approach called EPruner. Inspired by the face recognition community, we use a message-passing algorithm Affinity Propagation on the weight matrices to obtain an adaptive number of exemplars, which then act as the preserved filters. EPruner breaks the dependence on the training data in determining the "important" filters and allows the CPU implementation in seconds, an order of magnitude faster than GPU-based SOTAs. Moreover, we show that the weights of exemplars provide a better initialization for the fine-tuning. On VGGNet-16, EPruner achieves a 76.34%-FLOPs reduction by removing 88.80% parameters, with 0.06% accuracy improvement on CIFAR-10. In ResNet-152, EPruner achieves a 65.12%-FLOPs reduction by removing 64.18% parameters, with only 0.71% top-5 accuracy loss on ILSVRC-2012. Our code is available at https://github.com/lmbxmu/EPruner. Mingbao Lin, Rongrong Ji, Yan Wang 0059, Yongjian Wu 0001, Feiyue Huang, Qixiang Ye |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2021 | Evolving Fully Automated Machine Learning via Life-Long Knowledge AnchorsabstractAutomated machine learning (AutoML) has achieved remarkable progress on various tasks, which is attributed to its minimal involvement of manual feature and model designs. However, most of existing AutoML pipelines only touch parts of the full machine learning pipeline, e.g., neural architecture search or optimizer selection. This leaves potentially important components such as data cleaning and model ensemble out of the optimization, and still results in considerable human involvement and suboptimal performance. The main challenges lie in the huge search space assembling all possibilities over all components, as well as the generalization ability over different tasks like image, text, and tabular etc. In this paper, we present a first-of-its-kind fully AutoML pipeline, to comprehensively automate data preprocessing, feature engineering, model generation/selection/training and ensemble for an arbitrary dataset and evaluation metric. Our innovation lies in the comprehensive scope of a learning pipeline, with a novel "life-long" knowledge anchor design to fundamentally accelerate the search over the full search space. Such knowledge anchors record detailed information of pipelines and integrates them with an evolutionary algorithm for joint optimization across components. Experiments demonstrate that the result pipeline achieves state-of-the-art performance on multiple datasets and modalities. Specifically, the proposed framework was extensively evaluated in the NeurIPS 2019 AutoDL challenge, and won the only champion with a significant gap against other approaches, on all the image, video, speech, text and tabular tracks. Xiawu Zheng, Yang Zhang 0079, Sirui Hong, Huixia Li, Lang Tang, Youcheng Xiong, Yan Wang 0059, Xiaoshuai Sun, Pengfei Zhu 0001, Chenglin Wu 0001, Rongrong Ji |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2020 | HRank: Filter Pruning Using High-Rank Feature MapabstractNeural network pruning offers a promising prospect to facilitate deploying deep neural networks on resource-limited devices. However, existing methods are still challenged by the training inefficiency and labor cost in pruning designs, due to missing theoretical guidance of non-salient network components. In this paper, we propose a novel filter pruning method by exploring the High Rank of feature maps (HRank). Our HRank is inspired by the discovery that the average rank of multiple feature maps generated by a single filter is always the same, regardless of the number of image batches CNNs receive. Based on HRank, we develop a method that is mathematically formulated to prune filters with low-rank feature maps. The principle behind our pruning is that low-rank feature maps contain less information, and thus pruned results can be easily reproduced. Besides, we experimentally show that weights with high-rank feature maps contain more important information, such that even when a portion is not updated, very little damage would be done to the model performance. Without introducing any additional constraints, HRank leads to significant improvements over the state-of-the-arts in terms of FLOPs and parameters reduction, with similar accuracies. For example, with ResNet-110, we achieve a 58.2%-FLOPs reduction by removing 59.2% of the parameters, with only a small loss of 0.14% in top-1 accuracy on CIFAR-10. With Res-50, we achieve a 43.8%-FLOPs reduction by removing 36.7% of the parameters, with only a loss of 1.17% in the top-1 accuracy on ImageNet. The codes can be available at https://github.com/lmbxmu/HRank. Mingbao Lin, Rongrong Ji, Yan Wang 0059, Yichen Zhang 0002, Baochang Zhang 0001, Yonghong Tian 0001, Ling Shao 0001 |
CVPR | 3 |
| 2020 | Enabling Deep Residual Networks for Weakly Supervised Object Detection
Yunhang Shen, Rongrong Ji, Yan Wang 0059, Feng Zheng 0001, Feiyue Huang, Yunsheng Wu |
ECCV (8) | 3 |
| 2020 | Rotated Binary Neural NetworkabstractBinary Neural Network (BNN) shows its predominance in reducing the complexity of deep neural networks. However, it suffers severe performance degradation. One of the major impediments is the large quantization error between the full-precision weight vector and its binary vector. Previous works focus on compensating for the norm gap while leaving the angular bias hardly touched. In this paper, for the first time, we explore the influence of angular bias on the quantization error and then introduce a Rotated Binary Neural Network (RBNN), which considers the angle alignment between the full-precision weight vector and its binarized version. At the beginning of each training epoch, we propose to rotate the full-precision weight vector to its binary vector to reduce the angular bias. To avoid the high complexity of learning a large rotation matrix, we further introduce a bi-rotation formulation that learns two smaller rotation matrices. In the training stage, we devise an adjustable rotated weight vector for binarization to escape the potential local optimum. Our rotation leads to around 50% weight flips which maximize the information gain. Finally, we propose a training-aware approximation of the sign function for the gradient backward. Experiments on CIFAR-10 and ImageNet demonstrate the superiorities of RBNN over many state-of-the-arts. Our source code, experimental settings, training logs and binary models are available at https://github.com/lmbxmu/RBNN. Mingbao Lin, Rongrong Ji, Baochang Zhang 0001, Yan Wang 0059, Yongjian Wu 0001, Feiyue Huang, Chia-Wen Lin |
NeurIPS | 5 |
| 2020 | Semi-Supervised Adversarial Monocular Depth EstimationabstractIn this paper, we address the problem of monocular depth estimation when only a limited number of training image-depth pairs are available. To achieve a high regression accuracy, the state-of-the-art estimation methods rely on CNNs trained with a large number of image-depth pairs, which are prohibitively costly or even infeasible to acquire. Aiming to break the curse of such expensive data collections, we propose a semi-supervised adversarial learning framework that only utilizes a small number of image-depth pairs in conjunction with a large number of easily-available monocular images to achieve high performance. In particular, we use one generator to regress the depth and two discriminators to evaluate the predicted depth, i.e., one inspects the image-depth pair while the other inspects the depth channel alone. These two discriminators provide their feedbacks to the generator as the loss to generate more realistic and accurate depth predictions. Experiments show that the proposed approach can (1) improve most state-of-the-art models on the NYUD v2 dataset by effectively leveraging additional unlabeled data sources; (2) reach state-of-the-art accuracy when the training set is small, e.g., on the Make3D dataset; (3) adapt well to an unseen new dataset (Make3D in our case) after training on an annotated dataset (KITTI in our case). Rongrong Ji, Ke Li 0015, Yan Wang 0059, Xiaoshuai Sun, Feng Guo 0005, Yongjian Wu 0001, Feiyue Huang, Jiebo Luo 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2019 | Cyclic Guidance for Weakly Supervised Joint Detection and SegmentationabstractWeakly supervised learning has attracted growing research attention due to the significant saving in annotation cost for tasks that require intra-image annotations, such as object detection and semantic segmentation. To this end, existing weakly supervised object detection and semantic segmentation approaches follow an iterative label mining and model training pipeline. However, such a self-enforcement pipeline makes both tasks easy to be trapped in local minimums. In this paper, we join weakly supervised object detection and segmentation tasks with a multi-task learning scheme for the first time, which uses their respective failure patterns to complement each other's learning. Such cross-task enforcement helps both tasks to leap out of their respective local minimums. In particular, we present an efficient and effective framework termed Weakly Supervised Joint Detection and Segmentation (WS-JDS). WS-JDS has two branches for the above two tasks, which share the same backbone network. In the learning stage, it uses the same cyclic training paradigm but with a specific loss function such that the two branches benefit each other. Extensive experiments have been conducted on the widely-used Pascal VOC and COCO benchmarks, which demonstrate that our model has achieved competitive performance with the state-of-the-art algorithms. Yunhang Shen, Rongrong Ji, Yan Wang 0059, Yongjian Wu 0001, Liujuan Cao |
CVPR | 3 |
| 2019 | Variational Structured Semantic Inference for Diverse Image CaptioningabstractDespite the exciting progress in image captioning, generating diverse captions for a given image remains as an open problem. Existing methods typically apply generative models such as Variational Auto-Encoder to diversify the captions, which however neglect two key factors of diverse expression, i.e., the lexical diversity and the syntactic diversity. To model these two inherent diversities in image captioning, we propose a Variational Structured Semantic Inferring model (termed VSSI-cap) executed in a novel structured encoder-inferer-decoder schema. VSSI-cap mainly innovates in a novel structure, i.e., Variational Multi-modal Inferring tree (termed VarMI-tree). In particular, conditioned on the visual-textual features from the encoder, the VarMI-tree models the lexical and syntactic diversities by inferring their latent variables (with variations) in an approximate posterior inference guided by a visual semantic prior. Then, a reconstruction loss and the posterior-prior KL-divergence are jointly estimated to optimize the VSSI-cap model. Finally, diverse captions are generated upon the visual features and the latent variables from this structured encoder-inferer-decoder model. Experiments on the benchmark dataset show that the proposed VSSI-cap achieves significant improvements over the state-of-the-arts. Fuhai Chen, Rongrong Ji, Jiayi Ji, Xiaoshuai Sun, Baochang Zhang 0001, Xuri Ge, Yongjian Wu 0001, Feiyue Huang, Yan Wang 0059 |
NeurIPS | 9 |
| 2018 | Generative Adversarial Learning Towards Fast Weakly Supervised DetectionabstractWeakly supervised object detection has attracted extensive research efforts in recent years. Without the need of annotating bounding boxes, the existing methods usually follow a two/multi-stage pipeline with an online compulsive stage to extract object proposals, which is an order of magnitude slower than fast fully supervised object detectors such as SSD [31] and YOLO [34]. In this paper, we speedup online weakly supervised object detectors by orders of magnitude by proposing a novel generative adversarial learning paradigm. In the proposed paradigm, the generator is a one-stage object detector to generate bounding boxes from images. To guide the learning of object-level generator, a surrogator is introduced to mine high-quality bounding boxes for training. We further adapt a structural similarity loss in combination with an adversarial loss into the training objective, which solves the challenge that the bounding boxes produced by the surrogator may not well capture their ground truth. Our one-stage detector outperforms all existing schemes in terms of detection accuracy, running at 118 frames per second, which is up to 438× faster than the state-of-the-art weakly supervised detectors [8, 30, 15, 27, 45]. The code will be available publicly soon. Yunhang Shen, Rongrong Ji, Shengchuan Zhang, Wangmeng Zuo, Yan Wang 0059 |
CVPR | 5 |
| 2018 | Model-Driven Feedforward Prediction for Manipulation of Deformable ObjectsabstractRobotic manipulation of deformable objects is a difficult problem especially because of the complexity of the many different ways an object can deform. Searching such a high-dimensional state space makes it difficult to recognize, track, and manipulate deformable objects. In this paper, we introduce a predictive, model-driven approach to address this challenge, using a precomputed, simulated database of deformable object models. Mesh models of common deformable garments are simulated with the garments picked up in multiple different poses under gravity, and stored in a database for fast and efficient retrieval. To validate this approach, we developed a comprehensive pipeline for manipulating clothing as in a typical laundry task. First, the database is used for category and the pose estimation is used for a garment in an arbitrary position. A fully featured 3-D model of the garment is constructed in real time, and volumetric features are then used to obtain the most similar model in the database to predict the object category and pose. Second, the database can significantly benefit the manipulation of deformable objects via nonrigid registration, providing accurate correspondences between the reconstructed object model and the database models. Third, the accurate model simulation can also be used to optimize the trajectories for the manipulation of deformable objects, such as the folding of garments. Extensive experimental results are shown for the above tasks using a variety of different clothings. Note to Practitioners-This paper provides an open source, extensible, 3-D database for dissemination to the robotics and graphics communities. Model-driven methods are proliferating, and they need to be applied, tested, and validated in real environments. A key idea we have exploited is to have an innovative and novel use of simulation. This database will serve as infrastructure for developing advanced robotic machine learning algorithms. We want to address this machine learning idea ourselves, but we expect the dissemination of the database to other researchers with different agendas and task applications, which will bring wide progress in this area. Our proposed methods, as mentioned earlier, can be easily applied to interrelated areas. One example is that the 3-D shape-based matching algorithm can be used for other objects, such as bottles, papers, and food. After integrating with other robotic systems, the use of the robot can be easily extended to other tasks, such as making food, cleaning room, and fetching objects, to assist our daily life. Yinxiao Li, Yan Wang 0059, Yonghao Yue, Danfei Xu, Michael Case, Shih-Fu Chang, Eitan Grinspun, Peter K. Allen |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2017 | Learning explicit video attributes from mid-level representation for video captioning
Fudong Nian, Teng Li 0001, Yan Wang 0059, Xinyu Wu 0001, Bingbing Ni, Changsheng Xu |
Comput. Vis. Image Underst. | 3 |
| 2017 | Contextual aerial image categorization using codebook
Yan Wang 0059, Xinyu Wu 0001, Yating Yin, Teng Li 0001 |
J. Vis. Commun. Image Represent. | 2 |
| 2017 | Vehicle counting in crowded scenes with multi-channel and multi-task convolutional neural networks
Maojin Sun, Yan Wang 0059, Teng Li 0001, Jing Lv |
J. Vis. Commun. Image Represent. | 2 |
| 2017 | Robust face anti-spoofing with depth information
Yan Wang 0059, Fudong Nian, Teng Li 0001, Kongqiao Wang |
J. Vis. Commun. Image Represent. | 1 |
| 2016 | Search-Based Depth Estimation via Coupled Dictionary Learning with Large-Margin Structure Inference
Yan Zhang 0109, Rongrong Ji, Xiaopeng Fan 0001, Yan Wang 0059, Feng Guo 0005, Yue Gao 0002, Debin Zhao |
ECCV (5) | 4 |
| 2016 | PanoSwarm: Collaborative and Synchronized Multi-Device Panoramic PhotographyabstractTaking a picture has been traditionally a one-person task. In this paper we present a novel system that allows multiple mobile devices to work collaboratively in a synchronized fashion to capture a panorama of a highly dynamic scene, creating an entirely new photography experience that encourages social interactions and teamwork. Our system contains two components: a client app that runs on all participating devices, and a server program that monitors and communicates with each device. In a capturing session, the server collects in realtime the viewfinder images of all devices and stitches them on-the-fly to create a panorama preview, which is then streamed to all devices as visual guidance. The system also allows one camera to be the host and send direct visual instructions to others to guide camera adjustment. When ready, all devices take pictures at the same time for panorama stitching. Our preliminary study suggests that the proposed system can help users capture high quality panoramas with an enjoyable teamwork experience. Yan Wang 0059, Sunghyun Cho, Jue Wang 0001, Shih-Fu Chang |
IUI | 1 |
| 2016 | 3D shape retrieval using a single depth image from low-cost sensorsabstractContent-based 3D shape retrieval is an important problem in computer vision. Traditional retrieval interfaces require a 2D sketch or a manually designed 3D model as the query, which is difficult to specify and thus not practical in real applications. With the recent advance in low-cost 3D sensors such as Microsoft Kinect and Intel Realsense, capturing depth images that carry 3D information is fairly simple, making shape retrieval more practical and user-friendly. In this paper, we study the problem of cross-domain 3D shape retrieval using a single depth image from low-cost sensors as the query to search for similar human designed CAD models. We propose a novel method using an ensemble of autoencoders in which each autoencoder is trained to learn a compressed representation of depth views synthesize d from each database object. By viewing each autoencoder as a probabilistic model, a likelihood score can be derived as a similarity measure. A domain adaptation layer is built on top of autoencoder outputs to explicitly address the cross-domain issue (between noisy sensory data and clean 3D models) by incorporating training data of sensor depth images and their category labels in a weakly supennsed learning formulation. Experiments using real-world depth images and a large-scale CAD dataset demonstrate the effectiveness of our approach, which offers significant improvements over state-of-the-art 3D shape retrieval methods. Yan Wang 0059, Shih-Fu Chang |
WACV | 2 |
| 2016 | Pornographic image detection utilizing deep convolutional neural networks
Fudong Nian, Teng Li 0001, Yan Wang 0059, Mingliang Xu 0001 |
Neurocomputing | 3 |
| 2016 | Dense crowd counting from still images with convolutional neural networks
Yaocong Hu, Fudong Nian, Yan Wang 0059, Teng Li 0001 |
J. Vis. Commun. Image Represent. | 4 |
| 2016 | Joint Depth and Semantic Inference from a Single Image via Elastic Conditional Random Field
Rongrong Ji, Liujuan Cao, Yan Wang 0059 |
Pattern Recognit. | 3 |
| 2016 | EI3D: Expression-invariant 3D face recognition based on feature and shape matching
Yulan Guo, Yinjie Lei, Li Liu 0002, Yan Wang 0059, Mohammed Bennamoun, Ferdous Sohel |
Pattern Recognit. Lett. | 4 |
| 2015 | Regrasping and unfolding of garments using predictive thin shell modelingabstractDeformable objects such as garments are highly unstructured, making them difficult to recognize and manipulate. In this paper, we propose a novel method to teach a two-arm robot to efficiently track the states of a garment from an unknown state to a known state by iterative regrasping. The problem is formulated as a constrained weighted evaluation metric for evaluating the two desired grasping points during regrasping, which can also be used for a convergence criterion The result is then adopted as an estimation to initialize a regrasping, which is then considered as a new state for evaluation. The process stops when the predicted thin shell conclusively agrees with reconstruction. We show experimental results for regrasping a number of different garments including sweater, knitwear, pants, and leggings, etc. Yinxiao Li, Danfei Xu, Yonghao Yue, Yan Wang 0059, Shih-Fu Chang, Eitan Grinspun, Peter K. Allen |
ICRA | 4 |
| 2014 | Discriminative Indexing for Probabilistic Image Patch Priors
Yan Wang 0059, Sunghyun Cho, Jue Wang 0001, Shih-Fu Chang |
ECCV (4) | 1 |
| 2014 | From Low-Cost Depth Sensors to CAD: Cross-Domain 3D Shape Retrieval via Regression Tree Fields
Yan Wang 0059, Jun Wang 0006, Shih-Fu Chang |
ECCV (1) | 1 |
| 2014 | Real-time pose estimation of deformable objects using a volumetric approachabstractPose estimation of deformable objects is a fundamental and challenging problem in robotics. We present a novel solution to this problem by first reconstructing a 3D model of the object from a low-cost depth sensor such as Kinect, and then searching a database of simulated models in different poses to predict the pose. Given noisy depth images from 360-degree views of the target object acquired from the Kinect sensor, we reconstruct a smooth 3D model of the object using depth image segmentation and volumetric fusion. Then with an efficient feature extraction and matching scheme, we search the database, which contains a large number of deformable objects in different poses, to obtain the most similar model, whose pose is then adopted as the prediction. Extensive experiments demonstrate better accuracy and orders of magnitude speed-up compared to our previous work. An additional benefit of our method is that it produces a high-quality mesh model and camera pose, which is necessary for other tasks such as regrasping and object manipulation. Yinxiao Li, Yan Wang 0059, Michael Case, Shih-Fu Chang, Peter K. Allen |
IROS | 2 |
| 2013 | Label Propagation from ImageNet to 3D Point CloudsabstractRecent years have witnessed a growing interest in understanding the semantics of point clouds in a wide variety of applications. However, point cloud labeling remains an open problem, due to the difficulty in acquiring sufficient 3D point labels towards training effective classifiers. In this paper, we overcome this challenge by utilizing the existing massive 2D semantic labeled datasets from decade-long community efforts, such as Image Net and Label Me, and a novel ``cross-domain'' label propagation approach. Our proposed method consists of two major novel components, Exemplar SVM based label propagation, which effectively addresses the cross-domain issue, and a graphical model based contextual refinement incorporating 3D constraints. Most importantly, the entire process does not require any training data from the target scenes, also with good scalability towards large scale applications. We evaluate our approach on the well-known Cornell Point Cloud Dataset, achieving much greater efficiency and comparable accuracy even without any 3D training data. Our approach shows further major gains in accuracy when the training data from the target scenes is used, outperforming state-of-the-art approaches with far better efficiency. Yan Wang 0059, Rongrong Ji, Shih-Fu Chang |
CVPR | 1 |
| 2011 | Community Discovery from Movie and Its Application to Poster Generation
Yan Wang 0059, Tao Mei 0001, Xian-Sheng Hua 0001 |
MMM (1) | 1 |
| 2010 | Dynamic Video Collage
Yan Wang 0059, Tao Mei 0001, Jingdong Wang 0001, Xian-Sheng Hua 0001 |
MMM | 1 |