EDBT 2026 Demo / reviewers in the wild / expert
Jingjing Wang 0005
dblp:62/2631-5
· DBLP profile ↗
26ranked-venue papers
4as first author
18since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 22 · 3 first-author · 16 since 2021Artificial intelligence and machine learning · 11 · 2 first-author · 9 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Arbitrary-Scale Point Cloud Upsampling by Voxel-Based Network with Latent Geometric-Consistent LearningabstractRecently, arbitrary-scale point cloud upsampling mechanism became increasingly popular due to its efficiency and convenience for practical applications. To achieve this, most previous approaches formulate it as a problem of surface approximation and employ point-based networks to learn surface representations. However, learning surfaces from sparse point clouds is more challenging, and thus they often suffer from the low-fidelity geometry approximation. To address it, we propose an arbitrary-scale Point cloud Upsampling framework using Voxel-based Network (PU-VoxelNet). Thanks to the completeness and regularity inherited from the voxel representation, voxel-based networks are capable of providing predefined grid space to approximate 3D surface, and an arbitrary number of points can be reconstructed according to the predicted density distribution within each grid cell. However, we investigate the inaccurate grid sampling caused by imprecise density predictions. To address this issue, a density-guided grid resampling method is developed to generate high-fidelity points while effectively avoiding sampling outliers. Further, to improve the fine-grained details, we present an auxiliary training supervision to enforce the latent geometric consistency among local surface patches. Extensive experiments indicate the proposed approach outperforms the state-of-the-art approaches not only in terms of fixed upsampling rates but also for arbitrary-scale upsampling. The code is available at https://github.com/hikvision-research/3DVision Jingjing Wang 0005, Di Xie, Shiliang Pu |
AAAI | 3 |
| 2023 | Rethinking the Approximation Error in 3D Surface Fitting for Point Cloud Normal EstimationabstractMost existing approaches for point cloud normal estimation aim to locally fit a geometric surface and calculate the normal from the fitted surface. Recently, learning-based methods have adopted a routine of predicting pointwise weights to solve the weighted least-squares surface fitting problem. Despite achieving remarkable progress, these methods overlook the approximation error of the fitting problem, resulting in a less accurate fitted surface. In this paper, we first carry out in-depth analysis of the approximation error in the surface fitting problem. Then, in order to bridge the gap between estimated and precise surface normals, we present two basic design principles: 1) applies the Z-direction Transform to rotate local patches for a better surface fitting with a lower approximation error; 2) models the error of the normal estimation as a learnable term. We implement these two principles using deep neural networks, and integrate them with the state-of-the-art (SOTA) normal estimation methods in a plug-and-play manner. Extensive experiments verify our approaches bring benefits to point cloud normal estimation and push the frontier of state-of-the-art performance on both synthetic and real-world datasets. The code is available at https://github.com/hikvision-research/3DVision. Jingjing Wang 0005, Di Xie, Shiliang Pu |
CVPR | 3 |
| 2023 | MDR-MFI:Multi-Branch Decoupled Regression and Multi-Scale Feature Interaction for Partial-to-Partial Cloud RegistrationabstractPoint cloud registration is a fundamental task in the 3D vision field. Many previous works adopt the regression model to estimate the transformation parameters. However, these methods couple the estimation of rotation and translation via a single regression branch, which suffers from the mutual interference among rotation and translation. In addition, previous methods extract and interact features in a single scale, which ignores the rich information from multiple scales. To address above issues, in this paper, we propose a multi-branch decoupled regression and multi-scale feature interaction (MDR-MFI) framework for point cloud registration. Firstly, we decouple the estimation of 7 transformation parameters via multiple regression branches. The decoupled structure effectively mitigates the mutual interference among 7 parameters, resulting in improved performance. Secondly, we propose a multi-scale feature extraction and interaction framework to encourage the network to learn more discriminative features. Experimental results demonstrate that our method achieves state-of-the-art performance on public datasets. The code is available at https://github.com/hikvision-research/3DVision. Weidong Dai, Jingjing Wang 0005, Di Xie, Shiliang Pu |
ICASSP | 3 |
| 2023 | Single Domain Dynamic Generalization for Iris Presentation Attack DetectionabstractIris presentation attack detection (PAD) has achieved great success under intra-domain settings but easily degrades on unseen domains. Conventional domain generalization methods mitigate the gap by learning domain-invariant features. However, they ignore the discriminative information in the domain-specific features. Moreover, we usually face a more realistic scenario with only one single domain available for training. To tackle the above issues, we propose a Single Domain Dynamic Generalization (SDDG) framework, which simultaneously exploits domain-invariant and domain-specific features on a per-sample basis and learns to generalize to various unseen domains with numerous natural images. Specifically, a dynamic block is designed to adaptively adjust the network with a dynamic adaptor. And an information maximization loss is further combined to increase diversity. The whole network is integrated into the meta-learning paradigm. We generate amplitude perturbed images and cover diverse domains with natural images. Therefore, the network can learn to generalize to the perturbed domains in the meta-test phase. Extensive experiments show the proposed method is effective and outperforms the state-of-the-art on LivDet-Iris 2017 dataset. Yachun Li, Jingjing Wang 0005, Yuhui Chen, Di Xie, Shiliang Pu |
ICASSP | 2 |
| 2023 | Learning Expressive And Generalizable Motion Features For Face Forgery DetectionabstractPrevious face forgery detection methods mainly focus on appearance features, which may be easily attacked by sophisticated manipulation. Considering the majority of current face manipulation methods generate fake faces based on a single frame, which do not take frame consistency and coordination into consideration, artifacts on frame sequences are more effective for face forgery detection. However, current sequence-based face forgery detection methods use general video classification networks directly, which discard the special and discriminative motion information for face manipulation detection. To this end, we propose an effective sequence-based forgery detection framework based on an existing video classification method. To make the motion features more expressive for manipulation detection, we propose an alternative motion consistency block instead of the original motion features module. To make the learned features more generalizable, we propose an auxiliary anomaly detection block. With these two specially designed improvements, we make a general video classification network achieve promising results on three popular face forgery datasets. Jingyi Zhang 0003, Peng Zhang 0075, Jingjing Wang 0005, Di Xie, Shiliang Pu |
ICASSP | 3 |
| 2023 | Weakly Supervised Regional and Temporal Learning for Facial Action Unit RecognitionabstractAutomatic facial action unit (AU) recognition is a challenging task due to the scarcity of manual annotations. To alleviate this problem, a large amount of efforts has been dedicated to exploiting various weakly supervised methods which leverage numerous unlabeled data. However, many aspects with regard to some unique properties of AUs, such as the regional and relational characteristics, are not sufficiently explored in previous works. Motivated by this, we take the AU properties into consideration and propose two auxiliary AU related tasks to bridge the gap between limited annotations and the model performance in a self-supervised manner via the unlabeled data. Specifically, to enhance the discrimination of regional features with AU relation embedding, we design a task of RoI inpainting to recover the randomly cropped AU patches. Meanwhile, a single image based optical flow estimation task is proposed to leverage the dynamic change of facial muscles and encode the motion information into the global feature representation. Based on these two self-supervised auxiliary tasks, local features, mutual relation and motion cues of AUs are better captured in the backbone network. Furthermore, by incorporating semi-supervised learning, we propose an end-to-end trainable framework named weakly supervised regional and temporal learning (WSRTL) for AU recognition. Extensive experiments on BP4D and DISFA demonstrate the superiority of our method and new state-of-the-art performances are achieved. Jingwei Yan, Jingjing Wang 0005, Qiang Li 0044, Chunmao Wang, Shiliang Pu |
IEEE Trans. Multim. | 2 |
| 2022 | Point Cloud Upsampling via Cascaded Refinement Network
Jingjing Wang 0005, Di Xie, Shiliang Pu |
ACCV (1) | 3 |
| 2022 | Multi-scale Wavelet Transformer for Face Forgery Detection
Jingjing Wang 0005, Peng Zhang 0075, Chunmao Wang, Di Xie, Shiliang Pu |
ACCV (6) | 2 |
| 2022 | Unimodal-Concentrated Loss: Fully Adaptive Label Distribution Learning for Ordinal RegressionabstractLearning from a label distribution has achieved promising results on ordinal regression tasks such as facial age and head pose estimation wherein, the concept of adaptive label distribution learning (ALDL) has drawn lots of attention recently for its superiority in theory. However, compared with the methods assuming fixed form label distribution, ALDL methods have not achieved better performance. We argue that existing ALDL algorithms do not fully exploit the intrinsic properties of ordinal regression. In this paper, we emphatically summarize that learning an adaptive label distribution on ordinal regression tasks should follow three principles. First, the probability corresponding to the ground-truth should be the highest in label distribution. Second, the probabilities of neighboring labels should decrease with the increase of distance away from the ground-truth, i.e., the distribution is unimodal. Third, the label distribution should vary with samples changing, and even be distinct for different instances with the same label, due to the different levels of difficulty and ambiguity. Under the premise of these principles, we propose a novel loss function for fully adaptive label distribution learning, namely unimodal-concentrated loss. Specifically, the unimodal loss derived from the learning to rank strategy constrains the distribution to be unimodal. Furthermore, the estimation error and the variance of the predicted distribution for a specific sample are integrated into the proposed concentrated loss to make the predicted distribution maximize at the ground-truth and vary according to the predicting uncertainty. Extensive experimental results on typical ordinal regression tasks including age and head pose estimation, show the superiority of our proposed unimodal-concentrated loss compared with existing loss functions. Qiang Li 0044, Jingjing Wang 0005, Zhaoliang Yao, Yachun Li, Pengju Yang 0001, Jingwei Yan, Chunmao Wang, Shiliang Pu |
CVPR | 2 |
| 2022 | FBNet: Feedback Network for Point Cloud Completion
Hongyu Yan, Jingjing Wang 0005, Di Xie, Shiliang Pu |
ECCV (2) | 3 |
| 2022 | Learning Multiple Explainable and Generalizable Cues for Face Anti-SpoofingabstractAlthough previous CNN based face anti-spoofing methods have achieved promising performance under intra-dataset testing, they suffer from poor generalization under cross-dataset testing. The main reason is that they learn the network with only binary supervision, which may learn arbitrary cues overfitting on the training dataset. To make the learned feature explainable and more generalizable, some researchers introduce facial depth and reflection map as the auxiliary supervision. However, many other generalizable cues are unexplored for face anti-spoofing, which limits their performance under cross-dataset testing. To this end, we propose a novel framework to learn multiple explainable and generalizable cues (MEGC) for face anti-spoofing. Specifically, inspired by the process of human decision, four mainly used cues by humans are introduced as auxiliary supervision including the boundary of spoof medium, moiré pattern, reflection artifacts and facial depth in addition to the binary supervision. To avoid extra labelling cost, corresponding synthetic methods are proposed to generate these auxiliary supervision maps. Extensive experiments on public datasets validate the effectiveness of these cues, and state-of-the-art performances are achieved by our proposed method. Ying Bian, Peng Zhang 0075, Jingjing Wang 0005, Chunmao Wang, Shiliang Pu |
ICASSP | 3 |
| 2022 | Few-Shot One-Class Domain Adaptation Based On Frequency For Iris Presentation Attack DetectionabstractIris presentation attack detection (PAD) has achieved remarkable success to ensure the reliability and security of iris recognition systems. Most existing methods exploit discriminative features in the spatial domain and report outstanding performance under intra-dataset settings. However, the degradation of performance is inevitable under cross-dataset settings, suffering from domain shift. In consideration of real-world applications, a small number of bonafide samples are easily accessible. We thus define a new domain adaptation setting called Few-shot One-class Domain Adaptation (FODA), where adaptation only relies on a limited number of target bonafide samples. To address this problem, we propose a novel FODA framework based on the expressive power of frequency information. Specifically, our method integrates frequency-related information through two proposed modules. Frequency-based Attention Module (FAM) aggregates frequency information into spatial attention and explicitly emphasizes high-frequency fine-grained features. Frequency Mixing Module (FMM) mixes certain frequency components to generate large-scale target-style samples for adaptation with limited target bonafide samples. Extensive experiments on LivDet-Iris 2017 dataset demonstrate the proposed method achieves state-of-the-art or competitive performance under both cross-dataset and intra-dataset settings. Yachun Li, Ying Lian, Jingjing Wang 0005, Yuhui Chen, Chunmao Wang, Shiliang Pu |
ICASSP | 3 |
| 2022 | Semi-Supervised Ranking for Object Image Blur AssessmentabstractAssessing the blurriness of an object image is fundamentally important to improve the performance for object recognition and retrieval. The main challenge lies in the lack of abundant images with reliable labels and effective learning strategies. Current datasets are labeled with limited and confused quality levels. To overcome this limitation, we propose to label the rank relationships between pairwise images rather their quality levels, since it is much easier for humans to label, and establish a large-scale realistic face image blur assessment dataset with reliable labels. Based on this dataset, we propose a method to obtain the blur scores only with the pairwise rank labels as supervision. Moreover, to further improve the performance, we propose a self-supervised method based on quadruplet ranking consistency to leverage the unlabeled data more effectively. The supervised and self-supervised methods constitute a final semi-supervised learning framework, which can be trained end-to-end. Experimental results demonstrate the effectiveness of our method. Source of labeled datasets: https://github.com/yzliangHIK2022/SSRanking-for-Object-BA Qiang Li 0044, Zhaoliang Yao, Jingjing Wang 0005, Pengju Yang 0001, Di Xie, Shiliang Pu |
ICIP | 3 |
| 2022 | High-Accuracy and Energy-Efficient Action Recognition with Deep Spiking Neural Network
Jingren Zhang, Jingjing Wang 0005, Di Xie, Shiliang Pu |
ICONIP (2) | 2 |
| 2021 | Self-Domain Adaptation for Face Anti-SpoofingabstractAlthough current face anti-spoofing methods achieve promising results under intra-dataset testing, they suffer from poor generalization to unseen attacks. Most existing works adopt domain adaptation (DA) or domain generalization (DG) techniques to address this problem. However, the target domain is often unknown during training which limits the utilization of DA methods. DG methods can conquer this by learning domain invariant features without seeing any target data. However, they fail in utilizing the information of target data. In this paper, we propose a self-domain adaptation framework to leverage the unlabeled test domain data at inference. Specifically, a domain adaptor is designed to adapt the model for test domain. In order to learn a better adaptor, a meta-learning based adaptor learning algorithm is proposed using the data of multiple source domains at the training step. At test time, the adaptor is updated using only the test domain data according to the proposed unsupervised adaptor loss to further improve the performance. Extensive experiments on four public datasets validate the effectiveness of the proposed method. Jingjing Wang 0005, Jingyi Zhang 0003, Ying Bian, Youyi Cai, Chunmao Wang, Shiliang Pu |
AAAI | 1 |
| 2021 | Multi-Level Adaptive Region of Interest and Graph Learning for Facial Action Unit RecognitionabstractIn facial action unit (AU) recognition tasks, regional feature learning and AU relation modeling are two effective aspects which are worth exploring. However, the limited representation capacity of regional features makes it difficult for relation models to embed AU relationship knowledge. In this paper, we propose a novel multi-level adaptive ROI and graph learning (MARGL) framework to tackle this problem. Specifically, an adaptive ROI learning module is designed to automatically adjust the location and size of the predefined AU regions. Meanwhile, besides relationship between AUs, there exists strong relevance between regional features across multiple levels of the backbone network as level-wise features focus on different aspects of representation. In order to incorporate the intra-level AU relation and inter-level AU regional relevance simultaneously, a multi-level AU relation graph is constructed and graph convolution is performed to further enhance AU regional features of each level. Experiments on BP4D and DISFA demonstrate the proposed MARGL significantly outperforms the previous state-of-the-art methods. Jingwei Yan, Boyuan Jiang, Jingjing Wang 0005, Qiang Li 0044, Chunmao Wang, Shiliang Pu |
ICASSP | 3 |
| 2021 | Deep Spiking Neural Network for High-Accuracy and Energy-Efficient Face Action Unit RecognitionabstractIn recent years, spiking neural networks (SNNs) have received significant attention as the third-generation of networks due to their event-driven and low-powered nature. However, their applications have been limited to relatively simple tasks such as image classification, since it is difficult to train SNNs and converting deep artificial neural networks (ANNs) into SNNs directly usually causes large accuracy degradation. In this paper, we employ an SNN to solve a more challenging multi-label classification task and propose the first spiking-based network for face action unit (AU) recognition. Specifically, a relation extracting module based on graph convolution network (GCN) is proposed to leverage AU regional features. Channel-wise normalization methods for residual blocks of the Resnet backbone and GCN blocks are proposed for ANN-to-SNN conversion to keep the high performance. Experiments on the BP4D dataset show that our proposed model achieves high-accuracy performance, and converges 3 times faster than previous methods. Jingren Zhang, Jingjing Wang 0005, Jingwei Yan, Chunmao Wang, Shiliang Pu |
IJCNN | 2 |
| 2021 | Self-Supervised Regional and Temporal Auxiliary Tasks for Facial Action Unit RecognitionabstractAutomatic facial action unit (AU) recognition is a challenging task due to the scarcity of manual annotations. To alleviate this problem, a large amount of efforts has been dedicated to exploiting various methods which leverage numerous unlabeled data. However, many aspects with regard to some unique properties of AUs, such as the regional and relational characteristics, are not sufficiently explored in previous works. Motivated by this, we take the AU properties into consideration and propose two auxiliary AU related tasks to bridge the gap between limited annotations and the model performance in a self-supervised manner via the unlabeled data. Specifically, to enhance the discrimination of regional features with AU relation embedding, we design a task of RoI inpainting to recover the randomly cropped AU patches. Meanwhile, a single image based optical flow estimation task is proposed to leverage the dynamic change of facial muscles and encode the motion information into the global feature representation. Based on these two self-supervised auxiliary tasks, local features, mutual relation and motion cues of AUs are better captured in the backbone network with the proposed regional and temporal based auxiliary task learning (RTATL) framework. Extensive experiments on BP4D and DISFA demonstrate the superiority of our method and new state-of-the-art performances are achieved. Jingwei Yan, Jingjing Wang 0005, Qiang Li 0044, Chunmao Wang, Shiliang Pu |
ACM Multimedia | 2 |
| 2017 | Robust visual tracking with deep feature fusionabstractRecently, CNN (Convolutional Neural Network) based trackers have achieved promising results benefited from their robust feature representation. However, most trackers only use features from a certain layer, which limits their performance. In this paper, we propose a novel CNN based tracker. Firstly, we use local detection and global detection network for target localization. In local detection network, we fuse features from different layers to train a fully convolutional neural network for target localization. In case the local detection network fails when the target disappear for a while and appears in another location, we train a global detection network to detect if the target appears again. Then, we employ a correlation filter to estimate accurate scale of the target using HOG features extracted around predicted location. Extensive experiments on various challenging video sequences demonstrate the effectiveness of our proposed algorithm compared with several state-of-the-art trackers. Guokun Wang, Jingjing Wang 0005, Wenyi Tang, Nenghai Yu |
ICASSP | 2 |
| 2016 | Part-based multi-graph ranking for visual trackingabstractRecently, graph ranking-based methods have been introduced to visual tracking and achieved promising results due to the local structure preserving property. However, existing graph ranking-based trackers use holistic templates to construct the graphs which makes the trackers sensitive to occlusions. In this paper, we propose a part-based multi-graph ranking algorithm for robust visual tracking. In our method, template samples are divided into local parts. Multiple graphs are constructed based on different part samples and different feature representations. Then, the multiple graphs are integrated into a regularization framework with each graph assigned a weight. Furthermore, by imposing the l2,1 norm on the weight matrix of graphs, the confident parts are selected to reduce the effects of occluded ones. An effective optimization scheme is proposed to learn the weight matrix and the rank scores jointly. Experimental results on various challenging video sequences demonstrate our proposed algorithm outperforms state-of-the-art trackers. Jingjing Wang 0005, Chi Fei, Liansheng Zhuang, Nenghai Yu |
ICIP | 1 |
| 2016 | Anomaly detection via 3D-HOF and fast double sparse representationabstractThis paper presents a framework for anomaly detection in videos which considers both motion and appearance features. For motion cues, we propose a new feature called 3D-HOF, which effectively extracts both velocity and orientation from the optical flow map. At the same time, we introduce the concept of “depth of field” problem to make the detection more accurate when the velocity of an object may seem to be different according to its distance to the camera. For appearance cues, we use 3D gradients of spatio-temporal cuboids as features. Next in the detection part, we propose a fast double sparse representation method in order to make the process faster in some actual scenes. Finally, we integrate both the outcomes of using motion and appearance cues as the final outcomes. Results on discriminative datasets show efficiency and effectiveness compared to the state-of-the-art methods. Ziping Zhu, Jingjing Wang 0005, Nenghai Yu |
ICIP | 2 |
| 2016 | Multi-level visual tracking with hierarchical tree structural constraint
Jingjing Wang 0005, Nenghai Yu, Feng Zhu 0006, Liansheng Zhuang |
Neurocomputing | 1 |
| 2016 | Locality-preserving low-rank representation for graph construction from nonlinear manifolds
Liansheng Zhuang, Jingjing Wang 0005, Zhouchen Lin, Allen Y. Yang, Yi Ma 0001, Nenghai Yu |
Neurocomputing | 2 |
| 2015 | Moving Object Segmentation by Length-Unconstrained Trajectory Analysis
Qiyu Liao, Bingbing Zhuang, Jingjing Wang 0005, Nenghai Yu |
ICIG (2) | 3 |
| 2015 | Cooperative Target Tracking in Dual-Camera System with Bidirectional Information Fusion
Jingjing Wang 0005, Nenghai Yu |
ICIG (2) | 1 |
| 2015 | Constructing a Nonnegative Low-Rank and Sparse Graph With Data-Adaptive FeaturesabstractThis paper aims at constructing a good graph to discover the intrinsic data structures under a semisupervised learning setting. First, we propose to build a nonnegative low-rank and sparse (referred to as NNLRS) graph for the given data representation. In particular, the weights of edges in the graph are obtained by seeking a nonnegative low-rank and sparse reconstruction coefficients matrix that represents each data sample as a linear combination of others. The so-obtained NNLRS-graph captures both the global mixture of subspaces structure (by the low-rankness) and the locally linear structure (by the sparseness) of the data, hence it is both generative and discriminative. Second, as good features are extremely important for constructing a good graph, we propose to learn the data embedding matrix and construct the graph simultaneously within one framework, which is termed as NNLRS with embedded features (referred to as NNLRS-EF). Extensive NNLRS experiments on three publicly available data sets demonstrate that the proposed method outperforms the state-of-the-art graph construction method by a large margin for both semisupervised classification and discriminative analysis, which verifies the effectiveness of our proposed method. Liansheng Zhuang, Shenghua Gao, Jinhui Tang 0001, Jingjing Wang 0005, Zhouchen Lin, Yi Ma 0001, Nenghai Yu |
IEEE Trans. Image Process. | 4 |