EDBT 2026 Demo / reviewers in the wild / expert
Qi She
dblp:171/7773
· DBLP profile ↗
25ranked-venue papers
4as first author
19since 2021 · last 2025
0000-0002-4490-2941ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 24 · 3 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 2 first-author · 12 since 2021Systems, architecture and hardware · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs
Qizhe Zhang, Aosong Cheng, Ming Lu 0002, Renrui Zhang, Zhiyong Zhuo, Jiajun Cao, Shaobo Guo, Qi She, Shanghang Zhang |
ICCV | 8 |
| 2025 | Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMsabstractIn multimodal large language models (MLLMs), the length of input visual tokens is often significantly greater than that of their textual counterparts, leading to a high inference cost. Many works aim to address this issue by removing redundant visual tokens. However, current approaches either rely on attention-based pruning, which retains numerous duplicate tokens, or use similarity-based pruning, overlooking the instruction relevance, consequently causing suboptimal performance. In this paper, we go beyond attention or similarity by proposing a novel visual token pruning method named **CDPruner**, which maximizes the conditional diversity of retained tokens. We first define the conditional similarity between visual tokens conditioned on the instruction, and then reformulate the token pruning problem with determinantal point process (DPP) to maximize the conditional diversity of the selected subset. The proposed CDPruner is training-free and model-agnostic, allowing easy application to various MLLMs. Extensive experiments across diverse MLLMs show that CDPruner establishes new state-of-the-art on various vision-language benchmarks. By maximizing conditional diversity through DPP, the selected subset better represents the input images while closely adhering to user instructions, thereby preserving strong performance even with high reduction ratios. When applied to LLaVA, CDPruner reduces FLOPs by **95\%** and CUDA latency by **78\%**, while maintaining **94\%** of the original accuracy. Our code is available at https://github.com/Theia-4869/CDPruner. Qizhe Zhang, Lichen Li 0001, Ming Lu 0002, Yuan Zhang 0020, Junwen Pan, Qi She, Shanghang Zhang |
NeurIPS | 7 |
| 2024 | Boosting Weakly Supervised Object Localization and Segmentation With Domain AdaptionabstractWeakly supervised object localization (WSOL), adopting only image-level annotations to learn the pixel-level localization model, can release human resources in the annotation process. Most one-stage WSOL methods learn the localization model with multi-instance learning, making them only activate discriminative object parts rather than the whole object. In our work, we attribute this problem to the domain shift between the training and test process of WSOL and provide a novel perspective that views WSOL as a domain adaption (DA) task. Under this perspective, a DA-WSOL pipeline is elaborated to better assist WSOL with DA approaches by considering the specificities for the adaption of WSOL. Our DA-WSOL pipeline can discern the source-related and the Universum samples from other target samples based on a proposed target sampling strategy and then utilize them to solve the sample unbalancing and label unmatching between the source and target domain of WSOL. Experiments show that our pipeline outperforms SOTA methods on three WSOL benchmarks and can improve the performance of downstream weakly supervised semantic segmentation tasks. Lei Zhu 0012, Qi She, Qiushi Ren, Yanye Lu |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | Power Law in Deep Neural Networks: Sparse Network Generation and Continual Learning With Preferential AttachmentabstractTraining deep neural networks (DNNs) typically requires massive computational power. Existing DNNs exhibit low time and storage efficiency due to the high degree of redundancy. In contrast to most existing DNNs, biological and social networks with vast numbers of connections are highly efficient and exhibit scale-free properties indicative of the power law distribution, which can be originated by preferential attachment in growing networks. In this work, we ask whether the topology of the best performing DNNs shows the power law similar to biological and social networks and how to use the power law topology to construct well-performing and compact DNNs. We first find that the connectivities of sparse DNNs can be modeled by truncated power law distribution, which is one of the variations of the power law. The comparison of different DNNs reveals that the best performing networks correlated highly with the power law distribution. We further model the preferential attachment in DNNs evolution and find that continual learning in networks with growth in tasks correlates with the process of preferential attachment. These identified power law dynamics in DNNs can lead to the construction of highly accurate and compact DNNs based on preferential attachment. Inspired by the discovered findings, two novel applications have been proposed, including evolving optimal DNNs in sparse network generation and continual learning tasks with efficient network growth using power law dynamics. Experimental results indicate that the proposed applications can speed up training, save storage, and learn with fewer samples than other well-established baselines. Our demonstration of preferential attachment and power law in well-performing DNNs offers insight into designing and constructing more efficient deep learning. Lu Hou 0002, Qi She, Rosa H. M. Chan, James T. Kwok |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | Background-Aware Classification Activation Map for Weakly Supervised Object LocalizationabstractWeakly supervised object localization (WSOL) relaxes the requirement of dense annotations for object localization by using image-level annotation to supervise the learning process. However, most WSOL methods only focus on forcing the object classifier to produce high activation score on object parts without considering the influence of background locations, causing excessive background activations and ill-pose background score searching. Based on this point, our work proposes a novel mechanism called the background-aware classification activation map (B-CAM) to add background awareness for WSOL training. Besides aggregating an object image-level feature for supervision, our B-CAM produces an additional background image-level feature to represent the pure-background sample. This additional feature can provide background cues for the object classifier to suppress the background activations on object localization maps. Moreover, our B-CAM also trained a background classifier with image-level annotation to produce adaptive background scores when determining the binary localization mask. Experiments indicate the effectiveness of the proposed B-CAM on four different types of WSOL benchmarks, including CUB-200, ILSVRC, OpenImages, and VOC2012 datasets. Lei Zhu 0012, Qi She, Xiangxi Meng 0001, Mufeng Geng, Lujia Jin, Yibao Zhang, Qiushi Ren, Yanye Lu |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | Learning from Temporal Gradient for Semi-supervised Action RecognitionabstractSemi-supervised video action recognition tends to enable deep neural networks to achieve remarkable performance even with very limited labeled data. However, existing methods are mainly transferred from current image-based methods (e.g., FixMatch). Without specifically utilizing the temporal dynamics and inherent multimodal attributes, their results could be suboptimal. To better leverage the encoded temporal information in videos, we introduce temporal gradient as an additional modality for more attentive feature extraction in this paper. To be specific, our method explicitly distills the fine-grained motion representations from temporal gradient (TG) and imposes consistency across different modalities (i.e., RGB and TG). The performance of semi-supervised action recognition is significantly improved without additional computation or parameters during inference. Our method achieves the state-of-the-art performance on three video action recognition benchmarks (i.e., Kinetics-400, UCF-101, and HMDB-51) under several typical semi-supervised settings (i.e., different ratios of labeled data). Code is made available at https://github.com/lambert-x/video-semisup. Junfei Xiao, Longlong Jing, Lin Zhang 0040, Ju He, Qi She, Zongwei Zhou, Alan L. Yuille, Yingwei Li 0002 |
CVPR | 5 |
| 2022 | On Learning Contrastive Representations for Learning with Noisy LabelsabstractDeep neural networks are able to memorize noisy labels easily with a softmax cross entropy (CE) loss. Previous studies attempted to address this issue focus on incorporating a noise-robust loss function to the CE loss. However, the memorization issue is alleviated but still remains due to the non-robust CE loss. To address this issue, we focus on learning robust contrastive representations of data on which the classifier is hard to memorize the label noise under the CE loss. We propose a novel contrastive regularization function to learn such representations over noisy data where label noise does not dominate the representation learning. By theoretically investigating the representations induced by the proposed regularization function, we reveal that the learned representations keep information related to true labels and discard information related to corrupted labels. Moreover, our theoretical results also indicate that the learned representations are robust to the label noise. The effectiveness of this method is demonstrated with experiments on benchmark datasets. Qi She, A. Ian McLeod, Boyu Wang 0004 |
CVPR | 3 |
| 2022 | Weakly Supervised Object Localization as Domain AdaptionabstractWeakly supervised object localization (WSOL) focuses on localizing objects only with the supervision of image-level classification masks. Most previous WSOL methods follow the classification activation map (CAM) that localizes objects based on the classification structure with the multi-instance learning (MIL) mechanism. However, the MIL mechanism makes CAM only activate discriminative object parts rather than the whole object, weakening its performance for localizing objects. To avoid this problem, this work provides a novel perspective that models WSOL as a domain adaption (DA) task, where the score estimator trained on the source/image domain is tested on the target/pixel domain to locate objects. Under this perspective, a DA-WSOL pipeline is designed to better engage DA approaches into WSOL to enhance localization performance. It utilizes a proposed target sampling strategy to select different types of target samples. Based on these types of target samples, domain adaption localization (DAL) loss is elaborated. It aligns the feature distribution between the two domains by DA and makes the estimator perceive target domain cues by Universum regularization. Experiments show that our pipeline outperforms SOTA methods on multi benchmarks. Code are released at https://github.com/zh460045050/DA-WSOL_CVPR2022. Lei Zhu 0012, Qi She, Yunfei You, Boyu Wang 0004, Yanye Lu |
CVPR | 2 |
| 2022 | PDO-s3DCNNs: Partial Differential Operator Based Steerable 3D CNNsabstractSteerable models can provide very general and flexible equivariance by formulating equivariance requirements in the language of representation theory and feature fields, which has been recognized to be effective for many vision tasks. However, deriving steerable models for 3D rotations is much more difficult than that in the 2D case, due to more complicated mathematics of 3D rotations. In this work, we employ partial differential operators (PDOs) to model 3D filters, and derive general steerable 3D CNNs, which are called PDO-s3DCNNs. We prove that the equivariant filters are subject to linear constraints, which can be solved efficiently under various conditions. As far as we know, PDO-s3DCNNs are the most general steerable CNNs for 3D rotations, in the sense that they cover all common subgroups of SO(3) and their representations, while existing methods can only be applied to specific groups and representations. Extensive experiments show that our models can preserve equivariance well in the discrete domain, and outperform previous works on SHREC’17 retrieval and ISBI 2012 segmentation tasks with a low network complexity. Zhengyang Shen, Qi She, Jinwen Ma, Zhouchen Lin |
ICML | 3 |
| 2022 | CVPR 2020 continual learning in computer vision competition: Approaches, results, current challenges and future directions
Vincenzo Lomonaco, Lorenzo Pellegrini, Pau Rodríguez, Massimo Caccia, Qi She, Quentin Jodelet, Ruiping Wang 0001, Zheda Mai, David Vázquez 0001, German Ignacio Parisi, Nikhil Churamani, Marc Pickett, Issam H. Laradji, Davide Maltoni |
Artif. Intell. | 5 |
| 2022 | Towards lifelong object recognition: A dataset and benchmark
Chuanlin Lan, Qi Liu 0042, Qi She, Qihan Yang, Xinyue Hao 0001, Ivan Mashkin, Ka Shun Kei, Dong Qiang, Vincenzo Lomonaco, Xuesong Shi, Yimin Zhang 0002, Fei Qiao, Rosa H. M. Chan |
Pattern Recognit. | 4 |
| 2021 | TEAM-Net: Multi-modal Learning for Video Action Recognition with Partial Decoding
Qi She, Aljoscha Smolic |
BMVC | 2 |
| 2021 | Inter-intra Variant Dual Representations for Self-supervised Video Recognition
Lin Zhang 0040, Qi She, Zhengyang Shen, Changhu Wang |
BMVC | 2 |
| 2021 | Involution: Inverting the Inherence of Convolution for Visual RecognitionabstractConvolution has been the core ingredient of modern neural networks, triggering the surge of deep learning in vision. In this work, we rethink the inherent principles of standard convolution for vision tasks, specifically spatial-agnostic and channel-specific. Instead, we present a novel atomic operation for deep neural networks by inverting the aforementioned design principles of convolution, coined as involution. We additionally demystify the recent popular self-attention operator and subsume it into our involution family as an over-complicated instantiation. The proposed involution operator could be leveraged as fundamental bricks to build the new generation of neural networks for visual recognition, powering different deep learning models on several prevalent benchmarks, including ImageNet classification, COCO detection and segmentation, together with Cityscapes segmentation. Our involution-based models improve the performance of convolutional baselines using ResNet-50 by up to 1.6% top-1 accuracy, 2.5% and 2.4% bounding box AP, and 4.7% mean IoU absolutely while compressing the computational cost to 66%, 65%, 72%, and 57% on the above benchmarks, respectively. Code and pre-trained models for all the tasks are available at https://github.com/d-li14/involution. Jie Hu 0019, Changhu Wang, Xiangtai Li, Qi She, Lei Zhu 0012, Tong Zhang 0001, Qifeng Chen 0001 |
CVPR | 5 |
| 2021 | ACTION-Net: Multipath Excitation for Action RecognitionabstractSpatial-temporal, channel-wise, and motion patterns are three complementary and crucial types of information for video action recognition. Conventional 2D CNNs are computationally cheap but cannot catch temporal relationships; 3D CNNs can achieve good performance but are computationally intensive. In this work, we tackle this dilemma by designing a generic and effective module that can be embedded into 2D CNNs. To this end, we propose a spAtio-temporal, Channel and moTion excitatION (ACTION) module consisting of three paths: Spatio-Temporal Excitation (STE) path, Channel Excitation (CE) path, and Motion Excitation (ME) path. The STE path employs one channel 3D convolution to characterize spatio-temporal representation. The CE path adaptively recalibrates channel-wise feature responses by explicitly modeling interdependencies between channels in terms of the temporal aspect. The ME path calculates feature-level temporal differences, which is then utilized to excite motion-sensitive channels. We equip 2D CNNs with the proposed ACTION module to form a simple yet effective ACTION-Net with very limited extra computational cost. ACTION-Net is demonstrated by consistently outperforming 2D CNN counterparts on three backbones (i.e., ResNet-50, MobileNet V2 and BNInception) employing three datasets (i.e., Something-Something V2, Jester, and EgoGesture). Code is provided at https://github.com/V-Sense/ACTION-Net. Qi She, Aljoscha Smolic |
CVPR | 2 |
| 2021 | Learning the Superpixel in a Non-Iterative and Lifelong MannerabstractSuperpixel is generated by automatically clustering pixels in an image into hundreds of compact partitions, which is widely used to perceive the object contours for its excel-lent contour adherence. Although some works use the Convolution Neural Network (CNN) to generate high-quality superpixel, we challenge the design principles of these net-works, specifically for their dependence on manual labels and excess computation resources, which limits their flexibility compared with the traditional unsupervised segmentation methods. We target at redefining the CNN-based superpixel segmentation as a lifelong clustering task and pro-pose an unsupervised CNN-based method called LNS-Net. The LNS-Net can learn superpixel in a non-iterative and lifelong manner without any manual labels. Specifically, a lightweight feature embedder is proposed for LNS-Net to efficiently generate the cluster-friendly features. With those features, seed nodes can be automatically assigned to cluster pixels in a non-iterative way. Additionally, our LNS-Net can adapt the sequentially lifelong learning by rescaling the gradient of weight based on both channel and spatial context to avoid overfitting. Experiments show that the proposed LNS-Net achieves significantly better performance on three benchmarks with nearly ten times lower complexity compared with other state-of-the-art methods. Lei Zhu 0012, Qi She, Yanye Lu, Zhilin Lu 0002, Jie Hu 0019 |
CVPR | 2 |
| 2021 | MT-ORL: Multi-Task Occlusion Relationship LearningabstractRetrieving occlusion relation among objects in a single image is challenging due to sparsity of boundaries in image. We observe two key issues in existing works: firstly, lack of an architecture which can exploit the limited amount of coupling in the decoder stage between the two subtasks, namely occlusion boundary extraction and occlusion orientation prediction, and secondly, improper representation of occlusion orientation. In this paper, we propose a novel architecture called Occlusion-shared and Path-separated Network (OPNet), which solves the first issue by exploiting rich occlusion cues in shared high-level features and structured spatial information in task-specific low-level features. We then design a simple but effective orthogonal occlusion representation (OOR) to tackle the second issue. Our method surpasses the state-of-the-art methods by 6.1%/8.3% Boundary-AP and 6.5%/10% Orientation-AP on standard PIOD/BSDS ownership datasets. Code is available at https://github.com/fengpanhe/MT-ORL. Panhe Feng, Qi She, Lei Zhu 0012, Lin Zhang 0040, Zijian Feng, Changhu Wang, Chunpeng Li, Xuejing Kang, Anlong Ming |
ICCV | 2 |
| 2021 | MINE: Towards Continuous Depth MPI with NeRF for Novel View SynthesisabstractIn this paper, we propose MINE to perform novel view synthesis and depth estimation via dense 3D reconstruction from a single image. Our approach is a continuous depth generalization of the Multiplane Images (MPI) by introducing the NEural radiance fields (NeRF). Given a single image as input, MINE predicts a 4-channel image (RGB and volume density) at arbitrary depth values to jointly reconstruct the camera frustum and fill in occluded contents. The reconstructed and inpainted frustum can then be easily rendered into novel RGB or depth views using differentiable rendering. Extensive experiments on RealEstate10K, KITTI and Flowers Light Fields show that our MINE outperforms state-of-the-art by a large margin in novel view synthesis. We also achieve competitive results in depth estimation on iBims-1 and NYU-v2 without annotated depth supervision. Our source code is available at https://github.com/vincentfung13/MINE. Zijian Feng, Qi She, Henghui Ding, Changhu Wang, Gim Hee Lee |
ICCV | 3 |
| 2021 | Unifying Nonlocal Blocks for Neural NetworksabstractThe nonlocal-based blocks are designed for capturing long-range spatial-temporal dependencies in computer vision tasks. Although having shown excellent performance, they still lack the mechanism to encode the rich, structured information among elements in an image or video. In this paper, to theoretically analyze the property of these nonlocal-based blocks, we provide a new perspective to interpret them, where we view them as a set of graph filters generated on a fully-connected graph. Specifically, when choosing the Chebyshev graph filter, a unified formulation can be derived for explaining and analyzing the existing nonlocal-based blocks (e.g., nonlocal block, nonlocal stage, double attention block). Furthermore, by concerning the property of spectral, we propose an efficient and robust spectral nonlocal block, which can be more robust and flexible to catch long-range dependencies when inserted into deep neural networks than the existing nonlocal blocks. Experimental results demonstrate the clear-cut improvements and practical applicabilities of our method on image classification, action recognition, semantic segmentation, and person re-identification tasks. Code are available at https://github.com/zh460045050/SNL_ICCV2021. Lei Zhu 0012, Qi She, Yanye Lu, Xuejing Kang, Jie Hu 0019, Changhu Wang |
ICCV | 2 |
| 2020 | OpenLORIS-Object: A Robotic Vision Dataset and Benchmark for Lifelong Deep LearningabstractThe recent breakthroughs in computer vision have benefited from the availability of large representative datasets (e.g. ImageNet and COCO) for training. Yet, robotic vision poses unique challenges for applying visual algorithms developed from these standard computer vision datasets due to their implicit assumption over non-varying distributions for a fixed set of tasks. Fully retraining models each time a new task becomes available is infeasible due to computational, storage and sometimes privacy issues, while naïve incremental strategies have been shown to suffer from catastrophic forgetting. It is crucial for the robots to operate continuously under open-set and detrimental conditions with adaptive visual perceptual systems, where lifelong learning is a fundamental capability. However, very few datasets and benchmarks are available to evaluate and compare emerging techniques. To fill this gap, we provide a new lifelong robotic vision dataset ("OpenLORIS-Object") collected via RGB-D cameras. The dataset embeds the challenges faced by a robot in the real-life application and provides new benchmarks for validating lifelong object recognition algorithms. Moreover, we have provided a testbed of 9 state-of-the-art lifelong learning algorithms. Each of them involves 48 tasks with 4 evaluation metrics over the OpenLORIS-Object dataset. The results demonstrate that the object recognition task in the ever-changing difficulty environments is far from being solved and the bottlenecks are at the forward/backward transfer designs. Our dataset and benchmark are publicly available at https://lifelong-robotic-vision.github.io/dataset/object. Qi She, Xinyue Hao 0001, Qihan Yang, Chuanlin Lan, Vincenzo Lomonaco, Xuesong Shi, Yimin Zhang 0002, Fei Qiao, Rosa H. M. Chan |
ICRA | 1 |
| 2020 | Are We Ready for Service Robots? The OpenLORIS-Scene Datasets for Lifelong SLAMabstractService robots should be able to operate autonomously in dynamic and daily changing environments over an extended period of time. While Simultaneous Localization And Mapping (SLAM) is one of the most fundamental problems for robotic autonomy, most existing SLAM works are evaluated with data sequences that are recorded in a short period of time. In real-world deployment, there can be out-of-sight scene changes caused by both natural factors and human activities. For example, in home scenarios, most objects may be movable, replaceable or deformable, and the visual features of the same place may be significantly different in some successive days. Such out-of-sight dynamics pose great challenges to the robustness of pose estimation, and hence a robot’s long-term deployment and operation. To differentiate the forementioned problem from the conventional works which are usually evaluated in a static setting in a single run, the term lifelong SLAM is used here to address SLAM problems in an ever-changing environment over a long period of time. To accelerate lifelong SLAM research, we release the OpenLORIS-Scene datasets. The data are collected in real-world indoor scenes, for multiple times in each place to include scene changes in real life. We also design benchmarking metrics for lifelong SLAM, with which the robustness and accuracy of pose estimation are evaluated separately. The datasets and benchmark are available online at lifelong-robotic-vision.github.io/dataset/scene. Xuesong Shi, Dongjiang Li, Pengpeng Zhao 0005, Qinbin Tian, Qiwei Long, Chunhao Zhu, Jingwei Song, Fei Qiao, Yangquan Guo, Yimin Zhang 0002, Baoxing Qin, Wei Yang 0029, Fangshi Wang, Rosa H. M. Chan, Qi She |
ICRA | 18 |
| 2020 | Synthetic-Neuroscore: Using a neuro-AI interface for evaluating generative adversarial networks
Qi She, Alan F. Smeaton, Tomás Ward, Graham Healy |
Neurocomputing | 2 |
| 2019 | Neural Dynamics Discovery via Gaussian Process Recurrent Neural Networks
Qi She, Anqi Wu |
UAI | 1 |
| 2018 | Reduced-Rank Linear Dynamical SystemsabstractLinear Dynamical Systems are widely used to study the underlying patterns of multivariate time series. A basic assumption of these models is that high-dimensional time series can be characterized by some underlying, low-dimensional and time-varying latent states. However, existing approaches to LDS modeling mostly learn the latent space with a prescribed dimensionality. When dealing with short-length high- dimensional time series data, such models would be easily overfitted. We propose Reduced-Rank Linear Dynamical Systems (RRLDS), to automatically retrieve the intrinsic dimensionality of the latent space during model learning. Our key observation is that the rank of the dynamics matrix of LDS captures the intrinsic dimensionality, and the variational inference with a reduced-rank regularization finally leads to a concise, structured, and interpretable latent space. To enable our method to handle count-valued data, we introduce the dispersion-adaptive distribution to accommodate over-/ equal-/ and under-dispersion nature of such data. Results on both simulated and experimental data demonstrate our model can robustly learn latent space from short-length, noisy, count-valued data and significantly improve the prediction performance over the state-of-the-art methods. Qi She, Yuan Gao 0015, Rosa H. M. Chan |
AAAI | 1 |
| 2018 | Stochastic Dynamical Systems Based Latent Structure Discovery in High-Dimensional Time SeriesabstractThe brain encodes information by neural spiking activities, which can be described by time series data as spike counts. Latent Variable Models (LVMs) are widely used to study the unknown factors (i.e. the latent states) that are dependent in a network structure to modulate neural spiking activities. Yet, challenges in performing experiments to record on neuronal level commonly results in relatively short and noisy spike count data, which is insufficient to derive latent network structure by existing LVMs. Specifically, it is difficult to set the number of latent states. A small number of latent states may not be able to model the complexities of underlying systems, while a large number of latent states can lead to overfitting. Therefore, based on a specific LVMs called Linear Dynamical System (LDS), we propose a Reduced-Rank Linear Dynamical System (RRLDS) to estimate latent states and retrieve an optimal latent network structure from short, noisy spike count data. This framework estimates the model using Laplace approximation. To further handle count-valued data, we introduce the dispersion-adaptive distribution to accommodate over-/ equal-/ and under-dispersion nature of such data. Results on both simulated and experimental data demonstrate our model can robustly learn latent space from short-length, noisy, count-valued data and significantly improve the prediction performance over the state-of-the-art methods. Qi She, Rosa H. M. Chan |
ICASSP | 1 |