Di Xie

dblp:38/7733 · DBLP profile ↗
← Back
83ranked-venue papers
3as first author
55since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 50 · 1 first-author · 32 since 2021Artificial intelligence and machine learning · 43 · 1 first-author · 31 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 8 since 2021Computer networks · 5 · 1 first-authorSystems, architecture and hardware · 4 · 3 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Developing Dynamic Prediction Methods for Survival Time Lost in Chronic Kidney Disease Progression Under Competing Risks
abstract
Patients with chronic kidney disease (CKD) may progress to end-stage renal disease (ESRD) or die from other causes during long-term follow-up, making it essential to properly account for competing risks in prognostic modeling. However, most existing CKD prediction models rely on hazard-based measures, which primarily reflect relative effects, lack clinical interpretability, and cannot directly quantify survival time lost due to disease progression or death. To address this limitation, we adopted the restricted mean time lost (RMTL), an absolute and intuitive measure of survival time loss, and developed a dynamic prediction model under competing risks. Given the rich time-dependent covariate information in longitudinal CKD data and the clinical need for patients to understand their disease progression at different stages, we developed a dynamic RMTL prediction model that incorporates the landmark approach. The model captures how covariate effects evolve over time, offering insights into the dynamic impact of clinical variables, and enables individualized prediction of survival time loss over a future window from any given prediction time point. We evaluated its statistical properties through Monte Carlo simulations and demonstrated its practical utility using CKD patient data from the AASK cohort. The results showed that the proposed model yields accurate and robust estimates, captures time-varying covariate effects, and outperforms conventional static models in predictive performance. By quantifying survival time loss under competing risks, the dynamic RMTL model offers clinically interpretable and individualized risk estimates, supporting personalized risk assessment and intervention planning in chronic disease management.
Haoning Shen, Chengfeng Zhang, Xingzhi Wang, Di Xie, Shuyu Chen 0001, Pansheng Xue, Yuanying Chen, Yawen Hou, Zheng Chen 0024
IEEE J. Biomed. Health Informatics4
2026 EICSeg: Universal Medical Image Segmentation via Explicit In-Context Learning
abstract
Deep learning models for medical image segmentation often struggle with task-specific characteristics, limiting their generalization to unseen tasks with new anatomies, labels, or modalities. Retraining or fine-tuning these models requires substantial human effort and computational resources. To address this, in-context learning (ICL) has emerged as a promising paradigm, enabling query image segmentation by conditioning on example image-mask pairs provided as prompts. Unlike previous approaches that rely on implicit modeling or non-end-to-end pipelines, we redefine the core interaction mechanism in ICL as an explicit retrieval process, termed E-ICL, benefiting from the emergence of vision foundation models (VFMs). E-ICL captures dense correspondences between queries and prompts at minimal learning cost and leverages them to dynamically weight multi-class prompt masks. Built upon E-ICL, we propose EICSeg, the first end-to-end ICL framework that integrates complementary VFMs for universal medical image segmentation. Specifically, we introduce a lightweight SD-Adapter to bridge the distinct functionalities of the VFMs, enabling more accurate segmentation predictions. To fully exploit the potential of EICSeg, we further design a scalable self-prompt training strategy and an adaptive token-to-image prompt selection mechanism, facilitating both efficient training and inference. EICSeg is trained on 47 datasets covering diverse modalities and segmentation targets. Experiments on nine unseen datasets demonstrate its strong few-shot generalization ability, achieving an average Dice score of 74.0%, outperforming existing in-context and few-shot methods by 4.5%, and reducing the gap to task-specific models to 10.8%. Even with a single prompt, EICSeg achieves a competitive average Dice score of 60.1%. Notably, it performs automatic segmentation without manual prompt engineering, delivering results comparable to interactive models while requiring minimal labeled data. Source code will be available at https://github.com/zerone-fg/EICSeg.
Shiao Xie, Liangjun Zhang, Ziwei Niu, Fanfan Ye, Qiaoyong Zhong, Di Xie, Yen-Wei Chen 0001, Lanfen Lin
IEEE Trans. Medical Imaging6
2025 Gaze Label Alignment: Alleviating Domain Shift for Gaze Estimation
abstract
Gaze estimation methods encounter significant performance deterioration when being evaluated across different domains, because of the domain gap between the testing and training data. Existing methods try to solve this issue by reducing the deviation of data distribution, however, they ignore the existence of label deviation in the data due to the acquisition mechanism of the gaze label and the individual physiological differences. In this paper, we first point out that the influence brought by the label deviation cannot be ignored, and propose a gaze label alignment algorithm (GLA) to eliminate the label distribution deviation. Specifically, we first train the feature extractor on all domains to get domain invariant features, and then select an anchor domain to train the gaze regressor. We predict the gaze label on remaining domains and use a mapping function to align the labels. Finally, these aligned labels can be used to train gaze estimation models. Therefore, our method can be combined with any existing method. Experimental results show that our GLA method can effectively alleviate the label distribution shift, and SOTA gaze estimation methods can be further improved obviously.
Guanzhong Zeng, Zefu Xu, Pengwei Yin, Wenqi Ren, Di Xie
AAAI6
2025 A Real-Time Animation Blending Scheme for Cavalry Combat Skills via Interpolation Algorithms
abstract
With the vigorous development of computer graphics, real-time animation technology, and artificial intelligence, motion animation blending technology has been widely applied in multiple fields. Aiming at the problems of motion blending and animation transition, this study proposes a real-time and highly scalable animation blending scheme. This scheme deeply explores the martial arts culture of ancient cavalry, integrates multiple technologies, uses interpolation algorithms to regulate animation blending, and realizes collision detection and feedback simulation through the construction of an RPG-based combat system for cavalry in the early Ming Dynasty. Experiments show that the scheme significantly optimizes the movement effect of mounts, has more advantages than traditional methods, verifies the feasibility of the scheme, enhances the simulation realism and immersion of the combat system, and provides new ideas for related research.
Xiongjie Tao, Haixiao Gong, Bin Hu 0031, Yingli Zhao, Di Xie, Peiping Li
HPCC6
2025 From Decoupling to Adaptive Transformation: a Wider Optimization Space for PTQ
abstract
Post-Training low-bit Quantization (PTQ) is useful to accelerate DNNs due to its high efficiency, the current SOTAs of which mostly adopt feature reconstruction with self-distillation finetuning. However, when bitwidth goes to be extremely low, we find the current reconstruction optimization space is not optimal. Considering all possible parameters and the ignored fact that integer weight can be obtained early before actual inference, we thoroughly explore different optimization space by quant-step decoupling, where a wider PTQ optimization space, which consistently makes a better optimum, is found out. Based on these, we propose an Adaptive Quantization Transformation (AdaQTransform) for PTQ reconstruction, which makes the quantized output feature better fit the FP32 counterpart with adaptive per-channel transformation, thus achieves lower feature reconstruction error. In addition, it incurs negligible extra finetuning cost and no extra inference cost. Based on AdaQTransform, for the first time, we build a general quantization setting paradigm subsuming current PTQs, QATs and other potential forms. Experiments demonstrate AdaQTransform expands the optimization space for PTQ and helps current PTQs find a better optimum over CNNs, ViTs, LLMs and image super-resolution networks, e.g., it improves NWQ by 5.7% on ImageNet for W2A2-MobileNet-v2.
Zhaojing Wen, Qiulin Zhang, Rudan Chen, Xichao Yang, Di Xie
ICLR6
2025 Unbiased Evaluation of Large Language Models from a Causal Perspective
abstract
Benchmark contamination has become a significant concern in the LLM evaluation community. Previous Agents-as-an-Evaluator address this issue by involving agents in the generation of questions. Despite their success, the biases in Agents-as-an-Evaluator methods remain largely unexplored. In this paper, we present a theoretical formulation of evaluation bias, providing valuable insights into designing unbiased evaluation protocols. Furthermore, we identify two type of bias in Agents-as-an-Evaluator through carefully designed probing tasks on a minimal Agents-as-an-Evaluator setup. To address these issues, we propose the Unbiased Evaluator, an evaluation protocol that delivers a more comprehensive, unbiased, and interpretable assessment of LLMs. Extensive experiments reveal significant room for improvement in current LLMs. Additionally, we demonstrate that the Unbiased Evaluator not only offers strong evidence of benchmark contamination but also provides interpretable evaluation results.
Meilin Chen, Di Xie, Weijie Chen 0006
ICML4
2025 VQCounter: Designing Visual Prompt Queue for Accurate Open-World Counting
abstract
Class-agnostic counting enables enumerating arbitrary object classes beyond those seen during training. Recent studies attempted to exploit the potential of visual foundation models such as GroundingDINO. Despite the considerable progress, we observe certain shortcomings, including the limited diversity of visual prompts and suboptimal training regimen. To address these issues, we introduce VQCounter, which incorporates a visual prompt queue mechanism designed to enrich the diversity of visual prompts. A random modality switching strategy is proposed during training to strengthen both textual and visual modalities. Besides, in light of weak point supervision, a Voronoi diagram-based cost (VoronoiCost) is designed to improve Hungarian matching, leading to more stable and faster convergence. Building upon the Voronoi diagram, we also propose a novel set of more stringent evaluation metrics, which take point localization into account. Extensive experiments on the FSC-147 and CARPK datasets demonstrate that VQCounter achieves state-of-the-art performance in both zero-shot and few-shot settings, significantly outperforming existing methods across nearly all evaluations.
Fanfan Ye, Yiqi Fan, Qiaoyong Zhong, Shicai Yang, Di Xie, Jie Song 0011, Mingli Song
IJCAI5
2025 Training-Free Test-Time Adaptation via Shape and Style Guidance for Vision-Language Models
abstract
Test-time adaptation with pre-trained vision-language models shows impressive zero-shot classification abilities, and training-free methods further improve the performance without any optimization burden. However, existing training-free test-time adaptation methods typically rely on entropy criteria to select the visual features and update the visual caches, while ignoring the generalizable factors, such as shape-sensitive and style-insensitive factors. In this paper, we propose a novel shape and style guidance method (SSG) for training-free test-time adaptation in vision-language models, aiming to highlight the shape-sensitive (SHS) and style-insensitive (STI) factors in addition to entropy criteria. Specifically, SSG perturbs the raw test image with shape and style corruption operations, and measures the prediction difference between the raw and corrupted one as perturbed prediction difference (PPD). Based on the PPD measurement, SSG reweights the high-confidence visual features and corresponding predictions, aiming to highlight the effect of SHS and STI factors during the test-time procedure. Furthermore, SSG takes both PPD and entropy into consideration to update the visual cache, aiming to maintain the stored sample with high entropy and generalizable factors. Extensive experimental results on out-of-distribution and cross-domain benchmark datasets demonstrate that our proposed SSG consistently outperforms previous state-of-the-art methods while also exhibiting promising computational efficiency.
Shenglong Zhou 0002, Manjiang Yin, Leiyu Sun, Shicai Yang, Di Xie
NeurIPS5
2025 Semantic-aware contrastive learning via multi-prompt alignment
Ming Kong 0001, Luyuan Chen, Di Xie
Mach. Learn.5
2025 Adapt Anything: Tailor Any Image Classifier Across Domains and Categories Using Text-to-Image Diffusion Models
abstract
We study a novel problem in this paper, that is, if a modern text-to-image diffusion model can tailor any image classifier across domains and categories. Existing domain adaption works exploit both source and target data for domain alignment so as to transfer the knowledge from the labeled source data to the unlabeled target data. However, as the development of text-to-image diffusion models, we wonder if the high-fidelity synthetic data can serve as a surrogate of the source data in real world. In this way, we do not need to collect and annotate the source data for each image classification task in a one-for-one manner. Instead, we utilize only one off-the-shelf text-to-image model to synthesize images with labels derived from text prompts, and then leverage them as a bridge to dig out the knowledge from the task-agnostic text-to-image generator to the task-oriented image classifier via domain adaptation. Such a one-for-all adaptation paradigm allows us to adapt anything in the world using only one text-to-image generator as well as any unlabeled target data. Extensive experiments validate the feasibility of this idea, which even surprisingly surpasses the state-of-the-art domain adaptation works using the source data collected and annotated in real world.
Weijie Chen 0006, Haoyu Wang 0016, Shicai Yang, Lei Zhang 0054, Wei Wei 0008, Yanning Zhang 0001, Luojun Lin, Di Xie, Yueting Zhuang
IEEE Trans. Big Data8
2024 Arbitrary-Scale Point Cloud Upsampling by Voxel-Based Network with Latent Geometric-Consistent Learning
abstract
Recently, arbitrary-scale point cloud upsampling mechanism became increasingly popular due to its efficiency and convenience for practical applications. To achieve this, most previous approaches formulate it as a problem of surface approximation and employ point-based networks to learn surface representations. However, learning surfaces from sparse point clouds is more challenging, and thus they often suffer from the low-fidelity geometry approximation. To address it, we propose an arbitrary-scale Point cloud Upsampling framework using Voxel-based Network (PU-VoxelNet). Thanks to the completeness and regularity inherited from the voxel representation, voxel-based networks are capable of providing predefined grid space to approximate 3D surface, and an arbitrary number of points can be reconstructed according to the predicted density distribution within each grid cell. However, we investigate the inaccurate grid sampling caused by imprecise density predictions. To address this issue, a density-guided grid resampling method is developed to generate high-fidelity points while effectively avoiding sampling outliers. Further, to improve the fine-grained details, we present an auxiliary training supervision to enforce the latent geometric consistency among local surface patches. Extensive experiments indicate the proposed approach outperforms the state-of-the-art approaches not only in terms of fixed upsampling rates but also for arbitrary-scale upsampling. The code is available at https://github.com/hikvision-research/3DVision
Jingjing Wang 0005, Di Xie, Shiliang Pu
AAAI4
2024 CLIP-Gaze: Towards General Gaze Estimation via Visual-Linguistic Model
abstract
Gaze estimation methods often experience significant performance degradation when evaluated across different domains, due to the domain gap between the testing and training data. Existing methods try to address this issue using various domain generalization approaches, but with little success because of the limited diversity of gaze datasets, such as appearance, wearable, and image quality. To overcome these limitations, we propose a novel framework called CLIP-Gaze that utilizes a pre-trained vision-language model to leverage its transferable knowledge. Our framework is the first to leverage the vision-and-language cross-modality approach for gaze estimation task. Specifically, we extract gaze-relevant feature by pushing it away from gaze-irrelevant features which can be flexibly constructed via language descriptions. To learn more suitable prompts, we propose a personalized context optimization method for text prompt tuning. Furthermore, we utilize the relationship among gaze samples to refine the distribution of gaze-relevant features, thereby improving the generalization capability of the gaze estimation model. Extensive experiments demonstrate the excellent performance of CLIP-Gaze over existing methods on four cross-domain evaluations.
Pengwei Yin, Guanzhong Zeng, Di Xie
AAAI4
2024 LG-Gaze: Learning Geometry-Aware Continuous Prompts for Language-Guided Gaze Estimation
Pengwei Yin, Guanzhong Zeng, Di Xie
ECCV (83)4
2024 Optimization Method for Fractal Image Compression Based on Self-similarity Evaluation and Gradient Bisection Algorithm
Caixu Xu, Di Xie, Minglang Chen
ICIC (7)2
2024 Better Together: Data-Free Multi-Student Coevolved Distillation
Weijie Chen 0006, Yunyi Xuan, Shicai Yang, Di Xie, Luojun Lin, Yueting Zhuang
Knowl. Based Syst.4
2024 Palm Vein Recognition Under Unconstrained and Weak-Cooperative Conditions
abstract
Contactless palm vein has attracted significant attention for its high security, stability, and user-friendliness. However, current contactless palm vein recognition predominantly relies on databases collected from platforms with spatial and temporal constrained design, which inadequately reflect relaxed palm vein imaging circumstances. This paper proposes a novel manner called on-the-fly palm vein that frees the user’s palm from spatial and temporal constraints, enabling palm vein recognition under unconstrained and weak-cooperative conditions. Firstly, Designing efficient and user-friendly palm vein imaging and authentication via two dynamic palm motions is proposed, resulting in an on-the-fly palm vein recognition platform. Next, a large-scale and challenging palm vein database, SCUT Palm Vein Database Version 1 (SCUT_PV_v1), is constructed. It is the first palm vein database with images collected under unconstrained and weak-cooperative conditions, encompassing a wider range of palm pose variations, grayscale variations, and lower-quality images. Finally, a lightweight and efficient Adaptive Margin Palm Vein Authentication Network (AMPVNet) is proposed as a baseline for the SCUT_PV_v1, where a vein pattern-specific convolutional neural network (CNN) is designed to extract features and a tailored online data augmentation method, combining Random Perspective Transformation (RPT) with Random Grayscale Adjustment (RGA), is proposed to enrich the diversify of out-of-plane palm pose and grayscale variations. Extensive experimental results demonstrate the effectiveness of our proposed methods. As the first work for palm vein recognition under unconstrained and weak-cooperation conditions, the AMPVNet achieves a promising accuracy and computation result while maintaining robustness to palm pose and grayscale variations. The SCUT_PV_ v1 database will be public at https://github.com/SCUT-BIP-Lab/SCUT_PV_v1.
Dacan Luo, Yitao Qiao, Di Xie, Wenxiong Kang
IEEE Trans. Inf. Forensics Secur.3
2023 A Likelihood Probability-Based Online Summarization Ranking Model
Shuhao Yue, Dunhui Yu, Di Xie
ADMA (2)3
2023 Rethinking the Approximation Error in 3D Surface Fitting for Point Cloud Normal Estimation
abstract
Most existing approaches for point cloud normal estimation aim to locally fit a geometric surface and calculate the normal from the fitted surface. Recently, learning-based methods have adopted a routine of predicting pointwise weights to solve the weighted least-squares surface fitting problem. Despite achieving remarkable progress, these methods overlook the approximation error of the fitting problem, resulting in a less accurate fitted surface. In this paper, we first carry out in-depth analysis of the approximation error in the surface fitting problem. Then, in order to bridge the gap between estimated and precise surface normals, we present two basic design principles: 1) applies the Z-direction Transform to rotate local patches for a better surface fitting with a lower approximation error; 2) models the error of the normal estimation as a learnable term. We implement these two principles using deep neural networks, and integrate them with the state-of-the-art (SOTA) normal estimation methods in a plug-and-play manner. Extensive experiments verify our approaches bring benefits to point cloud normal estimation and push the frontier of state-of-the-art performance on both synthetic and real-world datasets. The code is available at https://github.com/hikvision-research/3DVision.
Jingjing Wang 0005, Di Xie, Shiliang Pu
CVPR4
2023 MDR-MFI:Multi-Branch Decoupled Regression and Multi-Scale Feature Interaction for Partial-to-Partial Cloud Registration
abstract
Point cloud registration is a fundamental task in the 3D vision field. Many previous works adopt the regression model to estimate the transformation parameters. However, these methods couple the estimation of rotation and translation via a single regression branch, which suffers from the mutual interference among rotation and translation. In addition, previous methods extract and interact features in a single scale, which ignores the rich information from multiple scales. To address above issues, in this paper, we propose a multi-branch decoupled regression and multi-scale feature interaction (MDR-MFI) framework for point cloud registration. Firstly, we decouple the estimation of 7 transformation parameters via multiple regression branches. The decoupled structure effectively mitigates the mutual interference among 7 parameters, resulting in improved performance. Secondly, we propose a multi-scale feature extraction and interaction framework to encourage the network to learn more discriminative features. Experimental results demonstrate that our method achieves state-of-the-art performance on public datasets. The code is available at https://github.com/hikvision-research/3DVision.
Weidong Dai, Jingjing Wang 0005, Di Xie, Shiliang Pu
ICASSP4
2023 Single Domain Dynamic Generalization for Iris Presentation Attack Detection
abstract
Iris presentation attack detection (PAD) has achieved great success under intra-domain settings but easily degrades on unseen domains. Conventional domain generalization methods mitigate the gap by learning domain-invariant features. However, they ignore the discriminative information in the domain-specific features. Moreover, we usually face a more realistic scenario with only one single domain available for training. To tackle the above issues, we propose a Single Domain Dynamic Generalization (SDDG) framework, which simultaneously exploits domain-invariant and domain-specific features on a per-sample basis and learns to generalize to various unseen domains with numerous natural images. Specifically, a dynamic block is designed to adaptively adjust the network with a dynamic adaptor. And an information maximization loss is further combined to increase diversity. The whole network is integrated into the meta-learning paradigm. We generate amplitude perturbed images and cover diverse domains with natural images. Therefore, the network can learn to generalize to the perturbed domains in the meta-test phase. Extensive experiments show the proposed method is effective and outperforms the state-of-the-art on LivDet-Iris 2017 dataset.
Yachun Li, Jingjing Wang 0005, Yuhui Chen, Di Xie, Shiliang Pu
ICASSP4
2023 PRIME: 3D Human Pose and Body Shape Recovery with Perspective Projection
abstract
Existing monocular 3D human pose and body shape (HPS) estimation methods make the coplanar assumption and use weak perspective projection in order to simplify the problem setting for images in the wild. However, weak perspective projection inevitably introduce prediction biases. To address this issue, we propose a plug-and-play Perspective Residual Log-likehood on Monocular 3D HPS Estimation (PRIME) module to significantly improve the accuracy of monocular 3D HPS estimation with trivial sacrifice on running time. PRIME applies full perspective projection to construct 2D re-projection loss or extract mesh-alignment features. Specifically, PRIME estimates the distribution of 2D joints and scale to calculate the perspective translation with the focal length. Further, we introduce side view constrain (SVC) of 2D joints to reduce the ambiguity of 3D HPS recovery. Experimental results demonstrate the effectiveness of our method.
Baobei Xu, Shukai Fang, Shicai Yang, Di Xie, Shiliang Pu
ICASSP5
2023 Learning Expressive And Generalizable Motion Features For Face Forgery Detection
abstract
Previous face forgery detection methods mainly focus on appearance features, which may be easily attacked by sophisticated manipulation. Considering the majority of current face manipulation methods generate fake faces based on a single frame, which do not take frame consistency and coordination into consideration, artifacts on frame sequences are more effective for face forgery detection. However, current sequence-based face forgery detection methods use general video classification networks directly, which discard the special and discriminative motion information for face manipulation detection. To this end, we propose an effective sequence-based forgery detection framework based on an existing video classification method. To make the motion features more expressive for manipulation detection, we propose an alternative motion consistency block instead of the original motion features module. To make the learned features more generalizable, we propose an auxiliary anomaly detection block. With these two specially designed improvements, we make a general video classification network achieve promising results on three popular face forgery datasets.
Jingyi Zhang 0003, Peng Zhang 0075, Jingjing Wang 0005, Di Xie, Shiliang Pu
ICASSP4
2023 HPFTN: Hierarchical Progressive Fusion Transformer Network for Video Denoising
abstract
This paper presents a simple yet effective approach to modeling space-time correspondences in the context of video denoising. Unlike most existing approaches, our method, namely HPFTN, can operate end-to-end on consecutive frames without motion estimation. To do so, the proposed hierarchical patch matching module uses a multiple scales correspondence matching scheme to effectively build correspondences between neighbor frames and the current frame, lowering the computational cost. The progressive feature fusion module further enhances the current frame representation ability by extensively exploiting spatial-temporal correlations from multiple frames on patch level. Finally, the pyramid transformer reconstruction module efficiently leverages both high-level semantic and low-level fine-grained detailed features to predict clean video frames. Extensive quantitative and qualitative experiments validate the effectiveness of our proposed model. Our source code will be released.
Shuaitao Zhang, Di Xie, Shiliang Pu
ICASSP4
2023 Unsupervised Prompt Tuning for Text-Driven Object Detection
abstract
Grounded language-image pre-trained models have shown strong zero-shot generalization to various downstream object detection tasks. Despite their promising performance, the models rely heavily on the laborious prompt engineering. Existing works typically address this problem by tuning text prompts using downstream training data in a few-shot or fully supervised manner. However, a rarely studied problem is to optimize text prompts without using any annotations. In this paper, we delve into this problem and propose an Unsupervised Prompt Tuning framework for text-driven object detection, which is composed of two novel mean teaching mechanisms. In conventional mean teaching, the quality of pseudo boxes is expected to optimize better as the training goes on, but there is still a risk of overfitting noisy pseudo boxes. To mitigate this problem, 1) we propose Nested Mean Teaching, which adopts nested-annotation to supervise teacher-student mutual learning in a bi-level optimization manner; 2) we propose Dual Complementary Teaching, which employs an offline pre-trained teacher and an online mean teacher via data-augmentation-based complementary labeling so as to ensure learning without accumulating confirmation bias. By integrating these two mechanisms, the proposed unsupervised prompt tuning framework achieves significant performance improvement on extensive object detection datasets.
Weizhen He, Weijie Chen 0006, Shicai Yang, Di Xie, Luojun Lin, Donglian Qi, Yueting Zhuang
ICCV5
2023 Distilling Vision-Language Foundation Models: A Data-Free Approach via Prompt Diversification
abstract
Data-Free Knowledge Distillation (DFKD) has shown great potential in creating a compact student model while alleviating the dependency on real training data by synthesizing surrogate data. However, prior arts are seldom discussed under distribution shifts, which may be vulnerable in real-world applications. Recent Vision-Language Foundation Models, e.g., CLIP, have demonstrated remarkable performance in zero-shot out-of-distribution generalization, yet consuming heavy computation resources. In this paper, we discuss the extension of DFKD to Vision-Language Foundation Models without access to the billion-level image-text datasets. The objective is to customize a student model for distribution-agnostic downstream tasks with given category concepts, inheriting the out-of-distribution generalization capability from the pre-trained foundation models. In order to avoid generalization degradation, the primary challenge of this task lies in synthesizing diverse surrogate images driven by text prompts. Since not only category concepts but also style information are encoded in text prompts, we propose three novel Prompt Diversification methods to encourage image synthesis with diverse styles, namely Mix-Prompt, Random-Prompt, and Contrastive-Prompt. Experiments on out-of-distribution generalization datasets demonstrate the effectiveness of the proposed methods, with Contrastive-Prompt performing the best.
Yunyi Xuan, Weijie Chen 0006, Shicai Yang, Di Xie, Luojun Lin, Yueting Zhuang
ACM Multimedia4
2023 Up to Thousands-fold Storage Saving: Towards Efficient Data-Free Distillation of Large-Scale Visual Classifiers
abstract
Data-Free Knowledge Distillation (DFKD) has started to make breakthroughs in classification tasks for large-scale datasets such as ImageNet-1k. Despite the encouraging results achieved, these modern DFKD methods still suffer from the massive waste of system storage and I/O resources. They either synthesize and store a vast amount of pseudo data or build thousands of generators. In this work, we introduce a storage-efficient scheme called Class-Expanding DFKD (CE-DFKD). It allows us to reduce storage costs by orders of magnitude in large-scale tasks using just one or a few generators without explicitly storing any data. The key to the success of our approach lies in alleviating the mode collapse of the generator by expanding its collapse range. Specifically, we first investigate and address the optimization conflict of previous single-generator-based DFKD methods by introducing conditional constraints. Then, we propose two class-expanding strategies to enrich the conditional information of the generator from both inter-class and intra-class perspectives. With the diversity of generated samples significantly enhanced, the proposed CE-DFKD outperforms existing methods by a large margin while achieving up to thousands of times storage savings. Besides the ImageNet-1k, the proposed CE-DFKD is compatible with widely used small-scale datasets and can be scaled to the more complex ImageNet-21k-P dataset, which was previously unreported in prior DFKD methods.
Fanfan Ye, Bingyi Lu, Qiaoyong Zhong, Di Xie
ACM Multimedia5
2023 Comparative validation of machine learning algorithms for surgical workflow and skill analysis with the HeiChole benchmark
abstract
PURPOSE: Surgical workflow and skill analysis are key technologies for the next generation of cognitive surgical assistance systems. These systems could increase the safety of the operation through context-sensitive warnings and semi-autonomous robotic assistance or improve training of surgeons via data-driven feedback. In surgical workflow analysis up to 91% average precision has been reported for phase recognition on an open data single-center video dataset. In this work we investigated the generalizability of phase recognition algorithms in a multicenter setting including more difficult recognition tasks such as surgical action and surgical skill. METHODS: To achieve this goal, a dataset with 33 laparoscopic cholecystectomy videos from three surgical centers with a total operation time of 22 h was created. Labels included framewise annotation of seven surgical phases with 250 phase transitions, 5514 occurences of four surgical actions, 6980 occurences of 21 surgical instruments from seven instrument categories and 495 skill classifications in five skill dimensions. The dataset was used in the 2019 international Endoscopic Vision challenge, sub-challenge for surgical workflow and skill analysis. Here, 12 research teams trained and submitted their machine learning algorithms for recognition of phase, action, instrument and/or skill assessment. RESULTS: F1-scores were achieved for phase recognition between 23.9% and 67.7% (n = 9 teams), for instrument presence detection between 38.5% and 63.8% (n = 8 teams), but for action recognition only between 21.8% and 23.3% (n = 5 teams). The average absolute error for skill assessment was 0.78 (n = 1 team). CONCLUSION: Surgical workflow and skill analysis are promising technologies to support the surgical team, but there is still room for improvement, as shown by our comparison of machine learning algorithms. This novel HeiChole benchmark can be used for comparable evaluation and validation of future work. In future studies, it is of utmost importance to create more open, high-quality datasets in order to allow the development of artificial intelligence and cognitive robotics in surgery.
Martin Wagner 0001, Beat P. Müller-Stich, Anna Kisilenko, Patrick Heger, Lars Mündermann, David M. Lubotsky, Tornike Davitashvili, Manuela Capek, Annika Reinke, Carissa Reid, Tong Yu 0009, Armine Vardazaryan, Chinedu Innocent Nwoye, Nicolas Padoy, Eungjoo Lee 0001, Constantin Disch, Hans Meine, Tong Xia, Fucang Jia, Satoshi Kondo, Wolfgang Reiter, Yueming Jin, Yonghao Long 0001, Meirui Jiang, Qi Dou 0001, Pheng-Ann Heng, Isabell Twick, Kadir Kirtaç, Enes Hosgor, Jon Lindström Bolmgren, Michael Stenzel, Björn von Siemens, Zhenxiao Ge, Haiming Sun, Di Xie, Mengqi Guo, Daochang Liu, Hannes Kenngott, Felix Nickel, Moritz von Frankenberg, Franziska Mathis-Ullrich, Annette Kopp-Schneider, Lena Maier-Hein, Stefanie Speidel, Sebastian Bodenstedt
Medical Image Anal.39
2023 Few-Shot Class-Incremental Learning by Sampling Multi-Phase Tasks
abstract
New classes arise frequently in our ever-changing world, e.g., emerging topics in social media and new types of products in e-commerce. A model should recognize new classes and meanwhile maintain discriminability over old classes. Under severe circumstances, only limited novel instances are available to incrementally update the model. The task of recognizing few-shot new classes without forgetting old classes is called few-shot class-incremental learning (FSCIL). In this work, we propose a new paradigm for FSCIL based on meta-learning by LearnIng Multi-phase Incremental Tasks (Limit), which synthesizes fake FSCIL tasks from the base dataset. The data format of fake tasks is consistent with the 'real' incremental tasks, and we can build a generalizable feature space for the unseen tasks through meta-learning. Besides, Limit also constructs a calibration module based on transformer, which calibrates the old class classifiers and new class prototypes into the same scale and fills in the semantic gap. The calibration module also adaptively contextualizes the instance-specific embedding with a set-to-set function. Limit efficiently adapts to new classes and meanwhile resists forgetting over old classes. Experiments on three benchmark datasets (CIFAR100, miniImageNet, and CUB200) and large-scale dataset, i.e., ImageNet ILSVRC2012 validate that Limit achieves state-of-the-art performance.
Da-Wei Zhou 0001, Han-Jia Ye, Di Xie, Shiliang Pu, De-Chuan Zhan
IEEE Trans. Pattern Anal. Mach. Intell.4
2022 Topology-Aware Convolutional Neural Network for Efficient Skeleton-Based Action Recognition
abstract
In the context of skeleton-based action recognition, graph convolutional networks (GCNs) have been rapidly developed, whereas convolutional neural networks (CNNs) have received less attention. One reason is that CNNs are considered poor in modeling the irregular skeleton topology. To alleviate this limitation, we propose a pure CNN architecture named Topology-aware CNN (Ta-CNN) in this paper. In particular, we develop a novel cross-channel feature augmentation module, which is a combo of map-attend-group-map operations. By applying the module to the coordinate level and the joint level subsequently, the topology feature is effectively enhanced. Notably, we theoretically prove that graph convolution is a special case of normal convolution when the joint dimension is treated as channels. This confirms that the topology modeling power of GCNs can also be implemented by using a CNN. Moreover, we creatively design a SkeletonMix strategy which mixes two persons in a unique manner and further boosts the performance. Extensive experiments are conducted on four widely used datasets, i.e. N-UCLA, SBU, NTU RGB+D and NTU RGB+D 120 to verify the effectiveness of Ta-CNN. We surpass existing CNN-based methods significantly. Compared with leading GCN-based methods, we achieve comparable performance with much less complexity in terms of the required GFLOPs and parameters.
Kailin Xu, Fanfan Ye, Qiaoyong Zhong, Di Xie
AAAI4
2022 Point Cloud Upsampling via Cascaded Refinement Network
Jingjing Wang 0005, Di Xie, Shiliang Pu
ACCV (1)4
2022 Multi-scale Wavelet Transformer for Face Forgery Detection
Jingjing Wang 0005, Peng Zhang 0075, Chunmao Wang, Di Xie, Shiliang Pu
ACCV (6)5
2022 Label Matching Semi-Supervised Object Detection
abstract
Semi-supervised object detection has made significant progress with the development of mean teacher driven self-training. Despite the promising results, the label mismatch problem is not yet fully explored in the previous works, leading to severe confirmation bias during self-training. In this paper, we delve into this problem and propose a simple yet effective LabelMatch framework from two different yet complementary perspectives, i.e., distribution-level and instance-level. For the former one, it is reasonable to approximate the class distribution of the unlabeled data from that of the labeled data according to Monte Carlo Sampling. Guided by this weakly supervision cue, we introduce a re-distribution mean teacher, which leverages adaptive label-distribution-aware confidence thresholds to generate unbiased pseudo labels to drive student learning. For the latter one, there exists an overlooked label assignment ambiguity problem across teacher-student models. To remedy this issue, we present a novel label assignment mechanism for self-training framework, namely proposal self-assignment, which injects the proposals from student into teacher and generates accurate pseudo labels to match each proposal in the student model accordingly. Experiments on both MS-COCO and PASCAL-VOC datasets demonstrate the considerable superiority of our proposed framework to other state-of-the-arts. Code will be available at https://github.com/HIK-LAB/SSOD.
Weijie Chen 0006, Shicai Yang, Yunyi Xuan, Jie Song 0011, Di Xie, Shiliang Pu, Mingli Song, Yueting Zhuang
CVPR6
2022 Slimmable Domain Adaptation
abstract
Vanilla unsupervised domain adaptation methods tend to optimize the model with fixed neural architecture, which is not very practical in real-world scenarios since the target data is usually processed by different resource-limited devices. It is therefore of great necessity to facilitate architecture adaptation across various devices. In this paper, we introduce a simple framework, Slimmable Domain Adaptation, to improve cross-domain generalization with a weight-sharing model bank, from which models of different capacities can be sampled to accommodate different accuracy-efficiency trade-offs. The main challenge in this frame-work lies in simultaneously boosting the adaptation performance of numerous models in the model bank. To tackle this problem, we develop a Stochastic EnsEmble Distillation method to fully exploit the complementary knowledge in the model bank for inter-model interaction. Nevertheless, considering the optimization conflict between inter-model interaction and intra-model adaptation, we augment the existing bi-classifier domain confusion architecture into an Optimization-Separated Tri-Classifier counterpart. After optimizing the model bank, architecture adaptation is leveraged via our proposed Unsupervised Performance Evaluation Metric. Under various resource constraints, our framework surpasses other competing approaches by a very large margin on multiple benchmarks. It is also worth emphasizing that our framework can preserve the performance improvement against the source-only model even when the computing complexity is reduced to 1/64. Code will be available at https://github.com/HIK-LAB/SlimDA.
Rang Meng, Weijie Chen 0006, Shicai Yang, Jie Song 0011, Luojun Lin, Di Xie, Shiliang Pu, Xinchao Wang, Mingli Song, Yueting Zhuang
CVPR6
2022 Attention Diversification for Domain Generalization
Rang Meng, Weijie Chen 0006, Shicai Yang, Jie Song 0011, Xinchao Wang, Lei Zhang 0038, Mingli Song, Di Xie, Shiliang Pu
ECCV (34)9
2022 FBNet: Feedback Network for Point Cloud Completion
Hongyu Yan, Jingjing Wang 0005, Di Xie, Shiliang Pu
ECCV (2)6
2022 Transductive Clip with Class-Conditional Contrastive Learning
abstract
Inspired by the remarkable zero-shot generalization capacity of vision-language pre-trained model, we seek to leverage the supervision from CLIP model to alleviate the burden of data labeling. However, such supervision inevitably contains the label noise, which significantly degrades the discriminative power of the classification model. In this work, we propose Transductive CLIP, a novel framework for learning a classification network with noisy labels from scratch. Firstly, a class-conditional contrastive learning mechanism is proposed to mitigate the reliance on pseudo labels and boost the tolerance to noisy labels. Secondly, ensemble labels is adopted as a pseudo label updating strategy to stabilize the training of deep neural networks with noisy labels. This framework can reduce the impact of noisy labels from CLIP model effectively by combining both techniques. Experiments on multiple benchmark datasets demonstrate the substantial improvements over other state-of-the-art methods.
Junchu Huang, Weijie Chen 0006, Shicai Yang, Di Xie, Shiliang Pu, Yueting Zhuang
ICASSP4
2022 Target-Aware Auto-Augmentation for Unsupervised Domain Adaptive Object Detection
abstract
Recent researches show that data auto-augmentation strategies can enhance the performance of object detection models. However, the existing works mainly focus on in-domain generalization. There is still a blank in out-of-domain generalization. In this paper, for the first time, we propose an auto-augmentation problem under unsupervised domain adaptation scenarios. To solve this problem, we propose a simple yet effective target-aware auto-augmentation technique to search for an optimal data augmentation strategy on labeled source data, so as to boost the detection ability on the given unlabeled target data. Our method can be easily plugged into the existing domain adaptation methods. Extensive experiments have been carried out to verify the effectiveness.
Weijie Chen 0006, Shicai Yang, Di Xie, Shiliang Pu
ICASSP5
2022 Simulation-and-Mining: Towards Accurate Source-Free Unsupervised Domain Adaptive Object Detection
abstract
Vanilla unsupervised domain adaptive (UDA) object detection typically requires the labeled source data for joint-training with the unlabeled target data, which is usually unavailable in real-world scenarios due to data privacy, leading to source data-free UDA object detection. Herein, we first analyze the phenomenon of cross-domain detection degradation varying from easy to hard samples (e.g. the objects with different scales or occlusion degrees), termed as domain generalization differentiation. In detail, the ability to detect easy samples is well transferred while the one to detect hard samples is dramatically degraded. To this end, we then revisit the existing self-training method, which is of great challenge to deal with the abundant false negatives (hard samples). Assumed that true positives (easy samples) labeled by the source model can be exploited as supervision cues. UDA is finally modeled into an unsupervised false negatives mining problem. Thus, we propose a Simulation-and-Mining (S&M) framework, which simulates false negatives by augmenting true positives and mines back false negatives alternatively and iteratively. Experimental results show the effectiveness.
Weijie Chen 0006, Shicai Yang, Yunyi Xuan, Di Xie, Yueting Zhuang, Shiliang Pu
ICASSP5
2022 Semi-Supervised Ranking for Object Image Blur Assessment
abstract
Assessing the blurriness of an object image is fundamentally important to improve the performance for object recognition and retrieval. The main challenge lies in the lack of abundant images with reliable labels and effective learning strategies. Current datasets are labeled with limited and confused quality levels. To overcome this limitation, we propose to label the rank relationships between pairwise images rather their quality levels, since it is much easier for humans to label, and establish a large-scale realistic face image blur assessment dataset with reliable labels. Based on this dataset, we propose a method to obtain the blur scores only with the pairwise rank labels as supervision. Moreover, to further improve the performance, we propose a self-supervised method based on quadruplet ranking consistency to leverage the unlabeled data more effectively. The supervised and self-supervised methods constitute a final semi-supervised learning framework, which can be trained end-to-end. Experimental results demonstrate the effectiveness of our method. Source of labeled datasets: https://github.com/yzliangHIK2022/SSRanking-for-Object-BA
Qiang Li 0044, Zhaoliang Yao, Jingjing Wang 0005, Pengju Yang 0001, Di Xie, Shiliang Pu
ICIP6
2022 Effcient Shift Network in Denoising-Friendly Space for Real Noise Removal
abstract
Recently, following the success of neural networks, image denoising has achieved great improvements. However, it is challenging to construct an efficient denoising model with excellent performance and less computation. In this paper, we propose an extremely lightweight framework to remove real image noise in denoising-friendly space. Specif-ically, we apply the wavelet transform to project the noisy image and feature maps into low and high frequency domain, which decouples the noise information from the clean ones to a certain extent, thereby reducing the difficulty of denoising task. In addition, we further introduce a lightweight op-erator called Grouped Shift Module (GSM) into our denoising network, hence much heavy computation can be saved. Experimental results on the current benchmark demonstrate that our Wavelet Shift Denoising Network (WSNet) even achieves PSNR 39.28 dB with only 3G FLOPs on the SIDD benchmark. Our source code and models are available at https://github.com/HIK-DLSlimIWSNet.
Shuaitao Zhang, Di Xie, Shiliang Pu
ICME5
2022 Learning Domain Adaptive Object Detection with Probabilistic Teacher
abstract
Self-training for unsupervised domain adaptive object detection is a challenging task, of which the performance depends heavily on the quality of pseudo boxes. Despite the promising results, prior works have largely overlooked the uncertainty of pseudo boxes during self-training. In this paper, we present a simple yet effective framework, termed as Probabilistic Teacher (PT), which aims to capture the uncertainty of unlabeled target data from a gradually evolving teacher and guides the learning of a student in a mutually beneficial manner. Specifically, we propose to leverage the uncertainty-guided consistency training to promote classification adaptation and localization adaptation, rather than filtering pseudo boxes via an elaborate confidence threshold. In addition, we conduct anchor adaptation in parallel with localization adaptation, since anchor can be regarded as a learnable parameter. Together with this framework, we also present a novel Entropy Focal Loss (EFL) to further facilitate the uncertainty-guided self-training. Equipped with EFL, PT outperforms all previous baselines by a large margin and achieve new state-of-the-arts.
Meilin Chen, Weijie Chen 0006, Shicai Yang, Jie Song 0011, Xinchao Wang, Lei Zhang 0038, Yunfeng Yan, Donglian Qi, Yueting Zhuang, Di Xie, Shiliang Pu
ICML10
2022 High-Accuracy and Energy-Efficient Action Recognition with Deep Spiking Neural Network
Jingren Zhang, Jingjing Wang 0005, Di Xie, Shiliang Pu
ICONIP (2)3
2022 KRNet: Towards Efficient Knowledge Replay
abstract
The knowledge replay technique has been widely used in many tasks such as continual learning and continuous domain adaptation. The key lies in how to effectively encode the knowledge extracted from previous data and replay them during current training procedure. A simple yet effective model to achieve knowledge replay is autoencoder. However, the number of stored latent codes in autoencoder increases linearly with the scale of data and the trained encoder is redundant for the replaying stage. In this paper, we propose a novel and efficient knowledge recording network (KRNet) which directly maps an arbitrary sample identity number to the corresponding datum. Compared with autoencoder, our KRNet requires significantly (400×) less storage cost for the latent codes and can be trained without the encoder sub-network. Extensive experiments validate the efficiency of KRNet, and as a showcase, it is successfully applied in the task of continual learning.
Qiaoyong Zhong, Di Xie, Shiliang Pu
ICPR3
2022 Self-distilled Knowledge Delegator for Exemplar-free Class Incremental Learning
abstract
Exemplar-free incremental learning is extremely challenging due to inaccessibility of data from old tasks. In this paper, we attempt to exploit the knowledge encoded in a previously trained classification model to handle the catas-trophic forgetting problem in continual learning. Specifically, we introduce a so-called knowledge delegator, which is capable of transferring knowledge from the trained model to a randomly re-initialized new model by generating informative samples. Given the previous model only, the delegator is effectively learned using a self-distillation mechanism in a data-free manner. The knowledge extracted by the delegator is then utilized to maintain the performance of the model on old tasks in incremental learning. This simple incremental learning framework surpasses existing exemplar-free methods by a large margin on four widely used class incremental benchmarks, namely CIFAR-100, ImageNet-Subset, Caltech-101 and Flowers-102. Notably, we achieve comparable performance to some exemplar-based methods without accessing any exemplars.
Fanfan Ye, Qiaoyong Zhong, Di Xie, Shiliang Pu
IJCNN4
2022 Self-Supervised Noisy Label Learning for Source-Free Unsupervised Domain Adaptation
abstract
Domain adaptation is an important property in robot vision, which enables the neural networks pre-trained on source domains to adapt target domains automatically without any annotation efforts. During this process, source data is not always accessible due to the constraints of expensive storage overhead and data privacy protection. Therefore, the source domain pre-trained model is expected to optimize with only unlabeled target data, termed as source-free unsupervised domain adaptation. In this paper, we view this problem as a special case of noisy label learning, since the given pre-trained model can generate noisy labels for unlabeled target data via network inference. The potential semantic cues for unsupervised domain adaptation exactly lie on these noisy labels. Inspired by this problem modeling, we propose a simple yet effective Self-Supervised Noisy Label Learning method, which injects self-supervised learning to impose the intrinsic data structure and facilitate label-denoising. Extensive experiments have been conducted on diverse benchmarks to validate the effectiveness. Our method achieves state-of-the-art performance.
Weijie Chen 0006, Luojun Lin, Shicai Yang, Di Xie, Shiliang Pu, Yueting Zhuang
IROS4
2022 "Lossless" Compression of Deep Neural Networks: A High-dimensional Neural Tangent Kernel Approach
abstract
Modern deep neural networks (DNNs) are extremely powerful; however, this comes at the price of increased depth and having more parameters per layer, making their training and inference more computationally challenging. In an attempt to address this key limitation, efforts have been devoted to the compression (e.g., sparsification and/or quantization) of these large-scale machine learning models, so that they can be deployed on low-power IoT devices.In this paper, building upon recent research advances in the neural tangent kernel (NTK) and random matrix theory, we provide a novel compression approach to wide and fully-connected \emph{deep} neural nets. Specifically, we demonstrate that in the high-dimensional regime where the number of data points $n$ and their dimension $p$ are both large, and under a Gaussian mixture model for the data, there exists \emph{asymptotic spectral equivalence} between the NTK matrices for a large family of DNN models. This theoretical result enables ''lossless'' compression of a given DNN to be performed, in the sense that the compressed network yields asymptotically the same NTK as the original (dense and unquantized) network, with its weights and activations taking values \emph{only} in $\{ 0, \pm 1 \}$ up to scaling. Experiments on both synthetic and real-world data are conducted to support the advantages of the proposed compression scheme, with code available at https://github.com/Model-Compression/Lossless_Compression.
Lingyu Gu, Yongqi Du, Di Xie, Shiliang Pu, Robert C. Qiu, Zhenyu Liao 0001
NeurIPS4
2021 A Free Lunch for Unsupervised Domain Adaptive Object Detection without Source Data
abstract
Unsupervised domain adaptation (UDA) assumes that source and target domain data are freely available and usually trained together to reduce the domain gap. However, considering the data privacy and the inefficiency of data transmission, it is impractical in real scenarios. Hence, it draws our eyes to optimize the network in the target domain without accessing labeled source data. To explore this direction in object detection, for the first time, we propose a source data-free domain adaptive object detection (SFOD) framework via modeling it into a problem of learning with noisy labels. Generally, a straightforward method is to leverage the pre-trained network from the source domain to generate the pseudo labels for target domain optimization. However, it is difficult to evaluate the quality of pseudo labels since no labels are available in target domain. In this paper, self-entropy descent (SED) is a metric proposed to search an appropriate confidence threshold for reliable pseudo label generation without using any handcrafted labels. Nonetheless, completely clean labels are still unattainable. After a thorough experimental analysis, false negatives are found to dominate in the generated noisy labels. Undoubtedly, false negatives mining is helpful for performance improvement, and we ease it to false negatives simulation through data augmentation like Mosaic. Extensive experiments conducted in four representative adaptation tasks have demonstrated that the proposed framework can easily achieve state-of-the-art performance. From another view, it also reminds the UDA community that the labeled source data are not fully exploited in the existing methods.
Weijie Chen 0006, Di Xie, Shicai Yang, Shiliang Pu, Yueting Zhuang
AAAI3
2021 TransForensics: Image Forgery Localization with Dense Self-Attention
abstract
Nowadays advanced image editing tools and technical skills produce tampered images more realistically, which can easily evade image forensic systems and make authenticity verification of images more difficult. To tackle this challenging problem, we introduce TransForensics, a novel image forgery localization method inspired by Transformers. The two major components in our framework are dense self-attention encoders and dense correction modules. The former is to model global context and all pairwise inter-actions between local patches at different scales, while the latter is used for improving the transparency of the hidden layers and correcting the outputs from different branches. Compared to previous traditional and deep learning methods, TransForensics not only can capture discriminative representations and obtain high-quality mask predictions but is also not limited by tampering types and patch sequence orders. By conducting experiments on main bench-marks, we show that TransForensics outperforms the state-of-the-art methods by a large margin.
Shicai Yang, Di Xie, Shiliang Pu
ICCV4
2021 Divide-and-Assemble: Learning Block-wise Memory for Unsupervised Anomaly Detection
abstract
Reconstruction-based methods play an important role in unsupervised anomaly detection in images. Ideally, we expect a perfect reconstruction for normal samples and poor reconstruction for abnormal samples. Since the generalizability of deep neural networks is difficult to control, existing models such as autoencoder do not work well. In this work, we interpret the reconstruction of an image as a divide-and-assemble procedure. Surprisingly, by varying the granularity of division on feature maps, we are able to modulate the reconstruction capability of the model for both normal and abnormal samples. That is, finer granularity leads to better reconstruction, while coarser granularity leads to poorer reconstruction. With proper granularity, the gap between the reconstruction error of normal and abnormal samples can be maximized. The divide-and-assemble framework is implemented by embedding a novel multi-scale block-wise memory module into an autoencoder network. Besides, we introduce adversarial learning and explore the semantic latent representation of the discriminator, which improves the detection of subtle anomaly. We achieve state-of-the-art performance on the challenging MVTec AD dataset. Remarkably, we improve the vanilla autoencoder model by 10.1% in terms of the AUROC score.
Jinlei Hou, Qiaoyong Zhong, Di Xie, Shiliang Pu
ICCV4
2021 Modulating Localization and Classification for Harmonized Object Detection
abstract
Object detection involves two sub-tasks, i.e. localizing objects in an image and classifying them into various categories. For existing CNN-based detectors, we notice the widespread divergence between localization and classification, which leads to degradation in performance. In this work, we propose a mutual learning framework to modulate the two tasks. In particular, the two tasks are forced to learn from each other with a novel mutual labeling strategy. Besides, we introduce a simple yet effective IoU rescoring scheme, which further reduces the divergence. Moreover, we define a Spearman rank correlation-based metric to quantify the divergence, which correlates well with the detection performance. The proposed approach is general-purpose and can be easily injected into existing detectors such as FCOS and RetinaNet. We achieve a significant performance gain over the baseline detectors on the COCO dataset.
Taiheng Zhang, Qiaoyong Zhong, Shiliang Pu, Di Xie
ICME4
2021 Look Before You Act: Boosting Pseudo-LiDAR with Online Semantic Embedding
abstract
Vision-based 3D object detection is a research focus in the field of autonomous driving system. While recently proposed pseudo-LiDAR is a promising solution, its performance is severely restricted by the image-based depth estimator, leading to a considerable performance gap against the LiDAR-based counterparts. In this paper, substantial advances are developed along an orthogonal direction to the previous efforts in the pseudo-LiDAR pipeline. Concretely, we propose a plug- and-play module, called Online Semantic Embedding (OSE), aligning image semantics with the pseudo-LiDAR detection in an end-to-end manner. On the KITTI object detection benchmark, existing stereo-based baselines integrated with our approach show impressive improvements without bells and whistles. Furthermore, we emphasize that OSE works in retrieving the performance under geometric imperfection conditions.
Liangjun Zhang, Di Xie, Shiliang Pu
IROS4
2021 MGD-GAN: Text-to-Pedestrian Generation Through Multi-grained Discrimination
Shengyu Zhang 0001, Zhou Zhao 0001, Siliang Tang, Kun Kuang 0001, Di Xie, Fei Wu 0001
PRCV (2)6
2021 Associative affinity network learning for multi-object tracking
abstract
We propose a joint feature and metric learning deep neural network architecture, called the associative affinity network (AAN), as an affinity model for multi-object tracking (MOT) in videos. The AAN learns the associative affinity between tracks and detections across frames in an end-to-end manner. Considering flawed detections, the AAN jointly learns bounding box regression, classification, and affinity regression via the proposed multi-task loss. Contrary to networks that are trained with ranking loss, we directly train a binary classifier to learn the associative affinity of each track-detection pair and use a matching cardinality loss to capture information among candidate pairs. The AAN learns a discriminative affinity model for data association to tackle MOT, and can also perform single-object tracking. Based on the AAN, we propose a simple multi-object tracker that achieves competitive performance on the public MOT16 and MOT17 test datasets.
Ma Liang, Qiaoyong Zhong, Di Xie, Shiliang Pu
Frontiers Inf. Technol. Electron. Eng.4
2021 Unsupervised object detection with scene-adaptive concept learning
abstract
Object detection is one of the hottest research directions in computer vision, has already made impressive progress in academia, and has many valuable applications in the industry. However, the mainstream detection methods still have two shortcomings: (1) even a model that is well trained using large amounts of data still cannot generally be used across different kinds of scenes; (2) once a model is deployed, it cannot autonomously evolve along with the accumulated unlabeled scene data. To address these problems, and inspired by visual knowledge theory, we propose a novel scene-adaptive evolution unsupervised video object detection algorithm that can decrease the impact of scene changes through the concept of object groups. We first extract a large number of object proposals from unlabeled data through a pre-trained detection model. Second, we build the visual knowledge dictionary of object concepts by clustering the proposals, in which each cluster center represents an object prototype. Third, we look into the relations between different clusters and the object information of different groups, and propose a graph-based group information propagation strategy to determine the category of an object concept, which can effectively distinguish positive and negative proposals. With these pseudo labels, we can easily fine-tune the pre-trained model. The effectiveness of the proposed method is verified by performing different experiments, and the significant improvements are achieved.
Shiliang Pu, Weijie Chen 0006, Shicai Yang, Di Xie, Yunhe Pan
Frontiers Inf. Technol. Electron. Eng.5
2021 Auxiliary diagnostic system for ADHD in children based on AI technology
abstract
Traditional diagnosis of attention deficit hyperactivity disorder (ADHD) in children is primarily through a questionnaire filled out by parents/teachers and clinical observations by doctors. It is inefficient and heavily depends on the doctor’s level of experience. In this paper, we integrate artificial intelligence (AI) technology into a software-hardware coordinated system to make ADHD diagnosis more efficient. Together with the intelligent analysis module, the camera group will collect the eye focus, facial expression, 3D body posture, and other children’s information during the completion of the functional test. Then, a multi-modal deep learning model is proposed to classify abnormal behavior fragments of children from the captured videos. In combination with other system modules, standardized diagnostic reports can be automatically generated, including test results, abnormal behavior analysis, diagnostic aid conclusions, and treatment recommendations. This system has participated in clinical diagnosis in Department of Psychology, The Children’s Hospital, Zhejiang University School of Medicine, and has been accepted and praised by doctors and patients.
Yanyi Zhang, Ming Kong 0001, Wenchen Hong, Di Xie, Chunmao Wang, Rongwang Yang
Frontiers Inf. Technol. Electron. Eng.5
2020 Neural Inheritance Relation Guided One-Shot Layer Assignment Search
abstract
Layer assignment is seldom picked out as an independent research topic in neural architecture search. In this paper, for the first time, we systematically investigate the impact of different layer assignments to the network performance by building an architecture dataset of layer assignment on CIFAR-100. Through analyzing this dataset, we discover a neural inheritance relation among the networks with different layer assignments, that is, the optimal layer assignments for deeper networks always inherit from those for shallow networks. Inspired by this neural inheritance relation, we propose an efficient one-shot layer assignment search approach via inherited sampling. Specifically, the optimal layer assignment searched in the shallow network can be provided as a strong sampling priori to train and search the deeper ones in supernet, which extremely reduces the network search space. Comprehensive experiments carried out on CIFAR-100 illustrate the efficiency of our proposed method. Our search results are strongly consistent with the optimal ones directly selected from the architecture dataset. To further confirm the generalization of our proposed method, we also conduct experiments on Tiny-ImageNet and ImageNet. Our searched results are remarkably superior to the handcrafted ones under the unchanged computational budgets. The neural inheritance relation discovered in this paper can provide insights to the universal neural architecture search.
Rang Meng, Weijie Chen 0006, Di Xie, Shiliang Pu
AAAI3
2020 Dynamic GCN: Context-enriched Topology Learning for Skeleton-based Action Recognition
abstract
raph Convolutional Networks (GCNs) have attracted increasing interests for the task of skeleton-based action recognition. The key lies in the design of the graph structure, which encodes skeleton topology information. In this paper, we propose Dynamic GCN, in which a novel convolutional neural network named Context-encoding Network (CeN) is introduced to learn skeleton topology automatically. In particular, when learning the dependency between two joints, contextual features from the rest joints are incorporated in a global manner. CeN is extremely lightweight yet effective, and can be embedded into a graph convolutional layer. By stacking multiple CeN-enabled graph convolutional layers, we build Dynamic GCN. Notably, as a merit of CeN, dynamic graph topologies are constructed for different input samples as well as graph convolutional layers of various depths. Besides, three alternative context modeling architectures are well explored, which may serve as a guideline for future research on graph topology learning. CeN brings only ~7% extra FLOPs for the baseline model, and Dynamic GCN achieves better performance with 2x ~4x fewer FLOPs than existing methods. By further combining static physical body connections and motion modalities, we achieve state-of-the-art performance on three large-scale benchmarks, namely NTU-RGB+D, NTU-RGB+D 120 and Skeleton-Kinetics.
Fanfan Ye, Shiliang Pu, Qiaoyong Zhong, Chao Li 0064, Di Xie, Huiming Tang
ACM Multimedia5
2020 Cascade region proposal and global context for deep object detection
Qiaoyong Zhong, Chao Li 0064, Di Xie, Shicai Yang, Shiliang Pu
Neurocomputing4
2020 Deep Reinforcement Learning for Smart Home Energy Management
abstract
We investigate an energy cost minimization problem for a smart home in the absence of a building thermal dynamics model with the consideration of a comfortable temperature range. Due to the existence of model uncertainty, parameter uncertainty (e.g., renewable generation output, nonshiftable power demand, outdoor temperature, and electricity price), and temporally coupled operational constraints, it is very challenging to design an optimal energy management algorithm for scheduling heating, ventilation, and air conditioning systems and energy storage systems in the smart home. To address the challenge, we first formulate the above problem as a Markov decision process, and then propose an energy management algorithm based on deep deterministic policy gradients. It is worth mentioning that the proposed algorithm does not require the prior knowledge of uncertain parameters and building the thermal dynamics model. The simulation results based on real-world traces demonstrate the effectiveness and robustness of the proposed algorithm.
Liang Yu 0001, Weiwei Xie, Di Xie, YuLong Zou, Dengyin Zhang, Zhixin Sun, Linghua Zhang, Yue Zhang 0011, Tao Jiang 0002
IEEE Internet Things J.3
2020 Advancing Image Understanding in Poor Visibility Environments: A Collective Benchmark Study
abstract
Existing enhancement methods are empirically expected to help the high-level end computer vision task: however, that is observed to not always be the case in practice. We focus on object or face detection in poor visibility enhancements caused by bad weathers (haze, rain) and low light conditions. To provide a more thorough examination and fair comparison, we introduce three benchmark sets collected in real-world hazy, rainy, and low-light conditions, respectively, with annotated objects/faces. We launched the UG2+ challenge Track 2 competition in IEEE CVPR 2019, aiming to evoke a comprehensive discussion and exploration about whether and how low-level vision techniques can benefit the high-level automatic visual recognition in various scenarios. To our best knowledge, this is the first and currently largest effort of its kind. Baseline results by cascading existing enhancement and detection models are reported, indicating the highly challenging nature of our new data as well as the large room for further technical innovations. Thanks to a large participation from the research community, we are able to analyze representative team solutions, striving to better identify the strengths and limitations of existing mindsets as well as the future directions.
Wenhan Yang, Ye Yuan 0012, Wenqi Ren, Jiaying Liu 0001, Walter J. Scheirer, Zhangyang Wang, Taiheng Zhang, Qiaoyong Zhong, Di Xie, Shiliang Pu, Yuqiang Zheng, Yanyun Qu, Yuhong Xie, Hao Jiang 0014, Siyuan Yang 0001, Yan Liu 0041, Xiaochao Qu, Pengfei Wan 0001, Shuai Zheng 0005, Minhui Zhong, Taiyi Su, Lingzhi He, Yandong Guo, Yao Zhao 0001, Zhenfeng Zhu, Jinxiu Liang, Jingwen Wang 0003, Yuhui Quan, Yong Xu 0007, Bo Liu 0112, Xin Liu 0012, Tingyu Lin 0003, Xiaochuan Li 0001, Feng Lu 0005, Lin Gu 0003, Shengdi Zhou, Cong Cao 0005, Cheng Chi 0003, Chubin Zhuang, Zhen Lei 0001, Stan Z. Li, Shizheng Wang, Ruizhe Liu, Dong Yi, Zheming Zuo, Jianning Chi, Huan Wang 0014, Kai Wang 0036, Yixiu Liu, Xingyu Gao 0001, Zhenyu Chen 0003, Yongzhou Li, Huicai Zhong, Jing Huang 0017, Heng Guo 0003, Jianfei Yang 0001, Wenjuan Liao, Jiangang Yang, Liguo Zhou, Mingyue Feng, Likun Qin
IEEE Trans. Image Process.9
2020 An Attentive Sequence to Sequence Translator for Localizing Video Clips by Natural Language
abstract
We propose a novel attentive sequence to sequence translator (ASST) for localizing video clips by natural language descriptions. We make two contributions. First, we propose an attentive mechanism that aligns natural language descriptions and video content. A bi-directional Recurrent Neural Network (RNN) parses natural language descriptions in two directions. Given a video-description pair, ASST generates a vector sequence representation. Each vector represents a video frame, conditioned by the description. The vector sequence representation not only preserves the temporal dependencies between the frames, but also provides an effective way to perform frame-level videolanguage matching. The attentive model then aligns words to each frame, thereby resulting in a more detailed understanding of video content and description semantics. Second, we design a hierarchical architecture for the network to jointly model language descriptions and video content. The hierarchical architecture exploits video content with multiple granularities, ranging from subtle details to global context. The integration of the multiple granularities yields a robust representation for multi-level videolanguage abstraction. We validate the effectiveness of our ASST on two large-scale datasets. Our ASST outperforms the state-ofthe-art by 4.28% in Rank@1 on the DiDeMo dataset. On the Charades-STA dataset, we significantly improve the state-of-theart by 13.41% in Recall@1,IoU = 0.5.
Di Xie, Fei Wu 0001
IEEE Trans. Multim.3
2019 A Layer Decomposition-Recomposition Framework for Neuron Pruning towards Accurate Lightweight Networks
abstract
Neuron pruning is an efficient method to compress the network into a slimmer one for reducing the computational cost and storage overhead. Most of state-of-the-art results are obtained in a layer-by-layer optimization mode. It discards the unimportant input neurons and uses the survived ones to reconstruct the output neurons approaching to the original ones in a layer-by-layer manner. However, an unnoticed problem arises that the information loss is accumulated as layer increases since the survived neurons still do not encode the entire information as before. A better alternative is to propagate the entire useful information to reconstruct the pruned layer instead of directly discarding the less important neurons. To this end, we propose a novel Layer DecompositionRecomposition Framework (LDRF) for neuron pruning, by which each layer’s output information is recovered in an embedding space and then propagated to reconstruct the following pruned layers with useful information preserved. We mainly conduct our experiments on ILSVRC-12 benchmark with VGG-16 and ResNet-50. What should be emphasized is that our results before end-to-end fine-tuning are significantly superior owing to the information-preserving property of our proposed framework. With end-to-end fine-tuning, we achieve state-of-the-art results of 5.13× and 3× speed-up with only 0.5% and 0.65% top-5 accuracy drop respectively, which outperform the existing neuron pruning methods.
Weijie Chen 0006, Di Xie, Shiliang Pu
AAAI3
2019 Learning Incremental Triplet Margin for Person Re-Identification
abstract
Person re-identification (ReID) aims to match people across multiple non-overlapping video cameras deployed at different locations. To address this challenging problem, many metric learning approaches have been proposed, among which triplet loss is one of the state-of-the-arts. In this work, we explore the margin between positive and negative pairs of triplets and prove that large margin is beneficial. In particular, we propose a novel multi-stage training strategy which learns incremental triplet margin and improves triplet loss effectively. Multiple levels of feature maps are exploited to make the learned features more discriminative. Besides, we introduce global hard identity searching method to sample hard identities when generating a training batch. Extensive experiments on Market-1501, CUHK03, and DukeMTMCreID show that our approach yields a performance boost and outperforms most existing state-of-the-art methods.
Qiaoyong Zhong, Di Xie, Shiliang Pu
AAAI4
2019 All You Need Is a Few Shifts: Designing Efficient Convolutional Neural Networks for Image Classification
abstract
Shift operation is an efficient alternative over depthwise separable convolution. However, it is still bottlenecked by its implementation manner, namely memory movement. To put this direction forward, a new and novel basic component named Sparse Shift Layer (SSL) is introduced in this paper to construct efficient convolutional neural networks. In this family of architectures, the basic block is only composed by 1x1 convolutional layers with only a few shift operations applied to the intermediate feature maps. To make this idea feasible, we introduce shift operation penalty during optimization and further propose a quantization-aware shift learning method to impose the learned displacement more friendly for inference. Extensive ablation studies indicate that only a few shift operations are sufficient to provide spatial information communication. Furthermore, to maximize the role of SSL, we redesign an improved network architecture to Fully Exploit the limited capacity of neural Network (FE-Net). Equipped with SSL, this network can achieve 75.0% top-1 accuracy on ImageNet with only 563M M-Adds. It surpasses other counterparts constructed by depthwise separable convolution and the networks searched by NAS in terms of accuracy and practical speed.
Weijie Chen 0006, Di Xie, Shiliang Pu
CVPR2
2019 Collaborative Spatiotemporal Feature Learning for Video Action Recognition
abstract
Spatiotemporal feature learning is of central importance for action recognition in videos. Existing deep neural network models either learn spatial and temporal features independently (C2D) or jointly with unconstrained parameters (C3D). In this paper, we propose a novel neural operation which encodes spatiotemporal features collaboratively by imposing a weight-sharing constraint on the learnable parameters. In particular, we perform 2D convolution along three orthogonal views of volumetric video data, which learns spatial appearance and temporal motion cues respectively. By sharing the convolution kernels of different views, spatial and temporal features are collaboratively learned and thus benefit from each other. The complementary features are subsequently fused by a weighted summation whose coefficients are learned end-to-end. Our approach achieves state-of-the-art performance on large-scale benchmarks and won the 1st place in the Moments in Time Challenge 2018. Moreover, based on the learned coefficients of different views, we are able to quantify the contributions of spatial and temporal features. This analysis sheds light on interpretability of the model and may also guide the future design of algorithm for video recognition.
Chao Li 0064, Qiaoyong Zhong, Di Xie, Shiliang Pu
CVPR3
2019 Self-Supervised Spatiotemporal Learning via Video Clip Order Prediction
abstract
We propose a self-supervised spatiotemporal learning technique which leverages the chronological order of videos. Our method can learn the spatiotemporal representation of the video by predicting the order of shuffled clips from the video. The category of the video is not required, which gives our technique the potential to take advantage of infinite unannotated videos. There exist related works which use frames, while compared to frames, clips are more consistent with the video dynamics. Clips can help to reduce the uncertainty of orders and are more appropriate to learn a video representation. The 3D convolutional neural networks are utilized to extract features for clips, and these features are processed to predict the actual order. The learned representations are evaluated via nearest neighbor retrieval experiments. We also use the learned networks as the pre-trained models and finetune them on the action recognition task. Three types of 3D convolutional neural networks are tested in experiments, and we gain large improvements compared to existing self-supervised methods.
Dejing Xu, Jun Xiao 0001, Zhou Zhao 0001, Jian Shao 0001, Di Xie, Yueting Zhuang
CVPR5
2019 An End-to-End Audio Classification System Based on Raw Waveforms and Mix-Training Strategy
abstract
Audio classification can distinguish different kinds of sounds, which is helpful for intelligent applications in daily life.However, it remains a challenging task since the sound events in an audio clip is probably multiple, even overlapping.This paper introduces an end-to-end audio classification system based on raw waveforms and mix-training strategy.Compared to human-designed features which have been widely used in existing research, raw waveforms contain more complete information and are more appropriate for multi-label classification.Taking raw waveforms as input, our network consists of two variants of ResNet structure which can learn a discriminative representation.To explore the information in intermediate layers, a multi-level prediction with attention structure is applied in our model.Furthermore, we design a mix-training strategy to break the performance limitation caused by the amount of training data.Experiments show that the mean average precision of the proposed audio classification system on Audio Set dataset is 37.2%.Without using extra training data, our system exceeds the state-of-the-art multi-level attention model.
Di Xie, Shicai Yang, Shiliang Pu
INTERSPEECH4
2019 MicroRNAs and complex diseases: from experimental results to computational models
abstract
Circular RNAs (circRNAs) are a class of single-stranded, covalently closed RNA molecules with a variety of biological functions. Studies have shown that circRNAs are involved in a variety of biological processes and play an important role in the development of various complex diseases, so the identification of circRNA-disease associations would contribute to the diagnosis and treatment of diseases. In this review, we summarize the discovery, classifications and functions of circRNAs and introduce four important diseases associated with circRNAs. Then, we list some significant and publicly accessible databases containing comprehensive annotation resources of circRNAs and experimentally validated circRNA-disease associations. Next, we introduce some state-of-the-art computational models for predicting novel circRNA-disease associations and divide them into two categories, namely network algorithm-based and machine learning-based models. Subsequently, several evaluation methods of prediction performance of these computational models are summarized. Finally, we analyze the advantages and disadvantages of different types of computational models and provide some suggestions to promote the development of circRNA-disease association identification from the perspective of the construction of new computational models and the accumulation of circRNA-related data.
Xing Chen 0001, Di Xie, Qi Zhao 0010, Zhu-Hong You
Briefings Bioinform.2
2019 Adversarial learning for viewpoints invariant 3D human pose estimation
Jun Xiao 0001, Di Xie, Jian Shao 0001
J. Vis. Commun. Image Represent.3
2018 Extreme Network Compression via Filter Group Approximation
Wenming Tan, Zheyang Li, Di Xie, Shiliang Pu
ECCV (8)5
2018 Small-Scale Pedestrian Detection Based on Topological Line Localization and Temporal Feature Aggregation
Leiyu Sun, Di Xie, Haiming Sun, Shiliang Pu
ECCV (7)3
2018 A Practical Convolutional Neural Network as Loop Filter for Intra Frame
abstract
Loop filters are used in video coding to remove artifacts or improve performance. Recent advances in deploying convolutional neural network (CNN) to replace traditional loop filters show large gains but with problems for practical application. First, different model is used for frames encoded with different quantization parameter (QP), respectively. It is expensive for hardware. Second, float points operation in CNN leads to inconsistency between encoding and decoding across different platforms. Third, redundancy within CNN model consumes precious computational resources. This paper proposes a CNN as the loop filter for intra frames and proposes a scheme to solve the above problems. It aims to design a single CNN model with low redundancy to adapt to decoded frames with different qualities and ensure consistency. To adapt to reconstructions with different qualities, both reconstruction and QP are taken as inputs. After training, the obtained model is compressed to reduce redundancy. To ensure consistency, dynamic fixed points (DFP) are adopted in testing CNN. Parameters in the compressed model are first quantized to DFP and then used for inference of CNN. Outputs of each layer in CNN are computed by DFP operations. Experimental results on JEM 7.0 report 3.14%,5.21 %, 6.28% BD-rate savings for luma and two chroma components with all intra configuration when replacing all traditional filters.
Xiaodan Song, Jiabao Yao, Lulu Zhou, Xiaoyang Wu 0007, Di Xie, Shiliang Pu
ICIP6
2018 Co-occurrence Feature Learning from Skeleton Data for Action Recognition and Detection with Hierarchical Aggregation
abstract
Skeleton-based human action recognition has recently drawn increasing attentions with the availability of large-scale skeleton datasets. The most crucial factors for this task lie in two aspects: the intra-frame representation for joint co-occurrences and the inter-frame representation for skeletons' temporal evolutions. In this paper we propose an end-to-end convolutional co-occurrence feature learning framework. The co-occurrence features are learned with a hierarchical methodology, in which different levels of contextual information are aggregated gradually. Firstly point-level information of each joint is encoded independently. Then they are assembled into semantic representation in both spatial and temporal domains. Specifically, we introduce a global spatial aggregation scheme, which is able to learn superior joint co-occurrence features over local aggregation. Besides, raw skeleton coordinates as well as their temporal difference are integrated with a two-stream paradigm. Experiments show that our approach consistently outperforms other state-of-the-arts on action recognition and detection benchmarks like NTU RGB+D, SBU Kinect Interaction and PKU-MMD.
Chao Li 0064, Qiaoyong Zhong, Di Xie, Shiliang Pu
IJCAI3
2018 BNPMDA: Bipartite Network Projection for MiRNA-Disease Association prediction
abstract
Motivation: A large number of resources have been devoted to exploring the associations between microRNAs (miRNAs) and diseases in the recent years. However, the experimental methods are expensive and time-consuming. Therefore, the computational methods to predict potential miRNA-disease associations have been paid increasing attention. Results: In this paper, we proposed a novel computational model of Bipartite Network Projection for MiRNA-Disease Association prediction (BNPMDA) based on the known miRNA-disease associations, integrated miRNA similarity and integrated disease similarity. We firstly described the preference degree of a miRNA for its related disease and the preference degree of a disease for its related miRNA with the bias ratings. We constructed bias ratings for miRNAs and diseases by using agglomerative hierarchical clustering according to the three types of networks. Then, we implemented the bipartite network recommendation algorithm to predict the potential miRNA-disease associations by assigning transfer weights to resource allocation links between miRNAs and diseases based on the bias ratings. BNPMDA had been shown to improve the prediction accuracy in comparison with previous models according to the area under the receiver operating characteristics (ROC) curve (AUC) results of three typical cross validations. As a result, the AUCs of Global LOOCV, Local LOOCV and 5-fold cross validation obtained by implementing BNPMDA were 0.9028, 0.8380 and 0.8980 ± 0.0013, respectively. We further implemented two types of case studies on several important human complex diseases to confirm the effectiveness of BNPMDA. In conclusion, BNPMDA could effectively predict the potential miRNA-disease associations at a high accuracy level. Availability and implementation: BNPMDA is available via http://www.escience.cn/system/file?fileId=99559. Supplementary information: Supplementary data are available at Bioinformatics online.
Xing Chen 0001, Di Xie, Lei Wang 0121, Qi Zhao 0010, Zhu-Hong You, Hongsheng Liu 0001
Bioinform.2
2018 Distributed Real-Time HVAC Control for Cost-Efficient Commercial Buildings Under Smart Grid Environment
abstract
In this paper, we investigate the problem of minimizing the long-term total cost (i.e., the sum of energy cost and thermal discomfort cost) associated with a heating, ventilation, and air conditioning (HVAC) system of a multizone commercial building under smart grid environment. To be specific, we first formulate a stochastic program to minimize the time average expected total cost with the consideration of uncertainties in electricity price, outdoor temperature, the most comfortable temperature level, and external thermal disturbance. Due to the existence of temporally and spatially coupled constraints as well as unknown information about the future system parameters, it is very challenging to solve the formulated problem. To this end, we propose a real-time HVAC control algorithm based on the framework of Lyapunov optimization techniques without the need to predict any system parameters and know their stochastic information. The key idea of the proposed algorithm is to construct and stabilize virtual queues associated with indoor temperatures of all zones. Moreover, we provide a distributed implementation of the proposed real-time algorithm with the aim of protecting user privacy and enhancing algorithmic scalability. Extensive simulation results based on real-world traces show that the proposed algorithm could reduce energy cost effectively with small sacrifice in thermal comfort.
Liang Yu 0001, Di Xie, Tao Jiang 0002, YuLong Zou, Kun Wang 0005
IEEE Internet Things J.2
2018 Fusing Geometric Features for Skeleton-Based Action Recognition Using Multilayer LSTM Networks
abstract
Recent skeleton-based action recognition approaches achieve great improvement by using recurrent neural network (RNN) models. Currently, these approaches build an end-to-end network from coordinates of joints to class categories and improve accuracy by extending RNN to spatial domains. First, while such well-designed models and optimization strategies explore relations between different parts directly from joint coordinates, we provide a simple universal spatial modeling method perpendicular to the RNN model enhancement. Specifically, according to the evolution of previous work, we select a set of simple geometric features, and then separately feed each type of features to a three-layer LSTM framework. Second, we propose a multistream LSTM architecture with a new smoothed score fusion technique to learn classification from different geometric feature streams. Furthermore, we observe that the geometric relational features based on distances between joints and selected lines outperform other features and the fusion results achieve the state-of-the-art performance on four datasets. We also show the sparsity of input gate weights in the first LSTM layer trained by geometric features and demonstrate that utilizing joint-line distances as input require less data for training.
Songyang Zhang 0004, Yang Yang 0009, Jun Xiao 0001, Xiaoming Liu 0002, Yi Yang 0001, Di Xie, Yueting Zhuang
IEEE Trans. Multim.6
2017 All You Need is Beyond a Good Init: Exploring Better Solution for Training Extremely Deep Convolutional Neural Networks with Orthonormality and Modulation
abstract
Deep neural network is difficult to train and this predicament becomes worse as the depth increases. The essence of this problem exists in the magnitude of backpropagated errors that will result in gradient vanishing or exploding phenomenon. We show that a variant of regularizer which utilizes orthonormality among different filter banks can alleviate this problem. Moreover, we design a backward error modulation mechanism based on the quasi-isometry assumption between two consecutive parametric layers. Equipped with these two ingredients, we propose several novel optimization solutions that can be utilized for training a specific-structured (repetitively triple modules of Conv-BNReLU) extremely deep convolutional neural network (CNN) WITHOUT any shortcuts/ identity mappings from scratch. Experiments show that our proposed solutions can achieve distinct improvements for a 44-layer and a 110-layer plain networks on both the CIFAR-10 and ImageNet datasets. Moreover, we can successfully train plain CNNs to match the performance of the residual counterparts. Besides, we propose new principles for designing network structure from the insights evoked by orthonormality. Combined with residual structure, we achieve comparative performance on the ImageNet dataset.
Di Xie, Shiliang Pu
CVPR1
2016 NDCMC: A Hybrid Data Collection Approach for Large-Scale WSNs Using Mobile Element and Hierarchical Clustering
abstract
To collect data from large-scale wireless sensor networks (WSNs) is a challenging issue and there are mainly two approaches to increase the efficiency: 1) by hierarchical routing based on node clustering and 2) by mobile elements (MEs). Since either method has pros and cons, this paper presents a hybrid approach, called node-density-based clustering and mobile collection (NDCMC), to combine the hierarchical routing and ME data collection in WSNs. A number of cluster heads (CHs) gather information from cluster members and then an ME visits these CHs to collect data. First, for a randomly deployed WSN, a new CH selection scheme based on the node density is proposed. The advantage is that the nodes which are surrounded by more deployed nodes are more likely to be CHs. Thus, the efficiency of both intracluster routing and ME data collection is improved. Second, a low-complexity traveling track planning algorithm is designed for an ME to pass by all CHs. The analytical model of NDCMC is also developed and the expectation of the sensor power consumption and network lifetime are derived. In addition, a simple random clustering and mobile collection (RCMC) scheme is introduced by which a number of CHs are selected randomly in a WSN. Although RCMC yields performance degradation, it has much less complexity. Extensive simulations show that the proposed hybrid NDCMC scheme leads to not only remarkable performance improvement but also convenient tradeoff between the network energy saving and the data collection latency.
Ruonan Zhang 0001, Jianping Pan 0001, Di Xie, Fubao Wang
IEEE Internet Things J.3
2015 A hybrid approach using mobile element and hierarchical clustering for data collection in WSNs
abstract
How to minimize the energy dissipation and extend the lifetime of wireless sensor networks (WSNs) is still an active research topic nowadays. Hierarchical routing based on node clustering is an effective method, while using mobile elements (MEs) to gather data can prevent huge energy consumption of the sensors from long-distance transmission. Considering that both methods have pros and cons, this paper presents a hybrid approach, called Node Density based Clustering and Mobile Collection (NDCM), to combine the hierarchical routing and ME data collection in WSNs. A number of Cluster Heads (CHs) first gather information from the cluster members and then the ME visits these CHs to collect data. A new CH selection scheme based on the node density is proposed. Thus, a node at the center of an area where nodes are densely deployed is more likely to be a CH, which can improve the efficiency of both intra-cluster routing and ME data collection. We also introduce a simple Random Clustering and Mobile Collection (RCM) scheme according to which a number of CHs are selected randomly throughout the network. In addition, the nodes which are covered by the radio range of the ME, called Virtual Heads (VHs), can also send/relay packets directly to the ME. The different mobility schemes are compared through extensive simulations and the results show that the proposed hybrid NDCM scheme leads to remarkable improvement in network lifetime and convenient trade off between the network energy saving and packet latency.
Ruonan Zhang 0001, Jianping Pan 0001, Jiajia Liu 0001, Di Xie
WCNC4
2013 PIKACHU: How to Rebalance Load in Optimizing MapReduce On Heterogeneous Clusters
Rohan Gandhi, Di Xie, Y. Charlie Hu
USENIX ATC2
2013 Upper Body Human Detection and Segmentation in Low Contrast Video
abstract
In the application of extracting human regions from videos, many existing methods may lose their efficacy when illumination varies or the human remains still. To address this problem, we propose a method in this paper for human region detection and segmentation by constructing a generalized human upper body model. The method mainly consists of two main procedures. First, foreground connected regions are extracted by background subtraction from the current frame and classified through a human upper body model pretrained with a support vector machine to determine whether they are human regions. Second, we assign an energy function to the region contour and apply an energy minimization procedure to evolve the contour when human regions are polluted by background; for example, a change in lighting conditions. After finding the optimal contour, we update the background and repeat the procedures in next frame. This feedback strategy rectifies the mistaken background regions promptly and extracts human regions correctly. Our experimental results demonstrate that the proposed method is robust enough to handle videos of low contrast as well as normal conditions.
Ruofeng Tong 0001, Di Xie, Min Tang 0001
IEEE Trans. Circuits Syst. Video Technol.2
2012 On the performance projectability of MapReduce
abstract
A key challenge faced by users of public clouds today is how to request for the right amount of resources in the production datacenter that satisfies a target performance for a given cloud application. An obvious approach is to develop a performance model for a class of applications such as MapReduce. However, several recent studies have shown that even for the class of well-studied MapReduce jobs, their running times can be seriously affected by numerous external factors ranging from dozen or so configuration parameters, to the physical machine characteristics (CPU, memory, disk, and network bandwidth), to implementation deficiencies such as Java, garbage collection. These factors make direct performance modeling extremely difficult. In this paper, we propose a more practical systematic methodology to solve this problem. Our approach develops a projection model, based on insights into performance bottlenecks of MapReduce jobs and their scaling properties, and parameterized with component running times based on profiling on small clusters with sampled inputs. Evaluation results show our projection model can predict job running times with 2.7% of accuracy when scaling to 32 nodes.
Di Xie, Y. Charlie Hu, Ramana Rao Kompella
CloudCom1
2012 The only constant is change: incorporating time-varying network reservations in data centers
abstract
In multi-tenant datacenters, jobs of different tenants compete for the shared datacenter network and can suffer poor performance and high cost from varying, unpredictable network performance. Recently, several virtual network abstractions have been proposed to provide explicit APIs for tenant jobs to specify and reserve virtual clusters (VC) with both explicit VMs and required network bandwidth between the VMs. However, all of the existing proposals reserve a fixed bandwidth throughout the entire execution of a job.
Di Xie, Ning Ding 0004, Y. Charlie Hu, Ramana Rao Kompella
SIGCOMM1