EDBT 2026 Demo / reviewers in the wild / expert
Zhenxing Niu
dblp:13/7704
· DBLP profile ↗
44ranked-venue papers
12as first author
20since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 32 · 9 first-author · 14 since 2021Artificial intelligence and machine learning · 26 · 8 first-author · 12 since 2021Computer networks · 2 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Efficient LLM-Jailbreaking via Multimodal-LLM JailbreakabstractThis paper focuses on jailbreaking attacks against large language models (LLMs), eliciting them to generate objectionable content in response to harmful user queries. Unlike previous LLM-jailbreak methods that directly orient to LLMs, our approach begins by constructing a multimodal large language model (MLLM) built upon the target LLM. Subsequently, we perform an efficient MLLM jailbreak and obtain a jailbreaking embedding. Finally, we convert the embedding into a textual jailbreaking suffix to carry out the jailbreak of target LLM. Compared to the direct LLM-jailbreak methods, our indirect jailbreaking approach is more efficient, as MLLMs are more vulnerable to jailbreak than pure LLMs. Additionally, to improve the attack success rate of jailbreak, we propose an image-text semantic matching scheme to identify a suitable initial input. Extensive experiments demonstrate that our approach surpasses current state-of-the-art jailbreak methods in terms of both efficiency and effectiveness. Moreover, our approach exhibits superior cross-class generalization abilities. Haoxuan Ji, Zhenxing Niu, Xinbo Gao 0001, Gang Hua 0001 |
AAAI | 3 |
| 2025 | EvoFormer: Learning Dynamic Graph-Level Representations with Structural and Temporal Bias CorrectionabstractDynamic graph-level embedding aims to capture structural evolution in networks, which is essential for modeling real-world scenarios. However, existing methods face two critical yet under-explored issues: Structural Visit Bias, where random walk sampling disproportionately emphasizes high-degree nodes, leading to redundant and noisy structural representations; and Abrupt Evolution Blindness, the failure to effectively detect sudden structural changes due to rigid or overly simplistic temporal modeling strategies, resulting in inconsistent temporal embeddings. To overcome these challenges, we propose EvoFormer, an evolution-aware Transformer framework tailored for dynamic graph-level representation learning. To mitigate Structural Visit Bias, EvoFormer introduces a Structure-Aware Transformer Module that incorporates positional encoding based on node structural roles, allowing the model to globally differentiate and accurately represent node structures. To overcome Abrupt Evolution Blindness, EvoFormer employs an Evolution-Sensitive Temporal Module, which explicitly models temporal evolution through a sequential three-step strategy: (I) Random Walk Timestamp Classification, generating initial timestamp-aware graph-level embeddings; (II) Graph-Level Temporal Segmentation, partitioning the graph stream into segments reflecting structurally coherent periods; and (III) Segment-Aware Temporal Self-Attention combined with an Edge Evolution Prediction task, enabling the model to precisely capture segment boundaries and perceive structural evolution trends, effectively adapting to rapid temporal shifts. Extensive evaluations on five benchmark datasets confirm that EvoFormer achieves state-of-the-art performance in graph similarity ranking, temporal anomaly detection, and temporal segmentation tasks, validating its effectiveness in correcting structural and temporal biases. Code is available at https://github.com/zlx0823/EvoFormerCode. Haodi Zhong, Liuxin Zou, Di Wang 0011, Bo Wan 0002, Zhenxing Niu, Quan Wang 0006 |
CIKM | 5 |
| 2025 | Towards Aligned Data Forgetting via Twin Machine UnlearningabstractModern privacy regulations have spurred the evolution of machine unlearning, a technique enabling a trained model to efficiently forget specific training data. In prior unlearning methods, the concept of “data forgetting” is often interpreted and implemented as achieving zero classification accuracy on such data. Nevertheless, the authentic aim of machine unlearning is to achieve alignment between the unlearned model and the gold model, i.e., encouraging them to have identical classification accuracy. On the other hand, the gold model often exhibits non-zero classification accuracy due to its generalization ability. To achieve aligned data forgetting, we propose a Twin Machine Unlearning (TMU) approach, where a twin unlearning problem is defined corresponding to the original unlearning problem. Consequently, the generalization-label predictor trained on the twin problem can be transferred to the original problem, facilitating aligned data forgetting. Comprehensive empirical experiments illustrate that our approach significantly enhances the alignment between the unlearned model and the gold model. Haoxuan Ji, Yuyao Sun, Fei Gao 0006, Haichang Gao, Zhenxing Niu |
ICME | 7 |
| 2025 | Revisiting 3D point cloud analysis with Markov process
Chenru Jiang, Wuwei Ma, Kaizhu Huang, Qiufeng Wang 0001, Xi Yang 0008, Weiguang Zhao, Junwei Wu 0001, Xinheng Wang 0001, Jimin Xiao, Zhenxing Niu |
Pattern Recognit. | 10 |
| 2025 | A Three-Stage Decision-Making Method Based on Machine Learning for Preventive Maintenance of Airport PavementabstractThe goal of preventative maintenance (PM) decision-making on airport pavements is to deploy the appropriate maintenance countermeasures at the correct time. This paper proposed a three-stage method for maintenance based on machine learning, which further refined the PM decision-making process. First, a pavement maintenance level model was developed using the PCA and PSO algorithm optimized SVM model. The model was then used to separate pavement maintenance into three categories: daily, PM, and major. Second, the DBSCAN and OPTICS were utilized to further divide the PM requirements finely. In order to implement the scientific decision-making of PM, suitable maintenance procedures were ultimately chosen based on the predominant damage kinds of the pavement units. The results showed that, when compared to the original SVM model, the classification accuracy of the PCA-PSO-SVM model was greatly improved, with total accuracy and accuracy of each class increasing by 10%, 41.7%, 4.6%, and 7.8%, respectively. When clustering the airport pavement performance dataset, OPTICS outperformed the DBSCAN technique. Four groups of PM demands were discovered by visualizing the best grouping levels after dimensionality reduction. Zhenxing Niu, Yinzhang He, Qinshi Hu, Jiupeng Zhang |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2024 | Towards Unified Robustness Against Both Backdoor and Adversarial AttacksabstractDeep Neural Networks (DNNs) are known to be vulnerable to both backdoor and adversarial attacks. In the literature, these two types of attacks are commonly treated as distinct robustness problems and solved separately, since they belong to training-time and inference-time attacks respectively. However, this paper revealed that there is an intriguing connection between them: (1) planting a backdoor into a model will significantly affect the model's adversarial examples and (2) for an infected model, its adversarial examples have similar features as the triggered images. Based on these observations, a novel Progressive Unified Defense (PUD) algorithm is proposed to defend against backdoor and adversarial attacks simultaneously. Specifically, our PUD has a progressive model purification scheme to jointly erase backdoors and enhance the model's adversarial robustness. At the early stage, the adversarial examples of infected models are utilized to erase backdoors. With the backdoor gradually erased, our model purification can naturally turn into a stage to boost the model's robustness against adversarial attacks. Besides, our PUD algorithm can effectively identify poisoned images, which allows the initial extra dataset not to be completely clean. Extensive experimental results show that, our discovered connection between backdoor and adversarial attacks is ubiquitous, no matter what type of backdoor attack. The proposed PUD outperforms the state-of-the-art backdoor defense, including the model repairing-based and data filtering-based methods. Besides, it also has the ability to compete with the most advanced adversarial defense methods. The code is available here. Zhenxing Niu, Yuyao Sun, Qiguang Miao, Rong Jin 0001, Gang Hua 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | Adversarial Attack and Defense in Deep RankingabstractDeep Neural Network classifiers are vulnerable to adversarial attacks, where an imperceptible perturbation could result in misclassification. However, the vulnerability of DNN-based image ranking systems remains under-explored. In this paper, we propose two attacks against deep ranking systems, i.e., Candidate Attack and Query Attack, that can raise or lower the rank of chosen candidates by adversarial perturbations. Specifically, the expected ranking order is first represented as a set of inequalities. Then a triplet-like objective function is designed to obtain the optimal perturbation. Conversely, an anti-collapse triplet defense is proposed to improve the ranking model robustness against all proposed attacks, where the model learns to prevent the adversarial attack from pulling the positive and negative samples close to each other. To comprehensively measure the empirical adversarial robustness of a ranking model with our defense, we propose an empirical robustness score, which involves a set of representative attacks against ranking models. Our adversarial ranking attacks and defenses are evaluated on MNIST, Fashion-MNIST, CUB200-2011, CARS196, and Stanford Online Products datasets. Experimental results demonstrate that our attacks can effectively compromise a typical deep ranking system. Nevertheless, our defense can significantly improve the ranking system's robustness and simultaneously mitigate a wide range of attacks. Le Wang 0003, Zhenxing Niu, Qilin Zhang 0004, Nanning Zheng 0001, Gang Hua 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | ShiftAttack: Toward Attacking the Localization Ability of Object DetectorabstractState-of-the-art (SOTA) adversarial attacks expose vulnerabilities in object detectors, often resulting in erroneous predictions. However, existing adversarial attacks neglect the stealth and flexibility of adversarial examples, which are crucial for conducting contextually consistent and inconspicuous attacks. To address these issues, leveraging the observed phenomenon of predicted box offsets in real-world object detection scenarios, this paper presents a novel adversarial attack framework called ShiftAttack. It leverages the concept of dense detection in prevalent object detectors, by boosting the confidence of low Intersection over Union (IoU) predictions within the positive samples (the set of predicted boxes responsible for localizing the same target), which leads to the erroneous exclusion of true positive predictions during the post-processing stage. Such a paradigm is highly stealthy as the shifted predictions seem like natural detector mistakes rather than obvious manipulations. To enhance the flexibility of ShiftAttack this paper proposes a generative approach called ShiftAttack Generator (SAG), which can not only shift predicted boxes for any target in arbitrary directions and distances but also facilitate adaptive feature exchange between pre- and post-shift regions to optimize the attack. Additionally, the proposed SAG incorporates the Dynamic Hinge Loss (DHL) to ensure the imperceptibility of perturbations, effectively mitigating the Patch-Pattern associated with the use of$\mathcal {L}_{2}$norm. Extensive experiments confirm that SAG surpasses other SOTA adversarial attacks in effectiveness, speed and stealthiness. Hao Li 0009, Maoguo Gong, Shiguo Chen, A. K. Qin 0001, Zhenxing Niu, Yue Wu 0004, Yu Zhou 0051 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2023 | Progressive Backdoor Erasing via connecting Backdoor and Adversarial AttacksabstractDeep neural networks (DNNs) are known to be vulnera-ble to both backdoor attacks as well as adversarial attacks. In the literature, these two types of attacks are commonly treated as distinct problems and solved separately, since they belong to training-time and inference-time attacks respectively. However, in this paper we find an intriguing connection between them: for a model planted with backdoors, we observe that its adversarial examples have similar behaviors as its triggered images, i.e., both activate the same subset of DNN neurons. It indicates that planting a back-door into a model will significantly affect the model's adversarial examples. Based on these observations, a novel Progressive Backdoor Erasing (PBE) algorithm is proposed to progressively purify the infected model by leveraging un-targeted adversarial attacks. Different from previous back-door defense methods, one significant advantage of our approach is that it can erase backdoor even when the clean extra dataset is unavailable. We empirically show that, against 5 state-of-the-art backdoor attacks, our PBE can effectively erase the backdoor without obvious performance degradation on clean samples and outperforms existing de-fense methods. Bingxu Mu, Zhenxing Niu, Le Wang 0003, Xue Wang 0010, Qiguang Miao, Rong Jin 0001, Gang Hua 0001 |
CVPR | 2 |
| 2023 | Towards Multi-Person Gesture Recognition using Commodity Wi-FiabstractComparing the recognition of human gestures using cameras, radar, or LiDAR, a WiFi-based gesture recognition system has distinct advantages, such as being low-cost, being device-free, and having much less privacy leakage. Recently, there have been advancements in WiFi-based gesture recognition, but most of the research has primarily focused on single-person gesture recognition. However, in real-world scenarios like e-learning, it is common for multiple individuals to engage in different actions simultaneously. To this end, this paper focuses on multi-person gesture recognition, which presents two major challenges, that is, the recognition accuracy due to the WiFi signal interference, and the processing time for real-time applications. Multi-person gesture recognition is more challenging than single-person scenario due to the interference caused by the superposition of WiFi signals induced by multiple moving individuals. In this paper, we define a concept of super-gesture and propose a WiFi-based Super-Gesture recognition (WiSG) method. Through the decomposition of the super-gesture's DFS spectrogram by Multi-Motion Trajectory algorithm, we extract modified signals of each person. Moreover, a novel feature called Field Motion Velocity is proposed by fully exploiting the advantages of our multiple transmitter-receiver WiFi sensing system. The proposed feature is not significantly affected by domains such as position, orientation, and other factors irrelevant to gestures. As a result, our approach can effectively recognize gestures across different domains. Evaluation results show that the cross-domain recognition accuracy of our WiSG can achieve up to 89% in multi-person scenario. Moreover, our approach can reduce processing time by 20 times against Widar3.0, which satisfies the requirements of most real-time applications. Xiaozhuang Liu, Zhenxing Niu, Wenye Wang |
ICCCN | 2 |
| 2023 | Structure First Detail Next: Image Inpainting with Pyramid GeneratorabstractRecent deep generative models have achieved promising performance in image inpainting. However, it is still challenging for a neural network to generate realistic image details and textures due to its inherent spectral bias. We suggest adopting a ‘structure first detail next’ workflow for image inpainting by knowing how artists work. Thus, we propose to build a Pyramid Generator by stacking several sub-generators, where lower-layer sub-generators focus on restoring image structures. In contrast, the higher-layer sub-generators emphasize image details. Our model progressively restores the input through the entire pyramid in a bottom-up fashion. Notably, our approach has a learning scheme of progressively increasing hole size, which allows it to restore large-hole images. In addition, our method could fully exploit the benefits of learning with high-resolution images and hence is suitable for high-resolution image inpainting. Extensive experimental results on benchmark datasets have validated the effectiveness of our approach compared with state-of-the-arts. Shuyi Qu, Zhenxing Niu, Jianke Zhu, Bin Dong 0003, Kaizhu Huang |
ICME | 2 |
| 2023 | Adaptive Ladder Loss for Learning Coherent Visual-Semantic EmbeddingabstractFor visual-semantic embedding, the existing methods normally treat the relevance between queries and candidates in a bipolar way – relevant or irrelevant, and all “irrelevant” candidates are uniformly pushed away from the query by an equal margin in the embedding space, regardless of their various proximity to the query. This practice disregards relatively discriminative information and could lead to suboptimal ranking in the retrieval results and poorer user experience, especially in the long-tail query scenario where a matching candidate may not necessarily exist. In this paper, we introduce a continuous variable to model the relevance degree between queries and multiple candidates, and propose to learn a coherent embedding space, where candidates with higher relevance degrees are mapped closer to the query than those with lower relevance degrees. In particular, the new ladder loss is proposed by extending the triplet loss inequality to a more general inequality chain, which implements variable push-away margins according to respective relevance degrees. To adapt to the varying mini-batch statistics and improve the efficiency of the ladder loss, we also propose a Silhouette score-based method to adaptively decide the ladder level and hence the underlying inequality chain. In addition, a proper Coherent Score metric is proposed to better measure the ranking results including those “irrelevant” candidates. Extensive experiments on multiple datasets validate the efficacy of our proposed method, which achieves significant improvement over existing state-of-the-art methods. Le Wang 0003, Zhenxing Niu, Qilin Zhang 0004, Nanning Zheng 0001 |
IEEE Trans. Multim. | 3 |
| 2022 | Semantic-shape Adaptive Feature Modulation for Semantic Image SynthesisabstractRecent years have witnessed substantial progress in se-mantic image synthesis, it is still challenging in synthesizing photo-realistic images with rich details. Most previ-ous methods focus on exploiting the given semantic map, which just captures an object-level layout for an image. Obviously, a fine-grained part-level semantic layout will benefit object details generation, and it can be roughly in-ferred from an object's shape. In order to exploit the part-level layouts, we propose a Shape-aware Position Descrip-tor (SPD) to describe each pixel's positional feature, where object shape is explicitly encoded into the SP D feature. Fur-thermore, a Semantic-shape Adaptive Feature Modulation (SAFM) block is proposed to combine the given semantic map and our positional features to produce adaptively mod-ulated features. Extensive experiments demonstrate that the proposed SPD and SAFM significantly improve the gener-ation of objects with rich details. Moreover, our method performs favorably against the SOTA methods in terms of quantitative and qualitative evaluation. The source code and model are available at SAFM. Zhengyao Lv, Xiaoming Li 0002, Zhenxing Niu, Bing Cao 0002, Wangmeng Zuo |
CVPR | 3 |
| 2022 | Unidirectional Video Denoising by Mimicking Backward Recurrent Modules with Look-Ahead Forward Ones
Junyi Li 0005, Xiaohe Wu, Zhenxing Niu, Wangmeng Zuo |
ECCV (18) | 3 |
| 2022 | C$^{2}$MT: A Credible and Class-Aware Multi-Task Transformer for SR-IQAabstractIn this letter a novel credible and class-aware multi-task transformer abbreviated as C$^{2}$MT for SRIQA, is proposed. In the proposed C$^{2}$MT, a quality-aware task for the quality prediction and the other class-aware task for the classification of SR algorithms are jointly framed to mine mutual information between the quality of SR images and the class of SR algorithms for more discriminative perceptual representation. In the class-aware task, we develop a supervised contrastive learning strategy to learn embedding perceived features related to a class-specific SR algorithm. While in the other quality-aware task, we employ a novel credible pseudo quality label generation strategy to actively adjust the quality labels by ranking the pair-wise consistency between the predicted quality scores and subjective perceptual scores but keep the image-level quality labels unchanged. The developed supervised contrastive learning and the variant of active learning strategies benefit learning a more consistent quality predictor for SR images. Experiment results indicate that our proposed C$^{2}$MT achieves state-of-the-art results on five popular SRIQA benchmark databases. Kaibing Zhang, Zhenxing Niu |
IEEE Signal Process. Lett. | 3 |
| 2021 | SGCN: Sparse Graph Convolution Network for Pedestrian Trajectory PredictionabstractPedestrian trajectory prediction is a key technology in autopilot, which remains to be very challenging due to complex interactions between pedestrians. However, previous works based on dense undirected interaction suffer from modeling superfluous interactions and neglect of trajectory motion tendency, and thus inevitably result in a considerable deviance from the reality. To cope with these issues, we present a Sparse Graph Convolution Network (SGCN) for pedestrian trajectory prediction. Specifically, the SGCN explicitly models the sparse directed interaction with a sparse directed spatial graph to capture adaptive interaction pedestrians. Meanwhile, we use a sparse directed temporal graph to model the motion tendency, thus to facilitate the prediction based on the observed direction. Finally, parameters of a bi-Gaussian distribution for trajectory prediction are estimated by fusing the above two sparse graphs. We evaluate our proposed method on the ETH and UCY datasets, and the experimental results show our method outperforms comparative state-of-the-art methods by 9% in Average Displacement Error (ADE) and 13% in Final Displacement Error (FDE). Notably, visualizations indicate that our method can capture adaptive interactions between pedestrians and their effective motion tendencies. Liushuai Shi, Le Wang 0003, Chengjiang Long, Sanping Zhou, Zhenxing Niu, Gang Hua 0001 |
CVPR | 6 |
| 2021 | Boosting Weakly Supervised Object Detection via Learning Bounding Box AdjustersabstractWeakly-supervised object detection (WSOD) has emerged as an inspiring recent topic to avoid expensive instance-level object annotations. However, the bounding boxes of most existing WSOD methods are mainly determined by precomputed proposals, thereby being limited in precise object localization. In this paper, we defend the problem setting for improving localization performance by leveraging the bounding box regression knowledge from a well-annotated auxiliary dataset. First, we use the well-annotated auxiliary dataset to explore a series of learnable bounding box adjusters (LBBAs) in a multi-stage training manner, which is class-agnostic. Then, only LBBAs and a weakly-annotated dataset with non-overlapped classes are used for training LBBA-boosted WSOD. As such, our LBBAs are practically more convenient and economical to implement while avoiding the leakage of the auxiliary well-annotated dataset. In particular, we formulate learning bounding box adjusters as a bi-level optimization problem and suggest an EM-like multi-stage training algorithm. Then, a multi-stage scheme is further presented for LBBA-boosted WSOD. Additionally, a masking strategy is adopted to improve proposal classification. Experimental results verify the effectiveness of our method. Our method performs favorably against state-of-the-art WSOD methods and knowledge transfer model with similar problem setting. Code is publicly available at https://github.com/DongSky/lbba_boosted_wsod. Bowen Dong 0001, Zitong Huang, Yuelin Guo, Qilong Wang 0001, Zhenxing Niu, Wangmeng Zuo |
ICCV | 5 |
| 2021 | Unlimited Neighborhood Interaction for Heterogeneous Trajectory PredictionabstractUnderstanding complex social interactions among agents is a key challenge for trajectory prediction. Most existing methods consider the interactions between pairwise traffic agents or in a local area, while the nature of interactions is unlimited, involving an uncertain number of agents and non-local areas simultaneously. Besides, they treat heterogeneous traffic agents the same, namely those among agents of different categories, while neglecting people’s diverse reaction patterns toward traffic agents in different categories. To address these problems, we propose a simple yet effective Unlimited Neighborhood Interaction Network (UNIN), which predicts trajectories of heterogeneous agents in multiple categories. Specifically, the proposed unlimited neighborhood interaction module generates the fused-features of all agents involved in an interaction simultaneously, which is adaptive to any number of agents and any range of interaction area. Meanwhile, a hierarchical graph attention module is proposed to obtain category-to-category interaction and agent-to-agent interaction. Finally, parameters of a Gaussian Mixture Model are estimated for generating the future trajectories. Extensive experimental results on benchmark datasets demonstrate a significant performance improvement of our method over the state-of-the-art methods. Fang Zheng 0009, Le Wang 0003, Sanping Zhou, Wei Tang 0016, Zhenxing Niu, Nanning Zheng 0001, Gang Hua 0001 |
ICCV | 5 |
| 2021 | Practical Relative Order Attack in Deep RankingabstractRecent studies unveil the vulnerabilities of deep ranking models, where an imperceptible perturbation can trigger dramatic changes in the ranking result. While previous attempts focus on manipulating absolute ranks of certain candidates, the possibility of adjusting their relative order remains under-explored. In this paper, we formulate a new adversarial attack against deep ranking systems, i.e., the Order Attack, which covertly alters the relative order among a selected set of candidates according to an attacker-specified permutation, with limited interference to other unrelated candidates. Specifically, it is formulated as a triplet-style loss imposing an inequality chain reflecting the specified permutation. However, direct optimization of such white-box objective is infeasible in a real-world attack scenario due to various black-box limitations. To cope with them, we propose a Short-range Ranking Correlation metric as a surrogate objective for black-box Order Attack to approximate the white-box method. The Order Attack is evaluated on the Fashion-MNIST and Stanford-Online-Products datasets under both white-box and black-box threat models. The black-box attack is also successfully implemented on a major e-commerce platform. Comprehensive experimental evaluations demonstrate the effectiveness of the proposed methods, revealing a new type of ranking model vulnerability. Le Wang 0003, Zhenxing Niu, Qilin Zhang 0004, Nanning Zheng 0001, Gang Hua 0001 |
ICCV | 3 |
| 2021 | Object Cosegmentation in Noisy Videos With Multilevel HypergraphabstractWith the target of simultaneously segmenting semantically related videos to identify the common objects, video object cosegmentation has attracted the attention of researchers in recent years. Existing methods are primarily based on pair-wise relations between adjacent pixels and regions, which are susceptible to performance degradation from object entries/exists or occlusions. Specifically, we refer these video frames without the common objects present as the “empty” frames. In this paper, we propose a multilevel hypergraph-based full Video object CoSegmentation (VCS) method, which incorporates high-level semantics and low-level appearance/motion/saliency to construct the hyperedge among multiple spatially and temporally adjacent regions. Specifically, the high-level semantic model fuses multiple object proposals from each frame instead of relying on a single object proposal per frame. A hypergraph cut is subsequently utilized to calculate the object cosegmentation. Experiments on four video object segmentation/cosegmentation datasets against state-of-the-art methods with both objective and subjective results manifest the effectiveness of the proposed VCS method, including the SegTrack and VCoSeg datasets without “empty” frames, the XJTU-Stevens dataset with 3.7% “empty” frames, and the Noisy-ViCoSeg dataset proposed together with our method with 30.3% “empty” frames. Le Wang 0003, Qilin Zhang 0004, Zhenxing Niu, Nanning Zheng 0001, Gang Hua 0001 |
IEEE Trans. Multim. | 4 |
| 2020 | Ladder Loss for Coherent Visual-Semantic EmbeddingabstractFor visual-semantic embedding, the existing methods normally treat the relevance between queries and candidates in a bipolar way – relevant or irrelevant, and all “irrelevant” candidates are uniformly pushed away from the query by an equal margin in the embedding space, regardless of their various proximity to the query. This practice disregards relatively discriminative information and could lead to suboptimal ranking in the retrieval results and poorer user experience, especially in the long-tail query scenario where a matching candidate may not necessarily exist. In this paper, we introduce a continuous variable to model the relevance degree between queries and multiple candidates, and propose to learn a coherent embedding space, where candidates with higher relevance degrees are mapped closer to the query than those with lower relevance degrees. In particular, the new ladder loss is proposed by extending the triplet loss inequality to a more general inequality chain, which implements variable push-away margins according to respective relevance degrees. In addition, a proper Coherent Score metric is proposed to better measure the ranking results including those “irrelevant” candidates. Extensive experiments on multiple datasets validate the efficacy of our proposed method, which achieves significant improvement over existing state-of-the-art methods. Zhenxing Niu, Le Wang 0003, Zhanning Gao, Qilin Zhang 0004, Gang Hua 0001 |
AAAI | 2 |
| 2020 | Adversarial Ranking Attack and Defense
Zhenxing Niu, Le Wang 0003, Qilin Zhang 0004, Gang Hua 0001 |
ECCV (14) | 2 |
| 2020 | Fine-Grained Giant Panda IdentificationabstractThe image-based fine-grained identification of individual giant pandas (Ailuropoda melanoleuca) is an emerging technology, and it is extraordinarily challenging due to the extremely subtle visual differences between individual giant pandas and limited annotated training data. To address these challenges, we propose the Feature-Fusion Convolutional Neural Network with Patch Detector (FFCNN-PD) algorithm, which exploits the discriminative local patches and builds a hierarchical representation generated by fusing both global and local features. Specifically, an attentional cross-channel pooling is embedded in the FFCNN-PD to improve the class- specific patch detectors. In addition, we propose a new giant panda identification dataset (iPanda-30) to establish a benchmark. Experiments on the proposed iPanda-30 dataset and other fine-grained recognition datasets demonstrate the effectiveness of the FFCNN-PD algorithm against the existing state-of-the-arts. Rizhi Ding, Le Wang 0003, Qilin Zhang 0004, Zhenxing Niu, Nanning Zheng 0001, Gang Hua 0001 |
ICASSP | 4 |
| 2020 | Action Co-localization in an Untrimmed Video by Graph Neural Networks
Changbo Zhai, Le Wang 0003, Qilin Zhang 0004, Zhanning Gao, Zhenxing Niu, Nanning Zheng 0001, Gang Hua 0001 |
MMM (1) | 5 |
| 2020 | Incremental focal loss GANs
Fei Gao 0006, Jingjie Zhu, Hanliang Jiang, Zhenxing Niu, Weidong Han 0001, Jun Yu 0002 |
Inf. Process. Manag. | 4 |
| 2019 | Video Imprint Segmentation for Temporal Action Detection in Untrimmed VideosabstractWe propose a temporal action detection by spatial segmentation framework, which simultaneously categorize actions and temporally localize action instances in untrimmed videos. The core idea is the conversion of temporal detection task into a spatial semantic segmentation task. Firstly, the video imprint representation is employed to capture the spatial/temporal interdependences within/among frames and represent them as spatial proximity in a feature space. Subsequently, the obtained imprint representation is spatially segmented by a fully convolutional network. With such segmentation labels projected back to the video space, both temporal action boundary localization and per-frame spatial annotation can be obtained simultaneously. The proposed framework is robust to variable lengths of untrimmed videos, due to the underlying fixed-size imprint representations. The efficacy of the framework is validated in two public action detection datasets. Zhanning Gao, Le Wang 0003, Qilin Zhang 0004, Zhenxing Niu, Nanning Zheng 0001, Gang Hua 0001 |
AAAI | 4 |
| 2019 | Weakly Supervised Temporal Action Localization Through Contrast Based Evaluation NetworksabstractWeakly-supervised temporal action localization (WS-TAL) is a promising but challenging task with only video-level action categorical labels available during training. Without requiring temporal action boundary annotations in training data, WS-TAL could possibly exploit automatically retrieved video tags as video-level labels. However, such coarse video-level supervision inevitably incurs confusions, especially in untrimmed videos containing multiple action instances. To address this challenge, we propose the Contrast-based Localization EvaluAtioN Network (CleanNet) with our new action proposal evaluator, which provides pseudo-supervision by leveraging the temporal contrast in snippet-level action classification predictions. Essentially, the new action proposal evaluator enforces an additional temporal contrast constraint so that high-evaluation-score action proposals are more likely to coincide with true action instances. Moreover, the new action localization module is an integral part of CleanNet which enables end-to-end training. This is in contrast to many existing WS-TAL methods where action localization is merely a post-processing step. Experiments on THUMOS14 and ActivityNet datasets validate the efficacy of CleanNet against existing state-ofthe- art WS-TAL algorithms. Ziyi Liu 0001, Le Wang 0003, Qilin Zhang 0004, Zhanning Gao, Zhenxing Niu, Nanning Zheng 0001, Gang Hua 0001 |
ICCV | 5 |
| 2019 | Video ImprintabstractA new unified video analytics framework (ER3) is proposed for complex event retrieval, recognition and recounting, based on the proposed video imprint representation, which exploits temporal correlations among image features across video frames. With the video imprint representation, it is convenient to reverse map back to both temporal and spatial locations in video frames, allowing for both key frame identification and key areas localization within each frame. In the proposed framework, a dedicated feature alignment module is incorporated for redundancy removal across frames to produce the tensor representation, i.e., the video imprint. Subsequently, the video imprint is individually fed into both a reasoning network and a feature aggregation module, for event recognition/recounting and event retrieval tasks, respectively. Thanks to its attention mechanism inspired by the memory networks used in language modeling, the proposed reasoning network is capable of simultaneous event category recognition and localization of the key pieces of evidence for event recounting. In addition, the latent structure in our reasoning network highlights the areas of the video imprint, which can be directly used for event recounting. With the event retrieval task, the compact video representation aggregated from the video imprint contributes to better retrieval results than existing state-of-the-art methods. Zhanning Gao, Le Wang 0003, Nebojsa Jojic, Zhenxing Niu, Nanning Zheng 0001, Gang Hua 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2019 | Discovering Latent Topics by Gaussian Latent Dirichlet Allocation and Spectral ClusteringabstractToday, diversifying the retrieval results of a certain query will improve customers’ search efficiency. Showing the multiple aspects of information provides users an overview of the object, which helps them fast target their demands. To discover aspects, research focuses on generating image clusters from initially retrieved results. As an effective approach, latent Dirichlet allocation (LDA) has been proved to have good performance on discovering high-level topics. However, traditional LDA is designed to process textual words, and it needs the input as discrete data. When we apply this algorithm to process continuous visual images, a common solution is to quantize the continuous features into discrete form by a bag-of-visual-words algorithm. During this process, quantization error will lead to information that inevitably is lost. To construct a topic model with complete visual information, this work applies Gaussian latent Dirichlet allocation (GLDA) on the diversity issue of image retrieval. In this model, traditional multinomial distribution is substituted with Gaussian distribution to model continuous visual features. In addition, we propose a two-phase spectral clustering strategy, called dual spectral clustering , to generate clusters from region level to image level. The experiments on the challenging landmarks of the DIV400 database show that our proposal improves relevance and diversity by about 10% compared to traditional topic models. Xinbo Gao 0001, Zhenxing Niu, Qi Tian 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2018 | Joint Spatio-Temporal Action Localization in Untrimmed Videos with Per-Frame SegmentationabstractInspired by the recent spatio-temporal action localization efforts with tubelets (sequences of bounding boxes), we present a new spatio-temporal action detector Segment-tube, which consists of sequences of per-frame segmentation masks. The proposed Segment-tube detector can temporally pinpoint the starting/ending frame of each action class in the presence of preceding/subsequent interference actions in untrimmed videos. Simultaneously, the Segment-tube detector produces per-frame segmentation masks instead of bounding boxes, offering superior spatial accuracy to tubelets. This is achieved by alternating iterative optimization between temporal action localization and spatial action segmentation. Experimental results on multiple datasets validate the efficacy of the proposed detector. Xuhuan Duan, Le Wang 0003, Changbo Zhai, Nanning Zheng 0001, Qilin Zhang 0004, Zhenxing Niu, Gang Hua 0001 |
ICIP | 6 |
| 2018 | Joint Video Object Discovery and Segmentation by Coupled Dynamic Markov NetworksabstractIt is a challenging task to extract segmentation mask of a target from a single noisy video, which involves object discovery coupled with segmentation. To solve this challenge, we present a method to jointly discover and segment an object from a noisy video, where the target disappears intermittently throughout the video. Previous methods either only fulfill video object discovery, or video object segmentation presuming the existence of the object in each frame. We argue that jointly conducting the two tasks in a unified way will be beneficial. In other words, video object discovery and video object segmentation tasks can facilitate each other. To validate this hypothesis, we propose a principled probabilistic model, where two dynamic Markov networks are coupled-one for discovery and the other for segmentation. When conducting the Bayesian inference on this model using belief propagation, the bi-directional message passing reveals a clear collaboration between these two inference tasks. We validated our proposed method in five data sets. The first three video data sets, i.e., the SegTrack data set, the YouTube-objects data set, and the Davis data set, are not noisy, where all video frames contain the objects. The two noisy data sets, i.e., the XJTU-Stevens data set, and the Noisy-ViDiSeg data set, newly introduced in this paper, both have many frames that do not contain the objects. When compared with state of the art, it is shown that although our method produces inferior results on video data sets without noisy frames, we are able to obtain better results on video data sets with noisy frames. Ziyi Liu 0001, Le Wang 0003, Gang Hua 0001, Qilin Zhang 0004, Zhenxing Niu, Ying Wu 0001, Nanning Zheng 0001 |
IEEE Trans. Image Process. | 5 |
| 2018 | Knowledge-Based Topic Model for Unsupervised Object Discovery and LocalizationabstractUnsupervised object discovery and localization is to discover some dominant object classes and localize all of object instances from a given image collection without any supervision. Previous work has attempted to tackle this problem with vanilla topic models, such as latent Dirichlet allocation (LDA). However, in those methods no prior knowledge for the given image collection is exploited to facilitate object discovery. On the other hand, the topic models used in those methods suffer from the topic coherence issue-some inferred topics do not have clear meaning, which limits the final performance of object discovery. In this paper, prior knowledge in terms of the so-called must-links are exploited from Web images on the Internet. Furthermore, a novel knowledge-based topic model, called LDA with mixture of Dirichlet trees, is proposed to incorporate the must-links into topic modeling for object discovery. In particular, to better deal with the polysemy phenomenon of visual words, the must-link is re-defined as that one must-link only constrains one or some topic(s) instead of all topics, which leads to significantly improved topic coherence. Moreover, the must-links are built and grouped with respect to specific object classes, thus the must-links in our approach are semantic-specific, which allows to more efficiently exploit discriminative prior knowledge from Web images. Extensive experiments validated the efficiency of our proposed approach on several data sets. It is shown that our method significantly improves topic coherence and outperforms the unsupervised methods for object discovery and localization. In addition, compared with discriminative methods, the naturally existing object classes in the given image collection can be subtly discovered, which makes our approach well suited for realistic applications of unsupervised object discovery. Zhenxing Niu, Gang Hua 0001, Le Wang 0003, Xinbo Gao 0001 |
IEEE Trans. Image Process. | 1 |
| 2017 | Hierarchical Multimodal LSTM for Dense Visual-Semantic EmbeddingabstractWe address the problem of dense visual-semantic embedding that maps not only full sentences and whole images but also phrases within sentences and salient regions within images into a multimodal embedding space. Such dense embeddings, when applied to the task of image captioning, enable us to produce several region-oriented and detailed phrases rather than just an overview sentence to describe an image. Specifically, we present a hierarchical structured recurrent neural network (RNN), namely Hierarchical Multimodal LSTM (HM-LSTM). Compared with chain structured RNN, our proposed model exploits the hierarchical relations between sentences and phrases, and between whole images and image regions, to jointly establish their representations. Without the need of any supervised labels, our proposed model automatically learns the fine-grained correspondences between phrases and image regions towards the dense embedding. Extensive experiments on several datasets validate the efficacy of our method, which compares favorably with the state-of-the-art methods. Zhenxing Niu, Le Wang 0003, Xinbo Gao 0001, Gang Hua 0001 |
ICCV | 1 |
| 2017 | Video Object Discovery and Co-Segmentation with Extremely Weak SupervisionabstractWe present a spatio-temporal energy minimization formulation for simultaneous video object discovery and co-segmentation across multiple videos containing irrelevant frames. Our approach overcomes a limitation that most existing video co-segmentation methods possess, i.e., they perform poorly when dealing with practical videos in which the target objects are not present in many frames. Our formulation incorporates a spatio-temporal auto-context model, which is combined with appearance modeling for superpixel labeling. The superpixel-level labels are propagated to the frame level through a multiple instance boosting algorithm with spatial reasoning, based on which frames containing the target object are identified. Our method only needs to be bootstrapped with the frame-level labels for a few video frames (e.g., usually 1 to 3) to indicate if they contain the target objects or not. Extensive experiments on four datasets validate the efficacy of our proposed method: 1) object segmentation from a single video on the SegTrack dataset, 2) object co-segmentation from multiple videos on a video co-segmentation dataset, and 3) joint object discovery and co-segmentation from multiple videos containing irrelevant frames on the MOViCS dataset and XJTU-Stevens, a new dataset that we introduce in this paper. The proposed method compares favorably with the state-of-the-art in all of these experiments. Le Wang 0003, Gang Hua 0001, Rahul Sukthankar, Jianru Xue, Zhenxing Niu, Nanning Zheng 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2016 | Ordinal Regression with Multiple Output CNN for Age EstimationabstractTo address the non-stationary property of aging patterns, age estimation can be cast as an ordinal regression problem. However, the processes of extracting features and learning a regression model are often separated and optimized independently in previous work. In this paper, we propose an End-to-End learning approach to address ordinal regression problems using deep Convolutional Neural Network, which could simultaneously conduct feature learning and regression modeling. In particular, an ordinal regression problem is transformed into a series of binary classification sub-problems. And we propose a multiple output CNN learning algorithm to collectively solve these classification sub-problems, so that the correlation between these tasks could be explored. In addition, we publish an Asian Face Age Dataset (AFAD) containing more than 160K facial images with precise age ground-truths, which is the largest public age dataset to date. To the best of our knowledge, this is the first work to address ordinal regression problems by using CNN, and achieves the state-of-the-art performance on both the MORPH and AFAD datasets. Zhenxing Niu, Le Wang 0003, Xinbo Gao 0001, Gang Hua 0001 |
CVPR | 1 |
| 2015 | Visual Topic Network: Building better image representations for images in social media
Zhenxing Niu, Gang Hua 0001, Qi Tian 0001, Xinbo Gao 0001 |
Comput. Vis. Image Underst. | 1 |
| 2014 | Semi-supervised Relational Topic Model for Weakly Annotated Image Recognition in Social MediaabstractIn this paper, we address the problem of recognizing images with weakly annotated text tags. Most previous work either cannot be applied to the scenarios where the tags are loosely related to the images, or simply take a pre-fusion at the feature level or a post-fusion at the decision level to combine the visual and textual content. Instead, we first encode the text tags as the relations among the images, and then propose a semi-supervised relational topic model (ss-RTM) to explicitly model the image content and their relations. In such way, we can efficiently leverage the loosely related tags, and build an intermediate level representation for a collection of weakly annotated images. The intermediate level representation can be regarded as a mid-level fusion of the visual and textual content, which is able to explicitly model their intrinsic relationships. Moreover, image category labels are also modeled in the ss-RTM, and recognition can be conducted without training an additional discriminative classifier. Our extensive experiments on social multimedia datasets (images+tags) demonstrated the advantages of the proposed model. Zhenxing Niu, Gang Hua 0001, Xinbo Gao 0001, Qi Tian 0001 |
CVPR | 1 |
| 2014 | Personalized Visual Vocabulary Adaption for Social Image RetrievalabstractWith the popularity of mobile devices and social networks, users can easily build their personalized image sets. Thus, personalized image analysis, indexing, and retrieval have become important topics in social media analysis. Because of users' diverse preferences, their personalized image sets are usually related to specific topics and show large feature distribution bias from general Internet images. Therefore, the visual vocabulary trained on general Internet images may could not fit across users' personalized image sets very well. To improve the image retrieval performance on personalized image sets, we propose the personalized visual vocabulary adaption which removes non-discriminative visual words and replaces them with more exact and discriminative ones, i.e., adapt a general vocabulary toward a specific user's image set. The proposed algorithm updates the visual vocabulary during off-line feature quantization, and operates on a limited number of visual words, hence shows satisfying efficiency. Extensive experiments of image search on public datasets demonstrate the efficiency and superior performance of our approach. Zhenxing Niu, Shiliang Zhang, Xinbo Gao 0001, Qi Tian 0001 |
ACM Multimedia | 1 |
| 2012 | Context aware topic model for scene recognitionabstractWe present a discriminative latent topic model for scene recognition. The capacity of our model is originated from the modeling of two types of visual contexts, i.e., the category specific global spatial layout of different scene elements, and the reinforcement of the visual coherence in uniform local regions. In contrast, most previous methods for scene recognition either only modeled one of these two visual contexts, or just totally ignored both of them. We cast these two coupled visual contexts in a discriminative Latent Dirichlet Allocation framework, namely context aware topic model. Then scene recognition is achieved by Bayesian inference given a target image. Our experiments on several scene recognition benchmarks clearly demonstrated the advantages of the proposed model. Zhenxing Niu, Gang Hua 0001, Xinbo Gao 0001, Qi Tian 0001 |
CVPR | 1 |
| 2012 | Tactic analysis based on real-world ball trajectory in soccer video
Zhenxing Niu, Xinbo Gao 0001, Qi Tian 0001 |
Pattern Recognit. | 1 |
| 2011 | Spatial-DiscLDA for visual recognitionabstractTopic models such as pLSA, LDA and their variants have been widely adopted for visual recognition. However, most of the adopted models, if not all, are unsupervised, which neglected the valuable supervised labels during model training. In this paper, we exploit recent advancement in supervised topic modeling, more particularly, the DiscLDA model for object recognition. We extend it to a part based visual representation to automatically identify and model different object parts. We call the proposed model as Spatial-DiscLDA (S-DiscLDA). It models the appearances and locations of the object parts simultaneously, which also takes the supervised labels into consideration. It can be directly used as a classifier to recognize the object. This is performed by an approximate inference algorithm based on Gibbs sampling and bridge sampling methods. We examine the performance of our model by comparing its performance with another supervised topic model on two scene category datasets, i.e., LabelMe and UIUC-sport dataset. We also compare our approach with other approaches which model spatial structures of visual features on the popular Caltech-4 dataset. The experimental results illustrate that it provides competitive performance. Zhenxing Niu, Gang Hua 0001, Xinbo Gao 0001, Qi Tian 0001 |
CVPR | 1 |
| 2011 | Non-goal scene analysis for soccer video
Xinbo Gao 0001, Zhenxing Niu, Dacheng Tao, Xuelong Li 0001 |
Neurocomputing | 2 |
| 2010 | Real-world trajectory extraction for attack pattern analysis in soccer videoabstractMost existing approaches on tactic analysis of soccer video are based on mosaic trajectory analysis, which loses much semantic information comparing to the real-world trajectory. Without effective extraction of real-world trajectory, the tactic of soccer cannot be properly represented and analyzed from the perspective of soccer professionals. In this paper, a real-world trajectory extraction method is proposed. Moreover, six attack patterns are defined to represent tactic of soccer, and a novel attack pattern recognition algorithm is developed based on the analysis of the ball's state and real-world trajectory. To our best knowledge, this is the first work of its kind that systematically analyzes tactic of soccer based on real-world trajectory for broadcast soccer video. Our experiments demonstrate that the defined attack pattern can be effectively recognized. Zhenxing Niu, Qi Tian 0001, Xinbo Gao 0001 |
ACM Multimedia | 1 |
| 2008 | Semantic Video Shot Segmentation Based on Color Ratio Feature and SVMabstractWith the fast development of video semantic analysis, there has been increasing attention to the typical issue of the semantic analysis of soccer program. Based on the color feature analysis, this paper focuses on the video shot segmentation problem from the perspective of semantic analysis, i.e. the semantic shot segmentation. Most existing works segment and classify the shot by using the dominant color of the field in soccer video. In this paper, we extend the traditional dominant color to several semantic colors, and define the color ratio feature. Then, the support vector machine is used for shot classification. The experimental result shows that the color ratio feature is useful to improve the performance. Furthermore, considering the temporal variations in the semantic color due to environment, an adaptive semantic color extraction algorithm is proposed, and the influence of the number of semantic colors on classification precision is evaluated also. Zhenxing Niu, Xinbo Gao 0001, Dacheng Tao, Xuelong Li 0001 |
CW | 1 |