Weigang Zhang

dblp:93/4821 · DBLP profile ↗
← Back
82ranked-venue papers
7as first author
34since 2021 · last 2026
0000-0003-0042-7074ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 63 · 4 first-author · 25 since 2021Artificial intelligence and machine learning · 25 · 1 first-author · 10 since 2021Databases, data management, data science and information retrieval · 9 · 3 since 2021Computer networks · 4 · 1 first-author · 3 since 2021Systems, architecture and hardware · 2 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2026 AIFIND: Artifact-Aware Interpreting Fine-Grained Alignment for Incremental Face Forgery Detection
abstract
As forgery types continue to emerge consistently, Incremental Face Forgery Detection (IFFD) has become a crucial paradigm. However, existing methods typically rely on data replay or coarse binary supervision, which fails to explicitly constrain the feature space, leading to severe feature drift and catastrophic forgetting. To address this, we propose AIFIND, Artifact-Aware Interpreting Fine-Grained Alignment for Incremental Face Forgery Detection, which leverages semantic anchors to stabilize incremental learning. We design the Artifact-Driven Semantic Prior Generator to instantiate invariant semantic anchors, establishing a fixed coordinate system from low-level artifact cues. These anchors are injected into the image encoder via Artifact-Probe Attention, which explicitly constrains volatile visual features to align with stable semantic anchors. Adaptive Decision Harmonizer harmonizes the classifiers by preserving angular relationships of semantic anchors, maintaining geometric consistency across tasks. Extensive experiments on multiple incremental protocols validate the superiority of AIFIND.
Hao Wang 0035, Beichen Zhang 0006, Yanpei Gong, Shaoyi Fang, Zhaobo Qi, Yuanrong Xu, Xinyan Liu 0008, Weigang Zhang
ICMR8
2026 Distinguishing semantically similar queries in temporal video grounding via LLM-generated query
Yibo Dang, Zhaobo Qi, Xinyan Liu 0008, Xinzhe Han, Weigang Zhang
Multim. Syst.6
2026 Multimodal-guided mixture-of-experts bias removal strategy for natural language video localization
Xiaowen Ruan, Zhaobo Qi, Ruisi Chen, Yuanrong Xu, Beichen Zhang 0006, Weigang Zhang
Multim. Syst.6
2026 Dynamic example network for class-agnostic object counting
Xinyan Liu 0008, Guorong Li, Yuankai Qi, Ziheng Yan, Weigang Zhang, Laiyun Qing, Qingming Huang
Pattern Recognit.5
2026 Compactness driven Co-learning for crowd counting and localization
Ziheng Yan, Xinyan Liu 0008, Guorong Li, Weigang Zhang, Fang Wan 0001, Qingming Huang
Pattern Recognit.4
2025 Procedure Knowledge Decoupled Distillation Strategy for Procedure Planning in Instructional Videos
abstract
Procedure planning in instructional videos, producing a structured and plannable action sequence facilitating the transition from the start to the goal states, has achieved significant progress. The dominant single-branch non-autoregressive planning paradigm guides action sequence generation through action labels, overlooking the limitation of the absence of intermediate visual information. Hence, we introduce the procedure knowledge decoupled distillation strategy to address the above issue. This innovative strategy deliberately lets the teacher model see the real visual information among the start and goal states to enhance its action semantic understanding and relationship modeling ability, producing the potential probability distribution containing the real action class and other action classes that may occur. Accordingly, we introduce a decoupled intermediate information knowledge distillation loss, which comprises single action knowledge distillation and sequence distribution knowledge distillation for the student model. The former improves the student model's precise inference ability for individual actions by transferring knowledge of a single action target category using binary classification loss. Conversely, the latter uses MSE loss to constrain the student model to learn the action sequence probability distribution from the teacher model, thereby enhancing the student model's global planning capability. Extensive experiments on three datasets demonstrate that our strategy can improve the performance of multiple weakly supervised models, achieving promising procedure knowledge modeling ability and plug-and-play flexibility.
Xiaotian Pan, Zhaobo Qi, Yuanrong Xu, Weigang Zhang
AAAI5
2025 AdaPPA: Adaptive Position Pre-Fill Jailbreak Attack Approach Targeting LLMs
abstract
Jailbreak vulnerabilities in Large Language Models (LLMs) refer to methods that extract malicious content from the model by carefully crafting prompts or suffixes, which has garnered significant attention from the research community. However, traditional attack methods, which primarily focus on the semantic level, are easily detected by the model. These methods overlook the difference in the model’s alignment protection capabilities at different output stages. To address this issue, we propose an adaptive position pre-fill jailbreak attack approach for executing jailbreak attacks on LLMs. Our method leverages the model’s instruction-following capabilities to first output pre-filled safe content, then exploits its narrative-shifting abilities to generate harmful content. Extensive black-box experiments demonstrate our method can improve the attack success rate by 47% on the widely recognized secure model (Llama2) compared to existing approaches. Our code can be found at: https://github.com/Yummy416/AdaPPA.
Lijia Lv, Weigang Zhang, Xuehai Tang, Jie Wen 0007, Feng Liu 0001, Jizhong Han, Songlin Hu 0001
ICASSP2
2025 Masked Temporal Interpolation Diffusion for Procedure Planning in Instructional Videos
abstract
In this paper, we address the challenge of procedure planning in instructional videos, aiming to generate coherent and task-aligned action sequences from start and end visual observations. Previous work has mainly relied on text-level supervision to bridge the gap between observed states and unobserved actions, but it struggles with capturing intricate temporal relationships among actions. Building on these efforts, we propose the Masked Temporal Interpolation Diffusion (MTID) model that introduces a latent space temporal interpolation module within the diffusion model. This module leverages a learnable interpolation matrix to generate intermediate latent features, thereby augmenting visual supervision with richer mid-state details. By integrating this enriched supervision into the model, we enable end-to-end training tailored to task-specific requirements, significantly enhancing the model's capacity to predict temporally coherent action sequences. Additionally, we introduce an action-aware mask projection mechanism to restrict the action generation space, combined with a task-adaptive masked proximity loss to prioritize more accurate reasoning results close to the given start and end states over those in intermediate steps. Simultaneously, it filters out task-irrelevant action predictions, leading to contextually aware action sequences. Experimental results across three widely used benchmark datasets demonstrate that our MTID achieves promising action planning performance on most metrics.
Zhaobo Qi, Lingshuai Lin, Junqi Jing, Tingting Chai, Beichen Zhang 0006, Shuhui Wang, Weigang Zhang
ICLR8
2025 Combatting Data Imbalance and Noise in Micro-Action Recognition
Weidong Chen 0010, Zhaobo Qi, Pengqi Huang, Xinyan Liu 0008, Weigang Zhang
ACM Multimedia8
2025 Dual-guided multi-modal bias removal strategy for temporal sentence grounding in video
Xiaowen Ruan, Zhaobo Qi, Yuanrong Xu, Weigang Zhang
Multim. Syst.4
2025 KN-VLM: KNowledge-guided Vision-and-Language Model for visual abductive reasoning
Kuo Tan, Zhaobo Qi, Jianping Zhong, Yuanrong Xu, Weigang Zhang
Multim. Syst.5
2025 Uncertainty-Aware Mixture of Experts for Video Action Anticipation
abstract
Anticipating future actions in daily life videos is crucial for seamless human-machine collaboration. However, accurately predicting these actions is challenging due to the inherent uncertainty and non-determinism of future events. To address this, we propose the uncertainty-aware mixture-of-experts framework for action anticipation (AntMoE), which employs multiple anticipation experts to model diverse video evolution patterns through learnable expert embeddings. These anticipation experts generate diverse predictions by integrating the top-k semantically similar observed video frames related to the current predicted feature representation, along with their corresponding expert embeddings. An anticipation router then aggregates these predictions based on the relationship between the current feature representation and all expert embeddings. To enhance the effectiveness of AntMoE, we introduce an expert regularization loss with three components: orthogonal loss promotes orthogonality among expert embeddings; expert balance loss ensures equal activation of all experts during training; and stability loss encourages the generation of numerically stable aggregation weights. Additionally, we incorporate an anticipation ranking loss function that aligns the model’s confidence across varying anticipation time durations with the ground-truth ranking order, where a shorter anticipation time length corresponds to a higher confidence level. Experimental results across multiple benchmarks demonstrate that our method achieves remarkable anticipation performance.
Zhaobo Qi, Shuhui Wang, Weigang Zhang, Qingming Huang
IEEE Trans. Circuits Syst. Video Technol.3
2025 VPA: Multi-Modal Virtual Point Augmentation for 3D Object Detection
abstract
Integrating LiDAR and camera data is crucial for precise 3D object detection. Existing methods resort to augmenting virtual points from 2D image space in a random manner to complete the appearance of 3D objects with sparse points. However, these augmented virtual points have unreasonable 3D positions and representations, which brings serious negative effects on accurate detection. To this end, we introduce a general 3D object detection framework called Virtual Point Augmenting (VPA) to enrich the 3D point cloud by controllably generating virtual points with accurate depth and position information as well as domain-gap-eliminated multi-modal representations from image and point cloud spaces. VPA contains two core designs, namely Hybrid Sampling Method (HSM) and Fine-Grained Cross-modal Fusion (FGCF). HSM uses the constructed seed point distribution map based on the edge score and mask score map to sample high-quality seed points, and employs a feature similarity function to sample withkneighbors’ depth to obtain more accurate depth for the seed points, thereby enhancing the quality of the virtual points’ 3D positions. FGCF fuses the multi-modal features,i.e., the semantic feature, the geometric feature from the image space, and the 3D position feature in an adaptive manner using self-attention mechanism, thereby further improving the representation of the virtual points. We apply VPA to the LiDAR-based method CenterPoint and fusion-based method Cross-modal transformer. Experimental results on the nuScenes, KITTI, and Waymo benchmarks validate the efficiency of our VPA, which achieves promising performance with 72.9% mAP and 74.8% NDS without using test-time augmentation and model ensemble techniques on the nuScenes test set. Code is available at https://github.com/jianpingZhonggit/vpa.git.
Jianping Zhong, Zhaobo Qi, Kaiwen Duan, Yuanrong Xu, Weigang Zhang, Qingming Huang
IEEE Trans. Circuits Syst. Video Technol.5
2025 Multi-Modal 3D Object Detector with Object-Guided Fusion and Hierarchical Sample Selection
abstract
Accurately detecting objects in 3D scenes is crucial for autonomous driving. Although existing voxel-based methods have achieved remarkable progress, their performance on tail objects remains unsatisfactory. We identify two core issues contributing to this phenomenon: the detectors frequently misidentify some background elements as foreground objects, and there is a misalignment between the classification score and detection quality. To tackle these challenges, we introduce an object-level – guided multi-modal 3D object detector with an object-guided feature fusion (OFF) module and a hierarchical sample selection (HSS) strategy, named OGMMDet. Specifically, OFF introduces rich image features to enhance the representation of objects while using an object distribution heatmap to suppress the background. This approach provides geometry clues for tail objects while providing category priors to filter out the background. HSS uses a local-to-global ranking approach to calculate the relative classification loss weights of all proposals. It assigns higher weights to proposals with higher IoU when optimizing classification branches. This ensures that the model focuses its optimization on these higher-quality proposals. Consequently, there is a positive correlation between the classification score and IoU. This method alleviates the misalignment between the classification score and detection quality. Extensive experiments on the KITTI and nuScenes benchmarks demonstrate the effectiveness of our OGMMDet, which achieves 45.61% and 68.96% mean average precision (mAP) on pedestrians and cyclists on the KITTI benchmark, respectively. Code is available at https://github.com/ZhongJianPing1/ogmmdet.git .
Jianping Zhong, Zhaobo Qi, Kaiwen Duan, Yuanrong Xu, Weigang Zhang, Qingming Huang
ACM Trans. Multim. Comput. Commun. Appl.5
2024 Bias-Conflict Sample Synthesis and Adversarial Removal Debias Strategy for Temporal Sentence Grounding in Video
abstract
Temporal Sentence Grounding in Video (TSGV) is troubled by dataset bias issue, which is caused by the uneven temporal distribution of the target moments for samples with similar semantic components in input videos or query texts. Existing methods resort to utilizing prior knowledge about bias to artificially break this uneven distribution, which only removes a limited amount of significant language biases. In this work, we propose the bias-conflict sample synthesis and adversarial removal debias strategy (BSSARD), which dynamically generates bias-conflict samples by explicitly leveraging potentially spurious correlations between single-modality features and the temporal position of the target moments. Through adversarial training, its bias generators continuously introduce biases and generate bias-conflict samples to deceive its grounding model. Meanwhile, the grounding model continuously eliminates the introduced biases, which requires it to model multi-modality alignment information. BSSARD will cover most kinds of coupling relationships and disrupt language and visual biases simultaneously. Extensive experiments on Charades-CD and ActivityNet-CD demonstrate the promising debiasing capability of BSSARD. Source codes are available at https://github.com/qzhb/BSSARD.
Zhaobo Qi, Yibo Yuan, Xiaowen Ruan, Shuhui Wang, Weigang Zhang, Qingming Huang
AAAI5
2024 Quartet: A Holistic Hybrid Parallel Framework for Training Large Language Models
Weigang Zhang, Biyu Zhou, Xing Wu 0002, Chaochen Gao, Xuehai Tang, Ruixuan Li 0001, Jizhong Han, Songlin Hu 0001
Euro-Par (2)1
2024 ProFetch: Accelerate Deep Recommendation System Training with Proactively Designed Data Layout and Dynamic Prefetching
Biyu Zhou, Weigang Zhang, Xuehai Tang, Ruixuan Li 0001, Songlin Hu 0001
ICONIP (5)3
2024 Improving Sequential DeepFake Detection with Local information enhancement
Longyun Dong, Yuanrong Xu, Jianping Zhong, Zhaobo Qi, Weigang Zhang
MMAsia5
2024 Two-stream bolt preload prediction network using hydraulic pressure and nut angle signals
Lingchao Xu, Yongsheng Xu 0004, Weigang Zhang
Eng. Appl. Artif. Intell.5
2024 Uncertainty-Boosted Robust Video Activity Anticipation
abstract
Video activity anticipation aims to predict what will happen in the future, embracing a broad application prospect ranging from robot vision and autonomous driving. Despite the recent progress, the data uncertainty issue, reflected as the content evolution process and dynamic correlation in event labels, has been somehow ignored. This reduces the model generalization ability and deep understanding on video content, leading to serious error accumulation and degraded performance. In this paper, we address the uncertainty learning problem and propose an uncertainty-boosted robust video activity anticipation framework, which generates uncertainty values to indicate the credibility of the anticipation results. The uncertainty value is used to derive a temperature parameter in the softmax function to modulate the predicted target activity distribution. To guarantee the distribution adjustment, we construct a reasonable target activity label representation by incorporating the activity evolution from the temporal class correlation and the semantic relationship. Moreover, we quantify the uncertainty into relative values by comparing the uncertainty among sample pairs and their temporal-lengths. This relative strategy provides a more accessible way in uncertainty modeling than quantifying the absolute uncertainty values on the whole dataset. Experiments on multiple backbones and benchmarks show our framework achieves promising performance and better robustness/interpretability.
Zhaobo Qi, Shuhui Wang, Weigang Zhang, Qingming Huang
IEEE Trans. Pattern Anal. Mach. Intell.3
2024 Collaborative Debias Strategy for Temporal Sentence Grounding in Video
abstract
Temporal sentence grounding in video has witnessed significant advancements, but suffers from substantial dataset bias, which undermines its generalization ability. Existing debias approaches primarily concentrate on well-known distribution and linguistic biases, while overlooking the relationship among different biases, limiting their debias capability. In this work, we delve into the existence of visual bias and combinatorial bias in the widely used datasets, and introduce a collaborative debias structure that can be seamlessly integrated into present methods. It encompasses four low-capacity models, a re-label module, and a main model. Each biased model deliberately leverages bias as shortcut information to accurately perform grounding, achieved by customizing the appropriate model structure and input data format to align with the bias characteristics. During the training phase, the gradient descent direction for optimizing the main model should align with the negative gradient descent direction of the biased model that is optimized by utilizing ground truth labels. Subsequently, the re-label module introduces a gradient aggregation function, consolidating the gradient descent direction from these biased models and constructing new labels to compel the main model to effectively capture multi-modality alignment features instead of relying on shortcut contents for grounding. Finally, we design two debias structures, P-Debias and C-Debias, to exploit the independence and inclusion relationships between different types of biases. Extensive experiments on multiple span-based models over Charades-CD and ActivityNet-CD demonstrate the exceptional debias capability of our strategy (https://github.com/qzhb/CDS).
Zhaobo Qi, Yibo Yuan, Xiaowen Ruan, Shuhui Wang, Weigang Zhang, Qingming Huang
IEEE Trans. Circuits Syst. Video Technol.5
2024 Progressive Multi-Resolution Loss for Crowd Counting
abstract
Crowd counting is usually handled in a density map regression fashion, which is supervised via an L2 loss between the predicted density map and ground truth. To effectively regulate models, various improved L2 loss functions have been developed to find a better correspondence between predicted density and annotation positions. In this paper, we propose to predict the density map at one resolution but measure its quality via a derived log-formed loss at multiple resolutions. Unlike existing methods that assume density maps at different resolutions are independent, our loss is obtained by modeling the likelihood function inspired by the relationship of density maps across multi-resolutions. We find that the traditional single-resolution L2 loss is a particular case of our derived log-likelihood. We mathematically prove it is superior to a single-resolution L2 loss. Without bells and whistles, the proposed loss substantially improves several baselines and performs favorably compared to state-of-the-art methods on five crowd counting datasets: NWPU-Crowd, ShanghaiTech A & B, UCF-QNRF, and JHU-Crowd++. The source code and trained models are released athttps://github.com/streamer-AP/PML_Loss.git.
Ziheng Yan, Yuankai Qi, Guorong Li, Xinyan Liu 0008, Weigang Zhang, Ming-Hsuan Yang 0001, Qingming Huang
IEEE Trans. Circuits Syst. Video Technol.5
2024 Limb-Aware Virtual Try-On Network With Progressive Clothing Warping
abstract
Image-based virtual try-on aims to transfer an in-shop clothing image to a person image. Most existing methods adopt a single global deformation to perform clothing warping directly, which lacks fine-grained modeling of in-shop clothing and leads to distorted clothing appearance. In addition, existing methods usually fail to generate limb details well because they are limited by the used clothing-agnostic person representation without referring to the limb textures of the person image. To address these problems, we propose Limb-aware Virtual Try-on Network named PL-VTON, which performs fine-grained clothing warping progressively and generates high-quality try-on results with realistic limb details. Specifically, we present Progressive Clothing Warping (PCW) that explicitly models the location and size of in-shop clothing and utilizes a two-stage alignment strategy to progressively align the in-shop clothing with the human body. Moreover, a novel gravity-aware loss that considers the fit of the person wearing clothing is adopted to better handle the clothing edges. Then, we design Person Parsing Estimator (PPE) with a non-limb target parsing map to semantically divide the person into various regions, which provides structural constraints on the human body and therefore alleviates texture bleeding between clothing and body regions. Finally, we introduce Limb-aware Texture Fusion (LTF) that focuses on generating realistic details in limb regions, where a coarse try-on result is first generated by fusing the warped clothing image with the person image, then limb textures are further fused with the coarse result under limb-aware guidance to refine limb details. Extensive experiments demonstrate that our PL-VTON outperforms the state-of-the-art methods both qualitatively and quantitatively.
Shengping Zhang, Weigang Zhang, Xiangyuan Lan, Hongxun Yao, Qingming Huang
IEEE Trans. Multim.3
2023 MixPipe: Efficient Bidirectional Pipeline Parallelism for Training Large-Scale Models
abstract
The rapid development of large-scale deep neural networks has put forward an urgent demand for the efficiency of parallel training. Recently, bidirectional pipeline parallelism has been recognized as an effective approach for improving training throughput. This paper proposes MixPipe, a novel bidirectional pipeline parallelism for efficiently training large-scale models in synchronous scenarios. Compared with previous proposals, MixPipe achieves a better balance between pipeline utilization and device utilization, which benefits from the flexible regulating for the number of micro-batches injected into the bidirectional pipelines at the beginning. MixPipe also features a mixed schedule to balance memory usage and further reduce the bubble ratio. Evaluation results show that: for Transformer based language models (i.e., Bert and GPT-2 models), MixPipe improves the training throughput by up to 2.39× over the state-of-the-art synchronous pipeline approaches.
Weigang Zhang, Biyu Zhou, Xuehai Tang, Zhaoxing Wang, Songlin Hu 0001
DAC1
2023 Semantic-Aware Dynamic Feature Selection and Fusion for Object Detection in UAV Videos
abstract
Keypoint-based detectors perform well in surveillance videos but face challenges in detecting objects in UAV videos due to missed corners and mismatches. To address this, we propose a semantic-aware module with a feature fusion sub-module and a feature selection sub-module. The feature fusion module adaptively combines low-level and high-level features, enhancing corner recall. The feature selection module determines spatial location importance, improving discriminative capabilities and reducing background interference, resulting in better precision. Experiments on the UAVDT benchmark show our method achieves competitive results. Notably, our method improves corner recall by 4.0% and reduces the mismatch rate by 2.9% compared to the baseline. Code is available at https://github.com/jianpingZhonggit/SemanticAwareModule.
Jianping Zhong, Zhaobo Qi, Weigang Zhang, Qingming Huang
MMAsia3
2023 An efficiency control strategy of dual-motor multi-gear drive algorithm
abstract
The Dual-motor multi-gear coupling powertrain (DMCP) has the potential to improve transmission system efficiency and driving comfort, but its complex structure and multiple working modes present challenges. The switching between different modes is easy to cause longitudinal biggish vehicle jerk. To address these issues,this paper introduces the Deep Deterministic Policy Gradient (DDPG) algorithm in the design of an Energy Management Strategy (EMS) that minimises total drive power consumption. And the number of working modes is divided and simplified. The process of switching dual motor and single motor to single motor is introduced in detail. The simulation results using AMESim and MATLAB show that the energy management strategy can effectively improve the economy, achieve no power interruption during mode switching, shift impact is less than 8m/s3, and output torque is remains stable.
Wei Liang 0005, Jiahong Cai, Jiahong Xiao, Yinyan Gong, Weigang Zhang
Connect. Sci.7
2023 Unsupervised Low-Light Video Enhancement With Spatial-Temporal Co-Attention Transformer
abstract
Existing low-light video enhancement methods are dominated by Convolution Neural Networks (CNNs) that are trained in a supervised manner. Due to the difficulty of collecting paired dynamic low/normal-light videos in real-world scenes, they are usually trained on synthetic, static, and uniform motion videos, which undermines their generalization to real-world scenes. Additionally, these methods typically suffer from temporal inconsistency (e.g., flickering artifacts and motion blurs) when handling large-scale motions since the local perception property of CNNs limits them to model long-range dependencies in both spatial and temporal domains. To address these problems, we propose the first unsupervised method for low-light video enhancement to our best knowledge, named LightenFormer, which models long-range intra- and inter-frame dependencies with a spatial-temporal co-attention transformer to enhance brightness while maintaining temporal consistency. Specifically, an effective but lightweight S-curve Estimation Network (SCENet) is first proposed to estimate pixel-wise S-shaped non-linear curves (S-curves) to adaptively adjust the dynamic range of an input video. Next, to model the temporal consistency of the video, we present a Spatial-Temporal Refinement Network (STRNet) to refine the enhanced video. The core module of STRNet is a novel Spatial-Temporal Co-attention Transformer (STCAT), which exploits multi-scale self- and cross-attention interactions to capture long-range correlations in both spatial and temporal domains among frames for implicit motion estimation. To achieve unsupervised training, we further propose two non-reference loss functions based on the invertibility of the S-curve and the noise independence among frames. Extensive experiments on the SDSD and LLIV-Phone datasets demonstrate that our LightenFormer outperforms state-of-the-art methods.
Xiaoqian Lv, Shengping Zhang, Chenyang Wang 0002, Weigang Zhang, Hongxun Yao, Qingming Huang
IEEE Trans. Image Process.4
2023 Automatic Shadow Generation via Exposure Fusion
abstract
Shadow generation aims to generate a plausible shadow for the inserted foreground object in a composite image. Besides the composite image and the associated mask of the inserted foreground object, existing methods also require a mask of all background objects as well as their shadows as an auxiliary input, which is laborious in practical applications. Meanwhile, most existing methods use a linear illumination transformation to darken the shadow region, which is prone to produce unrealistic shadows especially when background illumination is complex. To address these problems, this paper proposes an automatic shadow generation method, which avoids the laborious acquisition of the background object masks while harmonizing the shadow region to achieve plausible shadow effects. Specifically, to implicitly exploit background illumination to infer the shadow shape of the inserted foreground object, we first propose a Hierarchy Attention U-Net (HAU-Net) to sequentially build global interactions between the foreground object and background across spatial and channel dimensions. Since the spatial-variant property of the shadow, we formulate shadow harmonization as an exposure fusion problem and propose an Illumination-Aware Fusion Network (IFNet), which uses an improved illumination model with a double linear transformation to produce multiple under-exposure images of the shadow region. IFNet then learns pixel-wise fusion kernels that consider the local smoothness of the shadow to fuse the composite image with these under-exposure images to generate the realistic shadow of the foreground object. Extensive experiments on the DESOBA and Shadow-AR datasets demonstrate that our method achieves state-of-the-art performance for shadow generation on both the BOS and BOS-free test images.
Quanling Meng, Shengping Zhang, Zonglin Li 0004, Chenyang Wang 0002, Weigang Zhang, Qingming Huang
IEEE Trans. Multim.5
2023 Temporal Dynamic Concept Modeling Network for Explainable Video Event Recognition
abstract
Recently, with the vigorous development of deep learning and multimedia technology, intelligent urban computing has received more and more extensive attention from academia and industry. Unfortunately, most of the related technologies are black-box paradigms that lack interpretability. Among them, video event recognition is a basic technology. Event contains multiple concepts and their rich interactions, which can assist us to construct explainable event recognition methods. However, the crucial concepts needed to recognize events have various temporal existing patterns, and the relationship between events and the temporal characteristics of concepts has not been fully exploited. This brings great challenges for concept-based event categorization. To address the above issues, we introduce the temporal concept receptive field, which is the length of the temporal window size required to capture key concepts for concept-based event recognition methods. Accordingly, we introduce the temporal dynamic convolution (TDC) to model the temporal concept receptive field dynamically according to different events. Its core idea is to combine the results of multiple convolution layers with the learned coefficients from two complementary perspectives. These convolution layers contain a variety of kernel sizes, which can provide temporal concept receptive fields of different lengths. Similarly, we also propose the cross-domain temporal dynamic convolution (CrTDC) with the help of the rich relationship between different concepts. Different coefficients can help us to capture suitable temporal concept receptive field sizes and highlight crucial concepts to obtain accurate and complete concept representations for event analysis. Based on the TDC and CrTDC, we introduce the temporal dynamic concept modeling network (TDCMN) for explainable video event recognition. We evaluate TDCMN on large-scale and challenging datasets FCVID, ActivityNet, and CCV. Experimental results show that TDCMN significantly improves the event recognition performance of concept-based methods, and the explainability of our method inspires us to construct more explainable models from the perspective of the temporal concept receptive field.
Weigang Zhang, Zhaobo Qi, Shuhui Wang, Chi Su, Li Su 0003, Qingming Huang
ACM Trans. Multim. Comput. Commun. Appl.1
2022 Introduction to the Special Issue on Fine-Grained Visual Recognition and Re-Identification
abstract
introduction Share on Introduction to the Special Issue on Fine-Grained Visual Recognition and Re-Identification Authors: Shiliang Zhang Peking University Peking UniversityView Profile , Guorong Li University of Chinese Academy of Sciences University of Chinese Academy of SciencesView Profile , Weigang Zhang Harbin Institute of Technology Harbin Institute of TechnologyView Profile , Qingming Huang University of Chinese Academy of Sciences University of Chinese Academy of SciencesView Profile , Tiejun Huang Peking University Peking UniversityView Profile , Mubarak Shah University of Central Florida University of Central FloridaView Profile , Nicu Sebe University of Trento University of TrentoView Profile Authors Info & Claims ACM Transactions on Multimedia Computing, Communications, and ApplicationsVolume 18Issue 1sFebruary 2022 Article No.: 24pp 1–3https://doi.org/10.1145/3505280Online:25 January 2022Publication History 0citation169DownloadsMetricsTotal Citations0Total Downloads169Last 12 Months169Last 6 weeks22 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Shiliang Zhang, Guorong Li, Weigang Zhang, Qingming Huang, Tiejun Huang 0001, Mubarak Shah, Nicu Sebe
ACM Trans. Multim. Comput. Commun. Appl.3
2021 Exploiting sample correlation for crowd counting with multi-expert network
abstract
Crowd counting is a difficult task because of the diversity of scenes. Most of the existing crowd counting methods adopt complex structures with massive backbones to enhance the generalization ability. Unfortunately, the performance of existing methods on large-scale data sets is not satisfactory. In order to handle various scenarios with less complex network, we explored how to efficiently use the multi-expert model for crowd counting tasks. We mainly focus on how to train more efficient expert networks and how to choose the most suitable expert. Specifically, we propose a task-driven similarity metric based on sample’s mutual enhancement, referred as co-fine-tune similarity, which can find a more efficient subset of data for training the expert network. Similar samples are considered as a cluster which is used to obtain parameters of an expert. Besides, to make better use of the proposed method, we design a simple network called FPN with Deconvolution Counting Network, which is a more suitable base model for the multi-expert counting network. Experimental results show that multiple experts FDC (MFDC) achieves the best performance on four public data sets, including the large scale NWPU-Crowd data set. Furthermore, the MFDC trained on an extensive dense crowd data set can generalize well on the other data sets without extra training or fine-tuning.1
Xinyan Liu 0008, Guorong Li, Zhenjun Han, Weigang Zhang, Qingming Huang, Nicu Sebe
ICCV4
2021 DBAM: Dense Boundary and Actionness Map for Action Localization in Videos via Sentence Query
Weigang Zhang, Yushu Liu, Jianping Zhong, Guorong Li, Qingming Huang
ICIG (3)1
2021 Action Category and Phase Consistency Regularization for High-Quality Temporal Action Proposal Generation
abstract
Temporal action detection is a fundamental yet challenging task in video content analysis. The performance of existing methods still remains far from satisfactory as the mAP reduces dramatically at high tIoU threshold. With the goal of predicting the starting and ending points more precisely, this work first introduces the action category label into the temporal proposal generation stage of the training process. Specifically, with the category information, we proposed two extra constrains, i.e, action based constraint and action-class agnostic constraints. The former aims at minimizing the discrepancy inside the same action category while the latter forces the feature of the samples aggregates in the same phase. Comprehensive experiments are conducted on the THUMOS’14 benchmark. A remarkable improvement of average recall is attained especially when the number of proposals is small. And our approach achieves 29.0% mAP at a strict [email protected].
Yushu Liu, Weigang Zhang, Guorong Li, Qingming Huang
ICME2
2021 Graph Regularized Encoder-Decoder Networks for Image Representation Learning
abstract
Image representation learning with encoder-decoder networks plays a fundamental role in multimedia processing. Recent findings show that traditional encoder-decoders can be negatively affected by small visual perturbations. The learned non-smooth feature embedding cannot guarantee to capture semantic-meaningful geometric distance between visually-similar image samples. Inspired by manifold learning, we propose a graph regularized encoder-decoder network, which can preserve local geometric information of the code embedding space. More discriminative feature embedding is learnt to attain both high-level image semantic and neighbor relationship of image clusters. The proposed graph regularizer is formulated upon multi-layer perceptions. It uses the local invariance principle to explicitly reconstruct the geometric similarity graph. Theoretical analysis is provided to show the connection between our deep regularizer and traditional graph Laplacian regularizer. Practically, the network complexity is alleviated by anchor based bipartite graph, and this leverages our method into large scale scenario. Experimental evaluations show the comparable results of the proposed method with state-of-the-art models on different tasks.
Liang Li 0003, Shuhui Wang, Weigang Zhang, Qingming Huang, Qi Tian 0001
IEEE Trans. Multim.4
2020 Interpretable Visual Reasoning via Probabilistic Formulation Under Natural Supervision
Xinzhe Han, Shuhui Wang, Chi Su, Weigang Zhang, Qingming Huang, Qi Tian 0001
ECCV (9)4
2020 A Structured Latent Variable Recurrent Network With Stochastic Attention For Generating Weibo Comments
abstract
Building intelligent agents to generate realistic Weibo comments is challenging. For such realistic Weibo comments, the key criterion is improving diversity while maintaining coherency. Considering that the variability of linguistic comments arises from multi-level sources, including both discourse-level properties and word-level selections, we improve the comment diversity by leveraging such inherent hierarchy. In this paper, we propose a structured latent variable recurrent network, which exploits the hierarchical-structured latent variables with stochastic attention to model the variations of comments. First, we endow both discourse-level and word-level latent variables with hierarchical and temporal dependencies for constructing multi-level hierarchy. Second, we introduce a stochastic attention to infer the key-words of interest in the input post. As a result, diverse comments can be generated with both discourse-level properties and local-word selections. Experiments on open-domain Weibo data show that our model generates more diverse and realistic comments.
Liang Li 0003, Shuhui Wang, Weigang Zhang, Qingming Huang, Qi Tian 0001
IJCAI4
2020 Modeling Temporal Concept Receptive Field Dynamically for Untrimmed Video Analysis
abstract
Event analysis in untrimmed videos has attracted increasing attention due to the application of cutting-edge techniques such as CNN. As a well studied property for CNN-based models, the receptive field is a measurement for measuring the spatial range covered by a single feature response, which is crucial in improving the image categorization accuracy. In video domain, video event semantics are actually described by complex interaction among different concepts, while their behaviors vary drastically from one video to another, leading to the difficulty in concept-based analytics for accurate event categorization. To model the concept behavior, we study temporal concept receptive field of concept-based event representation, which encodes the temporal occurrence pattern of different mid-level concepts. Accordingly, we introduce temporal dynamic convolution (TDC) to give stronger flexibility to concept-based event analytics. TDC can adjust the temporal concept receptive field size dynamically according to different inputs. Notably, a set of coefficients are learned to fuse the results of multiple convolutions with different kernel widths that provide various temporal concept receptive field sizes. Different coefficients can generate appropriate and accurate temporal concept receptive field size according to input videos and highlight crucial concepts. Based on TDC, we propose the temporal dynamic concept modeling network~(TDCMN) to learn an accurate and complete concept representation for efficient untrimmed video analysis. Experiment results on FCVID and ActivityNet show that TDCMN demonstrates adaptive event recognition ability conditioned on different inputs, and improve the event recognition performance of Concept-based methods by a large margin. Code is available at https://github.com/qzhb/TDCMN.
Zhaobo Qi, Shuhui Wang, Chi Su, Li Su 0003, Weigang Zhang, Qingming Huang
ACM Multimedia5
2020 Fixation guided network for salient object detection
abstract
Convolutional neural network (CNN) based salient object detection (SOD) has achieved great development in recent years. However, in some challenging cases, i.e. small-scale salient object, low contrast salient object and cluttered background, existing salient object detect methods are still not satisfying. In order to accurately detect salient objects, SOD networks need to fix the position of most salient part. Fixation prediction (FP) focuses on the most visual attractive regions, so we think it could assist in locating salient objects. As far as we know, there are few methods jointly consider SOD and FP tasks. In this paper, we propose a fixation guided salient object detection network (FGNet) to leverage the correlation between SOD and FP. FGNet consists of two branches to deal with fixation prediction and salient object detection respectively. Further, an effective feature cooperation module (FCM) is proposed to fuse complementary information between the two branches. Extensive experiments on four popular datasets and comparisons with twelve state-of-the-art methods show that the proposed FGNet well captures the main context of images and locates salient objects more accurately.
Li Su 0003, Weigang Zhang, Qingming Huang
MMAsia3
2020 Video Anomaly Detection Using Open Data Filter and Domain Adaptation
abstract
Video anomaly detection is a very challenging task because of the rarity, openness, and the definition of the anomalies. Researchers pay more attention to the characteristics of anomalies and have proposed a variety of anomaly detection models. However, most existing methods only use normal events to construct anomaly detection models and ignore the diversity and openness of normal events. Actually, because real-world video data often have an open-ended distribution, some normal patterns hardly ever appeared in the training data. In addition, analogous to human experience in identifying anomalies, rare abnormal events can play a certain role in the detection of similar abnormal events in the dataset. Therefore, assuming that a small number of abnormal events are known, we propose a novel supervised anomaly detection model which explicitly detects open normal events and open abnormal events in the dataset and treats open data and seen data with different classifiers. First, we use the training video to train an imbalanced classifier as the seen data classifier. Then, during the testing phase, an open data filter module isused to divide the test data into seen data and open data. Finally, we directly use the seen data classifier to generate anomaly scores for the seen test data. For the open test data, we adopt a domain adaptation method to reduce the distribution difference between it and the training data and train a new classifier to score for it. Extensive experimental results prove the effectiveness of our model.
Chen Zhang 0013, Guorong Li, Li Su 0003, Weigang Zhang, Qingming Huang
VCIP4
2020 Two-stream deep sparse network for accurate and efficient image restoration
Shuhui Wang, Liang Li 0003, Weigang Zhang, Qingming Huang
Comput. Vis. Image Underst.4
2020 The Unmanned Aerial Vehicle Benchmark: Object Detection, Tracking and Baseline
Hongyang Yu 0001, Guorong Li, Weigang Zhang, Qingming Huang, Dawei Du, Qi Tian 0001, Nicu Sebe
Int. J. Comput. Vis.3
2020 Action Recognition Using Form and Motion Modalities
abstract
Action recognition has attracted increasing interest in computer vision due to its potential applications in many vision systems. One of the main challenges in action recognition is to extract powerful features from videos. Most existing approaches exploit either hand-crafted techniques or learning-based methods to extract features from videos. However, these methods mainly focus on extracting the dynamic motion features, which ignore the static form features. Therefore, these methods cannot fully capture the underlying information in videos accurately. In this article, we propose a novel feature representation method for action recognition, which exploits hierarchical sparse coding to learn the underlying features from videos. The learned features characterize the form and motion simultaneously and therefore provide more accurate and complete feature representation. The learned form and motion features are considered as two modalities, which are used to represent both the static and motion features. These modalities are further encoded into a global representation via a pairwise dictionary learning and then fed to an SVM classifier for action classification. Experimental results on several challenging datasets validate that the proposed method is superior to several state-of-the-art methods.
Quanling Meng, Heyan Zhu, Weigang Zhang, Xuefeng Piao, Aijie Zhang
ACM Trans. Multim. Comput. Commun. Appl.3
2019 Learning Attribute-Specific Representations for Visual Tracking
abstract
In recent years, convolutional neural networks (CNNs) have achieved great success in visual tracking. Most of existing methods train or fine-tune a binary classifier to distinguish the target from its background. However, they may suffer from the performance degradation due to insufficient training data. In this paper, we show that attribute information (e.g., illumination changes, occlusion and motion) in the context facilitates training an effective classifier for visual tracking. In particular, we design an attribute-based CNN with multiple branches, where each branch is responsible for classifying the target under a specific attribute. Such a design reduces the appearance diversity of the target under each attribute and thus requires less data to train the model. We combine all attributespecific features via ensemble layers to obtain more discriminative representations for the final target/background classification. The proposed method achieves favorable performance on the OTB100 dataset compared to state-of-the-art tracking methods. After being trained on the VOT datasets, the proposed network also shows a good generalization ability on the UAV-Traffic dataset, which has significantly different attributes and target appearances with the VOT datasets.
Yuankai Qi, Shengping Zhang, Weigang Zhang, Li Su 0003, Qingming Huang, Ming-Hsuan Yang 0001
AAAI3
2019 Multi-Label Image Classification with Attention Mechanism and Graph Convolutional Networks
abstract
The task of multi-label image classification is to predict a set of proper labels for an input image. To this end, it is necessary to strengthen the association between the labels and the image regions, and utilize the relationship between the labels. In this paper, we propose a novel framework for multi-label image classification, which uses attention mechanism and Graph Convolutional Network (GCN) simultaneously. The attention mechanism can focus on specific target regions while ignoring other useless information around, thereby enhancing the association of the labels with the image regions. By constructing a directed graph over the labels, GCN can learn the relationship between the labels from a global perspective and map this label graph to a set of inter-dependent object classifiers. The framework first uses ResNet to extract features while using attention mechanism to generate attention maps for all labels and obtain weighted features. GCN uses weighted fusion features from the output of the resnet and attention mechanism to achieve classification. Experimental results show that both the attention mechanism and GCN can effectively improve the classification performance, and the proposed framework is competitive with the state-of-the-art methods.
Quanling Meng, Weigang Zhang
MMAsia2
2019 Self-balance Motion and Appearance Model for Multi-object Tracking in UAV
abstract
Under the tracking-by-detection framework, multi-object tracking methods try to connect object detections with target trajectories by reasonable policy. Most methods represent objects by the appearance and motion. The inference of the association is mostly judged by a fusion of appearance similarity and motion consistency. However, the fusion ratio between appearance and motion are often determined by subjective setting. In this paper, we propose a novel self-balance method fusing appearance similarity and motion consistency. Extensive experimental results on public benchmarks demonstrate the effectiveness of the proposed method with comparisons to several state-of-the-art trackers.
Hongyang Yu 0001, Guorong Li, Weigang Zhang, Hongxun Yao, Qingming Huang
MMAsia3
2019 Improving multi-label classification with missing labels by learning label-specific features
Jun Huang 0003, Zekai Cheng, Zhixiang Yuan, Weigang Zhang, Qingming Huang
Inf. Sci.6
2019 Beyond global fusion: A group-aware fusion approach for multi-view image clustering
Zhe Xue, Guorong Li, Shuhui Wang, Jun Huang 0003, Weigang Zhang, Qingming Huang
Inf. Sci.5
2019 Split Multiplicative Multi-View Subspace Clustering
abstract
Various subspace clustering methods have been successively developed to process multi-view datasets. Most of the existing methods try to obtain a consensus structure coefficient matrix based on view-specific subspace recoveries. However, since view-specific structures contain individualized components that are intrinsically different from the consensus structure, directly adopting view-specific subspace structures might not be a reasonable choice. With this concern in mind, our goal in this paper is to seek novel strategies to extract valuable components from view-specific structures that are consistent with the consensus subspace structure. To this end, we propose a novel multi-view subspace clustering method named Split Multiplicative Multi-view Subspace Clustering (SM2SC) with the joint strength of a multiplicative decomposition scheme and a variable splitting scheme. Specifically, the multiplicative decomposition scheme effectively guarantees the structural consistency of the extracted components. Then the variable splitting scheme takes a step further via extracting the structural consistent components from view-specific structures. Furthermore, an alternating optimization algorithm is proposed to optimize the resulting optimization problem, which is non-convex and constrained. We prove that this algorithm could converge to a critical point. Finally, we provide empirical studies on real-world datasets that speak to the practical efficacy of our proposed method. The source code is released on GitHub.
Zhiyong Yang 0001, Qianqian Xu 0001, Weigang Zhang, Xiaochun Cao, Qingming Huang
IEEE Trans. Image Process.3
2019 SkeletonNet: A Hybrid Network With a Skeleton-Embedding Process for Multi-View Image Representation Learning
abstract
Multi-view representation learning plays a fundamental role in multimedia data analysis. Some specific inter-view alignment principles are adopted in conventional models, where there is an assumption that different views share a common latent subspace. However, when dealing views on diverse semantic levels, the view-specific characteristics are neglected, and the divergent inconsistency of similarity measurements hinders sufficient information sharing. This paper proposes a hybrid deep network by introducing tensor factorization into the multi-view deep auto-encoder. The network adopts skeleton-embedding process for unsupervised multi-view subspace learning. It takes full consideration of view-specific characteristics, and leverages the strength of both shallow and deep architectures for modeling low- and high-level views, respectively. We first formulate the high-level-view semantic distribution as the underlying skeleton structure of the learned subspace, and then infer the local tangent structures according to the affinity propagation of low-level-view geometric correlations. As a consequence, more discriminative subspace representation can be learned from global semantic pivots to local geometric details. Experimental comparisons on three benchmark image datasets show the promising performance and flexibility of our model.
Liang Li 0003, Shuhui Wang, Weigang Zhang, Qingming Huang, Qi Tian 0001
IEEE Trans. Multim.4
2018 Reverse Densely Connected Feature Pyramid Network for Object Detection
Yongjian Xin, Shuhui Wang, Liang Li 0003, Weigang Zhang, Qingming Huang
ACCV (5)4
2018 Less Is More: Picking Informative Frames for Video Captioning
Shuhui Wang, Weigang Zhang, Qingming Huang
ECCV (13)3
2018 The Unmanned Aerial Vehicle Benchmark: Object Detection and Tracking
Dawei Du, Yuankai Qi, Hongyang Yu 0001, Kaiwen Duan, Guorong Li, Weigang Zhang, Qingming Huang, Qi Tian 0001
ECCV (10)7
2018 Semantic Manifold Alignment in Visual Feature Space for Zero-Shot Learning
abstract
Zero-Shot Learning (ZSL) is getting more attention for its potential to solve a task without training examples, such as to recognize a category of unseen object in computer vision task. Most existing methods are suffered from hubness problem and semantic gap problem. In this paper, we propose a novel strategy based on Aligning Semantic Manifolds in Feature Space (ASMFS) to boost the performance of ZSL. Considering that the semantic representations must be predicted in the location of their corresponding visual instances, we adjust the predicted unseen semantic representations by the average of their K nearest neighbors (K-NN). The experimental results over two basic ZSL models and four public datasets demonstrate the universal enhancement performance of the proposed strategy. It significantly boosts the existing ZSL approaches with low over cost and outperforms eight state-of-the-art methods.
Changsu Liao, Li Su 0003, Weigang Zhang, Qingming Huang
ICME3
2018 Edge Guided Generation Network for Video Prediction
abstract
Video prediction is a challenging problem due to the highly complex variation of video appearance and motions. Traditional methods that directly predict pixel values often result in blurring and artifacts. Furthermore, cumulative errors can lead to a sharp drop of prediction quality in long-term prediction. To alleviate the above problems, we propose a novel edge guided video prediction network, which firstly models the dynamic of frame edges and predicts the future frame edges, then generates the future frames under the guidance of the obtained future frame edges. Specifically, our network consists of two modules that are ConvLSTM based edge prediction module and the edge guided frames generation module. The whole network is differentiable and can be trained end-to-end without any supervision effort. Extensive experiments on KTH human action dataset and challenging autonomous driving KITTI dataset demonstrate that our method achieves better results than state-of-the-art methods especially in long-term video predictions.
Kai Xu 0013, Guorong Li, Huijuan Xu 0001, Weigang Zhang, Qingming Huang
ICME4
2018 Bilevel Multiview Latent Space Learning
abstract
Different kinds of features describe different aspects of image data, and each feature can be treated as a view when we take it as a particular understanding of images. Leveraging multiple views provides a richer and comprehensive description than using only a single view. However, multiview data are often represented by high-dimensional heterogeneous features, so it is meaningful to find a low-dimensional consensus representation from multiple views. In this paper, we propose an unsupervised multiview dimensionality reduction method for images based on bilevel latent space learning. As different views have different physical meanings and statistical properties, they are not directly comparable. Therefore, we learn the comparable representation for each view in the first level. The shared and the private nature of multiview data are exploited to accurately preserve the information of each view. Then, we fuse different views into a low-dimensional representation by conducting joint matrix factorization in the second level. To guarantee the low-dimensional representation to be compact and discriminative, the intrinsic geometric structure of data is utilized. Besides, our method considers resisting the outliers and noise contained in multiview data, which may influence the learned representation and deteriorate its semantic consistency. We design appropriate optimization objectives to learn the latent spaces in different levels. Compared with the existing methods, our method could provide a more flexible multiview learning strategy that not only accurately captures the information of each view but also is robust to outliers and noise, which can obtain a more discriminative and compact low-dimensional representation. Experiments on two real-world image data sets demonstrate the advantages of our method over the existing multiview dimensionality reduction methods.
Zhe Xue, Guorong Li, Shuhui Wang, Weigang Zhang, Qingming Huang
IEEE Trans. Circuits Syst. Video Technol.4
2017 A Graph Regularized Deep Neural Network for Unsupervised Image Representation Learning
abstract
Deep Auto-Encoder (DAE) has shown its promising power in high-level representation learning. From the perspective of manifold learning, we propose a graph regularized deep neural network (GR-DNN) to endue traditional DAEs with the ability of retaining local geometric structure. A deep-structured regularizer is formulated upon multi-layer perceptions to capture this structure. The robust and discriminative embedding space is learned to simultaneously preserve the high-level semantics and the geometric structure within local manifold tangent space. Theoretical analysis presents the close relationship between the proposed graph regularizer and the graph Laplacian regularizer in terms of the optimization objective. We also alleviate the growth of the network complexity by introducing the anchor-based bipartite graph, which guarantees the good scalability for large scale data. The experiments on four datasets show the comparable results of the proposed GR-DNN with the state-of-the-art methods.
Liang Li 0003, Shuhui Wang, Weigang Zhang, Qingming Huang
CVPR4
2017 Online low-rank similarity function learning with adaptive relative margin for cross-modal retrieval
abstract
This paper presents a Cross-Modal Online Low-Rank Similarity function learning method (CMOLRS) for cross-modal retrieval, which learns a low-rank bilinear similarity measure on data from different modalities. CMOLRS models the cross-modal relations by relative similarities on a set of training data triplets and formulates the relative relations as convex hinge loss functions. By adapting the margin of hinge loss using information from feature space and label space for each triplet, CMOLRS effectively captures the multi-level semantic correlation among cross-modal data. The similarity function is learned by online learning in the manifold of low-rank matrices, thus good scalability is gained when processing large scale datasets. Extensive experiments are conducted on three public datasets. Comparisons with the state-of-the-art methods show the effectiveness and efficiency of our approach.
Yiling Wu, Shuhui Wang, Weigang Zhang, Qingming Huang
ICME3
2017 Dependency Exploitation: A Unified CNN-RNN Approach for Visual Emotion Recognition
abstract
Visual emotion recognition aims to associate images with appropriate emotions. There are different visual stimuli that can affect human emotion from low-level to high-level, such as color, texture, part, object, etc. However, most existing methods treat different levels of features as independent entity without having effective method for feature fusion. In this paper, we propose a unified CNN-RNN model to predict the emotion based on the fused features from different levels by exploiting the dependency among them. Our proposed architecture leverages convolutional neural network (CNN) with multiple layers to extract different levels of features with in a multi-task learning framework, in which two related loss functions are introduced to learn the feature representation. Considering the dependencies within the low-level and high-level features, a new bidirectional recurrent neural network (RNN) is proposed to integrate the learned features from different layers in the CNN model. Extensive experiments on both Internet images and art photo datasets demonstrate that our method outperforms the state-of-the-art methods with at least 7% performance improvement.
Xinge Zhu, Liang Li 0003, Weigang Zhang, Tianrong Rao, Min Xu 0001, Qingming Huang, Dong Xu 0001
IJCAI3
2017 Deep Unsupervised Convolutional Domain Adaptation
abstract
In multimedia analysis, the task of domain adaptation is to adapt the feature representation learned in the source domain with rich label information to the target domain with less or even no label information. Significant research endeavors have been devoted to aligning the feature distributions between the source and the target domains in the top fully connected layers based on unsupervised DNN-based models. However, the domain adaptation has been arbitrarily constrained near the output ends of the DNN models, which thus brings about inadequate knowledge transfer in DNN-based domain adaptation process, especially near the input end. We develop an attention transfer process for convolutional domain adaptation. The domain discrepancy, measured in correlation alignment loss, is minimized on the second-order correlation statistics of the attention maps for both source and target domains. Then we propose Deep Unsupervised Convolutional Domain Adaptation DUCDA method, which jointly minimizes the supervised classification loss of labeled source data and the unsupervised correlation alignment loss measured on both convolutional layers and fully connected layers. The multi-layer domain adaptation process collaborately reinforces each individual domain adaptation component, and significantly enhances the generalization ability of the CNN models. Extensive cross-domain object classification experiments show DUCDA outperforms other state-of-the-art approaches, and validate the promising power of DUCDA towards large scale real world application.
Junbao Zhuo, Shuhui Wang, Weigang Zhang, Qingming Huang
ACM Multimedia3
2017 Rotative maximal pattern: A local coloring descriptor for object classification and recognition
Junbiao Pang, Weigang Zhang, Laiyun Qing, Qingming Huang
Inf. Sci.4
2017 Justify role of Similarity Diffusion Process in cross-media topic ranking: an empirical evaluation
Junbiao Pang, Weigang Zhang, Qingming Huang
Multim. Tools Appl.3
2016 Accelerate convolutional neural networks for binary classification via cascading cost-sensitive feature
abstract
Convolutional Neural Networks (CNNs) have delivered impressive state-of-the-art performances for many vision tasks, while the computation costs of these networks during test-time are notorious. Empirical results have discovered that CNNs have learned the redundant representations both within and across different layers. When CNNs are applied for binary classification, we investigate a method to exploit this redundancy across layers, and construct a cascade of classifiers which explicitly balances classification accuracy and hierarchical feature extraction costs. Our method cost-sensitively selects feature points across several layers from trained networks and embeds non-expensive yet discriminative features into a cascade. Experiments on binary classification demonstrate that our framework leads to drastic test-time improvements, e.g., possible 47.2x speedup for TRECVID upper body detection, 2.82x speedup for Pascal VOC2007 People detection, 3.72x for INRIA Person detection with less than 0.5% drop in accuracies of the original networks.
Junbiao Pang, Huihuang Lin, Li Su 0003, Chunjie Zhang 0001, Weigang Zhang, Lijuan Duan, Qingming Huang
ICIP5
2016 Robust latent poisson deconvolution from multiple imperfect features for web topic detection
abstract
In web topic detection, detecting “hot” topics from enormous User-Generated Content (UGC) on web data poses two main difficulties that conventional approaches can barely handle: 1) poor feature representations from noisy images and short texts; and 2) uncertain roles of modalities where visual content is either highly or weakly relevant to textual cues due to less-constrained data. In this paper, following the detection by ranking approach, we address the problem by learning a robust shared representation from multiple, noisy and complementary features, and integrating both textual and visual graphs into a k-Nearest Neighbor Similarity Graph (k-N2SG). Then Non-negative Matrix Factorization using Random walk (NMFR) is introduced to generate topic candidates. An efficient fusion of multiple graphs is then done by a Latent Poisson Deconvolution (LPD) which consists of a poisson deconvolution with sparse basis similarities for each edge. Experiments show significantly improved accuracy of the proposed approach in comparison with the state-of-the-art methods on two public data sets.
Junbiao Pang, Chunjie Zhang 0001, Liang Li 0003, Li Su 0003, Weigang Zhang, Qingming Huang, Guiping Su
ICME6
2016 Beyond appearance model: Learning appearance variations for object tracking
Guorong Li, Bingpeng Ma, Jun Huang 0003, Qingming Huang, Weigang Zhang
Neurocomputing5
2016 Online web video topic detection and tracking with semi-supervised learning
Guorong Li, Shuqiang Jiang, Weigang Zhang, Junbiao Pang, Qingming Huang
Multim. Syst.3
2016 Effective Multimodality Fusion Framework for Cross-Media Topic Detection
abstract
Due to the prevalence of We-Media, information is quickly published and received in various forms anywhere and anytime through the Internet. The rich cross-media information carried by the multimodal data in multiple media has a wide audience, deeply reflects the social realities, and brings about much greater social impact than any single media information. Therefore, automatically detecting topics from cross media is of great benefit for the organizations (i.e., advertising agencies and governments) that care about the social opinions. However, cross-media topic detection is challenging from the following aspects: 1) the multimodal data from different media often involve distinct characteristics and 2) topics are presented in an arbitrary manner among the noisy web data. In this paper, we propose a multimodality fusion framework and a topic recovery (TR) approach to effectively detect topics from cross-media data. The multimodality fusion framework flexibly incorporates the heterogeneous multimodal data into a multimodality graph, which takes full advantage from the rich cross-media information to effectively detect topic candidates (T.C.). The TR approach solidly improves the entirety and purity of detected topics by: 1) merging the T.C. that are highly relevant themes of the same real topic and 2) filtering out the less-relevant noise data in the merged T.C. Extensive experiments on both single-media and cross-media data sets demonstrate the promising flexibility and effectiveness of our method in detecting topics from cross media.
Lingyang Chu, Guorong Li, Shuhui Wang, Weigang Zhang, Qingming Huang
IEEE Trans. Circuits Syst. Video Technol.5
2016 Robust Latent Poisson Deconvolution From Multiple Features for Web Topic Detection
abstract
Detecting “hot” topics from the enormous usergenerated content (UGC) data on web poses two main difficulties that the conventional approaches can barely handle:1) poor feature representations from noisy images or short texts, and 2) uncertain roles of modalities where the visual content is either highly or weakly relevant to the textual cues due to the less-constrained UGC. In this paper, following the detection-by-ranking approach, we address above challenges by learning a robust latent representation from multiple, noisy and a high probability of the complementary features. Both the textual features and the visual ones are encoded into a k-nearest neighbor hybrid similarity graph (HSG), where nonnegative matrix factorization using random walk is introduced to generate topic candidates. An efficient fusion of multiple HSGs is then done by a latent poisson deconvolution, which consists of a poisson deconvolution with sparse basis similarity for each edge. Experiments show significantly improved accuracy of the proposed approach in comparison with the state-of-the-art methods on two public datasets.
Junbiao Pang, Chunjie Zhang 0001, Weigang Zhang, Qingming Huang
IEEE Trans. Multim.4
2015 Group sensitive Classifier Chains for multi-label classification
abstract
In multi-label classification, labels often have correlations with each other. Exploiting label correlations can improve the performances of classifiers. Current multi-label classification methods mainly consider the global label correlations. However, the label correlations may be different over different data groups. In this paper, we propose a simple and efficient framework for multi-label classification, called Group sensitive Classifier Chains. We assume that similar examples not only share the same label correlations, but also tend to have similar labels. We augment the original feature space with label space and cluster them into groups, then learn the label dependency graph in each group respectively and build the classifier chains on each group specific label dependency graph. The group specific classifier chains which are built on the nearest group of the test example are used for prediction. Comparison results with the state-of-the-art approaches manifest competitive performances of our method.
Jun Huang 0003, Guorong Li, Shuhui Wang, Weigang Zhang, Qingming Huang
ICME4
2015 GOMES: A group-aware multi-view fusion approach towards real-world image clustering
abstract
Different features describe different views of visual appearance, multi-view based methods can integrate the information contained in each view and improve the image clustering performance. Most of the existing methods assume that the importance of one type of feature is the same to all the data. However, the visual appearance of images are different, so the description abilities of different features vary with different images. To solve this problem, we propose a group-aware multi-view fusion approach. Images are partitioned into groups which consist of several images sharing similar visual appearance. We assign different weights to evaluate the pairwise similarity between different groups. Then the clustering results and the fusion weights are learned by an iterative optimization procedure. Experimental results indicate that our approach achieves promising clustering performance compared with the existing methods.
Zhe Xue, Guorong Li, Shuhui Wang, Chunjie Zhang 0001, Weigang Zhang, Qingming Huang
ICME5
2015 Online learning affinity measure with CovBoost for multi-target tracking
Guorong Li, Qingming Huang, Shuqiang Jiang, Yingkun Xu, Weigang Zhang
Neurocomputing5
2015 Online dictionary learning for Local Coordinate Coding with Locality Coding Adaptors
Junbiao Pang, Chunjie Zhang 0001, Weigang Zhang, Laiyun Qing, Qingming Huang
Neurocomputing4
2015 Fusing cross-media for topic detection by dense keyword groups
Weigang Zhang, Tianlong Chen 0003, Guorong Li, Junbiao Pang, Qingming Huang, Wen Gao 0001
Neurocomputing1
2015 Local Laplacian Coding From Theoretical Analysis of Local Coding Schemes for Locally Linear Classification
abstract
Local coordinate coding (LCC) is a framework to approximate a Lipschitz smooth function by combining linear functions into a nonlinear one. For locally linear classification, LCC requires a coding scheme that heavily determines the nonlinear approximation ability, posing two main challenges: 1) the locality making faraway anchors have smaller influences on current data and 2) the flexibility balancing well between the reconstruction of current data and the locality. In this paper, we address the problem from the theoretical analysis of the simplest local coding schemes, i.e., local Gaussian coding and local student coding, and propose local Laplacian coding (LPC) to achieve the locality and the flexibility. We apply LPC into locally linear classifiers to solve diverse classification tasks. The comparable or exceeded performances of state-of-the-art methods demonstrate the effectiveness of the proposed method.
Junbiao Pang, Chunjie Zhang 0001, Weigang Zhang, Qingming Huang
IEEE Trans. Cybern.4
2015 Unsupervised Web Topic Detection Using A Ranked Clustering-Like Pattern Across Similarity Cascades
abstract
Despite the massive growth of social media on the Internet, the process of organizing, understanding, and monitoring user generated content (UGC) has become one of the most pressing problems in today's society. Discovering topics on the web from a huge volume of UGC is one of the promising approaches to achieve this goal. Compared with classical topic detection and tracking in news articles, identifying topics on the web is by no means easy due to the noisy, sparse, and less- constrained data on the Internet. In this paper, we investigate methods from the perspective of similarity diffusion, and propose a clustering-like pattern across similarity cascades (SCs). SCs are a series of subgraphs generated by truncating a similarity graph with a set of thresholds, and then maximal cliques are used to capture topics. Finally, a topic-restricted similarity diffusion process is proposed to efficiently identify real topics from a large number of candidates. Experiments demonstrate that our approach outperforms the state-of-the-art methods on three public data sets.
Junbiao Pang, Fei Jia, Chunjie Zhang 0001, Weigang Zhang, Qingming Huang
IEEE Trans. Multim.4
2014 DA-CCD: A novel action representation by Deep Architecture of local depth feature
abstract
With the widespread use of depth sensors, it is crucial to provide an effective and efficient solution for human action analysis applications upon the informative depth data. In this paper, we present a generic framework of modeling the human action by deep architecture enhanced local features with depth data. To introduce robust higher-level representations, we augment the adaptive and scalable local depth features in a deep feature learning manner. Specifically, a Deep Architecture of Comparative Coding Descriptor (DA-CCD) is proposed to represent the depth action data. Our approach obtains consistently superior recognition precisions on view specific/non-specific scenarios compared with other leading action representations of depth data on the Huawei/3DLife Dataset.
Zhongwei Cheng, Yanhao Zhang 0001, Weigang Zhang, Qingming Huang
ICIP5
2014 Weakly supervised cross-view action recognition via sequential motion accumulation
abstract
In real application scenarios, the visual observations of the same type of action vary significantly from one view to another. This paper addresses the action recognition problem under the view changes, especially when no labels are available in the target view. A novel feature, called Sequential Motion Accumulation (SMA), is proposed to characterize actions. The SMA descriptor depicts the temporal structure of motion property to explore the distinguishing action characteristics and their invariances across views. Moreover, we propose a weakly supervised categorization approach to generate target-view categorical prior for learning a cross-view metric, which can further improve the recognition accuracy of the SMA descriptor. Our method is verified on the multiview IXMAS dataset, and it achieves superior performance compared with the state-of-the-art methods.
Zhongwei Cheng, Yanhao Zhang 0001, Weigang Zhang, Qingming Huang
ICIP5
2014 Web topic detection using a ranked clustering-like pattern across similarity cascades
abstract
In multi-media and social media communities, web topic detection poses two main difficulties that conventional approaches can barely handle: 1) there are large inter-topic variations among web topics; 2) supervised information is rare to identify the real topics. In this paper, we address these problems from the similarity diffusion perspective among objects on web, and present a clustering-like pattern across similarity cascades (SCs). SCs are a series of subgraphs generated by truncating a weighted graph with a set of thresholds, and then maximal cliques are used to describe the topic candidates. Poisson deconvolution is adopted to efficiently identify the real topics from these topic candidates. Experiments demonstrate that our approach outperforms the state-of-the-arts on two datasets. In addition, we report accuracy v.s. false positives per topic (FPPT) curves for performance evaluation. To our knowledge, this is the first complete evaluation of web topic detection at the topic-wise level, and it establishes a new benchmark for this problem.
Fei Jia, Junbiao Pang, Weigang Zhang, Guorong Li, Chunjie Zhang 0001, Qingming Huang, Yugui Liu
ICME3
2014 Web video thumbnail recommendation with content-aware analysis and query-sensitive matching
Weigang Zhang, Chunxi Liu, Zhenjun Wang, Guorong Li, Qingming Huang, Wen Gao 0001
Multim. Tools Appl.1
2013 Discriminative Spatial Codebook Generation for Image Classification
abstract
Codebook plays an important role in the bag-of-visual-words (BoW) model for image classification. However, the traditional codebook generation procedure ignores the spatial information. Although a lot of works have been done to consider the spatial information for codebook generation, most of them rely on fixed region selection or partition of images, hence are not able to cope with the variations of images. To solve this problem, in this paper, we propose a novel discriminative spatial coding algorithm which can automatically generate and select the most representative codebooks for image representation. This is achieved by first generate a number of spatial codebooks through over-complete image partition with overlap. Second, for each local feature to be encoded, the most discriminative codebook is selected by jointly minimizing the encoding error and the codebook's spatial distance. Experimental results on several public image datasets show the effectiveness of the proposed discriminative spatial coding method for efficient image classification.
Chunjie Zhang 0001, Jing Liu 0001, Weigang Zhang, Qingming Huang
ICIG5
2013 Cross-media topic detection: A multi-modality fusion framework
abstract
Detecting topics from Web data attracts increasing attention in recent years. Most previous works on topic detection mainly focus on the data from single medium, however, the rich and complementary information carried by multiple media can be used to effectively enhance the topic detection performance. In this paper, we propose a flexible data fusion framework to detect topics that simultaneously exist in different mediums. The framework is based on a multi-modality graph (MMG), which is obtained by fusing two single-modality graphs together: a text graph and a visual graph. Each node of MMGrepresents a multi-modal data and the edge weight between two nodes jointly measures their content and upload-time similarities. Since the data about the same topic often have similar content and are usually uploaded in a similar period of time, they would naturally form a dense (namely, strongly connected) subgraph in MMG. Such dense subgraph is robust to noise and can be efficiently detected by pair-wise clustering methods. The experimental results on single-medium and cross-media datasets demonstrate the flexibility and effectiveness of our method.
Guorong Li, Lingyang Chu, Shuhui Wang, Weigang Zhang, Qingming Huang
ICME5
2006 Extracting Story Units in Sports Video Based on Unsupervised Video Scene Clustering
abstract
Many sports videos such as archery, diving and tennis have repetitive structure patterns. They are reliable clues to generate highlights, summarization and automatic annotation. In this paper, we present a novel approach to analyze these structure patterns in sports video to extract story units. First, an unsupervised scene clustering method for sports video is adopted to automatically categorize the video shots into several disparate scenes. Then, the clustering results are modeled by a transition matrix. Finally, the key scene shots are detected to analyze the structure patterns and extract the story units. Experimental results on several types of broadcast sports video demonstrate that our approach is effective
Chunxi Liu, Qingming Huang, Shuqiang Jiang, Weigang Zhang
ICME4
2005 A System for Automatic Generation of Music Sports-Video
abstract
In this paper, we present a new representation of sports video abstract — Music Sports-Video (MSV), which provides exciting sports content accompanied with high quality background music for audiences and is available for high-quality audio-visual entertainment. We also propose a system generating MSV from user-provided sports video and music automatically. Firstly, the given sports video is segmented into a series of story units. Then all the story units are ordered by the predefined Exciting Degree (ED) and some high ED or user preferred story units are selected for MSV generation. Secondly, the ED of the given music is estimated by energy analysis on music beat. Thirdly, the selected story units are matched with music by their ED corresponding. Finally, the output MSV is rendered by connecting the selected exciting story units with appropriate transition effects, accompanied with the music. Experiments show encouraging results.
Weigang Zhang, Liyuan Xing, Qingming Huang, Wen Gao 0001
ICME1