Jiansheng Chen 0001

dblp:87/274-1 · DBLP profile ↗
← Back
59ranked-venue papers
6as first author
43since 2021 · last 2026
0000-0002-2040-7938ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 38 · 5 first-author · 23 since 2021Artificial intelligence and machine learning · 30 · 4 first-author · 22 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021
YearPublicationVenuePosition
2026 Object Detection Data Synthesis via Box-to-Image Generation Based on Diffusion Models
abstract
Modern diffusion-based image generative models have made significant progress and become promising to enrich training data for the object detection task. However, the generation quality and the controllability for complex scenes containing multi-class objects and dense objects with occlusions remain limited. This paper presents ODGEN, a novel method to generate high-quality images conditioned on bounding boxes, thereby facilitating data synthesis for object detection. Given a domain-specific object detection dataset, we first fine-tune a pre-trained diffusion model on both cropped foreground objects and entire images to fit target distributions. Then we propose to control the diffusion model using synthesized visual prompts with spatial constraints and object-wise textual descriptions. ODGEN exhibits robustness in handling complex scenes and specific domains. Further, we design a dataset synthesis pipeline to evaluate ODGEN on 7 domain-specific benchmarks to demonstrate its effectiveness. Adding training data generated by ODGEN improves up to 25.3% [email protected]:.95 with object detectors like YOLOv5 and YOLOv7, outperforming prior controllable generative methods. We also design an evaluation protocol based on COCO-2014 to validate the synthetic data of ODGEN in general domains and observe an advantage up to 5.6% in [email protected]:.95 against existing methods. In addition, we employ a series of large-scale object detection datasets to train a general model named Stable Box Diffusion, which covers thousands of object categories in most common scenes.
Huimin Ma 0001, Jiansheng Chen 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2026 Implicit alignment and query refinement for RGB-T semantic segmentation
Chang Liu 0136, Haizhuang Liu, Junbao Zhuo, Bochao Zou, Jiansheng Chen 0001, Qianchuan Zhao, Huimin Ma 0001
Pattern Recognit.5
2026 Multi-Scale Spatial Channel Joint Representation for General Multi-Modality Image Fusion With Self-Supervision
abstract
The rapid advancement of multi-modality image fusion technology enables researchers to simultaneously acquire information from different modalities within a single fused image. In existing methods, some general approaches can implement both infrared and visible image fusion (IVIF) and medical image fusion (MIF) in the same framework. Nevertheless, these methods often ignore the learning of specific features in different modalities, resulting in unsatisfactory performance in fused results. To overcome this issue, we propose a multi-scale joint framework with self-supervision for general multi-modality image fusion, abbreviated as SCSFusion. It enables more targeted and robust implementation of IVIF and MIF. Specifically, in the fusion network, a joint attention module is employed to parallelly capture self-attention features in spatial and channel domains, which can keep fused results accurate in visual representation. Meanwhile, we utilize source images of different modalities to generate visual-focused maps as pseudo labels for self-supervised training of the fusion results. It effectively preserves the salient details in each fused image from being disrupted by other extracted information. Moreover, a medical dataset with segmentation labels, termed M2DF, is reorganized for fusion and down-stream tasks in MIF. With the help of M2DF, a pre-trained segmentation model can be cascaded with the fusion network, aiming to obtain high-level semantic features from inputs and enhance the data generalization in our general framework. We have conducted extensive experiments and analyses on SCSFusion in M$\rm ^{3}$FD, FMB, and M2DF datasets, respectively. The results indicate that the fused images generated by SCSFusion can not only achieve visually appealing results and superior performance metrics in MIF and IVIF, but also exhibit satisfactory performance in down-stream tasks.
Jiawei Li 0016, Jiansheng Chen 0001, Jinyuan Liu 0001, Xinlong Ding, Huimin Ma 0001
IEEE Trans. Multim.2
2025 Enhancing Contrastive Learning Inspired by the Philosophy of "The Blind Men and the Elephant"
abstract
Contrastive learning is a prevalent technique in self-supervised vision representation learning, typically generating positive pairs by applying two data augmentations to the same image. Designing effective data augmentation strategies is crucial for the success of contrastive learning. Inspired by the story of the blind men and the elephant, we introduce JointCrop and JointBlur. These methods generate more challenging positive pairs by leveraging the joint distribution of the two augmentation parameters, thereby enabling contrastive learning to acquire more effective feature representations. To the best of our knowledge, this is the first effort to explicitly incorporate the joint distribution of two data augmentation parameters into contrastive learning. As a plug-and-play framework without additional computational overhead, JointCrop and JointBlur enhance the performance of SimCLR, BYOL, MoCo v1, MoCo v2, MoCo v3, SimSiam, and Dino baselines with notable improvements.
Yudong Zhang 0008, Ruobing Xie, Jiansheng Chen 0001, Xingwu Sun, Zhanhui Kang, Yu Wang 0002
AAAI3
2025 SAM2Object: Consolidating View Consistency via SAM2 for Zero-Shot 3D Instance Segmentation
abstract
In the field of zero-shot 3D instance segmentation, existing 2D-to-3D lifting methods typically obtain 2D segmentation across multiple RGB frames using vision foundation models, which are then projected and merged into 3D space. However, since the inference of vision foundation models on a single frame is not integrated with adjacent frames, the masks of the same object may vary across different frames, leading to a lack of view consistency in the 2D segmentation. Furthermore, current lifting methods average the 2D segmentation from multiple views during the projection into 3D space, causing low-quality masks and high-quality masks to share the same weight. These factors can lead to fragmented 3D segmentation. In this paper, we present SAM2Object, a novel zero-shot 3D instance segmentation method that effectively utilizes the Segment Anything Model 2 to segment and track objects, consolidating view consistency across frames. Our approach combines these consistent 2D masks with 3D geometric priors, improving the robustness of 3D segmentation. Additionally, we introduce mask consolidation module to filter out low-quality masks across frames, which enables more precise 2D-to-3D matching. Comprehensive evaluations on Scan-NetV2, ScanNet++ and ScanNet200 demonstrate the robustness and effectiveness of SAM2Object, showcasing its ability to outperform previous methods. Our project page is at https://jihuaizhaohd.github.io/SAM2Object.
Jihuai Zhao, Junbao Zhuo, Jiansheng Chen 0001, Huimin Ma 0001
CVPR3
2025 DADet: Safeguarding Image Conditional Diffusion Models Against Adversarial and Backdoor Attacks via Diffusion Anomaly Detection
Xinlong Ding, Jiawei Li 0016, Yudong Zhang 0008, Rongquan Wang, Huimin Ma 0001, Jiansheng Chen 0001
ICCV8
2025 From Black Boxes to Transparent Minds: Evaluating and Enhancing the Theory of Mind in Multimodal Large Language Models
abstract
As large language models evolve, there is growing anticipation that they will emulate human-like Theory of Mind (ToM) to assist with routine tasks. However, existing methods for evaluating machine ToM focus primarily on unimodal models and largely treat these models as black boxes, lacking an interpretative exploration of their internal mechanisms. In response, this study adopts an approach based on internal mechanisms to provide an interpretability-driven assessment of ToM in multimodal large language models (MLLMs). Specifically, we first construct a multimodal ToM test dataset, GridToM, which incorporates diverse belief testing tasks and perceptual information from multiple perspectives. Next, our analysis shows that attention heads in multimodal large models can distinguish cognitive information across perspectives, providing evidence of ToM capabilities. Furthermore, we present a lightweight, training-free approach that significantly enhances the model’s exhibited ToM by adjusting in the direction of the attention head.
Siqi Liu 0010, Bochao Zou, Jiansheng Chen 0001, Huimin Ma 0001
ICML4
2025 Fighting Fire with Fire (F3): A Training-free and Efficient Visual Adversarial Example Purification Method in LVLMs
abstract
Recent advances in large vision-language models (LVLMs) have showcased their remarkable capabilities across a wide range of multimodal vision-language tasks. However, these models remain vulnerable to visual adversarial attacks, which can substantially compromise their performance. In this paper, we introduce F3, a novel adversarial purification framework that employs a counterintuitive ''fighting fire with fire'' strategy: intentionally introducing simple perturbations to adversarial examples to mitigate their harmful effects. Specifically, F3 leverages cross-modal attentions derived from randomly perturbed adversary examples as reference targets. By injecting noise into these adversarial examples, F3 effectively refines their attention, resulting in cleaner and more reliable model outputs. Remarkably, this seemingly paradoxical approach of employing noise to counteract adversarial attacks yields impressive purification results. Furthermore, F3 offers several distinct advantages: it is training-free and straightforward to implement, and exhibits significant computational efficiency improvements compared to existing purification methods. These attributes render F3 particularly suitable for large-scale industrial applications where both robust performance and operational efficiency are critical priorities. The code is available at https://github.com/btzyd/F3.
Yudong Zhang 0008, Ruobing Xie, Jiansheng Chen 0001, Xingwu Sun, Zhanhui Kang, Di Wang 0052, Yu Wang 0002
ACM Multimedia4
2025 DHCP: Detecting Hallucinations by Cross-modal Attention Pattern in Large Vision-Language Models
abstract
Large vision-language models (LVLMs) have demonstrated exceptional performance on complex multimodal tasks. However, they continue to suffer from significant hallucination issues, including object, attribute, and relational hallucinations. To accurately detect these hallucinations, we investigated the variations in cross-modal attention patterns between hallucination and non-hallucination states. Leveraging these distinctions, we developed a lightweight detector capable of identifying hallucinations. Our proposed method, Detecting Hallucinations by Cross-modal Attention Patterns (DHCP), is straightforward and does not require additional LVLM training or extra LVLM inference steps. Experimental results show that DHCP achieves remarkable performance in hallucination detection. By offering novel insights into the identification and analysis of hallucinations in LVLMs, DHCP contributes to advancing the reliability and trustworthiness of these models. The code is available at https://github.com/btzyd/DHCP.
Yudong Zhang 0008, Ruobing Xie, Xingwu Sun, Jiansheng Chen 0001, Zhanhui Kang, Di Wang 0052, Yu Wang 0002
ACM Multimedia5
2025 QAVA: Query-Agnostic Visual Attack to Large Vision-Language Models
abstract
Yudong Zhang, Ruobing Xie, Jiansheng Chen, Xingwu Sun, Zhanhui Kang, Yu Wang. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Yudong Zhang 0008, Ruobing Xie, Jiansheng Chen 0001, Xingwu Sun, Zhanhui Kang, Yu Wang 0002
NAACL (Long Papers)3
2025 AGC-Drive: A Large-Scale Dataset for Real-World Aerial-Ground Collaboration in Driving Scenarios
abstract
By sharing information across multiple agents, collaborative perception helps autonomous vehicles mitigate occlusions and improve overall perception accuracy. While most previous work focus on vehicle-to-vehicle and vehicle-to-infrastructure collaboration, with limited attention to aerial perspectives provided by UAVs, which uniquely offer dynamic, top-down views to alleviate occlusions and monitor large-scale interactive environments. A major reason for this is the lack of high-quality datasets for aerial-ground collaborative scenarios. To bridge this gap, we present AGC-Drive, the first large-scale real-world dataset for Aerial-Ground Cooperative 3D perception. The data collection platform consists of two vehicles, each equipped with five cameras and one LiDAR sensor, and one UAV carrying a forward-facing camera and a LiDAR sensor, enabling comprehensive multi-view and multi-agent perception. Consisting of approximately 80K LiDAR frames and 360K images, the dataset covers 14 diverse real-world driving scenarios, including urban roundabouts, highway tunnels, and on/off ramps. Notably, 17\% of the data comprises dynamic interaction events, including vehicle cut-ins, cut-outs, and frequent lane changes. AGC-Drive contains 350 scenes, each with approximately 100 frames and fully annotated 3D bounding boxes covering 13 object categories. We provide benchmarks for two 3D perception tasks: vehicle-to-vehicle collaborative perception and vehicle-to-UAV collaborative perception. Additionally, we release an open-source toolkit, including spatiotemporal alignment verification tools, multi-agent visualization systems, and collaborative annotation utilities. The dataset and code are available at https://github.com/PercepX/AGC-Drive.
Yunhao Hou, Bochao Zou, Shangdong Yang, Junbao Zhuo, Siheng Chen, Jiansheng Chen 0001, Huimin Ma 0001
NeurIPS9
2025 Few Annotated Pixels and Point Cloud Based Weakly Supervised Semantic Segmentation of Driving Scenes
Huimin Ma 0001, Jiansheng Chen 0001, Yu Wang 0002
Int. J. Comput. Vis.4
2025 Correction: Few Annotated Pixels and Point Cloud Based Weakly Supervised Semantic Segmentation of Driving Scenes
Huimin Ma 0001, Jiansheng Chen 0001, Yu Wang 0002
Int. J. Comput. Vis.4
2025 Occlusion-guided multi-modal fusion for vehicle-infrastructure cooperative 3D object detection
Huazhen Chu, Haizhuang Liu, Junbao Zhuo, Jiansheng Chen 0001, Huimin Ma 0001
Pattern Recognit.4
2025 SparseComm: An Efficient Sparse Communication Framework for Vehicle-Infrastructure Cooperative 3D Detection
Haizhuang Liu, Huazhen Chu, Junbao Zhuo, Bochao Zou, Jiansheng Chen 0001, Huimin Ma 0001
Pattern Recognit.5
2025 RhythmFormer: Extracting patterned rPPG signals based on periodic sparse attention
Bochao Zou, Zizheng Guo 0002, Jiansheng Chen 0001, Junbao Zhuo, Weiran Huang 0001, Huimin Ma 0001
Pattern Recognit.3
2025 GET3DGS: Generate 3D Gaussians Based on Points Deformation Fields
abstract
The 3D Gaussian Splatting method has recently shown significant advancements in rendering speed and scene composition quality, enhancing its industrial applications and boosting the demand for 3D Gaussian asset generation. However, existing mature 3D generation technologies predominantly rely on implicit representations, which often struggle to balance geometric quality with editability. The production of 3D Gaussian assets generally involves diffusion models that require a dual-stage process of reconstruction and generation, resulting in substantial training and inference costs. To overcome these challenges, we introduce GET3DGS, an innovative approach that combines 3D-aware GANs with 3D Gaussian Splatting representations. This method facilitates the manipulation of the physical attributes of 3D Gaussians, such as geometry and texture, via point deformation fields. Offering faster inference speeds and end-to-end training capabilities, our model outperforms existing diffusion model-based methods. By deriving high-quality Gaussian point cloud geometric representations from 2D images, our approach reduces material accumulation costs and produces data compatible with 3D Gaussian rendering engines. We have evaluated the generative performance of our model on ShapeNet and OmniObject3D and demonstrate competitive results in terms of image and geometric quality relative to previous methods.
Haochen Yu, Weixi Gong, Jiansheng Chen 0001, Huimin Ma 0001
IEEE Trans. Circuits Syst. Video Technol.3
2025 HomuGAN: A 3D-Aware GAN With the Method of Cylindrical Spatial-Constrained Sampling
abstract
Controllable 3D-aware scene synthesis seeks to disentangle the various latent codes in the implicit space enabling the generation network to create highly realistic images with 3D consistency. Recent approaches often integrate Neural Radiance Fields with the upsampling method of StyleGAN2, employing Convolutions with style modulation to transform spatial coordinates into frequency domain representations. Our analysis indicates that this approach can give rise to a bubble phenomenon in StyleNeRF. We argue that the style modulation introduces extraneous information into the implicit space, disrupting 3D implicit modeling and degrading image quality. We introduce HomuGAN, incorporating two key improvements. First, we disentangle the style modulation applied to implicit modeling from that utilized for super-resolution, thus alleviating the bubble phenomenon. Second, we introduce Cylindrical Spatial-Constrained Sampling and Parabolic Sampling. The latter sampling method, as an alternative method to the former, specifically contributes to the performance of foreground modeling of vehicles. We evaluate HomuGAN on publicly available datasets, comparing its performance to existing methods. Empirical results demonstrate that our model achieves the best performance, exhibiting relatively outstanding disentanglement capability. Moreover, HomuGAN addresses the training instability problem observed in StyleNeRF and reduces the bubble phenomenon.
Haochen Yu, Weixi Gong, Jiansheng Chen 0001, Huimin Ma 0001
IEEE Trans. Image Process.3
2025 Individualized Driving Intention Prediction With Inverse Reinforcement Learning
abstract
Advanced Driver Assistance Systems (ADAS) are designed to prevent collisions, identify the condition of drivers while operating vehicles, and provide additional information to enhance drivers’ awareness of potential hazards on the road. Today, ADAS are capable of predicting drivers’ actions several seconds in advance, preparing for potential future hazards to prevent accidents or reduce injuries to occupants. Most previous works have achieved prediction results by analyzing and processing a vast amount of driving data from multiple drivers, based on the macro intention preferences of multiple drivers and external environmental features, collectively referred to as the generalized intention prediction network. This network utilizes extensive driving data to predict the common driving intentions of the overall driving population, without considering individualized driving styles. However, according to our research, different drivers exhibit distinct latent preferences in real-world driving scenarios. The generalized intention prediction network is influenced by these latent preferences, resulting in poor generalization capabilities and inaccurate predictions across different drivers. In this study, we propose a individualized driver intention prediction network. Based on Inverse Reinforcement Learning (IRL), it extracts individualized driving intention feature preferences that influence driving intentions from the driver’s historical behavior to improve generalized prediction results and achieve individualized driving intention prediction. We demonstrate that preferences vary among different drivers in the driving domain, leading to biases in model predictions. Upon experimental validation, the method we have proposed demonstrates remarkable efficacy on both the Brain4Cars and IESDD datasets, thereby showcasing its enhanced applicability in real-world scenarios.
Siqi Liu 0010, Jiansheng Chen 0001, Chenghao Guo, Jiehui Wu, Qifeng Luo, Huimin Ma 0001
IEEE Trans. Intell. Transp. Syst.3
2025 Improving Adversarial Robustness Against Universal Patch Attacks Through Feature Norm Suppressing
abstract
Universal adversarial patch attacks, which are readily implemented, have been validated to be able to fool real-world deep convolutional neural networks (CNNs), posing a serious threat to practical computer vision systems based on CNNs. Unfortunately, current defending approaches are severely understudied facing the following problems. Patch detection-based methods suffer from dramatic performance drops against white-box or adaptive attacks since they rely heavily on empirical clues. Methods based on adversarial training or certified defense are difficult to be scaled up to large-scale datasets or complex practical networks due to prohibitively high computational overhead or over strong assumptions on the network structure. In this article, we focus on two cases of widely adopted universal adversarial patch attacks, namely the universal targeted attack on image classifiers and the universal vanishing attack on object detectors. We find that, for popular CNNs, the attacking success of the adversarial patch relies on feature vectors centered at the patch location with large norm in classifiers and large channel-aware norm (CA-Norm) in detectors, and further present a mathematical explanation for this phenomenon. Based on this, we propose a simple but effective defending method using the feature norm suppressing (FNS) layer, which can renormalize the feature norm by nonincreasing functions. As a differentiable module, FNS can be adaptively inserted in various CNN architectures to achieve multistage suppression of the generation of large norm feature vectors. Moreover, FNS is efficient with no trainable parameters and very low computational overhead. We evaluate our proposed defending method across multiple CNN architectures and datasets against the strong adaptive white-box attacks in both visual classification and detection tasks. In both tasks, FNS significantly outperforms previous defending methods on adversarial robustness with a relatively low influence on the performance of benign images. Code is available at https://github.com/jschenthu/FNS.
Jiansheng Chen 0001, Yu Wang 0002, Youze Xue, Huimin Ma 0001
IEEE Trans. Neural Networks Learn. Syst.2
2025 Isolated Diffusion: Optimizing Multi-Concept Text-to-Image Generation Training-Freely With Isolated Diffusion Guidance
abstract
Large-scale text-to-image diffusion models have achieved great success in synthesizing high-quality and diverse images given target text prompts. Despite the revolutionary image generation ability, current state-of-the-art models still struggle to deal with multi-concept generation accurately in many cases. This phenomenon is known as "concept bleeding" and displays as the unexpected overlapping or merging of various concepts. This paper presents a general approach for text-to-image diffusion models to address the mutual interference between different subjects and their attachments in complex scenes, pursuing better text-image consistency. The core idea is to isolate the synthesizing processes of different concepts. We propose to bind each attachment to corresponding subjects separately with split text prompts. Besides, we introduce a revision method to fix the concept bleeding problem in multi-subject synthesis. We first depend on pre-trained object detection and segmentation models to obtain the layouts of subjects. Then we isolate and resynthesize each subject individually with corresponding text prompts to avoid mutual interference. Overall, we achieve a training-free strategy, named Isolated Diffusion, to optimize multi-concept text-to-image synthesis. It is compatible with the latest Stable Diffusion XL (SDXL) and prior Stable Diffusion (SD) models. We compare our approach with alternative methods using a variety of multi-concept text prompts and demonstrate its effectiveness with clear advantages in text-image consistency and user study.
Huimin Ma 0001, Jiansheng Chen 0001
IEEE Trans. Vis. Comput. Graph.3
2024 Step Vulnerability Guided Mean Fluctuation Adversarial Attack against Conditional Diffusion Models
abstract
The high-quality generation results of conditional diffusion models have brought about concerns regarding privacy and copyright issues. As a possible technique for preventing the abuse of diffusion models, the adversarial attack against diffusion models has attracted academic attention recently. In this work, utilizing the phenomenon that diffusion models are highly sensitive to the mean value of the input noise, we propose the Mean Fluctuation Attack (MFA) to introduce mean fluctuations by shifting the mean values of the estimated noises during the reverse process. In addition, we reveal that the vulnerability of different reverse steps against adversarial attacks actually varies significantly. By modeling the step vulnerability and using it as guidance to sample the target steps for generating adversarial examples, the effectiveness of adversarial attacks can be substantially enhanced. Extensive experiments show that our algorithm can steadily cause the mean shift of the predicted noises so as to disrupt the entire reverse generation process and degrade the generation results significantly. We also demonstrate that the step vulnerability is intrinsic to the reverse process by verifying its effectiveness in an attack method other than MFA. Code and Supplementary is available at https://github.com/yuhongwei22/MFA
Jiansheng Chen 0001, Xinlong Ding, Yudong Zhang 0008, Ting Tang, Huimin Ma 0001
AAAI2
2024 Upper-Body Hierarchical Graph for Skeleton Based Emotion Recognition in Assistive Driving
Jiehui Wu, Jiansheng Chen 0001, Qifeng Luo, Siqi Liu 0010, Youze Xue, Huimin Ma 0001
ECCV (26)2
2024 PLS: Unsupervised Domain Adaptation for 3d Object Detection Via Pseudo-Label Sizes
abstract
3D object detection has gained increasing attention in modern autonomous driving systems. However, the performance of the detector significantly degrades during cross-domain deployment due to domain shift. The detector is inevitably biased towards its training dataset when employed on a target dataset, particularly towards object sizes. State-of-the-art unsupervised domain adaptation approaches explicitly address the variation in object sizes by appropriately scaling the source data. However, such methods require additional target domain statistics information, which contradicts the original unsupervised assumption. In this work, we present PLS, a novel unsupervised domain adaptation method for 3D object detection to overcome the object sizes bias via Pseudo-Label Sizes, which utilizes only source domain annotations. PLS alternates between generating high-quality pseudo-label sizes through the detector and model training with the pseudo-label sizes to scale and augment the source data. This iterative process enables the detector to be trained with augmented data that resembles the target domain sizes, thereby improving the performance of detector in cross-domain scenarios. Our experimental results show the outstanding performance of our PLS in various scenarios. In addition, PLS is a plug-and-play module that can be used to directly replace existing weakly-supervised scaling methods. Experimental results show that existing excellent architectures with PLS are able to achieve better performance, and making them completely unsupervised.
Rongquan Wang, Xin Li 0034, Haizhuang Liu, Jiansheng Chen 0001, Huimin Ma 0001
ICASSP6
2024 CMT: Co-training Mean-Teacher for Unsupervised Domain Adaptation on 3D Object Detection
Junbao Zhuo, Xin Li 0034, Haizhuang Liu, Rongquan Wang, Jiansheng Chen 0001, Huimin Ma 0001
ACM Multimedia6
2024 Affinity3D: Propagating Instance-Level Semantic Affinity for Zero-Shot Point Cloud Semantic Segmentation
abstract
Zero-shot point cloud semantic segmentation aims to recognize novel classes at the point level. Previous methods mainly transfer excellent zero-shot generalization capabilities from images to point clouds. However, directly transferring knowledge from images to point clouds faces two ambiguous problems. On the one hand, 2D models will generate wrong predictions when the image changes. On the other hand, directly mapping 3D points to 2D pixels by perspective projection fails to consider the visibility of 3D points in camera view. The wrong geometric alignment of 3D points and 2D pixels causes semantic ambiguity. To tackle these two problems, we propose a framework named Affinity3D that intends to empower 3D semantic segmentation models to perceive novel samples. Our framework aggregates instances in 3D and recognizes them in 2D, leveraging the excellent geometric separation in 3D and the zero-shot capabilities of 2D models. Affinity3D involves an affinity module that rectifies the wrong predictions by comparing them with similar instances and a visibility module preventing knowledge transfer from visible 2D pixels to invisible 3D points. Extensive experiments have been conducted on the SemanticKITTI and nuScenes datasets. Our framework achieves state-of-the-art performance on both two datasets. Code is available at https://github.com/opjang5/Affinity3D.
Haizhuang Liu, Junbao Zhuo, Jiansheng Chen 0001, Huimin Ma 0001
ACM Multimedia4
2024 PIP: Detecting Adversarial Examples in Large Vision-Language Models via Attention Patterns of Irrelevant Probe Questions
abstract
Large Vision-Language Models (LVLMs) have demonstrated their powerful multimodal capabilities. However, they also face serious safety problems, as adversaries can induce robustness issues in LVLMs through the use of well-designed adversarial examples. Therefore, LVLMs are in urgent need of detection tools for adversarial examples to prevent incorrect responses. In this work, we first discover that LVLMs exhibit regular attention patterns for clean images when presented with probe questions. We propose an unconventional method named PIP, which utilizes the attention patterns of one randomly selected irrelevant probe question (e.g., "Is there a clock''') to distinguish adversarial examples from clean examples. Regardless of the image to be tested and its corresponding question, PIP only needs to perform one additional inference of the image to be tested and the probe question, and then achieves successful detection of adversarial examples. Even under black-box attacks and open dataset scenarios, our PIP, coupled with a simple SVM, still achieves more than 98% recall and a precision of over 90%. Our PIP is the first attempt to detect adversarial attacks on LVLMs via simple irrelevant probe questions, shedding light on deeper understanding and introspection within LVLMs. The code is available at https://github.com/btzyd/pip.
Yudong Zhang 0008, Ruobing Xie, Jiansheng Chen 0001, Xingwu Sun, Yu Wang 0002
ACM Multimedia3
2024 Improving Item-side Fairness of Multimodal Recommendation via Modality Debiasing
abstract
Multimodal recommender systems have acquired applications in broad web scenarios such as e-commerce businesses and short-video platforms. Existing multimodal recommendation methods generally boost performance by introducing item-side multimodal content as supplement information. However, the common training paradigm, i.e., encoding unimodal content respectively and fusing them to fit user preference scores, makes the model biased towards items with prevailing modality content under non-uniform training data. This results in a serious item-side unfairness issue, i.e., some items with prevailing modality content are over-recommended while a large number of items don't receive adequate recommendation opportunities, leaving corresponding content providers at great disadvantage. Aiming to eliminate such modality bias and promote item-side fairness, we propose a fairness-aware modality debiasing framework based on counterfactual inference. In the training stage, we additionally introduce unimodal prediction branches to capture the modality bias. In the inference stage, we conduct a fairness-aware counterfactual inference to adaptively eliminate the modality bias. The proposed framework is model-agnostic and flexible to be implemented in various multimodal recommendation models. Extensive experiments on two datasets demonstrate that the proposed method can significantly enhance item-side fairness while providing competitive recommendation accuracy. Our proposed framework is expected to help mitigate the unfair treatment experienced by vulnerable content providers on multimedia web platforms. Codes are available in https://github.com/tsinghua-fib-lab-WWW2024-Modality-Debiasing.
Chen Gao 0001, Jiansheng Chen 0001, Depeng Jin, Yong Li 0008
WWW3
2024 Image paragraph captioning with topic clustering and topic shift prediction
Ting Tang, Jiansheng Chen 0001, Huimin Ma 0001, Yudong Zhang 0008
Knowl. Based Syst.2
2024 High-Quality and Diverse Few-Shot Image Generation via Masked Discrimination
abstract
Few-shot image generation aims to generate images of high quality and great diversity with limited data. However, it is difficult for modern GANs to avoid overfitting when trained on only a few images. The discriminator can easily remember all the training samples and guide the generator to replicate them, leading to severe diversity degradation. Several methods have been proposed to relieve overfitting by adapting GANs pre-trained on large source domains to target domains using limited real samples. This work presents masked discrimination to realize few-shot GAN adaptation, which is the first feature-level augmentation method for generative tasks. Random masks are applied to features extracted by the discriminator from input images. We aim to encourage the discriminator to judge various images that share partially common features with training samples as realistic. Correspondingly, the generator is guided to generate diverse images instead of replicating training samples. In addition, we employ a cross-domain consistency loss for the discriminator to keep relative distances between generated samples in its feature space. It strengthens global image discrimination and guides adapted GANs to preserve more information learned from source domains for higher image quality, resulting in better cross-domain correspondence. The effectiveness of our approach is demonstrated both qualitatively and quantitatively with higher quality and greater diversity on a series of few-shot image generation tasks than prior methods.
Huimin Ma 0001, Jiansheng Chen 0001
IEEE Trans. Image Process.3
2023 Transferable Structure-based Adversarial Attack of Heterogeneous Graph Neural Network
abstract
Heterogeneous graph neural networks (HGNNs) have achieved remarkable development recently and exhibited superior performance in various tasks. However, recently HGNNs have been shown to have robustness weakness towards adversarial perturbations, which brings critical pitfalls for real applications, e.g. node classification and recommender systems. In particular, the transfer-based black-box attack is the most practical method to attack unknown models and poses a great threat to the reliability of HGNNs. In this work, we take the first step to explore the transferability of adversarial examples of HGNNs. Due to the overfitting of the source model, the adversarial perturbations generated by traditional methods usually exhibit unpromising transferability. To address this problem and boost adversarial transferability, we expect to seek common vulnerable directions of different models to attack. Inspired by the observation of the notable commonality of edge attention distribution between different HGNNs, we propose to guide the perturbation generation toward disrupting edge attention distribution. This edge attention-guided attack prioritizes the perturbation on edges that are more likely to be given common attention by different models, which benefits the transferability of adversarial perturbations. Finally, we develop two edge attention-guided attack methods towards heterogeneous relations tailored for HGNNs, called EA-FGSM and EA-PGD. Extensive experiments on six representative models and two datasets verify the effectiveness of our methods and form an unprecedented transfer robustness benchmark for HGNNs.
Yudong Zhang 0008, Jiansheng Chen 0001, Depeng Jin, Yong Li 0008
CIKM3
2023 Enhancing Adversarial Robustness of Multi-modal Recommendation via Modality Balancing
abstract
Recently multi-modal recommender systems have been widely applied in real scenarios such as e-commerce businesses. Existing multi-modal recommendation methods exploit the multi-modal content of items as auxiliary information and fuse them to boost performance. Despite the superior performance achieved by multi-modal recommendation models, there's currently no understanding of their robustness to adversarial attacks. In this work, we first identify the vulnerability of existing multi-modal recommendation models. Next, we show the key reason for such vulnerability is modality imbalance, i.e., the prediction score margin between positive and negative samples in the sensitive modality will drop dramatically facing adversarial attacks and fail to be compensated by other modalities. Finally, based on this finding we propose a novel defense method to enhance the robustness of multi-modal recommendation models through modality balancing. Specifically, we first adopt an embedding distillation to obtain a pair of content-similar but prediction-different item embeddings in the sensitive modality and calculate the score margin reflecting the modality vulnerability. Then we optimize the model to utilize the score margin between positive and negative samples in other modalities to compensate for the vulnerability. The proposed method can serve as a plug-and-play module and is flexible to be applied to a wide range of multi-modal recommendation models. Extensive experiments on two real-world datasets demonstrate that our method significantly improves the robustness of multi-modal recommendation models with nearly no performance degradation on clean data.
Chen Gao 0001, Jiansheng Chen 0001, Depeng Jin, Huimin Ma 0001, Yong Li 0008
ACM Multimedia3
2023 Learning Fine-grained User Interests for Micro-video Recommendation
abstract
Recent years have witnessed the rapid development of online micro-video platforms, in which the recommender system plays an essential role in overcoming the information overloading problem and providing personalized content for users. Although some progress has been achieved in the micro-video recommendation, there are still some limitations in learning the representations of user interests and video features. Specifically, the user modeling in existing works is performed at a coarse-grained level, i.e., video level. However, in micro-video recommendation, the user feedback is at a continuous form---users can skip over a video at each frame---which reveals fine-grained user preferences. In this work, we approach the problem of learning fine-grained user preferences for micro-video recommendation by first collecting two real-world datasets. To address the challenges of preference modeling and weak supervision signal, we propose a solution named FRAME (short for Fine-gRAined preference-modeling for Micro-video rEcommendation). Specifically, we first adopt visual feature extraction and transformation to maintain the fine-grained video embeddings. We then propose graph convolution layers to learn the user preference from complex and fine-grained user-clip relations, and hybrid-supervision objectives for enhancing the supervision signal. The experimental results on two collected real-world datasets demonstrate the effectiveness of our proposed model. We release the datasets and codes in https://github.com/tsinghua-fib-lab/FRAME, which we believe can benefit the community.
Chen Gao 0001, Jiansheng Chen 0001, Depeng Jin, Meng Wang 0001, Yong Li 0008
SIGIR3
2023 Shaping Deep Feature Space Towards Gaussian Mixture for Visual Classification
abstract
The softmax cross-entropy loss function has been widely used to train deep models for various tasks. In this work, we propose a Gaussian mixture (GM) loss function for deep neural networks for visual classification. Unlike the softmax cross-entropy loss, our method explicitly shapes the deep feature space towards a Gaussian Mixture distribution. With a classification margin and a likelihood regularization, the GM loss facilitates both high classification performance and accurate modeling of the feature distribution. The GM loss can be readily used to distinguish the adversarial examples based on the discrepancy between feature distributions of clean and adversarial examples. Furthermore, theoretical analysis shows that a symmetric feature space can be achieved by using the GM loss, which enables the models to perform robustly against adversarial attacks. The proposed model can be implemented easily and efficiently without introducing more trainable parameters. Extensive evaluations demonstrate that the method with the GM loss performs favorably on image classification, face recognition, and detection as well as recognition of adversarial examples generated by various attacks.
Weitao Wan, Jiansheng Chen 0001, Yuanyi Zhong, Ming-Hsuan Yang 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2023 Temporal Information Fusion Network for Driving Behavior Prediction
abstract
Since enormous hazards are caused by traffic crashes every year, ensuring safe driving is a hot topic in transportation. Technologies related to the Advanced Driver Assistance System (ADAS) are evolving rapidly. But without an adequate understanding of driving intention, ADAS usually can’t help the driver prepare for the danger in advance. This paper focuses on the fusion strategy of driver and environment information and proposes a lightweight end-to-end model, temporal information fusion network (TIFN). Driving behavior is the interactive result of the driver and the external world. To better understand the driver’s intention, the state update cell (STU) is proposed to introduce the influence of environment information into the driver’s state modeling, inspired by the selective attention of the human cognition process. Meanwhile, semantic segmentation features are extracted to offer clear clues affecting driver attention in place of motion optical flow images and binary value vectors. Finally, the driver’s intention and environment state are combined to make a joint prediction. The experiments evaluated on Brain4cars and IESDD show that the proposed approach has superior performance than other approaches that only use camera data.
Chenghao Guo, Haizhuang Liu, Jiansheng Chen 0001, Huimin Ma 0001
IEEE Trans. Intell. Transp. Syst.3
2023 MotionVideoGAN: A Novel Video Generator Based on the Motion Space Learned From Image Pairs
abstract
Video generation has achieved rapid progress benefiting from high-quality renderings provided by powerful image generators. We regard the video synthesis task as generating a sequence of images sharing the same contents but varying in motions. However, most previous video synthesis frameworks based on pre-trained image generators treat content and motion generation separately, leading to unrealistic generated videos. Therefore, we design a novel framework to build the motion space, aiming to achieve content consistency and fast convergence for video generation. We present MotionVideoGAN, a novel video generator synthesizing videos based on the motion space learned by pre-trained image pair generators. Firstly, we propose an image pair generator named MotionStyleGAN to generate image pairs sharing the same contents and producing various motions. Then we manage to acquire motion codes to edit one image in the generated image pairs and keep the other unchanged. The motion codes help us edit images within the motion space since the edited image shares the same contents with the other unchanged one in image pairs. Finally, we introduce a latent code generator to produce latent code sequences using motion codes for video generation. Our approach achieves state-of-the-art performance on the most complex video dataset ever used for unconditional video generation evaluation, UCF101.
Huimin Ma 0001, Jiansheng Chen 0001
IEEE Trans. Multim.3
2022 Gestalt-Guided Image Understanding for Few-Shot Learning
Kun Song 0004, Jiansheng Chen 0001, Huimin Ma 0001
ACCV (2)3
2022 Eliminating Spatial Ambiguity for Weakly Supervised 3D Object Detection without Spatial Labels
abstract
Previous weakly-supervised methods of 3D object detection in driving scenes mainly rely on spatial labels, which provide the location, dimension, or orientation information. The annotation of 3D spatial labels is time-consuming. There also exist methods that do not require spatial labels, but their detections may fall on object parts rather than entire objects or backgrounds. In this paper, a novel cross-modal weakly-supervised 3D progressive refinement framework (WS3DPR) for 3D object detection that only needs image-level class annotations is introduced. The proposed framework consists of two stages: 1) classification refinement for potential objects localization and 2) regression refinement for spatial pseudo labels reasoning. In the first stage, a region proposal network is trained by cross-modal class knowledge transferred from 2D image to 3D point cloud and class information propagation. In the second stage, the locations, dimensions, and orientations of 3D bounding boxes are further refined with geometric reasoning based on 2D frustum and 3D region. When only image-level class labels are available, proposals with different 3D locations become overlapped in 2D, leading to the misclassification of foreground objects. Therefore, a 2D-3D semantic consistency block is proposed to disentangle different 3D proposals after projection. The overall framework progressively learns features in a coarse to fine manner. Comprehensive experiments on the KITTI3D dataset demonstrate that our method achieves competitive performance compared with previous methods with a lightweight labeling process.
Haizhuang Liu, Huimin Ma 0001, Bochao Zou, Rongquan Wang, Jiansheng Chen 0001
ACM Multimedia7
2022 3D Human Mesh Reconstruction by Learning to Sample Joint Adaptive Tokens for Transformers
abstract
Reconstructing 3D human mesh from a single RGB image is a challenging task due to the inherent depth ambiguity. Researchers commonly use convolutional neural networks to extract features and then apply spatial aggregation on the feature maps to explore the embedded 3D cues in the 2D image. Recently, two methods of spatial aggregation, the transformers and the spatial attention, are adopted to achieve the state-of-the-art performance, whereas they both have limitations. The use of transformers helps modelling long-term dependency across different joints whereas the grid tokens are not adaptive for the positions and shapes of human joints in different images. On the contrary, the spatial attention focuses on joint-specific features. However, the non-local information of the body is ignored by the concentrated attention maps. To address these issues, we propose a Learnable Sampling module to generate joint adaptive tokens and then use transformers to aggregate global information. Feature vectors are sampled accordingly from the feature maps to form the tokens of different joints. The sampling weights are predicted by a learnable network so that the model can learn to sample joint-related features adaptively. Our adaptive tokens are explicitly correlated with human joints, so that more effective modeling of global dependency among different human joints can be achieved. To validate the effectiveness of our method, we conduct experiments on several popular datasets including Human3.6M and 3DPW. Our method achieves lower reconstruction errors in terms of both the vertex-based metric and the joint-based metric compared to previous state of the arts. The codes and the trained models are released at https://github.com/thuxyz19/Learnable-Sampling.
Youze Xue, Jiansheng Chen 0001, Yudong Zhang 0008, Huimin Ma 0001, Hongbing Ma
ACM Multimedia2
2022 Attribute assisted teacher-critical training strategies for image captioning
Jiansheng Chen 0001, Huimin Ma 0001, Hongbing Ma, Wanli Ouyang
Neurocomputing2
2022 Co-attention dictionary network for weakly-supervised semantic segmentation
Weitao Wan, Jiansheng Chen 0001, Ming-Hsuan Yang 0001, Huimin Ma 0001
Neurocomputing2
2022 Boosting Monocular 3D Human Pose Estimation With Part Aware Attention
abstract
Monocular 3D human pose estimation is challenging due to depth ambiguity. Convolution-based and Graph-Convolution-based methods have been developed to extract 3D information from temporal cues in motion videos. Typically, in the lifting-based methods, most recent works adopt the transformer to model the temporal relationship of 2D keypoint sequences. These previous works usually consider all the joints of a skeleton as a whole and then calculate the temporal attention based on the overall characteristics of the skeleton. Nevertheless, the human skeleton exhibits obvious part-wise inconsistency of motion patterns. It is therefore more appropriate to consider each part's temporal behaviors separately. To deal with such part-wise motion inconsistency, we propose the Part Aware Temporal Attention module to extract the temporal dependency of each part separately. Moreover, the conventional attention mechanism in 3D pose estimation usually calculates attention within a short time interval. This indicates that only the correlation within the temporal context is considered. Whereas, we find that the part-wise structure of the human skeleton is repeating across different periods, actions, and even subjects. Therefore, the part-wise correlation at a distance can be utilized to further boost 3D pose estimation. We thus propose the Part Aware Dictionary Attention module to calculate the attention for the part-wise features of input in a dictionary, which contains multiple 3D skeletons sampled from the training set. Extensive experimental results show that our proposed part aware attention mechanism helps a transformer-based model to achieve state-of-the-art 3D pose estimation performance on two widely used public datasets. The codes and the trained models are released at https://github.com/thuxyz19/3D-HPE-PAA.
Youze Xue, Jiansheng Chen 0001, Xiangming Gu, Huimin Ma 0001, Hongbing Ma
IEEE Trans. Image Process.2
2021 Enhancing Adversarial Robustness For Image Classification By Regularizing Class Level Feature Distribution
abstract
Recent researches have shown that deep neural networks (DNNs) are vulnerable to adversarial examples. Adversarial training is practically the most effective approach to improve the robustness of DNNs against adversarial examples. However, conventional adversarial training methods only focus on the classification results or the instance level relationship on feature representations for adversarial examples. Inspired by the fact that adversarial examples break the distinguishability of the feature representations of DNNs for different classes, we propose Intra and Inter Class Feature Regularization $(\mathrm{I}^{2}$ FR) to make the feature distribution of adversarial examples maintain the same classification property as clean examples. On the one hand, the intra-class regularization restricts the distance of features between adversarial examples and both the corresponding clean data and samples for the same class. On the other hand, the inter-class regularization prevents the feature of adversarial examples from getting close to other classes. By adding $\mathrm{I}^{2}$ FR in both adversarial example generation and model training steps in adversarial training, we can get stronger and more diverse adversarial examples, and the neural network learns a more distinguishable and reasonable feature distribution. Experiments on various adversarial training frameworks demonstrate that $\mathrm{I}^{2}$ FR is adaptive for multiple training frameworks and outperforms the state-of-the-art methods for classification of both clean data and adversarial examples.
Youze Xue, Jiansheng Chen 0001, Yu Wang 0002, Huimin Ma 0001
ICIP3
2020 Show, Conceive and Tell: Image Captioning with Prospective Linguistic Information
Jiansheng Chen 0001
ACCV (6)2
2020 Adversarial Training with Bi-directional Likelihood Regularization for Visual Classification
Weitao Wan, Jiansheng Chen 0001, Ming-Hsuan Yang 0001
ECCV (24)2
2020 Image Captioning With End-to-End Attribute Detection and Subsequent Attributes Prediction
abstract
Semantic attention has been shown to be effective in improving the performance of image captioning. The core of semantic attention based methods is to drive the model to attend to semantically important words, or attributes. In previous works, the attribute detector and the captioning network are usually independent, leading to the insufficient usage of the semantic information. Also, all the detected attributes, no matter whether they are appropriate for the linguistic context at the current step, are attended to through the whole caption generation process. This may sometimes disrupt the captioning model to attend to incorrect visual concepts. To solve these problems, we introduce two end-to-end trainable modules to closely couple attribute detection with image captioning as well as prompt the effective uses of attributes by predicting appropriate attributes at each time step. The multimodal attribute detector (MAD) module improves the attribute detection accuracy by using not only the image features but also the word embedding of attributes already existing in most captioning models. MAD models the similarity between the semantics of attributes and the image object features to facilitate accurate detection. The subsequent attribute predictor (SAP) module dynamically predicts a concise attribute subset at each time step to mitigate the diversity of image attributes. Compared to previous attribute based methods, our approach enhances the explainability in how the attributes affect the generated words and achieves a state-of-the-art single model performance of 128.8 CIDEr-D on the MSCOCO dataset. Extensive experiments on the MSCOCO dataset show that our proposal actually improves the performances in both image captioning and attribute detection simultaneously. The codes are available at: https://github.com/ RubickH/Image-Captioning-with-MAD-and-SAP.
Jiansheng Chen 0001, Wanli Ouyang, Weitao Wan, Youze Xue
IEEE Trans. Image Process.2
2019 sEMG-Based Tremor Severity Evaluation for Parkinson's Disease Using a Light-Weight CNN
abstract
We propose a deep learning based approach for quantifying the tremor severity of Parkinson's disease (PD) based on surface electromyography (sEMG). We design the S-Net, a light weight and computational efficient convolutional neural network that learns the similarity between sEMG signals in terms of the tremor severity. Labeled sEMG samples are used for jointly voting for the final results. Experiments on 147 PD patients demonstrate that our approach outperforms traditional methods by a significant margin. In addition, our approach is simple and has potentials in real applications.
Zengyi Qin, Zhenyu Jiang 0002, Jiansheng Chen 0001
IEEE Signal Process. Lett.3
2017 Linear Spectral Clustering Superpixel
abstract
In this paper, we present a superpixel segmentation algorithm called linear spectral clustering (LSC), which is capable of producing superpixels with both high boundary adherence and visual compactness for natural images with low computational costs. In LSC, a normalized cuts-based formulation of image segmentation is adopted using a distance metric that measures both the color similarity and the space proximity between image pixels. However, rather than directly using the traditional eigen-based algorithm, we approximate the similarity metric through a deliberately designed kernel function such that pixel values can be explicitly mapped to a high-dimensional feature space. We then apply the conclusion that by appropriately weighting each point in this feature space, the objective functions of the weighted K-means and the normalized cuts share the same optimum points. Consequently, it is possible to optimize the cost function of the normalized cuts by iteratively applying simple K-means clustering in the proposed feature space. LSC possesses linear computational complexity and high memory efficiency, since it avoids both the decomposition of the affinity matrix and the generation of the large kernel matrix. By utilizing the underlying mathematical equivalence between the two types of seemingly different methods, LSC successfully preserves global image structures through efficient local operations. Experimental results show that LSC performs as well as or even better than the state-of-the-art superpixel segmentation algorithms in terms of several commonly used evaluation metrics in image segmentation. The applicability of LSC is further demonstrated in two related computer vision tasks.
Jiansheng Chen 0001, Zhengqin Li
IEEE Trans. Image Process.1
2015 Superpixel segmentation using Linear Spectral Clustering
abstract
We present in this paper a superpixel segmentation algorithm called Linear Spectral Clustering (LSC), which produces compact and uniform superpixels with low computational costs. Basically, a normalized cuts formulation of the superpixel segmentation is adopted based on a similarity metric that measures the color similarity and space proximity between image pixels. However, instead of using the traditional eigen-based algorithm, we approximate the similarity metric using a kernel function leading to an explicitly mapping of pixel values and coordinates into a high dimensional feature space. We revisit the conclusion that by appropriately weighting each point in this feature space, the objective functions of weighted K-means and normalized cuts share the same optimum point. As such, it is possible to optimize the cost function of normalized cuts by iteratively applying simple K-means clustering in the proposed feature space. LSC is of linear computational complexity and high memory efficiency and is able to preserve global properties of images. Experimental results show that LSC performs equally well or better than state of the art superpixel segmentation algorithms in terms of several commonly used evaluation metrics in image segmentation.
Zhengqin Li, Jiansheng Chen 0001
CVPR2
2010 CPGL: A classification method combining PCA and the Group Lasso method
abstract
Sparse representation based optimization has emerged as a new paradigm for solving classification problems and has achieved satisfactory performances. Recent research, however, has revealed its noteworthy limitation in handling samples with high intra-class pair-wise correlations. In this paper, we study this problem from a novel perspective of de-correlating the input data. A new method is proposed by combining Principle Component Analysis (PCA) and the Group Lasso method. The highly correlated training samples are first orthogonalized using PCA, and then the Group Lasso algorithm is adopted for performing the classification. Experimental results show that our proposed method over-performs the Group Lasso method in the face recognition application on two public databases.
Jing Wang 0194, Guangda Su, Jiansheng Chen 0001, Yiu Sang Moon
ICIP3
2010 Restoration of low resolution car plate images using PCA based image super-resolution
abstract
In this paper, a car plate image restoration system is presented, in which the PCA (Principle Component Analysis) based super-resolution method is applied to restore very low resolution car plate images (height <; 10 pixels) from mainland China and Hong Kong. Several algorithms are proposed to improve the robustness and performance of this system. The core part of this system is a dynamic PCA algorithm, in which the training set is adjusted adaptively for achieving better restoration results. Extensive experiments on both simulated test images and practical test images show that our system performances satisfactorily in terms of the restoration error rate and the flexibility in practical applications.
Guangda Su, Jiansheng Chen 0001, Yiu Sang Moon
ICIP3
2010 Palmprint authentication using a symbolic representation of images
Jiansheng Chen 0001, Yiu Sang Moon, Ming-Fai Wong, Guangda Su
Image Vis. Comput.1
2009 Piecewise linear aging function for facial age estimation
abstract
Instead of constructing a complex aging model on the whole age space, a piecewise linear aging function is proposed to approximate the ground truth aging function locally. To handle the `regression toward the mean' problem of local aging function regression, a weighting strategy is used to assign larger weights to the samples near typical aging appearance of each age. In age estimation step, a global aging function is used to predict the test sample's rough age range and the local linear aging function on that range is then used to give a final result. Experimental results show that our method can get more accuracy results in facial age estimation.
Shenglan Ben, Jiansheng Chen 0001, Guangda Su
ICIP2
2008 The statistical modelling of fingerprint minutiae distribution with implications for fingerprint individuality studies
abstract
The spatial distribution of fingerprint minutiae is a core problem in the fingerprint individuality study, the cornerstone of the fingerprint authentication technology. Previously, the assumption in most research that minutiae distribution is random has been proved to be inaccurate and may lead to significant overestimates of fingerprint uniqueness. In this paper, we propose a stochastic model for describing and simulating fingerprint minutiae patterns. Through coupling a pair potential Markov point process with a thinned process, this model successfully depicts the complex statistical behavior of fingerprint minutiae. Parameters of this model can be determined by nonlinear minimization. Furthermore, experiment results show that the statistical properties of our proposed model dovetails nicely with real minutiae data in terms of the false fingerprint correspondence probability. Such evidences indicate that the proposed model is a more accurate foundation for minutiae based fingerprint individuality studies as well as the artificial fingerprint synthesis when compared to the model of random distribution.
Jiansheng Chen 0001, Yiu Sang Moon
CVPR1
2008 Towards more accurate 3D face registration under the guidance of prior anatomical knowledge on human faces
abstract
Three-dimensional face registration is a critical step in 3D face recognition. A fully automatic registration method for aligning frontal 3D face data is presented in this paper with high accuracy and robustness to facial expressions. In our method, the nose region, which is relatively more rigid than other facial regions in Anatomy sense, is automatically located and analyzed for computing the precise location of a symmetry plane. We then proceed on to find a stable reference point and a nose line from the global information of the nose region. In this way, the six degrees of freedom as well as a unified coordinate system can be determined for each face. Extensive experiments have been conducted on the FRGC V1.0 benchmark face dataset to evaluate the accuracy and robustness of our registration method. Firstly, we compare its results with two other registration methods. One of such methods employs manually marked points on visualized face data and the other is based on the use of a symmetry plane analysis obtained from the whole face region. Secondly, we test its application in a 3-D face verification system. Preliminary experiment results show that this approach can efficiently reduce the intra-class distance and performs better than the other two registration methods in face recognition.
X. M. Tang, Jiansheng Chen 0001, Yiu Sang Moon
FG2
2008 Real-time object correspondence in stereo camera system
abstract
In this paper, we address the problem of object correspondence construction in stereo camera systems by using a real-time algorithm adopting reverse stereo triangulation. This algorithm is based on a belief that any incorrect object-pair will eventually demonstrate inconsistency in its spatial location calculated from stereo triangulation, so that correct object-pairs can be identified from all possible object-pairs. We present experimental results from a dual camera human face capturing system in which more than 99% object correspondences can be accurately identified, while 100% of falsely detected objects are eliminated. Besides, our proposed method can handle no less than 100 object-pairs within 1 ms in a P4 1.5 GHz desktop PC.
Fai Chan, Jiansheng Chen 0001, Yiu Sang Moon
ICASSP2
2008 Using SIFT features in palmprint authentication
abstract
As a new branch of biometrics, palmprint authentication has attracted increasing amount of attention because palmprints are abundant of line features so that low resolution images can be used. In this paper, we present two novel approaches for palmprints authentication. Firstly, we employ the SIFT (Scale Invariant Feature Transformation) for palmprint authentication. Point-wise matching is used to match SIFT key points extracted form palmprint images. Secondly, we extend a time series technology, SAX (Symbolic Aggregate approximation), to 2D data for the palmprint representation and matching. Using a public palmprint database, we demonstrate that the two proposed approaches, when combined together, can achieve the palmprint authentication accuracy comparable to that of the state of the art algorithms.
Jiansheng Chen 0001, Yiu Sang Moon
ICPR1
2007 A Minutiae-based Fingerprint Individuality Model
abstract
Fingerprint individuality study deals with the crucial problem of the discriminative power of fingerprints for recognizing people. In this paper, we present a novel fingerprint individuality model based on minutiae, the most commonly used fingerprint feature. The probability of the false correspondence among fingerprints from different fingers is calculated by combining the distinctiveness of the spatial locations and directions of the minutiae. To validate our model, experiments were performed using different fingerprint databases. The matching score distribution predicted by our model actually fits the observed experimental results satisfactorily. Comparing to most previous fingerprint individuality models, our model makes more reasonably conservative estimate of the fingerprint discriminative power, making it a powerful tool for studying the fingerprint individuality as well as the performance evaluation of fingerprint verification systems.
Jiansheng Chen 0001, Yiu Sang Moon
CVPR1
2006 A Statistical Study on the Fingerprint Minutiae Distribution
abstract
Fingerprint minutiae distribution is the key issue of fingerprint individuality study. A method for studying fingerprint minutiae distribution by analyzing their second order statistical properties is proposed in this paper. Experiments have been performed on 467 different fingerprints selected from three major fingerprint databases. Results show that fingerprint minutiae tend to overdisperse on a small scale; and cluster on a large scale. Our findings which have successfully explained and unified various previous research observations should enlighten the study of fingerprint minutiae pattern modeling, an important foundation for boosting improvement in the fingerprint authentication technology.
Jiansheng Chen 0001, Yiu Sang Moon
ICASSP (2)1