VLDB 2026 Research / reviewers in the wild / expert
Xiangyang Xue 0001
dblp:84/3791
· DBLP profile ↗
319ranked-venue papers
2as first author
142since 2021 · last 2026
0000-0002-4897-9209ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 195 · 1 first-author · 99 since 2021Graphics, computer vision, multimedia, augmented reality and games · 188 · 2 first-author · 79 since 2021Databases, data management, data science and information retrieval · 17 · 3 since 2021Systems, architecture and hardware · 13 · 10 since 2021Computer networks · 11 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 8 since 2021Security and privacy · 4 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Theory of computation · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LOG-Nav: Efficient Layout-Aware Object-Goal Navigation with Hierarchical PlanningabstractWe introduce LOG-Nav, an efficient layout-aware object-goal navigation approach designed for complex multi-room indoor environments. By planning hierarchically leveraging a global topologigal map with layout information and local imperative approach with detailed scene representation memory, LOG-Nav achieves both efficient and effective navigation. The process is managed by an LLM-powered agent, ensuring seamless effective planning and navigation, without the need for human interaction, complex rewards, or costly training. Our experimental results on the MP3D benchmark achieves 85% object navigation success rate (SR) and 79% success rate weighted by path length (SPL) (over 40% point improvement in SR and 60% improvement in SPL compared to exsisting methods). Furthermore, we validate the robustness of our approach through virtual agent and real-world robotic deployment, showcasing its capability in practical scenarios. Xiangyang Xue 0001, Taiping Zeng |
AAAI | 3 |
| 2026 | From Intent to Execution: Multimodal Chain-of-Thought Reinforcement Learning for Precise CAD Code GenerationabstractComputer-Aided Design (CAD) plays a vital role in engineering and manufacturing, yet current CAD workflows require extensive domain expertise and manual modeling effort. Recent advances in large language models (LLMs) have made it possible to generate code from natural language, opening new opportunities for automating parametric 3D modeling. However, directly translating human design intent into executable CAD code remains highly challenging, due to the need for logical reasoning, syntactic correctness, and numerical precision. In this work, we propose CAD-RL, a multimodal Chain-of-Thought (CoT) guided reinforcement learning post training framework for CAD modeling code generation. Our method combines CoT-based Cold Start with goal-driven reinforcement learning post training using three task-specific rewards: executability reward, geometric accuracy reward, and external evaluation reward. To ensure stable policy learning under sparse and high-variance reward conditions, we introduce three targeted optimization strategies: Trust Region Stretch for improved exploration, Precision Token Loss for enhanced dimensions parameter accuracy, and Overlong Filtering to reduce noisy supervision. To support training and benchmarking, we release ExeCAD, a noval dataset comprising 16,540 real-world CAD examples with paired natural language and structured design language descriptions, executable CADQuery scripts, and rendered 3D models. Experiments demonstrate that CAD-RL achieves significant improvements in reasoning quality, output precision, and code executability over existing VLMs. Ke Niu 0004, Haiyang Yu 0004, Mengyang Zhao 0002, Teng Fu 0001, Bin Li 0015, Xiangyang Xue 0001 |
AAAI | 7 |
| 2026 | SCoUT: A Framework for Structured Stereotype Analysis in Language ModelsabstractExisting stereotype auditing methods for Large Language Models (LLMs) typically rely on isolated rating schemes or task-specific probes, lacking theoretical grounding and failing to reveal internal organization beyond surface-level output patterns. In this paper, we introduce SCoUT (Stereotype Content-oriented Utility structure via Thurstonian modeling), a closed-loop framework that structurally models, explicitly probes, and functionally steers stereotype dimensions (warmth and competence) in LLMs. SCoUT first reconstructs a global stereotype utility structure aligned with Stereotype Content Model theory via Thurstonian comparative judgments. Across multiple open-source LLMs, this modeling achieves high pairwise-preference prediction accuracy (≥ 0.90 on larger-scale models) and exhibits strong cross-model consistency. Probing internal attention mechanisms localizes this structure to specific heads (Spearman’s ρ up to 0.83 for warmth and 0.90 for competence) and surfaces a salient asymmetry between warmth and competence. Further, targeted inference-time activation modifications on these dimension-sensitive heads consistently steer model outputs along the intended axes. By bridging behavioral measurement with internal representation and controllable steering, SCoUT offers an end-to-end framework that uncovers and interprets the latent structure of stereotypes, advancing stereotype auditing from surface detection to structural analysis. Jinxuan Wu, Bin Li 0015, Xiangyang Xue 0001 |
AAAI | 3 |
| 2026 | Temporal-filter enhanced prediction of whole-brain neural activity using the physiology-aligned latent variable model
Chong Li 0007, Jianfeng Ma 0001, Xiangyang Xue 0001, Yuguo Yu |
Neurocomputing | 4 |
| 2026 | FDGReID: Federated Domain Generalization for Person Re-identification
Ke Niu 0004, Haiyang Yu 0004, Teng Fu 0001, Mengyang Zhao 0002, Bin Li 0015, Xuelin Qian, Xiangyang Xue 0001 |
Mach. Learn. | 7 |
| 2026 | Low-Light Scene Text Image Enhancement in the Wild
Haiyang Yu 0004, Yinglian Zhu, Ke Niu 0004, Xiangyang Xue 0001, Bin Li 0015 |
Mach. Learn. | 5 |
| 2026 | Abstracting Concept-Changing Rules for Solving Raven's Progressive Matrix Problems
Bin Li 0015, Xiangyang Xue 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2026 | GS-Net: Point cloud sampling with graph neural networks
Xiaolei Chen 0002, Shoumeng Qiu, Xiangyang Xue 0001, Jian Pu |
Pattern Recognit. | 4 |
| 2026 | Instruction-guided fusion of multi-layer visual features in Large Vision-Language Models
Xu Li 0016, Yi Zheng 0003, Haotian Chen 0003, Xiaolei Chen 0002, Yuxuan Liang 0004, Chenghang Lai, Bin Li 0015, Xiangyang Xue 0001 |
Pattern Recognit. | 8 |
| 2026 | Pyramid Token Pruning for High-Resolution Large Vision-Language Models via Region, Token, and Instruction-Guided Importance
Yuxuan Liang 0004, Xu Li 0016, Xiaolei Chen 0002, Haotian Chen 0003, Yi Zheng 0003, Rui Zhu 0013, Bin Li 0015, Xiangyang Xue 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 9 |
| 2026 | Unleashing the Potential of Tracklets for Unsupervised Video Person Re-Identification
Nanxing Meng, Qizao Wang, Bin Li 0015, Xiangyang Xue 0001 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2026 | A New Semi-Supervised Video Anomaly Detection Baseline in Lack of Anomalous SamplesabstractVideo anomaly detection (VAD) has been widely studied for its important applications in multimedia community. Recently, many Weakly Supervised VAD (WS-VAD) methods have been proposed, which tend to treat VAD as a classification task through multiple instance learning and result in the need to collect sufficient anomaly classes and samples to be used for training a classifier. However, anomaly events tend to be open-set and rare in real-world applications, so we often have difficulty collecting all anomaly classes and enough sample anomalies, which is a difficult situation for WS-VAD to cope with. To this end, we consider to treat VAD as an out-of-distribution detection task rather than a classification task and propose a simple but effective semi-supervised baseline method. First, we leverage the powerful zero-shot capability of large visual language models to generate summary text descriptions for videos and extract visual features as intermediates for subsequent use. Next, we use a text encoder to extract language features and combine them with visual features to obtain robust multimodal features. Finally, we introduce an out-of-distribution detection method that learns the center of normality in multimodal space from normal and unlabeled samples, while deviating abnormal samples from the center to cope with the scarcity of abnormal samples. To implement our baseline method, we also provide a new semi-supervised dataset by reorganizing an existing benchmark, which is the first available dataset in the VAD community that provides trimmed videos consisting of complete abnormal events. Experiments demonstrate that our method performs more robustly when fewer anomaly classes and anomaly samples collected. Mengyang Zhao 0002, Haiyang Yu 0004, Teng Fu 0001, Yang Liu 0246, Wei Zhou 0021, Bin Li 0015, Xiangyang Xue 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 7 |
| 2026 | Natural Language to Code for Automated Annotation in Autonomous DrivingabstractThe fast expansion of deep learning models has led to an increasing need for well-annotated datasets, while traditional manual annotation cannot meet this requirement. Current research on annotation mainly focuses on automating the annotation process. These studies typically rely on a set of predefined functionalities. However, in complex scenarios, for example, autonomous driving, annotation workflow, and postprocessing functions must be tailored to specific tasks. The challenge here lies in ensuring that newly generated functions integrate with the existing function set of the system, which requires the pipeline to understand the user requirement and real-time system context to generate appropriate input-output data structures. Previous annotation methods relied on human programming to meet this requirement. This dependency on professional assistance restricts the generalizability of annotation methods. Drawing on modern software engineering principles, we introduce an interactive code generation pipeline based on natural language input to address this challenge. Our approach supports real-time code generation by natural language input. To the best of our knowledge, this is one of the latest applications that apply customizable functional extensions in the annotation pipeline. Our evaluations on public datasets and a self-built real-world dataset demonstrate that our method significantly enhances the range of application scenarios for annotation tools while reducing manual intervention. Upon acceptance, the code will be open source. Aoxiang Qin, Xiangyang Xue 0001, Jian Pu |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2025 | Foundation Model Driven Appearance Extraction for Robust Multiple Object TrackingabstractMultiple Object Tracking (MOT) is a fundamental task in computer vision. Existing methods utilize motion information or appearance information to perform object tracking. However, these algorithms still struggle with special circumstances, such as occlusion and blurring in complex scenes. Inspired by the fact that people can pinpoint objects through verbal descriptions, we explore performing long-term robust tracking using semantic features of objects. Motivated by the success of the multimodal foundation model in text-image alignment, we reconsider the appearance feature extraction module in MOT and propose a Foundation model Driven multi-object tracker (FDTracker). Specifically, we propose a two-stage trained appearance feature extractor. In the first stage, using a single image of the object as input, the model could capture the attributes of objects with the assistance of natural language instructions. In the second stage, using a sequence of images of objects as input, the model learns how to use these attributes to distinguish between different objects and connect the same object at different times. Finally, for coordinating appearance and motion information, we propose a reasonable combined strategy, which better facilitates trajectory assignment and reconnection. Extensive experiments on benchmarks demonstrate the robustness of FDTracker. Teng Fu 0001, Haiyang Yu 0004, Ke Niu 0004, Bin Li 0015, Xiangyang Xue 0001 |
AAAI | 5 |
| 2025 | PC-BEV: An Efficient Polar-Cartesian BEV Fusion Framework for LiDAR Semantic SegmentationabstractAlthough multiview fusion has demonstrated potential in LiDAR segmentation, its dependence on computationally intensive point-based interactions, arising from the lack of fixed correspondences between views such as range view and Bird's-Eye View (BEV), hinders its practical deployment. This paper challenges the prevailing notion that multiview fusion is essential for achieving high performance. We demonstrate that significant gains can be realized by directly fusing Polar and Cartesian partitioning strategies within the BEV space. Our proposed BEV-only segmentation model leverages the inherent fixed grid correspondences between these partitioning schemes, enabling a fusion process that is orders of magnitude faster (170x speedup) than conventional point-based methods. Furthermore, our approach facilitates dense feature fusion, preserving richer contextual information compared to sparse point-based alternatives. To enhance scene understanding while maintaining inference efficiency, we also introduce a hybrid Transformer-CNN architecture. Extensive evaluation on the SemanticKITTI and nuScenes datasets provides compelling evidence that our method outperforms previous multiview fusion approaches in terms of both performance and inference speed, highlighting the potential of BEV-based fusion for LiDAR segmentation. Shoumeng Qiu, Xinrun Li, Xiangyang Xue 0001, Jian Pu |
AAAI | 3 |
| 2025 | Towards a Vision-Language Episodic Memory Framework: Large-scale Pretrained Model-Augmented Hippocampal Attractor Dynamics
Chong Li 0007, Taiping Zeng, Xiangyang Xue 0001, Jianfeng Feng |
CogSci | 3 |
| 2025 | Towards Enhanced Image Inpainting: Mitigating Unwanted Object Insertion and Preserving Color ConsistencyabstractRecent advances in image inpainting increasingly use generative models to handle large irregular masks. However, these models can create unrealistic inpainted images due to two main issues: (1) Unwanted object insertion: Even with unmasked areas as context, generative models may still generate arbitrary objects in the masked region that don’t align with the rest of the image. (2) Color inconsistency: Inpainted regions often have color shifts that causes a smeared appearance, reducing image quality. Retraining the generative model could help solve these issues, but it’s costly since state-of-the-art latent-based diffusion and rectified flow models require a three-stage training process: training a VAE, training a generative U-Net or transformer, and fine-tuning for inpainting. Instead, this paper proposes a post-processing approach, dubbed as ASUKA (Aligned Stable inpainting with UnKnown Areas prior), to improve inpainting models. To address unwanted object insertion, we leverage a Masked Auto-Encoder (MAE) for reconstruction-based priors. This mitigates object hallucination while maintaining the model’s generation capabilities. To address color inconsistency, we propose a specialized VAE decoder that treats latent-to-image decoding as a local harmonization task, significantly reducing color shifts for color-consistent inpainting. We validate ASUKA on SD 1.5 and FLUX inpainting variants with Places2 and MISATO, our proposed diverse collection of datasets. Results show that ASUKA mitigates object hallucination and improves color consistency over standard diffusion and rectified flow models and other inpainting methods. Yikai Wang 0002, Chenjie Cao, Junqiu Yu, Xiangyang Xue 0001, Yanwei Fu 0001 |
CVPR | 5 |
| 2025 | MVGenMaster: Scaling Multi-View Generation from Any Image via 3D Priors Enhanced Diffusion ModelabstractWe introduce MVGenMaster, a multi-view diffusion model enhanced with 3D priors to address versatile Novel View Synthesis (NVS) tasks. MVGenMaster leverages 3D priors that are warped using metric depth and camera poses, significantly enhancing both generalization and 3D consistency in NVS. Our model features a simple yet effective pipeline that can generate up to 100 novel views conditioned on variable reference views and camera poses with a single forward process. Additionally, we have developed a comprehensive large-scale multi-view image dataset called MvD-1M, comprising up to 1.6 million scenes, equipped with well-aligned metric depth to train MVGenMaster. Moreover, we present several training and model modifications to strengthen the model with scaled-up datasets. Extensive evaluations across in- and out- of- domain benchmarks demonstrate the effectiveness of our proposed method and data formulation. Chenjie Cao, Chaohui Yu, Shang Liu 0002, Fan Wang 0019, Xiangyang Xue 0001, Yanwei Fu 0001 |
CVPR | 5 |
| 2025 | CAP-Net: A Unified Network for 6D Pose and Size Estimation of Categorical Articulated Parts from a Single RGB-D ImageabstractThis paper tackles category-level pose estimation of articulated objects in robotic manipulation tasks and introduces a new benchmark dataset. While recent methods estimate part poses and sizes at the category level, they often rely on geometric cues and complex multi-stage pipelines that first segment parts from the point cloud, followed by Normalized Part Coordinate Space (NPCS) estimation for 6D poses. These approaches overlook dense semantic cues from RGB images, leading to suboptimal accuracy, particularly for objects with small parts. To address these limitations, we propose a single-stage Network, CAP-Net, for estimating the 6D poses and sizes of Categorical Articulated Parts. This method combines RGB-D features to generate instance segmentation and NPCS representations for each part in an end-to-end manner. CAP-Net uses a unified network to simultaneously predict point-wise class labels, centroid offsets, and NPCS maps. A clustering algorithm then groups points of the same predicted class based on their estimated centroid distances to isolate each part. Finally, the NPCS region of each part is aligned with the point cloud to recover its final pose and size. To bridge the sim-to-real domain gap, we introduce the RGBD-Art dataset, the largest RGB-D articulated dataset to date, featuring photorealistic RGB images and depth noise simulated from real sensors. Experimental evaluations on the RGBD-Art dataset demonstrate that our method significantly outperforms the state-of-the-art approach. Real-world deployments of our model in robotic tasks underscore its robustness and exceptional sim-to-real transfer capabilities, confirming its substantial practical utility. Our dataset, code and pre-trained models are available on the project page2. Jingshun Huang, Yanwei Fu 0001, Xiangyang Xue 0001, Yi Zhu 0001 |
CVPR | 5 |
| 2025 | ReasonGrounder: LVLM-Guided Hierarchical Feature Splatting for Open-Vocabulary 3D Visual Grounding and ReasoningabstractOpen-vocabulary 3D visual grounding and reasoning aim to localize objects in a scene based on implicit language descriptions, even when they are occluded. This ability is crucial for tasks such as vision-language navigation and autonomous robotics. However, current methods struggle because they rely heavily on fine-tuning with 3D annotations and mask proposals, which limits their ability to handle diverse semantics and common knowledge required for effective reasoning. In this work, we propose ReasonGrounder, an LVLM-guided framework that uses hierarchical 3D feature Gaussian fields for adaptive grouping based on physical scale, enabling open-vocabulary 3D grounding and reasoning. ReasonGrounder interprets implicit instructions using large vision-language models (LVLM) and localizes occluded objects through 3D Gaussian splatting. By incorporating 2D segmentation masks from the SAM and multi-view CLIP embeddings, ReasonGrounder selects Gaussian groups based on object scale, enabling accurate localization through both explicit and implicit language understanding, even in novel, occluded views. We also contribute ReasoningGD, a new dataset containing over 10K scenes and 2 million annotations for evaluating open-vocabulary 3D grounding and amodal perception under occlusion. Experiments show that ReasonGrounder significantly improves 3D grounding accuracy in real-world scenarios. Zhenyang Liu, Yikai Wang 0002, Sixiao Zheng, Tongying Pan, Longfei Liang, Yanwei Fu 0001, Xiangyang Xue 0001 |
CVPR | 7 |
| 2025 | CocoER: Aligning Multi-Level Feature by Competition and Coordination for Emotion RecognitionabstractWith the explosion of human-machine interaction, emotion recognition has reignited attention. Previous works focus on improving visual feature fusion and reasoning from multiple image levels. Although it is non-trivial to deduce a person’s emotion by integrating multi-level feature (head, body and context), the emotion recognition results of each level is usually different from one another, which creates inconsistency in the prevailing feature alignment method and decrease recognition performance. In this work, we propose a multi-level image feature refinement method for emotion recognition (CocoER) to mitigate the impact caused by conflicting results from multi-level recognition. First, we leverage cross-level attention to improve visual feature consistency between hierarchically cropped head, body and context windows. Then, vocabulary informed alignment is incorporated into the recognition framework to produce pseudo label and guide hierarchical visual feature refinement. To effectively fuse multi-level feature, we elaborate on a competition process of eliminating irrelevant image level predictions and a coordination process to enhance the feature across all levels. Extensive experiments are executed on two popular datasets, and our method achieves state-of-the-art performance with multi-level interpretation results. Code is available at: https://github.com/bisno/CocoER. Xuli Shen, Hua Cai, Weilin Shen, Qing Xu 0017, Dingding Yu, Weifeng Ge, Xiangyang Xue 0001 |
CVPR | 7 |
| 2025 | Content and Salient Semantics Collaboration for Cloth-Changing Person Re-IdentificationabstractCloth-changing person re-identification aims at recognizing the same person with clothing changes across non-overlapping cameras. Advanced methods either resort to identity-related auxiliary modalities (e.g., sketches, silhouettes, and keypoints) or clothing labels to mitigate the impact of clothes. However, relying on unpractical and inflexible auxiliary modalities or annotations limits their real-world applicability. In this paper, we promote cloth-changing person re-identification by leveraging abundant semantics present within pedestrian images, without the need for any auxiliaries. Specifically, we first propose a unified Semantics Mining and Refinement (SMR) module to extract robust identity-related content and salient semantics, mitigating interference from clothing appearances effectively. We further propose the Content and Salient Semantics Collaboration (CSSC) framework to collaborate and leverage various semantics, facilitating cross-parallel semantic interaction and refinement. Our proposed method achieves state-of-the-art performance on three cloth-changing benchmarks, demonstrating its superiority over advanced competitors. The code is available at https://github.com/QizaoWang/CSSC-CCReID. Qizao Wang, Xuelin Qian, Bin Li 0015, Lifeng Chen, Yanwei Fu 0001, Xiangyang Xue 0001 |
ICASSP | 6 |
| 2025 | Spatial-Temporal Aware Visuomotor Diffusion Policy LearningabstractVisual imitation learning is effective for robots to learn versatile tasks. However, many existing methods rely on behavior cloning with supervised historical trajectories, limiting their 3D spatial and 4D spatiotemporal awareness. Consequently, these methods struggle to capture the 3D structures and 4D spatiotemporal relationships necessary for real-world deployment. In this work, we propose 4D Diffusion Policy (DP4), a novel visual imitation learning method that incorporates spatiotemporal awareness into diffusion-based policies. Unlike traditional approaches that rely on trajectory cloning, DP4 leverages a dynamic Gaussian world model to guide the learning of 3D spatial and 4D spatiotemporal perceptions from interactive environments. Our method constructs the current 3D scene from a single-view RGB-D observation and predicts the future 3D scene, optimizing trajectory generation by explicitly modeling both spatial and temporal dependencies. Extensive experiments across 17 simulation tasks with 173 variants and 3 real-world robotic tasks demonstrate that the 4D Diffusion Policy (DP4) outperforms baseline methods, improving the average simulation task success rate by 16.4% (Adroit), 14% (DexArt), and 6.45% (RLBench), and the average real-world robotic task success rate by 8.6%. Zhenyang Liu, Yikai Wang 0002, Kuanning Wang, Longfei Liang, Xiangyang Xue 0001, Yanwei Fu 0001 |
ICCV | 5 |
| 2025 | ChatReID: Open-Ended Interactive Person Retrieval via Hierarchical Progressive Tuning for Vision Language ModelsabstractPerson re-identification (Re-ID) is a crucial task in computer vision, aiming to recognize individuals across non-overlapping camera views. While recent advanced vision-language models (VLMs) excel in logical reasoning and multi-task generalization, their applications in Re-ID tasks remain limited. They either struggle to perform accurate matching based on identity-relevant features or assist image-dominated branches as auxiliary semantics. In this paper, we propose a novel framework ChatReID, that shifts the focus towards a text-side-dominated retrieval paradigm, enabling flexible and interactive re-identification. To integrate the reasoning abilities of language models into Re-ID pipelines, We first present a large-scale instruction dataset, which contains more than 8 million prompts to promote the model fine-tuning. Next. we introduce a hierarchical progressive tuning strategy, which endows Re-ID ability through three stages of tuning, i.e., from person attribute understanding to fine-grained image retrieval and to multi-modal task reasoning. Extensive experiments across ten popular benchmarks demonstrate that ChatReID outperforms existing methods, achieving state-of-the-art performance in all Re-ID tasks. More experiments demonstrate that ChatReID not only has the ability to recognize fine-grained details but also to integrate them into a coherent reasoning process. Ke Niu 0004, Haiyang Yu 0004, Mengyang Zhao 0002, Teng Fu 0001, Siyang Yi, Bin Li 0015, Xuelin Qian, Xiangyang Xue 0001 |
ICCV | 9 |
| 2025 | From Sky to Site: A Unified Framework for Static and Dynamic 3D Reconstruction in Construction Sites
Tonglin Chen, Xinlin Ren, Jiangyu Feng, Bin Li 0015, Xiangyang Xue 0001 |
ICIG (2) | 6 |
| 2025 | EmoHead: Emotional Talking Head via Manipulating Semantic Expression ParametersabstractGenerating emotion-specific talking head videos from audio input is an important and complex challenge for human-machine interaction. However, emotion is highly abstract concept with ambiguous boundaries, and it necessitates disentangled expression parameters to generate emotionally expressive talking head videos. In this work, we present EmoHead to synthesize talking head videos via semantic expression parameters. To predict expression parameter for arbitrary audio input, we apply an audio-expression module that can be specified by an emotion tag. This module aims to enhance correlation from audio input across various emotions. Furthermore, we leverage pre-trained hyperplane to refine facial movements by probing along the vertical direction. Finally, the refined expression parameters regularize neural radiance fields and facilitate the emotion-consistent generation of talking head videos. Experimental results demonstrate that semantic expression parameters lead to better reconstruction quality and controllability. Xuli Shen, Hua Cai, Dingding Yu, Weilin Shen, Qing Xu 0017, Xiangyang Xue 0001 |
ICME | 6 |
| 2025 | Beyond Task-Specific Reasoning: A Unified Conditional Generative Framework for Abstract Visual ReasoningabstractAbstract visual reasoning (AVR) enables humans to quickly discover and generalize abstract rules to new scenarios. Designing intelligent systems with human-like AVR abilities has been a long-standing topic in the artificial intelligence community. Deep AVR solvers have recently achieved remarkable success in various AVR tasks. However, they usually use task-specific designs or parameters in different tasks. In such a paradigm, solving new tasks often means retraining the model, and sometimes retuning the model architectures, which increases the cost of solving AVR problems. In contrast to task-specific approaches, this paper proposes a novel Unified Conditional Generative Solver (UCGS), aiming to address multiple AVR tasks in a unified framework. First, we prove that some well-known AVR tasks can be reformulated as the problem of estimating the predictability of target images in problem panels. Then, we illustrate that, under the proposed framework, training one conditional generative model can solve various AVR tasks. The experiments show that with a single round of multi-task training, UCGS demonstrates abstract reasoning ability across various AVR tasks. Especially, UCGS exhibits the ability of zero-shot reasoning, enabling it to perform abstract reasoning on problems from unseen AVR tasks in the testing phase. Bin Li 0015, Xiangyang Xue 0001 |
ICML | 3 |
| 2025 | One-Shot Heterogeneous Federated Learning with Local Model-Guided Diffusion ModelsabstractIn recent years, One-shot Federated Learning (OSFL) methods based on Diffusion Models (DMs) have garnered increasing attention due to their remarkable performance. However, most of these methods require the deployment of foundation models on client devices, which significantly raises the computational requirements and reduces their adaptability to heterogeneous client models. In this paper, we propose FedLMG, a heterogeneous one-shot Federated learning method with Local Model-Guided diffusion models. In our method, clients do not need access to any foundation models but only train and upload their local models, which is consistent with traditional FL methods. On the clients, we employ classification loss and batch normalization loss to capture the broad category features and detailed contextual features of the client distributions. On the server, based on the uploaded client models, we utilize backpropagation to guide the server’s DM in generating synthetic datasets that comply with the client distributions, which are then used to train the aggregated model. By using the local models as a medium to transfer client knowledge, our method significantly reduces the computational requirements on client devices and effectively adapts to scenarios with heterogeneous clients. Extensive quantitation and visualization experiments on three large-scale real-world datasets, along with theoretical analysis, demonstrate that the synthetic datasets generated by FedLMG exhibit comparable quality and diversity to the client datasets, which leads to an aggregated model that outperforms all compared methods and even the performance ceiling, further elucidating the significant potential of utilizing DMs in FL. Mingzhao Yang, Shangchao Su, Bin Li 0015, Xiangyang Xue 0001 |
ICML | 4 |
| 2025 | You Only Estimate Once: Unified, One-stage, Real-Time Category-Level Articulated Object 6D Pose Estimation for Robotic GraspingabstractThis paper addresses the problem of category-level pose estimation for articulated objects in robotic manipulation tasks. Recent works have shown promising results in estimating part pose and size at the category level. However, these approaches primarily follow a complex multi-stage pipeline that first segments part instances in the point cloud and then estimates the Normalized Part Coordinate Space (NPCS) representation for 6D poses. These approaches suffer from high computational costs and low performance in real-time robotic tasks. To address these limitations, we propose YOEO, a single-stage method that simultaneously outputs instance segmentation and NPCS representations in an end-to-end manner. We use a unified network to generate point-wise semantic labels and centroid offsets, allowing points from the same part instance to vote for the same centroid. We further utilize a clustering algorithm to distinguish points based on their estimated centroid distances. Finally, we first separate the NPCS region of each instance. Then, we align the separated regions with the real point cloud to recover the final pose and size. Experimental results on the GAPart dataset demonstrate the pose estimation capabilities of our proposed single-shot method. We also deploy our synthetically-trained model in a real-world setting, providing real-time visual feedback at 200Hz, enabling a physical Kinova robot to interact with unseen articulated objects. This showcases the utility and effectiveness of our proposed method2. Jingshun Huang, Yanwei Fu 0001, Yu-Gang Jiang 0001, Xiangyang Xue 0001 |
ICRA | 6 |
| 2025 | RAG-6DPose: Retrieval-Augmented 6D Pose Estimation via Leveraging CAD as Knowledge BaseabstractAccurate 6D pose estimation is key for robotic manipulation, enabling precise object localization for tasks like grasping. We present RAG-6DPose, a retrieval-augmented approach that leverages 3D CAD models as a knowledge base by integrating both visual and geometric cues. Our RAG-6DPose roughly contains three stages: 1) Building a Multi-Modal CAD Knowledge Base by extracting 2D visual features from multi-view CAD rendered images and also attaching 3D points; 2) Retrieving relevant CAD features from the knowledge base based on the current query image via our ReSPC module; and 3) Incorporating retrieved CAD information to refine pose predictions via retrieval-augmented decoding. Experimental results on standard benchmarks and real-world robotic tasks demonstrate the effectiveness and robustness of our approach, particularly in handling occlusions and novel viewpoints. Supplementary material is available on our project website: https://sressers.github.io/RAG-6DPose. Kuanning Wang, Yuqian Fu, Yanwei Fu 0001, Longfei Liang, Yu-Gang Jiang 0001, Xiangyang Xue 0001 |
IROS | 7 |
| 2025 | Ali-UI: Enhancing Complex Vision-Language Navigation with Alignment of Unified Map and Instruction ParsingabstractVisual language navigation (VLN) poses challenges in guiding agents through unseen environments based on natural language instructions. Existing methods either rely on imitation learning, for which training across various complex scenarios remains challenging, or leverage large visual language models (LVLMs) for zero-shot object recognition and expert iterative reasoning for improved scene understanding. Although LVLMs enhance target detection generalization, current VLN methods lack robustness in terms of environmental generalization and struggle with multi-step, coarsely directed instructions. Addressing these challenges, we introduce Ali-UI, a novel vision-language navigation approach that enables agents to navigate from random starting points in unvisited scenes and handle complex multi-step instructions. Specifically, we incorporate continuously accumulating global grid maps and local semantic maps as scene memory by employing frontier-based exploration. Multi-step coarsely directed commands are broken down with the assistance of LLaVA and matched with the scene, considering temporal and spatial alignment. Panoramic data are saved in topological form and queried by instruction segments for sequential navigation. Extensive experiments carried out in simulated environments demonstrate that Ali-UI outperforms existing state-of-the-art methods in terms of flexible human instructions and scene generalization, with the success rate improved by 23.37% and the SPL increased by 19.27% in R2R dataset. Da Huang 0006, Yanwei Fu 0001, Xiangyang Xue 0001 |
ACM Multimedia | 5 |
| 2025 | A Neural Representation Framework with LLM-Driven Spatial Reasoning for Open-Vocabulary 3D Visual GroundingabstractOpen-vocabulary 3D visual grounding aims to localize target objects based on free-form language queries, which is crucial for embodied AI applications such as autonomous navigation, robotics, and augmented reality. Learning 3D language fields through neural representations enables accurate understanding of 3D scenes from limited viewpoints and facilitates the localization of target objects in complex environments. However, existing language field methods struggle to accurately localize instances using spatial relations in language queries, such as ''the book on the chair.'' This limitation mainly arises from inadequate reasoning about spatial relations in both language queries and 3D scenes. In this work, we propose SpatialReasoner, a novel neural representation-based framework with large language model (LLM)-driven spatial reasoning that constructs a visual properties-enhanced hierarchical feature field for open-vocabulary 3D visual grounding. To enable spatial reasoning in language queries, SpatialReasoner fine-tunes an LLM to capture spatial relations and explicitly infer instructions for the target, anchor, and spatial relation. To enable spatial reasoning in 3D scenes, SpatialReasoner incorporates visual properties (opacity and color) to construct a hierarchical feature field. This field represents language and instance features using distilled CLIP features and masks extracted via the Segment Anything Model (SAM). The field is then queried using the inferred instructions in a hierarchical manner to localize the target 3D instance based on the spatial relation in the language query. Notably, SpatialReasoner is not limited to a specific 3D neural representation; it serves as a framework adaptable to various representations, such as Neural Radiance Fields (NeRF) or 3D Gaussian Splatting (3DGS). Extensive experiments show that our framework can be seamlessly integrated into different neural representations, outperforming baseline models in 3D visual grounding while empowering their spatial reasoning capability. Project Homepage:ZhenyangLiu.github.io/SpatialReasoner. Zhenyang Liu, Sixiao Zheng, Siyu Chen 0023, Cairong Zhao, Longfei Liang, Xiangyang Xue 0001, Yanwei Fu 0001 |
ACM Multimedia | 6 |
| 2025 | TP-MDDN: Task-Preferenced Multi-Demand-Driven Navigation with Autonomous Decision-MakingabstractIn daily life, people often move through spaces to find objects that meet their needs, posing a key challenge in embodied AI. Traditional Demand-Driven Navigation (DDN) handles one need at a time but does not reflect the complexity of real-world tasks involving multiple needs and personal choices. To bridge this gap, we introduce Task-Preferenced Multi-Demand-Driven Navigation (TP-MDDN), a new benchmark for long-horizon navigation involving multiple sub-demands with explicit task preferences. To solve TP-MDDN, we propose AWMSystem, an autonomous decision-making system composed of three key modules: BreakLLM (instruction decomposition), LocateLLM (goal selection), and StatusMLLM (task monitoring). For spatial memory, we design MASMap, which combines 3D point cloud accumulation with 2D semantic mapping for accurate and efficient environmental understanding. Our Dual-Tempo action generation framework integrates zero-shot planning with policy-based fine control, and is further supported by an Adaptive Error Corrector that handles failure cases in real time. Experiments demonstrate that our approach outperforms state-of-the-art baselines in both perception accuracy and navigation robustness. Yanwei Fu 0001, Yu-Gang Jiang 0001, Xiangyang Xue 0001 |
NeurIPS | 6 |
| 2025 | CReFT-CAD: Boosting Orthographic Projection Reasoning for CAD via Reinforcement Fine-TuningabstractComputer-Aided Design (CAD) is pivotal in industrial manufacturing, with orthographic projection reasoning foundational to its entire workflow—encompassing design, manufacturing, and simulation. However, prevailing deep-learning approaches employ standard 3D reconstruction pipelines as an alternative, which often introduce imprecise dimensions and limit the parametric editability required for CAD workflows. Recently, some researchers adopt vision–language models (VLMs), particularly supervised fine-tuning (SFT), to tackle CAD-related challenges. SFT shows promise but often devolves into pattern memorization, resulting in poor out-of-distribution (OOD) performance on complex reasoning tasks. To tackle these limitations, we introduce CReFT-CAD, a two-stage fine-tuning paradigm: first, a curriculum-driven reinforcement learning stage with difficulty-aware rewards to steadily build reasoning abilities; second, supervised post-tuning to refine instruction following and semantic extraction. Complementing this, we release TriView2CAD, the first large-scale, open-source benchmark for orthographic projection reasoning, comprising 200,000 synthetic and 3,000 real-world orthographic projections with precise dimensional annotations and six interoperable data modalities. Benchmarking leading VLMs on orthographic projection reasoning, we show that CReFT-CAD significantly improves reasoning accuracy and OOD generalizability in real-world scenarios, providing valuable insights to advance CAD reasoning research. The code and adopted datasets are available at \url{https://github.com/KeNiu042/CReFT-CAD}. Ke Niu 0004, Haiyang Yu 0004, Teng Fu 0001, Mengyang Zhao 0002, Bin Li 0015, Xiangyang Xue 0001 |
NeurIPS | 8 |
| 2025 | Uni3C: Unifying Precisely 3D-Enhanced Camera and Human Motion Controls for Video GenerationabstractCamera and human motion controls have been extensively studied for video generation, but existing approaches typically address them separately, suffering from limited data with high-quality annotations for both aspects. To overcome this, we present Uni3C, a unified 3D-enhanced framework for precise control of both camera and human motion in video generation. Uni3C includes two key contributions. First, we propose a plug-and-play control module trained with a frozen video generative backbone, PCDController, which utilizes unprojected point clouds from monocular depth to achieve accurate camera control. By leveraging the strong 3D priors of point clouds and the powerful capacities of video foundational models, PCDController shows impressive generalization, performing well regardless of whether the inference backbone is frozen or fine-tuned. This flexibility enables different modules of Uni3C to be trained in specific domains, i.e., either camera control or human motion control, reducing the dependency on jointly annotated data. Second, we propose a jointly aligned 3D world guidance for the inference phase that seamlessly integrates both scenic point clouds and SMPL-X characters to unify the control signals for camera and human motion, respectively. Extensive experiments confirm that PCDController enjoys strong robustness in driving camera motion for fine-tuned backbones of video generation. Uni3C substantially outperforms competitors in both camera controllability and human motion quality. Additionally, we collect tailored validation sets featuring challenging camera movements and human actions to validate the effectiveness of our method. Codes are released at https://github.com/alibaba-damo-academy/Uni3C. Chenjie Cao, Jingkai Zhou, Shikai Li, Jingyun Liang, Chaohui Yu, Fan Wang 0019, Xiangyang Xue 0001, Yanwei Fu 0001 |
SIGGRAPH Asia | 7 |
| 2025 | Learning global object-centric representations via disentangled slot attention
Tonglin Chen, Yinxuan Huang, Zhimeng Shen, Bin Li 0015, Xiangyang Xue 0001 |
Mach. Learn. | 6 |
| 2025 | Synthesizing efficient data with diffusion models for person re-identification pre-training
Ke Niu 0004, Haiyang Yu 0004, Xuelin Qian, Teng Fu 0001, Bin Li 0015, Xiangyang Xue 0001 |
Mach. Learn. | 6 |
| 2025 | Distribution aligned semantics adaption for lifelong person re-identification
Qizao Wang, Xuelin Qian, Bin Li 0015, Xiangyang Xue 0001 |
Mach. Learn. | 4 |
| 2025 | Dynamic Routing and Knowledge Re-Learning for Data-Free Black-Box AttackabstractDeep learning models have emerged as strong and efficient tools that can be applied to a broad spectrum of complex learning problems and many real-world applications. However, more and more works show that deep models are vulnerable to adversarial examples. Compared to vanilla attack settings, this paper advocates a more practical setting of data-free black-box attack, for which the attackers can completely not access the structures and parameters of the target model, as well as the intermediate features and any training data associated with the model. To tackle this task, previous methods generate transferable adversarial examples from a transparent substitute model to the target model. However, we found that these works have the limitations of taking static substitute model structure for different targets, only using hard synthesized examples once, and still relying on data statistics of the target model. This may potentially harm the performance of attacking the target model. To this end, we propose a novel Dynamic Routing and Knowledge Re-Learning framework (DraKe) to effectively learn a dynamic substitute model from the target model. Specifically, given synthesized training samples, a dynamic substitute structure learning strategy is proposed to adaptively generate optimal substitute model structure via a policy network according to different target models and tasks. To facilitate the substitute training, we present a graph-based structure information learning to capture the structural knowledge learned from the target model. For the inherent limitation that online data generation can only be learned once, a dynamic knowledge re-learning strategy is proposed to adjust the weights of optimization objectives and re-learn hard samples. Extensive experiments on four public image classification datasets and one face recognition benchmark are conducted to evaluate the efficacy of our Drake. We can obtain significant improvement compared with state-of-the-art competitors. More importantly, our DraKe consistently achieves attack superiority for different target models (e.g., residual networks, and vision transformers), showing great potential for complex real-world applications. Xuelin Qian, Wenxuan Wang 0003, Yu-Gang Jiang 0001, Xiangyang Xue 0001, Yanwei Fu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | Image-Text-Image Knowledge Transfer for Lifelong Person Re-Identification With Hybrid Clothing StatesabstractWith the continuous expansion of intelligent surveillance networks, lifelong person re-identification (LReID) has received widespread attention, pursuing the need of self-evolution across different domains. However, existing LReID studies accumulate knowledge with the assumption that people would not change their clothes. In this paper, we propose a more practical task, namely lifelong person re-identification with hybrid clothing states (LReID-Hybrid), which takes a series of cloth-changing and same-cloth domains into account during lifelong learning. To tackle the challenges of knowledge granularity mismatch and knowledge presentation mismatch in LReID-Hybrid, we take advantage of the consistency and generalization capabilities of the text space, and propose a novel framework, dubbed Teata, to effectively align, transfer, and accumulate knowledge in an "image-text-image" closed loop. Concretely, to achieve effective knowledge transfer, we design a Structured Semantic Prompt (SSP) learning to decompose the text prompt into several structured pairs to distill knowledge from the image space with a unified granularity of text description. Then, we introduce a Knowledge Adaptation and Projection (KAP) strategy, which tunes text knowledge via a slow-paced learner to adapt to different tasks without catastrophic forgetting. Extensive experiments demonstrate the superiority of our proposed Teata for LReID-Hybrid as well as on conventional LReID benchmarks over advanced methods. Qizao Wang, Xuelin Qian, Bin Li 0015, Yanwei Fu 0001, Xiangyang Xue 0001 |
IEEE Trans. Image Process. | 5 |
| 2025 | Toward Camera Open-Set 3D Object Detection for Autonomous Driving ScenariosabstractConventional camera-based 3D object detectors in autonomous driving are limited to recognizing a predefined set of objects, which poses a safety risk when encountering novel or unseen objects in real-world scenarios. To address this limitation, we present OS-Det3D, a two-stage training framework designed for camera-based open-set 3D object detection. In the first stage, our proposed 3D object discovery network (ODN3D) uses geometric cues from LiDAR point clouds to generate class-agnostic 3D object proposals, each of which are assigned a 3D objectness score. This approach allows the network to discover objects beyond known categories, allowing for the detection of unfamiliar objects. However, due to the absence of class constraints, ODN3D-generated proposals may include noisy data, particularly in cluttered or dynamic scenes. To mitigate this issue, we introduce a joint selection (JS) module in the second stage. The JS module uses both camera bird’s eye view (BEV) feature responses and 3D objectness scores to filter out low-quality proposals, yielding high-quality pseudo ground truth for unknown objects. OS-Det3D significantly enhances the ability of camera 3D detectors to discover and identify unknown objects while also improving the performance on known objects, as demonstrated through extensive experiments on the nuScenes and KITTI datasets. Zhuolin He, Xinrun Li, Jiacheng Tang, Shoumeng Qiu, Wenfu Wang, Xiangyang Xue 0001, Jian Pu |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2025 | Cross-Modal Complementary Learning and Template-Based Reasoning Chains for Future Event Prediction in VideosabstractAlthough multi-modal large language models (MLLMs) have impressive cross-modal reasoning and prediction capabilities, a unified and rigorous evaluation standard is still lacking. In this paper, we propose a future event prediction task to evaluate their cross-modal temporal prediction capability. This task requires the model to generate descriptions of events that may occur in future based on the input premise video. We build a dataset on the existing datasets for model evaluation. This task faces many challenges, including the complexity of processing video data, such as understanding changes in objects, actions, and time dimensions within the video and the interference of redundant information. To address these challenges, we propose a novel cross-modal prediction framework that introduces cross-modal supplementary learning and template-based reasoning chains based on MLLMs. Cross-modal supplementary learning aims to promote visual and text information to supplement and mine their respective information, primarily to capture critical information in videos, relying on the adaptive temporal filter and casual Q-Former. The template-based reasoning chain drives GPT-4 to generate a series of template question pairs through design prompts, gradually guiding the model to perform hierarchical reasoning to support the final prediction. Through experimental evaluation, the performance of the current MLLMs may not meet the requirements, and our model outperforms all existing models in predicting future events. It shows that the capabilities of MLLMs can be further explored. Chenghang Lai, Weifeng Ge, Xiangyang Xue 0001 |
IEEE Trans. Multim. | 3 |
| 2025 | Unsupervised Learning of Global Object-Centric Representations for Compositional Scene UnderstandingabstractThe ability to extract invariant visual features of objects from complex scenes and identify the same objects in different scenes is inborn for humans. To endow AI systems with such capability, we introduce a novel compositional scene understanding method known as Compositional Scene understanding via Global Object-centric representations (CSGOs). CSGO achieves comprehensive scene understanding, including the discovery and identification of objects, by leveraging a set of learnable global object-centric representations in an unsupervised manner. CSGO comprises three components: 1) Local Object-Centric Learning, which is responsible for extracting localized and scene-specific object-centric representations to discover objects; 2) Image Decoding, facilitating the reconstruction of object and scene images using the obtained object-centric representation as input; and 3) Global Object-Centric Learning, identifying the object across diverse scenes according to a set of learnable global object-centric representations, which indicates the scene-free intrinsic attributes (i.e., appearance and shape) of objects. Experimental results on three synthetic datasets and one real-world scene dataset demonstrate that CSGO has excellent object identification and attribute disentanglement abilities. Furthermore, the scene decomposition performance (indicating object discovery performance) of CSGO is superior to comparison methods. Tonglin Chen, Yinxuan Huang, Bin Li 0015, Xiangyang Xue 0001 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2024 | Generating and Reweighting Dense Contrastive Patterns for Unsupervised Anomaly DetectionabstractRecent unsupervised anomaly detection methods often rely on feature extractors pretrained with auxiliary datasets or on well-crafted anomaly-simulated samples. However, this might limit their adaptability to an increasing set of anomaly detection tasks due to the priors in the selection of auxiliary datasets or the strategy of anomaly simulation. To tackle this challenge, we first introduce a prior-less anomaly generation paradigm and subsequently develop an innovative unsupervised anomaly detection framework named GRAD, grounded in this paradigm. GRAD comprises three essential components: (1) a diffusion model (PatchDiff) to generate contrastive patterns by preserving the local structures while disregarding the global structures present in normal images, (2) a self-supervised reweighting mechanism to handle the challenge of long-tailed and unlabeled contrastive patterns generated by PatchDiff, and (3) a lightweight patch-level detector to efficiently distinguish the normal patterns and reweighted contrastive patterns. The generation results of PatchDiff effectively expose various types of anomaly patterns, e.g. structural and logical anomaly patterns. In addition, extensive experiments on both MVTec AD and MVTec LOCO datasets also support the aforementioned observation and demonstrate that GRAD achieves competitive anomaly detection accuracy and superior inference speed. Songmin Dai, Yifan Wu 0011, Xiaoqiang Li 0002, Xiangyang Xue 0001 |
AAAI | 4 |
| 2024 | Federated Adaptive Prompt Tuning for Multi-Domain Collaborative LearningabstractFederated learning (FL) enables multiple clients to collaboratively train a global model without disclosing their data. Previous researches often require training the complete model parameters. However, the emergence of powerful pre-trained models makes it possible to achieve higher performance with fewer learnable parameters in FL. In this paper, we propose a federated adaptive prompt tuning algorithm, FedAPT, for multi-domain collaborative image classification with powerful foundation models, like CLIP. Compared with direct federated prompt tuning, our core idea is to adaptively unlock specific domain knowledge for each test sample in order to provide them with personalized prompts. To implement this idea, we design an adaptive prompt tuning module, which consists of a meta prompt, an adaptive network, and some keys. The server randomly generates a set of keys and assigns a unique key to each client. Then all clients cooperatively train the global adaptive network and meta prompt with the local datasets and the frozen keys. Ultimately, the global aggregation model can assign a personalized prompt to CLIP based on the domain features of each test sample. We perform extensive experiments on two multi-domain image classification datasets across two different settings -- supervised and unsupervised. The results show that FedAPT can achieve better performance with less than 10% of the number of parameters of the fully trained model, and the global model can perform well in diverse client domains simultaneously. Shangchao Su, Mingzhao Yang, Bin Li 0015, Xiangyang Xue 0001 |
AAAI | 4 |
| 2024 | Exploring One-Shot Semi-supervised Federated Learning with Pre-trained Diffusion ModelsabstractRecently, semi-supervised federated learning (semi-FL) has been proposed to handle the commonly seen real-world scenarios with labeled data on the server and unlabeled data on the clients. However, existing methods face several challenges such as communication costs, data heterogeneity, and training pressure on client devices. To address these challenges, we introduce the powerful diffusion models (DM) into semi-FL and propose FedDISC, a Federated Diffusion-Inspired Semi-supervised Co-training method. Specifically, we first extract prototypes of the labeled server data and use these prototypes to predict pseudo-labels of the client data. For each category, we compute the cluster centroids and domain-specific representations to signify the semantic and stylistic information of their distributions. After adding noise, these representations are sent back to the server, which uses the pre-trained DM to generate synthetic datasets complying with the client distributions and train a global model on it. With the assistance of vast knowledge within DM, the synthetic datasets have comparable quality and diversity to the client images, subsequently enabling the training of global models that achieve performance equivalent to or even surpassing the ceiling of supervised centralized training. FedDISC works within one communication round, does not require any local training, and involves very minimal information uploading, greatly enhancing its practicality. Extensive experiments on three large-scale datasets demonstrate that FedDISC effectively addresses the semi-FL problem on non-IID clients and outperforms the compared SOTA methods. Sufficient visualization experiments also illustrate that the synthetic dataset generated by FedDISC exhibits comparable diversity and quality to the original client dataset, with a neglectable possibility of leaking privacy-sensitive information of the clients. Mingzhao Yang, Shangchao Su, Bin Li 0015, Xiangyang Xue 0001 |
AAAI | 4 |
| 2024 | Enhancing Cross-Subject fMRI-to-Video Decoding with Global-Local Functional Alignment
Chong Li 0007, Xuelin Qian, Yun Wang 0021, Jingyang Huo, Xiangyang Xue 0001, Yanwei Fu 0001, Jianfeng Ma 0001 |
ECCV (83) | 5 |
| 2024 | Make a Strong Teacher with Label Assistance: A Novel Knowledge Distillation Approach for Semantic Segmentation
Shoumeng Qiu, Xinrun Li, Ru Wan, Xiangyang Xue 0001, Jian Pu |
ECCV (84) | 5 |
| 2024 | Improving Neural Surface Reconstruction with Feature Priors from Multi-view Images
Xinlin Ren, Chenjie Cao, Yanwei Fu 0001, Xiangyang Xue 0001 |
ECCV (58) | 4 |
| 2024 | FedRA: A Random Allocation Strategy for Federated Tuning to Unleash the Power of Heterogeneous Clients
Shangchao Su, Bin Li 0015, Xiangyang Xue 0001 |
ECCV (48) | 3 |
| 2024 | EAFormer: Scene Text Segmentation with Edge-Aware Transformers
Haiyang Yu 0004, Teng Fu 0001, Bin Li 0015, Xiangyang Xue 0001 |
ECCV (25) | 4 |
| 2024 | Towards Generative Abstract Reasoning: Completing Raven's Progressive Matrix via Rule Abstraction and SelectionabstractEndowing machines with abstract reasoning ability has been a long-term research topic in artificial intelligence. Raven's Progressive Matrix (RPM) is widely used to probe abstract visual reasoning in machine intelligence, where models will analyze the underlying rules and select one image from candidates to complete the image matrix. Participators of RPM tests can show powerful reasoning ability by inferring and combining attribute-changing rules and imagining the missing images at arbitrary positions of a matrix. However, existing solvers can hardly manifest such an ability in realistic RPM tests. In this paper, we propose a deep latent variable model for answer generation problems through Rule AbstractIon and SElection (RAISE). RAISE can encode image attributes into latent concepts and abstract atomic rules that act on the latent concepts. When generating answers, RAISE selects one atomic rule out of the global knowledge set for each latent concept to constitute the underlying rule of an RPM. In the experiments of bottom-right and arbitrary-position answer generation, RAISE outperforms the compared solvers in most configurations of realistic RPM datasets. In the odd-one-out task and two held-out configurations, RAISE can leverage acquired latent concepts and atomic rules to find the rule-breaking image in a matrix and handle problems with unseen combinations of rules and attributes. Bin Li 0015, Xiangyang Xue 0001 |
ICLR | 3 |
| 2024 | Unsupervised Object Discovery Via Object-Centric RepresentationabstractUnsupervised object discovery enables us to localize potential objects without any supervision, which has broad application prospects such as detecting skin disease and monitoring water waste, etc. However, existing research is based on the category-irrelevant object discovery, which are hard to discover specific set of categories without explicit supervised signal. To solve this problem, this paper proposes a novel Object-Centric Learning (OCL) framework, built upon pretrained Vision Transformer (ViT) model, to learn a set of latent representations of specified objects. These representations are gradually refined by slot-attention mechanism, which allows the model to further differentiate the representations of different categories of objects. A background completion self-supervised training task is further proposed to improve the generalization ability of model in real-world scenarios. Experimental results demonstrate that our OCL achieves state-of-the-art performance (77.29%, 29.88% and 57.96%) on skin cancer database, PH2 database and UAV-BD dataset. Bingfei Fu, Xiangyang Xue 0001 |
ICME | 2 |
| 2024 | Multi-LIO: A Lightweight Multiple LiDAR-Inertial Odometry SystemabstractThe integration of multiple LiDAR sensors has the potential to significantly enhance odometry systems by providing comprehensive environmental measurements. However, current multiple LiDAR-inertial odometry frameworks face challenges in real-time processing due to the voluminous data generated. This paper introduces a real-time, computationally efficient multiple LiDAR-inertial odometry system (Multi-LIO) that outperforms existing state-of-the-art solutions in accuracy and scalability. Utilizing a novel parallel strategy for state updates and a voxelized map format, Multi-LIO optimizes computational efficiency. Furthermore, we introduce a point-wise uncertainty estimation method to augment the accuracy of scan-to-map registration, particularly in large-scale and complex scenarios. We validate our system’s performance through extensive experiments on various challenging sequences. Multi-LIO emerges as a robust, scalable, and extensible solution, adaptable to various LiDAR configurations. Qi Chen 0025, Guanghao Li 0001, Xiangyang Xue 0001, Jian Pu |
ICRA | 3 |
| 2024 | FastOcc: Accelerating 3D Occupancy Prediction by Fusing the 2D Bird's-Eye View and Perspective ViewabstractIn autonomous driving, 3D occupancy prediction outputs voxel-wise status and semantic labels for more comprehensive understandings of 3D scenes compared with traditional perception tasks, such as 3D object detection and bird’s-eye view (BEV) semantic segmentation. Recent researchers have extensively explored various aspects of this task, including view transformation techniques, ground-truth label generation, and elaborate network design, aiming to achieve superior performance. However, the inference speed, crucial for running on an autonomous vehicle, is neglected. To this end, a new method, dubbed FastOcc, is proposed. By carefully analyzing the network effect and latency from four parts, including the input image resolution, image backbone, view transformation, and occupancy prediction head, it is found that the occupancy prediction head holds considerable potential for accelerating the model while keeping its accuracy. Targeted at improving this component, the time-consuming 3D convolution network is replaced with a novel residual-like architecture, where features are mainly digested by a lightweight 2D BEV convolution network and compensated by integrating the 3D voxel features interpolated from the original image features. Experiments on the Occ3D-nuScenes benchmark demonstrate that our FastOcc achieves state-of-the-art results with a fast inference speed. Wenhao Guan, Di Feng, Yuheng Du, Xiangyang Xue 0001, Jian Pu |
ICRA | 7 |
| 2024 | OpenAnnotate3D: Open-Vocabulary Auto-Labeling System for Multi-modal 3D DataabstractIn the era of big data and large models, automatic annotating functions for multi-modal data are of great significance for real-world AI-driven applications, such as autonomous driving and embodied AI. Unlike traditional closed-set annotation, open-vocabulary annotation is essential to achieve human-level cognition capability. However, there are few open-vocabulary auto-labeling systems for multi-modal 3D data. In this paper, we introduce OpenAnnotate3D, an open-source open-vocabulary auto-labeling system that can automatically generate 2D masks, 3D masks, and 3D bounding box annotations for vision and point cloud data. Our system integrates the chain-of-thought capabilities of Large Language Models (LLMs) and the cross-modality capabilities of vision-language models (VLMs). To the best of our knowledge, OpenAnnotate3D is one of the pioneering works for open-vocabulary multi-modal 3D auto-labeling. We conduct comprehensive evaluations on both public and in-house real-world datasets, which demonstrate that the system significantly improves annotation efficiency compared to manual annotation while providing accurate open-vocabulary auto-annotating results. Likun Cai, Xianhui Cheng, Zhongxue Gan 0001, Xiangyang Xue 0001, Wenchao Ding 0001 |
ICRA | 5 |
| 2024 | LAC-Net: Linear-Fusion Attention-Guided Convolutional Network for Accurate Robotic Grasping Under the OcclusionabstractThis paper addresses the challenge of perceiving complete object shapes through visual perception. While prior studies have demonstrated encouraging outcomes in segmenting the visible parts of objects within a scene, amodal segmentation, in particular, has the potential to allow robots to infer the occluded parts of objects. To this end, this paper introduces a new framework that explores amodal segmentation for robotic grasping in cluttered scenes, thus greatly enhancing robotic grasping abilities. Initially, we use a conventional segmentation algorithm to detect the visible segments of the target object, which provides shape priors for completing the full object mask. Particularly, to explore how to utilize semantic features from RGB images and geometric information from depth images, we propose a Linear-fusion Attention-guided Convolutional Network (LAC-Net). LAC-Net utilizes the linear-fusion strategy to effectively fuse this cross-modal data, and then uses the prior visible mask as attention map to guide the network to focus on target feature locations for further complete mask recovery. Using the amodal mask of the target object provides advantages in selecting more accurate and robust grasp points compared to relying solely on the visible segments. The results on different datasets show that our method achieves state-of-the-art performance. Furthermore, the robot experiments validate the feasibility and robustness of this method in the real world. Our code and demonstrations are available on the project page: https://jrryzh.github.io/LAC-Net. Yongchong Gu, Jianxiong Gao, Qiang Sun 0008, Xinwei Sun 0001, Xiangyang Xue 0001, Yanwei Fu 0001 |
IROS | 7 |
| 2024 | FedDEO: Description-Enhanced One-Shot Federated Learning with Diffusion ModelsabstractIn recent years, the attention towards One-Shot Federated Learning (OSFL) has been driven by its capacity to minimize communication. With the development of the diffusion model (DM), several methods employ the DM for OSFL, utilizing model parameters, image features, or textual prompts as mediums to transfer the local client knowledge to the server. However, these mediums often require public datasets or the uniform feature extractor, significantly limiting their practicality. In this paper, we propose FedDEO, a Description-Enhanced One-Shot Federated Learning Method with DMs, offering a novel exploration of utilizing the DM in OSFL. The core idea of our method involves training local descriptions on the clients, serving as the medium to transfer the knowledge of the distributed clients to the server. Firstly, we train local descriptions on the client data to capture the characteristics of client distributions, which are then uploaded to the server. On the server, the descriptions are used as conditions to guide the DM in generating synthetic datasets that comply with the distributions of various clients, enabling the training of the aggregated model. Theoretical analyses and sufficient quantitation and visualization experiments on three large-scale real-world datasets demonstrate that through the training of local descriptions, the server is capable of generating synthetic datasets with high quality and diversity. Consequently, with advantages in communication and privacy protection, the aggregated model outperforms compared FL or diffusion-based OSFL methods and, on some clients, outperforms the performance ceiling of centralized training. Mingzhao Yang, Shangchao Su, Bin Li 0015, Xiangyang Xue 0001 |
ACM Multimedia | 4 |
| 2024 | MVInpainter: Learning Multi-View Consistent Inpainting to Bridge 2D and 3D EditingabstractNovel View Synthesis (NVS) and 3D generation have recently achieved prominent improvements. However, these works mainly focus on confined categories or synthetic 3D assets, which are discouraged from generalizing to challenging in-the-wild scenes and fail to be employed with 2D synthesis directly. Moreover, these methods heavily depended on camera poses, limiting their real-world applications.
To overcome these issues, we propose MVInpainter, re-formulating the 3D editing as a multi-view 2D inpainting task. Specifically, MVInpainter partially inpaints multi-view images with the reference guidance rather than intractably generating an entirely novel view from scratch, which largely simplifies the difficulty of in-the-wild NVS and leverages unmasked clues instead of explicit pose conditions. To ensure cross-view consistency, MVInpainter is enhanced by video priors from motion components and appearance guidance from concatenated reference key\&value attention. Furthermore, MVInpainter incorporates slot attention to aggregate high-level optical flow features from unmasked regions to control the camera movement with pose-free training and inference. Sufficient scene-level experiments on both object-centric and forward-facing datasets verify the effectiveness of MVInpainter, including diverse tasks, such as multi-view object removal, synthesis, insertion, and replacement. The project page is https://ewrfcas.github.io/MVInpainter/. Chenjie Cao, Chaohui Yu, Fan Wang 0019, Xiangyang Xue 0001, Yanwei Fu 0001 |
NeurIPS | 4 |
| 2024 | Improving Viewpoint-Independent Object-Centric Representations through Active Viewpoint SelectionabstractGiven the complexities inherent in visual scenes, such as object occlusion, a comprehensive understanding often requires observation from multiple viewpoints. Existing multi-viewpoint object-centric learning methods typically employ random or sequential viewpoint selection strategies. While applicable across various scenes, these strategies may not always be ideal, as certain scenes could benefit more from specific viewpoints. To address this limitation, we propose a novel active viewpoint selection strategy. This strategy predicts images from unknown viewpoints based on information from observation images for each scene. It then compares the object-centric representations extracted from both viewpoints and selects the unknown viewpoint with the largest disparity, indicating the greatest gain in information, as the next observation viewpoint. Through experiments on various datasets, we demonstrate the effectiveness of our active viewpoint selection strategy, significantly enhancing segmentation and reconstruction performance compared to random viewpoint selection. Moreover, our method can accurately predict images from unknown viewpoints. Yinxuan Huang, Chengmin Gao, Bin Li 0015, Xiangyang Xue 0001 |
NeurIPS | 4 |
| 2024 | Automated Label Unification for Multi-Dataset Semantic Segmentation with GNNsabstractDeep supervised models possess significant capability to assimilate extensive training data, thereby presenting an opportunity to enhance model performance through training on multiple datasets. However, conflicts arising from different label spaces among datasets may adversely affect model performance. In this paper, we propose a novel approach to automatically construct a unified label space across multiple datasets using graph neural networks. This enables semantic segmentation models to be trained simultaneously on multiple datasets, resulting in performance improvements. Unlike existing methods, our approach facilitates seamless training without the need for additional manual reannotation or taxonomy reconciliation. This significantly enhances the efficiency and effectiveness of multi-dataset segmentation model training. The results demonstrate that our method significantly outperforms other multi-dataset training methods when trained on seven datasets simultaneously, and achieves state-of-the-art performance on the WildDash 2 benchmark. Our code can be found in https://github.com/Mrhonor/AutoUniSeg. Xiangyang Xue 0001, Jian Pu |
NeurIPS | 3 |
| 2024 | Learning a Mixture of Conditional Gating Blocks for Visual Question Answering
Qiang Sun 0008, Yanwei Fu 0001, Xiangyang Xue 0001 |
J. Comput. Sci. Technol. | 3 |
| 2024 | Compositional scene modeling with global object-centric representations
Tonglin Chen, Zhimeng Shen, Bin Li 0015, Xiangyang Xue 0001 |
Mach. Learn. | 4 |
| 2024 | Style spectroscope: improve interpretability and controllability through Fourier analysis
Zhiyu Jin, Xuli Shen, Bin Li 0015, Xiangyang Xue 0001 |
Mach. Learn. | 4 |
| 2024 | Chinese character recognition with radical-structured stroke trees
Haiyang Yu 0004, Jingye Chen, Bin Li 0015, Xiangyang Xue 0001 |
Mach. Learn. | 4 |
| 2024 | DeepSFM: Robust Deep Iterative Refinement for Structure From MotionabstractStructure from Motion (SfM) is a fundamental computer vision problem which has not been well handled by deep learning. One of the promising solutions is to apply explicit structural constraint, e.g., 3D cost volume, into the neural network. Obtaining accurate camera poses from images alone can be challenging, especially with complicated environmental factors. Existing methods usually assume accurate camera poses from GT or other methods, which is unrealistic in practice and additional sensors are needed. In this work, we design a physical driven architecture, namely DeepSFM, inspired by traditional Bundle Adjustment, which consists of two cost volume based architectures to iteratively refine depth and pose. The explicit constraints on both depth and pose, when combined with the learning components, bring merit from both traditional BA and emerging deep learning technology. To speed up the learning and inference efficiency, we apply the Gated Recurrent Units (GRUs)-based depth and pose update modules with coarse to fine cost volumes on the iterative refinements. In addition, with the extended residual depth prediction module, our model can be adapted to dynamic scenes effectively. Extensive experiments on various datasets show that our model achieves state-of-the-art performance with superior robustness against challenging inputs. Xinlin Ren, Xingkui Wei, Zhuwen Li, Yanwei Fu 0001, Yinda Zhang 0001, Xiangyang Xue 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2024 | Unsupervised Object-Centric Learning From Multiple Unspecified ViewpointsabstractVisual scenes are extremely diverse, not only because there are infinite possible combinations of objects and backgrounds but also because the observations of the same scene may vary greatly with the change of viewpoints. When observing a multi-object visual scene from multiple viewpoints, humans can perceive the scene compositionally from each viewpoint while achieving the so-called "object constancy" across different viewpoints, even though the exact viewpoints are untold. This ability is essential for humans to identify the same object while moving and to learn from vision efficiently. It is intriguing to design models that have a similar ability. In this article, we consider a novel problem of learning compositional scene representations from multiple unspecified (i.e., unknown and unrelated) viewpoints without using any supervision and propose a deep generative model which separates latent representations into a viewpoint-independent part and a viewpoint-dependent part to solve this problem. During the inference, latent representations are randomly initialized and iteratively updated by integrating the information in different viewpoints with neural networks. Experiments on several specifically designed synthetic datasets have shown that the proposed method can effectively learn from multiple unspecified viewpoints. Jinyang Yuan, Tonglin Chen, Zhimeng Shen, Bin Li 0015, Xiangyang Xue 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2024 | Guided contrastive boundary learning for semantic segmentation
Shoumeng Qiu, Haiqiang Zhang, Ru Wan, Xiangyang Xue 0001, Jian Pu |
Pattern Recognit. | 5 |
| 2024 | Embrace sustainable AI: Dynamic data subset selection for image classification
Zimo Yin, Jian Pu, Ru Wan, Xiangyang Xue 0001 |
Pattern Recognit. | 4 |
| 2024 | GAT-COBO: Cost-Sensitive Graph Neural Network for Telecom Fraud DetectionabstractAlong with the rapid evolution of mobile communication technologies, such as 5G, there has been a significant increase in telecom fraud, which severely dissipates individual fortune and social wealth. In recent years, graph mining techniques are gradually becoming a mainstream solution for detecting telecom fraud. However, the graph imbalance problem, caused by the Pareto principle, brings severe challenges to graph data mining. This emerging and complex issue has received limited attention in prior research. In this paper, we propose aGraphATtention network withCOst-sensitiveBOosting (GAT-COBO) for the graph imbalance problem. First, we design a GAT-based base classifier to learn the embeddings of all nodes in the graph. Then, we feed the embeddings into a well-designed cost-sensitive learner for imbalanced learning. Next, we update the weights according to the misclassification cost to make the model focus more on the minority class. Finally, we sum the node embeddings obtained by multiple cost-sensitive learners to obtain a comprehensive node representation, which is used for the downstream anomaly detection task. Extensive experiments on two real-world telecom fraud detection datasets demonstrate that our proposed method is effective for the graph imbalance problem, outperforming the state-of-the-art GNNs and GNN-based fraud detectors. In addition, our model is also helpful for solving the widespread over-smoothing problem in GNNs. The GAT-COBO code and datasets are available athttps://github.com/xxhu94/GAT-COBO. Xinxin Hu, Haotian Chen 0003, Hongchang Chen, Xing Li 0013, Xiangyang Xue 0001 |
IEEE Trans. Big Data | 8 |
| 2024 | Cost-Sensitive GNN-Based Imbalanced Learning for Mobile Social Network Fraud DetectionabstractIn recent years, the increasing prevalence of mobile social network fraud has led to significant distress and depletion of personal and social wealth, resulting in considerable economic harm. Graph neural networks (GNNs) have emerged as a popular approach to tackle this issue. However, the challenge of graph imbalance, which can greatly impede the effectiveness of GNN-based fraud detection methods, has received little attention in prior research. Thus, we are going to present a novel cost-sensitive graph neural network (CSGNN) in this article. Initially, reinforcement learning is utilized to train a suitable sampling threshold, followed by neighbor sampling based on node similarity, which helps to alleviate the graph imbalance issue preliminarily. Subsequently, message aggregation is executed on the sampled graph using GNN to obtain node embeddings. Concurrently, the optimization objective for the cost matrix is formulated using the sample histogram matrix, scatter matrix, and confusion matrix. The cost matrix and GNN are collaboratively optimized through the backpropagation algorithm. Ultimately, the derived cost-sensitive node embedding is employed for fraudulent node detection. Furthermore, this study provides a theoretical demonstration of the effectiveness of adaptive cost-sensitive learning in GNN. Extensive experiments are carried out on two publicly accessible real-world mobile network fraud datasets, revealing that the proposed CSGNN effectively addresses the graph imbalance issue while outperforming state-of-the-art algorithms in detection performance. The CSGNN code and datasets can be accessed at https://github.com/xxhu94/CSGNN. Xinxin Hu, Haotian Chen 0003, Hongchang Chen, Xing Li 0013, Xiangyang Xue 0001 |
IEEE Trans. Comput. Soc. Syst. | 8 |
| 2024 | Object-Centric Cross-Modal Knowledge Reasoning for Future Event Prediction in VideosabstractAlthough multi-modal large language models possess impressive cross-modal reasoning and prediction capabilities, they lack a unified and rigorous evaluation standard. In this paper, we introduce a future event prediction task to assess the cross-modal temporal prediction capabilities of these models. This task requires the model to generate descriptions of events that may occur in the future based on input video. To tackle this new task, we propose an object-centric cross-modal knowledge reasoning framework, which combines a basic information encoder, an adaptive multi-segment filter, a spatial-temporal relation encoder, a vision-text interaction module, and a pre-trained large language model decoder. The adaptive multi-segment filter captures selectively capture critical visual information in videos, enhancing the model’s focus on relevant features. The spatial-temporal relation encoder decomposes and associates the objects and scene information in the video. Additionally, the vision-text interaction module enhances the connection between visual sequences and their corresponding textual narratives, ensuring semantic coherence and consistency. To evaluate our framework, we constructed a dataset containing descriptions, dialogues of future events, and object-centric event reasoning chains. Experimental results indicate that the proposed framework outperforms all previous methods for future event prediction. Ablation studies further demonstrate the effectiveness of the designed modules. Chenghang Lai, Haibo Wang 0006, Weifeng Ge, Xiangyang Xue 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Exploring Fine-Grained Representation and Recomposition for Cloth-Changing Person Re-IdentificationabstractCloth-changing person Re-IDentification (Re-ID) is a particularly challenging task, suffering from two limitations of inferior discriminative features and limited training samples. Existing methods mainly leverage auxiliary information to facilitate identity-relevant feature learning, including soft-biometrics features of shapes or gaits, and additional labels of clothing. However, this information may be unavailable in real-world applications. In this paper, we propose a novel FIne-grained Representation and Recomposition (FIRe2) framework to tackle both limitations without any auxiliary annotation or data. Specifically, we first design a Fine-grained Feature Mining (FFM) module to separately cluster images of each person. Images with similar so-called fine-grained attributes (e.g., clothes and viewpoints) are encouraged to cluster together. An attribute-aware classification loss is introduced to perform fine-grained learning based on cluster labels, which are not shared among different people, promoting the model to learn identity-relevant features. Furthermore, to take full advantage of fine-grained attributes, we present a Fine-grained Attribute Recomposition (FAR) module by recomposing image features with different attributes in the latent space. It significantly enhances robust feature learning. Extensive experiments demonstrate that FIRe2 can achieve state-of-the-art performance on five widely-used cloth-changing person Re-ID benchmarks. The code is available athttps://github.com/QizaoWang/FIRe-CCReID. Qizao Wang, Xuelin Qian, Bin Li 0015, Xiangyang Xue 0001, Yanwei Fu 0001 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2024 | VPL-SLAM: A Vertical Line Supported Point Line Monocular SLAM SystemabstractTraditional monocular visual simultaneous localization and mapping (SLAM) systems rely on point features or line features to estimate and optimize the camera trajectory and build a map of the surrounding environment. However, in complex scenarios such as underground parking, the performance of traditional point-line SLAM systems tends to degrade due to mirror reflection, illumination change, poor texture, and other interference. This paper proposes VPL-SLAM, a structural vertical line supported point-line monocular SLAM system that works well in complex environments such as underground parking or campus. The proposed system leverages structural vertical lines at all instances of the process. With the assistance of the structural vertical lines and global vertical direction, our system can output a more accurate visual odometry result. Furthermore, the resulting map of our system is a more reasonable structural line feature map than the previous point-line-based monocular SLAM systems. Our system has been tested with the popular autonomous driving dataset Kitti Odometry. In addition, to fully test the proposed SLAM system, we also test our system using a self-collected dataset, including underground parking and campus scenarios. As a result, our proposal reveals a more accurate navigation result and a more reasonable structural resulting map compared to state-of-the-art point-line SLAM systems such as Structure PLP-SLAM. Qi Chen 0025, Yu Cao 0024, Guanghao Li 0001, Shoumeng Qiu, Xiangyang Xue 0001, Hong Lu 0001, Jian Pu |
IEEE Trans. Intell. Transp. Syst. | 7 |
| 2023 | Grad-PU: Arbitrary-Scale Point Cloud Upsampling via Gradient Descent with Learned Distance FunctionsabstractMost existing point cloud upsampling methods have roughly three steps: feature extraction, feature expansion and 3D coordinate prediction. However, they usually suffer from two critical issues: (1) fixed upsampling rate after one-time training, since the feature expansion unit is customized for each upsampling rate; (2) outliers or shrinkage artifact caused by the difficulty of precisely predicting 3D coordinates or residuals of upsampled points. To adress them, we propose a new framework for accurate point cloud upsampling that supports arbitrary upsampling rates. Our method first interpolates the low-res point cloud according to a given upsampling rate. And then refine the positions of the interpolated points with an iterative optimization process, guided by a trained model estimating the difference between the current point cloud and the high-res target. Extensive quantitative and qualitative results on benchmarks and downstream tasks demonstrate that our method achieves the state-of-the-art accuracy and efficiency. Danhang Tang, Yinda Zhang 0001, Xiangyang Xue 0001, Yanwei Fu 0001 |
CVPR | 4 |
| 2023 | PourIt!: Weakly-supervised Liquid Perception from a Single Image for Visual Closed-Loop Robotic PouringabstractLiquid perception is critical for robotic pouring tasks. It usually requires the robust visual detection of flowing liquid. However, while recent works have shown promising results in liquid perception, they typically require labeled data for model training, a process that is both time-consuming and reliant on human labor. To this end, this paper proposes a simple yet effective framework PourIt!, to serve as a tool for robotic pouring tasks. We design a simple data collection pipeline that only needs image-level labels to reduce the reliance on tedious pixel-wise annotations. Then, a binary classification model is trained to generate Class Activation Map (CAM) that focuses on the visual difference between these two kinds of collected data, i.e., the existence of liquid drop or not. We also devise a feature contrast strategy to improve the quality of the CAM, thus entirely and tightly covering the actual liquid regions. Then, the container pose is further utilized to facilitate the 3D point cloud recovery of the detected liquid region. Finally, the liquid-to-container distance is calculated for visual closed-loop control of the physical robot. To validate the effectiveness of our proposed method, we also contribute a novel dataset for our task and name it PourIt! dataset. Extensive results on this dataset and physical Franka robot have shown the utility and effectiveness of our method in the robotic pouring tasks. Our dataset, code and pre-trained models will be available on the project page1. Yanwei Fu 0001, Xiangyang Xue 0001 |
ICCV | 3 |
| 2023 | Learning Versatile 3D Shape Generation with Improved Auto-regressive ModelsabstractAuto-Regressive (AR) models have achieved impressive results in 2D image generation by modeling joint distributions in the grid space. While this approach has been extended to the 3D domain for powerful shape generation, it still has two limitations: expensive computations on volumetric grids and ambiguous auto-regressive order along grid dimensions. To overcome these limitations, we propose the Improved Auto-regressive Model (ImAM) for 3D shape generation, which applies discrete representation learning based on a latent vector instead of volumetric grids. Our approach not only reduces computational costs but also preserves essential geometric details by learning the joint distribution in a more tractable order. Moreover, thanks to the simplicity of our model architecture, we can naturally extend it from unconditional to conditional generation by concatenating various conditioning inputs, such as point clouds, categories, images, and texts. Extensive experiments demonstrate that ImAM can synthesize diverse and faithful shapes of multiple categories, achieving state-of-the-art performance. Simian Luo, Xuelin Qian, Yanwei Fu 0001, Yinda Zhang 0001, Ying Tai, Zhenyu Zhang 0005, Chengjie Wang 0001, Xiangyang Xue 0001 |
ICCV | 8 |
| 2023 | Chinese Text Recognition with A Pre-Trained CLIP-Like Model Through Image-IDS AligningabstractScene text recognition has been studied for decades due to its broad applications. However, despite Chinese characters possessing different characteristics from Latin characters, such as complex inner structures and large categories, few methods have been proposed for Chinese Text Recognition (CTR). Particularly, the characteristic of large categories poses challenges in dealing with zero-shot and few-shot Chinese characters. In this paper, inspired by the way humans recognize Chinese texts, we propose a two-stage framework for CTR. Firstly, we pre-train a CLIP-like model through aligning printed character images and Ideographic Description Sequences (IDS). This pre-training stage simulates humans recognizing Chinese characters and obtains the canonical representation of each character. Subsequently, the learned representations are employed to supervise the CTR model, such that traditional single-character recognition can be improved to text-line recognition through image-IDS matching. To evaluate the effectiveness of the proposed method, we conduct extensive experiments on both Chinese character recognition (CCR) and CTR. The experimental results demonstrate that the proposed method performs best in CCR and outperforms previous methods in most scenarios of the CTR benchmark. It is worth noting that the proposed method can recognize zero-shot Chinese characters in text images without fine-tuning, whereas previous methods require fine-tuning when new classes appear. The code is available at https://github.com/FudanVI/FudanOCR/tree/main/image-ids-CTR. Haiyang Yu 0004, Bin Li 0015, Xiangyang Xue 0001 |
ICCV | 4 |
| 2023 | Foreign Object Detection Based on Compositional Scene Modeling
Bingfei Fu, Xiangyang Xue 0001 |
ICIG (3) | 3 |
| 2023 | Compositional Law Parsing with Latent Random Functions
Bin Li 0015, Xiangyang Xue 0001 |
ICLR | 3 |
| 2023 | Cross-domain Federated Object DetectionabstractDetection models trained by one party (including server) may face severe performance degradation when distributed to other users (clients). Federated learning can enable multi-party collaborative learning without leaking client data. In this paper, we focus on a special cross-domain scenario in which the server has large-scale labeled data and multiple clients only have a small amount of labeled data; meanwhile, there exist differences in data distributions among the clients. In this case, traditional federated learning methods can’t help a client learn both the global knowledge of all participants and its own unique knowledge. To make up for this limitation, we propose a cross-domain federated object detection framework, named FedOD. The proposed framework first performs the federated training to obtain a public global aggregated model through multi-teacher distillation, and sends the aggregated model back to each client for fine-tuning its personalized local model. After a few rounds of communication, on each client we can perform weighted ensemble inference on the public global model and the personalized local model. We establish a federated object detection dataset which has significant background differences and instance differences based on multiple public autonomous driving datasets, and then conduct extensive experiments on the dataset. The experimental results validate the effectiveness of the proposed method. Shangchao Su, Bin Li 0015, Mingzhao Yang, Xiangyang Xue 0001 |
ICME | 5 |
| 2023 | TextFormer: Component-aware Text Segmentation with TransformerabstractIn recent years, deep learning techniques have made significant advancements in text segmentation. However, most existing methods do not take into account that characters are composed of smaller components, such as strokes and other local patterns. Furthermore, the similarities between text components are crucial for effective text segmentation. With this in mind, we propose a multi-level Transformer-based method for text segmentation that incorporates a recognition module. To enhance the interaction between text components and extract features at different granularities, we introduce Global and Local Self-Attention blocks. Our recognition module is trained jointly with the segmentation module to improve the model’s ability to focus on text details and improve its perception of texts. By aggregating features from multiple granularities, our segmentation module produces accurate pixel-level mask predictions. The experimental results demonstrate the effectiveness of our approach on several text segmentation benchmarks and show that it outperforms existing methods. Chaoyue Wu, Haiyang Yu 0004, Bin Li 0015, Xiangyang Xue 0001 |
ICME | 5 |
| 2023 | Multi-to-Single Knowledge Distillation for Point Cloud Semantic Segmentationabstract3D point cloud semantic segmentation is one of the fundamental tasks for environmental understanding. Although significant progress has been made in recent years, the performance of classes with few examples or few points is still far from satisfactory. In this paper, we propose a novel multi-to-single knowledge distillation framework for the 3D point cloud semantic segmentation task to boost the performance of those hard classes. Instead of fusing all the points of multi-scans directly, only the instances that belong to the previously defined hard classes are fused. To effectively and sufficiently distill valuable knowledge from multi-scans, we leverage a multilevel distillation framework, i.e., feature representation distillation, logit distillation, and affinity distillation. We further develop a novel instance-aware affinity distillation algorithm for capturing high-level structural knowledge to enhance the distillation efficacy for hard classes. Finally, we conduct experiments on the SemanticKITTI dataset, and the results on both the validation and test sets demonstrate that our method yields substantial improvements compared with the baseline method. The code is available at https://github.com/skyshoumeng/M2SKD. Shoumeng Qiu, Haiqiang Zhang, Xiangyang Xue 0001, Jian Pu |
ICRA | 4 |
| 2023 | Orientation-Independent Chinese Text Recognition in Scene ImagesabstractScene text recognition (STR) has attracted much attention due to its broad applications. The previous works pay more attention to dealing with the recognition of Latin text images with complex backgrounds by introducing language models or other auxiliary networks. Different from Latin texts, many vertical Chinese texts exist in natural scenes, which brings difficulties to current state-of-the-art STR methods. In this paper, we take the first attempt to extract orientation-independent visual features by disentangling content and orientation information of text images, thus recognizing both horizontal and vertical texts robustly in natural scenes. Specifically, we introduce a Character Image Reconstruction Network (CIRN) to recover corresponding printed character images with disentangled content and orientation information. We conduct experiments on a scene dataset for benchmarking Chinese text recognition, and the results demonstrate that the proposed method can indeed improve performance through disentangling content and orientation information. To further validate the effectiveness of our method, we additionally collect a Vertical Chinese Text Recognition (VCTR) dataset. The experimental results show that the proposed method achieves 45.63\% improvement on VCTR when introducing CIRN to the baseline model. Haiyang Yu 0004, Bin Li 0015, Xiangyang Xue 0001 |
IJCAI | 4 |
| 2023 | Towards Accurate Video Text Spotting with Text-wise Semantic ReasoningabstractVideo text spotting (VTS) aims at extracting texts from videos, where text detection, tracking and recognition are conducted simultaneously. There have been some works that can tackle VTS; however, they may ignore the underlying semantic relationships among texts within a frame. We observe that the texts within a frame usually share similar semantics, which suggests that, if one text is predicted incorrectly by a text recognizer, it still has a chance to be corrected via semantic reasoning. In this paper, we propose an accurate video text spotter, VLSpotter, that reads texts visually, linguistically, and semantically. For ‘visually’, we propose a plug-and-play text-focused super-resolution module to alleviate motion blur and enhance video quality. For ‘linguistically’, a language model is employed to capture intra-text context to mitigate wrongly spelled text predictions. For ‘semantically’, we propose a text-wise semantic reasoning module to model inter-text semantic relationships and reason for better results. The experimental results on multiple VTS benchmarks demonstrate that the proposed VLSpotter outperforms the existing state-of-the-art methods in end-to-end video text spotting. Xinyan Zu, Haiyang Yu 0004, Bin Li 0015, Xiangyang Xue 0001 |
IJCAI | 4 |
| 2023 | Understanding Depth Map Progressively: Adaptive Distance Interval Separation for Monocular 3d Object DetectionabstractMonocular 3D object detection aims to locate objects in different scenes with just a single image. Due to the absence of depth information, several monocular 3D detection techniques have emerged that rely on auxiliary depth maps from the depth estimation task. There are multiple approaches to understanding the representation of depth maps, including treating them as pseudo-LiDAR point clouds, leveraging implicit end-to-end learning of depth information, or considering them as an image input. However, these methods have certain drawbacks, such as their reliance on the accuracy of estimated depth maps and suboptimal utilization of depth maps due to their image-based nature. While LiDAR-based methods and convolutional neural networks (CNNs) can be utilized for pseudo point clouds and depth maps, respectively, it is always an alternative. In this paper, we propose a framework named the Adaptive Distance Interval Separation Network (ADISN) that adopts a novel perspective on understanding depth maps, as a form that lies between LiDAR and images. We utilize an adaptive separation approach that partitions the depth map into various subgraphs based on distance and treats each of these subgraphs as an individual image for feature extraction. After adaptive separations, each subgraph solely contains pixels within a learned interval range. If there is a truncated object within this range, an evident curved edge will appear, which we can leverage for texture extraction using CNNs to obtain rich depth information in pixels. Meanwhile, to mitigate the inaccuracy of depth estimation, we designed an uncertainty module. To take advantage of both images and depth maps, we use different branches to learn localization detection tasks and appearance tasks separately. Our approach significantly enhances the baseline and outperforms depth-assisted techniques, as shown by our extensive experiments on the KITTI monocular 3D object detection benchmark. Xianhui Cheng, Shoumeng Qiu, Zhikang Zou, Jian Pu, Xiangyang Xue 0001 |
IJCNN | 5 |
| 2023 | Language Guided Robotic Grasping with Fine-Grained InstructionsabstractGiven a single RGB image and the attribute-rich language instructions, this paper investigates the novel problem of using Fine-grained instructions for the Language guided robotic Grasping (FLarG). This problem is made challenging by learning fine-grained language descriptions to ground target objects. Recent advances have been made in visually grounding the objects simply by several coarse attributes [1]. However, these methods have poor performance as they cannot well align the multi-modal features, and do not make the best of recent powerful large pre-trained vision and language models, e.g., CLIP. To this end, this paper proposes a FLarG pipeline including stages of CLIP-guided object localization, and 6-DoF category-level object pose estimation for grasping. Specially, we first take the CLIP-based segmentation model CRIS as the backbone and propose an end-to-end DyCRIS model that uses a novel dynamic mask strategy to well fuse the multi-level language and vision features. Then, the well-trained instance segmentation backbone Mask R-CNN is adopted to further improve the predicted mask of our DyCRIS. Finally, the target object pose is inferred for the robotics grasping by using the recent 6-DoF object pose estimation method. To validate our CLIP-enhanced pipeline, we also construct a validation dataset for our FLarG task and name it RefNOCS. Extensive results on RefNOCS have shown the utility and effectiveness of our proposed method. The project homepage is available at https://sunqiang85.github.ioIFLarG/. Qiang Sun 0008, Yanwei Fu 0001, Xiangyang Xue 0001 |
IROS | 5 |
| 2023 | DeNoising-MOT: Towards Multiple Object Tracking with Severe OcclusionsabstractMultiple object tracking (MOT) tends to become more challenging when severe occlusions occur. In this paper, we analyze the limitations of traditional Convolutional Neural Network-based methods and Transformer-based methods in handling occlusions and propose DNMOT, an end-to-end trainable DeNoising Transformer for MOT. To address the challenge of occlusions, we explicitly simulate the scenarios when occlusions occur. Specifically, we augment the trajectory with noises during training and make our model learn the denoising process in an encoder-decoder architecture, so that our model can exhibit strong robustness and perform well under crowded scenes. Additionally, we propose a Cascaded Mask strategy to better coordinate the interaction between different types of queries in the decoder to prevent the mutual suppression between neighboring trajectories under crowded scenes. Notably, the proposed method requires no additional modules like matching strategy and motion state estimation in inference. We conduct extensive experiments on the MOT17, MOT20, and DanceTrack datasets, and the experimental results show that our method outperforms previous state-of-the-art methods by a clear margin. Teng Fu 0001, Haiyang Yu 0004, Ke Niu 0004, Bin Li 0015, Xiangyang Xue 0001 |
ACM Multimedia | 6 |
| 2023 | Scene Text Segmentation with Text-Focused TransformersabstractText segmentation is a crucial aspect of various text-related tasks, including text erasing, text editing, and font style transfer. In recent years, multiple text segmentation datasets, such as TextSeg focusing on Latin text segmentation and BTS on bilingual text segmentation, have been proposed. However, existing methods either disregard the annotations of text location or directly use pre-trained text detectors. In general, these methods cannot fully utilize the annotations of text location in the datasets. To explicitly incorporate text location information to guide text segmentation, we propose an end-to-end text-focused segmentation framework, where text detection and segmentation are jointly optimized. In the proposed framework, we first extract multi-level global visual features through residual convolution blocks and then predict the mask of text areas using a text detection head. Subsequently, we develop a text-focused module that compels the model to pay more attention to text areas. Specifically, we introduce two types of attention masks to extract corresponding features: text-aware and instance-aware features. Finally, we employ hierarchical Transformer encoders to fuse multi-level features and predict the text mask with a text segmentation head. To evaluate the effectiveness of our method, we conduct experiments on six text segmentation benchmarks. The experimental results demonstrate that the proposed method outperforms the previous state-of-the-art (SOTA) methods by a clear margin in most cases. The code and supplementary materials are available at https://github.com/FudanVI/FudanOCR/tree/main/text-focused-Transformers https://github.com/FudanVI/FudanOCR/tree/main/text-focused-Transformers. Haiyang Yu 0004, Ke Niu 0004, Bin Li 0015, Xiangyang Xue 0001 |
ACM Multimedia | 5 |
| 2023 | Weakly-Supervised Text Instance SegmentationabstractText segmentation is a challenging computer vision task with many downstream applications. Current text segmentation models need to be trained with pixel-level annotations, which requires a lot of labor cost. In this paper, we take the first attempt to perform weakly-supervised text instance segmentation through bridging text recognition and text segmentation. We observe that text recognition models are able to produce the attention localization of each text instance. Based on this observation, we propose a two-stage Text Adaptive Refinement (TAR) module to generate the pseudo labels based on the attention map of a text recognizer. Meanwhile, we develop a text segmentation module to take the rough attention location as input to predict segmentation masks, which are supervised by the aforementioned pseudo labels. In addition, we introduce a mask-augmented contrastive learning by treating the segmentation result as an augmented version of the input text image, thus improving the visual representation and further enhancing the performance of both recognition and segmentation. The experimental results demonstrate that the proposed method outperforms the state-of-the-art (SOTA) weakly-supervised generic segmentation methods by 18.95% and 17.80% in fgIoU on ICDAR13-FST and TextSeg. On MLT-S, COCO-TS and Total-Text, the proposed method achieves about 82% of the fully-supervised methods' performance. When evaluated on instance segmentation, the proposed method exceeds existing SOTA methods by 23.32% and 21.34% on ICDAR13-FST and TextSeg, respectively. Code and Supplementary Materials are available at https://github.com/FudanVI/FudanOCR/tree/main/weakly-text-segmentation. Xinyan Zu, Haiyang Yu 0004, Bin Li 0015, Xiangyang Xue 0001 |
ACM Multimedia | 4 |
| 2023 | Training-free Diffusion Model Adaptation for Variable-Sized Text-to-Image SynthesisabstractDiffusion models (DMs) have recently gained attention with state-of-the-art performance in text-to-image synthesis. Abiding by the tradition in deep learning, DMs are trained and evaluated on the images with fixed sizes. However, users are demanding for various images with specific sizes and various aspect ratio. This paper focuses on adapting text-to-image diffusion models to handle such variety while maintaining visual fidelity. First we observe that, during the synthesis, lower resolution images suffer from incomplete object portrayal, while higher resolution images exhibit repetitively disordered presentation. Next, we establish a statistical relationship indicating that attention entropy changes with token quantity, suggesting that models aggregate spatial information in proportion to image resolution. The subsequent interpretation on our observations is that objects are incompletely depicted due to limited spatial information for low resolutions, while repetitively disorganized presentation arises from redundant spatial information for high resolutions. From this perspective, we propose a scaling factor to alleviate the change of attention entropy and mitigate the defective pattern observed. Extensive experimental results validate the efficacy of the proposed scaling factor, enabling models to achieve better visual effects, image quality, and text alignment. Notably, these improvements are achieved without additional training or fine-tuning techniques. Zhiyu Jin, Xuli Shen, Bin Li 0015, Xiangyang Xue 0001 |
NeurIPS | 4 |
| 2023 | ImpDet: Exploring Implicit Fields for 3D Object DetectionabstractConventional 3D object detection approaches concentrate on bounding boxes representation learning with several parameters, i.e., localization, dimension, and orientation. Despite its popularity and universality, such a straightforward paradigm is sensitive to slight numerical deviations, especially in localization. By exploiting the property that point clouds are naturally captured on the surface of objects along with accurate location and intensity information, we introduce a new perspective that views bounding box regression as an implicit function. This leads to our proposed framework, termed Implicit Detection or ImpDet, which leverages implicit field learning for 3D object detection. Our ImpDet assigns specific values to points in different local 3D spaces, thereby high-quality boundaries can be generated by classifying points inside or outside the boundary. To solve the problem of sparsity on the object surface, we further present a simple yet efficient virtual sampling strategy to not only fill the empty region, but also learn rich semantic features to help refine the boundaries. Extensive experimental results on KITTI and Waymo benchmarks demonstrate the effectiveness and robustness of unifying implicit fields into object detection. Xuelin Qian, Li Wang 0033, Yi Zhu 0001, Li Zhang 0040, Yanwei Fu 0001, Xiangyang Xue 0001 |
WACV | 6 |
| 2023 | One-shot Federated Learning without server-side training
Shangchao Su, Bin Li 0015, Xiangyang Xue 0001 |
Neural Networks | 3 |
| 2023 | H4MER: Human 4D Modeling by Learning Neural Compositional Representation With TransformerabstractDespite the impressive results achieved by deep learning based 3D reconstruction, the techniques of directly learning to model 4D human captures with detailed geometry have been less studied. This work presents a novel neural compositional representation for Human 4D Modeling with transformER (H4MER). Specifically, our H4MER is a compact and compositional representation for dynamic human by exploiting the human body prior from the widely used SMPL parametric model. Thus, H4MER can represent a dynamic 3D human over a temporal span with the codes of shape, initial pose, motion and auxiliaries. A simple yet effective linear motion model is proposed to provide a rough and regularized motion estimation, followed by per-frame compensation for pose and geometry details with the residual encoded in the auxiliary codes. We present a novel Transformer-based feature extractor and conditional GRU decoder to facilitate learning and improve the representation capability. Extensive experiments demonstrate our method is not only effective in recovering dynamic human with accurate motion and detailed geometry, but also amenable to various 4D human related tasks, including monocular video fitting, motion retargeting, 4D completion, and future prediction. Boyan Jiang, Yinda Zhang 0001, Jingyang Huo, Xiangyang Xue 0001, Yanwei Fu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Pixel2Mesh++: 3D Mesh Generation and Refinement From Multi-View ImagesabstractWe study the problem of shape generation in 3D mesh representation from a small number of color images with or without camera poses. While many previous works learn to hallucinate the shape directly from priors, we adopt to further improve the shape quality by leveraging cross-view information with a graph convolution network. Instead of building a direct mapping function from images to 3D shape, our model learns to predict series of deformations to improve a coarse shape iteratively. Inspired by traditional multiple view geometry methods, our network samples nearby area around the initial mesh's vertex locations and reasons an optimal deformation using perceptual feature statistics built from multiple input images. Extensive experiments show that our model produces accurate 3D shapes that are not only visually plausible from the input perspectives, but also well aligned to arbitrary viewpoints. With the help of physically driven architecture, our model also exhibits generalization capability across different semantic categories, and the number of input images. Model analysis experiments show that our model is robust to the quality of the initial mesh and the error of camera pose, and can be combined with a differentiable renderer for test-time optimization. Chao Wen 0001, Yinda Zhang 0001, Chenjie Cao, Zhuwen Li, Xiangyang Xue 0001, Yanwei Fu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2023 | Compositional Scene Representation Learning via Reconstruction: A SurveyabstractVisual scenes are composed of visual concepts and have the property of combinatorial explosion. An important reason for humans to efficiently learn from diverse visual scenes is the ability of compositional perception, and it is desirable for artificial intelligence to have similar abilities. Compositional scene representation learning is a task that enables such abilities. In recent years, various methods have been proposed to apply deep neural networks, which have been proven to be advantageous in representation learning, to learn compositional scene representations via reconstruction, advancing this research direction into the deep learning era. Learning via reconstruction is advantageous because it may utilize massive unlabeled data and avoid costly and laborious data annotation. In this survey, we first outline the current progress on reconstruction-based compositional scene representation learning with deep neural networks, including development history and categorizations of existing methods from the perspectives of the modeling of visual scenes and the inference of scene representations; then provide benchmarks, including an open source toolbox to reproduce the benchmark experiments, of representative methods that consider the most extensively studied problem setting and form the foundation for other methods; and finally discuss the limitations of existing methods and future directions of this research topic. Jinyang Yuan, Tonglin Chen, Bin Li 0015, Xiangyang Xue 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Dynamic Graph Message Passing NetworksabstractModelling long-range dependencies is critical for scene understanding tasks in computer vision. Although convolution neural networks (CNNs) have excelled in many vision tasks, they are still limited in capturing long-range structured relationships as they typically consist of layers of local kernels. A fully-connected graph, such as the self-attention operation in Transformers, is beneficial for such modelling, however, its computational overhead is prohibitive. In this paper, we propose a dynamic graph message passing network, that significantly reduces the computational complexity compared to related works modelling a fully-connected graph. This is achieved by adaptively sampling nodes in the graph, conditioned on the input, for message passing. Based on the sampled nodes, we dynamically predict node-dependent filter weights and the affinity matrix for propagating information between them. This formulation allows us to design a self-attention module, and more importantly a new Transformer-based backbone network, that we use for both image classification pretraining, and for addressing various downstream tasks (e.g. object detection, instance and semantic segmentation). Using this model, we show significant improvements with respect to strong, state-of-the-art baselines on four different tasks. Our approach also outperforms fully-connected graphs while using substantially fewer floating-point operations and parameters. Code and models will be made publicly available at https://github.com/fudan-zvg/DGMN2. Li Zhang 0040, Mohan Chen 0001, Anurag Arnab, Xiangyang Xue 0001, Philip Torr 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Rethinking Local and Global Feature Representation for Dense Prediction
Mohan Chen 0001, Li Zhang 0040, Rui Feng 0001, Xiangyang Xue 0001, Jianfeng Feng |
Pattern Recognit. | 4 |
| 2023 | Multi-view Shape Generation for a 3D Human-like BodyabstractThree-dimensional (3D) human-like body reconstruction via a single RGB image has attracted significant research attention recently. Most of the existing methods rely on the Skinned Multi-Person Linear model and thus can only predict unified human bodies. Moreover, meshes reconstructed by current methods sometimes perform well from a canonical view but not from other views, as the reconstruction process is commonly supervised by only a single view. To address these limitations, this article proposes a multi-view shape generation network for a 3D human-like body. Particularly, we propose a coarse-to-fine learning model that gradually deforms a template body toward the ground truth body. Our model utilizes the information of multi-view renderings and corresponding 3D vertex transformation as supervision. Such supervision will help to generate 3D bodies well aligned to all views. To accurately operate mesh deformation, a graph convolutional network structure is introduced to support the shape generation from 3D vertex representation. Additionally, a graph up-pooling operation is designed over the intermediate representations of the graph convolutional network, and thus our model can generate 3D shapes with higher resolution. Novel loss functions are employed to help optimize the whole multi-view generation model, resulting in smoother surfaces. In addition, two multi-view human body datasets are produced and contributed to the community. Extensive experiments conducted on the benchmark datasets demonstrate the efficacy of our model over the competitors. Chilam Cheang, Yanwei Fu 0001, Xiangyang Xue 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2023 | When, Where and How Does it Fail? A Spatial-Temporal Visual Analytics Approach for Interpretable Object Detection in Autonomous DrivingabstractArguably the most representative application of artificial intelligence, autonomous driving systems usually rely on computer vision techniques to detect the situations of the external environment. Object detection underpins the ability of scene understanding in such systems. However, existing object detection algorithms often behave as a black box, so when a model fails, no information is available on When, Where and How the failure happened. In this paper, we propose a visual analytics approach to help model developers interpret the model failures. The system includes the micro- and macro-interpreting modules to address the interpretability problem of object detection in autonomous driving. The micro-interpreting module extracts and visualizes the features of a convolutional neural network (CNN) algorithm with density maps, while the macro-interpreting module provides spatial-temporal information of an autonomous driving vehicle and its environment. With the situation awareness of the spatial, temporal and neural network information, our system facilitates the understanding of the results of object detection algorithms, and helps the model developers better understand, tune and develop the models. We use real-world autonomous driving data to perform case studies by involving domain experts in computer vision and autonomous driving to evaluate our system. The results from our interviews with them show the effectiveness of our approach. Zhaoyu Zhou, Chengshun Wang, Yijie Hou, Li Zhang 0040, Xiangyang Xue 0001, Michael Kamp, Xiaolong Zhang 0001, Siming Chen 0001 |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2022 | Text Gestalt: Stroke-Aware Scene Text Image Super-resolutionabstractIn the last decade, the blossom of deep learning has witnessed the rapid development of scene text recognition. However, the recognition of low-resolution scene text images remains a challenge. Even though some super-resolution methods have been proposed to tackle this problem, they usually treat text images as general images while ignoring the fact that the visual quality of strokes (the atomic unit of text) plays an essential role for text recognition. According to Gestalt Psychology, humans are capable of composing parts of details into the most similar objects guided by prior knowledge. Likewise, when humans observe a low-resolution text image, they will inherently use partial stroke-level details to recover the appearance of holistic characters. Inspired by Gestalt Psychology, we put forward a Stroke-Aware Scene Text Image Super-Resolution method containing a Stroke-Focused Module (SFM) to concentrate on stroke-level internal structures of characters in text images. Specifically, we attempt to design rules for decomposing English characters and digits at stroke-level, then pre-train a text recognizer to provide stroke-level attention maps as positional clues with the purpose of controlling the consistency between the generated super-resolution image and high-resolution ground truth. The extensive experimental results validate that the proposed method can indeed generate more distinguishable images on TextZoom and manually constructed Chinese character dataset Degraded-IC13. Furthermore, since the proposed SFM is only used to provide stroke-level guidance when training, it will not bring any time overhead during the test phase. Code is available at https://github.com/FudanVI/FudanOCR/tree/main/text-gestalt. Jingye Chen, Haiyang Yu 0004, Jianqi Ma, Bin Li 0015, Xiangyang Xue 0001 |
AAAI | 5 |
| 2022 | Unsupervised Learning of Compositional Scene Representations from Multiple Unspecified ViewpointsabstractVisual scenes are extremely rich in diversity, not only because there are infinite combinations of objects and background, but also because the observations of the same scene may vary greatly with the change of viewpoints. When observing a visual scene that contains multiple objects from multiple viewpoints, humans are able to perceive the scene in a compositional way from each viewpoint, while achieving the so-called ``object constancy'' across different viewpoints, even though the exact viewpoints are untold. This ability is essential for humans to identify the same object while moving and to learn from vision efficiently. It is intriguing to design models that have the similar ability. In this paper, we consider a novel problem of learning compositional scene representations from multiple unspecified viewpoints without using any supervision, and propose a deep generative model which separates latent representations into a viewpoint-independent part and a viewpoint-dependent part to solve this problem. To infer latent representations, the information contained in different viewpoints is iteratively integrated by neural networks. Experiments on several specifically designed synthetic datasets have shown that the proposed method is able to effectively learn from multiple unspecified viewpoints. Jinyang Yuan, Bin Li 0015, Xiangyang Xue 0001 |
AAAI | 3 |
| 2022 | QS-Craft: Learning to Quantize, Scrabble and Craft for Conditional Human Motion Animation
Yuxin Hong, Xuelin Qian, Simian Luo, Guodong Guo, Xiangyang Xue 0001, Yanwei Fu 0001 |
ACCV (6) | 5 |
| 2022 | Co-attention Aligned Mutual Cross-Attention for Cloth-Changing Person Re-identification
Qizao Wang, Xuelin Qian, Yanwei Fu 0001, Xiangyang Xue 0001 |
ACCV (5) | 4 |
| 2022 | Density-preserving Deep Point Cloud CompressionabstractLocal density of point clouds is crucial for representing local details, but has been overlooked by existing point cloud compression methods. To address this, we propose a novel deep point cloud compression method that preserves local density information. Our method works in an auto-encoder fashion: the encoder downsamples the points and learns point-wise features, while the decoder upsamples the points using these features. Specifically, we propose to encode local geometry and density with three embeddings: density embedding, local position embedding and ancestor embedding. During the decoding, we explicitly predict the upsampling factor for each point, and the directions and scales of the upsampled points. To mitigate the clustered points issue in existing methods, we design a novel sub-point convolution layer, and an upsampling block with adaptive scale. Furthermore, our method can also compress point-wise attributes, such as normal. Extensive qualitative and quantitative results on SemanticKITTI and ShapeNet demonstrate that our method achieves the state-of-the-art rate-distortion trade-off. Xinlin Ren, Danhang Tang, Yinda Zhang 0001, Xiangyang Xue 0001, Yanwei Fu 0001 |
CVPR | 5 |
| 2022 | H4D: Human 4D Modeling by Learning Neural Compositional RepresentationabstractDespite the impressive results achieved by deep learning based 3D reconstruction, the techniques of directly learning to model 4D human captures with detailed geometry have been less studied. This work presents a novel framework that can effectively learn a compact and compositional representation for dynamic human by exploiting the human body prior from the widely used SMPL parametric model. Particularly, our representation, named H4D, represents a dynamic 3D human over a temporal span with the SMPL parameters of shape and initial pose, and latent codes encoding motion and auxiliary information. A simple yet effective linear motion model is proposed to provide a rough and regularized motion estimation, followed by perframe compensation for pose and geometry details with the residual encoded in the auxiliary code. Technically, we introduce novel GRU-based architectures to facilitate learning and improve the representation capability. Extensive experiments demonstrate our method is not only efficacy in recovering dynamic human with accurate motion and detailed geometry, but also amenable to various 4D human related tasks, including motion retargeting, motion completion and future prediction. Boyan Jiang, Yinda Zhang 0001, Xingkui Wei, Xiangyang Xue 0001, Yanwei Fu 0001 |
CVPR | 4 |
| 2022 | SAR-Net: Shape Alignment and Recovery Network for Category-level 6D Object Pose and Size EstimationabstractGiven a single scene image, this paper proposes a method of Category-level 6D Object Pose and Size Estimation (COPSE) from the point cloud of the target object, without external real pose-annotated training data. Specifically, beyond the visual cues in RGB images, we rely on the shape information predominately from the depth (D) channel. The key idea is to explore the shape alignment of each instance against its corresponding category-level template shape, and the symmetric correspondence of each object category for estimating a coarse 3D object shape. Our framework deforms the point cloud of the category-level template shape to align the observed instance point cloud for implicitly representing its 3D rotation. Then we model the symmetric correspondence by predicting symmetric point cloud from the partially observed point cloud. The concatenation of the observed point cloud and symmetric one reconstructs a coarse object shape, thus facilitating object center (3D translation) and 3D size estimation. Extensive experiments on the category-level NOCS benchmark demonstrate that our lightweight model still competes with state-of-the-art approaches that require labeled real-world images. We also deploy our approach to a physical Baxter robot to perform grasping tasks on unseen but category-known instances, and the results further validate the efficacy of our proposed model. Code and pre-trained models are available on the project webpage11Project webpage. https://hetolin.github.io/SAR-Net. Zichang Liu, Chilam Cheang, Yanwei Fu 0001, Guodong Guo, Xiangyang Xue 0001 |
CVPR | 6 |
| 2022 | DST: Dynamic Substitute Training for Data-free Black-box AttackabstractWith the wide applications of deep neural network models in various computer vision tasks, more and more works study the model vulnerability to adversarial examples. For data-free black box attack scenario, existing methods are inspired by the knowledge distillation, and thus usually train a substitute model to learn knowledge from the target model using generated data as input. However, the substitute model always has a static network structure, which limits the attack ability for various target models and tasks. In this paper, we propose a novel dynamic substitute training attack method to encourage substitute model to learn better and faster from the target model. Specifically, a dynamic substitute structure learning strategy is proposed to adaptively generate optimal substitute model structure via a dy-namic gate according to different target models and tasks. Moreover, we introduce a task-driven graph-based structure information learning constrain to improve the quality of generated training data, and facilitate the substitute model learning structural relationships from the target model multiple outputs. Extensive experiments have been conducted to verify the efficacy of the proposed attack method, which can achieve better performance compared with the state-of-the-art competitors on several datasets. Project page: https://wxwangiris.github.io/DST Wenxuan Wang 0003, Xuelin Qian, Yanwei Fu 0001, Xiangyang Xue 0001 |
CVPR | 4 |
| 2022 | LoRD: Local 4D Implicit Representation for High-Fidelity Dynamic Human Modeling
Boyan Jiang, Xinlin Ren, Mingsong Dou, Xiangyang Xue 0001, Yanwei Fu 0001, Yinda Zhang 0001 |
ECCV (26) | 4 |
| 2022 | RCLane: Relay Chain Prediction for Lane Detection
Shenghua Xu, Xinyue Cai, Li Zhang 0040, Hang Xu 0004, Yanwei Fu 0001, Xiangyang Xue 0001 |
ECCV (38) | 7 |
| 2022 | High-Fidelity Portrait Editing Via Exploring Differentiable Guided Sketches from the Latent SpaceabstractThis paper studies the task of sketch-guided high-fidelity portrait editing. Advanced unconditional generators, such as StyleGAN, can generate a high-quality portrait image with great diversity. In previous researches, StyleGAN has successfully been utilized for color-guided image editing through latent vector optimization. Nonetheless, passing sketch information to the generating model directly is nontrivial. To this end, we present an algorithm that addresses the problem of well controlling the generation process via differentiable guided sketches from latent space. Specifically, we re-purpose the classic operator – eXtended difference-of-Gaussians (XDoG) that derives differentiable sketches from images. We also propose a multi-scale sketch loss assisted with which can finally guide the model follow the guidance sketch to generate. Extensive experiments validate the efficacy of our model in sketch-guided editing. We show that the quality of produced images is better than that of competitors. Chengrong Wang, Chenjie Cao, Yanwei Fu 0001, Xiangyang Xue 0001 |
ICASSP | 4 |
| 2022 | Towards Scalable and Fast Distributionally Robust Optimization for Data-Driven Deep LearningabstractWe introduce a scalable and fast method for solving distributionally robust optimization (DRO). Previous works have demonstrated that DRO outperforms empirical risk on a collection of inconsistent distribution of test data (the property of “uncertainty set”). However, DRO is hard to be applied for large-scale datasets and large parameterized model, due to the datapoint-level and non-differentiable objective function. In this paper, we formalize the DRO problem with the supremum of a family of subgroup-level loss functions. Subgroup loss is the cost function of partitioned uncertainty set. Then we implement the maximum of subgroup loss as the objective function and update model parameters by reweighting the descent direction, calculated from a differentiable objective function. Experimental results unveil that large parameterized models with the proposed method successfully adapt to uncertainty set whether the distribution contains out-of-domain or imbalanced property. Remarkably, with the explored reweighting strategy, the proposed algorithm effectively achieves competitive performance and robustness. Xuli Shen, Qing Xu 0017, Weifeng Ge, Xiangyang Xue 0001 |
ICDM | 5 |
| 2022 | Learning 6-DoF Object Poses to Grasp Category-Level Objects by Language InstructionsabstractThis paper studies the task of any objects grasping from the known categories by free-form language instructions. This task demands the technique in computer vision, natural language processing, and robotics. We bring these disciplines together on this open challenge, which is essential to human-robot interaction. Critically, the key challenge lies in inferring the category of objects from linguistic instructions and accurately estimating the 6-DoF information of unseen objects from the known classes. In contrast, previous works focus on inferring the pose of object candidates at the instance level. This significantly limits its applications in real-world scenarios. In this paper, we propose a language-guided 6-DoF category-level object localization model to achieve robotic grasping by comprehending human intention. To this end, we propose a novel two-stage method. Particularly, the first stage grounds the target in the RGB image through language description of names, attributes, and spatial relations of objects. The second stage extracts and segments point clouds from the cropped depth image and estimates the full 6-DoF object pose at category-level. Under such a manner, our approach can locate the specific object by following human instructions, and estimate the full 6-DoF pose of a category-known but unseen instance which is not utilized for training the model. Extensive experimental results show that our method is competitive with the state-of-the-art language-conditioned grasp method. Importantly, we deploy our approach on a physical robot to validate the usability of our framework in real-world applications. Please refer to the supplementary for the demo videos of our robot experiments. Chilam Cheang, Yanwei Fu 0001, Xiangyang Xue 0001 |
ICRA | 4 |
| 2022 | I Know What You Draw: Learning Grasp Detection Conditioned on a Few Freehand SketchesabstractIn this paper, we are interested in the problem of generating target grasps by understanding freehand sketches. The sketch is useful for the persons who cannot formulate language and the cases where a textual description is not available on the fly. However, very few works are aware of the usability of this novel interactive way between humans and robots. To this end, we propose a method to generate a potential grasp configuration relevant to the sketch -depicted objects. Due to the inherent ambiguity of sketches with abstract details, we take the advantage of the graph by incorporating the structure of the sketch to enhance the representation ability. This graph-represented sketch is further validated to improve the generalization of the network, capable of learning the sketch-queried grasp detection by using a small collection (around 100 samples) of hand-drawn sketches. Additionally, our model is trained and tested in an end-to-end manner which is easy to be implemented in real-world applications. Experiments on the multi-object VMRD and GraspNet-1Billion datasets demonstrate the good generalization of the proposed method. The physical robot experiments confirm the utility of our method in object-cluttered scenes. Chilam Cheang, Yanwei Fu 0001, Xiangyang Xue 0001 |
ICRA | 4 |
| 2022 | Local Slot Attention for Vision and Language NavigationabstractVision-and-language navigation (VLN), a frontier study aiming to pave the way for general-purpose robots, has been a hot topic in the computer vision and natural language processing community. The VLN task requires an agent to navigate to a goal location following natural language instructions in unfamiliar environments. Yifeng Zhuang, Qiang Sun 0008, Yanwei Fu 0001, Lifeng Chen, Xiangyang Xue 0001 |
ICMR | 5 |
| 2022 | Chinese Character Recognition with Augmented Character Profile MatchingabstractChinese character recognition (CCR) has drawn continuous research interest due to its wide applications. After decades of study, there still exist several challenges,e.g., different characters with similar appearance and the one-to-many problem. There is no unified solution to the above challenges as previous methods tend to address these problems separately. In this paper, we propose a Chinese character recognition method named Augmented Character Profile Matching (ACPM), which utilizes a collection of character knowledge from three decomposition levels to recognize Chinese characters. Specifically, the feature maps of each character image are utilized as the character-level knowledge. In addition, we introduce a radical-stroke counting module (RSC) to help produce augmented character profiles, including the number of radicals, the number of strokes, and the total length of strokes, which characterize the character more comprehensively. The feature maps of the character image and the outputs of the RSC module are collected to constitute a character profile for selecting the closest candidate character through joint matching. The experimental results show that the proposed method outperforms the state-of-the-art methods on both the ICDAR 2013 and CTW datasets by 0.35% and 2.23%, respectively. Moreover, it also clearly outperforms the compared methods in the zero-shot settings. Code is available at https://github.com/FudanVI/FudanOCR/tree/main/character-profile-matching. Xinyan Zu, Haiyang Yu 0004, Bin Li 0015, Xiangyang Xue 0001 |
ACM Multimedia | 4 |
| 2022 | DEMoS: a deep learning-based ensemble approach for predicting the molecular subtypes of gastric adenocarcinomas from histopathological imagesabstractMOTIVATION: The molecular subtyping of gastric cancer (adenocarcinoma) into four main subtypes based on integrated multiomics profiles, as proposed by The Cancer Genome Atlas (TCGA) initiative, represents an effective strategy for patient stratification. However, this approach requires the use of multiple technological platforms, and is quite expensive and time-consuming to perform. A computational approach that uses histopathological image data to infer molecular subtypes could be a practical, cost- and time-efficient complementary tool for prognostic and clinical management purposes. RESULTS: Here, we propose a deep learning ensemble approach (called DEMoS) capable of predicting the four recognized molecular subtypes of gastric cancer directly from histopathological images. DEMoS achieved tile-level area under the receiver-operating characteristic curve (AUROC) values of 0.785, 0.668, 0.762 and 0.811 for the prediction of these four subtypes of gastric cancer [i.e. (i) Epstein-Barr (EBV)-infected, (ii) microsatellite instability (MSI), (iii) genomically stable (GS) and (iv) chromosomally unstable tumors (CIN)] using an independent test dataset, respectively. At the patient-level, it achieved AUROC values of 0.897, 0.764, 0.890 and 0.898, respectively. Thus, these four subtypes are well-predicted by DEMoS. Benchmarking experiments further suggest that DEMoS is able to achieve an improved classification performance for image-based subtyping and prevent model overfitting. This study highlights the feasibility of using a deep learning ensemble-based method to rapidly and reliably subtype gastric cancer (adenocarcinoma) solely using features from histopathological images. AVAILABILITY AND IMPLEMENTATION: All whole slide images used in this study was collected from the TCGA database. This study builds upon our previously published HEAL framework, with related documentation and tutorials available at http://heal.erc.monash.edu.au. The source code and related models are freely accessible at https://github.com/Docurdt/DEMoS.git. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yanan Wang 0003, Changyuan Hu, Terry Kwok, Christopher A. Bain, Xiangyang Xue 0001, Robin B. Gasser, Geoffrey I. Webb, Alex Boussioutas, Xian Shen, Roger J. Daly, Jiangning Song |
Bioinform. | 5 |
| 2022 | Learning the Compositional Domains for Generalized Zero-shot Learning
Hanze Dong, Yanwei Fu 0001, Sung Ju Hwang, Leonid Sigal, Xiangyang Xue 0001 |
Comput. Vis. Image Underst. | 5 |
| 2022 | HandO: a hybrid 3D hand-object reconstruction model for unknown objects
Chilam Cheang, Yanwei Fu 0001, Xiangyang Xue 0001 |
Multim. Syst. | 4 |
| 2022 | AGO-Net: Association-Guided 3D Point Cloud Object Detection NetworkabstractThe human brain can effortlessly recognize and localize objects, whereas current 3D object detection methods based on LiDAR point clouds still report inferior performance for detecting occluded and distant objects: The point cloud appearance varies greatly due to occlusion, and has inherent variance in point densities along the distance to sensors. Therefore, designing feature representations robust to such point clouds is critical. Inspired by human associative recognition, we propose a novel 3D detection framework that associates intact features for objects via domain adaptation. We bridge the gap between the perceptual domain, where features are derived from real scenes with sub-optimal representations, and the conceptual domain, where features are extracted from augmented scenes that consist of non-occlusion objects with rich detailed information. A feasible method is investigated to construct conceptual scenes without external datasets. We further introduce an attention-based re-weighting module that adaptively strengthens the feature adaptation of more informative regions. The network's feature enhancement ability is exploited without introducing extra cost during inference, which is plug-and-play in various 3D detection frameworks. We achieve new state-of-the-art performance on the KITTI 3D detection benchmark in both accuracy and speed. Experiments on nuScenes and Waymo datasets also validate the versatility of our method. Liang Du 0004, Xiaoqing Ye, Xiao Tan 0001, Edward Johns, Errui Ding, Xiangyang Xue 0001, Jianfeng Feng |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2022 | Visual Evaluation for Autonomous DrivingabstractAutonomous driving technologies often use state-of-the-art artificial intelligence algorithms to understand the relationship between the vehicle and the external environment, to predict the changes of the environment, and then to plan and control the behaviors of the vehicle accordingly. The complexity of such technologies makes it challenging to evaluate the performance of autonomous driving systems and to find ways to improve them. The current approaches to evaluating such autonomous driving systems largely use a single score to indicate the overall performance of a system, but domain experts have difficulties in understanding how individual components or algorithms in an autonomous driving system may contribute to the score. To address this problem, we collaborate with domain experts on autonomous driving algorithms, and propose a visual evaluation method for autonomous driving. Our method considers the data generated in all components during the whole process of autonomous driving, including perception results, planning routes, prediction of obstacles, various controlling parameters, and evaluation of comfort. We develop a visual analytics workflow to integrate an evaluation mathematical model with adjustable parameters, support the evaluation of the system from the level of the overall performance to the level of detailed measures of individual components, and to show both evaluation scores and their contributing factors. Our implemented visual analytics system provides an overview evaluation score at the beginning and shows the animation of the dynamic change of the scores at each period. Experts can interactively explore the specific component at different time periods and identify related factors. With our method, domain experts not only learn about the performance of an autonomous driving system, but also identify and access the problematic parts of each component. Our visual evaluation system can be applied to the autonomous driving simulation system and used for various evaluation cases. The results of using our system in some simulation cases and the feedback from involved domain experts confirm the usefulness and efficiency of our method in helping people gain in-depth insight into autonomous driving systems. Yijie Hou, Chengshun Wang, Xiangyang Xue 0001, Xiaolong Zhang 0001, Siming Chen 0001 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2022 | Metaverse: Perspectives from graphics, interactions and visualizationabstractThe metaverse is a visual world that blends the physical world and digital world. At present, the development of the metaverse is still in the early stage, and there lacks a framework for the visual construction and exploration of the metaverse. In this paper, we propose a framework that summarizes how graphics, interaction, and visualization techniques support the visual construction of the metaverse and user-centric exploration. We introduce three kinds of visual elements that compose the metaverse and the two graphical construction methods in a pipeline. We propose a taxonomy of interaction technologies based on interaction tasks, user actions, feedback and various sensory channels, and a taxonomy of visualization techniques that assist user awareness. Current potential applications and future opportunities are discussed in the context of visual construction and exploration of the metaverse. We hope this paper can provide a stepping stone for further research in the area of graphics, interaction and visualization in the metaverse. Yuheng Zhao, Jinjing Jiang, Yi Chen 0007, Richen Liu, Yalong Yang 0001, Xiangyang Xue 0001, Siming Chen 0001 |
Vis. Informatics | 6 |
| 2021 | Raven's Progressive Matrices Completion with Latent Gaussian Process PriorsabstractAbstract reasoning ability is fundamental to human intelligence. It enables humans to uncover relations among abstract concepts and further deduce implicit rules from the relations. As a well-known abstract visual reasoning task, Raven's Progressive Matrices (RPM) are widely used in human IQ tests. Although extensive research has been conducted on RPM solvers with machine intelligence, few studies have considered further advancing the standard answer-selection (classification) problem to a more challenging answer-painting (generating) problem, which can verify whether the model has indeed understood the implicit rules. In this paper we aim to solve the latter one by proposing a deep latent variable model, in which multiple Gaussian processes are employed as priors of latent variables to separately learn underlying abstract concepts from RPMs; thus the proposed model is interpretable in terms of concept-specific latent variables. The latent Gaussian process also provides an effective way of extrapolation for answer painting based on the learned concept-changing rules. We evaluate the proposed model on RPM-like datasets with multiple continuously-changing visual concepts. Experimental results demonstrate that our model requires only few training samples to paint high-quality answers, generate novel RPM panels, and achieve interpretability through concept-specific latent variables. Bin Li 0015, Xiangyang Xue 0001 |
AAAI | 3 |
| 2021 | Knowledge-Guided Object Discovery with Acquired Deep ImpressionsabstractWe present a framework called Acquired Deep Impressions (ADI) which continuously learns knowledge of objects as ``impressions'' for compositional scene understanding. In this framework, the model first acquires knowledge from scene images containing a single object in a supervised manner, and then continues to learn from novel multi-object scene images which may contain objects that have not been seen before without any further supervision, under the guidance of the learned knowledge as humans do. By memorizing impressions of objects into parameters of neural networks and applying the generative replay strategy, the learned knowledge can be reused to generate images with pseudo-annotations and in turn assist the learning of novel scenes. The proposed ADI framework focuses on the acquisition and utilization of knowledge, and is complementary to existing deep generative models proposed for compositional scene representation. We adapt a base model to make it fall within the ADI framework and conduct experiments on two types of datasets. Empirical results suggest that the proposed framework is able to effectively utilize the acquired impressions and improve the scene decomposition performance. Jinyang Yuan, Bin Li 0015, Xiangyang Xue 0001 |
AAAI | 3 |
| 2021 | Rethinking local and global feature representation for semantic segmentation
Mohan Chen 0001, Xinxuan Zhao, Bingfei Fu, Li Zhang 0040, Xiangyang Xue 0001 |
BMVC | 5 |
| 2021 | Learning Dynamic Alignment via Meta-Filter for Few-Shot LearningabstractFew-shot learning (FSL), which aims to recognise new classes by adapting the learned knowledge with extremely limited few-shot (support) examples, remains an important open problem in computer vision. Most of the existing methods for feature alignment in few-shot learning only consider image-level or spatial-level alignment while omitting the channel disparity. Our insight is that these methods would lead to poor adaptation with redundant matching, and leveraging channel-wise adjustment is the key to well adapting the learned knowledge to new classes. Therefore, in this paper, we propose to learn a dynamic alignment, which can effectively highlight both query regions and channels according to different local support information. Specifically, this is achieved by first dynamically sampling the neighbourhood of the feature position conditioned on the input few shot, based on which we further predict a both position-dependent and channel-dependent Dynamic Meta-filter. The filter is used to align the query feature with position-specific and channel-specific knowledge. Moreover, we adopt Neural Ordinary Differential Equation (ODE) to enable a more accurate control of the alignment. In such a sense our model is able to better capture fine-grained semantic context of the few-shot example and thus facilitates dynamical knowledge adaptation for few-shot learning. The resulting framework establishes the new state-of-the-arts on major few-shot visual recognition benchmarks, including miniImageNet and tieredImageNet. Chengming Xu 0001, Yanwei Fu 0001, Chen Liu 0030, Chengjie Wang 0001, Feiyue Huang, Li Zhang 0040, Xiangyang Xue 0001 |
CVPR | 8 |
| 2021 | Scene Text Telescope: Text-Focused Scene Image Super-ResolutionabstractImage super-resolution, which is often regarded as a preprocessing procedure of scene text recognition, aims to recover the realistic features from a low-resolution text image. It has always been challenging due to large variations in text shapes, fonts, backgrounds, etc. However, most existing methods employ generic super-resolution frameworks to handle scene text images while ignoring text-specific properties such as text-level layouts and character-level details. In this paper, we establish a text-focused super-resolution framework, called Scene Text Telescope (STT). In terms of text-level layouts, we propose a Transformer-Based Super-Resolution Network (TBSRN) containing a Self-Attention Module to extract sequential information, which is robust to tackle the texts in arbitrary orientations. In terms of character-level details, we propose a Position-Aware Module and a Content-Aware Module to highlight the position and the content of each character. By observing that some characters look indistinguishable in low-resolution conditions, we use a weighted cross-entropy loss to tackle this problem. We conduct extensive experiments, including text recognition with pre-trained recognizers and image quality evaluation, on TextZoom and several scene text recognition benchmarks to assess the super-resolution images. The experimental results show that our STT can indeed generate text-focused super-resolution images and outperform the existing methods in terms of recognition accuracy. Jingye Chen, Bin Li 0015, Xiangyang Xue 0001 |
CVPR | 3 |
| 2021 | Learning Compositional Representation for 4D Captures With Neural ODEabstractLearning based representation has become the key to the success of many computer vision systems. While many 3D representations have been proposed, it is still an unaddressed problem how to represent a dynamically changing 3D object. In this paper, we introduce a compositional representation for 4D captures, i.e. a deforming 3D object over a temporal span, that disentangles shape, initial state, and motion respectively. Each component is represented by a latent code via a trained encoder. To model the motion, a neural Ordinary Differential Equation (ODE) is trained to update the initial state conditioned on the learned motion code, and a decoder takes the shape code and the updated state code to reconstruct the 3D model at each time stamp. To this end, we propose an Identity Exchange Training (IET) strategy to encourage the network to learn effectively decoupling each component. Extensive experiments demonstrate that the proposed method outperforms existing state-of-the-art deep learning based methods on 4D reconstruction, and significantly improves on various tasks, including motion transfer and completion. Boyan Jiang, Yinda Zhang 0001, Xingkui Wei, Xiangyang Xue 0001, Yanwei Fu 0001 |
CVPR | 4 |
| 2021 | Depth-Conditioned Dynamic Message Propagation for Monocular 3D Object DetectionabstractThe objective of this paper is to learn context- and depth- aware feature representation to solve the problem of monocular 3D object detection. We make following contributions: (i) rather than appealing to the complicated pseudo-LiDAR based approach, we propose a depth-conditioned dynamic message propagation (DDMP) network to effectively integrate the multi-scale depth information with the image context; (ii) this is achieved by first adaptively sampling context-aware nodes in the image context and then dynamically predicting hybrid depth-dependent filter weights and affinity matrices for propagating information; (Hi) by augmenting a center-aware depth encoding (CDE) task, our method successfully alleviates the inaccurate depth prior; (iv) we thoroughly demonstrate the effectiveness of our proposed approach and show state-of-the-art results among the monocular-based approaches on the KITTI benchmark dataset. Particularly, we rank 1stin the highly competitive KITTI monocular 3D object detection track on the submission day (November 16th, 2020). Code and models are released at https: //github.com/fudan-zvg/DDMP Li Wang 0033, Liang Du 0004, Xiaoqing Ye, Yanwei Fu 0001, Guodong Guo, Xiangyang Xue 0001, Jianfeng Feng, Li Zhang 0040 |
CVPR | 6 |
| 2021 | Delving into Data: Effectively Substitute Training for Black-box AttackabstractDeep models have shown their vulnerability when processing adversarial samples. As for the black-box attack, without access to the architecture and weights of the attacked model, training a substitute model for adversarial attacks has attracted wide attention. Previous substitute training approaches focus on stealing the knowledge of the target model based on real training data or synthetic data, without exploring what kind of data can further improve the transferability between the substitute and target models. In this paper, we propose a novel perspective substitute training that focuses on designing the distribution of data used in the knowledge stealing process. More specifically, a diverse data generation module is proposed to synthesize large-scale data with wide distribution. And adversarial substitute training strategy is introduced to focus on the data distributed near the decision boundary. The combination of these two modules can further boost the consistency of the substitute model and target model, which greatly improves the effectiveness of adversarial attack. Extensive experiments demonstrate the efficacy of our method against state-of-the-art competitors under non-target and target at-tack settings. Detailed visualization and analysis are also provided to help understand the advantage of our method. Wenxuan Wang 0003, Bangjie Yin, Taiping Yao, Li Zhang 0040, Yanwei Fu 0001, Shouhong Ding, Feiyue Huang, Xiangyang Xue 0001 |
CVPR | 9 |
| 2021 | The Devil is in the Task: Exploiting Reciprocal Appearance-Localization Features for Monocular 3D Object DetectionabstractLow-cost monocular 3D object detection plays a fundamental role in autonomous driving, whereas its accuracy is still far from satisfactory. In this paper, we dig into the 3D object detection task and reformulate it as the sub-tasks of object localization and appearance perception, which benefits to a deep excavation of reciprocal information underlying the entire task. We introduce a Dynamic Feature Reflecting Network, named DFR-Net, which contains two novel standalone modules: (i) the Appearance-Localization Feature Reflecting module (ALFR) that first separates task-specific features and then self-mutually reflects the reciprocal features; (ii) the Dynamic Intra-Trading module (DIT) that adaptively realigns the training processes of various sub-tasks via a self-learning manner. Extensive experiments on the challenging KITTI dataset demonstrate the effectiveness and generalization of DFR-Net. We rank 1stamong all the monocular 3D object detectors in the KITTI test set (till March 16th, 2021). The proposed method is also easy to be plug-and-play in many cutting-edge 3D detection frameworks at negligible cost to boost performance. The code will be made publicly available. Zhikang Zou, Xiaoqing Ye, Liang Du 0004, Xianhui Cheng, Xiao Tan 0001, Li Zhang 0040, Jianfeng Feng, Xiangyang Xue 0001, Errui Ding |
ICCV | 8 |
| 2021 | Depth-Guided AdaIN and Shift Attention Network for Vision-And-Language NavigationabstractVisual Language Navigation (VLN) is the grand goal of AI, which enables the agent to act by the language instructions from humans. In VLN task, the agent learns to search for a specific region described by the instructions in the training environments, and performs the navigation in the unseen environments. Normally, there exists a large domain gap be-tween the seen and unseen environments. Numerous works have been put on data augmentation and designing new loss in such a multi-task navigation setting. However, as a spatial and temporal searching task, a valuable signal source for the navigation – depth has not yet fully explored and thus been ignored in previous efforts. Typically, the current models lack the ability to capture the relative spatial directions to the grounding view. To address these issues, we propose an environment adaptive method based on a Depth-guided Adaptive Instance Normalization (DG-AdaIN) module to adjust the RGB features in term of the depth features, and develop a shift attention module to model the relative direct information in the attention map. Extensive experiments have validated the efficacy of our method on the benchmark dataset. Qiang Sun 0008, Yifeng Zhuang, Zhengqing Chen, Yanwei Fu 0001, Xiangyang Xue 0001 |
ICME | 5 |
| 2021 | Distance Restricted Transformer Encoder for Multi-Label ClassificationabstractMulti-label image classification is a fundamental but challenging task in Multimedia community. It aims to predict a set of labels presented in an image. Great progress has been made by exploring convolutional neural network with binary cross-entropy loss recently. However, conventional approaches are limited to highlight the key visual contents associated with target labels and pay little attention to confining the distances between visual and positive/negative label representations. To target these aspects, we firstly introduce a variant transformer encoder model for acquiring the underlying and crucial visual information related to ground truth labels. Specifically, a novel primal feature guided net is designed to maintain the original visual features during encoding process. Secondly, we exploit a distance restricted learning strategy in a common semantic space to shrink the distances of images with positive labels while expand with the negative ones during training stage. Extensive experiments are executed on MSCOCO and WIDER Attribute datasets and outstanding performance is achieved compared with other state-of-the-art models. Yandong Guo, Yanwei Fu 0001, Xiangyang Xue 0001 |
ICME | 6 |
| 2021 | Zero-Shot Chinese Character Recognition with Stroke-Level DecompositionabstractChinese character recognition has attracted much research interest due to its wide applications. Although it has been studied for many years, some issues in this field have not been completely resolved yet, \textit{e.g.} the zero-shot problem. Previous character-based and radical-based methods have not fundamentally addressed the zero-shot problem since some characters or radicals in test sets may not appear in training sets under a data-hungry condition. Inspired by the fact that humans can generalize to know how to write characters unseen before if they have learned stroke orders of some characters, we propose a stroke-based method by decomposing each character into a sequence of strokes, which are the most basic units of Chinese characters. However, we observe that there is a one-to-many relationship between stroke sequences and Chinese characters. To tackle this challenge, we employ a matching-based strategy to transform the predicted stroke sequence to a specific character. We evaluate the proposed method on handwritten characters, printed artistic characters, and scene characters. The experimental results validate that the proposed method outperforms existing methods on both character zero-shot and radical zero-shot tasks. Moreover, the proposed method can be easily generalized to other languages whose characters can be decomposed into strokes. Jingye Chen, Bin Li 0015, Xiangyang Xue 0001 |
IJCAI | 3 |
| 2021 | Neural Symbolic Representation Learning for Image CaptioningabstractTraditional image captioning models mainly rely on one encoder-decoder architecture to generate one natural sentence for a given image. Such an architecture mostly uses deep neural networks to extract the neural representations of the image while ignoring the information of abstractive concepts as well as their intertwined relationships conveyed in the image. To this end, to comprehensively characterize the image content and bridge the gap between neural representations and high-level abstractive concepts, we make the first attempt to investigate the ability of neural symbolic representation of the image for the image captioning task. We first parse and convert a given image to neural symbolic representation in the form of an attributed relational graph, with the nodes denoting the abstractive concepts and the branches indicating the relationships between connected nodes, respectively. By performing computations over the attributed relational graph, the neural symbolic representation evolves step by step, with the node and branch representations as well as their corresponding importance weights transiting step by step. Empirically, extensive experiments validate the effectiveness of the proposed method. It enables a more comprehensive understanding of the given image by integrating the neural representation and neural symbolic representation, with the state-of-the-art results being achieved on both the MSCOCO and Flickr30k datasets. Besides, the proposed neural symbolic representation is demonstrated to better generalize to other domains with significant performance improvements compared with existing methods on the cross domain image captioning task. Lin Ma 0002, Yanwei Fu 0001, Xiangyang Xue 0001 |
ICMR | 4 |
| 2021 | The Image Local Autoregressive TransformerabstractRecently, AutoRegressive (AR) models for the whole image generation empowered by transformers have achieved comparable or even better performance compared to Generative Adversarial Networks (GANs). Unfortunately, directly applying such AR models to edit/change local image regions, may suffer from the problems of missing global information, slow inference speed, and information leakage of local guidance. To address these limitations, we propose a novel model -- image Local Autoregressive Transformer (iLAT), to better facilitate the locally guided image synthesis. Our iLAT learns the novel local discrete representations, by the newly proposed local autoregressive (LA) transformer of the attention mask and convolution mechanism. Thus iLAT can efficiently synthesize the local image regions by key guidance information. Our iLAT is evaluated on various locally guided image syntheses, such as pose-guided person image synthesis and face editing. Both quantitative and qualitative results show the efficacy of our model. Chenjie Cao, Yuxin Hong, Chengrong Wang, Chengming Xu 0001, Yanwei Fu 0001, Xiangyang Xue 0001 |
NeurIPS | 7 |
| 2021 | Progressive Coordinate Transforms for Monocular 3D Object DetectionabstractRecognizing and localizing objects in the 3D space is a crucial ability for an AI agent to perceive its surrounding environment. While significant progress has been achieved with expensive LiDAR point clouds, it poses a great challenge for 3D object detection given only a monocular image. While there exist different alternatives for tackling this problem, it is found that they are either equipped with heavy networks to fuse RGB and depth information or empirically ineffective to process millions of pseudo-LiDAR points. With in-depth examination, we realize that these limitations are rooted in inaccurate object localization. In this paper, we propose a novel and lightweight approach, dubbed {\em Progressive Coordinate Transforms} (PCT) to facilitate learning coordinate representations. Specifically, a localization boosting mechanism with confidence-aware loss is introduced to progressively refine the localization prediction. In addition, semantic image representation is also exploited to compensate for the usage of patch proposals. Despite being lightweight and simple, our strategy allows us to establish a new state-of-the-art among the monocular 3D detectors on the competitive KITTI benchmark. At the same time, our proposed PCT shows great generalization to most coordinate-based 3D detection frameworks. Li Wang 0033, Li Zhang 0040, Yi Zhu 0001, Zhi Zhang 0005, Tong He 0002, Mu Li 0003, Xiangyang Xue 0001 |
NeurIPS | 7 |
| 2021 | Temporal Context Aggregation for Video Retrieval with Contrastive LearningabstractThe current research focus on Content-Based Video Retrieval requires higher-level video representation describing the long-range semantic dependencies of relevant incidents, events, etc. However, existing methods commonly process the frames of a video as individual images or short clips, making the modeling of long-range semantic dependencies difficult. In this paper, we propose TCA (Temporal Context Aggregation for Video Retrieval), a video representation learning framework that incorporates longrange temporal information between frame-level features using the self-attention mechanism. To train it on video retrieval datasets, we propose a supervised contrastive learning method that performs automatic hard negative mining and utilizes the memory bank mechanism to increase the capacity of negative samples. Extensive experiments are conducted on multiple video retrieval tasks, such as CC WEB VIDEO, FIVR-200K, and EVVE. The proposed method shows a significant performance advantage (~ 17% mAP on FIVR-200K) over state-of-the-art methods with video-level features, and deliver competitive results with 22x faster inference time comparing with frame-level features. Jie Shao 0006, Xin Wen 0004, Bingchen Zhao, Xiangyang Xue 0001 |
WACV | 4 |
| 2021 | Syntax-guided text generation via graph neural network
Qipeng Guo, Xipeng Qiu, Xiangyang Xue 0001, Zheng Zhang 0001 |
Sci. China Inf. Sci. | 3 |
| 2021 | CDTD: A Large-Scale Cross-Domain Benchmark for Instance-Level Image-to-Image Translation and Domain Adaptive Object Detection
Mingyang Huang, Jianping Shi, Zechun Liu, Harsh Maheshwari, Yutong Zheng, Xiangyang Xue 0001, Marios Savvides, Thomas S. Huang |
Int. J. Comput. Vis. | 7 |
| 2021 | Pixel2Mesh: 3D Mesh Model Generation via Image Guided DeformationabstractIn this paper, we propose an end-to-end deep learning architecture that generates 3D triangular meshes from single color images. Restricted by the nature of prevalent deep learning techniques, the majority of previous works represent 3D shapes in volumes or point clouds. However, it is non-trivial to convert these representations to compact and ready-to-use mesh models. Unlike the existing methods, our network represents 3D shapes in meshes, which are essentially graphs and well suited for graph-based convolutional neural networks. Leveraging perceptual features extracted from an input image, our network produces the correct geometry by progressively deforming an ellipsoid. To make the whole deformation procedure stable, we adopt a coarse-to-fine strategy, and define various mesh/surface related losses to capture properties of various aspects, which benefits producing the visually appealing and physically accurate 3D geometry. In addition, our model by nature can be adapted to objects in specific domains, e.g., human faces, and be easily extended to learn per-vertex properties, e.g., color. Extensive experiments show that our method not only qualitatively produces the mesh model with better details, but also achieves the higher 3D shape estimation accuracy compared against the state-of-the-arts. Nanyang Wang, Yinda Zhang 0001, Zhuwen Li, Yanwei Fu 0001, Wei Liu 0005, Xiangyang Xue 0001, Yu-Gang Jiang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2021 | Periphery-aware COVID-19 diagnosis with contrastive representation enhancement
Junlin Hou, Jilan Xu, Longquan Jiang 0003, Shanshan Du, Rui Feng 0001, Yuejie Zhang, Xiangyang Xue 0001 |
Pattern Recognit. | 8 |
| 2020 | Feature Deformation Meta-Networks in Image Captioning of Novel ObjectsabstractThis paper studies the task of image captioning with novel objects, which only exist in testing images. Intrinsically, this task can reflect the generalization ability of models in understanding and captioning the semantic meanings of visual concepts and objects unseen in training set, sharing the similarity to one/zero-shot learning. The critical difficulty thus comes from that no paired images and sentences of the novel objects can be used to help train the captioning model. Inspired by recent work (Chen et al. 2019b) that boosts one-shot learning by learning to generate various image deformations, we propose learning meta-networks for deforming features for novel object captioning. To this end, we introduce the feature deformation meta-networks (FDM-net), which is trained on source data, and learn to adapt to the novel object features detected by the auxiliary detection model. FDM-net includes two sub-nets: feature deformation, and scene graph sentence reconstruction, which produce the augmented image features and corresponding sentences, respectively. Thus, rather than directly deforming images, FDM-net can efficiently and dynamically enlarge the paired images and texts by learning to deform image features. Extensive experiments are conducted on the widely used novel object captioning dataset, and the results show the effectiveness of our FDM-net. Ablation study and qualitative visualization further give insights of our model. Tingjia Cao, Lin Ma 0002, Yanwei Fu 0001, Yu-Gang Jiang 0001, Xiangyang Xue 0001 |
AAAI | 7 |
| 2020 | Multi-Scale Self-Attention for Text ClassificationabstractIn this paper, we introduce the prior knowledge, multi-scale structure, into self-attention modules. We propose a Multi-Scale Transformer which uses multi-scale multi-head self-attention to capture features from different scales. Based on the linguistic perspective and the analysis of pre-trained Transformer (BERT) on a huge corpus, we further design a strategy to control the scale distribution for each layer. Results of three different kinds of tasks (21 datasets) show our Multi-Scale Transformer outperforms the standard Transformer consistently and significantly on small and moderate size datasets. Qipeng Guo, Xipeng Qiu, Pengfei Liu 0003, Xiangyang Xue 0001, Zheng Zhang 0001 |
AAAI | 4 |
| 2020 | Joint Parsing and Generation for Abstractive SummarizationabstractSentences produced by abstractive summarization systems can be ungrammatical and fail to preserve the original meanings, despite being locally fluent. In this paper we propose to remedy this problem by jointly generating a sentence and its syntactic dependency parse while performing abstraction. If generating a word can introduce an erroneous relation to the summary, the behavior must be discouraged. The proposed method thus holds promise for producing grammatical sentences and encouraging the summary to stay true-to-original. Our contributions of this work are twofold. First, we present a novel neural architecture for abstractive summarization that combines a sequential decoder with a tree-based decoder in a synchronized manner to generate a summary sentence and its syntactic parse. Secondly, we describe a novel human evaluation protocol to assess if, and to what extent, a summary remains true to its original meanings. We evaluate our method on a number of summarization datasets and demonstrate competitive results against strong baselines. Kaiqiang Song, Logan Lebanoff, Qipeng Guo, Xipeng Qiu, Xiangyang Xue 0001, Chen Li 0003, Dong Yu 0001, Fei Liu 0004 |
AAAI | 5 |
| 2020 | Self-supervised Learning of Orc-Bert Augmentor for Recognizing Few-Shot Oracle Characters
Wenhui Han, Xinlin Ren, Yanwei Fu 0001, Xiangyang Xue 0001 |
ACCV (6) | 5 |
| 2020 | Long-Term Cloth-Changing Person Re-identification
Xuelin Qian, Wenxuan Wang 0003, Li Zhang 0040, Fangrui Zhu, Yanwei Fu 0001, Tao Xiang 0002, Yu-Gang Jiang 0001, Xiangyang Xue 0001 |
ACCV (3) | 8 |
| 2020 | Sketch-BERT: Learning Sketch Bidirectional Encoder Representation From Transformers by Self-Supervised Learning of Sketch GestaltabstractPrevious researches of sketches often considered sketches in pixel format and leveraged CNN based models in the sketch understanding. Fundamentally, a sketch is stored as a sequence of data points, a vector format representation, rather than the photo-realistic image of pixels. SketchRNN studied a generative neural representation for sketches of vector format by Long Short Term Memory networks (LSTM). Unfortunately, the representation learned by SketchRNN is primarily for the generation tasks, rather than the other tasks of recognition and retrieval of sketches. To this end and inspired by the recent BERT model, we present a model of learning Sketch Bidirectional Encoder Representation from Transformer (Sketch-BERT). We generalize BERT to sketch domain, with the novel proposed components and pre-training algorithms, including the newly designed sketch embedding networks, and the self-supervised learning of sketch gestalt. Particularly, towards the pre-training task, we present a novel Sketch Gestalt Model (SGM) to help train the Sketch-BERT. Experimentally, we show that the learned representation of Sketch-BERT can help and improve the performance of the downstream tasks of sketch recognition, sketch retrieval, and sketch gestalt. Yanwei Fu 0001, Xiangyang Xue 0001, Yu-Gang Jiang 0001 |
CVPR | 3 |
| 2020 | FM2u-Net: Face Morphological Multi-Branch Network for Makeup-Invariant Face VerificationabstractIt is challenging in learning a makeup-invariant face verification model, due to (1) insufficient makeup/non-makeup face training pairs, (2) the lack of diverse makeup faces, and (3) the significant appearance changes caused by cosmetics. To address these challenges, we propose a unified Face Morphological Multi-branch Network (FMMu-Net) for makeup-invariant face verification, which can simultaneously synthesize many diverse makeup faces through face morphology network (FM-Net) and effectively learn cosmetics-robust face representations using attention-based multi-branch learning network (AttM-Net). For challenges (1) and (2), FM-Net (two stacked auto-encoders) can synthesize realistic makeup face images by transferring specific regions of cosmetics via cycle consistent loss. For challenge (3), AttM-Net, consisting of one global and three local (task-driven on two eyes and mouth) branches, can effectively capture the complementary holistic and detailed information. Unlike DeepID2 which uses simple concatenation fusion, we introduce a heuristic method AttM-FM, attached to AttM-Net, to adaptively weight the features of different branches guided by the holistic information. We conduct extensive experiments on makeup face verification benchmarks (M-501, M-203, and FAM) and general face recognition datasets (LFW and IJB-A). Our framework FMMu-Net achieves state-of-the-art performances. Wenxuan Wang 0003, Yanwei Fu 0001, Xuelin Qian, Yu-Gang Jiang 0001, Qi Tian 0001, Xiangyang Xue 0001 |
CVPR | 6 |
| 2020 | Neural Pose Transfer by Spatially Adaptive Instance NormalizationabstractPose transfer has been studied for decades, in which the pose of a source mesh is applied to a target mesh. Particularly in this paper, we are interested in transferring the pose of source human mesh to deform the target human mesh, while the source and target meshes may have different identity information. Traditional studies assume that the paired source and target meshes are existed with the point-wise correspondences of user annotated landmarks/mesh points, which requires heavy labelling efforts. On the other hand, the generalization ability of deep models is limited, when the source and target meshes have different identities. To break this limitation, we proposes the first neural pose transfer model that solves the pose transfer via the latest technique for image style transfer, leveraging the newly proposed component -- spatially adaptive instance normalization. Our model does not require any correspondences between the source and target meshes. Extensive experiments show that the proposed model can effectively transfer deformation from source to target meshes, and has good generalization ability to deal with unseen identities or poses of meshes. Code is available at https://github.com/jiashunwang/Neural-Pose-Transfer. Jiashun Wang, Chao Wen 0001, Yanwei Fu 0001, Tianyun Zou, Xiangyang Xue 0001, Yinda Zhang 0001 |
CVPR | 6 |
| 2020 | DeepSFM: Structure from Motion via Deep Bundle Adjustment
Xingkui Wei, Yinda Zhang 0001, Zhuwen Li, Yanwei Fu 0001, Xiangyang Xue 0001 |
ECCV (1) | 5 |
| 2020 | BERT-ATTACK: Adversarial Attack Against BERT Using BERTabstractAdversarial attacks for discrete data (such as texts) have been proved significantly more challenging than continuous data (such as images) since it is difficult to generate adversarial samples with gradient-based methods.Current successful attack methods for texts usually adopt heuristic replacement strategies on the character or word level, which remains challenging to find the optimal solution in the massive space of possible combinations of replacements while preserving semantic consistency and language fluency.In this paper, we propose BERT-Attack, a high-quality and effective method to generate adversarial samples using pre-trained masked language models exemplified by BERT.We turn BERT against its fine-tuned models and other deep neural models in downstream tasks so that we can successfully mislead the target models to predict incorrectly.Our method outperforms state-of-theart attack strategies in both success rate and perturb percentage, while the generated adversarial samples are fluent and semantically preserved.Also, the cost of calculation is low, thus possible for large-scale generations.The code is available at https://github.com/ LinyangLee/BERT-Attack. Linyang Li, Ruotian Ma, Qipeng Guo, Xiangyang Xue 0001, Xipeng Qiu |
EMNLP (1) | 4 |
| 2020 | Towards Hierarchical Importance Attribution: Explaining Compositional Semantics for Neural Sequence Models
Xisen Jin, Zhongyu Wei, Junyi Du, Xiangyang Xue 0001, Xiang Ren 0001 |
ICLR | 4 |
| 2020 | Learnable Higher-Order Representation for Action Recognition
Jie Shao 0006, Xiangyang Xue 0001 |
ICPR | 2 |
| 2020 | 3DCFS: Fast and Robust Joint 3D Semantic-Instance Segmentation via Coupled Feature SelectionabstractWe propose a novel fast and robust 3D point clouds segmentation framework via coupled feature selection, named 3DCFS, that jointly performs semantic and instance segmentation. Inspired by the human scene perception process, we design a novel coupled feature selection module, named CFSM, that adaptively selects and fuses the reciprocal semantic and instance features from two tasks in a coupled manner. To further boost the performance of the instance segmentation task in our 3DCFS, we investigate a loss function that helps the model learn to balance the magnitudes of the output embedding dimensions during training, which makes calculating the Euclidean distance more reliable and enhances the generalizability of the model. Extensive experiments demonstrate that our 3DCFS outperforms state-of-the-art methods on benchmark datasets in terms of accuracy, speed and computational cost. Codes are available at: https://github.com/Biotan/3DCFS. Liang Du 0004, Jingang Tan, Xiangyang Xue 0001, Hongkai Wen 0001, Jianfeng Feng, Jiamao Li |
ICRA | 3 |
| 2020 | Is normalization indispensable for training deep neural network?abstractNormalization operations are widely used to train deep neural networks, and they can improve both convergence and generalization in most tasks. The theories for normalization's effectiveness and new forms of normalization have always been hot topics in research. To better understand normalization, one question can be whether normalization is indispensable for training deep neural network? In this paper, we study what would happen when normalization layers are removed from the network, and show how to train deep neural networks without normalization layers and without performance degradation. Our proposed method can achieve the same or even slightly better performance in a variety of tasks: image classification in ImageNet, object detection and segmentation in MS-COCO, video classification in Kinetics, and machine translation in WMT English-German, etc. Our study may help better understand the role of normalization layers and can be a competitive alternative to normalization layers. Codes are available. Jie Shao 0006, Kai Hu 0010, Changhu Wang, Xiangyang Xue 0001, Bhiksha Raj |
NeurIPS | 4 |
| 2020 | Vocabulary-Informed Zero-Shot and Open-Set LearningabstractDespite significant progress in object categorization, in recent years, a number of important challenges remain; mainly, the ability to learn from limited labeled data and to recognize object classes within large, potentially open, set of labels. Zero-shot learning is one way of addressing these challenges, but it has only been shown to work with limited sized class vocabularies and typically requires separation between supervised and unsupervised classes, allowing former to inform the latter but not vice versa. We propose the notion of vocabulary-informed learning to alleviate the above mentioned challenges and address problems of supervised, zero-shot, generalized zero-shot and open set recognition using a unified framework. Specifically, we propose a weighted maximum margin framework for semantic manifold-based recognition that incorporates distance constraints from (both supervised and unsupervised) vocabulary atoms. Distance constraints ensure that labeled samples are projected closer to their correct prototypes, in the embedding space, than to others. We illustrate that resulting model shows improvements in supervised, zero-shot, generalized zero-shot, and large open set recognition, with up to 310K class vocabulary on Animal with Attributes and ImageNet datasets. Yanwei Fu 0001, Hanze Dong, Yu-Gang Jiang 0001, Meng Wang 0001, Xiangyang Xue 0001, Leonid Sigal |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2020 | Leader-Based Multi-Scale Attention Deep Architecture for Person Re-IdentificationabstractPerson re-identification (re-id) aims to match people across non-overlapping camera views in a public space. This is a challenging problem because the people captured in surveillance videos often wear similar clothing. Consequently, the differences in their appearance are typically subtle and only detectable at particular locations and scales. In this paper, we propose a deep re-id network (MuDeep) that is composed of two novel types of layers - a multi-scale deep learning layer, and a leader-based attention learning layer. Specifically, the former learns deep discriminative feature representations at different scales, while the latter utilizes the information from multiple scales to lead and determine the optimal weightings for each scale. The importance of different spatial locations for extracting discriminative features is learned explicitly via our leader-based attention learning layer. Extensive experiments are carried out to demonstrate that the proposed MuDeep outperforms the state-of-the-art on a number of benchmarks and has a better generalization ability under a domain generalization setting. Xuelin Qian, Yanwei Fu 0001, Tao Xiang 0002, Yu-Gang Jiang 0001, Xiangyang Xue 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2020 | Object Detection from Scratch with Deep SupervisionabstractIn this paper, we propose Deeply Supervised Object Detectors (DSOD), an object detection framework that can be trained from scratch. Recent advances in object detection heavily depend on the off-the-shelf models pre-trained on large-scale classification datasets like ImageNet and OpenImage. However, one problem is that adopting pre-trained models from classification to detection task may incur learning bias due to the different objective function and diverse distributions of object categories. Techniques like fine-tuning on detection task could alleviate this issue to some extent but are still not fundamental. Furthermore, transferring these pre-trained models across discrepant domains will be more difficult (e.g., from RGB to depth images). Thus, a better solution to handle these critical problems is to train object detectors from scratch, which motivates our proposed method. Previous efforts on this direction mainly failed by reasons of the limited training data and naive backbone network structures for object detection. In DSOD, we contribute a set of design principles for learning object detectors from scratch. One of the key principles is the deep supervision, enabled by layer-wise dense connections in both backbone networks and prediction layers, plays a critical role in learning good detectors from scratch. After involving several other principles, we build our DSOD based on the single-shot detection framework (SSD). We evaluate our method on PASCAL VOC 2007, 2012 and COCO datasets. DSOD achieves consistently better results than the state-of-the-art methods with much more compact models. Specifically, DSOD outperforms baseline method SSD on all three benchmarks, while requiring only 1/2 parameters. We also observe that DSOD can achieve comparable/slightly better results than Mask RCNN [1] + FPN [2] (under similar input size) with only 1/3 parameters, using no extra data or pre-trained models. Zhuang Liu 0003, Yu-Gang Jiang 0001, Yurong Chen 0001, Xiangyang Xue 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2020 | Learning to Score Figure Skating Sport VideosabstractThis paper aims at learning to score the figure skating sports videos. To address this task, we propose a deep architecture that includes two complementary components, i.e., Self-Attentive LSTM and Multi-scale Convolutional Skip LSTM. These two components can efficiently learn the local and global sequential information in each video. Furthermore, we present a large-scale figure skating sports video dataset - FisV dataset. This dataset includes 500 figure skating videos with the average length of 2 minutes and 50 seconds. Each video is annotated by two scores of nine different referees, i.e., Total Element Score(TES) and Total Program Component Score (PCS). Our proposed model is validated on FisV and MIT-skate datasets. The experimental results show the effectiveness of our models in learning to score the figure skating videos. The codes and datasets would be downloaded from https://github.com/loadder/MS_LSTM.git. Chengming Xu 0001, Yanwei Fu 0001, Zitian Chen, Yu-Gang Jiang 0001, Xiangyang Xue 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2020 | Pose-Guided Person Image Synthesis in the Non-Iconic ViewsabstractGenerating realistic images with the guidance of reference images and human poses is challenging. Despite the success of previous works on synthesizing person images in the iconic views, no efforts are made towards the task of poseguided image synthesis in the non-iconic views. Particularly, we find that previous models cannot handle such a complex task, where the person images are captured in the non-iconic views by commercially-available digital cameras. To this end, we propose a new framework - Multi-branch Refinement Network (MR-Net), which utilizes several visual cues, including target person poses, foreground person body and scene images parsed. Furthermore, a novel Region of Interest (RoI) perceptual loss is proposed to optimize the MR-Net. Extensive experiments on two non-iconic datasets, Penn Action and BBC-Pose, as well as an iconic dataset - Market-1501, show the efficacy of the proposed model that can tackle the problem of pose-guided person image generation from the non-iconic views. The data, models, and codes are downloadable from https://github.com/loadder/MR-Net. Chengming Xu 0001, Yanwei Fu 0001, Chao Wen 0001, Yu-Gang Jiang 0001, Xiangyang Xue 0001 |
IEEE Trans. Image Process. | 6 |
| 2020 | M$^3$Lung-Sys: A Deep Learning System for Multi-Class Lung Pneumonia Screening From CT ImagingabstractTo counter the outbreak of COVID-19, the accurate diagnosis of suspected cases plays a crucial role in timely quarantine, medical treatment, and preventing the spread of the pandemic. Considering the limited training cases and resources (e.g, time and budget), we propose a Multi-task Multi-slice Deep Learning System (M3Lung-Sys) for multi-class lung pneumonia screening from CT imaging, which only consists of two 2D CNN networks, i.e., slice- and patient-level classification networks. The former aims to seek the feature representations from abundant CT slices instead of limited CT volumes, and for the overall pneumonia screening, the latter one could recover the temporal information by feature refinement and aggregation between different slices. In addition to distinguish COVID-19 from Healthy, H1N1, and CAP cases, our M3Lung-Sys also be able to locate the areas of relevant lesions, without any pixel-level annotation. To further demonstrate the effectiveness of our model, we conduct extensive experiments on a chest CT imaging dataset with a total of 734 patients (251 healthy people, 245 COVID-19 patients, 105 H1N1 patients, and 133 CAP patients). The quantitative results with plenty of metrics indicate the superiority of our proposed model on both slice- and patient-level classification tasks. More importantly, the generated lesion location maps make our system interpretable and more valuable to clinicians. Xuelin Qian, Huazhu Fu, Weiya Shi, Tao Chen 0003, Yanwei Fu 0001, Xiangyang Xue 0001 |
IEEE J. Biomed. Health Informatics | 7 |
| 2020 | A Multi-Task Neural Approach for Emotion Attribution, Classification, and SummarizationabstractEmotional content is a crucial ingredient in user-generated videos. However, the sparsity of emotional expressions in the videos poses an obstacle to visual emotion analysis. In this paper, we propose a new neural approach, Bi-stream Emotion Attribution-Classification Network (BEAC-Net), to solve three related emotion analysis tasks: emotion recognition, emotion attribution, and emotion-oriented summarization, in a single integrated framework. BEAC-Net has two major constituents, an attribution network and a classification network. The attribution network extracts the main emotional segment that classification should focus on in order to mitigate the sparsity issue. The classification network utilizes both the extracted segment and the original video in a bi-stream architecture. We contribute a new dataset for the emotion attribution task with human-annotated ground-truth labels for emotion segments. Experiments on two video datasets demonstrate superior performance of the proposed framework and the complementary nature of the dual classification streams. Guoyun Tu, Yanwei Fu 0001, Boyang Li 0001, Jiarui Gao, Yu-Gang Jiang 0001, Xiangyang Xue 0001 |
IEEE Trans. Multim. | 6 |
| 2019 | MEAL: Multi-Model Ensemble via Adversarial LearningabstractOften the best performing deep neural models are ensembles of multiple base-level networks. Unfortunately, the space required to store these many networks, and the time required to execute them at test-time, prohibits their use in applications where test sets are large (e.g., ImageNet). In this paper, we present a method for compressing large, complex trained ensembles into a single network, where knowledge from a variety of trained deep neural networks (DNNs) is distilled and transferred to a single DNN. In order to distill diverse knowledge from different trained (teacher) models, we propose to use adversarial-based learning strategy where we define a block-wise training loss to guide and optimize the predefined student network to recover the knowledge in teacher models, and to promote the discriminator network to distinguish teacher vs. student features simultaneously. The proposed ensemble method (MEAL) of transferring distilled knowledge with adversarial learning exhibits three important advantages: (1) the student network that learns the distilled knowledge with discriminators is optimized better than the original model; (2) fast inference is realized by a single forward pass, while the performance is even better than traditional ensembles from multi-original models; (3) the student network can learn the distilled knowledge from a teacher model that has arbitrary structures. Extensive experiments on CIFAR-10/100, SVHN and ImageNet datasets demonstrate the effectiveness of our MEAL method. On ImageNet, our ResNet-50 based MEAL achieves top-1/5 21.79%/5.99% val error, which outperforms the original model by 2.06%/1.14%. Zhankui He, Xiangyang Xue 0001 |
AAAI | 3 |
| 2019 | Spatial Mixture Models with Learnable Deep Priors for Perceptual GroupingabstractHumans perceive the seemingly chaotic world in a structured and compositional way with the prerequisite of being able to segregate conceptual entities from the complex visual scenes. The mechanism of grouping basic visual elements of scenes into conceptual entities is termed as perceptual grouping. In this work, we propose a new type of spatial mixture models with learnable priors for perceptual grouping. Different from existing methods, the proposed method disentangles the representation of an object into “shape” and “appearance” which are modeled separately by the mixture weights and the conditional probability distributions. More specifically, each object in the visual scene is modeled by one mixture component, whose mixture weights and the parameter of the conditional probability distribution are generated by two neural networks, respectively. The mixture weights focus on modeling spatial dependencies (i.e., shape) and the conditional probability distributions deal with intra-object variations (i.e., appearance). In addition, the background is separately modeled as a special component complementary to the foreground objects. Our extensive empirical tests on two perceptual grouping datasets demonstrate that the proposed method outperforms the stateof-the-art methods under most experimental configurations. The learned conceptual entities are generalizable to novel visual scenes and insensitive to the diversity of objects. Jinyang Yuan, Bin Li 0015, Xiangyang Xue 0001 |
AAAI | 3 |
| 2019 | Towards Instance-Level Image-To-Image TranslationabstractUnpaired Image-to-image Translation is a new rising and challenging vision problem that aims to learn a mapping between unaligned image pairs in diverse domains. Recent advances in this field like MUNIT and DRIT mainly focus on disentangling content and style/attribute from a given image first, then directly adopting the global style to guide the model to synthesize new domain images. However, this kind of approaches severely incurs contradiction if the target domain images are content-rich with multiple discrepant objects. In this paper, we present a simple yet effective instance-aware image-to-image translation approach (INIT), which employs the fine-grained local (instance) and global styles to the target image spatially. The proposed INIT exhibits three import advantages: (1) the instance-level objective loss can help learn a more accurate reconstruction and incorporate diverse attributes of objects; (2) the styles used for target domain of local/global areas are from corresponding spatial regions in source domain, which intuitively is a more reasonable mapping; (3) the joint training process can benefit both fine and coarse granularity and incorporates instance information to improve the quality of global translation. We also collect a large-scale benchmark for the new instance-level translation task. We observe that our synthetic images can even benefit real-world vision tasks like generic object detection. Mingyang Huang, Jianping Shi, Xiangyang Xue 0001, Thomas S. Huang |
CVPR | 4 |
| 2019 | SSF-DAN: Separated Semantic Feature Based Domain Adaptation Network for Semantic SegmentationabstractDespite the great success achieved by supervised fully convolutional models in semantic segmentation, training the models requires a large amount of labor-intensive work to generate pixel-level annotations. Recent works exploit synthetic data to train the model for semantic segmentation, but the domain adaptation between real and synthetic images remains a challenging problem. In this work, we propose a Separated Semantic Feature based domain adaptation network, named SSF-DAN, for semantic segmentation. First, a Semantic-wise Separable Discriminator (SS-D) is designed to independently adapt semantic features across the target and source domains, which addresses the inconsistent adaptation issue in the class-wise adversarial learning. In SS-D, a progressive confidence strategy is included to achieve a more reliable separation. Then, an efficient Class-wise Adversarial loss Reweighting module (CA-R) is introduced to balance the class-wise adversarial learning process, which leads the generator to focus more on poorly adapted classes. The presented framework demonstrates robust performance, superior to state-of-the-art methods on benchmark datasets. Liang Du 0004, Jingang Tan, Hongye Yang, Jianfeng Feng, Xiangyang Xue 0001, Qibao Zheng, Xiaoqing Ye |
ICCV | 5 |
| 2019 | Parasitic GAN for Semi-Supervised Brain Tumor SegmentationabstractIn semantic segmentation, researchers face the shortage of pixel-level annotated data. And it is particularly severe in the medical images. On the other hand, the unlabeled data are abundantly produced in the diagnosis routine. In the paper, we introduced the Parasitic GAN for the brain tumor segmentation to exploit the unlabeled data more efficiently. Parasitic GAN is composed of three parts: the segmentor S, the generator G, and the discriminator V. With the label maps produced by the segmentor and the supplementary label maps synthesized by the generator, the discriminator could learn a more precise boundary of ground truth. Thus, the segmentor benefits from the adversarial learning mechanism and the extra supervision provided by the discriminator. This parasitic relationship between the segmentor and the generative adversarial network (G and V) restricts the fitness ability of the segmentor and improves its generalization capacity. In practice, it definitely improved the performance of segmentor in brain tumor segmentation tasks, increasing the dice score 0.010-0.035. Chengfeng Zhou, Yanwei Fu 0001, Xiangyang Xue 0001 |
ICIP | 4 |
| 2019 | CODA: Counting Objects via Scale-Aware Adversarial Density AdaptionabstractRecent advances in crowd counting have achieved promising results with increasingly complex convolutional neural network designs. However, due to the unpredictable domain shift, generalizing trained model to unseen scenarios is often suboptimal. Inspired by the observation that density maps of different scenarios share similar local structures, we propose a novel adversarial learning approach in this paper, i.e., CODA (Counting Objects via scale-aware adversarial Density Adaption). To deal with different object scales and density distributions, we perform adversarial training with pyramid patches of multi-scales from both source-and target-domain. Along with a ranking constraint across levels of the pyramid input, consistent object counts can be produced for different scales. Extensive experiments demonstrate that our network produces much better results on unseen datasets compared with existing counting adaption models. Notably, the performance of our CODA is comparable with the state-of-the-art fully-supervised models that are trained on the target dataset. Further analysis indicates that our density adaption framework can effortlessly extend to scenarios with different objects. The code is available at https://github.com/Willy0919/CODA. Li Wang 0033, Xiangyang Xue 0001 |
ICME | 3 |
| 2019 | Generative Modeling of Infinite Occluded Objects for Compositional Scene RepresentationabstractWe present a deep generative model which explicitly models object occlusions for compositional scene representation. Latent representations of objects are disentangled into location, size, shape, and appearance, and the visual scene can be generated compositionally by integrating these representations and an infinite-dimensional binary vector indicating presences of objects in the scene. By training the model to learn spatial dependences of pixels in the unsupervised setting, the number of objects, pixel-level segregation of objects, and presences of objects in overlapping regions can be estimated through inference of latent variables. Extensive experiments conducted on a series of specially designed datasets demonstrate that the proposed method outperforms two state-of-the-art methods when object occlusions exist. Jinyang Yuan, Bin Li 0015, Xiangyang Xue 0001 |
ICML | 3 |
| 2019 | Embodied One-Shot Video Recognition: Learning from Actions of a Virtual Embodied AgentabstractOne-shot learning aims to recognize novel target classes from few examples by transferring knowledge from source classes, under a general assumption that the source and target classes are semantically related but not exactly the same. Based on this assumption, recent work has focused on image-based one-shot learning, while little work has addressed video-based one shot learning. One of the challenges lies in that it is difficult to maintain the disjoint-class assumption for videos, since video clips of target classes may potentially appear in the videos of source classes. To address this issue, we introduce a novel setting, termed as embodied agents based one-shot learning, which leverages synthetic videos produced in a virtual environment to understand realistic videos of target classes. In this setting, we further propose two types of learning tasks: embodied one-shot video domain adaptation and embodied one-shot video transfer recognition. These tasks serve as a testbed for evaluating video related one-shot learning tasks. In addition, we propose a general video segment augmentation method, which significantly facilitates a variety of one-shot learning tasks. Experimental results validate the soundness of our setting and learning tasks, and also show the effectiveness of our augmentation approach to video recognition in the small-sample size regime. Yuqian Fu, Chengrong Wang, Yanwei Fu 0001, Yu-Xiong Wang, Cong Bai, Xiangyang Xue 0001, Yu-Gang Jiang 0001 |
ACM Multimedia | 6 |
| 2019 | TC-Net for iSBIR: Triplet Classification Network for Instance-level Sketch Based Image RetrievalabstractSketch has been employed as an effective communication tool to express the abstract and intuitive meaning of object. While content-based sketch recognition has been studied for several decades, the instance-level Sketch Based Image Retrieval (iSBIR) task has attracted significant research attention recently. In many previous iSBIR works -- TripletSN, and DSSA, edge maps were employed as intermediate representations in bridging the cross-domain discrepancy between photos and sketches. However, it is nontrivial to efficiently train and effectively use the edge maps in an iSBIR system. Particularly, we find that such an edge map based iSBIR system has several major limitations. First, the system has to be pre-trained on a significant amount of edge maps, either from large-scale sketch datasets, e.g., TU-Berlin~\citeeitz2012hdhso, or converted from other large-scale image datasets, e.g., ImageNet-1K\citedeng2009imagenet dataset. Second, the performance of such an iSBIR system is very sensitive to the quality of edge maps. Third and empirically, the multi-cropping strategy is essentially very important in improving the performance of previous iSBIR systems. To address these limitations, this paper advocates an end-to-end iSBIR system without using the edge maps. Specifically, we present a Triplet Classification Network (TC-Net) for iSBIR which is composed of two major components: triplet Siamese network, and auxiliary classification loss. Our TC-Net can break the limitations existed in previous works. Extensive experiments on several datasets validate the efficacy of the proposed network and system. Yanwei Fu 0001, Shaogang Gong, Xiangyang Xue 0001, Yu-Gang Jiang 0001 |
ACM Multimedia | 5 |
| 2019 | Comp-GAN: Compositional Generative Adversarial Network in Synthesizing and Recognizing Facial ExpressionabstractFacial expression is important in understanding our social interaction. Thus the ability to recognize facial expression enables the novel multimedia applications. With the advance of recent deep architectures, research on facial expression recognition has achieved great progress. However, these models are still suffering from the problems of lacking sufficient and diverse high quality training faces, vulnerability to the facial variations, and recognizing a limited number of basic types of emotions. To tackle these problems, this paper proposes a novel end-to-end Compositional Generative Adversarial Network (Comp-GAN) that is able to synthesize new face images with specified poses and desired facial expressions; and such synthesized images can be further utilized to help train a robust and generalized expression recognition model. Essentially, Comp-GAN can dynamically change the expression and pose of faces according to the input images while keeping the identity information. Specifically, the generator has two major components: one for generating images with desired expression and the other for changing the pose of faces. Furthermore, a face reconstruction learning process is applied to re-generate the input image and constrains the generator for preserving the key information such as facial identity. For the first time, various one/zero-shot facial expression recognition tasks have been created. We conduct extensive experiments to show that the images generated by Comp-GAN are helpful to improve the performance of one/zero-shot facial expression recognition. Wenxuan Wang 0003, Qiang Sun 0007, Yanwei Fu 0001, Tao Chen 0003, Chenjie Cao, Ziqi Zheng, Han Qiu 0002, Yu-Gang Jiang 0001, Xiangyang Xue 0001 |
ACM Multimedia | 10 |
| 2019 | Low-Rank and Locality Constrained Self-Attention for Sequence ModelingabstractSelf-attention mechanism becomes more and more popular in natural language processing (NLP) applications. Recent studies show the Transformer architecture which relies mainly on the attention mechanism achieves much success on large datasets. But a raised problem is its generalization ability is weaker than CNN and RNN on many moderate-sized datasets. We think the reason can be attributed to its unsuitable inductive bias of the self-attention structure. In this paper, we regard the self-attention as matrix decomposition problem and propose an improved self-attention module by introducing two linguistic constraints: low-rank and locality. We further develop the low-rank attention and band attention to parameterize the self-attention mechanism under the low-rank and locality constraints. Experiments on several real NLP tasks show our model outperforms the vanilla Transformer and other self-attention models on moderate size datasets. Additionally, evaluation on a synthetic task gives us a more detailed understanding of working mechanisms of different architectures. Qipeng Guo, Xipeng Qiu, Xiangyang Xue 0001, Zheng Zhang 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2019 | Multi-Level Semantic Feature Augmentation for One-Shot LearningabstractThe ability to quickly recognize and learn new visual concepts from limited samples enables humans to quickly adapt to new tasks and environments. This ability is enabled by semantic association of novel concepts with those that have already been learned and stored in memory. Computers can start to ascertain similar abilities by utilizing a semantic concept space. A concept space is a high-dimensional semantic space in which similar abstract concepts appear close and dissimilar ones far apart. In this paper, we propose a novel approach to one-shot learning that builds on this core idea. Our approach learns to map a novel sample instance to a concept, relates that concept to the existing ones in the concept space and, using these relationships, generates new instances, by interpolating among the concepts, to help learning. Instead of synthesizing new image instance, we propose to directly synthesize instance features by leveraging semantics using a novel auto-encoder network we call dual TriNet. The encoder part of the TriNet learns to map multi-layer visual features from CNN to a semantic vector. In semantic space, we search for related concepts, which are then projected back into the image feature spaces by the decoder portion of the TriNet. Two strategies in the semantic space are explored. Notably, this seemingly simple strategy results in complex augmented feature distributions in the image feature space, leading to substantially better performance. The codes and models are released in the github: https://github.com/tankche1/ Semantic-Feature-Augmentation-in-Few-shot-Learning. Zitian Chen, Yanwei Fu 0001, Yinda Zhang 0001, Yu-Gang Jiang 0001, Xiangyang Xue 0001, Leonid Sigal |
IEEE Trans. Image Process. | 5 |
| 2018 | Dual Skipping NetworksabstractInspired by the recent neuroscience studies on the left-right asymmetry of the human brain in processing low and high spatial frequency information, this paper introduces a dual skipping network which carries out coarse-to-fine object categorization. Such a network has two branches to simultaneously deal with both coarse and fine-grained classification tasks. Specifically, we propose a layer-skipping mechanism that learns a gating network to predict which layers to skip in the testing stage. This layer-skipping mechanism endows the network with good flexibility and capability in practice. Evaluations are conducted on several widely used coarse-to-fine object categorization benchmarks, and promising results are achieved by our proposed network model. Changmao Cheng, Yanwei Fu 0001, Yu-Gang Jiang 0001, Wei Liu 0005, Wenlian Lu, Jianfeng Feng, Xiangyang Xue 0001 |
CVPR | 7 |
| 2018 | Pose-Normalized Image Generation for Person Re-identification
Xuelin Qian, Yanwei Fu 0001, Tao Xiang 0002, Wenxuan Wang 0003, Yang Wu 0001, Yu-Gang Jiang 0001, Xiangyang Xue 0001 |
ECCV (9) | 8 |
| 2018 | ExFuse: Enhancing Feature Fusion for Semantic Segmentation
Zhenli Zhang, Xiangyu Zhang 0005, Chao Peng 0001, Xiangyang Xue 0001, Jian Sun 0001 |
ECCV (10) | 4 |
| 2018 | Harnessing Synthesized Abstraction Images to Improve Facial Attribute RecognitionabstractFacial attribute recognition is an important and yet challenging research topic. Different from most previous approaches which predict attributes only based on the whole images, this paper leverages facial parts locations for better attribute prediction. A facial abstraction image which contains both local facial parts and facial texture information is introduced. This abstraction image is generated by a Generative Adversarial Network (GAN). Then we build a dual-path facial attribute recognition network to utilize features from the original face images and facial abstraction images. Empirically, the features of facial abstraction images are complementary to features of original face images. With the facial parts localized by the abstraction images, our method improves facial attributes recognition, especially the attributes located on small face regions. Extensive evaluations conducted on CelebA and LFWA benchmark datasets show that state-of-the-art performance is achieved. Keke He, Yanwei Fu 0001, Wuhao Zhang, Chengjie Wang 0001, Yu-Gang Jiang 0001, Feiyue Huang, Xiangyang Xue 0001 |
IJCAI | 7 |
| 2018 | Stacked multichannel autoencoder - an efficient way of learning from synthetic data
Yanwei Fu 0001, Xiangyang Xue 0001, Yu-Gang Jiang 0001, Gady Agam |
Multim. Tools Appl. | 4 |
| 2018 | Exploiting Feature and Class Relationships in Video Categorization with Regularized Deep Neural NetworksabstractIn this paper, we study the challenging problem of categorizing videos according to high-level semantics such as the existence of a particular human action or a complex event. Although extensive efforts have been devoted in recent years, most existing works combined multiple video features using simple fusion strategies and neglected the utilization of inter-class semantic relationships. This paper proposes a novel unified framework that jointly exploits the feature relationships and the class relationships for improved categorization performance. Specifically, these two types of relationships are estimated and utilized by imposing regularizations in the learning process of a deep neural network (DNN). Through arming the DNN with better capability of harnessing both the feature and the class relationships, the proposed regularized DNN (rDNN) is more suitable for modeling video semantics. We show that rDNN produces better performance over several state-of-the-art approaches. Competitive results are reported on the well-known Hollywood2 and Columbia Consumer Video benchmarks. In addition, to stimulate future research on large scale video categorization, we collect and release a new benchmark dataset, called FCVID, which contains 91,223 Internet videos and 239 manually annotated categories. Yu-Gang Jiang 0001, Zuxuan Wu, Jun Wang 0006, Xiangyang Xue 0001, Shih-Fu Chang |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2018 | Modeling Multimodal Clues in a Hybrid Deep Learning Framework for Video ClassificationabstractVideos are inherently multimodal. This paper studies the problem of exploiting the abundant multimodal clues for improved video classification performance. We introduce a novel hybrid deep learning framework that integrates useful clues from multiple modalities, including static spatial appearance information, motion patterns within a short time window, audio information, as well as long-range temporal dynamics. More specifically, we utilize three Convolutional Neural Networks (CNNs) operating on appearance, motion, and audio signals to extract their corresponding features. We then employ a feature fusion network to derive a unified representation with an aim to capture the relationships among features. Furthermore, to exploit the long-range temporal dynamics in videos, we apply two long short-term memory (LSTM) networks with extracted appearance and motion features as inputs. Finally, we also propose refining the prediction scores by leveraging contextual relationships among video semantics. The hybrid deep learning framework is able to exploit a comprehensive set of multimodal features for video classification. Through an extensive set of experiments, we demonstrate that: 1) LSTM networks that model sequences in an explicitly recurrent manner are highly complementary to the CNN models; 2) the feature fusion network that produces a fused representation through modeling feature relationships outperforms a large set of alternative fusion strategies; and 3) the semantic context of video classes can help further refine the predictions for improved performance. Experimental results on two challenging benchmarks-the UCF-101 and the Columbia Consumer Videos (CCV)-provide strong quantitative evidence that our framework can produce promising results: 93.1% on the UCF-101 and 84.5% on the CCV, outperforming several competing methods with clear margins. Yu-Gang Jiang 0001, Zuxuan Wu, Jinhui Tang 0001, Zechao Li, Xiangyang Xue 0001, Shih-Fu Chang |
IEEE Trans. Multim. | 5 |
| 2018 | Arbitrary-Oriented Scene Text Detection via Rotation ProposalsabstractThis paper introduces a novel rotation-based framework for arbitrary-oriented text detection in natural scene images. We present theRotation Region Proposal Networks, which are designed to generate inclined proposals with text orientation angle information. The angle information is then adapted for bounding box regression to make the proposals more accurately fit into the text region in terms of the orientation. TheRotation Region-of-Interestpooling layer is proposed to project arbitrary-oriented proposals to a feature map for a text region classifier. The whole framework is built upon a region-proposal-based architecture, which ensures the computational efficiency of the arbitrary-oriented text detection compared with previous text detection systems. We conduct experiments using the rotation-based framework on three real-world scene text detection datasets and demonstrate its superiority in terms of effectiveness and efficiency over previous approaches. Jianqi Ma, Weiyuan Shao, Hao Ye 0005, Li Wang 0033, Hong Wang 0014, Yingbin Zheng, Xiangyang Xue 0001 |
IEEE Trans. Multim. | 7 |
| 2017 | UA-DETRAC 2017: Report of AVSS2017 & IWT4S Challenge on Advanced Traffic MonitoringabstractThe rapid advances of transportation infrastructure have led to a dramatic increase in the demand for smart systems capable of monitoring traffic and street safety. Fundamental to these applications are a community-based evaluation platform and benchmark for object detection and multi-object tracking. To this end, we organize the AVSS2017 Challenge on Advanced Traffic Monitoring, in conjunction with the International Workshop on Traffic and Street Surveillance for Safety and Security (IWT4S), to evaluate the state-of-the-art object detection and multi-object tracking algorithms in the relevance of traffic surveillance. Submitted algorithms are evaluated using the large-scale UA-DETRAC benchmark and evaluation protocol. The benchmark, the evaluation toolkit and the algorithm performance are publicly available from the website http://detrac-db.rit.albany.edu. Siwei Lyu, Ming-Ching Chang, Dawei Du, Longyin Wen, Honggang Qi, Yuezun Li, Yi Wei 0006, Lipeng Ke, Tao Hu 0011, Marco Del Coco, Pierluigi Carcagnì, Dmitriy Anisimov, Erik Bochinski, Fabio Galasso, Filiz Bunyak, Hao Ye 0005, Hong Wang 0014, Kannappan Palaniappan, Koray Ozcan, Li Wang 0033, Liang Wang 0001, Martin Lauer, Nattachai Watcharapinchai, Nenghui Song, Noor Al-Shakarji, Sikandar Amin, Sitapa Watcharapinchai, Tatiana Khanova, Thomas Sikora, Tino Kutschbach, Volker Eiselein, Wei Tian 0001, Xiangyang Xue 0001, Xiaoyi Yu, Yao Lu 0028, Yingbin Zheng, Yongzhen Huang, Yuqi Zhang 0001 |
AVSS | 35 |
| 2017 | Weakly Supervised Dense Video CaptioningabstractThis paper focuses on a novel and challenging vision task, dense video captioning, which aims to automatically describe a video clip with multiple informative and diverse caption sentences. The proposed method is trained without explicit annotation of fine-grained sentence to video region-sequence correspondence, but is only based on weak video-level sentence annotations. It differs from existing video captioning systems in three technical aspects. First, we propose lexical fully convolutional neural networks (Lexical-FCN) with weakly supervised multi-instance multi-label learning to weakly link video regions with lexical labels. Second, we introduce a novel submodular maximization scheme to generate multiple informative and diverse region-sequences based on the Lexical-FCN outputs. A winner-takes-all scheme is adopted to weakly associate sentences to region-sequences in the training phase. Third, a sequence-to-sequence learning based language model is trained with the weakly supervised information obtained through the association process. We show that the proposed method can not only produce informative and diverse dense captions, but also outperform state-of-the-art single video captioning methods by a large margin. Minjun Li, Yurong Chen 0001, Yu-Gang Jiang 0001, Xiangyang Xue 0001 |
CVPR | 7 |
| 2017 | Multi-scale Deep Learning Architectures for Person Re-identificationabstractPerson Re-identification (re-id) aims to match people across non-overlapping camera views in a public space. It is a challenging problem because many people captured in surveillance videos wear similar clothes. Consequently, the differences in their appearance are often subtle and only detectable at the right location and scales. Existing re-id models, particularly the recently proposed deep learning based ones match people at a single scale. In contrast, in this paper, a novel multi-scale deep learning model is proposed. Our model is able to learn deep discriminative feature representations at different scales and automatically determine the most suitable scales for matching. The importance of different spatial locations for extracting discriminative features is also learned explicitly. Experiments are carried out to demonstrate that the proposed model outperforms the state-of-the art on a number of benchmarks. Xuelin Qian, Yanwei Fu 0001, Yu-Gang Jiang 0001, Tao Xiang 0002, Xiangyang Xue 0001 |
ICCV | 5 |
| 2017 | DSOD: Learning Deeply Supervised Object Detectors from ScratchabstractWe present Deeply Supervised Object Detector (DSOD), a framework that can learn object detectors from scratch. State-of-the-art object objectors rely heavily on the off the-shelf networks pre-trained on large-scale classification datasets like Image Net, which incurs learning bias due to the difference on both the loss functions and the category distributions between classification and detection tasks. Model fine-tuning for the detection task could alleviate this bias to some extent but not fundamentally. Besides, transferring pre-trained models from classification to detection between discrepant domains is even more difficult (e.g. RGB to depth images). A better solution to tackle these two critical problems is to train object detectors from scratch, which motivates our proposed DSOD. Previous efforts in this direction mostly failed due to much more complicated loss functions and limited training data in object detection. In DSOD, we contribute a set of design principles for training object detectors from scratch. One of the key findings is that deep supervision, enabled by dense layer-wise connections, plays a critical role in learning a good detector. Combining with several other principles, we develop DSOD following the single-shot detection (SSD) framework. Experiments on PASCAL VOC 2007, 2012 and MS COCO datasets demonstrate that DSOD can achieve better results than the state-of-the-art solutions with much more compact models. For instance, DSOD outperforms SSD on all three benchmarks with real-time detection speed, while requires only 1/2 parameters to SSD and 1/10 parameters to Faster RCNN. Zhuang Liu 0003, Yu-Gang Jiang 0001, Yurong Chen 0001, Xiangyang Xue 0001 |
ICCV | 6 |
| 2017 | Towards 3D Human Pose Estimation in the Wild: A Weakly-Supervised ApproachabstractIn this paper, we study the task of 3D human pose estimation in the wild. This task is challenging due to lack of training data, as existing datasets are either in the wild images with 2D pose or in the lab images with 3D pose.,, We propose a weakly-supervised transfer learning method that uses mixed 2D and 3D labels in a unified deep neutral network that presents two-stage cascaded structure. Our network augments a state-of-the-art 2D pose estimation sub-network with a 3D depth regression sub-network. Unlike previous two stage approaches that train the two sub-networks sequentially and separately, our training is end-to-end and fully exploits the correlation between the 2D pose and depth estimation sub-tasks. The deep features are better learnt through shared representations. In doing so, the 3D pose labels in controlled lab environments are transferred to in the wild images. In addition, we introduce a 3D geometric constraint to regularize the 3D pose prediction, which is effective in the absence of ground truth depth labels. Our method achieves competitive results on both 2D and 3D benchmarks. Xingyi Zhou, Qixing Huang, Xiao Sun 0001, Xiangyang Xue 0001 |
ICCV | 4 |
| 2017 | Iterative object and part transfer for fine-grained recognitionabstractThe aim of fine-grained recognition is to identify sub-ordinate categories in images like different species of birds. Existing works have confirmed that, in order to capture the subtle differences across the categories, automatic localization of objects and parts is critical. Most approaches for object and part localization rehed on the bottom-up pipeline, where thousands of region proposals are generated and then filtered by pre-trained object/part models. This is computationally expensive and not scalable once the number of objects/parts becomes large. In this paper, we propose a nonparametric data-driven method for object and part localization. Given an unlabeled test image, our approach transfers annotations from a few similar images retrieved in the training set. In particular, we propose an iterative transfer strategy that gradually refine the predicted bounding boxes. Based on the located objects and parts, deep convolutional features are extracted for recognition. We evaluate our approach on the widely-used CUB200-2011 dataset and a new and large dataset called Birdsnap. On both datasets, we achieve better results than many state-of-the-art approaches, including a few using oracle (manually annotated) bounding boxes in the test images. Yu-Gang Jiang 0001, Dequan Wang, Xiangyang Xue 0001 |
ICME | 4 |
| 2017 | Evolving boxes for fast vehicle detectionabstractWe perform fast vehicle detection from traffic surveillance cameras. A novel deep learning framework, namely Evolving Boxes, is developed that proposes and refines the object boxes under different feature representations. Specifically, our framework is embedded with a light-weight proposal network to generate initial anchor boxes as well as to early discard unlikely regions; a fine-turning network produces detailed features for these candidate boxes. We show intriguingly that by applying different feature fusion techniques, the initial boxes can be refined for both localization and recognition. We evaluate our network on the recent DETRAC benchmark and obtain a significant improvement over the state-of-the-art Faster RCNN by 9.5% mAP. Further, our network achieves 9–13 FPS detection speed on a moderate commercial GPU. Li Wang 0033, Yao Lu 0028, Hong Wang 0014, Yingbin Zheng, Hao Ye 0005, Xiangyang Xue 0001 |
ICME | 6 |
| 2017 | Frame-Transformer Emotion Classification NetworkabstractEmotional content is a key ingredient in user-generated videos. However, due to the emotion sparsely expressed in the user-generated video, it is very difficult to analayze emotions in videos. In this paper, we propose a new architecture--Frame-Transformer Emotion Classification Network (FT-EC-net) to solve three highly correlated emotion analysis tasks: emotion recognition, emotion attribution and emotion-oriented summarization. We also contribute a new dataset for emotion attribution task by annotating the ground-truth labels of attribution segments. A comprehensive set of experiments on two datasets demonstrate the effectiveness of our framework. Jiarui Gao, Yanwei Fu 0001, Yu-Gang Jiang 0001, Xiangyang Xue 0001 |
ICMR | 4 |
| 2017 | Multi-task Deep Neural Network for Joint Face Recognition and Facial Attribute PredictionabstractDeep neural networks have significantly improved the performance of face recognition and facial attribute prediction, which however are still very challenging on the million scale dataset, i.e. MegaFace. In this paper, we for the first time, advocate a multi-task deep neural network for jointly learning face recognition and facial attribute prediction tasks. Extensive experimental evaluation clearly demonstrates the effectiveness of our architecture. Remarkably, on the largest face recognition benchmark -- MegaFace dataset, our networks can achieve the Rank-1 identication accuracy of 77.74% and face verication accuracy 79.24% TAR at 10-6 FAR, which are the best performance on the small protocol among all the publicly released methods. Zhanxiong Wang, Keke He, Yanwei Fu 0001, Rui Feng 0001, Yu-Gang Jiang 0001, Xiangyang Xue 0001 |
ICMR | 6 |
| 2017 | Adaptively Weighted Multi-task Deep Network for Person Attribute ClassificationabstractMulti-task learning aims to boost the performance of multiple prediction tasks by appropriately sharing relevant information among them. However, it always suffers from the negative transfer problem. And due to the diverse learning difficulties and convergence rates of different tasks, jointly optimizing multiple tasks is very challenging. To solve these problems, we present a weighted multi-task deep convolutional neural network for person attribute analysis. A novel validation loss trend algorithm is, for the first time proposed to dynamically and adaptively update the weight for learning each task in the training process. Extensive experiments on CelebA, Market-1501 attribute and Duke attribute datasets clearly show that state-of-the-art performance is obtained; and this validates the effectiveness of our proposed framework. Keke He, Zhanxiong Wang, Yanwei Fu 0001, Rui Feng 0001, Yu-Gang Jiang 0001, Xiangyang Xue 0001 |
ACM Multimedia | 6 |
| 2017 | Learning to Generate and Edit HairstylesabstractModeling hairstyles for classification, synthesis and image editing has many practical applications. However, existing hairstyle datasets, such as the Beauty e-Expert dataset, are too small for developing and evaluating computer vision models, especially the recent deep generative models such as generative adversarial network (GAN). In this paper, we contribute a new large-scale hairstyle dataset called Hairstyle30k, which is composed of 30k images containing 64 different types of hairstyles. To enable automated generating and modifying hairstyles in images, we also propose a novel GAN model termed Hairstyle GAN (H-GAN) which can be learned efficiently. Extensive experiments on the new dataset as well as existing benchmark datasets demonstrate the effectiveness of proposed H-GAN model Weidong Yin, Yanwei Fu 0001, Yiqing Ma, Yu-Gang Jiang 0001, Tao Xiang 0002, Xiangyang Xue 0001 |
ACM Multimedia | 6 |
| 2016 | Regional Gating Neural Networks for Multi-label Image Classification
Yurong Chen 0001, Jia-Ming Liu, Yu-Gang Jiang 0001, Xiangyang Xue 0001 |
BMVC | 6 |
| 2016 | Robust online visual tracking via a temporal ensemble frameworkabstractIn this paper, we propose a robust visual tracking method based on a temporal ensemble framework. Different from conventional ensemble-based trackers, which combine weak classifiers into a strong one using AdBoost in spatial fusion manners, our method adopts a powerful and efficient tracker integrated with its snapshots in different temporal windows of online tracking process to construct a temporal ensemble framework. Specifically, an adaptive correlation filter classifier is employed as the base tracker. During online tracking, the ensemble model determines the output through fusion of the base tracker's snapshots based on their response scores. By the temporal ensemble, accumulated errors caused by undesirable update can be corrected which greatly improves the robustness of the tracking system. Encouraging experimental results on challenging benchmark video sequences demonstrate that the proposed tracking method outperforms several state-of-the-art trackers in terms of both precision and robustness. Xiangyang Xue 0001 |
ICME | 2 |
| 2016 | Online video tracking using collaborative convolutional networksabstractRecently, convolutional neural network (CNN) models have achieved great success in many vision tasks. However, few attempts have been made to explore CNN for online model-free object tracking without time-consuming offline training. In this paper, we propose an online convolutional network (OC-N) for visual object tracking. To make the network less dependent on labeled data, K-means is employed to learn multistage filter banks for hierarchical feature learning. To preserve more spatial information for tracking, down-sampling and pooling operations are eliminated, which enables our system more sensitive to spatial variations. A regression model is adopted as the output layer of OCN to predict the position changes of the target. To deal with the stability-plasticity dilemma, two OCNs with different update rates are integrated to construct an ensemble framework. Experiments on challenging benchmark video sequences demonstrate that the proposed tracker outperforms several state-of-the-art methods. Xiangyang Xue 0001, Zhiyong An |
ICME | 2 |
| 2016 | Model-Based Deep Hand Pose Estimation
Xingyi Zhou, Qingfu Wan, Wei Zhang 0016, Xiangyang Xue 0001 |
IJCAI | 4 |
| 2016 | Multi-Stream Multi-Class Fusion of Deep Networks for Video ClassificationabstractThis paper studies deep network architectures to address the problem of video classification. A multi-stream framework is proposed to fully utilize the rich multimodal information in videos. Specifically, we first train three Convolutional Neural Networks to model spatial, short-term motion and audio clues respectively. Long Short Term Memory networks are then adopted to explore long-term temporal dynamics. With the outputs of the individual streams on multiple classes, we propose to mine class relationships hidden in the data from the trained models. The automatically discovered relationships are then leveraged in the multi-stream multi-class fusion process as a prior, indicating which and how much information is needed from the remaining classes, to adaptively determine the optimal fusion weights for generating the final scores of each class. Our contributions are two-fold. First, the multi-stream framework is able to exploit multimodal features that are more comprehensive than those previously attempted. Second, our proposed fusion method not only learns the best weights of the multiple network streams for each class, but also takes class relationship into account, which is known as a helpful clue in multi-class visual classification tasks. Our framework produces significantly better results than the state of the arts on two popular benchmarks, 92.2% on UCF-101 (without using audio) and 84.9% on Columbia Consumer Videos. Zuxuan Wu, Yu-Gang Jiang 0001, Xi Wang 0008, Hao Ye 0005, Xiangyang Xue 0001 |
ACM Multimedia | 5 |
| 2016 | Face Recognition via Active Annotation and LearningabstractIn this paper, we introduce an active annotation and learning framework for the face recognition task. Starting with an initial label deficient face image training set, we iteratively train a deep neural network and use this model to choose the examples for further manual annotation. We follow the active learning strategy and derive the Value of Information criterion to actively select candidate annotation images. During these iterations, the deep neural network is incrementally updated. Experimental results conducted on LFW benchmark and MS-Celeb-1M challenge demonstrate the effectiveness of our proposed framework. Hao Ye 0005, Weiyuan Shao, Hong Wang 0014, Jianqi Ma, Li Wang 0033, Yingbin Zheng, Xiangyang Xue 0001 |
ACM Multimedia | 7 |
| 2016 | Multiple task learning with flexible structure regularization
Jian Pu, Jun Wang 0006, Yu-Gang Jiang 0001, Xiangyang Xue 0001 |
Neurocomputing | 4 |
| 2016 | Flexible multi-task learning with latent task grouping
Jian Pu, Yu-Gang Jiang 0001, Rui Feng 0001, Xiangyang Xue 0001 |
Neurocomputing | 5 |
| 2015 | Cross-Modal Image Clustering via Canonical Correlation AnalysisabstractA new algorithm via Canonical Correlation Analysis (CCA) is developed in this paper to support more effective cross-modal image clustering for large-scale annotated image collections. It can be treated as a bi-media multimodal mapping problem and modeled as a correlation distribution over multimodal feature representations. It integrates the multimodal feature generation with the Locality Linear Coding (LLC) and co-occurrence association network, multimodal feature fusion with CCA, and accelerated hierarchical k-means clustering, which aims to characterize the correlations between the inter-related visual features in images and semantic features in captions, and measure their association degree more precisely. Very positive results were obtained in our experiments using a large quantity of public data. Cheng Jin 0001, Wenhui Mao, Yuejie Zhang, Xiangyang Xue 0001 |
AAAI | 5 |
| 2015 | Weakly supervised semantic segmentation for social imagesabstractImage semantic segmentation is the task of partitioning image into several regions based on semantic concepts. In this paper, we learn a weakly supervised semantic segmentation model from social images whose labels are not pixel-level but image-level; furthermore, these labels might be noisy. We present a joint conditional random field model leveraging various contexts to address this issue. More specifically, we extract global and local features in multiple scales by convolutional neural network and topic model. Inter-label correlations are captured by visual contextual cues and label co-occurrence statistics. The label consistency between image-level and pixel-level is finally achieved by iterative refinement. Experimental results on two real-world image datasets PASCAL VOC2007 and SIFT-Flow demonstrate that the proposed approach outperforms state-of-the-art weakly supervised methods and even achieves accuracy comparable with fully supervised methods. Wei Zhang 0016, Sheng Zeng, Dequan Wang, Xiangyang Xue 0001 |
CVPR | 4 |
| 2015 | Multiple Granularity Descriptors for Fine-Grained CategorizationabstractFine-grained categorization, which aims to distinguish subordinate-level categories such as bird species or dog breeds, is an extremely challenging task. This is due to two main issues: how to localize discriminative regions for recognition and how to learn sophisticated features for representation. Neither of them is easy to handle if there is insufficient labeled data. We leverage the fact that a subordinate-level object already has other labels in its ontology tree. These "free" labels can be used to train a series of CNN-based classifiers, each specialized at one grain level. The internal representations of these networks have different region of interests, allowing the construction of multi-grained descriptors that encode informative and discriminative features covering all the grain levels. Our multiple granularity framework can be learned with the weakest supervision, requiring only image-level label and avoiding the use of labor-intensive bounding box or part annotations. Experimental results on three challenging fine-grained image datasets demonstrate that our approach outperforms state-of-the-art algorithms, including those requiring strong labels. Dequan Wang, Jie Shao 0006, Wei Zhang 0016, Xiangyang Xue 0001, Zheng Zhang 0001 |
ICCV | 5 |
| 2015 | Evaluating Two-Stream CNN for Video ClassificationabstractVideos contain very rich semantic information. Traditional hand-crafted features are known to be inadequate in analyzing complex video semantics. Inspired by the huge success of the deep learning methods in analyzing image, audio and text data, significant efforts are recently being devoted to the design of deep nets for video analytics. Among the many practical needs, classifying videos (or video clips) based on their major semantic categories (e.g.,"skiing") is useful in many applications. In this paper, we conduct an in-depth study to investigate important implementation options that may affect the performance of deep nets on video classification. Our evaluations are conducted on top of a recent two-stream convolutional neural network (CNN) pipeline, which uses both static frames and motion optical flows, and has demonstrated competitive performance against the state-of-the-art methods. In order to gain insights and to arrive at a practical guideline, many important options are studied, including network architectures, model fusion, learning parameters and the final prediction methods. Based on the evaluations, very competitive results are attained on two popular video classification benchmarks. We hope that the discussions and conclusions from this work can help researchers in related fields to quickly set up a good basis for further investigations along this very promising direction. Hao Ye 0005, Zuxuan Wu, Xi Wang 0008, Yu-Gang Jiang 0001, Xiangyang Xue 0001 |
ICMR | 6 |
| 2015 | Modeling Spatial-Temporal Clues in a Hybrid Deep Learning Framework for Video ClassificationabstractClassifying videos according to content semantics is an important problem with a wide range of applications. In this paper, we propose a hybrid deep learning framework for video classification, which is able to model static spatial information, short-term motion, as well as long-term temporal clues in the videos. Specifically, the spatial and the short-term motion features are extracted separately by two Convolutional Neural Networks (CNN). These two types of CNN-based features are then combined in a regularized feature fusion network for classification, which is able to learn and utilize feature relationships for improved performance. In addition, Long Short Term Memory (LSTM) networks are applied on top of the two features to further model longer-term temporal clues. The main contribution of this work is the hybrid learning framework that can model several important aspects of the video data. We also show that (1) combining the spatial and the short-term motion features in the regularized fusion network is better than direct classification and fusion using the CNN with a softmax layer, and (2) the sequence-based LSTM is highly complementary to the traditional classification strategy without considering the temporal frame orders. Extensive experiments are conducted on two popular and challenging benchmarks, the UCF-101 Human Actions and the Columbia Consumer Videos (CCV). On both benchmarks, our framework achieves very competitive performance: 91.3% on the UCF-101 and 83.5% on the CCV. Zuxuan Wu, Xi Wang 0008, Yu-Gang Jiang 0001, Hao Ye 0005, Xiangyang Xue 0001 |
ACM Multimedia | 5 |
| 2015 | Human Action Recognition in Unconstrained Videos by Explicit Motion ModelingabstractHuman action recognition in unconstrained videos is a challenging problem with many applications. Most state-of-the-art approaches adopted the well-known bag-of-features representations, generated based on isolated local patches or patch trajectories, where motion patterns, such as object-object and object-background relationships are mostly discarded. In this paper, we propose a simple representation aiming at modeling these motion relationships. We adopt global and local reference points to explicitly characterize motion information, so that the final representation is more robust to camera movements, which widely exist in unconstrained videos. Our approach operates on the top of visual codewords generated on dense local patch trajectories, and therefore, does not require foreground-background separation, which is normally a critical and difficult step in modeling object relationships. Through an extensive set of experimental evaluations, we show that the proposed representation produces a very competitive performance on several challenging benchmark data sets. Further combining it with the standard bag-of-features or Fisher vector representations can lead to substantial improvements. Yu-Gang Jiang 0001, Qi Dai 0001, Wei Liu 0005, Xiangyang Xue 0001, Chong-Wah Ngo |
IEEE Trans. Image Process. | 4 |
| 2014 | Predicting Emotions in User-Generated VideosabstractUser-generated video collections are expanding rapidly in recent years, and systems for automatic analysis of these collections are in high demands. While extensive research efforts have been devoted to recognizing semantics like "birthday party" and "skiing", little attempts have been made to understand the emotions carried by the videos, e.g., "joy" and "sadness". In this paper, we propose a comprehensive computational framework for predicting emotions in user-generated videos. We first introduce a rigorously designed dataset collected from popular video-sharing websites with manual annotations, which can serve as a valuable benchmark for future research. A large set of features are extracted from this dataset, ranging from popular low-level visual descriptors, audio features, to high-level semantic attributes. Results of a comprehensive set of experiments indicate that combining multiple types of features---such as the joint use of the audio and visual clues---is important, and attribute features such as those containing sentiment-level semantics are very effective. Yu-Gang Jiang 0001, Baohan Xu, Xiangyang Xue 0001 |
AAAI | 3 |
| 2014 | Semantic Segmentation Using Multiple Graphs with Block-Diagonal ConstraintsabstractIn this paper we propose a novel method for image semantic segmentation using multiple graphs. The multiview affinity graph is constructed by leveraging the consistency between semantic space and multiple visualspaces. With block-diagonal constraints, we enforce the affinity matrix to be sparse such that the pairwise potential for dissimilar superpixels is close to zero. By a divide-and-conquer strategy, the optimizationfor learning affinity matrix is decomposed into several subproblems that can be solved in parallel. Using the neighborhood relationship between superpixels and the consistency between affinity matrix and labelconfidencematrix, we infer the semantic label for each superpixel of unlabeled images by minimizing an objective whose closed form solution can be easily obtained. Experimental results on two real-world image datasetsdemonstrate the effectiveness of our method. Ke Zhang 0028, Wei Zhang 0016, Sheng Zeng, Xiangyang Xue 0001 |
AAAI | 4 |
| 2014 | Which Looks Like Which: Exploring Inter-class Relationships in Fine-Grained Visual Categorization
Jian Pu, Yu-Gang Jiang 0001, Jun Wang 0006, Xiangyang Xue 0001 |
ECCV (3) | 4 |
| 2014 | Exploring Inter-feature and Inter-class Relationships with Deep Neural Networks for Video ClassificationabstractVideos contain very rich semantics and are intrinsically multimodal. In this paper, we study the challenging task of classifying videos according to their high-level semantics such as human actions or complex events. Although extensive efforts have been paid to study this problem, most existing works combined multiple features using simple fusion strategies and neglected the exploration of inter-class semantic relationships. In this paper, we propose a novel unified framework that jointly learns feature relationships and exploits the class relationships for improved video classification performance. Specifically, these two types of relationships are learned and utilized by rigorously imposing regularizations in a deep neural network (DNN). Such a regularized DNN can be efficiently launched using a GPU implementation with an affordable training cost. Through arming the DNN with better capability of exploring both the inter-feature and the inter-class relationships, the proposed regularized DNN is more suitable for identifying video semantics. With extensive experimental evaluations, we demonstrate that the proposed framework exhibits superior performance over several state-of-the-art approaches. On the well-known Hollywood2 and Columbia Consumer Video benchmarks, we obtain to-date the best reported results: 65.7% and 70.6% respectively in terms of mean average precision. Zuxuan Wu, Yu-Gang Jiang 0001, Jun Wang 0006, Jian Pu, Xiangyang Xue 0001 |
ACM Multimedia | 5 |
| 2014 | Addressing cold start in recommender systems: a semi-supervised co-training algorithmabstractCold start is one of the most challenging problems in recommender systems. In this paper we tackle the cold-start problem by proposing a context-aware semi-supervised co-training method named CSEL. Specifically, we use a factorization model to capture fine-grained user-item context. Then, in order to build a model that is able to boost the recommendation performance by leveraging the context, we propose a semi-supervised ensemble learning algorithm. The algorithm constructs different (weak) prediction models using examples with different contexts and then employs the co-training strategy to allow each (weak) prediction model to learn from the other prediction models. The method has several distinguished advantages over the standard recommendation methods for addressing the cold-start problem. First, it defines a fine-grained context that is more accurate for modeling the user-item preference. Second, the method can naturally support supervised learning and semi-supervised learning, which provides a flexible way to incorporate the unlabeled data. Mi Zhang 0001, Jie Tang 0001, Xuchen Zhang, Xiangyang Xue 0001 |
SIGIR | 4 |
| 2014 | Cost-Sensitive Multi-View Learning MachineabstractMulti-view learning aims to effectively learn from data represented by multiple independent sets of attributes, where each set is taken as one view of the original data. In real-world application, each view should be acquired in unequal cost. Taking web-page classification for example, it is cheaper to get the words on itself (view one) than to get the words contained in anchor texts of inbound hyper-links (view two). However, almost all the existing multi-view learning does not consider the cost of acquiring the views or the cost of evaluating them. In this paper, we support that different views should adopt different representations and lead to different acquisition cost. Thus we develop a new view-dependent cost different from the existing both class-dependent cost and example-dependent cost. To this end, we generalize the framework of multi-view learning with the cost-sensitive technique and further propose a Cost-sensitive Multi-View Learning Machine named CMVLM for short. In implementation, we take into account and measure both the acquisition cost and the discriminant scatter of each view. Then through eliminating the useless views with a predefined threshold, we use the reserved views to train the final classifier. The experimental results on a broad range of data sets including the benchmark UCI, image, and bioinformatics data sets validate that the proposed algorithm can effectively reduce the total cost and have a competitive even better classification performance. The contributions of this paper are that: (1) first proposing a view-dependent cost; (2) establishing a cost-sensitive multi-view learning framework; (3) developing a wrapper technique that is universal to most multiple kernel based classifier. Zhe Wang 0002, Mingzhe Lu, Zengxin Niu, Xiangyang Xue 0001, Daqi Gao |
Int. J. Pattern Recognit. Artif. Intell. | 4 |
| 2014 | Bounding the Advantage of Multicast Network Coding in General Network ModelsabstractNetwork coding encourages information flow mixing in a network. It helps increase the throughput and reduce the cost of data transmission, especially for one-to-many multicast applications. An interesting problem is to understand and quantify the coding advantage and cost advantage, i.e., the potential benefits of network coding, as compared to routing, in terms of increasing throughput and reducing transmission cost, respectively. Two classic network models were considered in previous studies: directed networks and undirected networks. This work further focuses on two types of parameterized networks, including bidirected networks and hyper-networks, generalizing the directed and the undirected network models, respectively. We prove upper- and lower-bounds on multicast coding advantage and cost advantage in these models. Xunrui Yin, Yan Wang 0058, Zongpeng Li, Xin Wang 0002, Jin Zhao 0001, Xiangyang Xue 0001 |
IEEE Trans. Commun. | 6 |
| 2014 | A Graph Minor Perspective to Multicast Network CodingabstractNetwork coding encourages information coding across a communication network. While the necessity, benefit and complexity of network coding are sensitive to the underlying graph structure of a network, existing theory on network coding often treats the network topology as a black box, focusing on algebraic or information theoretic aspects of the problem. This paper aims at an in-depth examination of the relation between algebraic coding and network topologies. We mathematically establish a series of results along the direction of: if network coding is necessary/beneficial, or if a particular finite field is required for coding, then the network must have a corresponding hidden structure embedded in its underlying topology, and such embedding is computationally efficient to verify. Specifically, we first formulate a meta-conjecture, the NC-minor conjecture, that articulates such a connection between graph theory and network coding, in the language of graph minors. We next prove that the NC-minor conjecture for multicasting two information flows is almost equivalent to the Hadwiger conjecture, which connects graph minors with graph coloring. Such equivalence implies the existence of K4, K5, K6, and KO(q/log q) minors, for networks that require F3, F4, F5, and Fqto multicast two flows, respectively. We finally prove that, for the general case of multicasting arbitrary number of flows, network coding can make a difference from routing only if the network contains a K4minor, and this minor containment result is tight. Practical implications of the above results are discussed. Xunrui Yin, Yan Wang 0058, Zongpeng Li, Xin Wang 0002, Xiangyang Xue 0001 |
IEEE Trans. Inf. Theory | 5 |
| 2013 | Understanding and Predicting Interestingness of VideosabstractThe amount of videos available on the Web is growing explosively. While some videos are very interesting and receive high rating from viewers, many of them are less interesting or even boring. This paper conducts a pilot study on the understanding of human perception of video interestingness, and demonstrates a simple computational method to identify more interesting videos. To this end we first construct two datasets of Flickr and YouTube videos respectively. Human judgements of interestingness are collected and used as the ground-truth for training computational models. We evaluate several off-the-shelf visual and audio features that are potentially useful for predicting interestingness on both datasets. Results indicate that audio and visual features are equally important and the combination of both modalities shows very promising results. Yu-Gang Jiang 0001, Rui Feng 0001, Xiangyang Xue 0001, Yingbin Zheng, Hanfang Yang |
AAAI | 4 |
| 2013 | Multiple Task Learning Using Iteratively Reweighted Least Square
Jian Pu, Yu-Gang Jiang 0001, Jun Wang 0006, Xiangyang Xue 0001 |
IJCAI | 4 |
| 2013 | Automatic Name-Face Alignment to Enable Cross-Media News Retrieval
Yuejie Zhang, Cheng Jin 0001, Xiangyang Xue 0001, Jianping Fan 0001 |
IJCAI | 5 |
| 2013 | Multi-View Embedding Learning for Incompletely Labeled Data
Wei Zhang 0016, Ke Zhang 0028, Pan Gu, Xiangyang Xue 0001 |
IJCAI | 4 |
| 2013 | Sparse Reconstruction for Weakly Supervised Semantic Segmentation
Ke Zhang 0028, Wei Zhang 0016, Yingbin Zheng, Xiangyang Xue 0001 |
IJCAI | 4 |
| 2013 | A graph minor perspective to network coding: Connecting algebraic coding with network topologiesabstractNetwork Coding encourages information coding across a communication network. While the necessity, benefit and complexity of network coding are sensitive to the underlying graph structure of a network, existing theory on network coding often treats the network topology as a black box, focusing on algebraic or information theoretic aspects of the problem. This work aims at an in-depth examination of the relation between algebraic coding and network topologies. We mathematically establish a series of results along the direction of: if network coding is necessary/beneficial, or if a particular finite field is required for coding, then the network must have a corresponding hidden structure embedded in its underlying topology, and such embedding is computationally efficient to verify. Specifically, we first formulate a meta-conjecture, the NC-Minor Conjecture, that articulates such a connection between graph theory and network coding, in the language of graph minors. We next prove that the NC-Minor Conjecture is almost equivalent to the Hadwiger Conjecture, which connects graph minors with graph coloring. Such equivalence implies the existence of K4, K5, K6, and KO(q/ log q)minors, for networks requiring F3, F4, F5and Fq, respectively. We finally prove that network coding can make a difference from routing only if the network contains a K4minor, and this minor containment result is tight. Practical implications of the above results are discussed. Xunrui Yin, Yan Wang 0058, Xin Wang 0002, Xiangyang Xue 0001, Zongpeng Li |
INFOCOM | 4 |
| 2013 | An efficient Kernel-based matrixized least squares support vector machine
Zhe Wang 0002, Xisheng He, Daqi Gao, Xiangyang Xue 0001 |
Neural Comput. Appl. | 4 |
| 2013 | Multi-Stage Non-Negative Matrix Factorization for Monaural Singing Voice SeparationabstractSeparating singing voice from music accompaniment can be of interest for many applications such as melody extraction, singer identification, lyrics alignment and recognition, and content-based music retrieval. In this paper, a novel algorithm for singing voice separation in monaural mixtures is proposed. The algorithm consists of two stages, where non-negative matrix factorization (NMF) is applied to decompose the mixture spectrograms with long and short windows respectively. A spectral discontinuity thresholding method is devised for the long-window NMF to select out NMF components originating from pitched instrumental sounds, and a temporal discontinuity thresholding method is designed for the short-window NMF to pick out NMF components that are from percussive sounds. By eliminating the selected components, most pitched and percussive elements of the music accompaniment are filtered out from the input sound mixture, with little effect on the singing voice. Extensive testing on the MIR-1K public dataset of 1000 short audio clips and the Beach-Boys dataset of 14 full-track real-world songs showed that the proposed algorithm is both effective and efficient. Bilei Zhu, Wei Li 0012, Ruijiang Li, Xiangyang Xue 0001 |
IEEE Trans. Speech Audio Process. | 4 |
| 2013 | A Segmentation and Graph-Based Video Sequence Matching Method for Video Copy DetectionabstractWe propose in this paper a segmentation and graph-based video sequence matching method for video copy detection. Specifically, due to the good stability and discriminative ability of local features, we use SIFT descriptor for video content description. However, matching based on SIFT descriptor is computationally expensive for large number of points and the high dimension. Thus, to reduce the computational complexity, we first use the dual-threshold method to segment the videos into segments with homogeneous content and extract keyframes from each segment. SIFT features are extracted from the keyframes of the segments. Then, we propose an SVD-based method to match two video frames with SIFT point set descriptors. To obtain the video sequence matching result, we propose a graph-based method. It can convert the video sequence matching into finding the longest path in the frame matching-result graph with time constraint. Experimental results demonstrate that the segmentation and graph-based video sequence matching method can detect video copies effectively. Also, the proposed method has advantages. Specifically, it can automatically find optimal sequence matching result from the disordered matching results based on spatial feature. It can also reduce the noise caused by spatial feature matching. And it is adaptive to video frame rate changes. Experimental results also demonstrate that the proposed method can obtain a better tradeoff between the effectiveness and the efficiency of video copy detection. Hong Lu 0001, Xiangyang Xue 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2013 | Query-Adaptive Image Search With Hash CodesabstractScalable image search based on visual similarity has been an active topic of research in recent years. State-of-the-art solutions often use hashing methods to embed high-dimensional image features into Hamming space, where search can be performed in real-time based on Hamming distance of compact hash codes. Unlike traditional metrics (e.g., Euclidean) that offer continuous distances, the Hamming distances are discrete integer values. As a consequence, there are often a large number of images sharing equal Hamming distances to a query, which largely hurts search results where fine-grained ranking is very important. This paper introduces an approach that enables query-adaptive ranking of the returned images with equal Hamming distances to the queries. This is achieved by firstly offline learning bitwise weights of the hash codes for a diverse set of predefined semantic concept classes. We formulate the weight learning process as a quadratic programming problem that minimizes intra-class distance while preserving inter-class relationship captured by original raw image features. Query-adaptive weights are then computed online by evaluating the proximity between a query and the semantic concept classes. With the query-adaptive bitwise weights, returned images can be easily ordered by weighted Hamming distance at a finer-grained hash code level rather than the original Hamming distance level. Experiments on a Flickr image dataset show clear improvements from our proposed approach. Yu-Gang Jiang 0001, Jun Wang 0006, Xiangyang Xue 0001, Shih-Fu Chang |
IEEE Trans. Multim. | 3 |
| 2012 | Semantic context learning with large-scale weakly-labeled image setabstractThere are a large number of images available on the web; meanwhile, only a subset of web images can be labeled by professionals because manual annotation is time-consuming and labor-intensive. Although we can now use the collaborative image tagging system, e.g., Flickr, to get a lot of tagged images provided by Internet users, these labels may be incorrect or incomplete. Furthermore, semantics richness requires more than one label to describe one image in real applications, and multiple labels usually interact with each other in semantic space. It is of significance to learn semantic context with large-scale weakly-labeled image set in the task of multi-label annotation. In this paper, we develop a novel method to learn semantic context and predict the labels of web images in a semi-supervised framework. To address the scalability issue, a small number of exemplar images are first obtained to cover the whole data cloud; then the label vector of each image is estimated as a local combination of the exemplar label vectors. Visual context, semantic context, and neighborhood consistency in both visual and semantic spaces are sufficiently leveraged in the proposed framework. Finally, the semantic context and the label confidence vectors for exemplar images are both learned in an iterative way. Experimental results on the real-world image dataset demonstrate the effectiveness of our method. Yao Lu 0028, Wei Zhang 0016, Ke Zhang 0028, Xiangyang Xue 0001 |
CIKM | 4 |
| 2012 | Parallel proximal support vector machine for high-dimensional pattern classificationabstractProximal support vector machine (PSVM) is a simple but effective classifier, especially for solving large-scale data classification problems. An inherent deficiency of PSVM lies on its inefficiency for dealing with high-dimensional data. In this paper, we propose a parallel version of PSVM (PPSVM). Based on random dimensionality partitioning, PPSVM can obtain partitioned local model parameters in parallel, with combined parameters to form the final global solution. In fact, PPSVM enjoys two properties: 1) It can calculate model parameters in parallel and is therefore a fast learning method with theoretically proved convergence; and 2) It can avoid the inversion of large matrix, which makes it suitable for high-dimensional data. In the paper, we also propose a random PPSVM with randomly partitioned data in each iteration to improve the performance of PSVM. Experimental results on real-world data demonstrate that the proposed methods can obtain similar or even better prediction accuracy than PSVM with much better runtime efficiency. Zhenfeng Zhu, Xingquan Zhu 0001, Yangdong Ye, Yue-Fei Guo, Xiangyang Xue 0001 |
CIKM | 5 |
| 2012 | Learning attention map from imagesabstractWhile bottom-up and top-down processes have shown effectiveness during predicting attention and eye fixation maps on images, in this paper, inspired by the perceptual organization mechanism before attention selection, we propose to utilize figure-ground maps for the purpose. So as to take both pixel-wise and region-wise interactions into consideration when predicting label probabilities for each pixel, we develop a context-aware model based on multiple segmentation to obtain final results. The MIT attention dataset [14] is applied finally to evaluate both new features and model. Quantitative experiments demonstrate that figure-ground cues are valid in predicting attention selection, and our proposed model produces improvements over baseline method. Yao Lu 0028, Wei Zhang 0016, Cheng Jin 0001, Xiangyang Xue 0001 |
CVPR | 4 |
| 2012 | Trajectory-Based Modeling of Human Actions with Motion Reference Points
Yu-Gang Jiang 0001, Qi Dai 0001, Xiangyang Xue 0001, Wei Liu 0005, Chong-Wah Ngo |
ECCV (5) | 3 |
| 2012 | Learning Hybrid Part Filters for Scene Recognition
Yingbin Zheng, Yu-Gang Jiang 0001, Xiangyang Xue 0001 |
ECCV (5) | 3 |
| 2012 | Groupwise Constrained Reconstruction for Subspace Clustering
Ruijiang Li, Bin Li 0015, Cheng Jin 0001, Xiangyang Xue 0001 |
ICML | 4 |
| 2012 | On benefits of network coding in bidirected networks and hyper-networksabstractNetwork coding is a technique that allows information flows to be encoded while routed across a data network. It was shown that network coding helps increase the throughput and reduce the cost of data transmission, especially for one-to-many multicast applications. An important direction in network coding research is to understand and quantify the coding advantage and cost advantage, i.e., the potential benefits of network coding, as compared to routing, in terms of increasing throughput and reducing transmission cost, respectively. Two classic network models were considered in previous studies of coding advantage: directed networks and undirected networks. The study of coding advantage in this work further focuses on two types of parameterized networks, including bidirected networks and hyper-networks, which generalizes the directed and the undirected network models, respectively. With proper parameter setting, more realistic modeling of networks in practice can be achieved. We prove upper-bounds and lower-bounds on the coding advantage for multicast in these models. Some of our bounds are new and unknown before, some improve upon previously proven bounds, and some answer open questions in the literature. Xunrui Yin, Xin Wang 0002, Jin Zhao 0001, Xiangyang Xue 0001, Zongpeng Li |
INFOCOM | 4 |
| 2012 | Min-cost multicast networks in Euclidean spaceabstractSpace information flow is a new field of research recently proposed by Li and Wu [1], [2]. It studies the transmission of information in a geometric space, where information flows can be routed along any trajectories, and can be encoded wherever they meet. The goal is to satisfy given end-to-end unicast/multicast throughput demands, while minimizing a natural bandwidth-distance sum-product (network volume). Space information flow models the design of a blueprint for a minimum-cost network. We study the multicast version of the space information flow problem, in Euclidean spaces. We present a simple example that demonstrates the design of an information network is indeed different from that of a transportation network. We discuss properties of optimal multicast network embedding, prove that network coding does not make a difference in the basic case of 1-to-2 multicast, and prove upper-bounds on the number of relay nodes required in an optimal acyclic multicast network. Xunrui Yin, Yan Wang 0058, Xin Wang 0002, Xiangyang Xue 0001, Zongpeng Li |
ISIT | 4 |
| 2012 | A fast video event recognition system and its application to video searchabstractTechniques for recognizing complex events in diverse Internet videos are important in many applications. State-of-the-art video event recognition approaches normally involve modules that demand extensive computation, which prevents their application to large scale problems. In this demonstration, we present a fast video event recognition system, which requires just a few seconds to process a general YouTube video with a few minutes of duration. The development of this system is grounded on several important findings from a large set of empirical studies, where we systematically evaluated many technical options for each critical module of a present-day video event recognition framework. Pooling the insights gained from this study leads to a speeded-up event recognition system that is 220-times faster than a decent baseline while still has a high degree of recognition accuracy. We also demonstrate the technical feasibility of using event recognition results as the sole clue for video search, where the similarity of videos is determined based on the consistency of the event recognition confidence scores. We showcase this capability using an Internet video dataset containing about 10 thousands of YouTube videos. Very promising results were observed. Yu-Gang Jiang 0001, Qi Dai 0001, Yingbin Zheng, Xiangyang Xue 0001 |
ACM Multimedia | 4 |
| 2012 | Semi-supervised multi-instance multi-label learning for video annotation taskabstractTraditional approaches for automatic video annotation usually represent one video clip with a flat feature vector, neglecting the fact that video data contain natural structures. It is also noteworthy that a video clip is often relevant to multiple concepts. Indeed, the video annotation task is inherently a Multi-Instance Multi-Label learning (MIML) problem. Considering that manually annotating videos is labor-intensive and time-consuming, this paper proposes a semi-supervised MIML approach, SSMIML, which is able to exploit abundant unannotated videos to help improve the annotation performance. This approach takes label correlations into account, and enforces similar instances to share similar multi-labels. Evaluation on TREVID 2005 show that the proposed approach outperforms several state-of-the-art methods. Xin-Shun Xu, Yuan Jiang 0001, Xiangyang Xue 0001, Zhi-Hua Zhou |
ACM Multimedia | 3 |
| 2012 | A Double-Ranking Strategy for Long-Tail Product RecommendationabstractIn this paper we attempt to retrieve the items in the long-tail for top-N recommendation. That is, to recommend products that the end-user likes, but that are not generally popular, which has been getting more and more notice lately. By analysing the existing issue of current recommendation algorithms, a strategy is proposed that succeeds in maintaining recommendation accuracy while reducing the concentration of the recommendation on popular items in the system. Evaluating on the publicly available Movie lens and Yahoo! datasets, the results show the recommendation algorithm proposed in this work retrieves items in the users' relatively unpopular tastes without losing the performance in their popular tastes, which ultimately results in a better overall accuracy for the system. Mi Zhang 0001, Neil J. Hurley, Wei Li 0012, Xiangyang Xue 0001 |
Web Intelligence | 4 |
| 2012 | Inverse matrix-free incremental proximal support vector machine
Zhenfeng Zhu, Xingquan Zhu 0001, Yue-Fei Guo, Yangdong Ye, Xiangyang Xue 0001 |
Decis. Support Syst. | 5 |
| 2012 | A covariance-free iterative algorithm for distributed principal component analysis on vertically partitioned data
Yue-Fei Guo, Xiaodong Lin 0004, Zhou Teng, Xiangyang Xue 0001, Jianping Fan 0001 |
Pattern Recognit. | 4 |
| 2012 | A simplified multi-class support vector machine with reduced dual optimization
Xisheng He, Zhe Wang 0002, Cheng Jin 0001, Yingbin Zheng, Xiangyang Xue 0001 |
Pattern Recognit. Lett. | 5 |
| 2012 | Gradient Ordinal Signature and Fixed-Point Embedding for Efficient Near-Duplicate Video DetectionabstractIn order to meet the requirement of large scale real-time near-duplicate video detection, this paper has achieved two goals. First, this paper proposes a more compact local image descriptor which is termed as gradient ordinal signature (GOS). GOS not only has the advantages of low dimension, simplicity in computation, and high discrimination but also is invariant to mirror reflection, rotation, and scale changes. Second, applying the characteristics of the proposed GOS and combining with the embedding theory of metric spaces, this paper proposes an efficient similarity search method based on the fixed-point embedding (FE). A main advantage of FE is that its parameters have good controllability, and its performance is stable and not sensitive to dataset changes. On the whole, the goal of our approach focuses on the speed rather than the accuracy of near-duplicate video detection. We have evaluated our method on four different settings to verify the two goals. Specifically, the tests include image and video datasets, respectively, to evaluate the performance of GOS. Experimental results demonstrate the effectiveness, efficiency, and lower memory usage of GOS. Furthermore, the third test compares FE with locality sensitivity hashing. FE also shows a speed improvement of about ten times and saves more than 60% in memory usage. The fourth test demonstrates that the combination of GOS and FE for near-duplicate video detection can achieve better overall efficiency than the state-of-the-art methods. Hong Lu 0001, Zhaohui Wen, Xiangyang Xue 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2012 | Fast Semantic Diffusion for Large-Scale Context-Based Image and Video AnnotationabstractExploring context information for visual recognition has recently received significant research attention. This paper proposes a novel and highly efficient approach, which is named semantic diffusion, to utilize semantic context for large-scale image and video annotation. Starting from the initial annotation of a large number of semantic concepts (categories), obtained by either machine learning or manual tagging, the proposed approach refines the results using a graph diffusion technique, which recovers the consistency and smoothness of the annotations over a semantic graph. Different from the existing graph-based learning methods that model relations among data samples, the semantic graph captures context by treating the concepts as nodes and the concept affinities as the weights of edges. In particular, our approach is capable of simultaneously improving annotation accuracy and adapting the concept affinities to new test data. The adaptation provides a means to handle domain change between training and test data, which often occurs in practice. Extensive experiments are conducted to improve concept annotation results using Flickr images and TV program videos. Results show consistent and significant performance gain (10 +% on both image and video data sets). Source codes of the proposed algorithms are available online. Yu-Gang Jiang 0001, Qi Dai 0001, Jun Wang 0006, Chong-Wah Ngo, Xiangyang Xue 0001, Shih-Fu Chang |
IEEE Trans. Image Process. | 5 |
| 2011 | Tracking User-Preference Varying Speed in Collaborative FilteringabstractIn real-world recommender systems, some users are easily influenced by new products and whereas others are unwilling to change their minds. So the preference varying speeds for users are different. Based on this observation, we propose a dynamic nonlinear matrix factorization model for collaborative filtering, aimed to improve the rating prediction performance as well as track the preference varying speeds for different users. We assume that user-preference changes smoothly over time, and the preference varying speeds for users are different. These two assumptions are incorporated into the proposed model as prior knowledge on user feature vectors, which can be learned efficiently by MAP estimation. The experimental results show that our method not only achieves state-of-the-art performance in the rating prediction task, but also provides an effective way to track user-preference varying speed. Ruijiang Li, Bin Li 0015, Cheng Jin 0001, Xiangyang Xue 0001, Xingquan Zhu 0001 |
AAAI | 4 |
| 2011 | Transfer active learningabstractActive learning traditionally assumes that labeled and unlabeled samples are subject to the same distributions and the goal of an active learner is to label the most informative unlabeled samples. In reality, situations may exist that we may not have unlabeled samples from the same domain as the labeled samples (i.e. target domain), whereas samples from auxiliary domains might be available. Under such situations, an interesting question is whether an active learner can actively label samples from auxiliary domains to benefit the target domain. In this paper, we propose a transfer active learning method, namely Transfer Active SVM (TrAcSVM), which uses a limited number of target instances to iteratively discover and label informative auxiliary instances. TrAcSVM employs an extended sigmoid function as instance weight updating approach to adjust the models for prediction of (newly arrived) target data. Experimental results on real-world data sets demonstrate that TrAcSVM obtains better efficiency and prediction accuracy than its peers. Zhenfeng Zhu, Xingquan Zhu 0001, Yangdong Ye, Yue-Fei Guo, Xiangyang Xue 0001 |
CIKM | 5 |
| 2011 | Salient Object Detection using concavity contextabstractConvexity (concavity) is a bottom-up cue to assign figure-ground relation in the perceptual organization [18]. It suggests that region on the convex side of a curved boundary tend to be figural. To explore the validity of this cue in the task of salient object detection, we segment the images in a test dataset into superpixels, and then locate the concave arcs and their bounding boxes along boundary of superpixels. Ecological statistics indicate that such bounding box contains salient object with a large probability. To utilize this spatial context information, i.e. concavity context, we follow the multi-scale analysis of human visual perception and design a hierarchical model. The model yields an affinity graph over candidate superpixels, in which weights between vertices are determined by the summation of concavity context on different scales in the hierarchy. Finally a graph-cut algorithm is performed to separate the salient and background objects. Evaluation on MSRA Salient Object Detection (SOD) dataset shows that concavity context is effective, and our approach provides improvement over state-of-the-art feature-based algorithms. Yao Lu 0028, Wei Zhang 0016, Hong Lu 0001, Xiangyang Xue 0001 |
ICCV | 4 |
| 2011 | Correlative multi-label multi-instance image annotationabstractIn this paper, each image is viewed as a bag of local regions, as well as it is investigated globally. A novel method is developed for achieving multi-label multi-instance image annotation, where image-level (bag-level) labels and region-level (instance-level) labels are both obtained. The associations between semantic concepts and visual features are mined both at the image level and at the region level. Inter-label correlations are captured by a co-occurence matrix of concept pairs. The cross-level label coherence encodes the consistency between the labels at the image level and the labels at the region level. The associations between visual features and semantic concepts, the correlations among the multiple labels, and the cross-level label coherence are sufficiently leveraged to improve annotation performance. Structural max-margin technique is used to formulate the proposed model and multiple interrelated classifiers are learned jointly. To leverage the available image-level labeled samples for the model training, the region-level label identification on the training set is firstly accomplished by building the correspondences between the multiple bag-level labels and the image regions. JEC distance based kernels are employed to measure the similarities both between images and between regions. Experimental results on real image datasets MSRC and Corel demonstrate the effectiveness of our method. Xiangyang Xue 0001, Wei Zhang 0016, Jianping Fan 0001, Yao Lu 0028 |
ICCV | 1 |
| 2011 | Cross-Domain Collaborative Filtering over TimeabstractCollaborative filtering (CF) techniques recommend items to users based on their historical ratings. In real-world scenarios, user interests may drift over time since they are affected by moods, contexts, and pop culture trends. This leads to the fact that a user's historical ratings comprise many aspects of user interests spanning a long time period. However, at a certain time slice, one user's interest may only focus on one or a couple of aspects. Thus, CF techniques based on the entire historical ratings may recommend inappropriate items. In this paper, we consider modeling user-interest drift over time based on the assumption that each user has multiple counterparts over temporal domains and successive counterparts are closely related. We adopt the cross-domain CF framework to share the static group-level rating matrix across temporal domains, and let user-interest distribution over item groups drift slightly between successive temporal domains. The derived method is based on a Bayesian latent factor model which can be inferred using Gibbs sampling. Our experimental results show that our method can achieve state-of-the-art recommendation performance as well as explicitly track and visualize user-interest drift over time. Bin Li 0015, Xingquan Zhu 0001, Ruijiang Li, Chengqi Zhang, Xiangyang Xue 0001, Xindong Wu 0001 |
IJCAI | 5 |
| 2011 | Learning Inter-Related Statistical Query Translation Models for English-Chinese Bi-Directional CLIRabstractTo support more precise query translation for English-Chinese Bi-Directional Cross-Language Information Retrieval (CLIR), we have developed a novel framework by integrating a semantic network to characterize the correlations between multiple inter-related text terms of interest and learn their inter-related statistical query translation models. First, a semantic network is automatically generated from large-scale English-Chinese bilingual parallel corpora to characterize the correlations between a large number of text terms of interest. Second, the semantic network is exploited to learn the statistical query translation models for such text terms of interest. Finally, these inter-related query translation models are used to translate the queries more precisely and achieve more effective CLIR. Our experiments on a large number of official public data have obtained very positive results. Yuejie Zhang, Lei Cen, Cheng Jin 0001, Xiangyang Xue 0001, Jianping Fan 0001 |
IJCAI | 4 |
| 2011 | Fusion of Multiple Features and Supervised Learning for Chinese OOV Term Detection and POS GuessingabstractIn this paper, to support more precise Chinese Out-of-Vocabulary (OOV) term detection and Part-of-Speech (POS) guessing, a unified mechanism is proposed and formulated based on the fusion of multiple features and supervised learning. Besides all the traditional features, the new features for statistical information and global contexts are introduced, as well as some constraints and heuristic rules, which reveal the relationships among OOV term candidates. Our experiments on the Chinese corpora from both People’s Daily and SIGHAN 2005 have achieved the consistent results, which are better than those acquired by pure rule-based or statistics-based models. From the experimental results for combining our model with Chinese monolingual retrieval on the data sets of TREC-9, it is found that the obvious improvement for the retrieval performance can also be obtained. Yuejie Zhang, Lei Cen, Cheng Jin 0001, Xiangyang Xue 0001 |
IJCAI | 5 |
| 2011 | Multi-Kernel Multi-Label Learning with Max-Margin Concept Network
Wei Zhang 0016, Xiangyang Xue 0001, Jianping Fan 0001 |
IJCAI | 2 |
| 2011 | Level influence of spatial pyramid matching in object classificationabstractIn this paper we propose to effectively consider the shape and size variations for object classification. Specifically, a novel image matching method is proposed to incorporate the image segmentation with Spatial Pyramid Matching (SPM), and test our method on flower classification. A Level Influence Factor (LIF) is introduced to represent weights of different pyramid levels based on the statistical information of each segmented image. Then the images are classified based on the LIF weighted spatial pyramid bag-of-visual-words feature, and some levels with weight values zeros are not needed to be compared further. Also, in SPM matching stage, the block in one image is compared with not only its corresponding block in another image, but also the spatially neighboring blocks of the corresponding blocks to find the best match. This fuzzy matching method can incorporate some translation of objects. Experiments are performed on a flower dataset containing 1360 images from 17 different categories. And experimental results demonstrate that our proposed method has better time efficiency than traditional SPM and outperforms the state-of-art flower classification methods. Hong Lu 0001, Renzhong Wei, Yanran Shen, Xiangyang Xue 0001 |
ACM Multimedia | 4 |
| 2011 | Ensemble approach based on conditional random field for multi-label image and video annotationabstractMulti-label image/video annotation is a challenging task that allows to correlate more than one high-level semantic keyword with an image/video-clip. Previously, a single model is usually used for the annotation task, with relatively large variance in performance. The correlation among the annotation keywords should also be considered. In this paper, to reduce the performance variance and exploit the correlation between keywords, we propose the En-CRF (Ensemble based on Conditional Random Field) method. In this method, multiple models are first trained for each keyword, then the predictions of these models and the correlations between keywords are incorporated into a conditional random field. Experimental results on benchmark data set, including Corel5k and TRECVID 2005, show that the En-CRF method is superior or highly competitive to several state-of-the-art methods. Xin-Shun Xu, Yuan Jiang 0001, Xiangyang Xue 0001, Zhi-Hua Zhou |
ACM Multimedia | 4 |
| 2011 | Ensemble multi-instance multi-label learning approach for video annotation taskabstractAutomatic video annotation is an important ingredient for video indexing, browsing, and retrieval. Traditional studies represent one video clip with a flat feature vector; however, video data usually has natural structure. Moreover, a video clip is generally relevant to multiple concepts. Indeed, the video annotation task is inherently a Multi-Instance Multi-Label (MIML) learning problem. In this paper, we propose the En-MIMLSVM approach for the video annotation task. It considers the class imbalance and long time training problems of most video annotation tasks. In addition, a temporally consistent weighted multi-instance kernel is developed to take into account both the temporal consistency in video data and the significance of instances of different levels in pyramid representation. The En-MIMLSVM is evaluated on TRECVID 2005 data set, and the results show that it outperforms several state-of-the-art methods. Xin-Shun Xu, Xiangyang Xue 0001, Zhi-Hua Zhou |
ACM Multimedia | 2 |
| 2011 | Towards content-based audio fragment authenticationabstractAudio authentication is a technique to protect the integrity and originality of audio signals. Due to the long duration, it is often desirable to authenticate only a segment of audio signal. To the authors' knowledge, this important issue has not been seriously researched so far. In this paper, a novel authentication algorithm is proposed for the purpose of audio fragment authentication. SIFT descriptor originated from computer vision field is introduced and calculated on audio spectrogram to accomplish the tasks of fragment alignment, time-domain blocking, cropping and inserting identification etc. Experiments show reliable discrimination between admissible and malicious manipulations, precise tamper localization and classification. Xiangyang Xue 0001, Wei Li 0012 |
ACM Multimedia | 1 |
| 2011 | Automatic image annotation with weakly labeled datasetabstractIt is very attractive to exploit weakly-labeled image dataset for multi-label annotation applications. In our paper the meaning of the terminology weakly labeled is threefold: i) only a small subset of the available images are labeled; ii) even for the labeled image, the given labels may be uncorrect or incomplete; iii) the given labels do not provide the exact object locations in the images. A novel method is developed to predict the multiple labels for images and to provide region-level labels for the objects. We cluster the image regions to learn several region-exemplars and predict the label vector for each image region as a locally weighted average of the label vectors on exemplars. By investigating the label confidence matrix for the region-exemplars from different perspectives (column picture and row picture), we sufficiently leverage the visual contexts, the semantic contexts, and the consistency between similarities in the visual feature space and semantic label space. Experimental results on real web images demonstrate the effectiveness of the proposed method. Wei Zhang 0016, Yao Lu 0028, Xiangyang Xue 0001, Jianping Fan 0001 |
ACM Multimedia | 3 |
| 2011 | Refining local descriptors by embedding semantic information for visual categorizationabstractLocal descriptor extraction and vector quantization are the important components of widely-used Bag-of-Features (BoF) model for visual categorization. This paper proposes a simple and efficient approach to refine the local descriptors for vector quantization by embedding semantic information. The original local descriptors are integrated by a sequence of category-independent and category-dependent basis. Particularly, the category-dependent basis is learned by minimizing the joint loss minimization over local descriptors from different categories with a shared regularization penalty, which can be formulated as a linear programming problem. The transferred descriptors are further quantized and aggregated to the visual vocabulary. Experiments are performed on PASCAL VOC 2007 benchmark and the quantitative comparisons with several state-of-the-art approaches demonstrate the effectiveness of our proposed approach. Yingbin Zheng, Renzhong Wei, Hong Lu 0001, Xiangyang Xue 0001 |
ACM Multimedia | 4 |
| 2011 | A Hybrid Probabilistic Model for Unified Collaborative and Content-Based Image TaggingabstractThe increasing availability of large quantities of user contributed images with labels has provided opportunities to develop automatic tools to tag images to facilitate image search and retrieval. In this paper, we present a novel hybrid probabilistic model (HPM) which integrates low-level image features and high-level user provided tags to automatically tag images. For images without any tags, HPM predicts new tags based solely on the low-level image features. For images with user provided tags, HPM jointly exploits both the image features and the tags in a unified probabilistic framework to recommend additional tags to label the images. The HPM framework makes use of the tag-image association matrix (TIAM). However, since the number of images is usually very large and user-provided tags are diverse, TIAM is very sparse, thus making it difficult to reliably estimate tag-to-tag co-occurrence probabilities. We developed a collaborative filtering method based on nonnegative matrix factorization (NMF) for tackling this data sparsity issue. Also, an L1 norm kernel method is used to estimate the correlations between image features and semantic concepts. The effectiveness of the proposed approach has been evaluated using three databases containing 5,000 images with 371 tags, 31,695 images with 5,587 tags, and 269,648 images with 5,018 tags, respectively. William Kwok-Wai Cheung, Guoping Qiu, Xiangyang Xue 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2011 | Real-Time, Adaptive, and Locality-Based Graph Partitioning Method for Video Scene ClusteringabstractWe propose in this paper an efficient, adaptive, and locality-based graph partitioning method for video scene clustering. First, a graph partitioning method is proposed to group video shots into scenes, and a peer-group filtering (PGF) scheme is used to identify all the shots similar to each particular shot based on Fisher's discriminant analysis. To work with computable shot similarity measures that have only limited discriminating power, we develop a graph partitioning scheme to cluster the shots by maximizing the likeness of shots within the same cluster and minimizing that between different clusters. Second, considering that video data are normally obtained and viewed sequentially, we propose to perform a locality-based PGF and graph partitioning on video segments with 50 shots, 100 shots, and so on. This proposed locality-based method has the advantage that the number of scene clusters is not required to be known a priori, and it can achieve performance comparable to that processing on the whole video sequence. Experimental results are presented to demonstrate the effectiveness and efficiency of the proposed method. Hong Lu 0001, Yap-Peng Tan, Xiangyang Xue 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2010 | Transfer incremental learning for pattern classificationabstractTraditional machine learning methods, such as Support Vector Machines (SVMs), usually assume that training and test data share the same distributions. Due to the inherent dynamic data nature, it is often observed that (1) the volumes of the training data may gradually grow; and (2) the existing and the newly arrived samples may be subject to different distributions or learning tasks. In this paper, we propose a Transfer Incremental Support Vector Machine(TrISVM), with the objective of tackling changes in data volumes and learning tasks at the same time. By using new updating rules to calculate the inverse matrix, TrISVM solves the existing incremental learning problem more efficiently, especially for high dimensional data. Furthermore, when using new samples to update the existing models, TrISVM employs sample-based weight adjustment procedures to ensure that the concept transferring between auxiliary and target samples can be leveraged to fulfill the transfer learning goal. Experimental results on real-world data sets demonstrate that TrISVM achieves better efficiency and prediction accuracy than both incremental-learning and transfer-learning based methods. In addition, the results also show that TrISVM is able to achieve bidirectional knowledge transfer between two similar tasks. Zhenfeng Zhu, Xingquan Zhu 0001, Yue-Fei Guo, Xiangyang Xue 0001 |
CIKM | 4 |
| 2010 | Maximizing Growth Codes Utility in Large-Scale Wireless Sensor Networks
Xin Wang 0002, Jin Zhao 0001, Xiangyang Xue 0001 |
Euro-Par (2) | 4 |
| 2010 | Quick matting: A matting method based on pixel spread and propagationabstractThe problem of matting is always solved by finding the alpha value for each pixel in the image. Many recent methods combine color sampling and affinity definition in different steps, leading to large computational cost. In the proposed method, when the alpha value of a pixel Piis calculated, the pixel is regarded as a foreground pixel to help calculate its adjacent pixels' alpha values, resulted in a faster solution. This spreading way of traversal also ensures local continuity of foreground object and improves the visual result. Experiments show our Quick Matting can achieve comparable alpha mattes as Robust Matting, while the speed is enhanced by about 25 times. Yiyang Gu, Cheng Jin 0001, Xiangyang Xue 0001 |
ICIP | 3 |
| 2010 | SVD-SIFT for web near-duplicate image detectionabstractStable and high distinctive image features are the basis for web near-duplicate image detection. SIFT (scale invariant feature transform) not only has good scale and brightness invariance, also has a certain robustness to affine distortion, perspective change, and additive noise. However, to extract SIFT features to represent an image, hundreds or even thousands of SIFT key points need to be selected. And each key point needs to be described by using a 128-dimensional feature vector. Thus, the matching cost of detection method based on SIFT features is high. In this paper, we propose to apply the singular value decomposition (SVD) method for feature matching and extract the new features from the set of SIFT feature points. The extracted feature is termed as SVD-SIFT. Experimental results demonstrate that the method can obtain a better tradeoff between effectiveness and efficiency for detection. Hong Lu 0001, Xiangyang Xue 0001 |
ICIP | 3 |
| 2010 | How context helps: A discriminative codeword selection method for object detectionabstractWe first propose in this paper to localize objects in images based on the models learned from the weakly labeled images. This task is termed as region of interest (ROI) detection. Local features such as SIFT or HOG are extracted and the discriminative words from clustered codewords based on SIFT and HOG are selected to model the objects. Then how to find the discriminative words to model the object is important. Existing ROI detection methods consider the information from the foreground objects by selecting the words appearing more in the images belonging to one specific image class. Considering the information from background/context is also helpful for object detection and classification, we propose to select the discriminative words which appear more in the foreground/object and less in the background/context. Second, another task is to give the class label (object in this setting) for a given image and also give the position of the object appearing in the image. This task is termed as objection detection. A normal way for this task after ROI is to extract features from the detected regions and not from the whole image. Since the discriminative words extracted during ROI detection has good discriminative ability, we propose to use these words for object detection. Experimental results on PASCAL VOC 2006 dataset and a larger dataset containing 29 classes demonstrate the effectiveness of the proposed method. Renzhong Wei, Hong Lu 0001, Yingbin Zheng, Lei Cen, Cheng Jin 0001, Xiangyang Xue 0001, Weiguo Wu |
ICIP | 6 |
| 2010 | Robust hashing for music copyright protection by combining beat segmentation and chromaabstractTime-scale modification and pitching shifting are two recognized challenging attacks to music copyright protection. To resist them simultaneously, a novel robust hashing method is proposed by combining the strength of music beat segmentation and chroma-based music feature. These two measures are aimed at solving the problem of desynchronization and frequency shifting respectively. Moreover, two layers of scrambling are performed to ensure the security. Experiments exhibit remarkable robustness against various attacks including pitch [email protected]%, time-scale [email protected]%, and [email protected]/10 etc. Wei Li 0012, Zhurong Wang, Bilei Zhu, Xiangyang Xue 0001 |
ACM Multimedia | 4 |
| 2010 | Bilingual query translation and expansion for supporting more effective cross-language image retrievalabstractTo support more effective Cross-Language Image Retrieval (ImageCLIR), a novel algorithm is developed by integrating a bilingual semantic network to achieve more precise bilingual query translation and expansion. An English-Chinese bilingual parallel corpus is used to construct the bilingual semantic network for determining more meaningful text terms and characterizing the inter-term correlations and similarity contexts between multiple inter-related text terms more precisely. Our experiments on CWMT2009 and CLEF have provided very promising results. Yuejie Zhang, Lei Cen, Cheng Jin 0001, Xiangyang Xue 0001 |
ACM Multimedia | 4 |
| 2010 | Semantic video indexing by fusing explicit and implicit context spacesabstractThis paper addresses the problem of context-based concept fusion (CBCF) for concept detection and semantic video indexing. We introduce a novel framework based on constructing context spaces of concepts, such that the contextual correlations are used to improve the performance of concept detectors. Different from traditional CBCF approach, we present two kinds of such context spaces: explicit context space for modeling the correlation of pairwise concepts, and implicit context space for representing latent themes trained from a set of concepts. The final concept detection scores are then directly fused from explicit and implicit context spaces. Experiments are presented on TRECVid 2006 benchmark and the comparisons with several state-of-the-art approaches demonstrate the effectiveness of proposed framework. Yingbin Zheng, Renzhong Wei, Hong Lu 0001, Xiangyang Xue 0001 |
ACM Multimedia | 4 |
| 2010 | A novel audio fingerprinting method robust to time scale modification and pitch shiftingabstractA novel audio fingerprinting method that is highly robust to Time Scale Modification (TSM) and pitch shifting is proposed. Instead of simply employing spectral or tempo-related features, our system is based on computer-vision techniques. We transform each 1-D audio signal into a 2-D image and treat TSM and pitch shifting of the audio signal as stretch and translation of the corresponding image. Robust local descriptors are extracted from the image and matched against those of the reference audio signals. Experimental results show that our system is highly robust to various audio distortions, including the challenging TSM and pitch shifting. Bilei Zhu, Wei Li 0012, Zhurong Wang, Xiangyang Xue 0001 |
ACM Multimedia | 4 |
| 2010 | Achieving O(1) IP lookup on GPU-based software routersabstractIP address lookup is a challenging problem due to the increasing routing table size, and higher line rate. This paper investigates a new way to build an efficient IP lookup scheme using graphics processor units(GPU). Our contribution here is to design a basic architecture for high-performance IP lookup engine with GPU, and to develop efficient algorithms for routing prefix operations such as lookup, deletion, insertion, and modification. In particular, the IP lookup scheme can achieve O(1) time complexity. Our experimental results on real-world route traces show promising 6x gains in IP lookup throughput. Jin Zhao 0001, Xinya Zhang, Xin Wang 0002, Xiangyang Xue 0001 |
SIGCOMM | 4 |
| 2010 | Robust audio identification for MP3 popular musicabstractAudio identification via fingerprint has been an active research field with wide applications for years. Many technical papers were published and commercial software systems were also employed. However, most of these previously reported methods work on the raw audio format in spite of the fact that nowadays compressed format audio, especially MP3 music, has grown into the dominant way to store on personal computers and transmit on the Internet. It would be interesting if a compressed unknown audio fragment is able to be directly recognized from the database without the fussy and time-consuming decompression-identification-recompression procedure. So far, very few algorithms run directly in the compressed domain for music information retrieval, and most of them take advantage of MDCT coefficients or derived energy type of features. As a first attempt, we propose in this paper utilizing compressed-domain spectral entropy as the audio feature to implement a novel audio fingerprinting algorithm. The compressed songs stored in a music database and the possibly distorted compressed query excerpts are first partially decompressed to obtain the MDCT coefficients as the intermediate result. Then by grouping granules into longer blocks, remapping the MDCT coefficients into 192 new frequency lines to unify the frequency distribution of long and short windows, and defining 9 new subbands which cover the main frequency bandwidth of popular songs in accordance with the scale-factor bands of short windows, we calculate the spectral entropy of all consecutive blocks and come to the final fingerprint sequence by means of magnitude relationship modeling. Experiments show that such fingerprints exhibit strong robustness against various audio signal distortions like recompression, noise interference, echo addition, equalization, band-pass filtering, pitch shifting, and slight time-scale modification etc. For 5s-long query examples which might be severely degraded, an average top-five retrieval precision rate of more than 90% can be obtained in our test data set composed of 1822 popular songs. Wei Li 0012, Yaduo Liu, Xiangyang Xue 0001 |
SIGIR | 3 |
| 2010 | Robust music identification based on low-order zernike moment in the compressed domainabstractIn this paper, we devise a novel robust music identification algorithm utilizing compressed-domain audio Zernike moment adapted from image processing techniques as the pivotal feature. Audio fingerprint derived from this feature exhibits strong robustness against various audio signal distortions including the challenging pitch shifting and time-scale modification. Experiments show that in our test dataset composed of 1822 popular songs, a 5s music query example which might have been severely corrupted is still sufficient to identify its original near-duplicate copy, with more than 90% top five precision rate. Wei Li 0012, Yaduo Liu, Xiangyang Xue 0001 |
SIGIR | 3 |
| 2010 | Normalized dimensionality reduction using nonnegative matrix factorization
Zhenfeng Zhu, Yue-Fei Guo, Xingquan Zhu 0001, Xiangyang Xue 0001 |
Neurocomputing | 4 |
| 2010 | Semi-automatic dynamic auxiliary-tag-aided image annotation
Shile Zhang, Bin Li 0015, Xiangyang Xue 0001 |
Pattern Recognit. | 3 |
| 2010 | Constructions of cryptographically significant boolean functions using primitive polynomialsabstractIt is known that Boolean functions used in stream and block ciphers should have good cryptographic properties to resist algebraic attacks. Up until now, there have been several constructions of Boolean functions achieving optimum algebraic immunity. However, most of their nonlinearities are very low. Carlet and Feng studied a class of Boolean functions with optimum algebraic immunity and deduced the lower bound of its nonlinearity, which is good, but not very high. Moreover, the main practical problem with this construction is that it cannot be implemented efficiently. In this paper, we put forward a new method to construct cryptographically significant Boolean functions by using primitive polynomials, and construct three infinite classes of Boolean functions with good cryptographic properties: balancedness, optimum algebraic degree, optimum algebraic immunity, and a high nonlinearity. Qichun Wang, Jie Peng 0001, Haibin Kan, Xiangyang Xue 0001 |
IEEE Trans. Inf. Theory | 4 |
| 2009 | Incorporating Spatial Correlogram into Bag-of-Features Model for Scene Categorization
Yingbin Zheng, Hong Lu 0001, Cheng Jin 0001, Xiangyang Xue 0001 |
ACCV (1) | 4 |
| 2009 | Transfer learning for collaborative filtering via a rating-matrix generative modelabstractCross-domain collaborative filtering solves the sparsity problem by transferring rating knowledge across multiple domains. In this paper, we propose a rating-matrix generative model (RMGM) for effective cross-domain collaborative filtering. We first show that the relatedness across multiple rating matrices can be established by finding a shared implicit cluster-level rating matrix, which is next extended to a cluster-level rating model. Consequently, a rating matrix of any related task can be viewed as drawing a set of users and items from a user-item joint mixture model as well as drawing the corresponding ratings from the cluster-level rating model. The combination of these two models gives the RMGM, which can be used to fill the missing ratings for both existing and new users. A major advantage of RMGM is that it can share the knowledge by pooling the rating data from multiple tasks even when the users and items of these tasks do not overlap. We evaluate the RMGM empirically on three real-world collaborative filtering data sets to show that RMGM can outperform the individual models trained separately. Bin Li 0015, Qiang Yang 0001, Xiangyang Xue 0001 |
ICML | 3 |
| 2009 | Can Movies and Books Collaborate? Cross-Domain Collaborative Filtering for Sparsity Reduction
Bin Li 0015, Qiang Yang 0001, Xiangyang Xue 0001 |
IJCAI | 3 |
| 2009 | Matrix-based Kernel Principal Component analysis for large-scale data setabstractKernel Principal Component Analysis (KPCA) is a nonlinear feature extraction approach, which generally needs to eigen-decompose the kernel matrix. But the size of kernel matrix scales with the number of data points, it is infeasible to store and compute the kernel matrix when faced with the large-scale data set. To overcome computational and storage problem for large-scale data set, a new framework, Matrixbased Kernel Principal Component Analysis (M-KPCA), is proposed. By dividing the large scale data set into small subsets, we could treat the autocorrelation matrix of each subset as the special computational unit. A novel polynomial-matrix kernel function is adopted to compute the similarity between the data matrices in place of vectors. It is also proved that the polynomial kernel is the extreme case of the polynomial-matrix one. The proposed M-KPCA can greatly reduce the size of kernel matrix, which makes its computation possible. The effectiveness is demonstrated by the experimental results on the artificial and real data set. Weiya Shi, Yue-Fei Guo, Xiangyang Xue 0001 |
IJCNN | 3 |
| 2009 | Temporal context as cortical spatial codesabstractIt is largely unknown how the brain deals with time. The new field of research on autonomous development must enable machines to develop intelligent behaviors that respond not only to spatial features, but also temporal features. Hidden Markov Model (HMM) has a probability based mechanism to deal with time warping, but no effective online method exists that can deal with general temporal structure and temporal abstraction. By online, we mean that the agent must respond to spatial and temporal context immediately while the sensory stream flows in. By general temporal context, we mean various desirable temporal subsets, such as deletion (e.g., stop words) and variable temporal lengths (e.g., beyond bigrams and trigrams). By temporal abstraction, we mean using abstract meaning of context, instead of concrete forms. This paper proposes a brain inspired online scheme for making sequential decisions based on general temporal context. By sequential decisions, the action from the network depends on not only inputs and outputs but also emergent internal context states. In our neuromorphic scheme, the internal states are not predefined symbols, but distributed context depending on the internal attention. Our complexity analysis shows how this scheme greatly reduces the exponential time complexity O(2t) of all the possible number of contexts of length t down to linear time complexity O(cnt), where n is the number of neurons in the network and c is the average number of synapses of each neuron. In this paper, we concentrate on processing sequential text inputs by an online agent network under motor-supervised learning. Juyang Weng, Mingmin Chi, Xiangyang Xue 0001 |
IJCNN | 4 |
| 2009 | Tree-structured data regeneration with network coding in distributed storage systemsabstractDistributed storage systems, built on peer-to-peer networks, can provide large-scale data storage and high data reliability by redundant schemes, such as replica, erasure codes and linear network coding. Redundant data may get lost due to the instability of distributed systems, such as permanent node departures, hardware failures, and accidental deletions. In order to maintain data availability, it is necessary to regenerate new redundant data in another node, referred to as a newcomer. Regeneration is expected to be finished as soon as possible, because the regeneration time can influence the data reliability and availability of distributed storage systems. It has been acknowledged that linear network coding can regenerate redundant data with less network traffic than replica and erasure codes. However, previous regeneration schemes are all star-structured regeneration schemes, in which data are transferred directly from existing storage nodes, referred to as providers, to the newcomer, so the regeneration time is always limited by the path with the narrowest bandwidth between newcomer and provider, due to bandwidth heterogeneity. In this paper, we exploit the bandwidth between providers and propose a tree-structured regeneration scheme using linear network coding. In our scheme, data can be transferred from providers to the newcomer through a regeneration tree, defined as a spanning tree covering the newcomer and all the providers. In a regeneration tree, a provider can receive data from other providers, then encode the received data with the data this provider stores, and finally send the encoded data to another provider or to the newcomer. We prove that a maximum spanning tree is an optimal regeneration tree and analyze its performance. In a trace-based simulation, the results show the tree-structured scheme can reduce the regeneration time by 75%-82% and improve data availability by 73%-124%. Jun Li 0017, Xin Wang 0002, Xiangyang Xue 0001, Baochun Li |
IWQoS | 4 |
| 2009 | Web image retrieval reranking with multi-view clusteringabstractGeneral image retrieval is often carried out by a text-based search engine, such as Google Image Search. In this case, natural language queries are used as input to the search engine. Usually, the user queries are quite ambiguous and the returned results are not well-organized as the ranking often done by the popularity of an image. In order to address these problems, we propose to use both textual and visual contents of retrieved images to reRank web retrieved results. In particular, a machine learning technique, a multi-view clustering algorithm is proposed to reorganize the original results provided by the text-based search engine. Preliminary results validate the effectiveness of the proposed framework. Mingmin Chi, Peiwu Zhang, Yingbin Zhao, Rui Feng 0001, Xiangyang Xue 0001 |
WWW | 5 |
| 2008 | Scene segmentation based on video structure and spectral methodsabstractScene is an important semantic unit for video analysis, retrieval and browsing. However, due to the lack of a generic algorithm, many studies focus on specific methods for certain video genes, e.g., news, sports, etc. In this paper, we propose a general framework for scene segmentation. First, we construct a graph, in which the elements encode the shot-to-shot coherent characteristics of a video clip based on visual similarity and temporal relation between shots. In this step, we only exploit the inherent property of video itself and it is independent of video genres. Second, spectral method is applied on the graph to group shots into scenes. The proposed method is not only simple but also effective to deal with organized features. Experimental results validate the robustness of our method on different kinds of videos. Bin Li 0015, Hong Lu 0001, Xiangyang Xue 0001 |
ICARCV | 4 |
| 2008 | Swifter: Chunked Network Coding for Peer-to-Peer Content DistributionabstractThe benefit of network coding with respect to simplifying scheduling overhead for content distribution has been extensively studied in previous literature. However, the complexity of network coding increases as the content size scales up. In this paper, we study the tradeoff between scheduling overhead and coding overhead. To this end, we propose Swifter, a P2P content distribution scheme, which employs local-rarest-first segment scheduling and chunked network coding algorithms. In Swifter, content is divided into segments, which are further divided into blocks. Each peer schedules a local-rarest segment request from its neighbors. Network coding is then used for generating a reply block within the requested segment. Leveraging our real-world implementation and experiments, we find that Swifter has low coding overhead and can reduce average download time by up to 40% compared to existing work. Jinbiao Xu, Jin Zhao 0001, Xin Wang 0002, Xiangyang Xue 0001 |
ICC | 4 |
| 2008 | CODED IP: On the Feasibility of IP-Layer Network CodingabstractNowadays, the real practice of network coding in wireline networks is focused on the P2P overlay networks. Although it can help to utilize network resources more efficiently, P2P network coding does not exhibit benefits in terms of the maximum throughput. Our work aims to implement network coding at the IP layer, which is an idea not fundamentally new, but with little real practice because of the enormous difficulties involved. In this paper we propose CODED IP, a protocol framework that plugs network coding into the current IP stack. Experiments on a 22-node testbed show that CODED IP provides multicast traffic with not only a significantly higher throughput than overlay network coding and naive IP multicast, but also a more balanced load distribution as compared with overlay network coding. Xunrui Yin, Xin Wang 0002, Jin Zhao 0001, Xiangyang Xue 0001 |
ICCCN | 5 |
| 2008 | An Improved Generalized Discriminant Analysis for Large-Scale Data SetabstractIn order to overcome the computation and storage problem for large-scale data set, an efficient iterative method of generalized discriminant analysis is proposed. Because sample vectors cannot explicitly be denoted in kernel space, some mathematical tricks are firstly used to transform the kernel matrix. Then, the columns of transformed matrix are used for iterative algorithm to extract nonlinear discriminant vectors. The proposed method reduces space complexity from O(m2) to O(m) and its effectiveness is validated from experimental results. Weiya Shi, Yue-Fei Guo, Cheng Jin 0001, Xiangyang Xue 0001 |
ICMLA | 4 |
| 2008 | Collaborative and content-based image labelingabstractMany on-line photo sharing systems allow users to tag their images so as to support semantic image search. In this paper, we study how one can take advantages of the already-tagged images to (semi-)automate the labeling of newly uploaded ones. In particular, we propose a hybrid approach for the prediction where user-provided tags and image visual contents are fused under a unified probabilistic framework. Kernel smoothing and collaborative filtering techniques are explored for improving the accuracy of the probabilistic models estimation. By comparing with some state-of-the-art content-based image labeling methods, we have empirically shown that 1) the proposed method can achieve comparable tag prediction accuracy when there is no user-provided tag, and that 2) it can significantly boost the prediction accuracy if the user can provide just a few tags. William Kwok-Wai Cheung, Xiangyang Xue 0001, Guoping Qiu |
ICPR | 3 |
| 2008 | Multilayer in-place learning networks for modeling functional layers in the laminar cortex
Juyang Weng, Tianyu Luwang, Hong Lu 0001, Xiangyang Xue 0001 |
Neural Networks | 4 |
| 2008 | Metric learning by discriminant neighborhood embedding
Wei Zhang 0016, Xiangyang Xue 0001, Zichen Sun, Hong Lu 0001, Yue-Fei Guo |
Pattern Recognit. | 2 |
| 2008 | Incorporating feature hierarchy and boosting to achieve more effective classifier training and concept-oriented video summarization and skimmingabstractFor online medical education purposes, we have developed a novel scheme to incorporate the results of semantic video classification to select the most representative video shots for generating concept-oriented summarization and skimming ofsurgery education videos. First, salient objects are used as the video patterns for feature extraction to achieve a good representation of the intermediate video semantics. The salient objects are defined as the salient video compounds that can be used to characterize the most significant perceptual properties of the corresponding real world physical objects in a video, and thus the appearances of such salient objects can be used to predict the appearances of the relevant semantic video concepts in a specific video domain. Second, a novelmulti-modal boostingalgorithm is developed to achieve more reliable video classifier training by incorporating feature hierarchy and boosting to dramatically reduce both the training cost and the size of training samples, thus it can significantly speed up SVM (support vector machine) classifier training. In addition, the unlabeled samples are integrated to reduce the human efforts on labeling large amount of training samples. Finally, the results of semantic video classification are incorporated to enable concept-oriented video summarization and skimming. Experimental results in a specific domain ofsurgery education videosare provided. Hangzai Luo, Yuli Gao, Xiangyang Xue 0001, Jinye Peng 0001, Jianping Fan 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2007 | Salient Object Detection on Large-Scale Video DataabstractRecently more and more researches focus on the concept extraction from unstructured video data. To bridge the semantic gap between the low-level features and the high-level video concepts, a mid-level understanding of the video contents, i.e., salient object is detected based on the techniques of image segmentation and machine learning. Specifically, 21 salient object detectors are developed and tested on TRECVID 2005 development video corpus. In addition, a boosting method is proposed to select the most representative features to achieve a higher performance than only using single modality, and lower complexity than taking all features into account. Shile Zhang, Jianping Fan 0001, Hong Lu 0001, Xiangyang Xue 0001 |
CVPR | 4 |
| 2007 | Efficient Feature Extraction for Image ClassificationabstractIn many image classification applications, input feature space is often high-dimensional and dimensionality reduction is necessary to alleviate the curse of dimensionality or to reduce the cost of computation. In this paper, we extract discriminant features for image classification by learning a low-dimensional embedding from finite labeled samples. In the new feature space, intra-class compactness and extra-class separability are achieved simultaneously. Target dimensionality of the embedding is selected by spectral analysis. Our method is designed suitable for data with both uni- and multi-modal class distributions. We also develop its two-dimensional variant which makes use of the matrix representation of images. Experimental results on three real image datasets demonstrate the efficacy of our method compared to the state of the art. Wei Zhang 0016, Xiangyang Xue 0001, Zichen Sun, Yue-Fei Guo, Mingmin Chi, Hong Lu 0001 |
ICCV | 2 |
| 2007 | Support cluster machineabstractFor large-scale classification problems, the training samples can be clustered beforehand as a downsampling pre-process, and then only the obtained clusters are used for training. Motivated by such assumption, we proposed a classification algorithm, Support Cluster Machine (SCM), within the learning framework introduced by Vapnik. For the SCM, a compatible kernel is adopted such that a similarity measure can be handled not only between clusters in the training phase but also between a cluster and a vector in the testing phase. We also proved that the SCM is a general extension of the SVM with the RBF kernel. The experimental results confirm that the SCM is very effective for largescale classification problems due to significantly reduced computational costs for both training and testing and comparable classification accuracies. As a by-product, it provides a promising approach to dealing with privacy-preserving data mining problems. Bin Li 0015, Mingmin Chi, Jianping Fan 0001, Xiangyang Xue 0001 |
ICML | 4 |
| 2007 | Optimal dimensionality of metric space for classificationabstractIn many real-world applications, Euclidean distance in the original space is not good due to the curse of dimensionality. In this paper, we propose a new method, called Discriminant Neighborhood Embedding (DNE), to learn an appropriate metric space for classification given finite training samples. We define a discriminant adjacent matrix in favor of classification task, i.e., neighboring samples in the same class are squeezed but those in different classes are separated as far as possible. The optimal dimensionality of the metric space can be estimated by spectral analysis in the proposed method, which is of great significance for high-dimensional patterns. Experiments with various datasets demonstrate the effectiveness of our method. Wei Zhang 0016, Xiangyang Xue 0001, Zichen Sun, Yue-Fei Guo, Hong Lu 0001 |
ICML | 2 |
| 2007 | The Multilayer In-Place Learning Network for the Development of General Invariances and Multi-Task LearningabstractCurrently, there is a lack of general-purpose in-place learning engines that incrementally learn multiple tasks, to develop "soft" multi-task-shared invariances in the intermediate internal representation while a developmental robot interacts with its environment. Computationally, biologically inspired in-place learning provides unusually efficient learning algorithms whose simplicity, low computational complexity, and generality are set apart from typical conventional learning algorithms. We present in this paper the multiple-layer in-place learning network (MILN) for this ambitious goal. As a key requirement for autonomous mental development, the network enables both unsupervised and supervised learning to occur concurrently, depending on whether motor supervision signals are available or not at the motor end (the last layer) during the agent's interactions with the environment. We present principles based on which MILN automatically develops invariant neurons in different layers and why such invariant neuronal clusters are important for learning later tasks in open-ended development. Juyang Weng, Tianyu Luwang, Hong Lu 0001, Xiangyang Xue 0001 |
IJCNN | 4 |
| 2007 | Incremetal Spatio-Temporal Feature Extraction and Retrieval for Large Video DatabaseabstractIn this paper we present a novel framework for semantic retrieval of video database. Each frame of video clips, characterized by its HSV (hue-saturation-value) color feature, is first projected onto the spatial principle components via CCIPCA (candid covariance-free incremental principal component analysis). Temporal Chebyshev polynomials for video clips of various lengths are captured subsequently. The similarity of two video clips is finally presented in a reasonable and computable form. The framework works incrementally and is suitable for videos of data streams in sequential order. Extensive experiments demonstrate that the framework can obtain promising results on video similarity comparison, and also with a comparably computational speedup. Bo Geng, Hong Lu 0001, Xiangyang Xue 0001 |
ISCAS | 3 |
| 2007 | A robust incremental learning framework for accurate skin region segmentation in color images
Bin Li 0015, Xiangyang Xue 0001, Jianping Fan 0001 |
Pattern Recognit. | 2 |
| 2006 | An Efficient Early Termination Algorithm of Intra Prediction for H.264/AVCabstractWe propose in this paper an efficient early termination algorithm of intra prediction for H.264/AVC. It uses the spatial correlation after 16times16 inter prediction to make the judgement on whether to discard intra prediction or not. Experimental results demonstrate that the proposed algorithm can save the encoding time of H.264/AVC (JM98) between 25-45% with negligible degradation in the quality Hong Lu 0001, Xiangyang Xue 0001, Yap-Peng Tan |
ICARCV | 3 |
| 2006 | Quotient Set-based Nonlinear Manifold for Image RestorationabstractIn this paper we propose a patch-wise coarse-to-fine algorithm for image restoration using the manifold way of visual perception. All undistorted image patches are supposed to lie on a quotient set-based nonlinear manifold, and restoration of each degraded image patch can be implemented by projecting it to a locally linear region of such nonlinear manifold. The details of the original image can be learned from the undistorted training samples. Moreover, there is no need for us to assume that the degradation function is linear or to estimate some parameters of the blurs and noises beforehand. Experimental results demonstrate the effectiveness of the proposed method Wei Zhang 0016, Xiangyang Xue 0001, Hong Lu 0001, Yue-Fei Guo |
ICARCV | 3 |
| 2006 | In-Place Learning for Positional and Scale InvarianceabstractIn-place learning is a biologically inspired concept, meaning that the computational network is responsible for its own learning. With in-place learning, there is no need for a separate learning network. We present in this paper a multiple-layer in-place learning network (MILN) for learning positional and scale invariance. The network enables both unsupervised and supervised learning to occur concurrently. When supervision is available (e.g., from the environment during autonomous development), the network performs supervised learning through its multiple layers. When supervision is not available, the network practices while using its own practice motor signal as self-supervision (i.e., unsupervised per classical definition). We present principles based on which MILN automatically develops positional and scale invariant neurons in different layers. From sequentially sensed video streams, the proposed in-place learning algorithm develops a hierarchy of network representations. The global invariance was achieved through multi-layer quasi-invariances, with increasing invariance from early layers to the later layers. Experimental results are presented to show the effects of the principles. Juyang Weng, Hong Lu 0001, Tianyu Luwang, Xiangyang Xue 0001 |
IJCNN | 4 |
| 2006 | Automatic image annotation by incorporating feature hierarchy and boosting to scale up SVM classifiersabstractThe performance of image classifiers largely depends on two inter-related issues:(1)suitable frameworks for image content representation and automatic feature extraction;(2) effective algorithms for image classifier training and feature subset selection. To address the first issue, a multiresolution grid-based framework is proposed for image content representation and feature extraction to bypass the time-consuming and erroneous process for image segmentation. To address the second issue, a hierarchical boosting algorithm is proposed by incorporating feature hierarchy and boosting to scale up SVM image classifier training in high-dimensional feature space. The high-dimensional multi-modal heterogeneous visual features are partitioned into multiple low-dimensional single-modal homogeneous feature subsets and each of them characterizes certain visual property of images. For each homogeneous feature subset, principal component analysis (PCA)is performed to exploit the feature correlations and a weak classifier is learned simultaneously. After the weak classifiers for different feature subsets and grid sizes are available, they are combined to boost an optimal classifier for the given object class or image concept, and the most representative feature subsets and grid sizes are selected. Our experiments on a specific domain of natural images have obtained very positive results. Yuli Gao, Jianping Fan 0001, Xiangyang Xue 0001, Ramesh Jain 0001 |
ACM Multimedia | 3 |
| 2006 | Null Foley-Sammon transform
Yue-Fei Guo, Lide Wu, Hong Lu 0001, Zhe Feng 0001, Xiangyang Xue 0001 |
Pattern Recognit. | 5 |
| 2006 | Discriminant neighborhood embedding for classification
Wei Zhang 0016, Xiangyang Xue 0001, Hong Lu 0001, Yue-Fei Guo |
Pattern Recognit. | 2 |
| 2006 | Hierarchical Indexing Structure for Efficient Similarity Search in Video RetrievalabstractWith the rapid increase in both centralized video archives and distributed WWW video resources, content-based video retrieval is gaining its importance. To support such applications efficiently, content-based video indexing must be addressed. Typically, each video is represented by a sequence of frames. Due to the high dimensionality of frame representation and the large number of frames, video indexing introduces an additional degree of complexity. In this paper, we address the problem of content-based video indexing and propose an efficient solution, called the Ordered VA-File (OVA-File) based on the VA-file. OVA-File is a hierarchical structure and has two novel features: 1) partitioning the whole file into slices such that only a small number of slices are accessed and checked during k Nearest Neighbor (kNN) search and 2) efficient handling of insertions of new vectors into the OVA-File, such that the average distance between the new vectors and those approximations near that position is minimized. To facilitate a search, we present an efficient approximate kNN algorithm named Ordered VA-LOW (OVA-LOW) based on the proposed OVA-File. OVA-LOW first chooses possible OVA-Slices by ranking the distances between their corresponding centers and the query vector, and then visits all approximations in the selected OVA-Slices to work out approximate kNN. The number of possible OVA-Slices is controlled by a user-defined parameter \delta. By adjusting \delta, OVA-LOW provides a trade-off between the query cost and the result quality. Query by video clip consisting of multiple frames is also discussed. Extensive experimental studies using real video data sets were conducted and the results showed that our methods can yield a significant speed-up over an existing VA-file-based method and iDistance with high query result quality. Furthermore, by incorporating temporal correlation of video content, our methods achieved much more efficient performance. Hong Lu 0001, Beng Chin Ooi, Heng Tao Shen, Xiangyang Xue 0001 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2006 | Localized audio watermarking technique robust against time-scale modificationabstractSynchronization attacks like random cropping and time-scale modification are very challenging problems to audio watermarking techniques. To combat these attacks, a novel content-dependent localized robust audio watermarking scheme is proposed. The basic idea is to first select steady high-energy local regions that represent music edges like note attacks, transitions or drum sounds by using different methods, then embed the watermark in these regions. Such regions are of great importance to the understanding of music and will not be changed much for maintaining high auditory quality. In this way, the embedded watermark has the potential to escape all kinds of distortions. Experimental results show strong robustness against common audio signal processing, time-domain synchronization attacks, and most distortions introduced in Stirmark for Audio. Wei Li 0012, Xiangyang Xue 0001, Peizhong Lu |
IEEE Trans. Multim. | 2 |
| 2005 | Loop-based topology maintenance and route discovery for wireless sensor networksabstractClustering is an effective topology maintenance technology to provide better scalability for WSN. When loop structures are deployed to clustering, it can provide better robustness and convenient routing, because the topology information can be disseminated within a loop instead of a single clusterhead and the nature of the loop that there are two paths between each pair of nodes within a loop provides backup routes. In this paper, we propose a loop-based approach to combine clustering and routing for cost saving, which consists of setup procedure, regular procedure and recovery procedure. For the setup procedure, we propose a loop discovery algorithm, which is a fully distributed, on-demand and configurable algorithm achieving route discovery and clustering at the same time. The algorithm prefers a low-mobility network which is typically suitable for wireless sensor networks. Xin Wang 0002, Florian Baueregger, Xiangyang Xue 0001, Chai-Keong Toh |
GLOBECOM | 4 |
| 2005 | Spectral Images and Features Co-Clustering with Application to Content-based Image RetrievalabstractIn this paper, we present a spectral graph partitioning method for the co-clustering of images and features. We present experimental results, which show that spectral co-clustering has computational advantages over traditional k-means algorithm, especially when the dimensionalities of feature vectors are high. In the context of image clustering, we also show that spectral co-clustering gives better performances. We advocate that the images and features co-clustering framework offers new opportunities for developing advanced image database management technology and illustrate a possible scheme for exploiting the co-clustering results for developing a novel content-based image retrieval method Jian Guan 0004, Guoping Qiu, Xiangyang Xue 0001 |
MMSP | 3 |
| 2005 | Region-based Pornographic Image DetectionabstractRecent advantages in digital images and networks have made pornographic images more accessible than ever. Methods of detecting pornographic images have been proposed in some existing work. However, the skin detection modules adopted are all pixel-based and skin region shape features are rarely properly used. This paper presents a framework for pornographic image detection based on skin region information. Different from traditional works, our approach extracts color and texture features from arbitrary-shaped segmented regions. Then Gaussian mixture models are built for skin and non-skin region classification, and the skin map is produced based on the classification result. Finally, eigenregion features are used to describe the layout of skin regions on the whole image and pornographic images are detected according to the skin modality Bin Li 0015, Xiangyang Xue 0001, Hong Lu 0001 |
MMSP | 3 |
| 2005 | Efficient Video Clip Retrieval Using Index StructureabstractRetrieving similar video clips from large video database requires high query efficiency, precision and recall, which remains a challenging problem since the traditional query algorithms are inefficient and time-consuming. In this paper, we adopt the high-dimensional index structure vector-approximation file (VA-file) to organize the video database, and propose a new similarity measure which takes the temporal order among the video representations into account to improve the accuracy of query. Based on the VA-file and similarity measure, a new video clip retrieval algorithm is proposed in our method to achieve high query efficiency by using restricted sliding window to construct candidate video clips. Experimental results show that the proposed video retrieval method is efficient and effective Linjun Yang, Hong Lu 0001, Xiangyang Xue 0001, Yap-Peng Tan |
MMSP | 4 |
| 2005 | VoD Service Model and Performance Evaluation on the China's High Performance Broadband Information Network (3Tnet)abstractProviding large-scale true video-on-demand (VoD) service is a challenging problem over the Internet. Some characteristics of the Internet, such as packet switching, best effort mechanism, make it difficult to provide VoD services with enough Quality of Service (QoS) assurance. In this paper, we propose a VoD service model that uses a content delivery network approach on backbone and a peer-to-peer (p2p) approach on access over the China’s high performance broadband information network (named 3Tnet). We discuss the feasibility of our approach and evaluate the performance by using simulation. The results show that our proposed service model can support true VoD service on the 3Tnet in a scalable way. Xin Wang 0002, Xiangyang Xue 0001 |
PDCAT | 3 |
| 2005 | Simulating Large-Scale Traffic Aggregation in an Automatic Switched Optical NetworkabstractAutomatic Switched Optical Network (ASON), standardized by ITU-T, is emerging as a major technology choice for inter-operational and intervendor optical backbone transportation. With the ability to flexibly and automatically establish and maintain connection, ASON promises an IP-traffic tolerant, end-to-end QoS guaranteed transportation mechanism within which both bandwidth-consuming stream media traffic and traditional web traffic can be relayed smoothly. In this paper, we look into the traffic aggregation effect of TCP traffic in an optical switched network environment modeled under the 3TNet network architecture. We provide a simplified simulation model, which shows how large-scale TCP traffic aggregation can be leveraged by ASON switching ability and meanwhile proves the superiority of this architecture. Xin Wang 0002, Xiangyang Xue 0001 |
PDCAT | 3 |
| 2005 | A Novel Peer-to-Peer Intrusion Detection SystemabstractMANETs are composed of mobile nodes without any infrastructure and nodes cooperate to set up routes for network communications. Because of these characters, MANETs operate in open medium, so they are particularly vulnerable to intrusions. In this paper, we present an efficient intrusion detection system called MAPIDS (Mobile Agent-based Peer-to-peer Intrusion Detection System). On detecting a suspicious activity, MAPIDS initiates a voting approach to make a collective decision and take further action. In contrast to other intrusion detection system based on collective decision, MAPIDS saves more bandwidth and energy and it is immune to sparse nodes problem. Ji Zheng, Xin Wang 0002, Xiangyang Xue 0001 |
PDCAT | 4 |
| 2005 | InsightVideo: toward hierarchical video content organization for efficient browsing, summarization and retrievalabstractHierarchical video browsing and feature-based video retrieval are two standard methods for accessing video content. Very little research, however, has addressed the benefits of integrating these two methods for more effective and efficient video content access. In this paper, we introduce InsightVideo, a video analysis and retrieval system, which joins video content hierarchy, hierarchical browsing and retrieval for efficient video access. We propose several video processing techniques to organize the content hierarchy of the video. We first apply a camera motion classification and key-frame extraction strategy that operates in the compressed domain to extract video features. Then, shot grouping, scene detection and pairwise scene clustering strategies are applied to construct the video content hierarchy. We introduce a video similarity evaluation scheme at different levels (key-frame, shot, group, scene, and video.) By integrating the video content hierarchy and the video similarity evaluation scheme, hierarchical video browsing and retrieval are seamlessly integrated for efficient content access. We construct a progressive video retrieval scheme to refine user queries through the interactions of browsing and retrieval. Experimental results and comparisons of camera motion classification, key-frame extraction, scene detection, and video retrieval are presented to validate the effectiveness and efficiency of the proposed algorithms and the performance of the system. Xingquan Zhu 0001, Ahmed K. Elmagarmid, Xiangyang Xue 0001, Lide Wu, Ann Christine Catlin |
IEEE Trans. Multim. | 3 |
| 2004 | Improved Robust Watermarking in DCT Domain for Color ImagesabstractIn the field of color images watermarking, many methods are accomplished by marking the image luminance, or by processing each color channel separately. This paper proposes a new DCT domain watermarking expressly devised for RGB color images based on the diversity technique in communication system. The watermark is hidden within the data in the same sequence by modifying a subset of block DCT coefficients of each color channel. Detection is based on a combination method by taking into account the information conveyed by three color channels. Even if a particular channel is severely faded, we may still be able to recover a reliable estimated of transmitted watermark through other propagation channel. Experimental results, as well as theoretical analysis, are presented to demonstrate the validity of the new approach with respect to algorithm operating on image luminance only. Xiaoqiang Li 0002, Xiangyang Xue 0001 |
AINA (1) | 2 |
| 2004 | Efficient identification of speakers in news video based on shot segmentationabstractAn effective method for speaker identification in news video is presented in this paper, which is based on shot segmentation and exploits both audio and visual cues. Firstly, audio is segmented by shot segmentation based on the observation that there is only one speaker in a shot of news video in most cases. Furthermore, speech/non-speech discrimination is implemented on each shot. Finally, text-independent speaker identification is proposed using audio features on the discriminated speech shots. Experimental results show that our algorithm can obtain satisfactory performance in identifying speakers, so it can be used in real application. Xiangyang Xue 0001, Hong Lu 0001, You-san Nie |
ICARCV | 2 |
| 2004 | Improved shot boundary detection method based on text edgesabstractShot boundary detection is a pre-requisite technique for video indexing and retrieval. To avoid the influence of flashlight on abrupt shot detection, many edge-based techniques are studied thoroughly. However, these techniques are still susceptible to miss and mistake detecting the abrupt changes. Our observation shows that one of the reasons for these errors is the existence of superimposed text which has rich edges and is ever presented in video frames. To provide a solution, we present a novel method that utilizes the edge type, text edge (edge in text area) or non-text-edge (edge in other text area), reducing erroneous detection with the appearance of video text. Compared to other edge-based detection techniques, experimental results show that our proposed method achieves preferable performance. Liuhong Liang, Yang Liu 0246, Xiangyang Xue 0001, Hong Lu 0001, Yap-Peng Tan |
ICARCV | 3 |
| 2004 | Effective video text detection using line featuresabstractText superimposed on video frames provides synoptic or supplemental information on video semantics. In this paper, we propose a novel method to detect superimposed text effectively. First, we detect edges by an improved Canny edge detector. Then, a line-feature vector graph is generated based on the edge map and the stroke information is extracted. Finally text regions are generated and filtered according to line features. Experimental results show that, without much increasing the computational cost, our proposed method could suppress the false alarms notably. Furthermore, our method can be easily customized to applications with different tradeoffs in recall and precision. Yang Liu 0246, Hong Lu 0001, Xiangyang Xue 0001, Yap-Peng Tan |
ICARCV | 3 |
| 2003 | A Mobile Multicast Algorithm Using Agents for Mobile Ad-hoc UsersabstractWe describe an improved agent-based multicast algorithm for mobile ad-hoc users. Mobile multicast agents (MMAs) form a virtual backbone of an ad-hoc network and they provide multicast tree discovery and multicast tree maintenance. However, they do not complete datagram delivery. In order to ensure reliable multicast, MMAs are also in charge of datagram retransmission. The improved mobile multicast algorithm can simplify the multicast tree discovery, reduce control overhead of the network, and increase the total network throughput, in comparison with general AODV multicast operation. At the same time, it can also save battery power and transmission capabilities of those MMAs since they are only used to transfer control information and retransmitted packets. Hence we avoid draining the resources of some mobile nodes (MMAs), which is unfair. Fei Li 0005, Xin Wang 0002, Xiangyang Xue 0001 |
AINA | 3 |
| 2003 | An Improved Dynamic Priority Queue for Multimedia Network CommunicationsabstractWe propose an optimizing queue scheme for multimedia network communications with dynamic priority. Arriving users are classified into two types based on their transmission contents, and they are served by different priority degrees. The high priority user (e.g., control messages, video, voice) has higher access possibility than the low priority user (e.g., data). This is realized by adding a dynamic priority to the priority class, instead of a constant priority. After SOR (successive over-relaxation) simulation the results illustrate that the improved dynamic priority queue is efficient to improve the performance of the total system. Xin Wang 0002, Fei Li 0005, Xiangyang Xue 0001 |
AINA | 3 |
| 2003 | An Optimized Multi-bits Blind Watermarking Scheme
Xiaoqiang Li 0002, Xiangyang Xue 0001, Wei Li 0012 |
ICICS | 2 |
| 2003 | Audio Watermarking Based on Music Content Analysis: Robust against Time Scale Modification
Wei Li 0012, Xiangyang Xue 0001 |
IWDW | 2 |