VLDB 2026 Research / reviewers in the wild / expert
Wei Zhang 0016
dblp:10/4661-16
· DBLP profile ↗
55ranked-venue papers
10as first author
29since 2021 · last 2026
0000-0002-2358-8543ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 41 · 7 first-author · 25 since 2021Artificial intelligence and machine learning · 33 · 9 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 3Security and privacy · 1Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Seeing Is Believing: Rich-Context Hallucination Detection for MLLMs via Backward Visual GroundingabstractMultimodal Large Language Models (MLLMs) have unlocked powerful cross-modal capabilities, but still significantly suffer from hallucinations. As such, accurate detection of hallucinations in MLLMs is imperative for ensuring their reliability in practical applications. To this end, guided by the principle of “Seeing is Believing”, we introduce VBackChecker, a novel reference-free hallucination detection framework that verifies the consistency of MLLM-generated responses with visual inputs, by leveraging a pixel-level Grounding LLM equipped with reasoning and referring segmentation capabilities. This referencefree framework not only effectively handles rich-context scenarios, but also offers interpretability. To facilitate this, an innovative pipeline is accordingly designed for generating instruction-tuning data (R-Instruct), featuring richcontext descriptions, grounding masks, and hard negative samples. We further establish R 2 -HalBench, a new hallucination benchmark for MLLMs, which, unlike previous benchmarks, encompasses real-world, rich-context descriptions from 18 MLLMs with high-quality annotations, spanning diverse object-, attribute-, and relationship-level details. VBackChecker outperforms prior complex frameworks and achieves state-of-the-art performance on R^2 -HalBench, even rivaling GPT-4o’s capabilities in hallucination detection. It also surpasses prior methods in the pixel-level grounding task, achieving over a 10% improvement. Pinxue Guo, Chongruo Wu, Xinyu Zhou 0006, Lingyi Hong, Zhaoyu Chen 0001, Kaixun Jiang, Sen-Ching S. Cheung, Wei Zhang 0016 |
AAAI | 9 |
| 2026 | LVOS: A Benchmark for Large-Scale Long-Term Video Object SegmentationabstractVideo object segmentation (VOS) aims to distinguish and track target objects in a video. Despite the excellent performance achieved by off-the-shelf VOS models, part of the existing VOS benchmarks mainly focuses on short-term videos, where objects remain visible most of the time. However, these benchmarks may not fully capture challenges encountered in practical applications, and the absence of long-term datasets restricts further investigation of VOS in realistic scenarios. Thus, we propose a novel benchmark named LVOS, comprising 720 videos with 296,401 frames and 407,945 high-quality annotations. Videos in LVOS last 1.14 minutes on average. Each video includes various attributes, especially challenges encountered in the wild, such as long-term reappearing and cross-temporal similar objects. Compared to previous benchmarks, our LVOS better reflects VOS models' performance in real scenarios. Based on LVOS, we evaluate 15 existing VOS models under 3 different settings and conduct a comprehensive analysis. On LVOS, these models suffer a large performance drop, highlighting the challenge of achieving precise tracking and segmentation in real-world scenarios. Attribute-based analysis indicates that one of the significant factors contributing to accuracy decline is the increased video length, interacting with complex challenges such as long-term reappearance, cross-temporal confusion, and occlusion, which emphasize LVOS's crucial role. We hope our LVOS can advance development of VOS in real scenes. Lingyi Hong, Zhongying Liu, Chenzhi Tan, Yuang Feng, Xinyu Zhou 0006, Pinxue Guo, Zhaoyu Chen 0001, Shuyong Gao, Wei Zhang 0016 |
IEEE Trans. Pattern Anal. Mach. Intell. | 11 |
| 2026 | ClickVOS: Click Video Object SegmentationabstractVideo Object Segmentation (VOS) task aims to segment objects in videos. However, previous settings either require time-consuming manual masks of target objects at the first frame during inference or lack the flexibility to specify arbitrary objects of interest. To address these limitations, we propose the setting named Click Video Object Segmentation (ClickVOS) which segments objects of interest across the whole video according to a single click per object in the first frame. And we provide the extended datasets DAVIS-P and YouTubeVOS-P that with point annotations to support this task. ClickVOS is of significant practical applications and research implications due to its only 1-2 seconds interaction time for indicating an object, comparing annotating the mask of an object needs several minutes. However, ClickVOS also presents increased challenges. To address this task, we propose an end-to-end baseline approach named called Attention Before Segmentation (ABS), motivated by the attention process of humans. ABS utilizes the given point in the first frame to perceive the target object through a concise yet effective segmentation attention. Although the initial object mask is possibly inaccurate, in our ABS, as the video goes on, the initially imprecise object mask can self-heal instead of deteriorating due to error accumulation, which is attributed to our designed improvement memory that continuously records stable global object memory and updates detailed dense memory. In addition, we conduct various baseline explorations utilizing off-the-shelf algorithms from related fields, which could provide insights for the further exploration of ClickVOS. The experimental results demonstrate the superiority of the proposed ABS approach. Extended datasets and codes will be available at https://github.com/PinxueGuo/ClickVOS. Pinxue Guo, Lingyi Hong, Xinyu Zhou 0006, Shuyong Gao, Wanyun Li, Zhaoyu Chen 0001, Xiaoqiang Li 0002, Wei Zhang 0016 |
IEEE Trans. Circuits Syst. Video Technol. | 9 |
| 2025 | Less Attention is More: Prompt Transformer for Generalized Category DiscoveryabstractGeneralized Category Discovery (GCD) typically relies on the pre-trained Vision Transformer (ViT) to extract features from a global receptive field, followed by contrastive learning to simultaneously classify unlabeled known classes and unknown classes without priors. Owing to the deficiency in the modeling capacity for inner-patch local information within ViT, current methods primarily focus on discriminative features at global level. This results in a model with more yet scattered attention, where neither excessive nor insufficient focus can grasp subtle differences to classify fine-grained unknown and known categories. To address this issue, we propose the AptGCD to deliver apt attention for GCD. It mimics the human brain how leveraging visual perception to refine local attention and comprehend global context by proposing a Meta Visual Prompt (MVP) and Prompt Transformer (PT). MVP is introduced into GCD for the first time, refining channel-level attention, while adaptively self-learning unique inner-patch features as prompts to achieve local visual modeling for our prompt transformer. Yet, relying solely on detailed features can lead to skewed judgments. Hence, PT harmonizes local and global representations, guiding the model's interpretation of features through broader contexts, thereby capturing more useful details with less attention. Extensive experiments on seven datasets demonstrate that AptGCD outperforms current methods, it achieves an average proportional ‘New’ accuracy improvement of approximately 9.2% over SOTA method on the all four fine-grained datasets, establishing a new standard in the field. The code is available at https://github.com/wendy26zhang/AptGCD. Wei Zhang 0016, Baopeng Zhang, Zhu Teng, Wenxin Luo, Junnan Zou, Jianping Fan 0007 |
CVPR | 1 |
| 2025 | General Compression Framework for Efficient Transformer Object TrackingabstractPrevious works have attempted to improve tracking efficiency through lightweight architecture design or knowledge distillation from teacher models to compact student trackers. However, these solutions often sacrifice accuracy for speed to a great extent, and also have the problems of complex training process and structural limitations. Thus, we propose a general model compression framework for efficient transformer object tracking, named CompressTracker, to reduce model size while preserving tracking accuracy. Our approach features a novel stage division strategy that segments the transformer layers of the teacher model into distinct stages to break the limitation of model structure. Additionally, we also design a unique replacement training technique that randomly substitutes specific stages in the student model with those from the teacher model, as opposed to training the student model in isolation. Replacement training enhances the student model's ability to replicate the teacher model's behavior and simplifies the training process. To further forcing student model to emulate teacher model, we incorporate prediction guidance and stage-wise feature mimicking to provide additional supervision during the teacher model's compression process. CompressTracker is structurally agnostic, making it compatible with any transformer architecture. We conduct a series of experiment to verify the effectiveness and generalizability of our CompressTracker. Our CompressTracker-SUTrack, compressed from SUTrack, retains about 99 performance on LaSOT (72.2 AUC) while achieves 2.42x speed up. Code is available at https://github.com/LingyiHongfd/CompressTracker. Lingyi Hong, Xinyu Zhou 0006, Shilin Yan, Pinxue Guo, Kaixun Jiang, Zhaoyu Chen 0001, Shuyong Gao, Xingdong Sheng, Wei Zhang 0016, Hong Lu 0001 |
ICCV | 11 |
| 2025 | Shallow Features Matter: Hierarchical Memory with Heterogeneous Interaction for Unsupervised Video Object SegmentationabstractUnsupervised Video Object Segmentation (UVOS) aims to predict pixel-level masks for the most salient objects in videos without any prior annotations. While memory mechanisms have been proven critical in various video segmentation paradigms, their application in UVOS yield only marginal performance gains despite sophisticated design. Our analysis reveals a simple but fundamental flaw in existing methods: over-reliance on memorizing high-level semantic features. UVOS inherently suffers from the deficiency of lacking fine-grained information due to the absence of pixel-level prior knowledge. Consequently, memory design relying solely on high-level features, which predominantly capture abstract semantic cues, is insufficient to generate precise predictions. To resolve this fundamental issue, we propose a novel hierarchical memory architecture to incorporate both shallow- and high-level features for memory, which leverages the complementary benefits of pixel and semantic information. Furthermore, to balance the simultaneous utilization of the pixel and semantic memory features, we propose a heterogeneous interaction mechanism to perform pixel-semantic mutual interactions, which explicitly considers their inherent feature discrepancies. Through the design of Pixel-guided Local Alignment Module (PLAM) and Semantic-guided Global Integration Module (SGIM), we achieve delicate integration of the fine-grained details in shallow-level memory and the semantic representations in high-level memory. Our Hierarchical Memory with Heterogeneous Interaction Network (HMHI-Net) consistently achieves state-of-the-art performance across all UVOS and video saliency detection benchmarks. Moreover, HMHI-Net consistently exhibits high performance across different backbones, further demonstrating its superiority and robustness. Project page: https://github.com/ZhengxyFlow/HMHI-Net . Songcheng He, Wanyun Li, Xiaoqiang Li 0002, Wei Zhang 0016 |
ACM Multimedia | 5 |
| 2025 | Self-supervised video object segmentation via pseudo label rectification
Pinxue Guo, Wei Zhang 0016, Xiaoqiang Li 0002, Jianping Fan 0007 |
Pattern Recognit. | 2 |
| 2025 | Large Visual Language Models Continual Learning With Dynamic Mixture of ExpertsabstractIn dynamic and evolving application scenarios, the ability of visual language models to continuously learn from new data while preserving historical knowledge is critically important. Existing continual learning methods for large visual language models (LVLMs) often restrict the number of tasks they can handle, causing performance to decline as tasks continue to increase. In this paper, we propose a novel continual learning framework that adapts to the growing number of tasks, enabling visual language models to handle a dynamic range of open-set tasks while overcoming the catastrophic forgetting problem of learning new tasks at the expense of forgetting old ones. Our method builds on a pre-trained CLIP model and incorporates a dynamic mixture-of-experts (MoE) layer, enabling flexible adaptation to a wide range of open-set tasks. We design an elastic expert weight management strategy to effectively mitigate the catastrophic forgetting problem. Furthermore, we optimize the LoRA experts with adaptive ranks to achieve a balanced trade-off between model complexity and representational capacity. Extensive experiments across diverse settings demonstrate that our proposed method significantly reduces the number of tunable parameters while consistently surpassing state-of-the-art methods in new task learning capability and maintaining performance on historical tasks. Xihao Huang, Wei Zhang 0016 |
IEEE Trans. Image Process. | 3 |
| 2024 | Referred by Multi-Modality: A Unified Temporal Transformer for Video Object SegmentationabstractRecently, video object segmentation (VOS) referred by multi-modal signals, e.g., language and audio, has evoked increasing attention in both industry and academia. It is challenging for exploring the semantic alignment within modalities and the visual correspondence across frames. However, existing methods adopt separate network architectures for different modalities, and neglect the inter-frame temporal interaction with references. In this paper, we propose MUTR, a Multi-modal Unified Temporal transformer for Referring video object segmentation. With a unified framework for the first time, MUTR adopts a DETR-style transformer and is capable of segmenting video objects designated by either text or audio reference. Specifically, we introduce two strategies to fully explore the temporal relations between videos and multi-modal signals. Firstly, for low-level temporal aggregation before the transformer, we enable the multi-modal references to capture multi-scale visual cues from consecutive video frames. This effectively endows the text or audio signals with temporal knowledge and boosts the semantic alignment between modalities. Secondly, for high-level temporal interaction after the transformer, we conduct inter-frame feature communication for different object embeddings, contributing to better object-wise correspondence for tracking along the video. On Ref-YouTube-VOS and AVSBench datasets with respective text and audio references, MUTR achieves +4.2% and +8.7% J&F improvements to state-of-the-art methods, demonstrating our significance for unified multi-modal VOS. Code is released at https://github.com/OpenGVLab/MUTR. Shilin Yan, Renrui Zhang, Wei Zhang 0016, Hongyang Li 0001, Yu Qiao 0001, Hao Dong 0003, Zhongjiang He, Peng Gao 0007 |
AAAI | 5 |
| 2024 | OneVOS: Unifying Video Object Segmentation with All-in-One Transformer Framework
Wanyun Li, Pinxue Guo, Xinyu Zhou 0006, Lingyi Hong, Yangji He, Wei Zhang 0016 |
ECCV (58) | 7 |
| 2024 | PanoVOS: Bridging Non-panoramic and Panoramic Views with Transformer for Video Segmentation
Shilin Yan, Xiaohao Xu, Renrui Zhang, Lingyi Hong, Wei Zhang 0016 |
ECCV (9) | 7 |
| 2024 | X-Prompt: Multi-modal Visual Prompt for Video Object Segmentation
Pinxue Guo, Wanyun Li, Lingyi Hong, Xinyu Zhou 0006, Zhaoyu Chen 0001, Kaixun Jiang, Wei Zhang 0016 |
ACM Multimedia | 9 |
| 2024 | HFVOS: History-Future Integrated Dynamic Memory for Video Object SegmentationabstractMemory-based methods have substantially enhanced the precision of video object segmentation (VOS) by storing features in an expanding memory bank. However, this comes at the cost of increased computational demands and storage overhead. While recent methods have sought to alleviate this issue via compression or selection strategies, their reliance solely on history cues and simple memory structures result in precision degradation and intrinsic limitations, such as error accumulation and poor robustness. In this paper, we introduce HFVOS, an efficient yet effective framework to bolster VOS performance in both speed and precision by meticulously considering the memory design with low redundancy, high accuracy, and adaptability. First, we construct a novel hierarchical memory update pipeline with the proposed Buffered Memory Mechanism, which incorporates both future and history cues to reduce redundancy and improve the utility of memory. Second, we propose an Adaptive Dual-stream Selection Network (ADSN) to carry out the adaptive selection and drop operations of the memory update, and integrate an ADSN based long-term memory to enhance the robustness, especially for long videos. Furthermore, to further boost HFVOS, a progressive selection loss is designed to facilitate ADSN gradually adapt to fewer features while preserving high precision. Experiments show that HFVOS achieves the state-of-the-art segmentation precision and speed on both short-term datasets (DAVIS-17 val: 86.8%J&Fand 33.0 FPS, DAVIS-16 val: 92.0%J&Fand 42.0 FPS) and long-term datasets (LVOS val: 58.0%J&Fand 37.4 FPS). Code will be available at https://github.com/L599wy/HFVOS. Wanyun Li, Jack Fan, Pinxue Guo, Lingyi Hong, Wei Zhang 0016 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | LVOS: A Benchmark for Long-term Video Object SegmentationabstractExisting video object segmentation (VOS) benchmarks focus on short-term videos which just last about 3-5 seconds and where objects are visible most of the time. These videos are poorly representative of practical applications, and the absence of long-term datasets restricts further investigation of VOS on the application in realistic scenarios. So, in this paper, we present a new benchmark dataset named LVOS, which consists of 220 videos with a total duration of 421 minutes. To the best of our knowledge, LVOS is the first densely annotated long-term VOS dataset. The videos in our LVOS last 1.59 minutes on average, which is 20 times longer than videos in existing VOS datasets. Each video includes various attributes, especially challenges deriving from the wild, such as long-term reappearing and cross-temporal similar objeccts. Based on LVOS, we assess existing video object segmentation algorithms and propose a Diverse Dynamic Memory network (DDMemory) that consists of three complementary memory banks to exploit temporal information adequately. The experimental results demonstrate the strength and weaknesses of prior methods, pointing promising directions for further study. Data and code are available at https://lingyihongfd.github.io/lvos.github.io/. Lingyi Hong, Zhongying Liu, Wei Zhang 0016, Pinxue Guo, Zhaoyu Chen 0001 |
ICCV | 4 |
| 2023 | SimulFlow: Simultaneously Extracting Feature and Identifying Target for Unsupervised Video Object SegmentationabstractUnsupervised video object segmentation (UVOS) aims at detecting the primary objects in a given video sequence without any human interposing. Most existing methods rely on two-stream architectures that separately encode the appearance and motion information before fusing them to identify the target and generate object masks. However, this pipeline is computationally expensive and can lead to suboptimal performance due to the difficulty of fusing the two modalities properly. In this paper, we propose a novel UVOS model called SimulFlow that simultaneously performs feature extraction and target identification, enabling efficient and effective unsupervised video object segmentation. Concretely, we design a novel SimulFlow Attention mechanism to bridege the image and motion by utilizing the flexibility of attention operation, where coarse masks predicted from fused feature at each stage are used to constrain the attention operation within the mask area and exclude the impact of noise. Because of the bidirectional information flow between visual and optical flow features in SimulFlow Attention, no extra hand-designed fusing module is required and we only adopt a light decoder to obtain the final prediction. We evaluate our method on several benchmark datasets and achieve state-of-the-art results. Our proposed approach not only outperforms existing methods but also addresses the computational complexity and fusion difficulties caused by two-stream architectures. Our models achieve 87.4 ℐ&F on DAVIS-16 with the highest speed (63.7 FPS on a 3090) and the lowest parameters (13.7 M). Our SimulFlow also obtains competitive results on video salient object detection datasets. Lingyi Hong, Wei Zhang 0016, Shuyong Gao, Hong Lu 0001 |
ACM Multimedia | 2 |
| 2023 | Reading Relevant Feature from Global Representation Memory for Visual Object TrackingabstractReference features from a template or historical frames are crucial for visual object tracking. Prior works utilize all features from a fixed template or memory for visual object tracking. However, due to the dynamic nature of videos, the required reference historical information for different search regions at different time steps is also inconsistent. Therefore, using all features in the template and memory can lead to redundancy and impair tracking performance. To alleviate this issue, we propose a novel tracking paradigm, consisting of a relevance attention mechanism and a global representation memory, which can adaptively assist the search region in selecting the most relevant historical information from reference features. Specifically, the proposed relevance attention mechanism in this work differs from previous approaches in that it can dynamically choose and build the optimal global representation memory for the current frame by accessing cross-
frame information globally. Moreover, it can flexibly read the relevant historical information from the constructed memory to reduce redundancy and counteract the negative effects of harmful information. Extensive experiments validate the effectiveness of the proposed method, achieving competitive performance on five challenging datasets with 71 FPS. Xinyu Zhou 0006, Pinxue Guo, Lingyi Hong, Wei Zhang 0016, Weifeng Ge |
NeurIPS | 5 |
| 2023 | Dual Cross-Attention for Video Object Segmentation via Uncertainty RefinementabstractIn this paper, we propose a novel approach to video object segmentation where dual streams consisting of a shared network and a special network are designed to constitute the feature memory of history frames. Cues of spatial position and time stamp are explicitly explored to learn the context for each frame in the video sequence. Self-attention and cross-attention are simultaneously exploited to extract more powerful features for segmentation. In contrast to STM and its variants, the proposed dual cross-attention performs in both appearance space and semantic space such that the derived features are more distinctive and then robust to similar overlapping objects. During decoding for segmentation, a local refinement technique is designed for the uncertain boundaries to obtain more precise and smooth object contours. Experimental results on the challenging benchmark datasets DAVIS-2016, DAVIS-2017, and YouTube-VOS demonstrate the effectiveness of our proposed approach to video object segmentation. Jiahao Hong, Wei Zhang 0016 |
IEEE Trans. Multim. | 2 |
| 2022 | Weakly-Supervised Salient Object Detection Using Point SupervisonabstractCurrent state-of-the-art saliency detection models rely heavily on large datasets of accurate pixel-wise annotations, but manually labeling pixels is time-consuming and labor-intensive. There are some weakly supervised methods developed for alleviating the problem, such as image label, bounding box label, and scribble label, while point label still has not been explored in this field. In this paper, we propose a novel weakly-supervised salient object detection method using point supervision. To infer the saliency map, we first design an adaptive masked flood filling algorithm to generate pseudo labels. Then we develop a transformer-based point-supervised saliency detection model to produce the first round of saliency maps. However, due to the sparseness of the label, the weakly supervised model tends to degenerate into a general foreground detection model. To address this issue, we propose a Non-Salient Suppression (NSS) method to optimize the erroneous saliency maps generated in the first round and leverage them for the second round of training. Moreover, we build a new point-supervised dataset (P-DUTS) by relabeling the DUTS dataset. In P-DUTS, there is only one labeled point for each salient object. Comprehensive experiments on five largest benchmark datasets demonstrate our method outperforms the previous state-of-the-art methods trained with the stronger supervision and even surpass several fully supervised state-of-the-art models. The code is available at: https://github.com/shuyonggao/PSOD. Shuyong Gao, Wei Zhang 0016, Yan Wang 0068, Yangji He |
AAAI | 2 |
| 2022 | FERV39k: A Large-Scale Multi-Scene Dataset for Facial Expression Recognition in VideosabstractCurrent benchmarks for facial expression recognition (FER) mainly focus on static images, while there are limited datasets for FER in videos. It is still ambiguous to evaluate whether performances of existing methods remain satisfactory in real-world application-oriented scenes. For example, the “Happy” expression with high intensity in Talk-Show is more discriminating than the same expression with low intensity in Official-Event. To fill this gap, we build a large-scale multi-scene dataset, coined as FERV39k. We analyze the important ingredients of constructing such a novel dataset in three aspects: (1) multi-scene hierarchy and expression class, (2) generation of candidate video clips, (3) trusted manual labelling process. Based on these guidelines, we select 4 scenarios subdivided into 22 scenes, annotate 86k samples automatically obtained from 4k videos based on the well-designed workflow, and finally build 38,935 video clips labeled with 7 classic expressions. Experiment benchmarks on four kinds of baseline frame-works were also provided and further analysis on their performance across different scenes and some challenges for future research were given. Besides, we systematically investigate key components of DFER by ablation studies. The baseline framework and our project are available on https://github.com/wangyanckxx/FERV39k. Yan Wang 0068, Yixuan Sun, Zhongying Liu, Shuyong Gao, Wei Zhang 0016, Weifeng Ge |
CVPR | 6 |
| 2022 | Tokenizing Features for Fast Video Object SegmentationabstractThis paper investigates how to take full advantage of the tem-poral and spatial information in videos with minimal compu-tational cost in the semi-supervised video object segmentation (VOS) task. Current state-of-the-art methods have achieved remarkable performance by matching features of the current frame with those of past frames to propagate the past segmen-tation masks to the current. However, the inference speeds of such matching-based methods are limited due to the tremen-dous amount of computation on pixel-to-pixel matching. To address this problem, we propose a fast matching mechanism for VOS that extracts essential object information as a handful of token vectors for matching. By extracting succinct but suf-ficient information from the pixel-wise features, we develop a fast VOS model which achieves competitive segmentation performance (81.6%$J$&$F$on DAVIS-2017), maintaining a high inference speed (FPS = 42.1). Tianfang Meng, Wei Zhang 0016 |
ICME | 2 |
| 2022 | Weakly Supervised Video Salient Object Detection via Point SupervisionabstractFully supervised video salient object detection models have achieved excellent performance, yet obtaining pixel-by-pixel annotated datasets is laborious. Several works attempt to use scribble annotations to mitigate this problem, but point supervision as a more labor-saving annotation method (even the most labor-saving method among manual annotation methods for dense prediction), has not been explored. In this paper, we propose a strong baseline model based on point supervision. To infer saliency maps with temporal information, we mine inter-frame complementary information from short-term and long-term perspectives, respectively. Specifically, we propose a hybrid token attention module, which mixes optical flow and image information from orthogonal directions, adaptively highlighting critical optical flow information (channel dimension) and critical token information (spatial dimension). To exploit long-term cues, we develop the Long-term Cross-Frame Attention module (LCFA), which assists the current frame in inferring salient objects based on multi-frame tokens. Furthermore, we label two point-supervised datasets, P-DAVIS and P-DAVSOD, by relabeling the DAVIS and the DAVSOD dataset. Experiments on the six benchmark datasets illustrate our method outperforms the previous state-of-the-art weakly supervised methods and even is comparable with some fully supervised approaches. Our source code and datasets are available at: https://github.com/shuyonggao/PVSOD. Shuyong Gao, Haozhe Xing, Wei Zhang 0016, Yan Wang 0068 |
ACM Multimedia | 3 |
| 2022 | Designing intelligent self-checkup based technologies for everyday healthy living
Yanqi Jiang, Xianghua Ding, Xinning Gui, Wei Zhang 0016 |
Int. J. Hum. Comput. Stud. | 6 |
| 2022 | Adaptive Online Mutual Learning Bi-Decoders for Video Object SegmentationabstractOne of the major challenges facing video object segmentation (VOS) is the gap between the training and test datasets due to unseen category in test set, as well as object appearance change over time in the video sequence. To overcome such challenges, an adaptive online framework for VOS is developed with bi-decoders mutual learning. We learn object representation per pixel with bi-level attention features in addition to CNN features, and then feed them into mutual learning bi-decoders whose outputs are further fused to obtain the final segmentation result. We design an adaptive online learning mechanism via a deviation correcting trigger such that bi-decoders online mutual learning will be activated when the previous frame is segmented well meanwhile the current frame is segmented relatively worse. Knowledge distillation from the well segmented previous frames, along with mutual learning between bi-decoders, improves generalization ability and robustness of VOS model. Thus, the proposed model adapts to the challenging scenarios including unseen categories, object deformation, and appearance variation during inference. We extensively evaluate our model on widely-used VOS benchmarks including DAVIS-2016, DAVIS-2017, YouTubeVOS-2018, YouTubeVOS-2019, and UVO. Experimental results demonstrate the superiority of the proposed model over state-of-the-art methods. Pinxue Guo, Wei Zhang 0016, Xiaoqiang Li 0002 |
IEEE Trans. Image Process. | 2 |
| 2022 | Adaptive Selection of Reference Frames for Video Object SegmentationabstractVideo object segmentation is a challenging task in computer vision because the appearances of target objects might change drastically along the time in the video. To solve this problem, space-time memory (STM) networks are exploited to make use of the information from all the intermediate frames between the first frame and the current frame in the video. However, fully using the information from all the memory frames may make STM not practical for long videos. To overcome this issue, a novel method is developed in this paper to select the reference frames adaptively. First, an adaptive selection criterion is introduced to choose the reference frames with similar appearance and precise mask estimation, which can efficiently capture the rich information of the target object and overcome the challenges of appearance changes, occlusion, and model drift. Secondly, bi-matching (bi-scale and bi-direction) is conducted to obtain more robust correlations for objects of various scales and prevents multiple similar objects in the current frame from being mismatched with the same target object in the reference frame. Thirdly, a novel edge refinement technique is designed by using an edge detection network to obtain smooth edges from the outputs of edge confidence maps, where the edge confidence is quantized into ten sub-intervals to generate smooth edges step by step. Experimental results on the challenging benchmark datasets DAVIS-2016, DAVIS-2017, YouTube-VOS, and a Long-Video dataset have demonstrated the effectiveness of our proposed approach to video object segmentation. Lingyi Hong, Wei Zhang 0016, Liangyu Chen 0002, Jianping Fan 0001 |
IEEE Trans. Image Process. | 2 |
| 2022 | Guided Filter Network for Semantic Image SegmentationabstractThe existing publicly available datasets with pixel-level labels contain limited categories, and it is difficult to generalize to the real world containing thousands of categories. In this paper, we propose an approach to generate object masks with detailed pixel-level structures/boundaries automatically to enable semantic image segmentation of thousands of targets in the real world without manually labelling. A Guided Filter Network (GFN) is first developed to learn the segmentation knowledge from an existed dataset, and such GFN then transfers the learned segmentation knowledge to generate initial coarse object masks for the target images. These coarse object masks are treated as pseudo labels to self-optimize the GFN iteratively in the target images. Our experiments on six image sets have demonstrated that our proposed approach can generate object masks with detailed pixel-level structures/boundaries, whose quality is comparable to the manually-labelled ones. Our proposed approach also achieves better performance on semantic image segmentation than most existing weakly-supervised, semi-supervised, and domain adaptation approaches under the same experimental conditions. Xiang Zhang 0018, Wanqing Zhao, Wei Zhang 0016, Jinye Peng 0001, Jianping Fan 0001 |
IEEE Trans. Image Process. | 3 |
| 2021 | Points As Queries: Weakly Semi-Supervised Object Detection by PointsabstractWe propose a novel point annotated setting for the weakly semi-supervised object detection task, in which the dataset comprises small fully annotated images and large weakly annotated images by points. It achieves a balance between tremendous annotation burden and detection performance. Based on this setting, we analyze existing detectors and find that these detectors have difficulty in fully exploiting the power of the annotated points. To solve this, we introduce a new detector, Point DETR, which extends DETR by adding a point encoder. Extensive experiments conducted on MS-COCO dataset in various data settings show the effectiveness of our method. In particular, when using 20% fully labeled data from COCO, our detector achieves a promising performance, 33.3 AP, which outperforms a strong baseline (FCOS) by 2.0 AP, and we demonstrate the point annotations bring over 10 points in various AR metrics. Liangyu Chen 0002, Tong Yang 0005, Xiangyu Zhang 0005, Wei Zhang 0016, Jian Sun 0001 |
CVPR | 4 |
| 2021 | Dual-Stream Network Based On Global Guidance for Salient Object DetectionabstractHigh-level features can help low-level features eliminate semantic ambiguity, which is crucial for obtaining the precise salient object. Some methods use high-level features to provide global guidance for some layers of the network. However, there remain several problems: (1) the global guidance has not been fully mined, which leads to its limited capacity; (2) the semantic gap between global guidance and low-level features is ignored, and simple merging methods will cause feature aliasing. To remedy the problems, we propose a dual-stream network based on global guidance with two plug-ins, global attention based multi-scale high-level feature extraction module (GAMS) to mine global guidance and scale adaptive global guidance module (SAGG) to seamlessly integrate the global guidance into each decoding layer. Comprehensive experiments on the five largest benchmark datasets demonstrate our method outperforms previous state-of-the-art methods by a large margin. Code is available at https://github.com/shuyonggao/DSGGN. Shuyong Gao, Wei Zhang 0016, Zhongwei Ji |
ICASSP | 3 |
| 2021 | Adaptable Ensemble DistillationabstractOnline knowledge distillation (OKD), which simultaneously trains several peer networks to construct a powerful teacher on on-the-fly, has drawn much attention in recent years. OKD is designed to simplify the training procedure of conventional offline distillation. However, the ensemble strategy of existing OKD methods is inflexible and highly relies on random initializations. In this paper, we propose Adaptable Ensemble Distillation (AED) that inherits the merits of existing OKD methods while overcoming their major drawbacks. The novelty of our AED lies in three aspects: (1) an individual-regulated mechanism is proposed to flexibly regulate individual model and further generates an online ensemble with strong adaptability; (2) a diversity-aroused loss is designed to explicitly diversify individual models, which enhances the robustness of the ensemble; (3) an empirical distillation technique is adopted to directly promote knowledge transfer in OKD framework. Extensive experiments show that our proposed AED consistently outperforms the existing state-of-the-art OKD methods on various datasets. Wei Zhang 0016, Zhe Jiang 0004 |
ICASSP | 3 |
| 2021 | Global Cognition and Local Perception Network for Blind Image Deblurring
Chuanfa Zhang, Wei Zhang 0016, Yiting Cheng 0001, Shuyong Gao |
MMM (1) | 2 |
| 2020 | Automatic Tongue Crack Extraction For Real-Time DiagnosisabstractTongue crack segmentation is an essential component of computer-aided diagnosis applied in Traditional Chinese Medicine (TCM). However, existing methods are inadequate when dealing with the vague boundary of the foreground and the variation of tongue images. To this end, we propose a P-shaped neural network architecture based on the lightweight encoder-decoder structure: the encoder transforms pixel position information into channel information by aggregating adjacent pixel values; the decoder restores the image size and obtains the refined pixel-level extraction results by integrating the information of the corresponding layer in the encoder. To further improve the utilization of network parameters and the model's generalization ability, we design three novel sub-modules: (1) the phantom module utilizes cheap operations to generate feature maps, speeding up the calculation; (2) the dual-input module increases the original input information to enhance the model's foreground understanding; (3) the dual attention gate module strengthens the information fusion of high-level and low-level feature maps, retaining good boundary information while capturing detail information. Additionally, we propose a pre-training method based on cropped patch images, which makes the model sensitive to details of the foreground before formal training. We demonstrate the model's effectiveness on our constructed dataset, achieving 60.6% IoU accuracy, and the segmentation of a 513 × 513$image takes 390 ms on CPU. And our dataset is available at https://github.com/pengjianqiang/FDU-TC. Jianqiang Peng, Yingtao Zhang, Wei Zhang 0016, Yajie Kong, Fufeng Li |
BIBM | 5 |
| 2020 | Multi-scale Generative Adversarial Network for Automatic Sublingual Vein SegmentationabstractSublingual vein segmentation is an essential yet challenging task in computer-aided Traditional Chinese Medicine (TCM) tongue diagnosis. The most intricate part of sublingual vein segmentation is the strong diversity of sublingual vein images (e.g., various exposure of vein and complex background) and the natural connectivity of veins. To this end, we propose a novel and effective approach based on generative adversarial networks (GANs). Specifically, the framework comprises two modules: a segmentation generator network with multi-scale outputs and a scale-consistent discriminator. The former segmentation generator learns the overall structure of the segmentation, which reduces the training difficulties of stride 1 output by multi-scale outputs. The latter scale-consistent discriminator regularizes the segmentation maps from different output strides, which keeps the natural connectivity of veins and avoids generated noise. Additionally, we have constructed and publicized a well-annotated sublingual vein dataset (FDU-SV) which we believe will promote the significant development of this area. Extensive experimental results confirm that our method outperforms other representative segmentation models with a remarkable margin, achieving the state-of-the-art performance of 64.53% mIoU score. Code and dataset are available at https://github.com/echobear313/FDUVEIN. Qingyue Xiong, Wei Zhang 0016, Yajie Kong, Fufeng Li |
BIBM | 4 |
| 2020 | Maximal Information Complemented Refinement Network for Gland Instance SegmentationabstractThe analysis of glandular morphology is a crucial step to determine the presence and grade of cancer. The rise of computational pathology has led to the development of automated segmentation to overcome the time-consuming manual segmentation. Although the existing encoder-decoder networks haved made significant progress, the downsample operation causes fine-grain information loss. It deteriorates boundaries' localization especially in malignant cases. In this paper, we propose a maximal information complemented refinement network based on UNet. We extend the skip connection with two information complement, aggregate spatial detail information by reuse low-level features, and introduce semantic information by high-level feature guidance. Besides, a weighted cross-entropy loss and generalized dice loss is used to tackle the fuzzy boundary and class imbalance. We evaluated our model against a dozen recent deep learning models on the 2015 MICCAI Gland Segmentation challenge (GlaS) dataset. Extensive experiments show that our proposal achieves the best overall performance, immensely improves the performance of malignant cases. Wei Zhang 0016, Yajie Kong |
BIBM | 4 |
| 2020 | Adaptive Fractional Dilated Convolution Network for Image Aesthetics AssessmentabstractTo leverage deep learning for image aesthetics assessment, one critical but unsolved issue is how to seamlessly incorporate the information of image aspect ratios to learn more robust models. In this paper, an adaptive fractional dilated convolution (AFDC), which is aspect-ratio-embedded, composition-preserving and parameter-free, is developed to tackle this issue natively in convolutional kernel level. Specifically, the fractional dilated kernel is adaptively constructed according to the image aspect ratios, where the interpolation of nearest two integer dilated kernels are used to cope with the misalignment of fractional sampling. Moreover, we provide a concise formulation for mini-batch training and utilize a grouping strategy to reduce computational overhead. As a result, it can be easily implemented by common deep learning libraries and plugged into popular CNN architectures in a computation-efficient manner. Our experimental results demonstrate that our proposed method achieves state-of-the-art performance on image aesthetics assessment over the AVA dataset. Qiuyu Chen, Wei Zhang 0016, Yi Xu 0003, Yu Zheng 0006, Jianping Fan 0001 |
CVPR | 2 |
| 2020 | Towards Stabilizing Batch Statistics in Backward Propagation of Batch Normalization
Ruosi Wan, Xiangyu Zhang 0005, Wei Zhang 0016, Jian Sun 0001 |
ICLR | 4 |
| 2019 | Embedding Complementary Deep Networks for Image ClassificationabstractIn this paper, a deep embedding algorithm is developed to achieve higher accuracy rates on large-scale image classification. By adapting the importance of the object classes to their error rates, our deep embedding algorithm can train multiple complementary deep networks sequentially, where each of them focuses on achieving higher accuracy rates for different subsets of object classes in an easy-to-hard way. By integrating such complementary deep networks to generate an ensemble network, our deep embedding algorithm can improve the accuracy rates for the hard object classes (which initially have higher error rates) at certain degrees while effectively preserving high accuracy rates for the easy object classes. Our deep embedding algorithm has achieved higher overall accuracy rates on large scale image classification. Qiuyu Chen, Wei Zhang 0016, Jun Yu 0002, Jianping Fan 0001 |
CVPR | 2 |
| 2019 | Deep Mixture of Diverse Experts for Large-Scale Visual RecognitionabstractIn this paper, a deep mixture of diverse experts algorithm is developed to achieve more efficient learning of a huge (mixture) network for large-scale visual recognition application. First, a two-layer ontology is constructed to assign large numbers of atomic object classes into a set of task groups according to the similarities of their learning complexities, where certain degrees of inter-group task overlapping are allowed to enable sufficient inter-group message passing. Second, one particular base deep CNNs with M+1 outputs is learned for each task group to recognize its M atomic object classes and identify one special class of "not-in-group", where the network structure (numbers of layers and units in each layer) of the well-designed deep CNNs (such as AlexNet, VGG, GoogleNet, ResNet) is directly used to configure such base deep CNNs. For enhancing the separability of the atomic object classes in the same task group, two approaches are developed to learn more discriminative base deep CNNs: (a) our deep multi-task learning algorithm that can effectively exploit the inter-class visual similarities; (b) our two-layer network cascade approach that can improve the accuracy rates for the hard object classes at certain degrees while effectively maintaining the high accuracy rates for the easy ones. Finally, all these complementary base deep CNNs with diverse but overlapped outputs are seamlessly combined to generate a mixture network with larger outputs for recognizing tens of thousands of atomic object classes. Our experimental results have demonstrated that our deep mixture of diverse experts algorithm can achieve very competitive results on large-scale visual recognition. Qiuyu Chen, Zhenzhong Kuang, Jun Yu 0002, Wei Zhang 0016, Jianping Fan 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2018 | Leveraging Content Sensitiveness and User Trustworthiness to Recommend Fine-Grained Privacy Settings for Social Image SharingabstractTo configure successful privacy settings for social image sharing, two issues are inseparable: 1) content sensitiveness of the images being shared; and 2) trustworthiness of the users being granted to see the images. This paper aims to consider these two inseparable issues simultaneously to recommend fine-grained privacy settings for social image sharing. For achieving more compact representation of image content sensitiveness (privacy), two approaches are developed: 1) a deep network is adapted to extract 1024-D discriminative deep features; and 2) a deep multiple instance learning algorithm is adopted to identify 280 privacy-sensitive object classes and events. Second, users on the social network are clustered into a set of representative social groups to generate a discriminative dictionary for user trustworthiness characterization. Finally, both the image content sensitiveness and the user trustworthiness are integrated to train a tree classifier to recommend fine-grained privacy settings for social image sharing. Our experimental studies have demonstrated both the efficiency and the effectiveness of our proposed algorithms. Jun Yu 0002, Zhenzhong Kuang, Baopeng Zhang, Wei Zhang 0016, Dan Lin 0001, Jianping Fan 0001 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2018 | Embedding Visual Hierarchy With Deep Networks for Large-Scale Visual RecognitionabstractIn this paper, a layer-wise mixture model (LMM) is developed to support hierarchical visual recognition, where a Bayesian approach is used to automatically adapt the visual hierarchy to the progressive improvements of the deep network along the time. Our LMM algorithm can provide an end-to-end approach for jointly learning: (a) the deep network for achieving more discriminative deep representations for object classes and their inter-class visual similarities; (b) the tree classifier for recognizing large numbers of object classes hierarchically; and (c) the visual hierarchy adaptation for achieving more accurate assignment and organization of large numbers of object classes. By learning the tree classifier, the deep network and the visual hierarchy adaptation jointly in an end-to-end manner, our LMM algorithm can achieve higher accuracy rates on hierarchical visual recognition. Our experiments are carried on ImageNet1K and ImageNet10K image sets, which have demonstrated that our LMM algorithm can achieve very competitive results on the accuracy rates as compared with the baseline methods. Baopeng Zhang, Wei Zhang 0016, Jun Yu 0002, Jianping Fan 0001 |
IEEE Trans. Image Process. | 4 |
| 2016 | Model-Based Deep Hand Pose Estimation
Xingyi Zhou, Qingfu Wan, Wei Zhang 0016, Xiangyang Xue 0001 |
IJCAI | 3 |
| 2015 | Weakly supervised semantic segmentation for social imagesabstractImage semantic segmentation is the task of partitioning image into several regions based on semantic concepts. In this paper, we learn a weakly supervised semantic segmentation model from social images whose labels are not pixel-level but image-level; furthermore, these labels might be noisy. We present a joint conditional random field model leveraging various contexts to address this issue. More specifically, we extract global and local features in multiple scales by convolutional neural network and topic model. Inter-label correlations are captured by visual contextual cues and label co-occurrence statistics. The label consistency between image-level and pixel-level is finally achieved by iterative refinement. Experimental results on two real-world image datasets PASCAL VOC2007 and SIFT-Flow demonstrate that the proposed approach outperforms state-of-the-art weakly supervised methods and even achieves accuracy comparable with fully supervised methods. Wei Zhang 0016, Sheng Zeng, Dequan Wang, Xiangyang Xue 0001 |
CVPR | 1 |
| 2015 | Multiple Granularity Descriptors for Fine-Grained CategorizationabstractFine-grained categorization, which aims to distinguish subordinate-level categories such as bird species or dog breeds, is an extremely challenging task. This is due to two main issues: how to localize discriminative regions for recognition and how to learn sophisticated features for representation. Neither of them is easy to handle if there is insufficient labeled data. We leverage the fact that a subordinate-level object already has other labels in its ontology tree. These "free" labels can be used to train a series of CNN-based classifiers, each specialized at one grain level. The internal representations of these networks have different region of interests, allowing the construction of multi-grained descriptors that encode informative and discriminative features covering all the grain levels. Our multiple granularity framework can be learned with the weakest supervision, requiring only image-level label and avoiding the use of labor-intensive bounding box or part annotations. Experimental results on three challenging fine-grained image datasets demonstrate that our approach outperforms state-of-the-art algorithms, including those requiring strong labels. Dequan Wang, Jie Shao 0006, Wei Zhang 0016, Xiangyang Xue 0001, Zheng Zhang 0001 |
ICCV | 4 |
| 2014 | Semantic Segmentation Using Multiple Graphs with Block-Diagonal ConstraintsabstractIn this paper we propose a novel method for image semantic segmentation using multiple graphs. The multiview affinity graph is constructed by leveraging the consistency between semantic space and multiple visualspaces. With block-diagonal constraints, we enforce the affinity matrix to be sparse such that the pairwise potential for dissimilar superpixels is close to zero. By a divide-and-conquer strategy, the optimizationfor learning affinity matrix is decomposed into several subproblems that can be solved in parallel. Using the neighborhood relationship between superpixels and the consistency between affinity matrix and labelconfidencematrix, we infer the semantic label for each superpixel of unlabeled images by minimizing an objective whose closed form solution can be easily obtained. Experimental results on two real-world image datasetsdemonstrate the effectiveness of our method. Ke Zhang 0028, Wei Zhang 0016, Sheng Zeng, Xiangyang Xue 0001 |
AAAI | 2 |
| 2013 | Multi-View Embedding Learning for Incompletely Labeled Data
Wei Zhang 0016, Ke Zhang 0028, Pan Gu, Xiangyang Xue 0001 |
IJCAI | 1 |
| 2013 | Sparse Reconstruction for Weakly Supervised Semantic Segmentation
Ke Zhang 0028, Wei Zhang 0016, Yingbin Zheng, Xiangyang Xue 0001 |
IJCAI | 2 |
| 2012 | Semantic context learning with large-scale weakly-labeled image setabstractThere are a large number of images available on the web; meanwhile, only a subset of web images can be labeled by professionals because manual annotation is time-consuming and labor-intensive. Although we can now use the collaborative image tagging system, e.g., Flickr, to get a lot of tagged images provided by Internet users, these labels may be incorrect or incomplete. Furthermore, semantics richness requires more than one label to describe one image in real applications, and multiple labels usually interact with each other in semantic space. It is of significance to learn semantic context with large-scale weakly-labeled image set in the task of multi-label annotation. In this paper, we develop a novel method to learn semantic context and predict the labels of web images in a semi-supervised framework. To address the scalability issue, a small number of exemplar images are first obtained to cover the whole data cloud; then the label vector of each image is estimated as a local combination of the exemplar label vectors. Visual context, semantic context, and neighborhood consistency in both visual and semantic spaces are sufficiently leveraged in the proposed framework. Finally, the semantic context and the label confidence vectors for exemplar images are both learned in an iterative way. Experimental results on the real-world image dataset demonstrate the effectiveness of our method. Yao Lu 0028, Wei Zhang 0016, Ke Zhang 0028, Xiangyang Xue 0001 |
CIKM | 2 |
| 2012 | Learning attention map from imagesabstractWhile bottom-up and top-down processes have shown effectiveness during predicting attention and eye fixation maps on images, in this paper, inspired by the perceptual organization mechanism before attention selection, we propose to utilize figure-ground maps for the purpose. So as to take both pixel-wise and region-wise interactions into consideration when predicting label probabilities for each pixel, we develop a context-aware model based on multiple segmentation to obtain final results. The MIT attention dataset [14] is applied finally to evaluate both new features and model. Quantitative experiments demonstrate that figure-ground cues are valid in predicting attention selection, and our proposed model produces improvements over baseline method. Yao Lu 0028, Wei Zhang 0016, Cheng Jin 0001, Xiangyang Xue 0001 |
CVPR | 2 |
| 2011 | Salient Object Detection using concavity contextabstractConvexity (concavity) is a bottom-up cue to assign figure-ground relation in the perceptual organization [18]. It suggests that region on the convex side of a curved boundary tend to be figural. To explore the validity of this cue in the task of salient object detection, we segment the images in a test dataset into superpixels, and then locate the concave arcs and their bounding boxes along boundary of superpixels. Ecological statistics indicate that such bounding box contains salient object with a large probability. To utilize this spatial context information, i.e. concavity context, we follow the multi-scale analysis of human visual perception and design a hierarchical model. The model yields an affinity graph over candidate superpixels, in which weights between vertices are determined by the summation of concavity context on different scales in the hierarchy. Finally a graph-cut algorithm is performed to separate the salient and background objects. Evaluation on MSRA Salient Object Detection (SOD) dataset shows that concavity context is effective, and our approach provides improvement over state-of-the-art feature-based algorithms. Yao Lu 0028, Wei Zhang 0016, Hong Lu 0001, Xiangyang Xue 0001 |
ICCV | 2 |
| 2011 | Correlative multi-label multi-instance image annotationabstractIn this paper, each image is viewed as a bag of local regions, as well as it is investigated globally. A novel method is developed for achieving multi-label multi-instance image annotation, where image-level (bag-level) labels and region-level (instance-level) labels are both obtained. The associations between semantic concepts and visual features are mined both at the image level and at the region level. Inter-label correlations are captured by a co-occurence matrix of concept pairs. The cross-level label coherence encodes the consistency between the labels at the image level and the labels at the region level. The associations between visual features and semantic concepts, the correlations among the multiple labels, and the cross-level label coherence are sufficiently leveraged to improve annotation performance. Structural max-margin technique is used to formulate the proposed model and multiple interrelated classifiers are learned jointly. To leverage the available image-level labeled samples for the model training, the region-level label identification on the training set is firstly accomplished by building the correspondences between the multiple bag-level labels and the image regions. JEC distance based kernels are employed to measure the similarities both between images and between regions. Experimental results on real image datasets MSRC and Corel demonstrate the effectiveness of our method. Xiangyang Xue 0001, Wei Zhang 0016, Jianping Fan 0001, Yao Lu 0028 |
ICCV | 2 |
| 2011 | Multi-Kernel Multi-Label Learning with Max-Margin Concept Network
Wei Zhang 0016, Xiangyang Xue 0001, Jianping Fan 0001 |
IJCAI | 1 |
| 2011 | Automatic image annotation with weakly labeled datasetabstractIt is very attractive to exploit weakly-labeled image dataset for multi-label annotation applications. In our paper the meaning of the terminology weakly labeled is threefold: i) only a small subset of the available images are labeled; ii) even for the labeled image, the given labels may be uncorrect or incomplete; iii) the given labels do not provide the exact object locations in the images. A novel method is developed to predict the multiple labels for images and to provide region-level labels for the objects. We cluster the image regions to learn several region-exemplars and predict the label vector for each image region as a locally weighted average of the label vectors on exemplars. By investigating the label confidence matrix for the region-exemplars from different perspectives (column picture and row picture), we sufficiently leverage the visual contexts, the semantic contexts, and the consistency between similarities in the visual feature space and semantic label space. Experimental results on real web images demonstrate the effectiveness of the proposed method. Wei Zhang 0016, Yao Lu 0028, Xiangyang Xue 0001, Jianping Fan 0001 |
ACM Multimedia | 1 |
| 2008 | Metric learning by discriminant neighborhood embedding
Wei Zhang 0016, Xiangyang Xue 0001, Zichen Sun, Hong Lu 0001, Yue-Fei Guo |
Pattern Recognit. | 1 |
| 2007 | Efficient Feature Extraction for Image ClassificationabstractIn many image classification applications, input feature space is often high-dimensional and dimensionality reduction is necessary to alleviate the curse of dimensionality or to reduce the cost of computation. In this paper, we extract discriminant features for image classification by learning a low-dimensional embedding from finite labeled samples. In the new feature space, intra-class compactness and extra-class separability are achieved simultaneously. Target dimensionality of the embedding is selected by spectral analysis. Our method is designed suitable for data with both uni- and multi-modal class distributions. We also develop its two-dimensional variant which makes use of the matrix representation of images. Experimental results on three real image datasets demonstrate the efficacy of our method compared to the state of the art. Wei Zhang 0016, Xiangyang Xue 0001, Zichen Sun, Yue-Fei Guo, Mingmin Chi, Hong Lu 0001 |
ICCV | 1 |
| 2007 | Optimal dimensionality of metric space for classificationabstractIn many real-world applications, Euclidean distance in the original space is not good due to the curse of dimensionality. In this paper, we propose a new method, called Discriminant Neighborhood Embedding (DNE), to learn an appropriate metric space for classification given finite training samples. We define a discriminant adjacent matrix in favor of classification task, i.e., neighboring samples in the same class are squeezed but those in different classes are separated as far as possible. The optimal dimensionality of the metric space can be estimated by spectral analysis in the proposed method, which is of great significance for high-dimensional patterns. Experiments with various datasets demonstrate the effectiveness of our method. Wei Zhang 0016, Xiangyang Xue 0001, Zichen Sun, Yue-Fei Guo, Hong Lu 0001 |
ICML | 1 |
| 2006 | Quotient Set-based Nonlinear Manifold for Image RestorationabstractIn this paper we propose a patch-wise coarse-to-fine algorithm for image restoration using the manifold way of visual perception. All undistorted image patches are supposed to lie on a quotient set-based nonlinear manifold, and restoration of each degraded image patch can be implemented by projecting it to a locally linear region of such nonlinear manifold. The details of the original image can be learned from the undistorted training samples. Moreover, there is no need for us to assume that the degradation function is linear or to estimate some parameters of the blurs and noises beforehand. Experimental results demonstrate the effectiveness of the proposed method Wei Zhang 0016, Xiangyang Xue 0001, Hong Lu 0001, Yue-Fei Guo |
ICARCV | 1 |
| 2006 | Discriminant neighborhood embedding for classification
Wei Zhang 0016, Xiangyang Xue 0001, Hong Lu 0001, Yue-Fei Guo |
Pattern Recognit. | 1 |