VLDB 2026 Research / reviewers in the wild / expert
Zheng Yang 0008
dblp:59/5806-8
· DBLP profile ↗
26ranked-venue papers
2as first author
18since 2021 · last 2026
0009-0009-2840-2494ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 22 · 1 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 1 first-author · 13 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Delving into Dynamic Scene Cue-Consistency for Robust 3D Multi-Object Trackingabstract3D multi-object tracking is a critical and challenging task in the field of autonomous driving. A common paradigm relies on modeling individual object motion, e.g., Kalman filters, to predict trajectories. While effective in simple scenarios, this approach often struggles in crowded environments or with inaccurate detections, as it overlooks the rich geometric relationships between objects. This highlights the need to leverage spatial cues. However, existing geometry-aware methods can be susceptible to interference from irrelevant objects, leading to ambiguous features and incorrect associations. To address this, we propose focusing on cue-consistency: identifying and matching stable spatial patterns over time. We introduce the Dynamic Scene Cue-Consistency Tracker (DSC-Track) to implement this principle. Firstly, we design a unified spatiotemporal encoder using Point Pair Features (PPF) to learn discriminative trajectory embeddings while suppressing interference. Secondly, our cue-consistency transformer module explicitly aligns consistent feature representations between historical tracks and current detections. Finally, a dynamic update mechanism preserves salient spatiotemporal information for stable online tracking. Extensive experiments on the nuScenes and Waymo Open Datasets validate the effectiveness and robustness of our approach. On the nuScenes benchmark, for instance, our method achieves state-of-the-art performance, reaching 73.2% and 70.3% AMOTA on the validation and test sets, respectively. Boxi Wu 0001, Tu Zheng, Wang Yunhua, Zheng Yang 0008 |
AAAI | 6 |
| 2025 | STraj: Self-training for Bridging the Cross-Geography Gap in Trajectory PredictionabstractAccurate trajectory prediction has prominent significance in autonomous driving scenarios. Most existing methods predict the trajectory of an agent by learning its interaction with other agents and the map within the scenario. However, the heterogeneous distribution of these elements across different geographical scenarios is always ignored. Thus, trajectory predictors might struggle to generalize well when deployed in different geographical scenarios. To bridge the cross-geography gap, in this paper, we propose a plug-and-play self-training pipeline, termed STraj, for cross-geography trajectory prediction. STraj comprises three progressive steps: pseudo label (i.e., time-series trajectory) generation, update, and utilization. First, to generate pseudo labels that generalize to the cross-geography scenarios, STraj pre-trains the predictor through the complementary agent and map augmentations. Second, to facilitate the stable training of the predictor, we design a specific pseudo label update strategy. This strategy selects high-consistency pseudo trajectories from the current and historical epochs to supervise the target domain samples. Third, with generated pseudo trajectories, we introduce trajectory-induced contrastive learning to mitigate the representation bias of cross-geography agents. Extensive experiment results on various cross-geography trajectory prediction benchmarks demonstrate the effectiveness of STraj. Zhanwei Zhang, Minghao Chen 0001, Zhihong Gu, Xinkui Zhao, Zheng Yang 0008, Binbin Lin 0001, Deng Cai 0001, Wenxiao Wang 0001 |
AAAI | 5 |
| 2025 | Local Conditional Controlling for Text-to-Image Diffusion ModelsabstractDiffusion models have exhibited impressive prowess in the text-to-image task. Recent methods add image-level structure controls, e.g., edge and depth maps, to manipulate the generation process together with text prompts to obtain desired images. This controlling process is globally operated on the entire image, which limits the flexibility of control regions. In this paper, we explore a novel and practical task setting: local control. It focuses on controlling specific local region according to user-defined image conditions, while the remaining regions are only conditioned by the original text prompt. However, it is non-trivial to achieve it. The naive manner of directly adding local conditions may lead to the local control dominance problem, which forces the model to focus on the controlled region and neglect object generation in other regions. To mitigate this problem, we propose Regional Discriminate Loss to update the noised latents, aiming at enhanced object generation in non-control regions. Furthermore, the proposed Focused Token Response suppresses weaker attention scores which lack the strongest response to enhance object distinction and reduce duplication. Lastly, we adopt Feature Mask Constraint to reduce quality degradation in images caused by information differences across the local control region. All proposed strategies are operated at the inference stage. Extensive experiments demonstrate that our method can synthesize high-quality images aligned with the text prompt under local control conditions. Yang Yang 0002, Zekai Luo, Hengjia Li, Zheng Yang 0008, Xiaofei He 0001, Wei Zhao 0019, Qinglin Lu, Wei Liu 0005, Boxi Wu 0001 |
AAAI | 7 |
| 2025 | Object-level Data Augmentation for Visual 3D Object Detection in Autonomous DrivingabstractData augmentation plays an important role in visual-based 3D object detection. Existing detectors typically employ image/BEV-level data augmentation techniques, failing to utilize flexible object-level augmentations because of 2D-3D inconsistencies. This limitation hinders us from increasing the diversity of training data. To alleviate this issue, we propose an object-level data augmentation approach that incorporates scene reconstruction and neural scene rendering. Specifically, we reconstruct the scene and objects by extracting image features from sequences and aligning them with associated LiDAR point clouds. This approach is intended to conduct the editing process within a 3D space, allowing for flexible object manipulation. Additionally, we introduce a neural scene renderer to project the edited 3D scene onto a specified camera plane and render it onto a 2D image. Combined with scene reconstruction, it overcomes the challenges stemming from 2D/3D inconsistencies, enabling the generation of object-level augmented images with corresponding labels for model training. To validate the proposed method, we apply our method to various multi-camera 3D object detectors, consistently boosting the performance. Junkai Xu, Zheng Yang 0008, Xiaofei He 0001, Boxi Wu 0001 |
ICASSP | 4 |
| 2025 | SAMRefiner: Taming Segment Anything Model for Universal Mask RefinementabstractIn this paper, we explore a principal way to enhance the quality of widely pre-existing coarse masks, enabling them to serve as reliable training data for segmentation models to reduce the annotation cost. In contrast to prior refinement techniques that are tailored to specific models or tasks in a close-world manner, we propose SAMRefiner, a universal and efficient approach by adapting SAM to the mask refinement task. The core technique of our model is the noise-tolerant prompting scheme. Specifically, we introduce a multi-prompt excavation strategy to mine diverse input prompts for SAM (\ie, distance-guided points, context-aware elastic bounding boxes, and Gaussian-style masks) from initial coarse masks. These prompts can collaborate with each other to mitigate the effect of defects in coarse masks. In particular, considering the difficulty of SAM to handle the multi-object case in semantic segmentation, we introduce a split-then-merge (STM) pipeline. Additionally, we extend our method to SAMRefiner++ by introducing an additional IoU adaption step to further boost the performance of the generic SAMRefiner on the target dataset. This step is self-boosted and requires no additional annotation. The proposed framework is versatile and can flexibly cooperate with existing segmentation methods. We evaluate our mask framework on a wide range of benchmarks under different settings, demonstrating better accuracy and efficiency. SAMRefiner holds significant potential to expedite the evolution of refinement tools. Our code is available at https://github.com/linyq2117/SAMRefiner. Yuqi Lin, Hengjia Li, Wenqi Shao, Zheng Yang 0008, Jun Zhao 0009, Xiaofei He 0001, Ping Luo 0002, Kaipeng Zhang |
ICLR | 4 |
| 2025 | CLRNetV2: A Faster and Stronger Lane DetectorabstractLane is critical in the vision navigation system of intelligent vehicles. Naturally, the lane is a traffic sign with high-level semantics, whereas it owns the specific local pattern which needs detailed low-level features to localize accurately. Using different feature levels is of great importance for accurate lane detection, but it is still under-explored. On the other hand, current lane detection methods still struggle to detect complex dense lanes, such as Y-shape or fork-shape. In this work, we present Cross Layer Refinement Network aiming at fully utilizing both high-level and low-level features in lane detection. In particular, it first detects lanes with high-level semantic features and then performs refinement based on low-level features. In this way, we can exploit more contextual information to detect lanes while leveraging local-detailed features to improve localization accuracy. We present Fast-ROIGather to gather global context, which further enhances the representation of lane features. To detect dense lanes accurately, we propose Correlation Discrimination Module (CDM) to discriminate the correlation of dense lanes, enabling nearly cost-free high-quality dense lane prediction. In addition to our novel network design, we introduce LineIoU loss which regresses lanes as a whole unit to improve localization accuracy. Experiments demonstrate our approach significantly outperforms the state-of-the-art lane detection methods. Tu Zheng, Yifei Huang 0005, Yang Liu 0212, Binbin Lin 0001, Zheng Yang 0008, Deng Cai 0001, Xiaofei He 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2024 | TagCLIP: A Local-to-Global Framework to Enhance Open-Vocabulary Multi-Label Classification of CLIP without TrainingabstractContrastive Language-Image Pre-training (CLIP) has demonstrated impressive capabilities in open-vocabulary classification. The class token in the image encoder is trained to capture the global features to distinguish different text descriptions supervised by contrastive loss, making it highly effective for single-label classification. However, it shows poor performance on multi-label datasets because the global feature tends to be dominated by the most prominent class and the contrastive nature of softmax operation aggravates it. In this study, we observe that the multi-label classification results heavily rely on discriminative local features but are overlooked by CLIP. As a result, we dissect the preservation of patch-wise spatial information in CLIP and proposed a local-to-global framework to obtain image tags. It comprises three steps: (1) patch-level classification to obtain coarse scores; (2) dual-masking attention refinement (DMAR) module to refine the coarse scores; (3) class-wise reidentification (CWR) module to remedy predictions from a global perspective. This framework is solely based on frozen CLIP and significantly enhances its multi-label classification performance on various benchmarks without dataset-specific training. Besides, to comprehensively assess the quality and practicality of generated tags, we extend their application to the downstream task, i.e., weakly supervised semantic segmentation (WSSS) with generated tags as image-level pseudo labels. Experiments demonstrate that this classify-then-segment paradigm dramatically outperforms other annotation-free segmentation methods and validates the effectiveness of generated tags. Our code is available at https://github.com/linyq2117/TagCLIP. Yuqi Lin, Minghao Chen 0001, Kaipeng Zhang, Hengjia Li, Zheng Yang 0008, Dongqin Lv, Binbin Lin 0001, Haifeng Liu 0001, Deng Cai 0001 |
AAAI | 6 |
| 2024 | Towards Fine-Grained HBOE with Rendered Orientation Set and Laplace SmoothingabstractHuman body orientation estimation (HBOE) aims to estimate the orientation of a human body relative to the camera’s frontal view. Despite recent advancements in this field, there still exist limitations in achieving fine-grained results. We identify certain defects and propose corresponding approaches as follows: 1). Existing datasets suffer from non-uniform angle distributions, resulting in sparse image data for certain angles. To provide comprehensive and high-quality data, we introduce RMOS (Rendered Model Orientation Set), a rendered dataset comprising 150K accurately labeled human instances with a wide range of orientations. 2). Directly using one-hot vector as labels may overlook the similarity between angle labels, leading to poor supervision. And converting the predictions from radians to degrees enlarges the regression error. To enhance supervision, we employ Laplace smoothing to vectorize the label, which contains more information. For fine-grained predictions, we adopt weighted Smooth-L1-loss to align predictions with the smoothed-label, thus providing robust supervision. 3). Previous works ignore body-part-specific information, resulting in coarse predictions. By employing local-window self-attention, our model could utilize different body part information for more precise orientation estimations. We validate the effectiveness of our method in the benchmarks with extensive experiments and show that our method outperforms state-of-the-art. Project is available at: https://github.com/Whalesong-zrs/Towards-Fine-grained-HBOE. Ruisi Zhao, Zheng Yang 0008, Binbin Lin 0001, Xiaohui Zhong, Xiaobo Ren, Deng Cai 0001, Boxi Wu 0001 |
AAAI | 3 |
| 2024 | Learning Occupancy for Monocular 3D Object DetectionabstractMonocular 3D detection is a challenging task due to the lack of accurate 3D information. Existing approaches typically rely on geometry constraints and dense depth esti-mates to facilitate the learning, but often fail to fully ex-ploit the benefits of three-dimensional feature extraction in frustum and 3D space. In this paper, we propose Occu- pancyM3D, a method of learning occupancy for monocu-lar 3D detection. It directly learns occupancy in frustum and 3D space, leading to more discriminative and informative 3D features and representations. Specifically, by using synchronized raw sparse LiDAR point clouds, we define the space status and generate voxel-based occupancy labels. We formulate occupancy prediction as a simple classification problem and design associated occupancy losses. Re-sulting occupancy estimates are employed to enhance orig-inal frustum/3D features. As a result, experiments on KITTI and Waymo open datasets demonstrate that the proposed method achieves a new state of the art and surpasses other methods by a significant margin. Junkai Xu, Zheng Yang 0008, Xiaopei Wu, Wei Qian 0003, Wenxiao Wang 0001, Boxi Wu 0001, Deng Cai 0001 |
CVPR | 4 |
| 2024 | Few-shot Hybrid Domain Adaptation of Image GeneratorabstractCan a pre-trained generator be adapted to the hybrid of multiple target domains and generate images with integrated attributes of them? In this work, we introduce a new task -- Few-shot $\textit{Hybrid Domain Adaptation}$ (HDA). Given a source generator and several target domains, HDA aims to acquire an adapted generator that preserves the integrated attributes of all target domains, without overriding the source domain's characteristics. Compared with $\textit{Domain Adaptation}$ (DA), HDA offers greater flexibility and versatility to adapt generators to more composite and expansive domains. Simultaneously, HDA also presents more challenges than DA as we have access only to images from individual target domains and lack authentic images from the hybrid domain. To address this issue, we introduce a discriminator-free framework that directly encodes different domains' images into well-separable subspaces. To achieve HDA, we propose a novel directional subspace loss comprised of a distance loss and a direction loss. Concretely, the distance loss blends the attributes of all target domains by reducing the distances from generated images to all target subspaces. The direction loss preserves the characteristics from the source domain by guiding the adaptation along the perpendicular to subspaces. Experiments show that our method can obtain numerous domain-specific attributes in a single adapted generator, which surpasses the baseline methods in semantic similarity, image fidelity, and cross-domain consistency. Hengjia Li, Yang Liu 0212, Linxuan Xia, Yuqi Lin, Wenxiao Wang 0001, Tu Zheng, Zheng Yang 0008, Xiaohui Zhong, Xiaobo Ren, Xiaofei He 0001 |
ICLR | 7 |
| 2024 | Temporal Feature Fusion for 3D Detection in Monocular VideoabstractPrevious monocular 3D detection works focus on the single frame input in both training and inference. In real-world applications, temporal and motion information naturally exists in monocular video. It is valuable for 3D detection but under-explored in monocular works. In this paper, we propose a straightforward and effective method for temporal feature fusion, which exhibits low computation cost and excellent transferability, making it conveniently applicable to various monocular models. Specifically, with the help of optical flow, we transform the backbone features produced by prior frames and fuse them into the current frame. We introduce the scene feature propagating mechanism, which accumulates history scene features without extra time-consuming. In this process, occluded areas are removed via forward-backward scene consistency. Our method naturally introduces valuable temporal features, facilitating 3D reasoning in monocular 3D detection. Furthermore, accumulated history scene features via scene propagating mitigate heavy computation overheads for video processing. Experiments are conducted on variant baselines, which demonstrate that the proposed method is model-agonistic and can bring significant improvement to multiple types of single-frame methods. Zheng Yang 0008, Binbin Lin 0001, Xiaofei He 0001, Boxi Wu 0001 |
IEEE Trans. Image Process. | 3 |
| 2022 | CLRNet: Cross Layer Refinement Network for Lane DetectionabstractLane is critical in the vision navigation system of the intelligent vehicle. Naturally, lane is a traffic sign with high-level semantics, whereas it owns the specific local pattern which needs detailed low-level features to localize accurately. Using different feature levels is of great importance for accurate lane detection, but it is still under-explored. In this work, we present Cross Layer Refinement Network (CLRNet) aiming at fully utilizing both high-level and low-level features in lane detection. In particular, it first detects lanes with high-level semantic features then performs refinement based on low-level features. In this way, we can exploit more contextual information to detect lanes while leveraging local detailed lane features to improve localization accuracy. We present ROIGather to gather global context, which further enhances the feature representation of lanes. In addition to our novel network design, we introduce Line IoU loss which regresses the lane line as a whole unit to improve the localization accuracy. Experiments demonstrate that the proposed method greatly outperforms the state-of-the-art lane detection approaches. Code is available at: https://github.com/Turoad/CLRNet. Tu Zheng, Yifei Huang 0005, Yang Liu 0212, Wenjian Tang, Zheng Yang 0008, Deng Cai 0001, Xiaofei He 0001 |
CVPR | 5 |
| 2022 | Lidar Point Cloud Guided Monocular 3D Object Detection
Zhengxu Yu, Senbo Yan, Dan Deng, Zheng Yang 0008, Haifeng Liu 0001, Deng Cai 0001 |
ECCV (1) | 6 |
| 2022 | DID-M3D: Decoupling Instance Depth for Monocular 3D Object Detection
Xiaopei Wu, Zheng Yang 0008, Haifeng Liu 0001, Deng Cai 0001 |
ECCV (1) | 3 |
| 2022 | WeakM3D: Towards Weakly Supervised Monocular 3D Object Detection
Senbo Yan, Boxi Wu 0001, Zheng Yang 0008, Xiaofei He 0001, Deng Cai 0001 |
ICLR | 4 |
| 2022 | Domain Reconstruction and Resampling for Robust Salient Object DetectionabstractSalient Object Detection (SOD) aims at detecting the salient objects covering the whole natural scene. However, one of the main problems in SOD is data bias. Natural scenes vary greatly, while each image in the SOD dataset contains a specific scene. It means that each image is just a sampling point in a specific scene, which is not representative and causes serious sampling bias. Building larger datasets is one solution but costly to address the sampling bias. Our method regards the data distribution of natural scenes as a Gaussian Mixture Distribution, and each scene follows a sub-Gaussian distribution. Our main idea is to reconstruct the data distribution of each scene from the sampling images and then resample from the distribution domain. We represent a scene by a distribution instead of a fixed sampling image to reserve the sampling uncertainty in SOD. Specifically, we employ a Style Conditional Variational AutoEncoder (Style-CVAE) to reconstruct the data distribution from image styles and a Gaussian Randomize Attribute Filter (GRAF) to reconstruct data distribution from image attributes (such as lightness, saturation, hue, etc.). We resample the reconstructed data distribution according to the Gaussian probability density function and train the SOD model. Experimental results prove that our method outperforms 16 state-of-the-art methods on five benchmarks. Senbo Yan, Chuer Yu, Zheng Yang 0008, Haifeng Liu 0001, Deng Cai 0001 |
ACM Multimedia | 4 |
| 2021 | RESA: Recurrent Feature-Shift Aggregator for Lane DetectionabstractLane detection is one of the most important tasks in self-driving. Due to various complex scenarios (e.g., severe occlusion, ambiguous lanes, etc.) and the sparse supervisory signals inherent in lane annotations, lane detection task is still challenging. Thus, it is difficult for the ordinary convolutional neural network (CNN) to train in general scenes to catch subtle lane feature from the raw image. In this paper, we present a novel module named REcurrent Feature-Shift Aggregator (RESA) to enrich lane feature after preliminary feature extraction with an ordinary CNN. RESA takes advantage of strong shape priors of lanes and captures spatial relationships of pixels across rows and columns. It shifts sliced feature map recurrently in vertical and horizontal directions and enables each pixel to gather global information. RESA can conjecture lanes accurately in challenging scenarios with weak appearance clues by aggregating sliced feature map. Moreover, we propose a Bilateral Up-Sampling Decoder that combines coarse-grained and fine-detailed features in the up-sampling stage. It can recover the low-resolution feature map into pixel-wise prediction meticulously. Our method achieves state-of-the-art results on two popular lane detection benchmarks (CULane and Tusimple). Code has been made available at: https://github.com/ZJULearning/resa. Tu Zheng, Wenjian Tang, Zheng Yang 0008, Haifeng Liu 0001, Deng Cai 0001 |
AAAI | 5 |
| 2021 | TTFNeXt for real-time object detection
Tu Zheng, Zheng Yang 0008, Haifeng Liu 0001, Deng Cai 0001 |
Neurocomputing | 4 |
| 2020 | PI-RCNN: An Efficient Multi-Sensor 3D Object Detector with Point-Based Attentive Cont-Conv Fusion ModuleabstractLIDAR point clouds and RGB-images are both extremely essential for 3D object detection. So many state-of-the-art 3D detection algorithms dedicate in fusing these two types of data effectively. However, their fusion methods based on Bird's Eye View (BEV) or voxel format are not accurate. In this paper, we propose a novel fusion approach named Point-based Attentive Cont-conv Fusion(PACF) module, which fuses multi-sensor features directly on 3D points. Except for continuous convolution, we additionally add a Point-Pooling and an Attentive Aggregation to make the fused features more expressive. Moreover, based on the PACF module, we propose a 3D multi-sensor multi-task network called Pointcloud-Image RCNN(PI-RCNN as brief), which handles the image segmentation and 3D object detection tasks. PI-RCNN employs a segmentation sub-network to extract full-resolution semantic feature maps from images and then fuses the multi-sensor features via powerful PACF module. Beneficial from the effectiveness of the PACF module and the expressive semantic features from the segmentation module, PI-RCNN can improve much in 3D object detection. We demonstrate the effectiveness of the PACF module and PI-RCNN on the KITTI 3D Detection benchmark, and our method can achieve state-of-the-art on the metric of 3D AP. Liang Xie 0003, Chao Xiang, Zhengxu Yu, Zheng Yang 0008, Deng Cai 0001, Xiaofei He 0001 |
AAAI | 5 |
| 2020 | Training-Time-Friendly Network for Real-Time Object DetectionabstractModern object detectors can rarely achieve short training time, fast inference speed, and high accuracy at the same time. To strike a balance among them, we propose the Training-Time-Friendly Network (TTFNet). In this work, we start with light-head, single-stage, and anchor-free designs, which enable fast inference speed. Then, we focus on shortening training time. We notice that encoding more training samples from annotated boxes plays a similar role as increasing batch size, which helps enlarge the learning rate and accelerate the training process. To this end, we introduce a novel approach using Gaussian kernels to encode training samples. Besides, we design the initiative sample weights for better information utilization. Experiments on MS COCO show that our TTFNet has great advantages in balancing training time, inference speed, and accuracy. It has reduced training time by more than seven times compared to previous real-time detectors while maintaining state-of-the-art performances. In addition, our super-fast version of TTFNet-18 and TTFNet-53 can outperform SSD300 and YOLOv3 by less than one-tenth of their training time, respectively. The code has been made available at https://github.com/ZJULearning/ttfnet. Tu Zheng, Zheng Yang 0008, Haifeng Liu 0001, Deng Cai 0001 |
AAAI | 4 |
| 2019 | Region Mutual Information Loss for Semantic SegmentationabstractSemantic segmentation is a fundamental problem in computer vision. It is considered as a pixel-wise classification problem in practice, and most segmentation models use a pixel-wise loss as their optimization criterion. However, the pixel-wise loss ignores the dependencies between pixels in an image. Several ways to exploit the relationship between pixels have been investigated, \eg, conditional random fields (CRF) and pixel affinity based methods. Nevertheless, these methods usually require additional model branches, large extra memories, or more inference time. In this paper, we develop a region mutual information (RMI) loss to model the dependencies among pixels more simply and efficiently. In contrast to the pixel-wise loss which treats the pixels as independent samples, RMI uses one pixel and its neighbour pixels to represent this pixel. Then for each pixel in an image, we get a multi-dimensional point that encodes the relationship between pixels, and the image is cast into a multi-dimensional distribution of these high-dimensional points. The prediction and ground truth thus can achieve high order consistency through maximizing the mutual information (MI) between their multi-dimensional distributions. Moreover, as the actual value of the MI is hard to calculate, we derive a lower bound of the MI and maximize the lower bound to maximize the real value of the MI. RMI only requires a few extra computational resources in the training stage, and there is no overhead during testing. Experimental results demonstrate that RMI can achieve substantial and consistent improvements in performance on PASCAL VOC 2012 and CamVid datasets. The code is available at \url{https://github.com/ZJULearning/RMI}. Shuai Zhao 0006, Yang Wang 0030, Zheng Yang 0008, Deng Cai 0001 |
NeurIPS | 3 |
| 2016 | A-Optimal Non-negative Projection with Hessian regularization
Zheng Yang 0008, Haifeng Liu 0001, Deng Cai 0001, Zhaohui Wu 0001 |
Neurocomputing | 1 |
| 2014 | Matrix Completion for Cross-view Pairwise Constraint PropagationabstractAs pairwise constraints are usually easier to access than label information, pairwise constraint propagation attracts more and more attention in semi-supervised learning. Most existing pairwise constraint propagation methods are based on canonical graph propagation model, which heavily depends on the edge weights in the graph and cannot preserve local and global consistency simultaneously. In order to address this drawback, we cast cross-view pairwise constraint propagation into a problem of low rank matrix completion and propose a Matrix Completion method for cross-view Pairwise Constraint Propagation(MCPCP). With low rank requirement and graph regularization, our MCPCP can preserve local and global consistency simultaneously. We develop an algorithm based on alternating direction method of multipliers(ADMM) to solve the optimization problem. Finally, the effectiveness of MCPCP is demonstrated in cross-view multimedia retrieval. Zheng Yang 0008, Yao Hu 0002, Haifeng Liu 0001, Huajun Chen, Zhaohui Wu 0001 |
ACM Multimedia | 1 |
| 2014 | Local Coordinate Concept Factorization for Image RepresentationabstractLearning sparse representation of high-dimensional data is a state-of-the-art method for modeling data. Matrix factorization-based techniques, such as nonnegative matrix factorization and concept factorization (CF), have shown great advantages in this area, especially useful for image representation. Both of them are linear learning problems and lead to a sparse representation of the images. However, the sparsity obtained by these methods does not always satisfy locality conditions. For example, the learned new basis vectors may be relatively far away from the original data. Thus, we may not be able to achieve the optimal performance when using the new representation for other learning tasks, such as classification and clustering. In this paper, we introduce a locality constraint into the traditional CF. By requiring the concepts (basis vectors) to be as close to the original data points as possible, each datum can be represented by a linear combination of only a few basis concepts. Thus, our method is able to achieve sparsity and locality simultaneously. We analyze the complexity of our novel algorithm and demonstrate the effectiveness in comparison with the state-of-the-art approaches through a set of evaluations based on realworld applications. Haifeng Liu 0001, Zheng Yang 0008, Zhaohui Wu 0001, Xuelong Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2012 | A-Optimal Non-negative Projection for image representationabstractAs a central problem in computer vision and pattern recognition, data representation has attracted great attention in the past years. Non-negative matrix factorization (NMF) which is a useful data representation method makes great contribution on finding the latent structure of the data and leads to a parts-based representation by decomposing the data matrix into a few bases and encodings with nonnegative constraints. However, non-negative constraint is insufficient for getting more robust data representation. In this paper, we propose a novel method, called A-Optimal Non-negative Projection (ANP) for image data representation and further analysis. ANP imposes a constraint on the encoding factor as a regularizer during matrix factorization. In this way, the learned data representation leads to a stable linear model no matter what kind of data label is selected for further processing. Thus, it can preserve more intrinsic characteristics of the data regardless of any specific labels. We demonstrate the effectiveness of this novel algorithm through a set of evaluations on real world applications. Haifeng Liu 0001, Zheng Yang 0008, Zhaohui Wu 0001, Xuelong Li 0001 |
CVPR | 2 |
| 2011 | Locality-Constrained Concept FactorizationabstractMatrix factorization based techniques, such as nonnegative matrix factorization (NMF) and concept factorization (CF), have attracted great attention in dimension reduction and data clustering. Both of them are linear learning problems and lead to a sparse representation of the data. However, the sparsity obtained by these methods does not always satisfy locality conditions, thus the obtained data representation is not the best. This paper introduces a locality-constrained concept factorization method which imposes a locality constraint onto the traditional concept factorization. By requiring the concepts (basis vectors) to be as close to the original data points as possible, each data can be represented by a linear combination of only a few basis concepts. Thus our method is able to achieve sparsity and locality at the same time. We demonstrate the effectiveness of this novel algorithm through a set of evaluations on real world applications. Haifeng Liu 0001, Zheng Yang 0008, Zhaohui Wu 0001 |
IJCAI | 2 |