EDBT 2026 Demo / reviewers in the wild / expert
Deng Cai 0001
dblp:c/DCai
· DBLP profile ↗
271ranked-venue papers
35as first author
73since 2021 · last 2026
0000-0001-9817-4065ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 174 · 19 first-author · 53 since 2021Graphics, computer vision, multimedia, augmented reality and games · 128 · 13 first-author · 41 since 2021Databases, data management, data science and information retrieval · 65 · 19 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 2 since 2021Computer networks · 3 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Enhancing Spatial Reasoning Through Visual and Textual Thinking
Xin Guo 0006, Zhongming Jin 0001, Weihang Pan, Penghui Shang, Deng Cai 0001, Binbin Lin 0001, Jieping Ye |
AAAI | 6 |
| 2026 | Vidsketch: Hand-drawn sketch-driven video generation with diffusion control
Lifan Jiang, Boxi Wu 0001, Deng Cai 0001 |
Neural Networks | 4 |
| 2026 | Occluded person Re-Identification with noise injection
Can Yao, Xi Du, Deng Cai 0001, Shenqi Lai |
Pattern Recognit. | 5 |
| 2026 | 3 × 3 Kernel Is All You Need for VisionabstractMost modern Convolutional Neural Networks (CNNs) employ a multi-branch structure with various-sized convolutions to capture long- and short-range dependencies. However, these CNNs use large kernel convolutions (e.g., astonishingly 101 kernels) and specialized techniques (e.g., reparameterization and sparsity), increasing complexity in both training and inference stages. This paper focuses on designing an efficient CNN based on pure 3×3 convolutions without introducing complex operations and techniques. Specifically, we propose a Spatial Pyramid (SP) block, which consists of the Multi-branch Residual (MbR) module and the Gated-branch Residual (GbR) module. The MbR introduces multiscale pooling as the key component, thus capturing long-range visual cues through large down-sampling rates and shorter-range dependencies through low down-sampling rates while maintaining low computational complexity. Besides, the GbR uses one 3×3 convolution to refine dependencies along spatial and channel dimensions. Based on the SP block, we construct the Spatial Pyramid CNN (SPCNN), a model composed exclusively of Point-Wise Convolution and 3×3 Depth-Wise Convolution. Under comparable computational complexity, SPCNN significantly outperforms the state-of-the-art CNN PeLK (83.6% vs 82.6%) with only 3 × 3 kernels (compared to 101 × 101 kernels in PeLK). Besides, our SPCNN demonstrates comparability with state-of-the-art backbones in lightweight models, object detection, instance segmentation, and semantic segmentation. Moreover, evaluations of four image retrieval benchmarks also demonstrate the effectiveness. All codes are released at https://github.com/xiaolai-sqlai/SPCNN. Shenqi Lai, Mengjian Li, Haifeng Liu 0001, Xueming Qian, Deng Cai 0001, Yaxiong Wang |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | Beyond Fidelity: Diverse Image Synthesis via Retrieval-Augmented DiffusionabstractImage synthesis is a key application of generative AI. It can help reduce overfitting and the high cost of collecting real-world data for downstream discriminative models. However, current methods mainly focus on making images look realistic and ignore their true goal: improving downstream model generalization and robustness. We find that existing approaches tend to produce high-fidelity synthetic images that closely resemble the original data. This limits their value for improving downstream task performance. To overcome this, we argue that diversity, not just fidelity, must guide synthetic data generation if it is to truly complement human-collected datasets. In this paper, we introduce a Retrieval-Augmented Generation framework for diverse diffusion-based image synthesis. At each generation step, we retrieve the top-K most similar samples in feature space from both real and previously generated images. We then apply an Anti-Attention mechanism that actively pushes the new image away from these retrieved samples in feature space, maximizing dissimilarity. We propose novel evaluation metrics to assess image synthesis diversity and demonstrate significant improvements over existing benchmarks. Moreover, downstream models trained with our synthetic data achieved a 1.9% absolute accuracy gain on standard benchmarks, outperforming existing synthesis techniques. Linxuan Xia, Boxi Wu 0001, Tianrun Wu, Deng Cai 0001, Wei Liu 0005 |
IEEE Trans. Image Process. | 6 |
| 2026 | LEViT: Locally Enhanced Vision Transformer for Efficient Object Re-IdentificationabstractVision Transformer (ViT) on object re-identification (ReID) has attracted significant attention recently. However, ViT-based ReID substantially increases computational complexity, imposing significant burdens during training and inference. This paper presents an efficient and effective ViT-based backbone for ReID tasks, called the Locally Enhanced Vision Transformer (LEViT). ViT models typically emphasize global relationship modeling, yet ReID tasks are more sensitive to local information. To address this gap, we propose a Locally Enhanced (LE) block to enhance local information by performing self-attention within local split windows. Since part-based models dominate ReID, calculating self-attention across all patches is computationally inefficient. We also replace the traditional Query-Key-Value projector with the Group Convolution (G-Conv) projector, enabling the model to capture local details more efficiently. Furthermore, G-Conv is integrated into the channel MLP to strengthen local feature sensitivity. Using these components, we develop two LEViT variants: LEViT-S and LEViT-L. To our knowledge, LEViT is the first highly adaptable ViT backbone for ReID tasks. Experimental evaluations demonstrate the effectiveness in five ReID datasets: Market1501, DukeMTMC, MSMT17, VeRi-776, and VehicleID. Notably, LEViT-S outperforms TransReID while requiring less than 10% computational complexity. Furthermore, LEViT obtains the state-of-the-art on three deep metric learning datasets: CUB-200-2011, Cars196, and University-1652. Our code will be available athttps://github.com/YuhuiWang99/LEViT. Shenqi Lai, Mingyuan Fan 0002, Junshi Huang, Haifeng Liu 0001, Deng Cai 0001, Xueming Qian, Yaxiong Wang |
IEEE Trans. Multim. | 6 |
| 2025 | STraj: Self-training for Bridging the Cross-Geography Gap in Trajectory PredictionabstractAccurate trajectory prediction has prominent significance in autonomous driving scenarios. Most existing methods predict the trajectory of an agent by learning its interaction with other agents and the map within the scenario. However, the heterogeneous distribution of these elements across different geographical scenarios is always ignored. Thus, trajectory predictors might struggle to generalize well when deployed in different geographical scenarios. To bridge the cross-geography gap, in this paper, we propose a plug-and-play self-training pipeline, termed STraj, for cross-geography trajectory prediction. STraj comprises three progressive steps: pseudo label (i.e., time-series trajectory) generation, update, and utilization. First, to generate pseudo labels that generalize to the cross-geography scenarios, STraj pre-trains the predictor through the complementary agent and map augmentations. Second, to facilitate the stable training of the predictor, we design a specific pseudo label update strategy. This strategy selects high-consistency pseudo trajectories from the current and historical epochs to supervise the target domain samples. Third, with generated pseudo trajectories, we introduce trajectory-induced contrastive learning to mitigate the representation bias of cross-geography agents. Extensive experiment results on various cross-geography trajectory prediction benchmarks demonstrate the effectiveness of STraj. Zhanwei Zhang, Minghao Chen 0001, Zhihong Gu, Xinkui Zhao, Zheng Yang 0008, Binbin Lin 0001, Deng Cai 0001, Wenxiao Wang 0001 |
AAAI | 7 |
| 2025 | Enhancing Chain-of-Thought Reasoning with Critical Representation Fine-tuningabstractChenxi Huang, Shaotian Yan, Liang Xie, Binbin Lin, Sinan Fan, Yue Xin, Deng Cai, Chen Shen, Jieping Ye. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Chenxi Huang 0004, Shaotian Yan, Liang Xie 0003, Binbin Lin 0001, Sinan Fan, Deng Cai 0001, Chen Shen 0003, Jieping Ye |
ACL (1) | 7 |
| 2025 | Balance Discriminability and Integrality for Robust Salient Object Detection
Senbo Yan, Chuer Yu, Haifeng Liu 0001, Deng Cai 0001 |
ICANN (2) | 4 |
| 2025 | Magicid: Hybrid Preference Optimization for Id-Consistent and Dynamic-Preserved Video Customization
Hengjia Li, Lifan Jiang, Hongwei Yi, Boxi Wu 0001, Deng Cai 0001 |
ICCV | 7 |
| 2025 | PersonalVideo: High ID-Fidelity Video Customization without Dynamic and Semantic DegradationabstractThe current text-to-video (T2V) generation has made significant progress in synthesizing realistic general videos, but it is still under-explored in identity-specific human video generation with customized ID images. The key challenge lies in maintaining high ID fidelity consistently while preserving the original motion dynamic and semantic following after the identity injection. Current video identity customization methods mainly rely on reconstructing given identity images on text-to-image models, which have a divergent distribution with the T2V model. This process introduces a tuning-inference gap, leading to dynamic and semantic degradation. To tackle this problem, we propose a novel framework, dubbed $\textbf{PersonalVideo}$, that applies a mixture of reward supervision on synthesized videos instead of the simple reconstruction objective on images. Specifically, we first incorporate identity consistency reward to effectively inject the reference's identity without the tuning-inference gap. Then we propose a novel semantic consistency reward to align the semantic distribution of the generated videos with the original T2V model, which preserves its dynamic and semantic following capability during the identity injection. With the non-reconstructive reward training, we further employ simulated prompt augmentation to reduce overfitting by supervising generated results in more semantic scenarios, gaining good robustness even with only a single reference image. Extensive experiments demonstrate our method's superiority in delivering high identity faithfulness while preserving the inherent video generation qualities of the original T2V model, outshining prior methods. Hengjia Li, Haonan Qiu, Shiwei Zhang 0001, Xiang Wang 0012, Yujie Wei 0001, Yingya Zhang, Boxi Wu 0001, Deng Cai 0001 |
ICCV | 9 |
| 2025 | DSMRec: Dual-Spectrum Mamba Based Model for Sequential Recommendation
Ruming He, Penghui Shang, Deng Cai 0001 |
ICIC (8) | 5 |
| 2025 | Hybrid State Representation with Attention and Collision Time Prediction for Autonomous Vehicle Decision-Making at Unsignalized Intersections
Ruming He, Junran Xie, Penghui Shang, Deng Cai 0001 |
ICIC (20) | 6 |
| 2025 | FlexCAD: Unified and Versatile Controllable CAD Generation with Fine-tuned Large Language ModelsabstractRecently, there is a growing interest in creating computer-aided design (CAD) models based on user intent, known as controllable CAD generation. Existing work offers limited controllability and needs separate models for different types of control, reducing efficiency and practicality. To achieve controllable generation across all CAD construction hierarchies, such as sketch-extrusion, extrusion, sketch, face, loop and curve, we propose FlexCAD, a unified model by fine-tuning large language models (LLMs). First, to enhance comprehension by LLMs, we represent a CAD model as a structured text by abstracting each hierarchy as a sequence of text tokens. Second, to address various controllable generation tasks in a unified model, we introduce a hierarchy-aware masking strategy. Specifically, during training, we mask a hierarchy-aware field in the CAD text with a mask token. This field, composed of a sequence of tokens, can be set flexibly to represent various hierarchies. Subsequently, we ask LLMs to predict this masked field. During inference, the user intent is converted into a CAD text with a mask token replacing the part the user wants to modify, which is then fed into FlexCAD to generate new CAD models.
Comprehensive experiments on public dataset demonstrate the effectiveness of FlexCAD in both generation quality and controllability. Zhanwei Zhang, Shizhao Sun, Wenxiao Wang 0001, Deng Cai 0001, Jiang Bian 0002 |
ICLR | 4 |
| 2025 | Towards Label-Free 3D Visual Grounding with Vision Foundation Models
Xiaopei Wu, Yuenan Hou, Binbin Lin 0001, Xinge Zhu, Yuexin Ma, Haifeng Liu 0001, Deng Cai 0001, Xiao Sun 0001 |
IROS | 7 |
| 2025 | Controlling Thinking Speed in Reasoning ModelsabstractHuman cognition is theorized to operate in two modes: fast, intuitive System 1 thinking and slow, deliberate System 2 thinking.
While current Large Reasoning Models (LRMs) excel at System 2 thinking, their inability to perform fast thinking leads to high computational overhead and latency.
In this work, we enable LRMs to approximate human intelligence through dynamic thinking speed adjustment, optimizing accuracy-efficiency trade-offs.
Our approach addresses two key questions: (1) how to control thinking speed in LRMs, and (2) when to adjust it for optimal performance.
For the first question, we identify the steering vector that governs slow-fast thinking transitions in LRMs' representation space.
Using this vector, we achieve the first representation editing-based test-time scaling effect, outperforming existing prompt-based scaling methods.
For the second question, we apply real-time difficulty estimation to signal reasoning segments of varying complexity.
Combining these techniques, we propose the first reasoning strategy that enables fast processing of easy steps and deeper analysis for complex reasoning.
Without any training or additional cost, our plug-and-play method yields an average +1.3\% accuracy with -8.6\% token usage across leading LRMs and advanced reasoning benchmarks.
All of our algorithms are implemented based on vLLM and are expected to support broader applications and inspire future research. Zhengkai Lin, Zhihang Fu, Ze Chen 0001, Chao Chen 0026, Liang Xie 0003, Wenxiao Wang 0001, Deng Cai 0001, Jieping Ye |
NeurIPS | 7 |
| 2025 | GeoCAD: Local Geometry-Controllable CAD Generation with Large Language ModelsabstractLocal geometry-controllable computer-aided design (CAD) generation aims to modify local parts of CAD models automatically, enhancing design efficiency.
It also ensures that the shapes of newly generated local parts follow user-specific geometric instructions (e.g., an isosceles right triangle or a rectangle with one corner cut off).
However, existing methods encounter challenges in achieving this goal.
Specifically, they either lack the ability to follow textual instructions or are unable to focus on the local parts.
To address this limitation, we introduce GeoCAD, a user-friendly and local geometry-controllable CAD generation method.
Specifically, we first propose a complementary captioning strategy to generate geometric instructions for local parts.
This strategy involves vertex-based and VLLM-based captioning for systematically annotating simple and complex parts, respectively.
In this way, we caption $\sim$221k different local parts in total.
In the training stage, given a CAD model, we randomly mask a local part.
Then, using its geometric instruction and the remaining parts as input, we prompt large language models (LLMs) to predict the masked part.
During inference, users can specify any local part for modification while adhering to a variety of predefined geometric instructions.
Extensive experiments demonstrate the effectiveness of GeoCAD in generation quality, validity and text-to-CAD consistency. Zhanwei Zhang, Junjie Liu 0002, Wenxiao Wang 0001, Binbin Lin 0001, Liang Xie 0003, Chen Shen 0003, Deng Cai 0001 |
NeurIPS | 8 |
| 2025 | TokenSqueeze: Performance-Preserving Compression for Reasoning LLMsabstractEmerging reasoning LLMs such as OpenAI-o1 and DeepSeek-R1 have achieved strong performance on complex reasoning tasks by generating long chain-of-thought (CoT) traces. However, these long CoTs result in increased token usage, leading to higher inference latency and memory consumption. As a result, balancing accuracy and reasoning efficiency has become essential for deploying reasoning LLMs in practical applications. Existing long-to-short (Long2Short) methods aim to reduce inference length but often sacrifice accuracy, revealing a need for an approach that maintains performance while lowering token costs. To address this efficiency-accuracy tradeoff, we propose TokenSqueeze, a novel Long2Short method that condenses reasoning paths while preserving performance and relying exclusively on self-generated data. First, to prevent performance degradation caused by excessive compression of reasoning depth, we propose to select self-generated samples whose reasoning depth is adaptively matched to the complexity of the problem. To further optimize the linguistic expression without altering the underlying reasoning paths, we introduce a distribution-aligned linguistic refinement method that enhances the clarity and conciseness of the reasoning path while preserving its logical integrity. Comprehensive experimental results demonstrated the effectiveness of TokenSqueeze in reducing token usage while maintaining accuracy. Notably, DeepSeek‑R1‑Distill‑Qwen‑7B fine-tuned by using our proposed method achieved a 50\% average token reduction while preserving accuracy on the MATH500 benchmark. TokenSqueeze exclusively utilizes the model's self-generated data, enabling efficient and high-fidelity reasoning without relying on manually curated short-answer datasets across diverse applications. Our code is available at \url{https://github.com/zhangyx1122/TokenSqueeze}. Zhengxu Yu, Weihang Pan, Zhongming Jin 0001, Deng Cai 0001, Binbin Lin 0001, Jieping Ye |
NeurIPS | 6 |
| 2025 | CLRNetV2: A Faster and Stronger Lane DetectorabstractLane is critical in the vision navigation system of intelligent vehicles. Naturally, the lane is a traffic sign with high-level semantics, whereas it owns the specific local pattern which needs detailed low-level features to localize accurately. Using different feature levels is of great importance for accurate lane detection, but it is still under-explored. On the other hand, current lane detection methods still struggle to detect complex dense lanes, such as Y-shape or fork-shape. In this work, we present Cross Layer Refinement Network aiming at fully utilizing both high-level and low-level features in lane detection. In particular, it first detects lanes with high-level semantic features and then performs refinement based on low-level features. In this way, we can exploit more contextual information to detect lanes while leveraging local-detailed features to improve localization accuracy. We present Fast-ROIGather to gather global context, which further enhances the representation of lane features. To detect dense lanes accurately, we propose Correlation Discrimination Module (CDM) to discriminate the correlation of dense lanes, enabling nearly cost-free high-quality dense lane prediction. In addition to our novel network design, we introduce LineIoU loss which regresses lanes as a whole unit to improve localization accuracy. Experiments demonstrate our approach significantly outperforms the state-of-the-art lane detection methods. Tu Zheng, Yifei Huang 0005, Yang Liu 0212, Binbin Lin 0001, Zheng Yang 0008, Deng Cai 0001, Xiaofei He 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2025 | NeRF-Det++: Incorporating Semantic Cues and Perspective-Aware Depth Supervision for Indoor Multi-View 3D DetectionabstractNeRF-Det has achieved impressive performance in indoor multi-view 3D detection by innovatively utilizing NeRF to enhance representation learning. Despite its notable performance, we uncover three decisive shortcomings in its current design, including semantic ambiguity, inappropriate sampling, and insufficient utilization of depth supervision. To combat the aforementioned problems, we present three corresponding solutions: 1) Semantic Enhancement. We project the freely available 3D segmentation annotations onto the 2D plane and leverage the corresponding 2D semantic maps as the supervision signal, significantly enhancing the semantic awareness of multi-view detectors. 2) Perspective-Aware Sampling. Instead of employing the uniform sampling strategy, we put forward the perspective-aware sampling policy that samples densely near the camera while sparsely in the distance, more effectively collecting the valuable geometric clues. 3) Ordinal Residual Depth Supervision. As opposed to directly regressing the depth values that are difficult to optimize, we divide the depth range of each scene into a fixed number of ordinal bins and reformulate the depth prediction as the combination of the classification of depth bins as well as the regression of the residual depth values, thereby benefiting the depth learning process. The resulting algorithm, NeRF-Det++, has exhibited appealing performance in the ScanNetV2 and ARKITScenes datasets. Notably, in ScanNetV2, NeRF-Det++ outperforms the competitive NeRF-Det by +1.9% in mAP $\text{@}0.25$ and +3.5% in mAP $\text{@}0.50$ . The code will be publicly available at https://github.com/mrsempress/NeRF-Detplusplus. Chenxi Huang 0004, Yuenan Hou, Weicai Ye, Xiaoshui Huang, Binbin Lin 0001, Deng Cai 0001 |
IEEE Trans. Image Process. | 7 |
| 2025 | LoRA-Composer: Leveraging Low-Rank Adaptation for Multi-Concept Customization in Training-Free Diffusion ModelsabstractCustomization generation techniques have significantly advanced the synthesis of specific concepts across varied contexts. Multi-concept customization emerges as the challenging task within this domain. Existing approaches often rely on training a fusion matrix of multiple Low-Rank Adaptations (LoRAs) to merge various concepts into a single image. However, we identify this straightforward method faces two major challenges: 1) concept confusion, where the model struggles to preserve distinct individual characteristics, and 2) concept vanishing, where the model fails to generate the intended subjects. To address these issues, we introduce LoRA-Composer, a training-free framework designed for seamlessly integrating multiple LoRAs, thereby enhancing the harmony among different concepts within generated images. LoRA-Composer addresses concept vanishing through concept injection constraints, enhancing visibility via an expanded cross-attention mechanism. To combat concept confusion, concept isolation constraints are introduced, refining the self-attention computation. Furthermore, we propose two inference techniques to accelerate inference speed without performance degradation and enhance the accuracy of the generated region, respectively. Extensive experiments demonstrate that LoRA-Composer significantly outperforms standard baselines, especially in scenarios without image-based conditions such as canny edge or pose estimation. Yang Yang 0002, Chaotian Song, Hengjia Li, Qinglin Lu, Deng Cai 0001, Xiaofei He 0001, Boxi Wu 0001, Wei Liu 0005 |
IEEE Trans. Image Process. | 9 |
| 2024 | TeCH: Text-Guided Reconstruction of Lifelike Clothed HumansabstractDespite recent research advancements in reconstructing clothed humans from a single image, accurately restoring the “unseen regions” with high-level details remains an unsolved challenge that lacks attention. Existing methods often generate overly smooth back-side surfaces with a blurry texture. But how to effectively capture all visual attributes of an individual from a single image, which are sufficient to reconstruct unseen areas (e.g. the back view)? Motivated by the power of foundation models, TeCH reconstructs the 3D human by leveraging 1) descriptive text prompts (e.g. garments, colors, hairstyles) which are automatically generated via a garment parsing model and Visual Question Answering (VQA), 2) a personalized fine-tuned Text-to-Image diffusion model (T2I) which learns the “indescribable” appearance. To represent high-resolution 3D clothed humans at an affordable cost, we propose a hybrid 3D representation based on DMTet, which consists of an explicit body shape grid and an implicit distance field. Guided by the descriptive prompts + personalized T2I diffusion model, the geometry and texture of the 3D humans are optimized through multi-view Score Distillation Sampling (SDS) and reconstruction losses based on the original observation. TeCH produces high-fidelity 3D clothed humans with consistent & delicate texture, and detailed full-body geometry. Quantitative and qualitative experiments demonstrate that TeCH outperforms the state-of-the-art methods in terms of reconstruction accuracy and rendering quality. The code will be publicly available for research purposes at huangyangyi.github.io/TeCH Yangyi Huang, Hongwei Yi, Yuliang Xiu, Tingting Liao, Jiaxiang Tang, Deng Cai 0001, Justus Thies |
3DV | 6 |
| 2024 | TagCLIP: A Local-to-Global Framework to Enhance Open-Vocabulary Multi-Label Classification of CLIP without TrainingabstractContrastive Language-Image Pre-training (CLIP) has demonstrated impressive capabilities in open-vocabulary classification. The class token in the image encoder is trained to capture the global features to distinguish different text descriptions supervised by contrastive loss, making it highly effective for single-label classification. However, it shows poor performance on multi-label datasets because the global feature tends to be dominated by the most prominent class and the contrastive nature of softmax operation aggravates it. In this study, we observe that the multi-label classification results heavily rely on discriminative local features but are overlooked by CLIP. As a result, we dissect the preservation of patch-wise spatial information in CLIP and proposed a local-to-global framework to obtain image tags. It comprises three steps: (1) patch-level classification to obtain coarse scores; (2) dual-masking attention refinement (DMAR) module to refine the coarse scores; (3) class-wise reidentification (CWR) module to remedy predictions from a global perspective. This framework is solely based on frozen CLIP and significantly enhances its multi-label classification performance on various benchmarks without dataset-specific training. Besides, to comprehensively assess the quality and practicality of generated tags, we extend their application to the downstream task, i.e., weakly supervised semantic segmentation (WSSS) with generated tags as image-level pseudo labels. Experiments demonstrate that this classify-then-segment paradigm dramatically outperforms other annotation-free segmentation methods and validates the effectiveness of generated tags. Our code is available at https://github.com/linyq2117/TagCLIP. Yuqi Lin, Minghao Chen 0001, Kaipeng Zhang, Hengjia Li, Zheng Yang 0008, Dongqin Lv, Binbin Lin 0001, Haifeng Liu 0001, Deng Cai 0001 |
AAAI | 10 |
| 2024 | Semi-supervised 3D Object Detection with PatchTeacher and PillarMixabstractSemi-supervised learning aims to leverage numerous unlabeled data to improve the model performance. Current semi-supervised 3D object detection methods typically use a teacher to generate pseudo labels for a student, and the quality of the pseudo labels is essential for the final performance. In this paper, we propose PatchTeacher, which focuses on partial scene 3D object detection to provide high-quality pseudo labels for the student. Specifically, we divide a complete scene into a series of patches and feed them to our PatchTeacher sequentially. PatchTeacher leverages the low memory consumption advantage of partial scene detection to process point clouds with a high-resolution voxelization, which can minimize the information loss of quantization and extract more fine-grained features. However, it is non-trivial to train a detector on fractions of the scene. Therefore, we introduce three key techniques, i.e., Patch Normalizer, Quadrant Align, and Fovea Selection, to improve the performance of PatchTeacher. Moreover, we devise PillarMix, a strong data augmentation strategy that mixes truncated pillars from different LiDAR scans to generate diverse training samples and thus help the model learn more general representation. Extensive experiments conducted on Waymo and ONCE datasets verify the effectiveness and superiority of our method and we achieve new state-of-the-art results, surpassing existing methods by a large margin. Codes are available at https://github.com/LittlePey/PTPM. Xiaopei Wu, Liang Xie 0003, Yuenan Hou, Binbin Lin 0001, Xiaoshui Huang, Haifeng Liu 0001, Deng Cai 0001, Wanli Ouyang |
AAAI | 8 |
| 2024 | Regulating Intermediate 3D Features for Vision-Centric Autonomous DrivingabstractMulti-camera perception tasks have gained significant attention in the field of autonomous driving. However, existing frameworks based on Lift-Splat-Shoot (LSS) in the multi-camera setting cannot produce suitable dense 3D features due to the projection nature and uncontrollable densification process. To resolve this problem, we propose to regulate intermediate dense 3D features with the help of volume rendering. Specifically, we employ volume rendering to process the dense 3D features to obtain corresponding 2D features (e.g., depth maps, semantic maps), which are supervised by associated labels in the training. This manner regulates the generation of dense 3D features on the feature level, providing appropriate dense and unified features for multiple perception tasks. Therefore, our approach is termed Vampire, stands for ``Volume rendering As Multi-camera Perception Intermediate feature REgulator''. Experimental results on the Occ3D and nuScenes datasets demonstrate that Vampire facilitates fine-grained and appropriate extraction of dense 3D features, and is competitive with existing SOTA methods across diverse downstream perception tasks like 3D occupancy prediction, LiDAR segmentation and 3D objection detection, while utilizing moderate GPU resources. We provide a video demonstration in the supplementary materials and Codes are available at github.com/cskkxjk/Vampire. Junkai Xu, Linxuan Xia, Dan Deng, Wei Qian 0003, Wenxiao Wang 0001, Deng Cai 0001 |
AAAI | 9 |
| 2024 | Towards Fine-Grained HBOE with Rendered Orientation Set and Laplace SmoothingabstractHuman body orientation estimation (HBOE) aims to estimate the orientation of a human body relative to the camera’s frontal view. Despite recent advancements in this field, there still exist limitations in achieving fine-grained results. We identify certain defects and propose corresponding approaches as follows: 1). Existing datasets suffer from non-uniform angle distributions, resulting in sparse image data for certain angles. To provide comprehensive and high-quality data, we introduce RMOS (Rendered Model Orientation Set), a rendered dataset comprising 150K accurately labeled human instances with a wide range of orientations. 2). Directly using one-hot vector as labels may overlook the similarity between angle labels, leading to poor supervision. And converting the predictions from radians to degrees enlarges the regression error. To enhance supervision, we employ Laplace smoothing to vectorize the label, which contains more information. For fine-grained predictions, we adopt weighted Smooth-L1-loss to align predictions with the smoothed-label, thus providing robust supervision. 3). Previous works ignore body-part-specific information, resulting in coarse predictions. By employing local-window self-attention, our model could utilize different body part information for more precise orientation estimations. We validate the effectiveness of our method in the benchmarks with extensive experiments and show that our method outperforms state-of-the-art. Project is available at: https://github.com/Whalesong-zrs/Towards-Fine-grained-HBOE. Ruisi Zhao, Zheng Yang 0008, Binbin Lin 0001, Xiaohui Zhong, Xiaobo Ren, Deng Cai 0001, Boxi Wu 0001 |
AAAI | 7 |
| 2024 | Learning Occupancy for Monocular 3D Object DetectionabstractMonocular 3D detection is a challenging task due to the lack of accurate 3D information. Existing approaches typically rely on geometry constraints and dense depth esti-mates to facilitate the learning, but often fail to fully ex-ploit the benefits of three-dimensional feature extraction in frustum and 3D space. In this paper, we propose Occu- pancyM3D, a method of learning occupancy for monocu-lar 3D detection. It directly learns occupancy in frustum and 3D space, leading to more discriminative and informative 3D features and representations. Specifically, by using synchronized raw sparse LiDAR point clouds, we define the space status and generate voxel-based occupancy labels. We formulate occupancy prediction as a simple classification problem and design associated occupancy losses. Re-sulting occupancy estimates are employed to enhance orig-inal frustum/3D features. As a result, experiments on KITTI and Waymo open datasets demonstrate that the proposed method achieves a new state of the art and surpasses other methods by a significant margin. Junkai Xu, Zheng Yang 0008, Xiaopei Wu, Wei Qian 0003, Wenxiao Wang 0001, Boxi Wu 0001, Deng Cai 0001 |
CVPR | 9 |
| 2024 | TASeg: Temporal Aggregation Network for LiDAR Semantic SegmentationabstractTraining deep models for LiDAR semantic segmentation is challenging due to the inherent sparsity of point clouds. Utilizing temporal data is a natural remedy against the spar-sity problem as it makes the input signal denser. However, previous multi-frame fusion algorithms fall short in utilizing sufficient temporal information due to the memory constraint, and they also ignore the informative temporal images. To fully exploit rich information hidden in long-term temporal point clouds and images, we present the Temporal Aggre-gation Network, termed TASeg. Specifically, we propose a Temporal LiDAR Aggregation and Distillation (TLAD) algorithm, which leverages historical priors to assign dif-ferent aggregation steps for different classes. It can largely reduce memory and time overhead while achieving higher accuracy. Besides, TLAD trains a teacher injected with gt priors to distill the model, further boosting the performance. To make full use of temporal images, we design a Temporal Image Aggregation and Fusion (TIAF) module, which can greatly expand the camera FOVand enhance the present features. Temporal LiDAR points in the camera FOV are used as mediums to transform temporal image features to the present coordinate for temporal multi-modal fusion. Moreover, we develop a Static-Moving Switch Augmentation (SMSA) algorithm, which utilizes sufficient temporal information to enable objects to switch their motion states freely, thus greatly increasing static and moving training samples. Our TASeg ranks 1st††the date of CVPR deadline, i.e., 2023-11-18 07:59 AM UTC. on three challenging tracks, i.e., SemanticKITTI single-scan track, multi-scan track and nuScenes LiDAR segmentation track, strongly demonstrating the superiority of our method. Codes are available at https://github.com/LittlePey/TASeg. Xiaopei Wu, Yuenan Hou, Xiaoshui Huang, Binbin Lin 0001, Tong He 0001, Xinge Zhu, Yuexin Ma, Boxi Wu 0001, Haifeng Liu 0001, Deng Cai 0001, Wanli Ouyang |
CVPR | 10 |
| 2024 | Pseudo Label Refinery for Unsupervised Domain Adaptation on Cross-Dataset 3D Object DetectionabstractRecent self-training techniques have shown notable improvements in unsupervised domain adaptation for 3D object detection (3D UDA). These techniques typically select pseudo labels, i.e., 3D boxes, to supervise models for the target domain. However, this selection process inevitably introduces unreliable 3D boxes, in which 3D points cannot be definitively assigned as foreground or background. Previous techniques mitigate this by reweighting these boxes as pseudo labels, but these boxes can still poison the training process. To resolve this problem, in this paper, we propose a novel pseudo label refinery framework. Specifically, in the selection process, to improve the reliability of pseudo boxes, we propose a complementary augmentation strategy. This strategy involves either removing all points within an unre-liable box or replacing it with a high-confidence box. More-over, the point numbers of instances in high-beam datasets are considerably higher than those in low-beam datasets, also degrading the quality of pseudo labels during the training process. We alleviate this issue by generating additional proposals and aligning RoI features across different domains. Experimental results demonstrate that our method effectively enhances the quality of pseudo labels and consistently surpasses the state-of-the-art methods on six autonomous driving benchmarks. Code will be available at https://github.com/Zhanwei-Z/PERE. Zhanwei Zhang, Minghao Chen 0001, Hengjia Li, Binbin Lin 0001, Ping Li 0006, Wenxiao Wang 0001, Boxi Wu 0001, Deng Cai 0001 |
CVPR | 10 |
| 2024 | From Yes-Men to Truth-Tellers: Addressing Sycophancy in Large Language Models with Pinpoint TuningabstractLarge Language Models (LLMs) tend to prioritize adherence to user prompts over providing veracious responses, leading to the sycophancy issue. When challenged by users, LLMs tend to admit mistakes and provide inaccurate responses even if they initially provided the correct answer. Recent works propose to employ supervised fine-tuning (SFT) to mitigate the sycophancy issue, while it typically leads to the degeneration of LLMs' general capability. To address the challenge, we propose a novel supervised pinpoint tuning (SPT), where the region-of-interest modules are tuned for a given objective. Specifically, SPT first reveals and verifies a small percentage (<5%) of the basic modules, which significantly affect a particular behavior of LLMs. i.e., sycophancy. Subsequently, SPT merely fine-tunes these identified modules while freezing the rest. To verify the effectiveness of the proposed SPT, we conduct comprehensive experiments, demonstrating that SPT significantly mitigates the sycophancy issue of LLMs (even better than SFT). Moreover, SPT introduces limited or even no side effects on the general capability of LLMs. Our results shed light on how to precisely, effectively, and efficiently explain and improve the targeted ability of LLMs. Wei Chen 0005, Zhen Huang 0007, Liang Xie 0003, Binbin Lin 0001, Houqiang Li, Le Lu 0001, Xinmei Tian 0001, Deng Cai 0001, Yonggang Zhang 0003, Wenxiao Wang 0001, Xu Shen 0001, Jieping Ye |
ICML | 8 |
| 2024 | G2LTraj: A Global-to-Local Generation Approach for Trajectory Prediction
Zhanwei Zhang, Zishuo Hua, Minghao Chen 0001, Binbin Lin 0001, Deng Cai 0001, Wenxiao Wang 0001 |
IJCAI | 6 |
| 2024 | Delving into the Reversal Curse: How Far Can Large Language Models Generalize?abstractWhile large language models (LLMs) showcase unprecedented capabilities, they also exhibit certain inherent limitations when facing seemingly trivial tasks.
A prime example is the recently debated "reversal curse", which surfaces when models, having been trained on the fact "A is B", struggle to generalize this knowledge to infer that "B is A".
In this paper, we examine the manifestation of the reversal curse across various tasks and delve into both the generalization abilities and the problem-solving mechanisms of LLMs. This investigation leads to a series of significant insights:
(1) LLMs are able to generalize to "B is A" when both A and B are presented in the context as in the case of a multiple-choice question.
(2) This generalization ability is highly correlated to the structure of the fact "A is B" in the training documents. For example, this generalization only applies to biographies structured in "[Name] is [Description]" but not to "[Description] is [Name]".
(3) We propose and verify the hypothesis that LLMs possess an inherent bias in fact recalling during knowledge application, which explains and underscores the importance of the document structure to successful learning.
(4) The negative impact of this bias on the downstream performance of LLMs can hardly be mitigated through training alone.
Based on these intriguing findings, our work not only presents a novel perspective for interpreting LLMs' generalization abilities from their intrinsic working mechanism but also provides new insights for the development of more effective learning methods for LLMs. Zhengkai Lin, Zhihang Fu, Kai Liu 0023, Liang Xie 0003, Binbin Lin 0001, Wenxiao Wang 0001, Deng Cai 0001, Jieping Ye |
NeurIPS | 7 |
| 2023 | Towards In-Distribution Compatible Out-of-Distribution DetectionabstractDeep neural network, despite its remarkable capability of discriminating targeted in-distribution samples, shows poor performance on detecting anomalous out-of-distribution data. To address this defect, state-of-the-art solutions choose to train deep networks on an auxiliary dataset of outliers. Various training criteria for these auxiliary outliers are proposed based on heuristic intuitions. However, we find that these intuitively designed outlier training criteria can hurt in-distribution learning and eventually lead to inferior performance. To this end, we identify three causes of the in-distribution incompatibility: contradictory gradient, false likelihood, and distribution shift. Based on our new understandings, we propose a new out-of-distribution detection method by adapting both the top-design of deep models and the loss function. Our method achieves in-distribution compatibility by pursuing less interference with the probabilistic characteristic of in-distribution features. On several benchmarks, our method not only achieves the state-of-the-art out-of-distribution detection performance but also improves the in-distribution accuracy. Boxi Wu 0001, Jie Jiang 0015, Haidong Ren, Zifan Du, Wenxiao Wang 0001, Zhifeng Li 0001, Deng Cai 0001, Xiaofei He 0001, Binbin Lin 0001, Wei Liu 0005 |
AAAI | 7 |
| 2023 | One-shot Implicit Animatable Avatars with Model-based PriorsabstractExisting neural rendering methods for creating human avatars typically either require dense input signals such as video or multi-view images, or leverage a learned prior from large-scale specific 3D human datasets such that reconstruction can be performed with sparse-view inputs. Most of these methods fail to achieve realistic reconstruction when only a single image is available. To enable the data-efficient creation of realistic anima table 3D humans, we propose ELICIT, a novel method for learning human-specific neural radiance fields from a single image. Inspired by the fact that humans can effortlessly estimate the body geometry and imagine full-body clothing from a single image, we leverage two priors in ELICIT: 3D geometry prior and visual semantic prior. Specifically, ELICIT utilizes the 3D body shape geometry prior from a skinned vertex-based template model (i.e., SMPL) and implements the visual clothing semantic prior with the CLIP-based pre-trained models. Both priors are used to jointly guide the optimization for creating plausible content in the invisible areas. Taking advantage of the CLIP models, ELICIT can use text descriptions to generate text-conditioned unseen regions. In order to further improve visual details, we propose a segmentation-based sampling strategy that locally refines different parts of the avatar. Comprehensive evaluations on multiple popular benchmarks, including ZJU-MoCAP, Human3.6M, and DeepFashion, show that ELICIT outperforms strong baseline methods of avatar creation when only a single image is available. The code is public for research purposes at https://huangyangyi.github.io/ELICIT Yangyi Huang, Hongwei Yi, Weiyang Liu, Boxi Wu 0001, Wenxiao Wang 0001, Binbin Lin 0001, Debing Zhang, Deng Cai 0001 |
ICCV | 9 |
| 2023 | MonoNeRD: NeRF-like Representations for Monocular 3D Object DetectionabstractIn the field of monocular 3D detection, it is common practice to utilize scene geometric clues to enhance the detector’s performance. However, many existing works adopt these clues explicitly such as estimating a depth map and back-projecting it into 3D space. This explicit methodology induces sparsity in 3D representations due to the increased dimensionality from 2D to 3D, and leads to substantial information loss, especially for distant and occluded objects. To alleviate this issue, we propose MonoNeRD, a novel detection framework that can infer dense 3D geometry and occupancy. Specifically, we model scenes with Signed Distance Functions (SDF), facilitating the production of dense 3D representations. We treat these representations as Neural Radiance Fields (NeRF) and then employ volume rendering to recover RGB images and depth maps. To the best of our knowledge, this work is the first to introduce volume rendering for M3D, and demonstrates the potential of implicit reconstruction for image-based 3D perception. Extensive experiments conducted on the KITTI-3D benchmark and Waymo Open Dataset demonstrate the effectiveness of MonoNeRD. Codes are available at https://github.com/cskkxjk/MonoNeRD. Junkai Xu, Hao Li 0009, Wei Qian 0003, Wenxiao Wang 0001, Deng Cai 0001 |
ICCV | 8 |
| 2023 | LUT-NN: Empower Efficient Neural Network Inference with Centroid Learning and Table LookupabstractOn-device Deep Neural Network (DNN) inference consumes significant computing resources and development efforts. To alleviate that, we propose LUT-NN, the first system to empower inference by table lookup, to reduce inference cost. LUT-NN learns the typical features for each operator, named centroid, and precompute the results for these centroids to save in lookup tables. During inference, the results of the closest centroids with the inputs can be read directly from the table, as the approximated outputs without computations. Xiaohu Tang 0003, Yang Wang 0053, Ting Cao 0003, Li Lyna Zhang, Qi Chen 0009, Deng Cai 0001, Yunxin Liu 0001, Mao Yang 0004 |
MobiCom | 6 |
| 2023 | Neural collapse inspired attraction-repulsion-balanced loss for imbalanced learning
Liang Xie 0003, Deng Cai 0001, Xiaofei He 0001 |
Neurocomputing | 3 |
| 2023 | OBMO: One Bounding Box Multiple Objects for Monocular 3D Object DetectionabstractCompared to typical multi-sensor systems, monocular 3D object detection has attracted much attention due to its simple configuration. However, there is still a significant gap between LiDAR-based and monocular-based methods. In this paper, we find that the ill-posed nature of monocular imagery can lead to depth ambiguity. Specifically, objects with different depths can appear with the same bounding boxes and similar visual features in the 2D image. Unfortunately, the network cannot accurately distinguish different depths from such non-discriminative visual features, resulting in unstable depth training. To facilitate depth learning, we propose a simple yet effective plug-and-play module, One Bounding Box Multiple Objects (OBMO). Concretely, we add a set of suitable pseudo labels by shifting the 3D bounding box along the viewing frustum. To constrain the pseudo-3D labels to be reasonable, we carefully design two label scoring strategies to represent their quality. In contrast to the original hard depth labels, such soft pseudo labels with quality scores allow the network to learn a reasonable depth range, boosting training stability and thus improving final performance. Extensive experiments on KITTI and Waymo benchmarks show that our method significantly improves state-of-the-art monocular 3D detectors by a significant margin (The improvements under the moderate setting on KITTI validation set are 1.82 ~ 10.91% mAP in BEV and 1.18 ~ 9.36% mAP in 3D). Codes have been released at https://github.com/mrsempress/OBMO. Chenxi Huang 0004, Tong He 0001, Haidong Ren, Wenxiao Wang 0001, Binbin Lin 0001, Deng Cai 0001 |
IEEE Trans. Image Process. | 6 |
| 2023 | X-View: Non-Egocentric Multi-View 3D Object Detectorabstract3D object detection algorithms for autonomous driving reason about 3D obstacles either from 3D birds-eye view or perspective view or both. Recent works attempt to improve the detection performance via mining and fusing from multiple egocentric views. Although the egocentric perspective view alleviates some weaknesses of the birds-eye view, the sectored grid partition becomes so coarse in the distance that the targets and surrounding context mix together, which makes the features less discriminative. In this paper, we generalize the research on 3D multi-view learning and propose a novel multi-view-based 3D detection method, named X-view, to overcome the drawbacks of the multi-view methods. Specifically, X-view breaks through the traditional limitation about the perspective view whose original point must be consistent with the 3D Cartesian coordinate. X-view is designed as a general paradigm that can be applied on almost any 3D detectors based on LiDAR with only little increment of running time, no matter it is voxel/grid-based or raw-point-based. We conduct experiments on KITTI and NuScenes datasets to demonstrate the robustness and effectiveness of our proposed X-view. The results show that X-view obtains consistent improvements when combined with mainstream state-of-the-art 3D methods. Liang Xie 0003, Deng Cai 0001, Xiaofei He 0001 |
IEEE Trans. Image Process. | 3 |
| 2023 | Urban Traffic Light Control via Active Multi-Agent Communication and Supply-Demand ModelingabstractUrban traffic light control is an important and challenging real-world problem. By regarding intersections as agents, most of the reinforcement learning-based methods generate agents’ actions independently. They can cause action conflict and result in overflow or road resource waste in adjacent intersections. Recently, some collaborative methods have alleviated the above problems by extending the observable surroundings of agents, which can be considered inactive cross-agent communication methods. However, when agents act synchronously in these works, the perceived action value is biased, and the information exchanged is insufficient. In this work, we first propose a novel Multi-agent Communication and Action Rectification (MaCAR) framework. It enables active communication between agents by considering the impact of synchronous actions of agents. Another fundamental problem of traffic light control is the balance between traffic demand and road supply capacity. To fully describe the relation between traffic demand and road supply capacity (Supply-Demand modeling, SD), we further model and forecast the Supply-Demand relation to facilitating the effectiveness of the model’s action. The experiments show that our model outperforms state-of-the-art methods on both synthetic and real-world datasets. Combining the SD with MaCAR, SD-MaCAR can further boost the traffic light control performance even in traffic accident scenarios. Xin Guo 0006, Zhengxu Yu, Pengfei Wang 0008, Zhongming Jin 0001, Jianqiang Huang 0001, Deng Cai 0001, Xiaofei He 0001, Xian-Sheng Hua 0001 |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2022 | DMN4: Few-Shot Learning via Discriminative Mutual Nearest Neighbor Neural NetworkabstractFew-shot learning (FSL) aims to classify images under low-data regimes, where the conventional pooled global feature is likely to lose useful local characteristics. Recent work has achieved promising performances by using deep descriptors. They generally take all deep descriptors from neural networks into consideration while ignoring that some of them are useless in classification due to their limited receptive field, e.g., task-irrelevant descriptors could be misleading and multiple aggregative descriptors from background clutter could even overwhelm the object's presence. In this paper, we argue that a Mutual Nearest Neighbor (MNN) relation should be established to explicitly select the query descriptors that are most relevant to each task and discard less relevant ones from aggregative clutters in FSL. Specifically, we propose Discriminative Mutual Nearest Neighbor Neural Network (DMN4) for FSL. Extensive experiments demonstrate that our method outperforms the existing state-of-the-arts on both fine-grained and generalized datasets. Yang Liu 0212, Tu Zheng, Jie Song 0011, Deng Cai 0001, Xiaofei He 0001 |
AAAI | 4 |
| 2022 | SCALoss: Side and Corner Aligned Loss for Bounding Box RegressionabstractBounding box regression is an important component in object detection. Recent work achieves promising performance by optimizing the Intersection over Union (IoU). However, IoU-based loss has the gradient vanish problem in the case of low overlapping bounding boxes, and the model could easily ignore these simple cases. In this paper, we propose Side Overlap (SO) loss by maximizing the side overlap of two bounding boxes, which puts more penalty for low overlapping bounding box cases. Besides, to speed up the convergence, the Corner Distance (CD) is added into the objective function. Combining the Side Overlap and Corner Distance, we get a new regression objective function, Side and Corner Align Loss (SCALoss). The SCALoss is well-correlated with IoU loss, which also benefits the evaluation metric but produces more penalty for low-overlapping cases. It can serve as a comprehensive similarity measure, leading to better localization performance and faster convergence speed. Experiments on COCO, PASCAL VOC, and LVIS benchmarks show that SCALoss can bring consistent improvement and outperform ln loss and IoU based loss with popular object detectors such as YOLOV3, SSD, Faster-RCNN. Code is available at: https://github.com/Turoad/SCALoss. Tu Zheng, Shuai Zhao 0006, Yang Liu 0212, Deng Cai 0001 |
AAAI | 5 |
| 2022 | Frame-wise Action Representations for Long Videos via Sequence Contrastive LearningabstractPrior works on action representation learning mainly focus on designing various architectures to extract the global representations for short video clips. In contrast, many practical applications such as video alignment have strong demand for learning dense representations for long videos. In this paper, we introduce a novel contrastive action representation learning (CARL) framework to learn frame-wise action representations, especially for long videos, in a self-supervised manner. Concretely, we introduce a simple yet efficient video encoder that considers spatio-temporal context to extract frame-wise representations. Inspired by the recent progress of self-supervised learning, we present a novel sequence contrastive loss (SCL) applied on two correlated views obtained through a series of spatio-temporal data augmentations. SCL optimizes the embedding space by minimizing the KL-divergence between the sequence similarity of two augmented views and a prior Gaussian distribution of timestamp distance. Experiments on FineGym, PennAction and Pouring datasets show that our method outperforms previous state-of-the-art by a large margin for downstream fine-grained action classification. Surprisingly, although without training on paired videos, our approach also shows outstanding performance on video alignment and fine-grained frame retrieval tasks. Code and models are available at https://github.com/minghchen/CARL_code. Minghao Chen 0001, Fangyun Wei, Deng Cai 0001 |
CVPR | 4 |
| 2022 | Learning to Affiliate: Mutual Centralized Learning for Few-shot ClassificationabstractFew-shot learning (FSL) aims to learn a classifier that can be easily adapted to accommodate new tasks, given only a few examples. To handle the limited-data in few-shot regimes, recent methods tend to collectively use a set of local features to densely represent an image instead of using a mixed global feature. They generally explore a unidirectional paradigm, e.g., finding the nearest support feature for every query feature and aggregating local matches for a joint classification. In this paper, we propose a novel Mutual Centralized Learning (MCL) to fully affiliate these two disjoint dense features sets in a bidirectional paradigm. We first associate each local feature with a particle that can bidirectionally random walk in discrete feature space. To estimate the class probability, we propose the dense features' accessibility that measures the expected number of visits to the dense features of that class in a Markov process. We relate our method to learning a centrality on an affiliation network and demonstrate its capability to be plugged in existing methods by highlighting centralized local features. Experiments show that our method achieves the new state-of-the-art. Yang Liu 0212, Weifeng Zhang 0005, Chao Xiang, Tu Zheng, Deng Cai 0001, Xiaofei He 0001 |
CVPR | 5 |
| 2022 | Sparse Fuse Dense: Towards High Quality 3D Detection with Depth CompletionabstractCurrent LiDAR-only 3D detection methods inevitably suffer from the sparsity of point clouds. Many multi-modal methods are proposed to alleviate this issue, while different representations of images and point clouds make it difficult to fuse them, resulting in suboptimal performance. In this paper, we present a novel multi-modal framework SFD (Sparse Fuse Dense), which utilizes pseudo point clouds generated from depth completion to tackle the issues mentioned above. Different from prior works, we propose a new RoI fusion strategy 3D-GAF (3D Grid-wise Attentive Fusion) to make fuller use of information from different types of point clouds. Specifically, 3D-GAF fuses 3D RoI features from the pair of point clouds in a grid-wise attentive way, which is more fine- grained and more precise. In addition, we propose a SynAugment (Synchronized Augmentation) to enable our multi-modal framework to utilize all data augmentation approaches tailored to LiDAR-only methods. Lastly, we customize an effective and efficient feature extractor CPConv (Color Point Convolution) for pseudo point clouds. It can explore 2D image features and 3D geometric features of pseudo point clouds simultaneously. Our method holds the highest entry on the KITTI car 3D object detection leaderboard††On the date of CVPR deadline, i.e., Nov.16, 2021, demonstrating the effectiveness of our SFD. Code will be made publicly available. Xiaopei Wu, Honghui Yang, Liang Xie 0003, Chenxi Huang 0004, Chengqi Deng, Haifeng Liu 0001, Deng Cai 0001 |
CVPR | 8 |
| 2022 | CLRNet: Cross Layer Refinement Network for Lane DetectionabstractLane is critical in the vision navigation system of the intelligent vehicle. Naturally, lane is a traffic sign with high-level semantics, whereas it owns the specific local pattern which needs detailed low-level features to localize accurately. Using different feature levels is of great importance for accurate lane detection, but it is still under-explored. In this work, we present Cross Layer Refinement Network (CLRNet) aiming at fully utilizing both high-level and low-level features in lane detection. In particular, it first detects lanes with high-level semantic features then performs refinement based on low-level features. In this way, we can exploit more contextual information to detect lanes while leveraging local detailed lane features to improve localization accuracy. We present ROIGather to gather global context, which further enhances the feature representation of lanes. In addition to our novel network design, we introduce Line IoU loss which regresses the lane line as a whole unit to improve the localization accuracy. Experiments demonstrate that the proposed method greatly outperforms the state-of-the-art lane detection approaches. Code is available at: https://github.com/Turoad/CLRNet. Tu Zheng, Yifei Huang 0005, Yang Liu 0212, Wenjian Tang, Zheng Yang 0008, Deng Cai 0001, Xiaofei He 0001 |
CVPR | 6 |
| 2022 | Lidar Point Cloud Guided Monocular 3D Object Detection
Zhengxu Yu, Senbo Yan, Dan Deng, Zheng Yang 0008, Haifeng Liu 0001, Deng Cai 0001 |
ECCV (1) | 8 |
| 2022 | DID-M3D: Decoupling Instance Depth for Monocular 3D Object Detection
Xiaopei Wu, Zheng Yang 0008, Haifeng Liu 0001, Deng Cai 0001 |
ECCV (1) | 5 |
| 2022 | Towards Efficient Adversarial Training on Vision Transformers
Boxi Wu 0001, Jindong Gu, Zhifeng Li 0001, Deng Cai 0001, Xiaofei He 0001, Wei Liu 0005 |
ECCV (13) | 4 |
| 2022 | Graph R-CNN: Towards Accurate 3D Object Detection with Semantic-Decorated Local Graph
Honghui Yang, Xiaopei Wu, Wenxiao Wang 0001, Wei Qian 0003, Xiaofei He 0001, Deng Cai 0001 |
ECCV (8) | 7 |
| 2022 | CrossFormer: A Versatile Vision Transformer Hinging on Cross-scale Attention
Wenxiao Wang 0001, Long Chen 0016, Binbin Lin 0001, Deng Cai 0001, Xiaofei He 0001, Wei Liu 0005 |
ICLR | 5 |
| 2022 | WeakM3D: Towards Weakly Supervised Monocular 3D Object Detection
Senbo Yan, Boxi Wu 0001, Zheng Yang 0008, Xiaofei He 0001, Deng Cai 0001 |
ICLR | 6 |
| 2022 | Domain Reconstruction and Resampling for Robust Salient Object DetectionabstractSalient Object Detection (SOD) aims at detecting the salient objects covering the whole natural scene. However, one of the main problems in SOD is data bias. Natural scenes vary greatly, while each image in the SOD dataset contains a specific scene. It means that each image is just a sampling point in a specific scene, which is not representative and causes serious sampling bias. Building larger datasets is one solution but costly to address the sampling bias. Our method regards the data distribution of natural scenes as a Gaussian Mixture Distribution, and each scene follows a sub-Gaussian distribution. Our main idea is to reconstruct the data distribution of each scene from the sampling images and then resample from the distribution domain. We represent a scene by a distribution instead of a fixed sampling image to reserve the sampling uncertainty in SOD. Specifically, we employ a Style Conditional Variational AutoEncoder (Style-CVAE) to reconstruct the data distribution from image styles and a Gaussian Randomize Attribute Filter (GRAF) to reconstruct data distribution from image attributes (such as lightness, saturation, hue, etc.). We resample the reconstructed data distribution according to the Gaussian probability density function and train the SOD model. Experimental results prove that our method outperforms 16 state-of-the-art methods on five benchmarks. Senbo Yan, Chuer Yu, Zheng Yang 0008, Haifeng Liu 0001, Deng Cai 0001 |
ACM Multimedia | 6 |
| 2022 | High Dimensional Similarity Search With Satellite System Graph: Efficiency, Scalability, and Unindexed Query CompatibilityabstractApproximate nearest neighbor search (ANNS) in high-dimensional space is essential in database and information retrieval. Recently, there has been a surge of interest in exploring efficient graph-based indices for the ANNS problem. Among them, navigating spreading-out graph (NSG) provides fine theoretical analysis and achieves state-of-the-art performance. However, we find there are several limitations with NSG: 1) NSG has no theoretical guarantee on nearest neighbor search when the query is not indexed in the database; and 2) NSG is too sparse which harms the search performance. In addition, NSG suffers from high indexing complexity. To address above problems, we propose the satellite system graphs (SSG) and a practical variant NSSG. Specifically, we propose a novel pruning strategy to produce SSGs from the complete graph. SSGs define a new family of MSNETs in which the out-edges of each node are distributed evenly in all directions. Each node in the graph builds effective connections to its neighborhood omnidirectionally, whereupon we derive SSG's excellent theoretical properties for both indexed and unindexed queries. We can adaptively adjust the sparsity of an SSG with a hyper-parameter to optimize the search performance. Further, NSSG is proposed to reduce the indexing complexity of the SSG for large-scale applications. Both theoretical and extensive experimental analysis are provided to demonstrate the strengths of the proposed approach over the existing representative algorithms. Our code has been released at https://github.com/ZJULearning/SSG. Cong Fu 0001, Changxu Wang, Deng Cai 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | Progressive Transfer Learning
Zhengxu Yu, Dong Shen 0003, Zhongming Jin 0001, Jianqiang Huang 0001, Deng Cai 0001, Xian-Sheng Hua 0001 |
IEEE Trans. Image Process. | 5 |
| 2022 | Toward Better Accuracy-Efficiency Trade-Offs: Divide and Co-TrainingabstractThe width of a neural network matters since increasing the width will necessarily increase the model capacity. However, the performance of a network does not improve linearly with the width and soon gets saturated. In this case, we argue that increasing the number of networks (ensemble) can achieve better accuracy-efficiency trade-offs than purely increasing the width. To prove it, one large network is divided into several small ones regarding its parameters and regularization components. Each of these small networks has a fraction of the original one's parameters. We then train these small networks together and make them see various views of the same data to increase their diversity. During this co-training process, networks can also learn from each other. As a result, small networks can achieve better ensemble performance than the large one with few or no extra parameters or FLOPs, i. e., achieving better accuracy-efficiency trade-offs. Small networks can also achieve faster inference speed than the large one by concurrent running. All of the above shows that the number of networks is a new dimension of model scaling. We validate our argument with 8 different neural architectures on common benchmarks through extensive experiments. Shuai Zhao 0006, Liguang Zhou, Wenxiao Wang 0001, Deng Cai 0001, Tin Lun Lam, Yangsheng Xu |
IEEE Trans. Image Process. | 4 |
| 2022 | LookCom: Learning Optimal Network for Community DetectionabstractCommunity detection is one of the fundamental tasks in graph mining, which aims to identify group assignment of nodes in a complex network. Recently, network embedding techniques have demonstrated their strong power in advancing the community detection task and achieve better performance than various traditional methods. Despite their empirical success, most of the existing algorithms directly leverage the observed coarse network structure for community detection. Therefore, they often lead to suboptimal performance as the observed connections fail to capture the essential tie strength information among nodes precisely and account for the impact of noisy links. In this paper, an optimal network structure for community detection is introduced to characterize the fine-grained tie strength information between connected nodes and alleviate the adverse effects of noisy links. To obtain an expressive node representation for community detection, we learn the optimal network structure and network embeddings in a joint framework, instead of using a two-stage approach to derive the node embeddings from the coarse network topology. In particular, we formulate the joint framework as an optimization problem and an alternating optimization algorithm is exploited to solve the proposed optimization problem. Additionally, theoretical analyses regarding the computational complexity and the convergence of the optimization algorithm are also provided. Extensive experiments on both synthetic and real-world networks demonstrate the effectiveness and superiority of the proposed framework. Yixiang Dong, Minnan Luo, Jundong Li, Deng Cai 0001 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2022 | Apparel-Invariant Feature Learning for Person Re-IdentificationabstractWith the rise of deep learning methods, person Re-Identification (ReID) performance has been improved tremendously in many public datasets. However, most public ReID datasets are collected in a short time window in which persons’ appearance rarely changes. In real-world applications such as in a shopping mall, the same person may change their wearings, and different persons may wear similar apparel. It reveals a critical problem that current ReID models heavily rely on a person’s apparel, resulting in an inconsistent ReID performance. Therefore, it is crucial to learn an apparel-invariant person representation under clothes changing or several persons wearing similar clothes cases. In this work, we tackle this problem from the viewpoint of invariant feature representation learning. The main contributions of this work are as follows. (1) We propose the semi-supervised Apparel-invariant Feature Learning (AIFL) framework to learn an apparel-invariant pedestrian representation using images of the same person wearing different clothes. (2) To obtain images of the same person wearing different clothes, we propose an unsupervised apparel-simulation GAN (AS-GAN) to synthesize cloth-changing images according to the target cloth embedding. It is worth noting that the images used in ReID tasks were cropped from real-world low-quality CCTV videos, making it more challenging to synthesize cloth-changing images. Extensive experiments demonstrate that our proposal can improve the ReID performance of the baseline models. Zhengxu Yu, Yilun Zhao 0001, Zhongming Jin 0001, Jianqiang Huang 0001, Deng Cai 0001, Xian-Sheng Hua 0001 |
IEEE Trans. Multim. | 6 |
| 2021 | RESA: Recurrent Feature-Shift Aggregator for Lane DetectionabstractLane detection is one of the most important tasks in self-driving. Due to various complex scenarios (e.g., severe occlusion, ambiguous lanes, etc.) and the sparse supervisory signals inherent in lane annotations, lane detection task is still challenging. Thus, it is difficult for the ordinary convolutional neural network (CNN) to train in general scenes to catch subtle lane feature from the raw image. In this paper, we present a novel module named REcurrent Feature-Shift Aggregator (RESA) to enrich lane feature after preliminary feature extraction with an ordinary CNN. RESA takes advantage of strong shape priors of lanes and captures spatial relationships of pixels across rows and columns. It shifts sliced feature map recurrently in vertical and horizontal directions and enables each pixel to gather global information. RESA can conjecture lanes accurately in challenging scenarios with weak appearance clues by aggregating sliced feature map. Moreover, we propose a Bilateral Up-Sampling Decoder that combines coarse-grained and fine-detailed features in the up-sampling stage. It can recover the low-resolution feature map into pixel-wise prediction meticulously. Our method achieves state-of-the-art results on two popular lane detection benchmarks (CULane and Tusimple). Code has been made available at: https://github.com/ZJULearning/resa. Tu Zheng, Wenjian Tang, Zheng Yang 0008, Haifeng Liu 0001, Deng Cai 0001 |
AAAI | 7 |
| 2021 | EasyTransfer: A Simple and Scalable Deep Transfer Learning Platform for NLP ApplicationsabstractThe literature has witnessed the success of leveraging Pre-trained Language Models (PLMs) and Transfer Learning (TL) algorithms to a wide range of Natural Language Processing (NLP) applications, yet it is not easy to build an easy-to-use and scalable TL toolkit for this purpose. To bridge this gap, the EasyTransfer platform is designed to develop deep TL algorithms for NLP applications. EasyTransfer is backended with a high-performance and scalable engine for efficient training and inference, and also integrates comprehensive deep TL algorithms, to make the development of industrial-scale TL applications easier. In EasyTransfer, the built-in data and model parallelism strategies, combined with AI compiler optimization, show to be 4.0x faster than the community version of distributed training. EasyTransfer supports various NLP models in the ModelZoo, including mainstream PLMs and multi-modality models. It also features various in-house developed TL algorithms, together with the AppZoo for NLP applications. The toolkit is convenient for users to quickly start model training, evaluation, and online deployment. EasyTransfer is currently deployed at Alibaba to support a variety of business scenarios, including item recommendation, personalized search, conversational question answering, etc. Extensive experiments on real-world datasets and online applications show that EasyTransfer is suitable for online production with cutting-edge performance for various applications. The source code of EasyTransfer is released at Github1. Minghui Qiu, Peng Li 0056, Chengyu Wang 0001, Haojie Pan, Ang Wang, Cen Chen 0001, Xianyan Jia, Yaliang Li, Jun Huang 0007, Deng Cai 0001, Wei Lin 0016 |
CIKM | 10 |
| 2021 | Salient Object Ranking with Position-Preserved AttentionabstractInstance segmentation can detect where the objects are in an image, but hard to understand the relationship between them. We pay attention to a typical relationship, relative saliency. A closely related task, salient object detection, predicts a binary map highlighting a visually salient region while hard to distinguish multiple objects. Directly combining two tasks by post-processing also leads to poor performance. There is a lack of research on relative saliency at present, limiting the practical applications such as content-aware image cropping, video summary, and image labeling.In this paper, we study the Salient Object Ranking (SOR) task, which manages to assign a ranking order of each detected object according to its visual saliency. We propose the first end-to-end framework of the SOR task and solve it in a multi-task learning fashion. The framework handles instance segmentation and salient object ranking simultaneously. In this framework, the SOR branch is independent and flexible to cooperate with different detection methods, so that easy to use as a plugin. We also intro-duce a Position-Preserved Attention (PPA) module tailored for the SOR branch. It consists of the position embedding stage and feature interaction stage. Considering the importance of position in saliency comparison, we preserve absolute coordinates of objects in ROI pooling operation and then fuse positional information with semantic features in the first stage. In the feature interaction stage, we apply the attention mechanism to obtain proposals’ contextualized representations to predict their relative ranking orders. Extensive experiments have been conducted on the ASR dataset. Without bells and whistles, our proposed method outperforms the former state-of-the-art method significantly. The code will be released publicly available on https://github.com/EricFH/SOR. Daoxin Zhang, Minghao Chen 0001, Yao Hu 0002, Deng Cai 0001, Xiaofei He 0001 |
ICCV | 7 |
| 2021 | Accelerate CNNs from Three Dimensions: A Comprehensive Pruning FrameworkabstractMost neural network pruning methods, such as filter-level and layer-level prunings, prune the network model along one dimension (depth, width, or resolution) solely to meet a computational budget. However, such a pruning policy often leads to excessive reduction of that dimension, thus inducing a huge accuracy loss. To alleviate this issue, we argue that pruning should be conducted along three dimensions comprehensively. For this purpose, our pruning framework formulates pruning as an optimization problem. Specifically, it first casts the relationships between a certain model’s accuracy and depth/width/resolution into a polynomial regression and then maximizes the polynomial to acquire the optimal values for the three dimensions. Finally, the model is pruned along the three optimal dimensions accordingly. In this framework, since collecting too much data for training the regression is very time-costly, we propose two approaches to lower the cost: 1) specializing the polynomial to ensure an accurate regression even with less training data; 2) employing iterative pruning and fine-tuning to collect the data faster. Extensive experiments show that our proposed algorithm surpasses state-of-the-art pruning algorithms and even neural architecture search-based algorithms. Wenxiao Wang 0001, Minghao Chen 0001, Shuai Zhao 0006, Long Chen 0016, Jinming Hu, Haifeng Liu 0001, Deng Cai 0001, Xiaofei He 0001, Wei Liu 0005 |
ICML | 7 |
| 2021 | MeLL: Large-scale Extensible User Intent Classification for Dialogue Systems with Meta Lifelong LearningabstractUser intent detection is vital for understanding their demands in dialogue systems. Although the User Intent Classification (UIC) task has been widely studied, for large-scale industrial applications, the task is still challenging. This is because user inputs in distinct domains may have different text distributions and target intent sets. When the underlying application evolves, new UIC tasks continuously emerge in a large quantity. Hence, it is crucial to develop a framework for large-scale extensible UIC that continuously fits new tasks and avoids catastrophic forgetting with an acceptable parameter growth rate. In this paper, we introduce the Meta Lifelong Learning (MeLL) framework to address this task. In MeLL, a BERT-based text encoder is employed to learn robust text representations across tasks, which is slowly updated for lifelong learning. We design global and local memory networks to capture the cross-task prototype representations of different classes, treated as the meta-learner quickly adapted to different tasks. Additionally, the Least Recently Used replacement policy is applied to manage the global memory such that the model size does not explode through time. Finally, each UIC task has its own task-specific output layer, with the attentive summarization of various features. We have conducted extensive experiments on both open-source and real industry datasets. Results show that MeLL improves the performance compared with strong baselines and also reduces the number of total parameters. We have also deployed MeLL on a real-world e-commerce dialogue system AliMe and observed significant improvements in terms of both F1 and the resources usage. Chengyu Wang 0001, Haojie Pan, Minghui Qiu, Jun Huang 0007, Haiqing Chen, Wei Lin 0016, Deng Cai 0001 |
KDD | 10 |
| 2021 | TopNet: Learning from Neural Topic Model to Generate Long StoriesabstractLong story generation (LSG) is one of the coveted goals in natural language processing. Different from most text generation tasks, LSG requires to output a long story of rich content based on a much shorter text input, and often suffers from information sparsity. In this paper, we propose TopNet to alleviate this problem, by leveraging the recent advances in neural topic modeling to obtain high-quality skeleton words to complement the short input. In particular, instead of directly generating a story, we first learn to map the short text input to a low-dimensional topic distribution (which is pre-assigned by a topic model). Based on this latent topic distribution, we can use the reconstruction decoder of the topic model to sample a sequence of inter-related words as a skeleton for the story. Experiments on two benchmark datasets show that our proposed framework is highly effective in skeleton word selection and significantly outperforms the state-of-the-art models in both automatic evaluation and human evaluation. Yazheng Yang, Boyuan Pan, Deng Cai 0001, Huan Sun 0001 |
KDD | 3 |
| 2021 | Do Wider Neural Networks Really Help Adversarial Robustness?abstractAdversarial training is a powerful type of defense against adversarial examples. Previous empirical results suggest that adversarial training requires wider networks for better performances. However, it remains elusive how does neural network width affect model robustness. In this paper, we carefully examine the relationship between network width and model robustness. Specifically, we show that the model robustness is closely related to the tradeoff between natural accuracy and perturbation stability, which is controlled by the robust regularization parameter λ. With the same λ, wider networks can achieve better natural accuracy but worse perturbation stability, leading to a potentially worse overall model robustness. To understand the origin of this phenomenon, we further relate the perturbation stability with the network's local Lipschitzness. By leveraging recent results on neural tangent kernels, we theoretically show that wider networks tend to have worse perturbation stability. Our analyses suggest that: 1) the common strategy of first fine-tuning λ on small networks and then directly use it for wide model training could lead to deteriorated model robustness; 2) one needs to properly enlarge λ to unleash the robustness potential of wider models fully. Finally, we propose a new Width Adjusted Regularization (WAR) method that adaptively enlarges λ on wide models and significantly saves the tuning time. Boxi Wu 0001, Deng Cai 0001, Xiaofei He 0001, Quanquan Gu |
NeurIPS | 3 |
| 2021 | Learning to Caricature via Semantic Shape TransformabstractAbstract Caricature is an artistic drawing created to abstract or exaggerate facial features of a person. Rendering visually pleasing caricatures is a difficult task that requires professional skills, and thus it is of great interest to design a method to automatically generate such drawings. To deal with large shape changes, we propose an algorithm based on a semantic shape transform to produce diverse and plausible shape exaggerations. Specifically, we predict pixel-wise semantic correspondences and perform image warping on the input photo to achieve dense shape transformation. We show that the proposed framework is able to render visually pleasing shape exaggerations while maintaining their facial structures. In addition, our model allows users to manipulate the shape via the semantic map. We demonstrate the effectiveness of our approach on a large photograph-caricature benchmark dataset with comparisons to the state-of-the-art methods. Wenqing Chu, Wei-Chih Hung, Yi-Hsuan Tsai, Yu-Ting Chang, Yijun Li 0001, Deng Cai 0001, Ming-Hsuan Yang 0001 |
Int. J. Comput. Vis. | 6 |
| 2021 | TTFNeXt for real-time object detection
Tu Zheng, Zheng Yang 0008, Haifeng Liu 0001, Deng Cai 0001 |
Neurocomputing | 6 |
| 2021 | Self-supervised attention flow for dialogue state tracking
Boyuan Pan, Yazheng Yang, Bo Li 0026, Deng Cai 0001 |
Neurocomputing | 4 |
| 2021 | COP: customized correlation-based Filter level pruning method for deep CNN compression
Wenxiao Wang 0001, Zhengxu Yu, Cong Fu 0001, Deng Cai 0001, Xiaofei He 0001 |
Neurocomputing | 4 |
| 2021 | AdaDB: An adaptive gradient method with data-dependent bound
Deng Cai 0001 |
Neurocomputing | 2 |
| 2021 | Complementary Pseudo Labels for Unsupervised Domain Adaptation On Person Re-IdentificationabstractIn recent years, supervised person re-identification (re-ID) models have received increasing studies. However, these models trained on the source domain always suffer dramatic performance drop when tested on an unseen domain. Existing methods are primary to use pseudo labels to alleviate this problem. One of the most successful approaches predicts neighbors of each unlabeled image and then uses them to train the model. Although the predicted neighbors are credible, they always miss some hard positive samples, which may hinder the model from discovering important discriminative information of the unlabeled domain. In this paper, to complement these low recall neighbor pseudo labels, we propose a joint learning framework to learn better feature embeddings via high precision neighbor pseudo labels and high recall group pseudo labels. The group pseudo labels are generated by transitively merging neighbors of different samples into a group to achieve higher recall. However, the merging operation may cause subgroups in the group due to imperfect neighbor predictions. To utilize these group pseudo labels properly, we propose using a similarity-aggregating loss to mitigate the influence of these subgroups by pulling the input sample towards the most similar embeddings. Extensive experiments on three large-scale datasets demonstrate that our method can achieve state-of-the-art performance under the unsupervised domain adaptation re-ID setting. Minghao Chen 0001, Jinming Hu, Dong Shen 0003, Haifeng Liu 0001, Deng Cai 0001 |
IEEE Trans. Image Process. | 6 |
| 2021 | ES-Net: Erasing Salient Parts to Learn More in Re-IdentificationabstractAs an instance-level recognition problem, re-identification (re-ID) requires models to capture diverse features. However, with continuous training, re-ID models pay more and more attention to the salient areas. As a result, the model may only focus on few small regions with salient representations and ignore other important information. This phenomenon leads to inferior performance, especially when models are evaluated on small inter-identity variation data. In this paper, we propose a novel network, Erasing-Salient Net (ES-Net), to learn comprehensive features by erasing the salient areas in an image. ES-Net proposes a novel method to locate the salient areas by the confidence of objects and erases them efficiently in a training batch. Meanwhile, to mitigate the over-erasing problem, this paper uses a trainable pooling layer P-pooling that generalizes global max and global average pooling. Experiments are conducted on two specific re-identification tasks (i.e., Person re-ID, Vehicle re-ID). Our ES-Net outperforms state-of-the-art methods on three Person re-ID benchmarks and two Vehicle re-ID benchmarks. Specifically, mAP / Rank-1 rate: 88.6% / 95.7% on Market1501, 78.8% / 89.2% on DuckMTMC-reID, 57.3% / 80.9% on MSMT17, 81.9% / 97.0% on Veri-776, respectively. Rank-1 / Rank-5 rate: 83.6% / 96.9% on VehicleID (Small), 79.9% / 93.5% on VehicleID (Medium), 76.9% / 90.7% on VehicleID (Large), respectively. Moreover, the visualized salient areas show human-interpretable visual explanations for the ranking results. Dong Shen 0003, Shuai Zhao 0006, Jinming Hu, Deng Cai 0001, Xiaofei He 0001 |
IEEE Trans. Image Process. | 5 |
| 2021 | A Revisit of Hashing Algorithms for Approximate Nearest Neighbor SearchabstractApproximate Nearest Neighbor Search (ANNS) is a fundamental problem in many areas of machine learning and data mining. During the past decade, numerous hashing algorithms are proposed to solve this problem. Every proposed algorithm claims to outperform Locality Sensitive Hashing (LSH), which is the most popular hashing method. However, the evaluation of these hashing article was not thorough enough, and the claim should be re-examined. If implemented correctly, almost all the hashing methods will have their performance improved as the code length increases. However, many existing hashing article only report the performance with the code length shorter than 128. In this article, we carefully revisit the problem of search-with-a-hash-index and analyze the pros and cons of two popular hash index search procedures. Then we proposed a simple but effective novel hash index search approach and made a thorough comparison of eleven popular hashing algorithms. Surprisingly, the random-projection-based Locality Sensitive Hashing ranked the first, which is in contradiction to the claims in all the other 10 hashing article. Despite the extreme simplicity of random-projection-based LSH, our results show that the capability of this algorithm has been far underestimated. For the sake of reproducibility, all the codes used in the article are released on GitHub, which can be used as a testing platform for a fair comparison between various hashing algorithms. Deng Cai 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2020 | Adversarial-Learned Loss for Domain AdaptationabstractRecently, remarkable progress has been made in learning transferable representation across domains. Previous works in domain adaptation are majorly based on two techniques: domain-adversarial learning and self-training. However, domain-adversarial learning only aligns feature distributions between domains but does not consider whether the target features are discriminative. On the other hand, self-training utilizes the model predictions to enhance the discrimination of target features, but it is unable to explicitly align domain distributions. In order to combine the strengths of these two methods, we propose a novel method called Adversarial-Learned Loss for Domain Adaptation (ALDA). We first analyze the pseudo-label method, a typical self-training method. Nevertheless, there is a gap between pseudo-labels and the ground truth, which can cause incorrect training. Thus we introduce the confusion matrix, which is learned through an adversarial manner in ALDA, to reduce the gap and align the feature distributions. Finally, a new loss function is auto-constructed from the learned confusion matrix, which serves as the loss for unlabeled target samples. Our ALDA outperforms state-of-the-art approaches in four standard domain adaptation datasets. Our code is available at https://github.com/ZJULearning/ALDA. Minghao Chen 0001, Shuai Zhao 0006, Haifeng Liu 0001, Deng Cai 0001 |
AAAI | 4 |
| 2020 | PI-RCNN: An Efficient Multi-Sensor 3D Object Detector with Point-Based Attentive Cont-Conv Fusion ModuleabstractLIDAR point clouds and RGB-images are both extremely essential for 3D object detection. So many state-of-the-art 3D detection algorithms dedicate in fusing these two types of data effectively. However, their fusion methods based on Bird's Eye View (BEV) or voxel format are not accurate. In this paper, we propose a novel fusion approach named Point-based Attentive Cont-conv Fusion(PACF) module, which fuses multi-sensor features directly on 3D points. Except for continuous convolution, we additionally add a Point-Pooling and an Attentive Aggregation to make the fused features more expressive. Moreover, based on the PACF module, we propose a 3D multi-sensor multi-task network called Pointcloud-Image RCNN(PI-RCNN as brief), which handles the image segmentation and 3D object detection tasks. PI-RCNN employs a segmentation sub-network to extract full-resolution semantic feature maps from images and then fuses the multi-sensor features via powerful PACF module. Beneficial from the effectiveness of the PACF module and the expressive semantic features from the segmentation module, PI-RCNN can improve much in 3D object detection. We demonstrate the effectiveness of the PACF module and PI-RCNN on the KITTI 3D Detection benchmark, and our method can achieve state-of-the-art on the metric of 3D AP. Liang Xie 0003, Chao Xiang, Zhengxu Yu, Zheng Yang 0008, Deng Cai 0001, Xiaofei He 0001 |
AAAI | 6 |
| 2020 | Question-Driven Purchasing Propensity Analysis for RecommendationabstractMerchants of e-commerce Websites expect recommender systems to entice more consumption which is highly correlated with the customers' purchasing propensity. However, most existing recommender systems focus on customers' general preference rather than purchasing propensity often governed by instant demands which we deem to be well conveyed by the questions asked by customers. A typical recommendation scenario is: Bob wants to buy a cell phone which can play the game PUBG. He is interested in HUAWEI P20 and asks “can PUBG run smoothly on this phone?” under it. Then our system will be triggered to recommend the most eligible cell phones to him. Intuitively, diverse user questions could probably be addressed in reviews written by other users who have similar concerns. To address this recommendation problem, we propose a novel Question-Driven Attentive Neural Network (QDANN) to assess the instant demands of questioners and the eligibility of products based on user generated reviews, and do recommendation accordingly. Without supervision, QDANN can well exploit reviews to achieve this goal. The attention mechanisms can be used to provide explanations for recommendations. We evaluate QDANN in three domains of Taobao. The results show the efficacy of our method and its superiority over baseline methods. Long Chen 0007, Ziyu Guan, Qibin Xu, Huan Sun 0001, Guangyue Lu, Deng Cai 0001 |
AAAI | 7 |
| 2020 | Training-Time-Friendly Network for Real-Time Object DetectionabstractModern object detectors can rarely achieve short training time, fast inference speed, and high accuracy at the same time. To strike a balance among them, we propose the Training-Time-Friendly Network (TTFNet). In this work, we start with light-head, single-stage, and anchor-free designs, which enable fast inference speed. Then, we focus on shortening training time. We notice that encoding more training samples from annotated boxes plays a similar role as increasing batch size, which helps enlarge the learning rate and accelerate the training process. To this end, we introduce a novel approach using Gaussian kernels to encode training samples. Besides, we design the initiative sample weights for better information utilization. Experiments on MS COCO show that our TTFNet has great advantages in balancing training time, inference speed, and accuracy. It has reduced training time by more than seven times compared to previous real-time detectors while maintaining state-of-the-art performances. In addition, our super-fast version of TTFNet-18 and TTFNet-53 can outperform SSD300 and YOLOv3 by less than one-tenth of their training time, respectively. The code has been made available at https://github.com/ZJULearning/ttfnet. Tu Zheng, Zheng Yang 0008, Haifeng Liu 0001, Deng Cai 0001 |
AAAI | 6 |
| 2020 | Adversarial Mutual Information for Text GenerationabstractRecent advances in maximizing mutual information (MI) between the source and target have demonstrated its effectiveness in text generation. However, previous works paid little attention to modeling the backward network of MI (i.e., dependency from the target to the source), which is crucial to the tightness of the variational information maximization lower bound. In this paper, we propose Adversarial Mutual Information (AMI): a text generation framework which is formed as a novel saddle point (min-max) optimization aiming to identify joint interactions between the source and target. Within this framework, the forward and backward networks are able to iteratively promote or demote each other’s generated instances by comparing the real and synthetic data distributions. We also develop a latent noise sampling strategy that leverages random variations at the high-level semantic space to enhance the long term dependency in the generation process. Extensive experiments based on different text generation tasks demonstrate that the proposed AMI framework can significantly outperform several strong baselines, and we also show that AMI has potential to lead to a tighter lower bound of maximum mutual information for the variational information maximization problem. Boyuan Pan, Yazheng Yang, Kaizhao Liang, Bhavya Kailkhura, Zhongming Jin 0001, Xian-Sheng Hua 0001, Deng Cai 0001, Bo Li 0026 |
ICML | 7 |
| 2020 | MaCAR: Urban Traffic Light Control via Active Multi-agent Communication and Action RectificationabstractUrban traffic light control is an important and challenging real-world problem. By regarding intersections as agents, most of the Reinforcement Learning (RL) based methods generate actions of agents independently. They can cause action conflict and result in overflow or road resource waste in adjacent intersections. Recently, some collaborative methods have alleviated the above problems by extending the observable surroundings of agents, which can be considered as inactive cross-agent communication methods. However, when agents act synchronously in these works, the perceived action value is biased and the information exchanged is insufficient. In this work, we propose a novel Multi-agent Communication and Action Rectification (MaCAR) framework. It enables active communication between agents by considering the impact of synchronous actions of agents. MaCAR consists of two parts: (1) an active Communication Agent Network (CAN) involving a Message Propagation Graph Neural Network (MPGNN); (2) a Traffic Forecasting Network (TFN) which learns to predict the traffic after agents' synchronous actions and the corresponding action values. By using predicted information, we mitigate the action value bias during training to help rectify agents' future actions. In experiments, we show that our proposal can outperforms state-of-the-art methods on both synthetic and real-world datasets. Zhengxu Yu, Shuxian Liang, Zhongming Jin 0001, Jianqiang Huang 0001, Deng Cai 0001, Xiaofei He 0001, Xian-Sheng Hua 0001 |
IJCAI | 6 |
| 2020 | Large Scale Abstractive Multi-Review Summarization (LSARS) via Aspect AlignmentabstractIn an active e-commerce environment, customers process a large number of reviews when deciding on whether to buy a product or not. Abstractive Multi-Review Summarization aims to assist users to efficiently consume the reviews that are the most relevant to them. We propose the first large-scale abstractive multi-review summarization dataset that leverages more than 17.9 billion raw reviews and uses novel aspect-alignment techniques based on aspect annotations. Furthermore, we demonstrate that one can generate higher-quality review summaries by using a novel aspect-alignment-based model. Results from both automatic and human evaluation show that the proposed dataset plus the innovative aspect-alignment model can generate high-quality and trustful review summaries. Haojie Pan, Rongqin Yang, Rui Wang 0005, Deng Cai 0001, Xiaozhong Liu 0001 |
SIGIR | 5 |
| 2020 | Bi-Decoder Augmented Network for Neural Machine Translation
Boyuan Pan, Yazheng Yang, Zhou Zhao 0001, Yueting Zhuang, Deng Cai 0001 |
Neurocomputing | 5 |
| 2020 | Decouple co-adaptation: Classifier randomization for person re-identification
Zhenyong Wei, Zhongming Jin 0001, Jianqiang Huang 0001, Xian-Sheng Hua 0001, Deng Cai 0001, Xiaofei He 0001 |
Neurocomputing | 7 |
| 2020 | Video Dialog via Multi-Grained Convolutional Self-Attention Context Multi-Modal NetworksabstractVideo dialog is a new and challenging task, which requires an AI agent to maintain a meaningful dialog with humans in natural language about video contents. Specifically, given a video, a dialog history and a new question about the video, the agent has to combine video information with dialog history to infer the answer. However, the existing methods of image dialog and video question answering, which fail to process the complexity of video information and establish the logical dependency of history contexts, are inappropriate to be applied directly to video dialog. In this paper, we propose a novel approach for video dialog called multi-grained convolutional self-attention context network, which combines video information with dialog history. Instead of using RNN to encode the sequence information, we design a multi-grained convolutional self-attention mechanism to capture both element and segment level interactions that contain multi-grained sequence information. Moreover, a hierarchical dialog history encoder is designed to learn the context-aware question representation. Finally, we establish two decoders in multiple-choice and open-ended forms respectively, which utilize different strategies to get the multi-model context-aware video representation and to generate human-like answers. We evaluate our method on two large-scale datasets. Due to the flexibility and parallelism of the new attention mechanism, our method can achieve higher time efficiency, and the extensive experiments also show the effectiveness of our method. Mao Gu, Zhou Zhao 0001, Weike Jin, Deng Cai 0001, Fei Wu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2020 | Moment Retrieval via Cross-Modal Interaction Networks With Query ReconstructionabstractMoment retrieval aims to localize the most relevant moment in an untrimmed video according to the given natural language query. Existing works often only focus on one aspect of this emerging task, such as the query representation learning, video context modeling or multi-modal fusion, thus fail to develop a comprehensive system for further performance improvement. In this paper, we introduce a novel Cross-Modal Interaction Network (CMIN) to consider multiple crucial factors for this challenging task, including the syntactic dependencies of natural language queries, long-range semantic dependencies in video context and the sufficient cross-modal interaction. Specifically, we devise a syntactic GCN to leverage the syntactic structure of queries for fine-grained representation learning and propose a multi-head self-attention to capture long-range semantic dependencies from video context. Next, we employ a multi-stage cross-modal interaction to explore the potential relations of video and query contents, and we also consider query reconstruction from the cross-modal representations of target moment as an auxiliary task to strengthen the cross-modal representations. The extensive experiments on ActivityNet Captions and TACoS demonstrate the effectiveness of our proposed method. Zhijie Lin 0001, Zhou Zhao 0001, Zijian Zhang 0002, Deng Cai 0001 |
IEEE Trans. Image Process. | 5 |
| 2020 | SIF: Self-Inspirited Feature Learning for Person Re-IdentificationabstractThe re-identification (ReID) task has received increasing studies in recent years and its performance has gained significant improvement. The progress mainly comes from searching for new network structures to learn person representations. Most of these networks are trained using the classic stochastic gradient descent optimizer. However, limited efforts have been made to explore potential performance of existing ReID networks directly by better training scheme, which leaves a large space for ReID research. In this paper, we propose a Self-Inspirited Feature Learning (SIF) method to enhance performance of given ReID networks from the viewpoint of optimization. We design a simple adversarial learning scheme to encourage a network to learn more discriminative person representation. In our method, an auxiliary branch is added into the network only in the training stage, while the structure of the original network stays unchanged during the testing stage. In summary, SIF has three aspects of advantages: (1) it is designed under general setting; (2) it is compatible with many existing feature learning networks on the ReID task; (3) it is easy to implement and has steady performance. We evaluate the performance of SIF on three public ReID datasets: Market1501, DuckMTMC-reID, and CUHK03(both labeled and detected). The results demonstrate significant improvement in performance brought by SIF. We also apply SIF to obtain state-of-the-art results on all the three datasets. Specifically, mAP / Rank-1 accuracy are: 87.6% / 95.2% (without re-rank) on Market1501, 79.4% / 89.8% on DuckMTMC-reID, 77.0% / 79.5% on CUHK03 (labeled) and 73.9% / 76.6% on CUHK03 (detected), respectively. The code of SIF will be available soon. Zhenyong Wei, Zhongming Jin 0001, Zhengxu Yu, Jianqiang Huang 0001, Deng Cai 0001, Xiaofei He 0001, Xian-Sheng Hua 0001 |
IEEE Trans. Image Process. | 6 |
| 2020 | Query-Biased Self-Attentive Network for Query-Focused Video SummarizationabstractThis paper addresses the task of query-focused video summarization, which takes user queries and long videos as inputs and generates query-focused video summaries. Compared to video summarization, which mainly concentrates on finding the most diverse and representative visual contents as a summary, the task of query-focused video summarization considers the user's intent and the semantic meaning of generated summary. In this paper, we propose a method, named query-biased self-attentive network (QSAN) to tackle this challenge. Our key idea is to utilize the semantic information from video descriptions to generate a generic summary and then to combine the information from the query to generate a query-focused summary. Specifically, we first propose a hierarchical self-attentive network to model the relative relationship at three levels, which are different frames from a segment, different segments of the same video, textual information of video description and its related visual contents. We train the model on video caption dataset and employ a reinforced caption generator to generate a video description, which can help us locate important frames or shots. Then we build a query-aware scoring module to compute the query-relevant score for each shot and generate the query-focused summary. Extensive experiments on the benchmark dataset demonstrate the competitive performance of our approach compared to some methods. Shuwen Xiao, Zhou Zhao 0001, Zijian Zhang 0002, Ziyu Guan, Deng Cai 0001 |
IEEE Trans. Image Process. | 5 |
| 2020 | Addressing the Item Cold-Start Problem by Attribute-Driven Active LearningabstractIn recommender systems, cold-start issues are situations where no previous events, e.g., ratings, are known for certain users or items. In this paper, we focus on the item cold-start problem. Both content information (e.g., item attributes) and initial user ratings are valuable for seizing users' preferences on a new item. However, previous methods for the item cold-start problem either (1) incorporate content information into collaborative filtering to perform hybrid recommendation, or (2) actively select users to rate the new item without considering content information and then do collaborative filtering. In this paper, we propose a novel recommendation scheme for the item cold-start problem by leveraging both active learning and items' attribute information. Specifically, we design useful user selection criteria based on items' attributes and users' rating history, and combine the criteria in an optimization framework for selecting users. By exploiting the feedback ratings, users' previous ratings and items' attributes, we then generate accurate rating predictions for the other unselected users. Experimental results on two real-world datasets show the superiority of our proposed method over traditional methods. Yu Zhu 0007, Jinghao Lin, Shibi He, Beidou Wang, Ziyu Guan, Haifeng Liu 0001, Deng Cai 0001 |
IEEE Trans. Knowl. Data Eng. | 7 |
| 2019 | Reinforced Dynamic Reasoning for Conversational Question GenerationabstractThis paper investigates a new task named Conversational Question Generation (CQG) which is to generate a question based on a passage and a conversation history (i.e., previous turns of question-answer pairs).CQG is a crucial task for developing intelligent agents that can drive question-answering style conversations or test user understanding of a given passage.Towards that end, we propose a new approach named Reinforced Dynamic Reasoning (ReDR) network, which is based on the general encoder-decoder framework but incorporates a reasoning procedure in a dynamic manner to better understand what has been asked and what to ask next about the passage.To encourage producing meaningful questions, we leverage a popular question answering (QA) model to provide feedback and fine-tune the question generator using a reinforcement learning mechanism.Empirical results on the recently released CoQA dataset demonstrate the effectiveness of our method in comparison with various baselines and model variants.Moreover, to show the applicability of our method, we also apply it to create multiturn question-answering conversations for passages in SQuAD. * Work done while visiting the Ohio State University.Shelly is in second grade.She is a new student at her school.Shelly's family has lived in many different places.Shelly was born in Florida.Her family moved to Tennessee when she was two years old.When she was four years old, they moved to Texas.They moved from there to Arizona, where they now live.Q1: What grade is Shelly in ?A1: second R1: Shelly is in second grade.Q2: Was she a Boyuan Pan, Hao Li 0009, Ziyu Yao 0002, Deng Cai 0001, Huan Sun 0001 |
ACL (1) | 4 |
| 2019 | Cross-domain Attention Network with Wasserstein Regularizers for E-commerce SearchabstractProduct search and recommendation is a task that every e-commerce platform wants to outperform their peels on. However, training a good search or recommendation model often requires more data than what many platforms have. Fortunately, the search tasks on different platforms share the common underlying structure. Considering each platform as a domain, we propose a cross-domain learning approach to help the task on data-deficient platforms by leveraging the data from data-abundant platforms. In our solution, the importance of features in different domains is addressed by a domain-specific attention network. Meanwhile, a multi-task regularizer based on Wasserstein distance is introduced to help extract both domain-invariant and domain-specific features. Our model consistently outperforms the competing methods on both public and real-world industry datasets. Quantitative evaluation shows that our model can discover important features for different domains, which helps us better understand different user needs across platforms. Last but not least, we have deployed our model online in three big e-commerce platforms namely Taobao, Tmall, and Qintao, and observed better performance than the production models for all the platforms. Minghui Qiu, Cen Chen 0001, Xiaoyi Zeng, Jun Huang 0007, Deng Cai 0001, Jingren Zhou 0001, Forrest Sheng Bao |
CIKM | 6 |
| 2019 | Query-based Interactive Recommendation by Meta-Path and Adapted Attention-GRUabstractRecently, interactive recommender systems are becoming increasingly popular. The insight is that, with the interaction between users and the system, (1) users can actively intervene the recommendation results rather than passively receive them, and (2) the system learns more about users so as to provide better recommendation. Yu Zhu 0007, Qingwen Liu 0002, Yingcai Ma, Wenwu Ou, Junxiong Zhu, Beidou Wang, Ziyu Guan, Deng Cai 0001 |
CIKM | 9 |
| 2019 | Domain Adaptation for Semantic Segmentation With Maximum Squares LossabstractDeep neural networks for semantic segmentation always require a large number of samples with pixel-level labels, which becomes the major difficulty in their real-world applications. To reduce the labeling cost, unsupervised domain adaptation (UDA) approaches are proposed to transfer knowledge from labeled synthesized datasets to unlabeled real-world datasets. Recently, some semi-supervised learning methods have been applied to UDA and achieved state-of-the-art performance. One of the most popular approaches in semi-supervised learning is the entropy minimization method. However, when applying the entropy minimization to UDA for semantic segmentation, the gradient of the entropy is biased towards samples that are easy to transfer. To balance the gradient of well-classified target samples, we propose the maximum squares loss. Our maximum squares loss prevents the training process being dominated by easy-to-transfer samples in the target domain. Besides, we introduce the image-wise weighting ratio to alleviate the class imbalance in the unlabeled target domain. Both synthetic-to-real and cross-city adaptation experiments demonstrate the effectiveness of our proposed approach. The code is released at https://github. com/ZJULearning/MaxSquareLoss. Minghao Chen 0001, Hongyang Xue, Deng Cai 0001 |
ICCV | 3 |
| 2019 | Attribute Attention for Semantic Disambiguation in Zero-Shot LearningabstractZero-shot learning (ZSL) aims to accurately recognize unseen objects by learning mapping matrices that bridge the gap between visual information and semantic attributes. Previous works implicitly treat attributes equally in compatibility score while ignoring that they have different importance for discrimination, which leads to severe semantic ambiguity. Considering both low-level visual information and global class-level features that relate to this ambiguity, we propose a practical Latent Feature Guided Attribute Attention (LFGAA) framework to perform object-based attribute attention for semantic disambiguation. By distracting semantic activation in dimensions that cause ambiguity, our method outperforms existing state-of-the-art methods on AwA2, CUB and SUN datasets in both inductive and transductive settings. Yang Liu 0212, Jishun Guo, Deng Cai 0001, Xiaofei He 0001 |
ICCV | 3 |
| 2019 | Weakly-Supervised Caricature Face Parsing Through Domain AdaptationabstractA caricature is an artistic form of a person's picture in which certain striking characteristics are abstracted or exaggerated in order to create a humor or sarcasm effect. For numerous caricature related applications such as attribute recognition and caricature editing, face parsing is an essential pre-processing step that provides a complete facial structure understanding. However, current state-of-the-art face parsing methods require large amounts of labeled data on the pixel-level and such process for caricature is tedious and labor-intensive. For real photos, there are numerous labeled datasets for face parsing. Thus, we formulate caricature face parsing as a domain adaptation problem, where real photos play the role of the source domain, adapting to the target caricatures. Specifically, we first leverage a spatial transformer based network to enable shape domain shifts. A feed-forward style transfer network is then utilized to capture texture-level domain gaps. With these two steps, we synthesize face caricatures from real photos, and thus we can use parsing ground truths of the original photos to learn the parsing model. Experimental results on the synthetic and real caricatures demonstrate the effectiveness of the proposed domain adaptation algorithm. Code is available at: https://github.com/ZJULearning/CariFaceParsing. Wenqing Chu, Wei-Chih Hung, Yi-Hsuan Tsai, Deng Cai 0001, Ming-Hsuan Yang 0001 |
ICIP | 4 |
| 2019 | COP: Customized Deep Model Compression via Regularized Correlation-Based Filter-Level PruningabstractNeural network compression empowers the effective yet unwieldy deep convolutional neural networks (CNN) to be deployed in resource-constrained scenarios. Most state-of-the-art approaches prune the model in filter-level according to the "importance" of filters. Despite their success, we notice they suffer from at least two of the following problems: 1) The redundancy among filters is not considered because the importance is evaluated independently. 2) Cross-layer filter comparison is unachievable since the importance is defined locally within each layer. Consequently, we must manually specify layer-wise pruning ratios. 3) They are prone to generate sub-optimal solutions because they neglect the inequality between reducing parameters and reducing computational cost. Reducing the same number of parameters in different positions in the network may reduce different computational cost. To address the above problems, we develop a novel algorithm named as COP (correlation-based pruning), which can detect the redundant filters efficiently. We enable the cross-layer filter comparison through global normalization. We add parameter-quantity and computational-cost regularization terms to the importance, which enables the users to customize the compression according to their preference (smaller or faster). Extensive experiments have shown COP outperforms the others significantly. The code is released at https://github.com/ZJULearning/COP. Wenxiao Wang 0001, Cong Fu 0001, Jishun Guo, Deng Cai 0001, Xiaofei He 0001 |
IJCAI | 4 |
| 2019 | Progressive Transfer Learning for Person Re-identificationabstractModel fine-tuning is a widely used transfer learning approach in person Re-identification (ReID) applications, which fine-tuning a pre-trained feature extraction model into the target scenario instead of training a model from scratch. It is challenging due to the significant variations inside the target scenario, e.g., different camera viewpoint, illumination changes, and occlusion. These variations result in a gap between the distribution of each mini-batch and the distribution of the whole dataset when using mini-batch training. In this paper, we study model fine-tuning from the perspective of the aggregation and utilization of the global information of the dataset when using mini-batch training. Specifically, we introduce a novel network structure called Batch-related Convolutional Cell (BConv-Cell), which progressively collects the global information of the dataset into a latent state and uses this latent state to rectify the extracted feature. Based on BConv-Cells, we further proposed the Progressive Transfer Learning (PTL) method to facilitate the model fine-tuning process by joint training the BConv-Cells and the pre-trained ReID model. Empirical experiments show that our proposal can improve the performance of the ReID model greatly on MSMT17, Market-1501, CUHK03 and DukeMTMC-reID datasets. The code will be released later on at \url{https://github.com/ZJULearning/PTL} Zhengxu Yu, Zhongming Jin 0001, Jishun Guo, Jianqiang Huang 0001, Deng Cai 0001, Xiaofei He 0001, Xian-Sheng Hua 0001 |
IJCAI | 6 |
| 2019 | Localizing Unseen Activities in Video via Image QueryabstractAction localization in untrimmed videos is an important topic in the field of video understanding. However, existing action localization methods are restricted to a pre-defined set of actions and cannot localize unseen activities. Thus, we consider a new task to localize unseen activities in videos via image queries, named Image-Based Activity Localization. This task faces three inherent challenges: (1) how to eliminate the influence of semantically inessential contents in image queries; (2) how to deal with the fuzzy localization of inaccurate image queries; (3) how to determine the precise boundaries of target segments. We then propose a novel self-attention interaction localizer to retrieve unseen activities in an end-to-end fashion. Specifically, we first devise a region self-attention method with relative position encoding to learn fine-grained image region representations. Then, we employ a local transformer encoder to build multi-step fusion and reasoning of image and video contents. We next adopt an order-sensitive localizer to directly retrieve the target segment. Furthermore, we construct a new dataset ActivityIBAL by reorganizing the ActivityNet dataset. The extensive experiments show the effectiveness of our method. Zhou Zhao 0001, Zhijie Lin 0001, Jingkuan Song, Deng Cai 0001 |
IJCAI | 5 |
| 2019 | AtSNE: Efficient and Robust Visualization on GPU through Hierarchical OptimizationabstractVisualization of high-dimensional data is a fundamental yet challenging problem in data mining. These visualization techniques are commonly used to reveal the patterns in the high-dimensional data, such as clusters and the similarity among clusters. Recently, some successful visualization tools (e.g., BH-t-SNE and LargeVis) have been developed. However, there are two limitations with them : (1) they cannot capture the global data structure well. Thus, their visualization results are sensitive to initialization, which may cause confusions to the data analysis. (2) They cannot scale to large-scale datasets. They are not suitable to be implemented on the GPU platform because their complex algorithm logic, high memory cost, and random memory access mode will lead to low hardware utilization. To address the aforementioned problems, we propose a novel visualization approach named as Anchor-t-SNE (AtSNE), which provides efficient GPU-based visualization solution for large-scale and high-dimensional data. Specifically, we generate a number of anchor points from the original data and regard them as the skeleton of the layout, which holds the global structure information. We propose a hierarchical optimization approach to optimize the positions of the anchor points and ordinary data points in the layout simultaneously. Our approach presents much better and robust visual effects on 11 public datasets, and achieve 5 to 28 times speed-up on different datasets, compared with the current state-of-the-art methods. In particular, we deliver a high-quality 2-D layout for a 20 million and 96-dimension dataset within 5 hours, while the current methods fail to give results due to running out of the memory. Cong Fu 0001, Deng Cai 0001, Xiang Ren 0001 |
KDD | 3 |
| 2019 | A Minimax Game for Instance based Selective Transfer LearningabstractDeep neural network based transfer learning has been widely used to leverage information from the domain with rich data to help domain with insufficient data. When the source data distribution is different from the target data, transferring knowledge between these domains may lead to negative transfer. To mitigate this problem, a typical way is to select useful source domain data for transferring. However, limited studies focus on selecting high-quality source data to help neural network based transfer learning. To bridge this gap, we propose a general Minimax Game based model for selective Transfer Learning (MGTL). More specifically, we build a selector, a discriminator and a TL module in the proposed method. The discriminator aims to maximize the differences between selected source data and target data, while the selector acts as an attacker to selected source data that are close to the target to minimize the differences. The TL module trains on the selected data and provides rewards to guide the selector. Those three modules play a minimax game to help select useful source data for transferring. Our method is also shown to speed up the training process of the learning task in the target domain than traditional TL methods. To the best of our knowledge, this is the first to build a minimax game based model for selective transfer learning. To examine the generality of our method, we evaluate it on two different tasks: item recommendation and text retrieval. Extensive experiments over both public and real-world datasets demonstrate that our model outperforms the competing methods by a large margin. Meanwhile, the quantitative evaluation shows our model can select data which are close to target data. Our model is also deployed in a real-world system and significant improvement over the baselines is observed. Minghui Qiu, Xisen Wang, Yaliang Li, Xiaoyi Zeng, Jun Huang 0007, Bo Zheng 0007, Deng Cai 0001, Jingren Zhou 0001 |
KDD | 9 |
| 2019 | Region Mutual Information Loss for Semantic SegmentationabstractSemantic segmentation is a fundamental problem in computer vision. It is considered as a pixel-wise classification problem in practice, and most segmentation models use a pixel-wise loss as their optimization criterion. However, the pixel-wise loss ignores the dependencies between pixels in an image. Several ways to exploit the relationship between pixels have been investigated, \eg, conditional random fields (CRF) and pixel affinity based methods. Nevertheless, these methods usually require additional model branches, large extra memories, or more inference time. In this paper, we develop a region mutual information (RMI) loss to model the dependencies among pixels more simply and efficiently. In contrast to the pixel-wise loss which treats the pixels as independent samples, RMI uses one pixel and its neighbour pixels to represent this pixel. Then for each pixel in an image, we get a multi-dimensional point that encodes the relationship between pixels, and the image is cast into a multi-dimensional distribution of these high-dimensional points. The prediction and ground truth thus can achieve high order consistency through maximizing the mutual information (MI) between their multi-dimensional distributions. Moreover, as the actual value of the MI is hard to calculate, we derive a lower bound of the MI and maximize the lower bound to maximize the real value of the MI. RMI only requires a few extra computational resources in the training stage, and there is no overhead during testing. Experimental results demonstrate that RMI can achieve substantial and consistent improvements in performance on PASCAL VOC 2012 and CamVid datasets. The code is available at \url{https://github.com/ZJULearning/RMI}. Shuai Zhao 0006, Yang Wang 0030, Zheng Yang 0008, Deng Cai 0001 |
NeurIPS | 4 |
| 2019 | Scaling Up Sparse Support Vector Machines by Simultaneous Feature and Sample ReductionabstractSparse support vector machine (SVM) is a popular classification technique that can simultaneously learn a small set of the most interpretable features and identify the support vectors. It has achieved great successes in many real-world applications. However, for large-scale problems involving a huge number of samples and ultra-high dimensional features, solving sparse SVMs remains challenging. By noting that sparse SVMs induce sparsities in both feature and sample spaces, we propose a novel approach, which is based on accurate estimations of the primal and dual optima of sparse SVMs, to simultaneously identify the inactive features and samples that are guaranteed to be irrelevant to the outputs. Thus, we can remove the identified inactive samples and features from the training phase, leading to substantial savings in the computational cost without sacrificing the accuracy. Moreover, we show that our method can be extended to multi-class sparse support vector machines. To the best of our knowledge, the proposed method is the first static feature and sample reduction method for sparse SVMs and multi-class sparse SVMs. Experiments on both synthetic and real data sets demonstrate that our approach significantly outperforms state-of-the-art methods and the speedup gained by our approach can be orders of magnitude. Wei Liu 0005, Jieping Ye, Deng Cai 0001, Xiaofei He 0001, Jie Wang 0005 |
J. Mach. Learn. Res. | 5 |
| 2019 | Fast Approximate Nearest Neighbor Search With The Navigating Spreading-out GraphabstractApproximate nearest neighbor search (ANNS) is a fundamental problem in databases and data mining. A scalable ANNS algorithm should be both memory-efficient and fast. Some early graph-based approaches have shown attractive theoretical guarantees on search time complexity, but they all suffer from the problem of high indexing time complexity. Recently, some graph-based methods have been proposed to reduce indexing complexity by approximating the traditional graphs; these methods have achieved revolutionary performance on million-scale datasets. Yet, they still can not scale to billion-node databases. In this paper, to further improve the search-efficiency and scalability of graph-based methods, we start by introducing four aspects: (1) ensuring the connectivity of the graph; (2) lowering the average out-degree of the graph for fast traversal; (3) shortening the search path; and (4) reducing the index size. Then, we propose a novel graph structure called Monotonic Relative Neighborhood Graph (MRNG) which guarantees very low search complexity (close to logarithmic time). To further lower the indexing complexity and make it practical for billion-node ANNS problems, we propose a novel graph structure named Navigating Spreading-out Graph (NSG) by approximating the MRNG. The NSG takes the four aspects into account simultaneously. Extensive experiments show that NSG outperforms all the existing algorithms significantly. In addition, NSG shows superior performance in the E-commercial scenario of Taobao (Alibaba Group) and has been integrated into their billion-scale search engine. Cong Fu 0001, Chao Xiang, Changxu Wang, Deng Cai 0001 |
Proc. VLDB Endow. | 4 |
| 2019 | On the Diversity of Conditional Image Synthesis With Semantic LayoutsabstractMany image processing tasks can be formulated as translating images between two image domains, such as colorization, super-resolution, and conditional image synthesis. In most of these tasks, an input image may correspond to multiple outputs. However, current existing approaches only show minor stochasticity of the outputs. In this paper, we present a novel approach to synthesize diverse realistic images corresponding to a semantic layout. We introduce a diversity loss objective, which maximizes the distance between synthesized image pairs and relates the input noise to the semantic segments in the synthesized images. Thus, our approach can not only produce multiple diverse images but also allow users to manipulate the output images by adjusting the noise manually. Experimental results show that images synthesized by our approach are diverse than that of the current existing works and equipping our diversity loss does not degrade the reality of the base networks. Moreover, our approach can be applied to unpaired datasets. Zichen Yang, Haifeng Liu 0001, Deng Cai 0001 |
IEEE Trans. Image Process. | 3 |
| 2019 | Multi-Turn Video Question Answering via Hierarchical Attention Context Reinforced NetworksabstractMulti-turn video question answering is a challenging task in visual information retrieval, which generates the accurate answer from the referenced video contents according to the visual conversation context and given question. However, the existing visual question answering methods mainly tackle the problem of single-turn video question answering, which may be ineffectively applied for multi-turn video question answering directly, due to the insufficiency of modeling the sequential conversation context. In this paper, we study the problem of multi-turn video question answering from the viewpoint of multi-stream hierarchical attention context reinforced network learning. We first propose the hierarchical attention context network for context-aware question understanding by modeling the hierarchically sequential conversation context structure. We then develop the multi-stream spatio-temporal attention network for learning the joint representation of the dynamic video contents and context-aware question embedding. We next devise a multi-step reasoning process to enhance the multi-stream hierarchical attention context network learning method. We finally predict the multiple-choice answer from the candidate answer set and further develop the reinforced decoder network to generate the open-ended natural language answer for multi-turn video question answering. We construct two large-scale multi-turn video question answering datasets. The extensive experiments show the effectiveness of our method. Zhou Zhao 0001, Xinghua Jiang, Deng Cai 0001 |
IEEE Trans. Image Process. | 4 |
| 2019 | Long-Form Video Question Answering via Dynamic Hierarchical Reinforced NetworksabstractOpen-ended long-form video question answering is a challenging task in visual information retrieval, which automatically generates a natural language answer from the referenced long-form video contents according to a given question. However, the existing works mainly focus on short-form video question answering, due to the lack of modeling semantic representations from long-form video contents. In this paper, we introduce a dynamic hierarchical reinforced network for open-ended long-form video question answering, which employs an encoder-decoder architecture with a dynamic hierarchical encoder and a reinforced decoder. Concretely, we first propose a frame-level dynamic long-short term memory (LSTM) network with binary segmentation gate to learn frame-level semantic representations according to the given question. We then develop a segment-level highway LSTM network with a question-aware highway gate for segment-level semantic modeling. Furthermore, we devise the reinforced decoder with a hierarchical attention mechanism to generate natural language answers. We construct a large-scale long-form video question answering dataset. The extensive experiments on the long-form dataset and another public short-form dataset show the effectiveness of our method. Zhou Zhao 0001, Shuwen Xiao, Zhenxin Xiao, Jun Yu 0002, Deng Cai 0001, Fei Wu 0001 |
IEEE Trans. Image Process. | 7 |
| 2019 | Sparse Coding Guided Spatiotemporal Feature Learning for Abnormal Event Detection in Large VideosabstractAbnormal event detection in large videos is an important task in research and industrial applications, which has attracted considerable attention in recent years. Existing methods usually solve this problem by extracting local features and then learning an outlier detection model on training videos. However, most previous approaches merely employ hand-crafted visual features, which is a clear disadvantage due to their limited representation capacity. In this paper, we present a novel unsupervised deep feature learning algorithm for the abnormal event detection problem. To exploit the spatiotemporal information of the inputs, we utilize the deep three-dimensional convolutional network (C3D) to perform feature extraction. Then, the key problem is how to train the C3D network without any category labels. Here, we employ the sparse coding results of the hand-crafted features generated from the inputs to guide the unsupervised feature learning. Specifically, we define a multilevel similarity relationship between these inputs according to the statistical information of the shared atoms. In the following, we introduce the quadruplet concept to model the multilevel similarity structure, which could be used to construct a generalized triplet loss for training the C3D network. Furthermore, the C3D network could be utilized to generate the features for sparse coding again, and this pipeline could be iterated for several times. By jointly optimizing between the sparse coding and the unsupervised feature learning, we can obtain robust and rich feature representations. Based on the learned representations, the sparse reconstruction error is applied to predicting the anomaly score of each testing input. Experiments on several publicly available video surveillance datasets in comparison with a number of existing works demonstrate that the proposed algorithm performs favorably against the state-of-the-art methods. Wenqing Chu, Hongyang Xue, Chengwei Yao, Deng Cai 0001 |
IEEE Trans. Multim. | 4 |
| 2018 | PixelLink: Detecting Scene Text via Instance SegmentationabstractMost state-of-the-art scene text detection algorithms are deep learning based methods that depend on bounding box regression and perform at least two kinds of predictions: text/non-text classification and location regression. Regression plays a key role in the acquisition of bounding boxes in these methods, but it is not indispensable because text/non-text prediction can also be considered as a kind of semantic segmentation that contains full location information in itself. However, text instances in scene images often lie very close to each other, making them very difficult to separate via semantic segmentation. Therefore, instance segmentation is needed to address this problem. In this paper, PixelLink, a novel scene text detection algorithm based on instance segmentation, is proposed. Text instances are first segmented out by linking pixels within the same instance together. Text bounding boxes are then extracted directly from the segmentation result without location regression. Experiments show that, compared with regression-based methods, PixelLink can achieve better or comparable performance on several benchmarks, while requiring many fewer training iterations and less training data. Dan Deng, Haifeng Liu 0001, Xuelong Li 0001, Deng Cai 0001 |
AAAI | 4 |
| 2018 | Discourse Marker Augmented Network with Reinforcement Learning for Natural Language InferenceabstractNatural Language Inference (NLI), also known as Recognizing Textual Entailment (RTE), is one of the most important problems in natural language processing.It requires to infer the logical relationship between two given sentences.While current approaches mostly focus on the interaction architectures of the sentences, in this paper, we propose to transfer knowledge from some important discourse markers to augment the quality of the NLI model.We observe that people usually use some discourse markers such as "so" or "but" to represent the logical relationship between two sentences.These words potentially have deep connections with the meanings of the sentences, thus can be utilized to help improve the representations of them.Moreover, we use reinforcement learning to optimize a new objective function with a reward defined by the property of the NLI datasets to make full use of the labels information.Experiments show that our method achieves the state-of-the-art performance on several large-scale datasets. Boyuan Pan, Yazheng Yang, Zhou Zhao 0001, Yueting Zhuang, Deng Cai 0001, Xiaofei He 0001 |
ACL (1) | 5 |
| 2018 | Device-Aware Rule Recommendation for the Internet of ThingsabstractWith over 34 billion IoT devices to be installed by 2020, the Internet of Things (IoT) is fundamentally changing our lives. One of the greatest benefits of the IoT is the powerful automations achieved by applying rules to IoT devices. For instance, a rule named "Make me a cup of coffee when I wake up'' automatically turns on the coffee machine when the sensor in the bedroom detects motion in the morning. With large numbers of possible rules out there, a recommendation system is of great necessity to help users find rules they need. However, little effort has been made to design a model tailored for the IoT rule recommendation, which comes with lots of new challenges compared with traditional recommendation tasks. We not only need to re-define "users'' and "items'' in the recommendation task, but also have to consider a new type of entities, devices, and the extra information and constraints brought by them. To handle these challenges, we propose a novel efficient recommendation algorithm, which not only considers the implicit feedback of users on rules, but also takes user-rule-device interactions and the match between rule device requirements and user device possessions into account. In collaboration with Samsung, one of the leading companies in this field, we have designed an IoT rule recommendation framework and evaluated our algorithm on a real-life industry dataset. Experiments show the effectiveness and efficiency of our method. Beidou Wang, Xin Guo 0006, Martin Ester, Ziyu Guan, Bandeep Singh, Yu Zhu 0007, Jiajun Bu, Deng Cai 0001 |
CIKM | 8 |
| 2018 | Translating Embeddings for Knowledge Graph Completion with Relation Attention MechanismabstractKnowledge graph embedding is an essential problem in knowledge extraction. Recently, translation based embedding models (e.g., TransE) have received increasingly attentions. These methods try to interpret the relations among entities as translations from head entity to tail entity and achieve promising performance on knowledge graph completion. Previous researchers attempt to transform the entity embedding concerning the given relation for distinguishability. Also, they naturally think the relation-related transforming should reflect attention mechanism, which means it should focus on only a part of the attributes. However, we found previous methods are failed with creating attention mechanism, and the reason is that they ignore the hierarchical routine of human cognition. When predicting whether a relation holds between two entities, people first check the category of entities, then they focus on fined-grained relation-related attributes to make the decision. In other words, the attention should take effect on entities filtered by the right category. In this paper, we propose a novel knowledge graph embedding method named TransAt to learn the translation based embedding, relation-related categories of entities and relation-related attention simultaneously. Extensive experiments show that our approach outperforms state-of-the-art methods significantly on public datasets, and our method can learn the true attention varying among relations. Wei Qian 0003, Cong Fu 0001, Yu Zhu 0007, Deng Cai 0001, Xiaofei He 0001 |
IJCAI | 4 |
| 2018 | Multi-Turn Video Question Answering via Multi-Stream Hierarchical Attention Context NetworkabstractConversational video question answering is a challenging task in visual information retrieval, which generates the accurate answer from the referenced video contents according to the visual conversation context and given question. However, the existing visual question answering methods mainly tackle the problem of single-turn video question answering, which may be ineffectively applied for multi-turn video question answering directly, due to the insufficiency of modeling the sequential conversation context. In this paper, we study the problem of multi-turn video question answering from the viewpoint of multi-step hierarchical attention context network learning. We first propose the hierarchical attention context network for context-aware question understanding by modeling the hierarchically sequential conversation context structure. We then develop the multi-stream spatio-temporal attention network for learning the joint representation of the dynamic video contents and context-aware question embedding. We next devise the hierarchical attention context network learning method with multi-step reasoning process for multi-turn video question answering. We construct two large-scale multi-turn video question answering datasets. The extensive experiments show the effectiveness of our method. Zhou Zhao 0001, Xinghua Jiang, Deng Cai 0001, Jun Xiao 0001, Xiaofei He 0001, Shiliang Pu |
IJCAI | 3 |
| 2018 | Attentional Image Retweet Modeling via Multi-Faceted Ranking Network LearningabstractRetweet prediction is a challenging problem in social media sites (SMS). In this paper, we study the problem of image retweet prediction in social media, which predicts the image sharing behavior that the user reposts the image tweets from their followees. Unlike previous studies, we learn user preference ranking model from their past retweeted image tweets in SMS. We first propose heterogeneous image retweet modeling network (IRM) that exploits users' past retweeted image tweets with associated contexts, their following relations in SMS and preference of their followees. We then develop a novel attentional multi-faceted ranking network learning framework with multi-modal neural networks for the proposed heterogenous IRM network to learn the joint image tweet representations and user preference representations for prediction task. The extensive experiments on a large-scale dataset from Twitter site shows that our method achieves better performance than other state-of-the-art solutions to the problem. Zhou Zhao 0001, Lingtao Meng, Jun Xiao 0001, Min Yang 0007, Fei Wu 0001, Deng Cai 0001, Xiaofei He 0001, Yueting Zhuang |
IJCAI | 6 |
| 2018 | Open-Ended Long-form Video Question Answering via Adaptive Hierarchical Reinforced NetworksabstractOpen-ended long-form video question answering is challenging problem in visual information retrieval, which automatically generates the natural language answer from the referenced long-form video content according to the question. However, the existing video question answering works mainly focus on the short-form video question answering, due to the lack of modeling the semantic representation of long-form video contents. In this paper, we consider the problem of long-form video question answering from the viewpoint of adaptive hierarchical reinforced encoder-decoder network learning. We propose the adaptive hierarchical encoder network to learn the joint representation of the long-form video contents according to the question with adaptive video segmentation. we then develop the reinforced decoder network to generate the natural language answer for open-ended video question answering. We construct a large-scale long-form video question answering dataset. The extensive experiments show the effectiveness of our method. Zhou Zhao 0001, Shuwen Xiao, Zhou Yu 0001, Jun Yu 0002, Deng Cai 0001, Fei Wu 0001, Yueting Zhuang |
IJCAI | 6 |
| 2018 | A Brand-level Ranking System with the Customized Attention-GRU ModelabstractIn e-commerce websites like Taobao, brand is playing a more important role in influencing users' decision of click/purchase, partly because users are now attaching more importance to the quality of products and brand is an indicator of quality. However, existing ranking systems are not specifically designed to satisfy this kind of demand. Some design tricks may partially alleviate this problem, but still cannot provide satisfactory results or may create additional interaction cost. In this paper, we design the first brand-level ranking system to address this problem. The key challenge of this system is how to sufficiently exploit users' rich behavior in e-commerce websites to rank the brands. In our solution, we firstly conduct the feature engineering specifically tailored for the personalized brand ranking problem and then rank the brands by an adapted Attention-GRU model containing three important modifications. Note that our proposed modifications can also apply to many other machine learning models on various tasks. We conduct a series of experiments to evaluate the effectiveness of our proposed ranking model and test the response to the brand-level ranking system from real users on a large-scale e-commerce platform, i.e. Taobao. Yu Zhu 0007, Junxiong Zhu, Yongliang Li, Beidou Wang, Ziyu Guan, Deng Cai 0001 |
IJCAI | 7 |
| 2018 | MacNet: Transferring Knowledge from Machine Comprehension to Sequence-to-Sequence ModelsabstractMachine Comprehension (MC) is one of the core problems in natural language processing, requiring both understanding of the natural language and knowledge about the world. Rapid progress has been made since the release of several benchmark datasets, and recently the state-of-the-art models even surpass human performance on the well-known SQuAD evaluation. In this paper, we transfer knowledge learned from machine comprehension to the sequence-to-sequence tasks to deepen the understanding of the text. We propose MacNet: a novel encoder-decoder supplementary architecture to the widely used attention-based sequence-to-sequence models. Experiments on neural machine translation (NMT) and abstractive text summarization show that our proposed framework can significantly improve the performance of the baseline models, and our method for the abstractive text summarization achieves the state-of-the-art results on the Gigaword dataset. Boyuan Pan, Yazheng Yang, Hao Li 0009, Zhou Zhao 0001, Yueting Zhuang, Deng Cai 0001, Xiaofei He 0001 |
NeurIPS | 6 |
| 2018 | Dialogue Act Recognition via CRF-Attentive Structured NetworkabstractDialogue Act Recognition (DAR) is a challenging problem in dialogue interpretation, which aims to associate semantic labels to utterances and characterize the speaker's intention. Currently, many existing approaches formulate the DAR problem ranging from multi-classification to structured prediction, which suffer from handcrafted feature extensions and attentive contextual dependencies. In this paper, we tackle the problem of DAR from the viewpoint of extending richer Conditional Random Field (CRF) structured dependencies without abandoning end-to-end training. We incorporate hierarchical semantic inference with memory mechanism on the utterance modeling at multiple levels. We then utilize the structured attention network on the linear-chain CRF to dynamically separate the utterances into cliques. The extensive experiments on two primary benchmark datasets Switchboard Dialogue Act (SWDA) and Meeting Recorder Dialogue Act (MRDA) datasets show that our method achieves better performance than other state-of-the-art solutions to the problem. Zheqian Chen, Rongqin Yang, Zhou Zhao 0001, Deng Cai 0001, Xiaofei He 0001 |
SIGIR | 4 |
| 2018 | Question retrieval for community-based question answering via heterogeneous social influential network
Zheqian Chen, Zhou Zhao 0001, Chengwei Yao, Deng Cai 0001 |
Neurocomputing | 5 |
| 2018 | Deep feature based contextual model for object detection
Wenqing Chu, Deng Cai 0001 |
Neurocomputing | 2 |
| 2018 | The forgettable-watcher model for video question answering
Wenqing Chu, Hongyang Xue, Zhou Zhao 0001, Deng Cai 0001, Chengwei Yao |
Neurocomputing | 4 |
| 2018 | Deep Rotation Equivariant Network
Junying Li, Zichen Yang, Haifeng Liu 0001, Deng Cai 0001 |
Neurocomputing | 4 |
| 2018 | Improving face recognition with domain adaptation
Ge Wen, Huaguan Chen, Deng Cai 0001, Xiaofei He 0001 |
Neurocomputing | 3 |
| 2018 | Split-Net: Improving face recognition in one forwarding operation
Ge Wen, Deng Cai 0001, Xiaofei He 0001 |
Neurocomputing | 3 |
| 2018 | Multi-label active learning based on submodular functionsabstractIn the data collection task, it is more expensive to annotate the instance in multi-label learning problem, since each instance is associated with multiple labels. Therefore it is more important to adopt active learning method in multi-label learning to reduce the labeling cost. Recent researches indicate submodular function optimization works well on subset selection problem and provides theoretical performance guarantees while simultaneously retaining extremely fast optimization. In this paper, we propose a query strategy by constructing a submodular function for the selected instance-label pairs, which can measure and combine the informativeness and representativeness. Thus the active learning problem can be formulated as a submodular function maximization problem, which can be solved efficiently and effectively by a simple greedy lazy algorithm. Experimental results show that the proposed approach outperforms several state-of-the-art multi-label active learning methods. Kuoliang Wu, Deng Cai 0001, Xiaofei He 0001 |
Neurocomputing | 2 |
| 2018 | Multi-Task Vehicle Detection With Region-of-Interest VotingabstractVehicle detection is a challenging problem in autonomous driving systems, due to its large structural and appearance variations. In this paper, we propose a novel vehicle detection scheme based on multi-task deep convolutional neural networks (CNNs) and region-of-interest (RoI) voting. In the design of CNN architecture, we enrich the supervised information with subcategory, region overlap, bounding-box regression, and category of each training RoI as a multi-task learning framework. This design allows the CNN model to share visual knowledge among different vehicle attributes simultaneously, and thus, detection robustness can be effectively improved. In addition, most existing methods consider each RoI independently, ignoring the clues from its neighboring RoIs. In our approach, we utilize the CNN model to predict the offset direction of each RoI boundary toward the corresponding ground truth. Then, each RoI can vote those suitable adjacent bounding boxes, which are consistent with this additional information. The voting results are combined with the score of each RoI itself to find a more accurate location from a large number of candidates. Experimental results on the real-world computer vision benchmarks KITTI and the PASCAL2007 vehicle data set show that our approach achieves superior performance in vehicle detection compared with other existing published works. Wenqing Chu, Yao Liu 0014, Chen Shen 0003, Deng Cai 0001, Xian-Sheng Hua 0001 |
IEEE Trans. Image Process. | 4 |
| 2018 | A Better Way to Attend: Attention With Trees for Video Question AnsweringabstractWe propose a new attention model for video question answering. The main idea of the attention models is to locate on the most informative parts of the visual data. The attention mechanisms are quite popular these days. However, most existing visual attention mechanisms regard the question as a whole. They ignore the word-level semantics where each word can have different attentions and some words need no attention. Neither do they consider the semantic structure of the sentences. Although the Extended Soft Attention (E-SA) model for video question answering leverages the word-level attention, it performs poorly on long question sentences. In this paper, we propose the heterogeneous tree-structured memory network (HTreeMN) for video question answering. Our proposed approach is based upon the syntax parse trees of the question sentences. The HTreeMN treats the words differently where the visual words are processed with an attention module and the verbal ones not. It also utilizes the semantic structure of the sentences by combining the neighbors based on the recursive structure of the parse trees. The understandings of the words and the videos are propagated and merged from leaves to the root. Furthermore, we build a hierarchical attention mechanism to distill the attended features. We evaluate our approach on two datasets. The experimental results show the superiority of our HTreeMN model over the other attention models especially on complex questions. Hongyang Xue, Wenqing Chu, Zhou Zhao 0001, Deng Cai 0001 |
IEEE Trans. Image Process. | 4 |
| 2018 | Identifying Genetic Risk Factors for Alzheimer's Disease via Shared Tree-Guided Feature Learning Across Multiple TasksabstractThe genome-wide association study (GWAS) is a popular approach to identify disease-associated genetic factors for Alzhemer's Disease (AD). However, it remains challenging because of the small number of samples, very high feature dimensionality and complex structures. To accurately identify genetic risk factors for AD, we propose a novel method based on an in-depth exploration of the hierarchical structure among the features and the commonality across related tasks. Specifically, we first extract and encode the tree hierarchy among features; then, we integrate the tree structures with multi-task feature learning (MTFL) to learn the shared features-that are predictive of AD-among related tasks simultaneously. Thus, we can unify the strength of both the prior structure information and MTFL to boost the prediction performance. However, due to the highly complex regularizer that encodes the tree structure and the extremely high feature dimensionality, the learning process can be computationally prohibitive. To address this, we further develop a novel safe screening rule to quickly identify and remove the irrelevant features before training. Experiment results demonstrate that the proposed approach significantly outperforms the state-of-the-art in detecting genetic risk factors of AD and the speedup gained by the proposed screening can be several orders of magnitude. Tingjin Luo, Jieping Ye, Deng Cai 0001, Xiaofei He 0001, Jie Wang 0005 |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2018 | Weakly-Supervised Deep Embedding for Product Review Sentiment AnalysisabstractProduct reviews are valuable for upcoming buyers in helping them make decisions. To this end, different opinion mining techniques have been proposed, where judging a review sentence's orientation (e.g., positive or negative) is one of their key challenges. Recently, deep learning has emerged as an effective means for solving sentiment classification problems. A neural network intrinsically learns a useful representation automatically without human efforts. However, the success of deep learning highly relies on the availability of large-scale training data. We propose a novel deep learning framework for product review sentiment classification which employs prevalently available ratings as weak supervision signals. The framework consists of two steps: (1) learning a high level representation (an embedding space) which captures the general sentiment distribution of sentences through rating information; and (2) adding a classification layer on top of the embedding layer and use labeled sentences for supervised fine-tuning. We explore two kinds of low level network structure for modeling review sentences, namely, convolutional feature extractors and long short-term memory. To evaluate the proposed framework, we construct a dataset containing 1.1M weakly labeled review sentences and 11,754 labeled review sentences from Amazon. Experimental results show the efficacy of the proposed framework and its superiority over baselines. Wei Zhao 0019, Ziyu Guan, Long Chen 0007, Xiaofei He 0001, Deng Cai 0001, Beidou Wang, Quan Wang 0006 |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2018 | Social-Aware Movie Recommendation via Multimodal Network LearningabstractWith the rapid development of Internet movie industry social-aware movie recommendation systems (SMRs) have become a popular online web service that provide relevant movie recommendations to users. In this effort many existing movie recommendation approaches learn a user ranking model from user feedback with respect to the movie's content. Unfortunately this approach suffers from the sparsity problem inherent in SMR data. In the present work we address the sparsity problem by learning a multimodal network representation for ranking movie recommendations. We develop a heterogeneous SMR network for movie recommendation that exploits the textual description and movie-poster image of each movie as well as user ratings and social relationships. With this multimodal data we then present a heterogeneous information network learning framework called SMR-multimodal network representation learning (MNRL) for movie recommendation. To learn a ranking metric from the heterogeneous information network we also developed a multimodal neural network model. We evaluated this model on a large-scale dataset from a real world SMR Web site and we find that SMR-MNRL achieves better performance than other state-of-the-art solutions to the problem. Zhou Zhao 0001, Hanqing Lu, Tim Weninger, Deng Cai 0001, Xiaofei He 0001, Yueting Zhuang |
IEEE Trans. Multim. | 5 |
| 2017 | Community-Based Question Answering via Asymmetric Multi-Faceted Ranking Network LearningabstractNowadays the community-based question answering (CQA) sites become the popular Internet-based web service, which have accumulated millions of questions and their posted answers over time. Thus, question answering becomes an essential problem in CQA sites, which ranks the high-quality answers to the given question. Currently, most of the existing works study the problem of question answering based on the deep semantic matching model to rank the answers based on their semantic relevance, while ignoring the authority of answerers to the given question. In this paper, we consider the problem of community-based question answering from the viewpoint of asymmetric multi-faceted ranking network embedding. We propose a novel asymmetric multi-faceted ranking network learning framework for community-based question answering by jointly exploiting the deep semantic relevance between question-answer pairs and the answerers' authority to the given question. We then develop an asymmetric ranking network learning method with deep recurrent neural networks by integrating both answers' relative quality rank to the given question and the answerers' following relations in CQA sites. The extensive experiments on a large-scale dataset from a real world CQA site show that our method achieves better performance than other state-of-the-art solutions to the problem. Zhou Zhao 0001, Hanqing Lu, Vincent Wenchen Zheng, Deng Cai 0001, Xiaofei He 0001, Yueting Zhuang |
AAAI | 4 |
| 2017 | Reverse Top-k Geo-Social Keyword Queries in Road NetworksabstractIdentifying prospective customers is an important aspect of marketing research. In this paper, we provide support for a new type of query, the Reverse Top-k Geo-Social Keyword (RkGSK) query. This query takes into account spatial, textual, and social information, and finds prospective customers for geotagged objects. As an example, a restaurant manager might apply the query to find prospective customers. To address this, we propose a hybrid index, the GIM-tree, which indexes locations, keywords, and social information of geo-tagged users and objects, and then, using the GIM-tree, we present efficient RkGSK query processing algorithms that exploit several pruning strategies. The effectiveness of RkGSK retrieval is characterized via a case study, and extensive experiments using real datasets offer insight into the efficiency of the proposed index and algorithms. Yunjun Gao, Gang Chen 0001, Christian S. Jensen, Deng Cai 0001 |
ICDE | 6 |
| 2017 | Scaling Up Sparse Support Vector Machines by Simultaneous Feature and Sample ReductionabstractSparse support vector machine (SVM) is a popular classification technique that can simultaneously learn a small set of the most interpretable features and identify the support vectors. It has achieved great successes in many real-world applications. However, for large-scale problems involving a huge number of samples and extremely high-dimensional features, solving sparse SVMs remains challenging. By noting that sparse SVMs induce sparsities in both feature and sample spaces, we propose a novel approach, which is based on accurate estimations of the primal and dual optima of sparse SVMs, to simultaneously identify the features and samples that are guaranteed to be irrelevant to the outputs. Thus, we can remove the identified inactive samples and features from the training phase, leading to substantial savings in both the memory usage and computational cost without sacrificing accuracy. To the best of our knowledge, the proposed method is the first static feature and sample reduction method for sparse SVMs. Experiments on both synthetic and real datasets (e.g., the kddb dataset with about 20 million samples and 30 million features) demonstrate that our approach significantly outperforms state-of-the-art methods and the speedup gained by our approach can be orders of magnitude. Wei Liu 0005, Jieping Ye, Deng Cai 0001, Xiaofei He 0001, Jie Wang 0005 |
ICML | 5 |
| 2017 | Stacked Similarity-Aware AutoencodersabstractAs one of the most popular unsupervised learning approaches, the autoencoder aims at transforming the inputs to the outputs with the least discrepancy. The conventional autoencoder and most of its variants only consider the one-to-one reconstruction, which ignores the intrinsic structure of the data and may lead to overfitting. In order to preserve the latent geometric information in the data, we propose the stacked similarity-aware autoencoders. To train each single autoencoder, we first obtain the pseudo class label of each sample by clustering the input features. Then the hidden codes of those samples sharing the same category label will be required to satisfy an additional similarity constraint. Specifically, the similarity constraint is implemented based on an extension of the recently proposed center loss. With this joint supervision of the autoencoder reconstruction error and the center loss, the learned feature representations not only can reconstruct the original data, but also preserve the geometric structure of the data. Furthermore, a stacked framework is introduced to boost the representation capacity. The experimental results on several benchmark datasets show the remarkable performance improvement of the proposed algorithm compared with other autoencoder based approaches. Wenqing Chu, Deng Cai 0001 |
IJCAI | 2 |
| 2017 | Robust Asymmetric Bayesian Adaptive Matrix FactorizationabstractLow rank matrix factorizations(LRMF) have attracted much attention due to its wide range of applications in computer vision, such as image impainting and video denoising. Most of the existing methods assume that the loss between an observed measurement matrix and its bilinear factorization follows symmetric distribution, like gaussian or gamma families. However, in real-world situations, this assumption is often found too idealized, because pictures under various illumination and angles may suffer from multi-peaks, asymmetric and irregular noises. To address these problems, this paper assumes that the loss follows a mixture of Asymmetric Laplace distributions and proposes robust Asymmetric Laplace Adaptive Matrix Factorization model(ALAMF) under bayesian matrix factorization framework. The assumption of Laplace distribution makes our model more robust and the asymmetric attribute makes our model more flexible and adaptable to real-world noise. A variational method is then devised for model inference. We compare ALAMF with other state-of-the-art matrix factorization methods both on data sets ranging from synthetic and real-world application. The experimental results demonstrate the effectiveness of our proposed approach. Xin Guo 0006, Boyuan Pan, Deng Cai 0001, Xiaofei He 0001 |
IJCAI | 3 |
| 2017 | Link Prediction via Ranking Metric Dual-Level Attention Network LearningabstractLink prediction is a challenging problem for complex network analysis, arising in many disciplines such as social networks and telecommunication networks. Currently, many existing approaches estimate the proximity of the link endpoints for link prediction from their feature or the local neighborhood around them, which suffer from the localized view of network connections and insufficiency of discriminative feature representation. In this paper, we consider the problem of link prediction from the viewpoint of learning discriminative path-based proximity ranking metric embedding. We propose a novel ranking metric network learning framework by jointly exploiting both node-level and path-level attentional proximity of the endpoints for link prediction. We then develop the path-based dual-level reasoning attentional learning method with recurrent neural network for proximity ranking metric embedding. The extensive experiments on two large-scale datasets show that our method achieves better performance than other state-of-the-art solutions to the problem. Zhou Zhao 0001, Ben Gao, Vincent Wenchen Zheng, Deng Cai 0001, Xiaofei He 0001, Yueting Zhuang |
IJCAI | 4 |
| 2017 | Microblog Sentiment Classification via Recurrent Random Walk Network LearningabstractMicroblog Sentiment Classification (MSC) is a challenging task in microblog mining, arising in many applications such as stock price prediction and crisis management. Currently, most of the existing approaches learn the user sentiment model from their posted tweets in microblogs, which suffer from the insufficiency of discriminative tweet representation. In this paper, we consider the problem of microblog sentiment classification from the viewpoint of heterogeneous MSC network embedding. We propose a novel recurrent random walk network learning framework for the problem by exploiting both users’ posted tweets and their social relations in microblogs. We then introduce the deep recurrent neural networks with random-walk layer for heterogeneous MSC network embedding, which can be trained end-to-end from the scratch. Weemploytheback-propagationmethodfortraining the proposed recurrent random walk network model. The extensive experiments on the large-scale public datasets from Twitter show that our method achieves better performance than other state-of-the-art solutions to the problem. Zhou Zhao 0001, Hanqing Lu, Deng Cai 0001, Xiaofei He 0001, Yueting Zhuang |
IJCAI | 3 |
| 2017 | Video Question Answering via Hierarchical Spatio-Temporal Attention NetworksabstractOpen-ended video question answering is a challenging problem in visual information retrieval, which automatically generates the natural language answer from the referenced video content according to the question. However, the existing visual question answering works only focus on the static image, which may be ineffectively applied to video question answering due to the temporal dynamics of video contents. In this paper, we consider the problem of open-ended video question answering from the viewpoint of spatio-temporal attentional encoder-decoder learning framework. We propose the hierarchical spatio-temporal attention network for learning the joint representation of the dynamic video contents according to the given question. We then develop the encoder-decoder learning method with reasoning recurrent neural networks for open-ended video question answering. We construct a large-scale video question answering dataset. The extensive experiments show the effectiveness of our method. Zhou Zhao 0001, Deng Cai 0001, Xiaofei He 0001, Yueting Zhuang |
IJCAI | 3 |
| 2017 | What to Do Next: Modeling User Behaviors by Time-LSTMabstractRecently, Recurrent Neural Network (RNN) solutions for recommender systems (RS) are becoming increasingly popular. The insight is that, there exist some intrinsic patterns in the sequence of users' actions, and RNN has been proved to perform excellently when modeling sequential data. In traditional tasks such as language modeling, RNN solutions usually only consider the sequential order of objects without the notion of interval. However, in RS, time intervals between users' actions are of significant importance in capturing the relations of users' actions and the traditional RNN architectures are not good at modeling them. In this paper, we propose a new LSTM variant, i.e. Time-LSTM, to model users' sequential actions. Time-LSTM equips LSTM with time gates to model time intervals. These time gates are specifically designed, so that compared to the traditional RNN solutions, Time-LSTM better captures both of users' short-term and long-term interests, so as to improve the recommendation performance. Experimental results on two real-world datasets show the superiority of the recommendation method using Time-LSTM over the traditional methods. Yu Zhu 0007, Hao Li 0009, Yikang Liao, Beidou Wang, Ziyu Guan, Haifeng Liu 0001, Deng Cai 0001 |
IJCAI | 7 |
| 2017 | Video Question Answering via Hierarchical Dual-Level Attention Network LearningabstractVideo question answering is a challenging task in visual information retrieval, which provides the accurate answer from the referenced video contents according to the given question. However, the existing visual question answering approaches mainly tackle the problem of static image question answering, which may be ineffectively applied for video question answering directly, due to the insufficiency of modeling the video temporal dynamics. In this paper, we study the problem of video question answering from the viewpoint of hierarchical dual-level attention network learning. We obtain the object appearance and movement information in the video based on both frame-level and segment-level feature representation methods. We then develop the hierarchical duallevel attention networks to learn the question-aware video representations with word-level and question-level attention mechanisms. We next devise the question-level fusion attention mechanism for our proposed networks to learn the questionaware joint video representation for video question answering. We construct two large-scale video question answering datasets. The extensive experiments validate the effectiveness of our method. Zhou Zhao 0001, Jinghao Lin, Xinghua Jiang, Deng Cai 0001, Xiaofei He 0001, Yueting Zhuang |
ACM Multimedia | 4 |
| 2017 | User Personalized Satisfaction Prediction via Multiple Instance Deep LearningabstractCommunity question answering(CQA) services have arisen as a popular knowledge sharing pattern for netizens. With abundant interactions among users, individuals are capable of obtaining satisfactory information. However, it is not effective for users to attain satisfying answers within minutes. Users have to check the progress over time until the appropriate answers submitted. We address this problem as a user personalized satisfaction prediction task. Existing methods usually exploit manual feature selection. It is not desirable as it requires careful design and is labor intensive. In this paper, we settle this issue by developing a new multiple instance deep learning framework. Specifically, in our settings, each question follows a multiple instance learning assumption, where its obtained answers can be regarded as instance sets in a bag and we define the question resolved with at least one satisfactory answer. We design an efficient framework exploiting multiple instance learning property with deep learning tactic to model the question-answer pairs relevance and rank the asker's satisfaction possibility. Extensive experiments on large-scale datasets from different forums of Stack Exchange demonstrate the feasibility of our proposed framework in predicting asker personalized satisfaction. Zheqian Chen, Ben Gao, Zhou Zhao 0001, Haifeng Liu 0001, Deng Cai 0001 |
WWW | 6 |
| 2017 | Special issue on selected and extended papers from the 2015 International Conference on Intelligence Science and Big Data Engineering (IScIDE 2015)
Shiguang Shan, Deng Cai 0001, Cheng Deng 0002, Hong Chang 0001 |
Neurocomputing | 2 |
| 2017 | Pre-training the deep generative models with adaptive hyperparameter optimization
Chengwei Yao, Deng Cai 0001, Jiajun Bu, Gencai Chen |
Neurocomputing | 2 |
| 2017 | Sparse Learning with Stochastic Composite OptimizationabstractIn this paper, we study Stochastic Composite Optimization (SCO) for sparse learning that aims to learn a sparse solution from a composite function. Most of the recent SCO algorithms have already reached the optimal expected convergence rate O(1/λT), but they often fail to deliver sparse solutions at the end either due to the limited sparsity regularization during stochastic optimization (SO) or due to the limitation in online-to-batch conversion. Even when the objective function is strongly convex, their high probability bounds can only attain O(√{log(1/δ)/T}) with δ is the failure probability, which is much worse than the expected convergence rate. To address these limitations, we propose a simple yet effective two-phase Stochastic Composite Optimization scheme by adding a novel powerful sparse online-to-batch conversion to the general Stochastic Optimization algorithms. We further develop three concrete algorithms, OptimalSL, LastSL and AverageSL, directly under our scheme to prove the effectiveness of the proposed scheme. Both the theoretical analysis and the experiment results show that our methods can really outperform the existing methods at the ability of sparse learning and at the meantime we can improve the high probability bound to approximately O(log(log(T)/δ)/λT). Lijun Zhang 0005, Zhongming Jin 0001, Rong Jin 0001, Deng Cai 0001, Xuelong Li 0001, Ronghua Liang, Xiaofei He 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2017 | Constrained Low-Rank Learning Using Least Squares-Based RegularizationabstractLow-rank learning has attracted much attention recently due to its efficacy in a rich variety of real-world tasks, e.g., subspace segmentation and image categorization. Most low-rank methods are incapable of capturing low-dimensional subspace for supervised learning tasks, e.g., classification and regression. This paper aims to learn both the discriminant low-rank representation (LRR) and the robust projecting subspace in a supervised manner. To achieve this goal, we cast the problem into a constrained rank minimization framework by adopting the least squares regularization. Naturally, the data label structure tends to resemble that of the corresponding low-dimensional representation, which is derived from the robust subspace projection of clean data by low-rank learning. Moreover, the low-dimensional representation of original data can be paired with some informative structure by imposing an appropriate constraint, e.g., Laplacian regularizer. Therefore, we propose a novel constrained LRR method. The objective function is formulated as a constrained nuclear norm minimization problem, which can be solved by the inexact augmented Lagrange multiplier algorithm. Extensive experiments on image classification, human pose estimation, and robust face recovery have confirmed the superiority of our method. Ping Li 0006, Jun Yu 0002, Meng Wang 0001, Deng Cai 0001, Xuelong Li 0001 |
IEEE Trans. Cybern. | 5 |
| 2017 | Depth Image Inpainting: Improving Low Rank Matrix Completion With Low Gradient RegularizationabstractWe address the task of single depth image inpainting. Without the corresponding color images, previous or next frames, depth image inpainting is quite challenging. One natural solution is to regard the image as a matrix and adopt the low rank regularization just as color image inpainting. However, the low rank assumption does not make full use of the properties of depth images. A shallow observation inspires us to penalize the nonzero gradients by sparse gradient regularization. However, statistics show that though most pixels have zero gradients, there is still a non-ignorable part of pixels, whose gradients are small but nonzero. Based on this property of depth images, we propose a low gradient regularization method in which we reduce the penalty for small gradients while penalizing the nonzero gradients to allow for gradual depth changes. The proposed low gradient regularization is integrated with the low rank regularization into the low rank low gradient approach for depth image inpainting. We compare our proposed low gradient regularization with the sparse gradient regularization. The experimental results show the effectiveness of our proposed approach. Hongyang Xue, Deng Cai 0001 |
IEEE Trans. Image Process. | 3 |
| 2017 | Unifying the Video and Question Attentions for Open-Ended Video Question AnsweringabstractVideo question answering is an important task toward scene understanding and visual data retrieval. However, current visual question answering works mainly focus on a single static image, which is distinct from the dynamic and sequential visual data in the real world. Their approaches cannot utilize the temporal information in videos. In this paper, we introduce the task of free-form open-ended video question answering. The open-ended answers enable wider applications compared with the common multiple-choice tasks in Visual-QA. We first propose a data set for open-ended Video-QA with the automatic question generation approaches. Then, we propose our sequential video attention and temporal question attention models. These two models apply the attention mechanism on videos and questions, while preserving the sequential and temporal structures of the guides. The two models are integrated into the model of unified attention. After the video and the question are encoded, the answers are generated wordwisely from our models by a decoder. In the end, we evaluate our models on the proposed data set. The experimental results demonstrate the effectiveness of our proposed model. Hongyang Xue, Zhou Zhao 0001, Deng Cai 0001 |
IEEE Trans. Image Process. | 3 |
| 2017 | On efficiently finding reverse k-nearest neighbors over uncertain graphs
Yunjun Gao, Xiaoye Miao, Gang Chen 0001, Baihua Zheng, Deng Cai 0001, Huiyong Cui |
VLDB J. | 5 |
| 2016 | Accelerated Sparse Linear Regression via Random ProjectionabstractIn this paper, we present an accelerated numerical method based on random projection for sparse linear regression. Previous studies have shown that under appropriate conditions, gradient-based methods enjoy a geometric convergence rate when applied to this problem. However, the time complexity of evaluating the gradient is as large as $\mathcal{O}(nd)$, where $n$ is the number of data points and $d$ is the dimensionality, making those methods inefficient for large-scale and high-dimensional dataset. To address this limitation, we first utilize random projection to find a rank-$k$ approximator for the data matrix, and reduce the cost of gradient evaluation to $\mathcal{O}(nk+dk)$, a significant improvement when $k$ is much smaller than $d$ and $n$. Then, we solve the sparse linear regression problem via a proximal gradient method with a homotopy strategy to generate sparse intermediate solutions. Theoretical analysis shows that our method also achieves a global geometric convergence rate, and moreover the sparsity of all the intermediate solutions are well-bounded over the iterations. Finally, we conduct experiments to demonstrate the efficiency of the proposed method. Lijun Zhang 0005, Rong Jin 0001, Deng Cai 0001, Xiaofei He 0001 |
AAAI | 4 |
| 2016 | Weakly-Supervised Deep Learning for Customer Review Sentiment Classification
Ziyu Guan, Long Chen 0007, Wei Zhao 0019, Shulong Tan, Deng Cai 0001 |
IJCAI | 6 |
| 2016 | Non-Negative Matrix Factorization with Sinkhorn Distance
Wei Qian 0003, Deng Cai 0001, Xiaofei He 0001, Xuelong Li 0001 |
IJCAI | 3 |
| 2016 | Expert Finding for Community-Based Question Answering via Ranking Metric Network Learning
Zhou Zhao 0001, Deng Cai 0001, Xiaofei He 0001, Yueting Zhuang |
IJCAI | 3 |
| 2016 | The Million Domain Challenge: Broadcast Email Prioritization by Cross-domain RecommendationabstractWith email overload becoming a billion-level drag on the economy, personalized email prioritization is of urgent need to help predict the importance level of an email. Despite lots of previous effort on the topic, broadcast email, an important type of emails with its unique challenges and intriguing opportunities, has been overlooked. The most salient opportunity lies in that effective collaborative filtering can be exploited due to thousands of receivers of a typical broadcast email. However, every broadcast email is completely cold and it is very costly to obtain users' preference feedback. Fortunately, there exist up to million-level broadcast mailing lists in a real life email system. Similar mailing lists can provide useful extra information for broadcast email prioritization in a target mailing list. How to mine such useful extra information is a challenging problem that has never been touched. In this work, we propose the first broadcast email prioritization framework considering large numbers of mailing lists by formulating this problem as a cross domain recommendation problem. An optimization framework is proposed to select the optimal set of source domains considering multiple criteria including overlap of users, feedback pattern similarity and coverage of users. Our method is thoroughly evaluated on a real world industrial dataset from Samsung Electronics and is proved highly effective and outperforms all the baselines. Beidou Wang, Martin Ester, Yikang Liao, Jiajun Bu, Yu Zhu 0007, Ziyu Guan, Deng Cai 0001 |
KDD | 7 |
| 2016 | Partial Multi-Modal Sparse Coding via Adaptive Similarity Structure RegularizationabstractMulti-modal sparse coding has played an important role in many multimedia applications, where data are usually with multiple modalities. Recently, various multi-modal sparse coding approaches have been proposed to learn sparse codes of multi-modal data, which assume that data appear in all modalities, or at least there is one modality containing all data. However, in real applications, it is often the case that some modalities of the data may suffer from missing information and thus result in partial multi-modality data. In this paper, we propose to solve the partial multi-modal sparse coding problem via multi-modal similarity structure regularization. Specifically, we propose a partial multi-modal sparse coding framework termed Adaptive Partial Multi-Modal Similarity Structure Regularization for Sparse Coding (AdaPM2SC), which preserves the similarity structure within the same modality and between different modalities. Experimental results conducted on two real-world datasets demonstrate that AdaPM2SC significantly outperforms the state-of-the-art methods under partial multi-modality scenario. Zhou Zhao 0001, Hanqing Lu, Deng Cai 0001, Xiaofei He 0001, Yueting Zhuang |
ACM Multimedia | 3 |
| 2016 | Distributed Representations of ExpertiseabstractCollaborative networks are common in real life, where domain experts work together to solve tasks issued by customers. How to model the proficiency of experts is critical for us to understand and optimize collaborative networks. Traditional expertise models, such as topic model based methods, cannot capture two aspects of human expertise simultaneously: Specialization (what area an expert is good at?) and Proficiency Level (to what degree?). In this paper, we propose new models to overcome this problem. We embed all historical task data in a lower dimension space and learn vector representations of expertise based on both solved and unsolved tasks. Specifically, in our first model, we assume that each expert will only handle tasks whose difficulty level just matches his/her proficiency level, while experts in the second model accept tasks whose levels are equal to or lower than his/her proficiency level. Experiments on real world datasets show that both models outperform topic model based approaches and standard classifiers such as logistic regression and support vector machine in terms of prediction accuracy. The learnt vector representations can be used to compare expertise in a large organization and optimize expert allocation. Fangqiu Han, Shulong Tan, Huan Sun 0001, Mudhakar Srivatsa, Deng Cai 0001, Xifeng Yan |
SDM | 5 |
| 2016 | Which to View: Personalized Prioritization for Broadcast EmailsabstractEmail is one of the most important communication tools today, but email overload resulting from the large number of unimportant or irrelevant emails is causing trillion-level economy loss every year. Thus personalized email prioritization algorithms are of urgent need. Despite lots of previous effort on this topic, broadcast email, an important type of email, is overlooked in previous literature. Broadcast emails are significantly different from normal emails, introducing both new challenges and opportunities. On one hand, lack of real senders and limited user interactions invalidate the key features exploited by traditional email prioritization algorithms; on the other hand, thousands of receivers for one broadcast email bring us the opportunity to predict importance through collaborative filtering. However, broadcast emails face a severe cold-start problem which hinders the direct application of collaborative filtering. In this paper, we propose the first framework for broadcast email prioritization by designing a novel active learning model that considers the collaborative filtering, implicit feedback and time sensitive responsiveness features of broadcast emails. Our method is thoroughly evaluated on a large scale real world industrial dataset from Samsung Electronics. Our method is proved highly effective and outperforms state-of-the-art personalized email prioritization methods. Beidou Wang, Martin Ester, Jiajun Bu, Yu Zhu 0007, Ziyu Guan, Deng Cai 0001 |
WWW | 6 |
| 2016 | Atom Decomposition Based Subgradient Descent for matrix classification
Wenqing Chu, Yao Hu 0002, Chen Zhao 0009, Haifeng Liu 0001, Deng Cai 0001 |
Neurocomputing | 5 |
| 2016 | Online robust principal component analysis via truncated nuclear norm regularization
Yao Hu 0002, Deng Cai 0001, Xiaofei He 0001 |
Neurocomputing | 4 |
| 2016 | Tracking people in RGBD videos using deep learning and motion clues
Hongyang Xue, Yao Liu 0014, Deng Cai 0001, Xiaofei He 0001 |
Neurocomputing | 3 |
| 2016 | A-Optimal Non-negative Projection with Hessian regularization
Zheng Yang 0008, Haifeng Liu 0001, Deng Cai 0001, Zhaohui Wu 0001 |
Neurocomputing | 3 |
| 2016 | Orthogonal Projective Sparse Coding for image representation
Wei Zhao 0019, Zheng Liu 0015, Ziyu Guan, Binbin Lin 0001, Deng Cai 0001 |
Neurocomputing | 5 |
| 2016 | Heterogeneous hypergraph embedding for document recommendation
Yu Zhu 0007, Ziyu Guan, Shulong Tan, Haifeng Liu 0001, Deng Cai 0001, Xiaofei He 0001 |
Neurocomputing | 5 |
| 2016 | Improving Collaborative Recommendation via User-Item SubgroupsabstractCollaborative filtering (CF) is out of question the most widely adopted and successful recommendation approach. A typical CF-based recommender system associates a user with a group of like-minded users based on their individual preferences over all the items, either explicit or implicit, and then recommends to the user some unobserved items enjoyed by the group. However, we find that two users with similar tastes on one item subset may have totally different tastes on another set. In other words, there exist many user-item subgroups each consisting of a subset of items and a group of like-minded users on these items. It is more reasonable to predict preferences through one user's correlated subgroups, but not the entire user-item matrix. In this paper, to find meaningful subgroups, we formulate a new Multiclass Co-Clustering (MCoC) model, which captures relations of user-to-item, user-to-user, and item-to-item simultaneously. Then, we combine traditional CF algorithms with subgroups for improving their top-$N$recommendation performance. Our approach can be seen as a new extension of traditional clustering CF models. Systematic experiments on several real data sets have demonstrated the effectiveness of our proposed approach. Jiajun Bu, Bin Xu 0005, Chun Chen 0001, Xiaofei He 0001, Deng Cai 0001 |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2016 | Graph Regularized Feature Selection with Data ReconstructionabstractFeature selection is a challenging problem for high dimensional data processing, which arises in many real applications such as data mining, information retrieval, and pattern recognition. In this paper, we study the problem of unsupervised feature selection. The problem is challenging due to the lack of label information to guide feature selection. We formulate the problem of unsupervised feature selection from the viewpoint of graph regularized data reconstruction. The underlying idea is that the selected features not only preserve the local structure of the original data space via graph regularization, but also approximately reconstruct each data point via linear combination. Therefore, the graph regularized data reconstruction error becomes a natural criterion for measuring the quality of the selected features. By minimizing the reconstruction error, we are able to select the features that best preserve both the similarity and discriminant information in the original data. We then develop an efficient gradient algorithm to solve the corresponding optimization problem. We evaluate the performance of our proposed algorithm on text clustering. The extensive experiments demonstrate the effectiveness of our proposed approach. Zhou Zhao 0001, Xiaofei He 0001, Deng Cai 0001, Lijun Zhang 0005, Wilfred Ng, Yueting Zhuang |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2016 | User Preference Learning for Online Social RecommendationabstractA social recommendation system has attracted a lot of attention recently in the research communities of information retrieval, machine learning, and data mining. Traditional social recommendation algorithms are often based on batch machine learning methods which suffer from several critical limitations, e.g., extremely expensive model retraining cost whenever new user ratings arrive, unable to capture the change of user preferences over time. Therefore, it is important to make social recommendation system suitable for real-world online applications where data often arrives sequentially and user preferences may change dynamically and rapidly. In this paper, we present a new framework of online social recommendation from the viewpoint of online graph regularized user preference learning (OGRPL), which incorporates both collaborative user-item relationship as well as item content features into an unified preference learning process. We further develop an efficient iterative procedure, OGRPL-FW which utilizes the Frank-Wolfe algorithm, to solve the proposed online optimization problem. We conduct extensive experiments on several large-scale datasets, in which the encouraging results demonstrate that the proposed algorithms obtain significantly lower errors (in terms of both RMSE and MAE) than the state-of-the-art online recommendation methods when receiving the same amount of training data in the online learning process. Zhou Zhao 0001, Hanqing Lu, Deng Cai 0001, Xiaofei He 0001, Yueting Zhuang |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2016 | Atom Decomposition with Adaptive Basis Selection Strategy for Matrix CompletionabstractEstimating missing entries in matrices has attracted much attention due to its wide range of applications like image inpainting and video denoising, which are usually considered as low-rank matrix completion problems theoretically. It is common to consider nuclear norm as a surrogate of the rank operator since it is the tightest convex lower bound of the rank operator under certain conditions. However, most approaches based on nuclear norm minimization involve a number of singular value decomposition (SVD) operations. Given a matrix X ∈ R m × n , the time complexity of the SVD operation is O ( mn 2 ), which brings prohibitive computational burden on large-scale matrices, limiting the further usage of these methods in real applications. Motivated by this observation, a series of atom-decomposition-based matrix completion methods have been studied. The key to these methods is to reconstruct the target matrix by pursuit methods in a greedy way, which only involves the computation of the top SVD and has great advantages in efficiency compared with the SVD-based matrix completion methods. However, due to gradually serious accumulation errors, atom-decomposition-based methods usually result in unsatisfactory reconstruction accuracy. In this article, we propose a new efficient and scalable atom decomposition algorithm for matrix completion called Adaptive Basis Selection Strategy ( ABSS ). Different from traditional greedy atom decomposition methods, a two-phase strategy is conducted to generate the basis separately via different strategies according to their different nature. At first, we globally prune the basis space to eliminate the unimportant basis as much as possible and locate the probable subspace containing the most informative basis. Then, another group of basis spaces are learned to improve the recovery accuracy based on local information. In this way, our proposed algorithm breaks through the accuracy bottleneck of traditional atom-decomposition-based matrix completion methods; meanwhile, it reserves the innate efficiency advantages over SVD-based matrix completion methods. We empirically evaluate the proposed algorithm ABSS on real visual image data and large-scale recommendation datasets. Results have shown that ABSS has much better reconstruction accuracy with comparable cost to atom-decomposition-based methods. At the same time, it outperforms the state-of-the-art SVD-based matrix completion algorithms by similar or better reconstruction accuracy with enormous advantages on efficiency. Yao Hu 0002, Chen Zhao 0013, Deng Cai 0001, Xiaofei He 0001, Xuelong Li 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2015 | Compressed Spectral Regression for Efficient Nonlinear Dimensionality Reduction
Deng Cai 0001 |
IJCAI | 1 |
| 2015 | Multi-view based multi-label propagation for image annotation
Zhanying He, Chun Chen 0001, Jiajun Bu, Ping Li 0006, Deng Cai 0001 |
Neurocomputing | 5 |
| 2015 | Unsupervised document summarization from data reconstruction perspective
Zhanying He, Chun Chen 0001, Jiajun Bu, Can Wang 0001, Lijun Zhang 0005, Deng Cai 0001, Xiaofei He 0001 |
Neurocomputing | 6 |
| 2015 | Large scale multi-class classification with truncated nuclear norm regularization
Yao Hu 0002, Zhongming Jin 0001, Debing Zhang, Deng Cai 0001, Xiaofei He 0001 |
Neurocomputing | 5 |
| 2015 | Opinions matter: a general approach to user profile modeling for contextual suggestion
Hongning Wang, Hui Fang 0001, Deng Cai 0001 |
Inf. Retr. J. | 4 |
| 2015 | Sparse fixed-rank representation for robust visual analysis
Ping Li 0006, Jiajun Bu, Bin Xu 0005, Zhanying He, Chun Chen 0001, Deng Cai 0001 |
Signal Process. | 6 |
| 2015 | Large Scale Spectral Clustering Via Landmark-Based Sparse RepresentationabstractSpectral clustering is one of the most popular clustering approaches. However, it is not a trivial task to apply spectral clustering to large-scale problems due to its computational complexity of O(n(3)), where n is the number of samples. Recently, many approaches have been proposed to accelerate the spectral clustering. Unfortunately, these methods usually sacrifice quite a lot information of the original data, thus result in a degradation of performance. In this paper, we propose a novel approach, called landmark-based spectral clustering, for large-scale clustering problems. Specifically, we select p ( << n) representative data points as the landmarks and represent the original data points as sparse linear combinations of these landmarks. The spectral embedding of the data can then be efficiently computed with the landmark-based representation. The proposed algorithm scales linearly with the problem size. Extensive experiments show the effectiveness and efficiency of our approach comparing to the state-of-the-art methods. Deng Cai 0001, Xinlei Chen |
IEEE Trans. Cybern. | 1 |
| 2015 | Treelets Binary Feature Retrieval for Fast Keypoint RecognitionabstractFast keypoint recognition is essential to many vision tasks. In contrast to the classification-based approaches, we directly formulate the keypoint recognition as an image patch retrieval problem, which enjoys the merit of finding the matched keypoint and its pose simultaneously. To effectively extract the binary features from each patch surrounding the keypoint, we make use of treelets transform that can group the highly correlated data together and reduce the noise through the local analysis. Treelets is a multiresolution analysis tool, which provides an orthogonal basis to reflect the geometry of the noise-free data. To facilitate the real-world applications, we have proposed two novel approaches. One is the convolutional treelets that capture the image patch information locally and globally while reducing the computational cost. The other is the higher-order treelets that reflect the relationship between the rows and columns within image patch. An efficient sub-signature-based locality sensitive hashing scheme is employed for fast approximate nearest neighbor search in patch retrieval. Experimental evaluations on both synthetic data and the real-world Oxford dataset have shown that our proposed treelets binary feature retrieval methods outperform the state-of-the-art feature descriptors and classification-based approaches. Jianke Zhu, Chenxia Wu, Chun Chen 0001, Deng Cai 0001 |
IEEE Trans. Cybern. | 4 |
| 2015 | EMR: A Scalable Graph-Based Ranking Model for Content-Based Image RetrievalabstractGraph-based ranking models have been widely applied in information retrieval area. In this paper, we focus on a well known graph-based model - the Ranking on Data Manifold model, or Manifold Ranking (MR). Particularly, it has been successfully applied to content-based image retrieval, because of its outstanding ability to discover underlying geometrical structure of the given image database. However, manifold ranking is computationally very expensive, which significantly limits its applicability to large databases especially for the cases that the queries are out of the database (new samples). We propose a novel scalable graph-based ranking model called Efficient Manifold Ranking (EMR), trying to address the shortcomings of MR from two main perspectives: scalable graph construction and efficient ranking computation. Specifically, we build an anchor graph on the database instead of a traditional$k$-nearest neighbor graph, and design a new form of adjacency matrix utilized to speed up the ranking. An approximate method is adopted for efficient out-of-sample retrieval. Experimental results on some large scale image databases demonstrate that EMR is a promising method for real world retrieval applications. Bin Xu 0005, Jiajun Bu, Chun Chen 0001, Can Wang 0001, Deng Cai 0001, Xiaofei He 0001 |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2014 | Mapping Users across Networks by Manifold Alignment on HypergraphabstractNowadays many people are members of multiple online social networks simultaneously, such as Facebook, Twitter and some other instant messaging circles. But these networks are usually isolated from each other. Mapping common users across these social networks will benefit many applications. Methods based on username comparison perform well on parts of users, however they can not work in the following situations: (a) users choose different usernames in different networks; (b) a unique username corresponds to different individuals. In this paper, we propose to utilize social structures to improve the mapping performance. Specifically, a novel subspace learning algorithm, Manifold Alignment on Hypergraph (MAH), is proposed. Different from traditional semi-supervised manifold alignment methods, we use hypergraph to model high-order relations here. For a target user in one network, the proposed algorithm ranks all users in the other network by their possibilities of being the corresponding user. Moreover, methods based on username comparison can be incorporated into our algorithm easily to further boost the mapping accuracy. Experimental results have demonstrated the effectiveness of our proposed algorithm in mapping users across networks. Shulong Tan, Ziyu Guan, Deng Cai 0001, Xuzhen Qin, Jiajun Bu, Chun Chen 0001 |
AAAI | 3 |
| 2014 | Who Also Likes It? Generating the Most Persuasive Social Explanations in Recommender SystemsabstractSocial explanation, the statement with the form of "A and B also like the item", is widely used in almost all the major recommender systems in the web and effectively improves the persuasiveness of the recommendation results by convincing more users to try. This paper presents the first algorithm to generate the most persuasive social explanation by recommending the optimal set of users to be put in the explanation. New challenges like modeling persuasiveness of multiple users, different types of users in social network, sparsity of likes, are discussed in depth and solved in our algorithm. The extensive evaluation demonstrates the advantage of our proposed algorithm compared with traditional methods. Beidou Wang, Martin Ester, Jiajun Bu, Deng Cai 0001 |
AAAI | 4 |
| 2014 | Sparse Learning for Stochastic Composite OptimizationabstractIn this paper, we focus on Stochastic Composite Optimization (SCO) for sparse learning that aims to learn a sparse solution. Although many SCO algorithms have been developed for sparse learning with an optimal convergence rate $O(1/T)$, they often fail to deliver sparse solutions at the end either because of the limited sparsity regularization during stochastic optimization or due to the limitation in online-to-batch conversion. To improve the sparsity of solutions obtained by SCO, we propose a simple but effective stochastic optimization scheme that adds a novel sparse online-to-batch conversion to the traditional SCO algorithms. The theoretical analysis shows that our scheme can find a solution with better sparse patterns without affecting the convergence rate. Experimental results on both synthetic and real-world data sets show that the proposed methods are more effective in recovering the sparse solution and have comparable convergence rate as the state-of-the-art SCO algorithms for sparse learning. Lijun Zhang 0005, Yao Hu 0002, Rong Jin 0001, Deng Cai 0001, Xiaofei He 0001 |
AAAI | 5 |
| 2014 | Iterative Multi-View Hashing for Cross Media IndexingabstractCross media retrieval engines have gained massive popularity with rapid development of the Internet. Users may perform queries in a corpus consisting of audio, video, and textual information. To make such systems practically possible for large mount of multimedia data, two critical issues must be carefully considered: (a) reduce the storage as much as possible; (b) model the relationship of the heterogeneous media data. Recently academic community have proved that encoding the data into compact binary codes can drastically reduce the storage and computational cost. However, it is still unclear how to integrate multiple information sources properly into the binary code encoding scheme. Yao Hu 0002, Zhongming Jin 0001, Hongyi Ren, Deng Cai 0001, Xiaofei He 0001 |
ACM Multimedia | 4 |
| 2014 | Discriminative Orthogonal Nonnegative matrix factorization with flexibility for data representation
Ping Li 0006, Jiajun Bu, Yi Yang 0001, Rongrong Ji, Chun Chen 0001, Deng Cai 0001 |
Expert Syst. Appl. | 6 |
| 2014 | Spatially correlated nonnegative matrix factorization
Xinlei Chen, Haifeng Liu 0001, Deng Cai 0001 |
Neurocomputing | 4 |
| 2014 | Manifold optimal experimental design via dependence maximization for active learning
Ping Li 0006, Jiajun Bu, Chun Chen 0001, Deng Cai 0001 |
Neurocomputing | 4 |
| 2014 | Active learning on manifolds
Haifeng Liu 0001, Deng Cai 0001 |
Neurocomputing | 3 |
| 2014 | Cross domain recommendation based on multi-type media fusion
Shulong Tan, Jiajun Bu, Xuzhen Qin, Chun Chen 0001, Deng Cai 0001 |
Neurocomputing | 5 |
| 2014 | Density Sensitive HashingabstractNearest neighbor search is a fundamental problem in various research fields like machine learning, data mining and pattern recognition. Recently, hashing-based approaches, for example, locality sensitive hashing (LSH), are proved to be effective for scalable high dimensional nearest neighbor search. Many hashing algorithms found their theoretic root in random projection. Since these algorithms generate the hash tables (projections) randomly, a large number of hash tables (i.e., long codewords) are required in order to achieve both high precision and recall. To address this limitation, we propose a novel hashing algorithm called density sensitive hashing (DSH) in this paper. DSH can be regarded as an extension of LSH. By exploring the geometric structure of the data, DSH avoids the purely random projections selection and uses those projective functions which best agree with the distribution of the data. Extensive experimental results on real-world data sets have shown that the proposed method achieves better performance compared to the state-of-the-art hashing approaches. Zhongming Jin 0001, Yue Lin 0003, Deng Cai 0001 |
IEEE Trans. Cybern. | 4 |
| 2014 | Fast and Accurate Hashing Via Iterative Nearest Neighbors ExpansionabstractRecently, the hashing techniques have been widely applied to approximate the nearest neighbor search problem in many real applications. The basic idea of these approaches is to generate binary codes for data points which can preserve the similarity between any two of them. Given a query, instead of performing a linear scan of the entire data base, the hashing method can perform a linear scan of the points whose hamming distance to the query is not greater than rh , where rh is a constant. However, in order to find the true nearest neighbors, both the locating time and the linear scan time are proportional to O(∑i=0(rh)(c || i)) ( c is the code length), which increase exponentially as rh increases. To address this limitation, we propose a novel algorithm named iterative expanding hashing in this paper, which builds an auxiliary index based on an offline constructed nearest neighbor table to avoid large rh . This auxiliary index can be easily combined with all the traditional hashing methods. Extensive experimental results over various real large-scale datasets demonstrate the superiority of the proposed approach. Zhongming Jin 0001, Debing Zhang, Yao Hu 0002, Shiding Lin, Deng Cai 0001, Xiaofei He 0001 |
IEEE Trans. Cybern. | 5 |
| 2014 | Constrained Concept Factorization for Image RepresentationabstractMatrix factorization based techniques, such as nonnegative matrix factorization and concept factorization, have attracted great attention in dimensionality reduction and data clustering. Previous studies show that both of them yield impressive results on image processing and document clustering. However, both of them are essentially unsupervised methods and cannot incorporate label information. In this paper, we propose a novel semisupervised matrix decomposition method for extracting the image concepts that are consistent with the known label information. With this constraint, we call the new approach constrained concept factorization. By requiring that the data points sharing the same label have the same coordinate in the new representation space, this approach has more discriminating power. The experimental results on several corpora show good performance of our novel algorithm in terms of clustering accuracy and mutual information. Haifeng Liu 0001, Genmao Yang, Zhaohui Wu 0001, Deng Cai 0001 |
IEEE Trans. Cybern. | 4 |
| 2014 | Feature Correlation Hypergraph: Exploiting High-order Potentials for Multimodal RecognitionabstractIn computer vision and multimedia analysis, it is common to use multiple features (or multimodal features) to represent an object. For example, to well characterize a natural scene image, we typically extract a set of visual features to represent its color, texture, and shape. However, it is challenging to integrate multimodal features optimally. Since they are usually high-order correlated, e.g., the histogram of gradient (HOG), bag of scale invariant feature transform descriptors, and wavelets are closely related because they collaboratively reflect the image texture. Nevertheless, the existing algorithms fail to capture the high-order correlation among multimodal features. To solve this problem, we present a new multimodal feature integration framework. Particularly, we first define a new measure to capture the high-order correlation among the multimodal features, which can be deemed as a direct extension of the previous binary correlation. Therefore, we construct a feature correlation hypergraph (FCH) to model the high-order relations among multimodal features. Finally, a clustering algorithm is performed on FCH to group the original multimodal features into a set of partitions. Moreover, a multiclass boosting strategy is developed to obtain a strong classifier by combining the weak classifiers learned from each partition. The experimental results on seven popular datasets show the effectiveness of our approach. Yinfu Feng, Jianke Zhu, Deng Cai 0001 |
IEEE Trans. Cybern. | 6 |
| 2013 | Compressed HashingabstractRecent studies have shown that hashing methods are effective for high dimensional nearest neighbor search. A common problem shared by many existing hashing methods is that in order to achieve a satisfied performance, a large number of hash tables (i.e., long code-words) are required. To address this challenge, in this paper we propose a novel approach called Compressed Hashing by exploring the techniques of sparse coding and compressed sensing. In particular, we introduce as parse coding scheme, based on the approximation theory of integral operator, that generate sparse representation for high dimensional vectors. We then project s-parse codes into a low dimensional space by effectively exploring the Restricted Isometry Property (RIP), a key property in compressed sensing theory. Both of the theoretical analysis and the empirical studies on two large data sets show that the proposed approach is more effective than the state-of-the-art hashing algorithms. Yue Lin 0003, Rong Jin 0001, Deng Cai 0001, Shuicheng Yan, Xuelong Li 0001 |
CVPR | 3 |
| 2013 | Complementary Projection HashingabstractRecently, hashing techniques have been widely applied to solve the approximate nearest neighbors search problem in many vision applications. Generally, these hashing approaches generate 2^c buckets, where c is the length of the hash code. A good hashing method should satisfy the following two requirements: 1) mapping the nearby data points into the same bucket or nearby (measured by the Hamming distance) buckets. 2) all the data points are evenly distributed among all the buckets. In this paper, we propose a novel algorithm named Complementary Projection Hashing (CPH) to find the optimal hashing functions which explicitly considers the above two requirements. Specifically, CPH aims at sequentially finding a series of hyper planes (hashing functions) which cross the sparse region of the data. At the same time, the data points are evenly distributed in the hyper cubes generated by these hyper planes. The experiments comparing with the state-of-the-art hashing methods demonstrate the effectiveness of the proposed method. Zhongming Jin 0001, Yao Hu 0002, Yue Lin 0003, Debing Zhang, Shiding Lin, Deng Cai 0001, Xuelong Li 0001 |
ICCV | 6 |
| 2013 | Active Learning Based on Local Representation
Yao Hu 0002, Debing Zhang, Zhongming Jin 0001, Deng Cai 0001, Xiaofei He 0001 |
IJCAI | 4 |
| 2013 | Harmonious Hashing
Bin Xu 0005, Jiajun Bu, Yue Lin 0003, Chun Chen 0001, Xiaofei He 0001, Deng Cai 0001 |
IJCAI | 6 |
| 2013 | Bilevel Visual Words Coding for Image Classification
Jiemi Zhang, Chenxia Wu, Deng Cai 0001, Jianke Zhu |
IJCAI | 3 |
| 2013 | A Unified Approximate Nearest Neighbor Search Scheme by Combining Data Structure and Hashing
Debing Zhang, Genmao Yang, Yao Hu 0002, Zhongming Jin 0001, Deng Cai 0001, Xiaofei He 0001 |
IJCAI | 5 |
| 2013 | Parallel field alignment for cross media retrievalabstractCross media retrieval systems have received increasing interest in recent years. Due to the semantic gap between low-level features and high-level semantic concepts of multimedia data, many researchers have explored joint-model techniques in cross media retrieval systems. Previous joint-model approaches usually focus on two traditional ways to design cross media retrieval systems: (a) fusing features from different media data; (b) learning different models for different media data and fusing their outputs. However, the process of fusing features or outputs will lose both low- and high-level abstraction information of media data. Hence, both ways do not really reveal the semantic correlations among the heterogeneous multimedia data. In this paper, we introduce a novel method for the cross media retrieval task, named Parallel Field Alignment Retrieval (PFAR), which integrates a manifold alignment framework from the perspective of vector fields. Instead of fusing original features or outputs, we consider the cross media retrieval as a manifold alignment problem using parallel fields. The proposed manifold alignment algorithm can effectively preserve the metric of data manifolds, model heterogeneous media data and project their relationship into intermediate latent semantic spaces during the process of manifold alignment. After the alignment, the semantic correlations are also determined. In this way, the cross media retrieval task can be resolved by the determined semantic correlations. Comprehensive experimental results have demonstrated the effectiveness of our approach. Xiangbo Mao, Binbin Lin 0001, Deng Cai 0001, Xiaofei He 0001, Jian Pei 0001 |
ACM Multimedia | 3 |
| 2013 | Whom to mention: expand the diffusion of tweets by @ recommendation on micro-blogging systemsabstractNowadays, micro-blogging systems like Twitter have become one of the most important ways for information sharing. In Twitter, a user posts a message (tweet) and the others can forward the message (retweet). Mention is a new feature in micro-blogging systems. By mentioning users in a tweet, they will receive notifications and their possible retweets may help to initiate large cascade diffusion of the tweet. To enhance a tweet's diffusion by finding the right persons to mention, we propose in this paper a novel recommendation scheme named as whom-to-mention. Specifically, we present an in-depth study of mention mechanism and propose a recommendation scheme to solve the essential question of whom to mention in a tweet. In this paper, whom-to-mention is formulated as a ranking problem and we try to address several new challenges which are not well studied in the traditional information retrieval tasks. By adopting features including user interest match, content-dependent user relationship and user influence, a machine learned ranking function is trained based on newly defined information diffusion based relevance. The extensive evaluation using data gathered from real users demonstrates the advantage of our proposed algorithm compared with the traditional recommendation methods. Beidou Wang, Can Wang 0001, Jiajun Bu, Chun Chen 0001, Wei Vivian Zhang, Deng Cai 0001, Xiaofei He 0001 |
WWW | 6 |
| 2013 | Subspace learning via Locally Constrained A-optimal nonnegative projection
Ping Li 0006, Jiajun Bu, Chun Chen 0001, Can Wang 0001, Deng Cai 0001 |
Neurocomputing | 5 |
| 2013 | Relational Multimanifold CoclusteringabstractCoclustering targets on grouping the samples (e.g.,documents and users) and the features (e.g., words and ratings) simultaneously. It employs the dual relation and the bilateral information between the samples and features. In many real-world applications, data usually reside on a submanifold of the ambient Euclidean space, but it is nontrivial to estimate the intrinsic manifold of the data space in a principled way. In this paper, we focus on improving the coclustering performance via manifold ensemble learning, which is able to maximally approximate the intrinsic manifolds of both the sample and feature spaces. To achieve this, we develop a novel coclustering algorithm called relational multimanifold coclustering based on symmetric nonnegative matrix trifactorization, which decomposes the relational data matrix into three submatrices. This method considers the intertype relationship revealed by the relational data matrix and also the intratype information reflected by the affinity matrices encoded on the sample and feature data distributions. Specifically, we assume that the intrinsic manifold of the sample or feature space lies in a convex hull of some predefined candidate manifolds. We want to learn a convex combination of them to maximally approach the desired intrinsic manifold. To optimize the objective function, the multiplicative rules are utilized to update the submatrices alternatively. In addition, both the entropic mirror descent algorithm and the coordinate descent algorithm are exploited to learn the manifold coefficient vector. Extensive experiments on documents, images, and gene expression data sets have demonstrated the superiority of the proposed algorithm compared with other well-established methods. Ping Li 0006, Jiajun Bu, Chun Chen 0001, Zhanying He, Deng Cai 0001 |
IEEE Trans. Cybern. | 5 |
| 2013 | Nonnegative Local Coordinate Factorization for Image RepresentationabstractRecently, nonnegative matrix factorization (NMF) has become increasingly popular for feature extraction in computer vision and pattern recognition. NMF seeks two nonnegative matrices whose product can best approximate the original matrix. The nonnegativity constraints lead to sparse parts-based representations that can be more robust than nonsparse global features. To obtain more accurate control over the sparseness, in this paper, we propose a novel method called nonnegative local coordinate factorization (NLCF) for feature extraction. NLCF adds a local coordinate constraint into the standard NMF objective function. Specifically, we require that the learned basis vectors be as close to the original data points as possible. In this way, each data point can be represented by a linear combination of only a few nearby basis vectors, which naturally leads to sparse representation. Extensive experimental results suggest that the proposed approach provides a better representation and achieves higher accuracy in image clustering. Jiemi Zhang, Deng Cai 0001, Wei Liu 0005, Xiaofei He 0001 |
IEEE Trans. Image Process. | 3 |
| 2013 | Parallel Field RankingabstractRecently, ranking data with respect to the intrinsic geometric structure (manifold ranking) has received considerable attentions, with encouraging performance in many applications in pattern recognition, information retrieval and recommendation systems. Most of the existing manifold ranking methods focus on learning a ranking function that varies smoothly along the data manifold. However, beyond smoothness, a desirable ranking function should vary monotonically along the geodesics of the data manifold, such that the ranking order along the geodesics is preserved. In this article, we aim to learn a ranking function that varies linearly and therefore monotonically along the geodesics of the data manifold. Recent theoretical work shows that the gradient field of a linear function on the manifold has to be a parallel vector field. Therefore, we propose a novel ranking algorithm on the data manifolds, called Parallel Field Ranking. Specifically, we try to learn a ranking function and a vector field simultaneously. We require the vector field to be close to the gradient field of the ranking function, and the vector field to be as parallel as possible. Moreover, we require the value of the ranking function at the query point to be the highest, and then decrease linearly along the manifold. Experimental results on both synthetic data and real data demonstrate the effectiveness of our proposed algorithm. Ming Ji, Binbin Lin 0001, Xiaofei He 0001, Deng Cai 0001, Jiawei Han 0001 |
ACM Trans. Knowl. Discov. Data | 4 |
| 2013 | Co-Occurrence-Based Diffusion for Expert Search on the WebabstractExpert search has been studied in different contexts, e.g., enterprises, academic communities. We examine a general expert search problem: searching experts on the web, where millions of webpages and thousands of names are considered. It has mainly two challenging issues: 1) webpages could be of varying quality and full of noises; 2) The expertise evidences scattered in webpages are usually vague and ambiguous. We propose to leverage the large amount of co-occurrence information to assess relevance and reputation of a person name for a query topic. The co-occurrence structure is modeled using a hypergraph, on which a heat diffusion based ranking algorithm is proposed. Query keywords are regarded as heat sources, and a person name which has strong connection with the query (i.e., frequently co-occur with query keywords and co-occur with other names related to query keywords) will receive most of the heat, thus being ranked high. Experiments on the ClueWeb09 web collection show that our algorithm is effective for retrieving experts and outperforms baseline algorithms significantly. This work would be regarded as one step toward addressing the more general entity search problem without sophisticated NLP techniques. Ziyu Guan, Gengxin Miao, Russell McLoughlin, Xifeng Yan, Deng Cai 0001 |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2013 | Semi-Supervised Nonlinear Hashing Using Bootstrap Sequential Projection LearningabstractIn this paper, we study the effective semi-supervised hashing method under the framework of regularized learning-based hashing. A nonlinear hash function is introduced to capture the underlying relationship among data points. Thus, the dimensionality of the matrix for computation is not only independent from the dimensionality of the original data space but also much smaller than the one using linear hash function. To effectively deal with the error accumulated during converting the real-value embeddings into the binary code after relaxation, we propose a semi-supervised nonlinear hashing algorithm using bootstrap sequential projection learning which effectively corrects the errors by taking into account of all the previous learned bits holistically without incurring the extra computational overhead. Experimental results on the six benchmark data sets demonstrate that the presented method outperforms the state-of-the-art hashing algorithms at a large margin. Chenxia Wu, Jianke Zhu, Deng Cai 0001, Chun Chen 0001, Jiajun Bu |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2012 | Document Summarization Based on Data ReconstructionabstractDocument summarization is of great value to many real world applications, such as snippets generation for search results and news headlines generation. Traditionally, document summarization is implemented by extracting sentences that cover the main topics of a document with a minimum redundancy. In this paper, we take a different perspective from data reconstruction and propose a novel framework named Document Summarization based on Data Reconstruction (DSDR). Specifically, our approach generates a summary which consist of those sentences that can best reconstruct the original document. To model the relationship among sentences, we introduce two objective functions: (1) linear reconstruction, which approximates the document by linear combinations of the selected sentences; (2) nonnegative linear reconstruction, which allows only additive, not subtractive, linear combinations. In this framework, the reconstruction error becomes a natural criterion for measuring the quality of the summary. For each objective function, we develop an efficient algorithm to solve the corresponding optimization problem. Extensive experiments on summarization benchmark data sets DUC 2006 and DUC 2007 demonstrate the effectiveness of our proposed approach. Zhanying He, Chun Chen 0001, Jiajun Bu, Can Wang 0001, Lijun Zhang 0005, Deng Cai 0001, Xiaofei He 0001 |
AAAI | 6 |
| 2012 | Random Projection with Filtering for Nearly Duplicate SearchabstractHigh dimensional nearest neighbor search is a fundamental problem and has found applications in many domains. Although many hashing based approaches have been proposed for approximate nearest neighbor search in high dimensional space, one main drawback is that they often return many false positives that need to be filtered out by a post procedure. We propose a novel method to address this limitation in this paper. The key idea is to introduce a filtering procedure within the search algorithm, based on the compressed sensing theory, that effectively removes the false positive answers. We first obtain a sparse representation for each data point by the landmark based approach, after which we solve the nearly duplicate search that the difference between the query and its nearest neighbors forms a sparse vector living in a small ℓp ball, where p ≤ 1. Our empirical study on real-world datasets demonstrates the effectiveness of the proposed approach compared to thestate-of-the-art hashing methods. Yue Lin 0003, Rong Jin 0001, Deng Cai 0001, Xiaofei He 0001 |
AAAI | 3 |
| 2012 | A Bregman Divergence Optimization Framework for Ranking on Data Manifold and Its New ExtensionsabstractRecently, graph-based ranking algorithms have received considerable interests in machine learning, computer vision and information retrieval communities. Ranking on data manifold (or manifold ranking, MR) is one of the representative approaches. One of the limitations of manifold ranking is its high computational complexity (O(n3), where n is the number of samples in database). In this paper, we cast the manifold ranking into a Bregman divergence optimization framework under which we transform the original MR to an equivalent optimal kernel matrix learning problem.With this new formulation, two effective and efficient extensions are proposed to enhance the ranking performance. Extensive experimental results on two real world image databases show the effectiveness of the proposed approach. Bin Xu 0005, Jiajun Bu, Chun Chen 0001, Deng Cai 0001 |
AAAI | 4 |
| 2012 | Accelerating locality preserving nonnegative matrix factorizationabstractMatrix factorization techniques have been frequently applied in information retrieval, computer vision and pattern recognition. Among them, Non-negative Matrix Factorization (NMF) has received considerable attention due to its psychological and physiological interpretation of naturally occurring data whose representation may be parts-based in the human brain. Locality Preserving Non-negative Matrix Factorization (LPNMF) is a recently proposed graph-based NMF extension which tries to preserves the intrinsic geometric structure of the data. Compared with the original NMF, LPNMF has more discriminating power on data representa- tion thanks to its geometrical interpretation and outstanding ability to discover the hidden topics. However, the computa- tional complexity of LPNMF is O(n3), where n is the number of samples. In this paper, we propose a novel approach called Accelerated LPNMF (A-LPNMF) to solve the com- putational issue of LPNMF. Specifically, A-LPNMF selects p (p j n) landmark points from the data and represents all the samples as the sparse linear combination of these landmarks. The non-negative factors which incorporates the geometric structure can then be efficiently computed. Experimental results on the real data sets demonstrate the effectiveness and efficiency of our proposed method. Guanhong Yao, Deng Cai 0001 |
CIKM | 2 |
| 2012 | Metric learning with two-dimensional smoothness for visual analysisabstractIn recent years, metric learning methods based on pairwise side information have attracted considerable interests, and lots of efforts have been devoted to utilize these methods for visual analysis like content based image retrieval and face identification. When applied to image analysis, these methods merely look on an n1× n2image as a vector in Rn1×n2space and the pixels of the image are considered as independent. They fail to consider the fact that an image represented in the plane is intrinsically a matrix, and pixels spatially close to each other may probably be correlated. Even though we have n1× n2pixels per image, this spatial correlation suggests the real number of freedom is far less. In this paper, we introduce a regularized metric learning framework, Two-Dimensional Smooth Metric Learning (2DSML), which uses a discretized Laplacian penalty to restrict the coefficients to be two-dimensional smooth. Many existing metric learning algorithms can fit into this framework and learn a spatially smooth metric which is better for image applications than their original version. Recognition, clustering and retrieval can be then performed based on the learned metric. Experimental results on benchmark image datasets demonstrate the effectiveness of our method. Xinlei Chen, Zifei Tong, Haifeng Liu 0001, Deng Cai 0001 |
CVPR | 4 |
| 2012 | A Convolutional Treelets Binary Feature Approach to Fast Keypoint Recognition
Chenxia Wu, Jianke Zhu, Jiemi Zhang, Chun Chen 0001, Deng Cai 0001 |
ECCV (5) | 5 |
| 2012 | Parallel field rankingabstractRecently, ranking data with respect to the intrinsic geometric structure (manifold ranking) has received considerable attentions, with encouraging performance in many applications in pattern recognition, information retrieval and recommendation systems. Most of the existing manifold ranking methods focus on learning a ranking function that varies smoothly along the data manifold. However, beyond smoothness, a desirable ranking function should vary monotonically along the geodesics of the data manifold, such that the ranking order along the geodesics is preserved. In this paper, we aim to learn a ranking function that varies linearly and therefore monotonically along the geodesics of the data manifold. Recent theoretical work shows that the gradient field of a linear function on the manifold has to be a parallel vector field. Therefore, we propose a novel ranking algorithm on the data manifolds, called Parallel Field Ranking. Specifically, we try to learn a ranking function and a vector field simultaneously. We require the vector field to be close to the gradient field of the ranking function, and the vector field to be as parallel as possible. Moreover, we require the value of the ranking function at the query point to be the highest, and then decrease linearly along the manifold. Experimental results on both synthetic data and real data demonstrate the effectiveness of our proposed algorithm. Ming Ji, Binbin Lin 0001, Xiaofei He 0001, Deng Cai 0001, Jiawei Han 0001 |
KDD | 4 |
| 2012 | Unsupervised face-name association via commute distanceabstractRecently, the task of unsupervised face-name association has received a considerable interests in multimedia and information retrieval communities. It is quite different with the generic facial image annotation problem because of its unsupervised and ambiguous assignment properties. Specifically, the task of face-name association should obey the following three constraints: (1) a face can only be assigned to a name appearing in its associated caption or to null; (2) a name can be assigned to at most one face; and (3) a face can be assigned to at most one name. Many conventional methods have been proposed to tackle this task while suffering from some common problems, eg, many of them are computational expensive and hard to make the null assignment decision. In this paper, we design a novel framework named face-name association via commute distance (FACD), which judges face-name and face-null assignments under a unified framework via commute distance (CD) algorithm. Then, to further speed up the on-line processing, we propose a novel anchor-based commute distance (ACD) algorithm whose main idea is using the anchor point representation structure to accelerate the eigen-decomposition of the adjacency matrix of a graph. Systematic experiment results on a large scale and real world image-caption database with a total of 194,046 detected faces and 244,725 names show that our proposed approach outperforms many state-of-the-art methods in performance. Our framework is appropriate for a large scale and real-time system. Jiajun Bu, Bin Xu 0005, Chenxia Wu, Chun Chen 0001, Jianke Zhu, Deng Cai 0001, Xiaofei He 0001 |
ACM Multimedia | 6 |
| 2012 | Search web images using objects, backgrounds and conditionsabstractAs the volumes of web images have grown rapidly in the last decade, Content-Based Image Retrieval (CBIR) has attracted substantial interests as an effective tool to manage the images. Most existing CBIR systems focus on the object in the image, while ignoring the conditions (day/night, sunny/rain, etc) and the backgrounds, both of which are very helpful to meet the user's information need. To overcome this shortcoming, in this paper, we present a novel CBIR system depending on a novel query formulation considering three aspects: Object, Background and Condition. Specifically, we design a user-friendly interface to help the user formulate a query. The interface can allow the user to give the percentage, relative position and size of each object in the background. Moreover, a corresponding effective ranking method is proposed to return the desirable search results. Experimental results demonstrate that our proposed system improves the searching performance and the user experience compared with the existing searching systems. Jiemi Zhang, Chenxia Wu, Deng Cai 0001 |
ACM Multimedia | 3 |
| 2012 | An exploration of improving collaborative recommender systems via user-item subgroupsabstractCollaborative filtering (CF) is one of the most successful recommendation approaches. It typically associates a user with a group of like-minded users based on their preferences over all the items, and recommends to the user those items enjoyed by others in the group. However we find that two users with similar tastes on one item subset may have totally different tastes on another set. In other words, there exist many user-item subgroups each consisting of a subset of items and a group of like-minded users on these items. It is more natural to make preference predictions for a user via the correlated subgroups than the entire user-item matrix. In this paper, to find meaningful subgroups, we formulate the Multiclass Co-Clustering (MCoC) problem and propose an effective solution to it. Then we propose an unified framework to extend the traditional CF algorithms by utilizing the subgroups information for improving their top-N recommendation performance. Our approach can be seen as an extension of traditional clustering CF models. Systematic experiments on three real world data sets have demonstrated the effectiveness of our proposed approach. Bin Xu 0005, Jiajun Bu, Chun Chen 0001, Deng Cai 0001 |
WWW | 4 |
| 2012 | Constrained Nonnegative Matrix Factorization for Image RepresentationabstractNonnegative matrix factorization (NMF) is a popular technique for finding parts-based, linear representations of nonnegative data. It has been successfully applied in a wide range of applications such as pattern recognition, information retrieval, and computer vision. However, NMF is essentially an unsupervised method and cannot make use of label information. In this paper, we propose a novel semi-supervised matrix decomposition method, called Constrained Nonnegative Matrix Factorization (CNMF), which incorporates the label information as additional constraints. Specifically, we show how explicitly combining label information improves the discriminating power of the resulting matrix decomposition. We explore the proposed CNMF method with two cost function formulations and provide the corresponding update solutions for the optimization problems. Empirical experiments demonstrate the effectiveness of our novel algorithm in comparison to the state-of-the-art approaches through a set of evaluations based on real-world applications. Haifeng Liu 0001, Zhaohui Wu 0001, Xuelong Li 0001, Deng Cai 0001, Thomas S. Huang |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2012 | Locally discriminative topic modeling
Jiajun Bu, Chun Chen 0001, Jianke Zhu, Lijun Zhang 0005, Haifeng Liu 0001, Can Wang 0001, Deng Cai 0001 |
Pattern Recognit. | 8 |
| 2012 | Manifold Adaptive Experimental Design for Text CategorizationabstractIn many information processing tasks, labels are usually expensive and the unlabeled data points are abundant. To reduce the cost on collecting labels, it is crucial to predict which unlabeled examples are the most informative, i.e., improve the classifier the most if they were labeled. Many active learning techniques have been proposed for text categorization, such as SVMActiveand Transductive Experimental Design. However, most of previous approaches try to discover the discriminant structure of the data space, whereas the geometrical structure is not well respected. In this paper, we propose a novel active learning algorithm which is performed in the data manifold adaptive kernel space. The manifold structure is incorporated into the kernel space by using graph Laplacian. This way, the manifold adaptive kernel space reflects the underlying geometry of the data. By minimizing the expected error with respect to the optimal classifier, we can select the most representative and discriminative data points for labeling. Experimental results on text categorization have demonstrated the effectiveness of our proposed approach. Deng Cai 0001, Xiaofei He 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2012 | Locally Discriminative CoclusteringabstractDifferent from traditional one-sided clustering techniques, coclustering makes use of the duality between samples and features to partition them simultaneously. Most of the existing co-clustering algorithms focus on modeling the relationship between samples and features, whereas the intersample and interfeature relationships are ignored. In this paper, we propose a novel coclustering algorithm named Locally Discriminative Coclustering (LDCC) to explore the relationship between samples and features as well as the intersample and interfeature relationships. Specifically, the sample-feature relationship is modeled by a bipartite graph between samples and features. And we apply local linear regression to discovering the intrinsic discriminative structures of both sample space and feature space. For each local patch in the sample and feature spaces, a local linear function is estimated to predict the labels of the points in this patch. The intersample and interfeature relationships are thus captured by minimizing the fitting errors of all the local linear functions. In this way, LDCC groups strongly associated samples and features together, while respecting the local structures of both sample and feature spaces. Our experimental results on several benchmark data sets have demonstrated the effectiveness of the proposed method. Lijun Zhang 0005, Chun Chen 0001, Jiajun Bu, Zhengguang Chen, Deng Cai 0001, Jiawei Han 0001 |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2011 | Large Scale Spectral Clustering with Landmark-Based RepresentationabstractSpectral clustering is one of the most popular clustering approaches. Despite its good performance, it is limited in its applicability to large-scale problems due to its high computational complexity. Recently, many approaches have been proposed to accelerate the spectral clustering. Unfortunately, these methods usually sacrifice quite a lot information of the original data, thus result in a degradation of performance. In this paper, we propose a novel approach, called Landmark-based Spectral Clustering (LSC), for large scale clustering problems. Specifically, we select $p\ (\ll n)$ representative data points as the landmarks and represent the original data points as the linear combinations of these landmarks. The spectral embedding of the data can then be efficiently computed with the landmark-based representation. The proposed algorithm scales linearly with the problem size. Extensive experiments show the effectiveness and efficiency of our approach comparing to the state-of-the-art methods. Xinlei Chen, Deng Cai 0001 |
AAAI | 2 |
| 2011 | Sparse concept coding for visual analysisabstractWe consider the problem of image representation for visual analysis. When representing images as vectors, the feature space is of very high dimensionality, which makes it difficult for applying statistical techniques for visual analysis. To tackle this problem, matrix factorization techniques, such as Singular Vector Decomposition (SVD) and Non-negative Matrix Factorization (NMF), received an increasing amount of interest in recent years. Matrix factorization is an unsupervised learning technique, which finds a basis set capturing high-level semantics in the data and learns coordinates in terms of the basis set. However, the representations obtained by them are highly dense and can not capture the intrinsic geometric structure in the data. In this paper, we propose a novel method, called Sparse Concept Coding (SCC), for image representation and analysis. Inspired from the recent developments on manifold learning and sparse coding, SCC provides a sparse representation which can capture the intrinsic geometric structure of the image space. Extensive experimental results on image clustering have shown that the proposed approach provides a better representation with respect to the semantic structure. Deng Cai 0001, Hujun Bao, Xiaofei He 0001 |
CVPR | 1 |
| 2011 | Efficient manifold ranking for image retrievalabstractManifold Ranking (MR), a graph-based ranking algorithm, has been widely applied in information retrieval and shown to have excellent performance and feasibility on a variety of data types. Particularly, it has been successfully applied to content-based image retrieval, because of its outstanding ability to discover underlying geometrical structure of the given image database. However, manifold ranking is computationally very expensive, both in graph construction and ranking computation stages, which significantly limits its applicability to very large data sets. In this paper, we extend the original manifold ranking algorithm and propose a new framework named Efficient Manifold Ranking (EMR). We aim to address the shortcomings of MR from two perspectives: scalable graph construction and efficient computation. Specifically, we build an anchor graph on the data set instead of the traditional k-nearest neighbor graph, and design a new form of adjacency matrix utilized to speed up the ranking computation. The experimental results on a real world image database demonstrate the effectiveness and efficiency of our proposed method. With a comparable performance to the original manifold ranking, our method significantly reduces the computational time, makes it a promising method to large scale real world retrieval problems. Bin Xu 0005, Jiajun Bu, Chun Chen 0001, Deng Cai 0001, Xiaofei He 0001, Wei Liu 0005, Jiebo Luo 0001 |
SIGIR | 4 |
| 2011 | Kernel approximately harmonic projection
Guanhong Yao, Binbin Lin 0001, Deng Cai 0001 |
Neurocomputing | 4 |
| 2011 | Graph Regularized Nonnegative Matrix Factorization for Data RepresentationabstractMatrix factorization techniques have been frequently applied in information retrieval, computer vision, and pattern recognition. Among them, Nonnegative Matrix Factorization (NMF) has received considerable attention due to its psychological and physiological interpretation of naturally occurring data whose representation may be parts based in the human brain. On the other hand, from the geometric perspective, the data is usually sampled from a low-dimensional manifold embedded in a high-dimensional ambient space. One then hopes to find a compact representation,which uncovers the hidden semantics and simultaneously respects the intrinsic geometric structure. In this paper, we propose a novel algorithm, called Graph Regularized Nonnegative Matrix Factorization (GNMF), for this purpose. In GNMF, an affinity graph is constructed to encode the geometrical information and we seek a matrix factorization, which respects the graph structure. Our empirical study shows encouraging results of the proposed algorithm in comparison to the state-of-the-art algorithms on real-world problems. Deng Cai 0001, Xiaofei He 0001, Jiawei Han 0001, Thomas S. Huang |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2011 | Active Learning Based on Locally Linear ReconstructionabstractWe consider the active learning problem, which aims to select the most representative points. Out of many existing active learning techniques, optimum experimental design (OED) has received considerable attention recently. The typical OED criteria minimize the variance of the parameter estimates or predicted value. However, these methods see only global euclidean structure, while the local manifold structure is ignored. For example, I-optimal design selects those data points such that other data points can be best approximated by linear combinations of all the selected points. In this paper, we propose a novel active learning algorithm which takes into account the local structure of the data space. That is, each data point should be approximated by the linear combination of only its neighbors. Given the local reconstruction coefficients for every data point and the coordinates of the selected points, a transductive learning algorithm called Locally Linear Reconstruction (LLR) is proposed to reconstruct every other point. The most representative points are thus defined as those whose coordinates can be used to best reconstruct the whole data set. The sequential and convex optimization schemes are also introduced to solve the optimization problem. The experimental results have demonstrated the effectiveness of our proposed method. Lijun Zhang 0005, Chun Chen 0001, Jiajun Bu, Deng Cai 0001, Xiaofei He 0001, Thomas S. Huang |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2011 | Graph Regularized Sparse Coding for Image RepresentationabstractSparse coding has received an increasing amount of interest in recent years. It is an unsupervised learning algorithm, which finds a basis set capturing high-level semantics in the data and learns sparse coordinates in terms of the basis set. Originally applied to modeling the human visual cortex, sparse coding has been shown useful for many applications. However, most of the existing approaches to sparse coding fail to consider the geometrical structure of the data space. In many real applications, the data is more likely to reside on a low-dimensional submanifold embedded in the high-dimensional ambient space. It has been shown that the geometrical information of the data is important for discrimination. In this paper, we propose a graph based algorithm, called graph regularized sparse coding, to learn the sparse representations that explicitly take into account the local manifold structure of the data. By using graph Laplacian as a smooth operator, the obtained sparse representations vary smoothly along the geodesics of the data manifold. The extensive experimental results on image classification and clustering have demonstrated the effectiveness of our proposed algorithm. Miao Zheng, Jiajun Bu, Chun Chen 0001, Can Wang 0001, Lijun Zhang 0005, Guang Qiu, Deng Cai 0001 |
IEEE Trans. Image Process. | 7 |
| 2011 | Variational inference with graph regularization for image annotationabstractImage annotation is a typical area where there are multiple types of attributes associated with each individual image. In order to achieve better performance, it is important to develop effective modeling by utilizing prior knowledge. In this article, we extend the graph regularization approaches to a more general case where the regularization is imposed on the factorized variational distributions, instead of posterior distributions implicitly involved in EM-like algorithms. In this way, the problem modeling can be more flexible, and we can choose any factor in the problem domain to impose graph regularization wherever there are similarity constraints among the instances. We formulate the problem formally and show its geometrical background in manifold learning. We also design two practically effective algorithms and analyze their properties such as the convergence. Finally, we apply our approach to image annotation and show the performance improvement of our algorithm. Yuanlong Shao, Deng Cai 0001 |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2011 | Locally Consistent Concept Factorization for Document ClusteringabstractPrevious studies have demonstrated that document clustering performance can be improved significantly in lower dimensional linear subspaces. Recently, matrix factorization-based techniques, such as Nonnegative Matrix Factorization (NMF) and Concept Factorization (CF), have yielded impressive results. However, both of them effectively see only the global euclidean geometry, whereas the local manifold geometry is not fully considered. In this paper, we propose a new approach to extract the document concepts which are consistent with the manifold geometry such that each concept corresponds to a connected component. Central to our approach is a graph model which captures the local geometry of the document submanifold. Thus, we call it Locally Consistent Concept Factorization (LCCF). By using the graph Laplacian to smooth the document-to-concept mapping, LCCF can extract concepts with respect to the intrinsic manifold structure and thus documents associated with the same concept can be well clustered. The experimental results on TDT2 and Reuters-21578 have shown that the proposed approach provides a better representation and achieves better clustering results in terms of accuracy and mutual information. Deng Cai 0001, Xiaofei He 0001, Jiawei Han 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2011 | Laplacian Regularized Gaussian Mixture Model for Data ClusteringabstractGaussian Mixture Models (GMMs) are among the most statistically mature methods for clustering. Each cluster is represented by a Gaussian distribution. The clustering process thereby turns to estimate the parameters of the Gaussian mixture, usually by the Expectation-Maximization algorithm. In this paper, we consider the case where the probability distribution that generates the data is supported on a submanifold of the ambient space. It is natural to assume that if two points are close in the intrinsic geometry of the probability distribution, then their conditional probability distributions are similar. Specifically, we introduce a regularized probabilistic model based on manifold structure for data clustering, called Laplacian regularized Gaussian Mixture Model (LapGMM). The data manifold is modeled by a nearest neighbor graph, and the graph structure is incorporated in the maximum likelihood objective function. As a result, the obtained conditional probability distribution varies smoothly along the geodesics of the data manifold. Experimental results on real data sets demonstrate the effectiveness of the proposed approach. Xiaofei He 0001, Deng Cai 0001, Yuanlong Shao, Hujun Bao, Jiawei Han 0001 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2011 | Speed up kernel discriminant analysis
Deng Cai 0001, Xiaofei He 0001, Jiawei Han 0001 |
VLDB J. | 1 |
| 2010 | Gaussian Mixture Model with Local ConsistencyabstractGaussian Mixture Model (GMM) is one of the most popular data clustering methods which can be viewed as a linear combination of different Gaussian components. In GMM, each cluster obeys Gaussian distribution and the task of clustering is to group observations into different components through estimating each cluster's own parameters. The Expectation-Maximization algorithm is always involved in such estimation problem. However, many previous studies have shown naturally occurring data may reside on or close to an underlying submanifold. In this paper, we consider the case where the probability distribution is supported on a submanifold of the ambient space. We take into account the smoothness of the conditional probability distribution along the geodesics of data manifold. That is, if two observations are close in intrinsic geometry, their distributions over different Gaussian components are similar. Simply speaking, we introduce a novel method based on manifold structure for data clustering, called Locally Consistent Gaussian Mixture Model (LCGMM). Specifically, we construct a nearest neighbor graph and adopt Kullback-Leibler Divergence as the distance measurement to regularize the objective function of GMM. Experiments on several data sets demonstrate the effectiveness of such regularization. Deng Cai 0001, Xiaofei He 0001 |
AAAI | 2 |
| 2010 | Laplacian Co-hashing of Terms and Documents
Dell Zhang, Jun Wang 0012, Deng Cai 0001, Jinsong Lu |
ECIR | 3 |
| 2010 | Unsupervised feature selection for multi-cluster dataabstractIn many data analysis tasks, one is often confronted with very high dimensional data. Feature selection techniques are designed to find the relevant feature subset of the original features which can facilitate clustering, classification and retrieval. In this paper, we consider the feature selection problem in unsupervised learning scenario, which is particularly difficult due to the absence of class labels that would guide the search for relevant information. The feature selection problem is essentially a combinatorial optimization problem which is computationally expensive. Traditional unsupervised feature selection methods address this issue by selecting the top ranked features based on certain scores computed independently for each feature. These approaches neglect the possible correlation between different features and thus can not produce an optimal feature subset. Inspired from the recent developments on manifold learning and L1-regularized models for subset selection, we propose in this paper a new approach, called Multi-Cluster Feature Selection (MCFS), for unsupervised feature selection. Specifically, we select those features such that the multi-cluster structure of the data can be best preserved. The corresponding optimization problem can be efficiently solved since it only involves a sparse eigen-problem and a L1-regularized least squares problem. Extensive experimental results over various real-life data sets have demonstrated the superiority of the proposed algorithm. Deng Cai 0001, Chiyuan Zhang, Xiaofei He 0001 |
KDD | 1 |
| 2010 | Self-taught hashing for fast similarity searchabstractThe ability of fast similarity search at large scale is of great importance to many Information Retrieval (IR) applications. A promising way to accelerate similarity search is semantic hashing which designs compact binary codes for a large number of documents so that semantically similar documents are mapped to similar codes (within a short Hamming distance). Although some recently proposed techniques are able to generate high-quality codes for documents known in advance, obtaining the codes for previously unseen documents remains to be a very challenging problem. In this paper, we emphasise this issue and propose a novel Self-Taught Hashing (STH) approach to semantic hashing: we first find the optimal l-bit binary codes for all documents in the given corpus via unsupervised learning, and then train l classifiers via supervised learning to predict the l-bit code for any query document unseen before. Our experiments on three real-world text datasets show that the proposed approach using binarised Laplacian Eigenmap (LapEig) and linear Support Vector Machine (SVM) outperforms state-of-the-art techniques significantly. Dell Zhang, Jun Wang 0012, Deng Cai 0001, Jinsong Lu |
SIGIR | 3 |
| 2010 | Document recommendation in social tagging servicesabstractSocial tagging services allow users to annotate various on-line resources with freely chosen keywords (tags). They not only facilitate the users in finding and organizing online re-sources, but also provide meaningful collaborative semantic data which can potentially be exploited by recommender systems. Traditional studies on recommender systems fo-cused on user rating data, while recently social tagging data is becoming more and more prevalent. How to perform re-source recommendation based on tagging data is an emerg-ing research topic. In this paper we consider the problem of document (e.g. Web pages, research papers) recommen-dation using purely tagging data. That is, we only have data containing users, tags, documents and the relation-ships among them. We propose a novel graph-based rep-resentation learning algorithm for this purpose. The users, tags and documents are represented in the same semantic space in which two related objects are close to each other. For a given user, we recommend those documents that are sufficiently close to him/her. Experimental results on two data sets crawled from Del.icio.us and CiteULike show that our algorithm can generate promising recommendations and outperforms traditional recommendation algorithms. Ziyu Guan, Can Wang 0001, Jiajun Bu, Chun Chen 0001, Deng Cai 0001, Xiaofei He 0001 |
WWW | 6 |
| 2009 | Active subspace learningabstractMany previous studies have shown that naturally occurring data cannot possibly fill up the high dimensional space uniformly, rather it must concentrate around lower dimensional structure. The typical supervised subspace learning algorithms to discover this low dimensional structure include Linear Discriminant Analysis (LDA). For LDA, the training data points are usually pre-given. However, in some real world applications like relevance feedback image retrieval, there is opportunity to interact with the user and actively select the training points for labeling. In this paper, we propose a novel active subspace learning algorithm which selects the most informative data points and uses them for learning an optimal subspace. Using techniques from experimental design, we discuss how to perform data selection in supervised or semi-supervised subspace learning by minimizing the expected error. Experiments on image retrieval show improvement over state-of-the-art methods. Xiaofei He 0001, Deng Cai 0001 |
ICCV | 2 |
| 2009 | Probabilistic dyadic data analysis with local and global consistencyabstractDyadic data arises in many real world applications such as social network analysis and information retrieval. In order to discover the underlying or hidden structure in the dyadic data, many topic modeling techniques were proposed. The typical algorithms include Probabilistic Latent Semantic Analysis (PLSA) and Latent Dirichlet Allocation (LDA). The probability density functions obtained by both of these two algorithms are supported on the Euclidean space. However, many previous studies have shown naturally occurring data may reside on or close to an underlying submanifold. We introduce a probabilistic framework for modeling both the topical and geometrical structure of the dyadic data that explicitly takes into account the local manifold structure. Specifically, the local manifold structure is modeled by a graph. The graph Laplacian, analogous to the Laplace-Beltrami operator on manifolds, is applied to smooth the probability density functions. As a result, the obtained probabilistic distributions are concentrated around the data manifold. Experimental results on real data sets demonstrate the effectiveness of the proposed approach. Deng Cai 0001, Xuanhui Wang, Xiaofei He 0001 |
ICML | 1 |
| 2009 | Locality Preserving Nonnegative Matrix Factorization
Deng Cai 0001, Xiaofei He 0001, Xuanhui Wang, Hujun Bao, Jiawei Han 0001 |
IJCAI | 1 |
| 2009 | Semi-supervised topic modeling for image annotationabstractWe propose a novel technique for semi-supervised image annotation which introduces a harmonic regularizer based on the graph Laplacian of the data into the probabilistic semantic model for learning latent topics of the images. By using a probabilistic semantic model, we connect visual features and textual annotations of images by their latent topics. Meanwhile, we incorporate the manifold assumption into the model to say that the probabilities of latent topics of images are drawn from a manifold, so that for images sharing similar visual features or the same annotations, their probability distribution of latent topics should also be similar. We create a nearest neighbor graph to model the manifold and propose a regularized EM algorithm to simultaneously learn a generative model and assign probability density of latent topics to images discriminatively. In this way, databases with very few labeled images can be annotated better than previous works. Yuanlong Shao, Xiaofei He 0001, Deng Cai 0001, Hujun Bao |
ACM Multimedia | 4 |
| 2009 | Convex experimental design using manifold structure for image retrievalabstractContent Based Image Retrieval (CBIR) has become one of the most active research areas in computer science. Relevance feedback is often used in CBIR systems to bridge the semantic gap. Typically, users are asked to make relevance judgements on some query results, and the feedback information is then used to re-rank the images in the database. An effective relevance feedback algorithm must provide the users with the most informative images with respect to the ranking function. In this paper, we propose a novel active learning algorithm, called Convex Laplacian Regularized I-optimal Design (CLapRID), for relevance feedback image retrieval. Our algorithm is based on a regression model which minimizes the least square error on the labeled images and simultaneously preserves the intrinsic geometrical structure of the image space. It selects the most informative images which minimize the average predictive variance. The optimization problem of CLapRID can be cast as a semidefinite programming (SDP) problem, and solved via interior-point methods. Experimental results on COREL database have demonstrate the effectiveness of the proposed algorithm for relevance feedback image retrieval. Categories and Subject Descriptors H.3.3 [Information storage and retrieval]: Information search and retrieval—Relevance feedback; G.3 [Mathematics of Computing]: Probability and Statistics—Experimental design Lijun Zhang 0005, Chun Chen 0001, Wei Chen 0005, Jiajun Bu, Deng Cai 0001, Xiaofei He 0001 |
ACM Multimedia | 5 |
| 2008 | Sparse Projections over Graph
Deng Cai 0001, Xiaofei He 0001, Jiawei Han 0001 |
AAAI | 1 |
| 2008 | Modeling hidden topics on document manifoldabstractTopic modeling has been a key problem for document analysis. One of the canonical approaches for topic modeling is Probabilistic Latent Semantic Indexing, which maximizes the joint probability of documents and terms in the corpus. The major disadvantage of PLSI is that it estimates the probability distribution of each document on the hidden topics independently and the number of parameters in the model grows linearly with the size of the corpus, which leads to serious problems with overfitting. Latent Dirichlet Allocation (LDA) is proposed to overcome this problem by treating the probability distribution of each document over topics as a hidden random variable. Both of these two methods discover the hidden topics in the Euclidean space. However, there is no convincing evidence that the document space is Euclidean, or flat. Therefore, it is more natural and reasonable to assume that the document space is a manifold, either linear or nonlinear. In this paper, we consider the problem of topic modeling on intrinsic document manifold. Specifically, we propose a novel algorithm called Laplacian Probabilistic Latent Semantic Indexing (LapPLSI) for topic modeling. LapPLSI models the document space as a submanifold embedded in the ambient space and directly performs the topic modeling on this document manifold in question. We compare the proposed LapPLSI approach with PLSI and LDA on three text data sets. Experimental results show that LapPLSI provides better representation in the sense of semantic structure. Deng Cai 0001, Qiaozhu Mei, Jiawei Han 0001, ChengXiang Zhai |
CIKM | 1 |
| 2008 | Training Linear Discriminant Analysis in Linear TimeabstractLinear Discriminant Analysis (LDA) has been a popular method for extracting features which preserve class separability. It has been widely used in many fields of information processing, such as machine learning, data mining, information retrieval, and pattern recognition. However, the computation of LDA involves dense matrices eigen-decomposition which can be computationally expensive both in time and memory. Specifically, LDA has O(mnt + t3) time complexity and requires O(mn + mt + nt) memory, where m is the number of samples, n is the number of features and t = min (m,n). When both m and n are large, it is infeasible to apply LDA. In this paper, we propose a novel algorithm for discriminant analysis, called Spectral Regression Discriminant Analysis (SRDA). By using spectral graph analysis, SRDA casts discriminant analysis into a regression framework which facilitates both efficient computation and the use of regularization techniques. Our theoretical analysis shows that SRDA can be computed with O(ms) time and O(ms) memory, where s(les n) is the average number of non-zero features in each sample. Extensive experimental results on four real world data sets demonstrate the effectiveness and efficiency of our algorithm. Deng Cai 0001, Xiaofei He 0001, Jiawei Han 0001 |
ICDE | 1 |
| 2008 | Non-negative Matrix Factorization on ManifoldabstractRecently non-negative matrix factorization (NMF) has received a lot of attentions in information retrieval, computer vision and pattern recognition. NMF aims to find two non-negative matrices whose product can well approximate the original matrix. The sizes of these two matrices are usually smaller than the original matrix. This results in a compressed version of the original data matrix. The solution of NMF yields a natural parts-based representation for the data. When NMF is applied for data representation, a major disadvantage is that it fails to consider the geometric structure in the data. In this paper, we develop a graph based approach for parts-based data representation in order to overcome this limitation. We construct an affinity graph to encode the geometrical information and seek a matrix factorization which respects the graph structure. We demonstrate the success of this novel algorithm by applying it on real world problems. Deng Cai 0001, Xiaofei He 0001, Xiaoyun Wu, Jiawei Han 0001 |
ICDM | 1 |
| 2008 | Topic modeling with network regularizationabstractIn this paper, we formally define the problem of topic modeling with network structure (TMN). We propose a novel solution to this problem, which regularizes a statistical topic model with a harmonic regularizer based on a graph structure in the data. The proposed method bridges topic modeling and social network analysis, which leverages the power of both statistical topic models and discrete regularization. The output of this model well summarizes topics in text, maps a topic on the network, and discovers topical communities. With concrete selection of a topic model and a graph-based regularizer, our model can be applied to text mining problems such as author-topic analysis, community discovery, and spatial text mining. Empirical experiments on two different genres of data show that our approach is effective, which improves text-oriented methods as well as network-oriented methods. The proposed model is general; it can be applied to any text collections with a mixture of topics and an associated network structure. Qiaozhu Mei, Deng Cai 0001, Duo Zhang 0001, ChengXiang Zhai |
WWW | 2 |
| 2008 | SRDA: An Efficient Algorithm for Large-Scale Discriminant AnalysisabstractLinear Discriminant Analysis (LDA) has been a popular method for extracting features that preserves class separability. The projection functions of LDA are commonly obtained by maximizing the between-class covariance and simultaneously minimizing the within-class covariance. It has been widely used in many fields of information processing, such as machine learning, data mining, information retrieval, and pattern recognition. However, the computation of LDA involves dense matrices eigendecomposition, which can be computationally expensive in both time and memory. Specifically, LDA has O(mnt + t3) time complexity and requires O(mn + mt + nt) memory, where m is the number of samples, n is the number of features, and t = min(m,n). When both m and n are large, it is infeasible to apply LDA. In this paper, we propose a novel algorithm for discriminant analysis, calledSpectral Regression Discriminant Analysis(SRDA). By using spectral graph analysis, SRDA casts discriminant analysis into a regression framework that facilitates both efficient computation and the use of regularization techniques. Specifically, SRDA only needs to solve a set of regularized least squares problems, and there is no eigenvector computation involved, which is a huge save of both time and memory. Our theoretical analysis shows that SRDA can be computed with O(mn) time and O(ms) memory, where .s(les n) is the average number of nonzero features in each sample. Extensive experimental results on four real-world data sets demonstrate the effectiveness and efficiency of our algorithm. Deng Cai 0001, Xiaofei He 0001, Jiawei Han 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2008 | Learning a Maximum Margin Subspace for Image RetrievalabstractOne of the fundamental problems in Content-Based Image Retrieval (CBIR) has been the gap between low-level visual features and high-level semantic concepts. To narrow down this gap, relevance feedback is introduced into image retrieval. With the user-provided information, a classifier can be learned to distinguish between positive and negative examples. However, in real-world applications, the number of user feedbacks is usually too small compared to the dimensionality of the image space. In order to cope with the high dimensionality, we propose a novel semisupervised method for dimensionality reduction called Maximum Margin Projection (MMP). MMP aims at maximizing the margin between positive and negative examples at each local neighborhood. Different from traditional dimensionality reduction algorithms such as Principal Component Analysis (PCA) and Linear Discriminant Analysis (LDA), which effectively see only the global euclidean structure, MMP is designed for discovering the local manifold structure. Therefore, MMP is likely to be more suitable for image retrieval, where nearest neighbor search is usually involved. After projecting the images into a lower dimensional subspace, the relevant images get closer to the query image; thus, the retrieval performance can be enhanced. The experimental results on Corel image database demonstrate the effectiveness of our proposed algorithm. Xiaofei He 0001, Deng Cai 0001, Jiawei Han 0001 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2007 | Isometric Projection
Deng Cai 0001, Xiaofei He 0001, Jiawei Han 0001 |
AAAI | 1 |
| 2007 | Regularized locality preserving indexing via spectral regressionabstractWe consider the problem of document indexing and representation. Recently, Locality Preserving Indexing (LPI) was proposed for learning a compact document subspace. Different from Latent Semantic Indexing (LSI) which is optimal in the sense of global Euclidean structure, LPI is optimal in the sense of local manifold structure. However, LPI is not efficient in time and memory which makes it difficult to be applied to very large data set. Specifically, the computation of LPI involves eigen-decompositions of two dense matrices which is expensive. In this paper, we propose a new algorithm called Regularized Locality Preserving Indexing (RLPI). Benefit from recent progresses on spectral graph analysis, we cast the original LPI algorithm into a regression framework which enable us to avoid eigen-decomposition of dense matrices. Also, with the regression based framework, different kinds of regularizers can be naturally incorporated into our algorithm which makes it more flexible. Extensive experimental results show that RLPI obtains similar or better results comparing to LPI and it is significantly faster, which makes it an efficient and effective data preprocessing method for large scale text clustering, classification and retrieval. Deng Cai 0001, Xiaofei He 0001, Wei Vivian Zhang, Jiawei Han 0001 |
CIKM | 1 |
| 2007 | Learning a Spatially Smooth Subspace for Face RecognitionabstractSubspace learning based face recognition methods have attracted considerable interests in recently years, including principal component analysis (PCA), linear discriminant analysis (LDA), locality preserving projection (LPP), neighborhood preserving embedding (NPE), marginal fisher analysis (MFA) and local discriminant embedding (LDE). These methods consider an n1timesn2image as a vector in Rn1timesn2and the pixels of each image are considered as independent. While an image represented in the plane is intrinsically a matrix. The pixels spatially close to each other may be correlated. Even though we have n1xn2pixels per image, this spatial correlation suggests the real number of freedom is far less. In this paper, we introduce a regularized subspace learning model using a Laplacian penalty to constrain the coefficients to be spatially smooth. All these existing subspace learning algorithms can fit into this model and produce a spatially smooth subspace which is better for image representation than their original version. Recognition, clustering and retrieval can be then performed in the image subspace. Experimental results on face recognition demonstrate the effectiveness of our method. Deng Cai 0001, Xiaofei He 0001, Yuxiao Hu 0001, Jiawei Han 0001, Thomas S. Huang |
CVPR | 1 |
| 2007 | Spectral Regression for Efficient Regularized Subspace LearningabstractSubspace learning based face recognition methods have attracted considerable interests in recent years, including principal component analysis (PCA), linear discriminant analysis (LDA), locality preserving projection (LPP), neighborhood preserving embedding (NPE) and marginal Fisher analysis (MFA). However, a disadvantage of all these approaches is that their computations involve eigen- decomposition of dense matrices which is expensive in both time and memory. In this paper, we propose a novel dimensionality reduction framework, called spectral regression (SR), for efficient regularized subspace learning. SR casts the problem of learning the projective functions into a regression framework, which avoids eigen-decomposition of dense matrices. Also, with the regression based framework, different kinds of regularizes can be naturally incorporated into our algorithm which makes it more flexible. Computational analysis shows that SR has only linear-time complexity which is a huge speed up comparing to the cubic-time complexity of the ordinary approaches. Experimental results on face recognition demonstrate the effectiveness and efficiency of our method. Deng Cai 0001, Xiaofei He 0001, Jiawei Han 0001 |
ICCV | 1 |
| 2007 | Semi-supervised Discriminant AnalysisabstractLinear Discriminant Analysis (LDA) has been a popular method for extracting features which preserve class separability. The projection vectors are commonly obtained by maximizing the between class covariance and simultaneously minimizing the within class covariance. In practice, when there is no sufficient training samples, the covariance matrix of each class may not be accurately estimated. In this paper, we propose a novel method, called Semi- supervised Discriminant Analysis (SDA), which makes use of both labeled and unlabeled samples. The labeled data points are used to maximize the separability between different classes and the unlabeled data points are used to estimate the intrinsic geometric structure of the data. Specifically, we aim to learn a discriminant function which is as smooth as possible on the data manifold. Experimental results on single training image face recognition and relevance feedback image retrieval demonstrate the effectiveness of our algorithm. Deng Cai 0001, Xiaofei He 0001, Jiawei Han 0001 |
ICCV | 1 |
| 2007 | Spectral Regression: A Unified Approach for Sparse Subspace LearningabstractRecently the problem of dimensionality reduction (or, subspace learning) has received a lot of interests in many fields of information processing, including data mining, information retrieval, and pattern recognition. Some popular methods include principal component analysis (PCA), linear discriminant analysis (LDA) and locality preserving projection (LPP). However, a disadvantage of all these approaches is that the learned projective functions are linear combinations of all the original features, thus it is often difficult to interpret the results. In this paper, we propose a novel dimensionality reduction framework, calledUnifiedSparseSubspaceLearning(USSL), for learning sparse projections. USSL casts the problem of learning the projective functions into a regression framework, which facilitates the use of different kinds of regularizes. By using a L1-norm regularizer (lasso), the sparse projections can be efficiently computed. Experimental results on real world classification and clustering problems demonstrate the effectiveness of our method. Deng Cai 0001, Xiaofei He 0001, Jiawei Han 0001 |
ICDM | 1 |
| 2007 | Efficient Kernel Discriminant Analysis via Spectral RegressionabstractLinear discriminant analysis (LDA) has been a popular method for extracting features which preserve class separability. The projection vectors are commonly obtained by maximizing the between class covariance and simultaneously minimizing the within class covariance. LDA can be performed either in the original input space or in the reproducing kernel Hilbert space (RKHS) into which data points are mapped, which leads to Kernel Discriminant Analysis (KDA). When the data are highly nonlinear distributed, KDA can achieve better performance than LDA. However, computing the projective functions in KDA involves eigen-decomposition of kernel matrix, which is very expensive when a large number of training samples exist. In this paper, we present a new algorithm for kernel discriminant analysis, called spectral regression kernel discriminant analysis (SRKDA). By using spectral graph analysis, SRKDA casts discriminant analysis into a regression framework which facilitates both efficient computation and the use of regularization techniques. Specifically, SRKDA only needs to solve a set of regularized regression problems and there is no eigenvector computation involved, which is a huge save of computational cost. Our computational analysis shows that SRKDA is 27 times faster than the ordinary KDA. Moreover, the new formulation makes it very easy to develop incremental version of the algorithm which can fully utilize the computational results of the existing training samples. Experiments on face recognition demonstrate the effectiveness and efficiency of the proposed algorithm. Deng Cai 0001, Xiaofei He 0001, Jiawei Han 0001 |
ICDM | 1 |
| 2007 | Locality Sensitive Discriminant Analysis
Deng Cai 0001, Xiaofei He 0001, Kun Zhou 0001, Jiawei Han 0001, Hujun Bao |
IJCAI | 1 |
| 2007 | Spectral regression: a unified subspace learning framework for content-based image retrievalabstractRelevance feedback is a well established and effective framework for narrowing down the gap between low-level visual features and high-level semantic concepts in content-based image retrieval. In most of traditional implementations of relevance feedback, a distance metric or a classifier is usually learned from user's provided negative and positive examples. However, due to the limitation of the user's feedbacks and the high dimensionality of the feature space, one is often confront with the issue of the curse of the dimensionality. Recently, several researchers have considered manifold ways to address this issue, such as Locality Preserving Projections, Augmented Relation Embedding, and Semantic Subspace Projection. In this paper, by using techniques from spectral graph embedding and regression, we propose a unified framework, called spectral regression, for learning an image subspace. This framework facilitates the analysis of the differences and connections between the algorithms mentioned above. And more crucially, it provides much faster computation and therefore makes the retrieval system capable of responding to the user's query more efficiently. Deng Cai 0001, Xiaofei He 0001, Jiawei Han 0001 |
ACM Multimedia | 1 |
| 2007 | Laplacian optimal design for image retrievalabstractRelevance feedback is a powerful technique to enhance Content-Based Image Retrieval (CBIR) performance. It solicits the user’s relevance judgments on the retrieved images returned by the CBIR systems. The user’s labeling is then used to learn a classifier to distinguish between relevant and irrelevant images. However, the top returned images may not be the most informative ones. The challenge is thus to determine which unlabeled images would be the most informative (i.e., improve the classifier the most) if they were labeled and used as training samples. In this paper, we propose a novel active learning algorithm, called Laplacian Optimal Design (LOD), for relevance feedback image retrieval. Our algorithm is based on a regression model which minimizes the least square error on the measured (or, labeled) images and simultaneously preserves the local geometrical structure of the image space. Specifically, we assume that if two images are sufficiently close to each other, then their measurements (or, labels) are close as well. By constructing a nearest neighbor graph, the geometrical structure of the image space can be described by the graph Laplacian. We discuss how results from the field of optimal experimental design may be used to guide our selection of a subset of images, which gives us the most amount of information. Experimental results on Corel database suggest that the proposed approach achieves higher precision in relevance feedback image retrieval. Xiaofei He 0001, Wanli Min, Deng Cai 0001, Kun Zhou 0001 |
SIGIR | 3 |
| 2007 | Clustering and searching WWW images using link and page layout analysisabstractDue to the rapid growth of the number of digital images on the Web, there is an increasing demand for an effective and efficient method for organizing and retrieving the available images. This article describes iFind, a system for clustering and searching WWW images. By using a vision-based page segmentation algorithm, a Web page is partitioned into blocks, and the textual and link information of an image can be accurately extracted from the block containing that image. The textual information is used for image indexing. By extracting the page-to-block, block-to-image, block-to-page relationships through link structure and page layout analysis, we construct an image graph. Our method is less sensitive to noisy links than previous methods like PageRank, HITS, and PicASHOW, and hence the image graph can better reflect the semantic relationship between images. Using the notion of Markov Chain, we can compute the limiting probability distributions of the images, ImageRanks, which characterize the importance of the images. The ImageRanks are combined with the relevance scores to produce the final ranking for image search. With the graph models, we can also use techniques from spectral graph theory for image clustering and embedding, or 2-D visualization. Some experimental results on 11.6 million images downloaded from the Web are provided in the article. Xiaofei He 0001, Deng Cai 0001, Ji-Rong Wen, Wei-Ying Ma, HongJiang Zhang |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2006 | Tensor space model for document analysisabstractVector Space Model (VSM) has been at the core of information retrieval for the past decades. VSM considers the documents as vectors in high dimensional space.In such a vector space, techniques like Latent Semantic Indexing (LSI), Support Vector Machines (SVM), Naive Bayes, etc., can be then applied for indexing and classification. However, in some cases, the dimensionality of the document space might be extremely large, which makes these techniques infeasible due to the curse of dimensionality. In this paper, we propose a novel Tensor Space Model for document analysis. We represent documents as the second order tensors, or matrices. Correspondingly, a novel indexing algorithm called Tensor Latent Semantic Indexing (TensorLSI) is developed in the tensor space. Our theoretical analysis shows that TensorLSI is much more computationally efficient than the conventional Latent Semantic Indexing, which makes it applicable for extremely large scale data set. Several experimental results on standard document data sets demonstrate the efficiency and effectiveness of our algorithm. Deng Cai 0001, Xiaofei He 0001, Jiawei Han 0001 |
SIGIR | 1 |
| 2006 | An algorithm for semi-supervised learning in image retrieval
Ke Lu 0001, Jidong Zhao, Deng Cai 0001 |
Pattern Recognit. | 3 |
| 2006 | Orthogonal Laplacianfaces for Face RecognitionabstractFollowing the intuition that the naturally occurring face data may be generated by sampling a probability distribution that has support on or near a submanifold of ambient space, we propose an appearance-based face recognition method, called orthogonal Laplacianface. Our algorithm is based on the locality preserving projection (LPP) algorithm, which aims at finding a linear approximation to the eigenfunctions of the Laplace Beltrami operator on the face manifold. However, LPP is nonorthogonal, and this makes it difficult to reconstruct the data. The orthogonal locality preserving projection (OLPP) method produces orthogonal basis functions and can have more locality preserving power than LPP. Since the locality preserving power is potentially related to the discriminating power, the OLPP is expected to have more discriminating power than LPP. Experimental results on three face databases demonstrate the effectiveness of our proposed algorithm. Deng Cai 0001, Xiaofei He 0001, Jiawei Han 0001, HongJiang Zhang |
IEEE Trans. Image Process. | 1 |
| 2005 | Neighborhood Preserving EmbeddingabstractRecently there has been a lot of interest in geometrically motivated approaches to data analysis in high dimensional spaces. We consider the case where data is drawn from sampling a probability distribution that has support on or near a submanifold of Euclidean space. In this paper, we propose a novel subspace learning algorithm called neighborhood preserving embedding (NPE). Different from principal component analysis (PCA) which aims at preserving the global Euclidean structure, NPE aims at preserving the local neighborhood structure on the data manifold. Therefore, NPE is less sensitive to outliers than PCA. Also, comparing to the recently proposed manifold learning algorithms such as Isomap and locally linear embedding, NPE is defined everywhere, rather than only on the training data points. Furthermore, NPE may be conducted in the original space or in the reproducing kernel Hilbert space into which data points are mapped. This gives rise to kernel NPE. Several experiments on face database demonstrate the effectiveness of our algorithm. Xiaofei He 0001, Deng Cai 0001, Shuicheng Yan, HongJiang Zhang |
ICCV | 2 |
| 2005 | Statistical and computational analysis of locality preserving projectionabstractRecently, several manifold learning algorithms have been proposed, such as ISOMAP (Tenenbaum et al., 2000), Locally Linear Embedding (Roweis & Saul, 2000), Laplacian Eigenmap (Belkin & Niyogi, 2001), Locality Preserving Projection (LPP) (He & Niyogi, 2003), etc. All of them aim at discovering the meaningful low dimensional structure of the data space. In this paper, we present a statistical analysis of the LPP algorithm. Different from Principal Component Analysis (PCA) which obtains a subspace spanned by the largest eigenvectors of the global covariance matrix, we show that LPP obtains a subspace spanned by the smallest eigenvectors of the local covariance matrix. We applied PCA and LPP to real world document clustering task. Experimental results show that the performance can be significantly improved in the subspace, and especially LPP works much better than PCA. Xiaofei He 0001, Deng Cai 0001, Wanli Min |
ICML | 2 |
| 2005 | Image clustering with tensor representationabstractWe consider the problem of image representation and clustering. Traditionally, an n1 x n2 image is represented by a vector in the Euclidean space ℝ n1 x n2. Some learning algorithms are then applied to these vectors in such a high dimensional space for dimensionality reduction, classification, and clustering. However, an image is intrinsically a matrix, or the second order tensor. The vector representation of the images ignores the spatial relationships between the pixels in an image. In this paper, we introduce a tensor framework for image analysis. We represent the images as points in the tensor space Rn1 mathcal Rn2 which is a tensor product of two vector spaces. Based on the tensor representation, we propose a novel image representation and clustering algorithm which explicitly considers the manifold structure of the tensor space. By preserving the local structure of the data manifold, we can obtain a tensor subspace which is optimal for data representation in the sense of local isometry. We call it TensorImage approach. Traditional clustering algorithm such as k-means is then applied in the tensor subspace. Our algorithm shares many of the data representation and clustering properties of other techniques such as Locality Preserving Projections, Laplacian Eigenmaps, and spectral clustering, yet our algorithm is much more computationally efficient. Experimental results show the efficiency and effectiveness of our algorithm. Xiaofei He 0001, Deng Cai 0001, Haifeng Liu 0001, Jiawei Han 0001 |
ACM Multimedia | 2 |
| 2005 | Laplacian Score for Feature SelectionabstractIn supervised learning scenarios, feature selection has been studied widely in the literature. Selecting features in unsupervised learning scenarios is a much harder problem, due to the absence of class labels that would guide the search for relevant information. And, almost all of previous unsupervised feature selection methods are "wrapper" techniques that require a learning algorithm to evaluate the candidate feature subsets. In this paper, we propose a "filter" method for feature selection which is independent of any learning algorithm. Our method can be performed in either supervised or unsupervised fashion. The proposed method is based on the observation that, in many real world classification problems, data from the same class are often close to each other. The importance of a feature is evaluated by its power of locality preserving, or, Laplacian Score. We compare our method with data variance (unsupervised) and Fisher score (supervised) on two data sets. Experimental results demonstrate the effectiveness and efficiency of our algorithm. Xiaofei He 0001, Deng Cai 0001, Partha Niyogi |
NIPS | 2 |
| 2005 | Tensor Subspace AnalysisabstractPrevious work has demonstrated that the image variations of many ob- jects (human faces in particular) under variable lighting can be effec- tively modeled by low dimensional linear spaces. The typical linear sub- space learning algorithms include Principal Component Analysis (PCA), Linear Discriminant Analysis (LDA), and Locality Preserving Projec- tion (LPP). All of these methods consider an n1 × n2 image as a high dimensional vector in Rn1×n2, while an image represented in the plane is intrinsically a matrix. In this paper, we propose a new algorithm called Tensor Subspace Analysis (TSA). TSA considers an image as the sec- ond order tensor in Rn1 ⊗ Rn2, where Rn1 and Rn2 are two vector spaces. The relationship between the column vectors of the image ma- trix and that between the row vectors can be naturally characterized by TSA. TSA detects the intrinsic local geometrical structure of the tensor space by learning a lower dimensional tensor subspace. We compare our proposed approach with PCA, LDA and LPP methods on two standard databases. Experimental results demonstrate that TSA achieves better recognition rate, while being much more efficient. Xiaofei He 0001, Deng Cai 0001, Partha Niyogi |
NIPS | 2 |
| 2005 | Community Mining from Multi-relational Networks
Deng Cai 0001, Zheng Shao, Xiaofei He 0001, Xifeng Yan, Jiawei Han 0001 |
PKDD | 1 |
| 2005 | Orthogonal locality preserving indexingabstractWe consider the problem of document indexing and representation. Recently, Locality Preserving Indexing (LPI) was proposed for learning a compact document subspace. Different from Latent Semantic Indexing which is optimal in the sense of global Euclidean structure, LPI is optimal in the sense of local manifold structure. However, LPI is extremely sensitive to the number of dimensions. This makes it difficult to estimate the intrinsic dimensionality, while inaccurately estimated dimensionality would drastically degrade its performance. One reason leading to this problem is that LPI is non-orthogonal. Non-orthogonality distorts the metric structure of the document space. In this paper, we propose a new algorithm called Orthogonal LPI. Orthogonal LPI iteratively computes the mutually orthogonal basis functions which respect the local geometrical structure. Moreover, our empirical study shows that OLPI can have more locality preserving power than LPI. We compare the new algorithm to LSI and LPI. Extensive experimental results show that Orthogonal LPI obtains better performance than both LSI and LPI. More crucially, it is insensitive to the number of dimensions, which makes it an efficient data preprocessing method for text clustering, classification, retrieval, etc. Deng Cai 0001, Xiaofei He 0001 |
SIGIR | 1 |
| 2005 | Document Clustering Using Locality Preserving IndexingabstractWe propose a novel document clustering method which aims to cluster the documents into different semantic classes. The document space is generally of high dimensionality and clustering in such a high dimensional space is often infeasible due to the curse of dimensionality. By using locality preserving indexing (LPI), the documents can be projected into a lower-dimensional semantic space in which the documents related to the same semantics are close to each other. Different from previous document clustering methods based on latent semantic indexing (LSI) or nonnegative matrix factorization (NMF), our method tries to discover both the geometric and discriminating structures of the document space. Theoretical analysis of our method shows that LPI is an unsupervised approximation of the supervised linear discriminant analysis (LDA) method, which gives the intuitive motivation of our method. Extensive experimental evaluations are performed on the Reuters-21578 and TDT2 data sets. Deng Cai 0001, Xiaofei He 0001, Jiawei Han 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2004 | Organizing WWW images based on the analysis of page layout and Web link structureabstractDue to the rapid growth of the number of digital images on the Web, there is an increasing demand for an effective and efficient method of organizing and retrieving the images available. This paper describes a method for clustering and embedding WWW images. By using a vision-based page segmentation algorithm, a Web page is partitioned into blocks, and the textual and link information of an image can be accurately extracted from the block containing that image. By extracting the page-to-block, block-to-image, block-to-page relationships through a link structure and page layout analysis, we construct an image graph. With the image graph model, we use techniques from spectral graph theory for image clustering and embedding. Some experimental results are given in the paper. Deng Cai 0001, Xiaofei He 0001, Wei-Ying Ma, Ji-Rong Wen, HongJiang Zhang |
ICME | 1 |
| 2004 | Hierarchical clustering of WWW image search results using visual, textual and link informationabstractWe consider the problem of clustering Web image search results. Generally, the image search results returned by an image search engine contain multiple topics. Organizing the results into different semantic clusters facilitates users' browsing. In this paper, we propose a hierarchical clustering method using visual, textual and link analysis. By using a vision-based page segmentation algorithm, a web page is partitioned into blocks, and the textual and link information of an image can be accurately extracted from the block containing that image. By using block-level link analysis techniques, an image graph can be constructed. We then apply spectral techniques to find a Euclidean embedding of the images which respects the graph structure. Thus for each image, we have three kinds of representations, i.e. visual feature based representation, textual feature based representation and graph based representation. Using spectral clustering techniques, we can cluster the search results into different semantic clusters. An image search example illustrates the potential of these techniques. Deng Cai 0001, Xiaofei He 0001, Zhiwei Li 0006, Wei-Ying Ma, Ji-Rong Wen |
ACM Multimedia | 1 |
| 2004 | Locality preserving clustering for image databaseabstractIt is important and challenging to make the growing image repositories easy to search and browse. Image clustering is a technique that helps in several ways, including image data preprocessing, user interface designing, and search result representation. Spectral clustering method has been one of the most promising clustering methods in the last few years, because it can cluster data with complex structure, and the (near) global optimum is guaranteed. However, existing spectral clustering algorithms, like Normalized Cut, are difficult to handle data points out of training set. In this paper, we propose a clustering algorithm named Locality Preserving Clustering (LPC), which shares many of the data representation properties of nonlinear spectral method. Yet LPC provides an explicit mapping function which is defined everywhere, both on training data points and testing points. Experimental results show that LPC is more accurate than both "direct Kmeans" and "PCA + Kmeans". We also show that LPC produces in general comparable results with Normalized Cut, yet is more efficient than Normalized Cut. Deng Cai 0001, Xiaofei He 0001, Wei-Ying Ma, Xueyin Lin |
ACM Multimedia | 2 |
| 2004 | Block-level link analysisabstractLink Analysis has shown great potential in improving the performance of web search. PageRank and HITS are two of the most popular algorithms. Most of the existing link analysis algorithms treat a web page as a single node in the web graph. However, in most cases, a web page contains multiple semantics and hence the web page might not be considered as the atomic node. In this paper, the web page is partitioned into blocks using the vision-based page segmentation algorithm. By extracting the page-to-block, block-to-page relationships from link structure and page layout analysis, we can construct a semantic graph over the WWW such that each node exactly represents a single semantic topic. This graph can better describe the semantic structure of the web. Based on block-level link analysis, we proposed two new algorithms, Block Level PageRank and Block Level HITS, whose performances we study extensively using web data. Deng Cai 0001, Xiaofei He 0001, Ji-Rong Wen, Wei-Ying Ma |
SIGIR | 1 |
| 2004 | Block-based web searchabstractMultiple-topic and varying-length of web pages are two negative factors significantly affecting the performance of web search. In this paper, we explore the use of page segmentation algorithms to partition web pages into blocks and investigate how to take advantage of block-level evidence to improve retrieval performance in the web context. Because of the special characteristics of web pages, different page segmentation method will have different impact on web search performance. We compare four types of methods, including fixed-length page segmentation, DOM-based page segmentation, vision-based page segmentation, and a combined method which integrates both semantic and fixed-length properties. Experiments on block-level query expansion and retrieval are performed. Among the four approaches, the combined method achieves the best performance for web search. Our experimental results also show that such a semantic partitioning of web pages effectively deals with the problem of multiple drifting topics and mixed lengths, and thus has great potential to boost up the performance of current web search engines. Deng Cai 0001, Shipeng Yu, Ji-Rong Wen, Wei-Ying Ma |
SIGIR | 1 |
| 2004 | Locality preserving indexing for document representationabstractDocument representation and indexing is a key problem for document analysis and processing, such as clustering, classification and retrieval. Conventionally, Latent Semantic Indexing (LSI) is considered effective in deriving such an indexing. LSI essentially detects the most representative features for document representation rather than the most discriminative features. Therefore, LSI might not be optimal in discriminating documents with different semantics. In this paper, a novel algorithm called Locality Preserving Indexing (LPI) is proposed for document indexing. Each document is represented by a vector with low dimensionality. In contrast to LSI which discovers the global structure of the document space, LPI discovers the local structure and obtains a compact document representation subspace that best detects the essential semantic structure. We compare the proposed LPI approach with LSI on two standard databases. Experimental results show that LPI provides better representation in the sense of semantic structure. Xiaofei He 0001, Deng Cai 0001, Haifeng Liu 0001, Wei-Ying Ma |
SIGIR | 2 |
| 2003 | Extracting Content Structure for Web Pages Based on Visual Representation
Deng Cai 0001, Shipeng Yu, Ji-Rong Wen, Wei-Ying Ma |
APWeb | 1 |
| 2003 | Improving pseudo-relevance feedback in web information retrieval using web page segmentationabstractIn contrast to traditional document retrieval, a web page as a whole is not a good information unit to search because it often contains multiple topics and a lot of irrelevant information from navigation, decoration, and interaction part of the page. In this paper, we propose a VIsion-based Page Segmentation (VIPS) algorithm to detect the semantic content structure in a web page. Compared with simple DOM based segmentation method, our page segmentation scheme utilizes useful visual cues to obtain a better partition of a page at the semantic level. By using our VIPS algorithm to assist the selection of query expansion terms in pseudo-relevance feedback in web information retrieval, we achieve 27% performance improvement on Web Track dataset. Shipeng Yu, Deng Cai 0001, Ji-Rong Wen, Wei-Ying Ma |
WWW | 2 |