VLDB 2026 Research / reviewers in the wild / expert
Lin Ma 0002
dblp:74/3608-2
· DBLP profile ↗
188ranked-venue papers
23as first author
84since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 139 · 14 first-author · 55 since 2021Artificial intelligence and machine learning · 107 · 5 first-author · 60 since 2021Systems, architecture and hardware · 7 · 5 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-authorComputer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DBGroup: Dual-Branch Point Grouping for Weakly Supervised 3D Semantic Instance SegmentationabstractWeakly supervised 3D instance segmentation is essential for 3D scene understanding, especially as the growing scale of data and high annotation costs associated with fully supervised approaches. Existing methods primarily rely on two forms of weak supervision: one-thing-one-click annotations and bounding box annotations, both of which aim to reduce labeling efforts. However, these approaches still encounter limitations, including labor-intensive annotation processes, high complexity, and reliance on expert annotators. To address these challenges, we propose DBGroup, a two-stage weakly supervised 3D instance segmentation framework that leverages scene-level annotations as a more efficient and scalable alternative. In the first stage, we introduce a Dual-Branch Point Grouping module to generate pseudo labels guided by semantic and mask cues extracted from multi-view images. To further improve label quality, we develop two refinement strategies: Granularity-Aware Instance Merging and Semantic Selection and Propagation. The second stage involves multi-round self-training on an end-to-end instance segmentation network using the refined pseudo-labels. Additionally, we introduce an Instance Mask Filter strategy to address inconsistencies within the pseudo labels. Extensive experiments demonstrate that DBGroup achieves competitive performance compared to sparse-point-level supervised 3D instance segmentation methods, while surpassing state-of-the-art scene-level supervised 3D semantic segmentation approaches. Xuexun Liu, Xiaoxu Xu, Qiudan Zhang, Lin Ma 0002, Xu Wang 0006 |
AAAI | 4 |
| 2026 | X-SAM: From Segment Anything to Any SegmentationabstractLarge Language Models (LLMs) demonstrate strong capabilities in broad knowledge representation, yet they are inherently deficient in pixel-level perceptual understanding. Although the Segment Anything Model (SAM) represents a significant advancement in visual-prompt-driven image segmentation, it exhibits notable limitations in multi-mask prediction and category-specific segmentation tasks, and it cannot integrate all segmentation tasks within a unified model architecture. To address these limitations, we present X-SAM, a streamlined Multimodal Large Language Model (MLLM) framework that extends the segmentation paradigm from segment anything to any segmentation. Specifically, we introduce a novel unified framework that enables more advanced pixel-level perceptual comprehension for MLLMs. Furthermore, we propose a new segmentation task, termed Visual GrounDed (VGD) segmentation, which segments all instance objects with interactive visual prompts and empowers MLLMs with visual grounded, pixel-wise interpretative capabilities. To enable effective training on diverse data sources, we present a unified training strategy that supports co-training across multiple datasets. Experimental results demonstrate that X-SAM achieves state-of-the-art performance on a wide range of image segmentation benchmarks, highlighting its efficiency for multimodal, pixel-level visual understanding. Hao Wang 0050, Limeng Qiao, Zequn Jie, Chengjian Feng, Lin Ma 0002, Xiangyuan Lan, Xiaodan Liang |
AAAI | 7 |
| 2026 | Contrastive Mean Teacher for Robust Low-Light Image Enhancement
Zhangkai Ni, Menglin Han, Wenhan Yang, Hanli Wang, Lin Ma 0002, Sam Kwong |
Int. J. Comput. Vis. | 5 |
| 2026 | A Visual-Linguistic Approach for Robust RGB-Thermal Tracking With Dynamic Template AdaptationabstractRGB-T tracking seeks to improve tracking robust ness in complex environments by exploiting the complementary advantages of RGB and thermal infrared (TIR) modalities. Nevertheless, current RGB-T tracking methods encounter two critical limitations stemming from their reliance on fixed template images for search region matching. First, the resemblance between the initial fixed template and the target in later frames diminishes as time goes on, causing tracking reliability to decline. Second, background clutter within template bounding boxes introduces disruptive noise that compromises tracking accuracy. To address these challenges, we propose a new referring RGB-T tracking approach that integrates visual-linguistic cues to enhance multimodal data, enabling background noise elimination in the template and dynamic template image updating through adaptive mechanisms. Additionally, to facilitate research in language guided multimodal tracking, we construct a large-scale dataset named Refer-RGBT, which comprises 1508 multimodal video pairs with synchronized RGB/TIR sequences of diverse objects. Each sequence is annotated by descriptive textual captions that detail their appearance and actions. Extensive evaluations on the Refer-RGBT dataset demonstrate that our proposed approach achieves cutting-edge performance compared to the state-of-the art (SOTA) methods. Xu Wang 0006, Huanxin Zheng, Haohong Liao, Qiudan Zhang, Lin Ma 0002, Jianmin Jiang |
IEEE Trans. Multim. | 5 |
| 2025 | Affordances-Oriented Planning Using Foundation Models for Continuous Vision-Language NavigationabstractLLM-based agents have demonstrated impressive zero-shot performance in vision-language navigation (VLN) task. However, existing LLM-based methods often focus only on solving high-level task planning by selecting nodes in predefined navigation graphs for movements, overlooking low-level control in navigation scenarios. To bridge this gap, we propose AO-Planner, a novel Affordances-Oriented Planner for continuous VLN task. Our AO-Planner integrates various foundation models to achieve affordances-oriented low-level motion planning and high-level decision-making, both performed in a zero-shot setting. Specifically, we employ a Visual Affordances Prompting (VAP) approach, where the visible ground is segmented by SAM to provide navigational affordances, based on which the LLM selects potential candidate waypoints and plans low-level paths towards selected waypoints. We further propose a high-level PathAgent which marks planned paths into the image input and reasons the most probable path by comprehending all environmental information. Finally, we convert the selected path into 3D coordinates using camera intrinsic parameters and depth information, avoiding challenging 3D predictions for LLMs. Experiments on the challenging R2R-CE and RxR-CE datasets show that AO-Planner achieves state-of-the-art zero-shot performance (8.8% improvement on SPL). Our method can also serve as a data annotator to obtain pseudo-labels, distilling its waypoint prediction ability into a learning-based predictor. This new predictor does not require any waypoint data from the simulator and achieves 47% SR competing with supervised methods. We establish an effective connection between LLM and 3D world, presenting novel prospects for employing foundation models in low-level motion control. Bingqian Lin, Xinmin Liu, Lin Ma 0002, Xiaodan Liang, Kwan-Yee Kenneth Wong |
AAAI | 4 |
| 2025 | Towards Efficient Foundation Model for Zero-shot Amodal SegmentationabstractAiming to predict the complete shape of partially occluded objects, amodal segmentation is an important capacity towards visual intelligence. In order to promote the practicability, zero-shot foundation model competent for the open world gains growing attention in this field. Nevertheless, prior models exhibit deficiencies in efficiency and stability. To address this problem, utilizing the implicit prior knowledge, we propose the first SAM-based amodal segmentation foundation model, SAMBA. Methodologically, a novel framework with multilevel facilitation is designed to better adapt the task characteristics and unleash the potential capabilities of SAM. In the modality level, a separation-to-fusion structure is employed that jointly learns modal and amodal segmentation to enhance mutual coordination. In the instance level, to ease the complexity of amodal feature extraction, we introduce a principal focusing mechanism to indicate objects of interest. In the pixel level, mixture-of-experts is incorporated with a specialized distribution loss, by which distinct occlusion rates correspond to different experts to improve the accuracy. Experiments are conducted on several eminent datasets, and the results show that the performance of SAMBA is superior to existing zero-shot and even supervised approaches. Furthermore, our proposed model has notable advantages in terms of speed and size. Zhaochen Liu, Limeng Qiao, Xiangxiang Chu, Lin Ma 0002, Tingting Jiang 0001 |
CVPR | 4 |
| 2025 | RoboTron-Drive: All-in-One Large Multimodal Model for Autonomous Driving
Chengjian Feng, Baihui Xiao, Zequn Jie, Xiaodan Liang, Lin Ma 0002 |
ICCV | 8 |
| 2025 | RoboTron-Sim: Improving Real-World Driving via Simulated Hard-Case
Baihui Xiao, Chengjian Feng, Lin Ma 0002 |
ICCV | 6 |
| 2025 | RoboTron-Mani: All-in-One Multimodal Large Model for Robotic Manipulation
Fanfan Liu, Yiyang Huang 0002, Zechao Guan, Yufeng Zhong 0001, Chengjian Feng, Lin Ma 0002 |
ICCV | 8 |
| 2025 | DisTime: Distribution-Based Time Representation for Video Large Language ModelsabstractDespite advances in general video understanding, Video Large Language Models (Video-LLMs) face challenges in precise temporal localization due to discrete time representations and limited temporally aware datasets. Existing methods for temporal expression either conflate time with text-based numerical values, add a series of dedicated temporal tokens, or regress time using specialized temporal grounding heads. To address these issues, we introduce DisTime, a lightweight framework designed to enhance temporal comprehension in Video-LLMs. DisTime employs a learnable token to create a continuous temporal embedding space and incorporates a Distribution-based Time Decoder that generates temporal probability distributions, effectively mitigating boundary ambiguities and maintaining temporal continuity. Additionally, the Distribution-based Time Encoder re-encodes timestamps to provide time markers for Video-LLMs. To overcome temporal granularity limitations in existing datasets, we propose an automated annotation paradigm that combines the captioning capabilities of Video-LLMs with the localization expertise of dedicated temporal models. This leads to the creation of InternVid-TG, a substantial dataset with 1.25M temporally grounded events across 179k videos, surpassing ActivityNet-Caption by 55 times. Extensive experiments demonstrate that DisTime achieves state-of-the-art performance across benchmarks in three time-sensitive tasks while maintaining competitive performance in Video QA tasks. Code and data are released at https://github.com/josephzpng/DisTime. Yingsen Zeng, Zepeng Huang, Chengjian Feng, Lin Ma 0002 |
ICCV | 6 |
| 2025 | RoboTron-Nav: A Unified Framework for Embodied Navigation Integrating Perception, Planning, and PredictionabstractIn language-guided visual navigation, agents locate target objects in unseen environments using natural language instructions. For reliable navigation in unfamiliar scenes, agents should possess strong perception, planning, and prediction capabilities. Additionally, when agents revisit previously explored areas during long-term navigation, they may retain irrelevant and redundant historical perceptions, leading to suboptimal results. In this work, we propose RoboTron-Nav, a unified framework that integrates perception, planning, and prediction capabilities through multitask collaborations on navigation and embodied question answering tasks, thereby enhancing navigation performances. Furthermore, RoboTron-Nav employs an adaptive 3D-aware history sampling strategy to effectively and efficiently utilize historical observations. By leveraging large language model, RoboTron-Nav comprehends diverse commands and complex visual scenes, resulting in appropriate navigation actions. RoboTron-Nav achieves an 81.1% success rate in object goal navigation on the $\mathrm{CHORES}$-$\mathbb{S}$ benchmark, setting a new state-of-the-art performance. Project page: https://yvfengzhong.github.io/RoboTron-Nav Yufeng Zhong 0001, Chengjian Feng, Fanfan Liu, Lin Ma 0002 |
ICCV | 6 |
| 2025 | CO-MOT: Boosting End-to-end Transformer-based Multi-Object Tracking via Coopetition Label Assignment and Shadow SetsabstractExisting end-to-end Multi-Object Tracking (e2e-MOT) methods have not surpassed non-end-to-end tracking-by-detection methods. One possible reason lies in the training label assignment strategy that consistently binds the tracked objects with tracking queries and assigns few newborns to detection queries. Such an assignment, with one-to-one bipartite matching, yields an unbalanced training, _i.e._, scarce positive samples for detection queries, especially for an enclosed scene with the majority of the newborns at the beginning of videos. As such, e2e-MOT will incline to generate a tracking terminal without renewal or re-initialization, compared to other tracking-by-detection methods.
To alleviate this problem, we propose **Co-MOT**, a simple yet effective method to facilitate e2e-MOT by a novel coopetition label assignment with a shadow concept. Specifically, we add tracked objects to the matching targets for detection queries when performing the label assignment for training the intermediate decoders. For query initialization, we expand each query by a set of shadow counterparts with limited disturbance to itself.
With extensive ablation studies, Co-MOT achieves superior performances without extra costs, _e.g._, 69.4% HOTA on DanceTrack and 52.8% TETA on BDD100K. Impressively, Co-MOT only requires 38% FLOPs of MOTRv2 with comparable performances, resulting in the 1.4× faster inference speed. Source code is publicly available at [GitHub](https://github.com/BingfengYan/CO-MOT). Weixin Luo, Yiyang Gan, Lin Ma 0002 |
ICLR | 5 |
| 2025 | Dadu-Corki: Algorithm-Architecture Co-Design for Embodied AI-powered Robotic ManipulationabstractEmbodied AI robots have the potential to fundamentally improve the way human beings live and manufacture.Continued progress in the burgeoning field of using large language models to control robots depends critically on an efficient computing substrate, and this trend is strongly evident in manipulation tasks.In particular, today's computing systems for embodied AI robots for manipulation tasks are designed purely based on the interest of algorithm developers, where robot actions are divided into a discrete frame basis.Such an execution pipeline creates high latency and energy consumption.This paper proposes Corki, an algorithm-architecture co-design framework for real-time embodied AI-powered robotic manipulation applications.We aim to decouple LLM inference, robotic control, and data communication in the embodied AI robots' compute pipeline.Instead of predicting action for one single frame, * equal contribution. Yiyang Huang 0002, Yuhui Hao, Bo Yu 0014, Yuxin Yang 0002, Feng Min, Yinhe Han 0001, Lin Ma 0002, Shaoshan Liu, Qiang Liu 0011, Yiming Gan |
ISCA | 8 |
| 2025 | FlexVAR: Flexible Visual Autoregressive Modeling without Residual PredictionabstractThis work challenges the residual prediction paradigm in visual autoregressive modeling and presents FlexVAR, a new Flexible Visual AutoRegressive image generation paradigm. FlexVAR facilitates autoregressive learning with ground-truth prediction, enabling each step to independently produce plausible images. This simple, intuitive approach swiftly learns visual distributions and makes the generation process more flexible and adaptable. Trained solely on low-resolution images (< 256px), FlexVAR can: (1) Generate images of various resolutions and aspect ratios, even exceeding the resolution of the training images. (2) Support various image-to-image tasks, including image refinement, in/out-painting, and image expansion. (3) Adapt to various autoregressive steps, allowing for faster inference with fewer steps or enhancing image quality with more steps. Our 1.0B model outperforms its VAR counterpart on the ImageNet 256 × 256 benchmark. Moreover, when zero-shot transfer the image generation process with 13 steps, the performance further improves to 2.08 FID, outperforming state-of-the-art autoregressive models AiM/VAR by 0.25/0.28 FID and popular diffusion models LDM/DiT by 1.52/0.19 FID, respectively. When transferring our 1.0B model to the ImageNet 512 × 512 benchmark in a zero-shot manner, FlexVAR achieves competitive results compared to the VAR 2.3B model, which is a fully supervised model trained at 512 × 512 resolution. Siyu Jiao, Gengwei Zhang, Yinlong Qian, Jiancheng Huang, Yao Zhao 0001, Humphrey Shi, Lin Ma 0002, Yunchao Wei, Zequn Jie |
NeurIPS | 7 |
| 2025 | GUIDED: Granular Understanding via Identification, Detection, and Discrimination for Fine-Grained Open-Vocabulary Object DetectionabstractFine-grained open-vocabulary object detection (FG-OVD) aims to detect novel object categories described by attribute-rich texts. While existing open-vocabulary detectors show promise at the base-category level, they underperform in fine-grained settings due to the semantic entanglement of subjects and attributes in pretrained vision-language model (VLM) embeddings -- leading to over-representation of attributes, mislocalization, and semantic drift in embedding space. We propose GUIDED, a decomposition framework specifically designed to address the semantic entanglement between subjects and attributes in fine-grained prompts. By separating object localization and fine-grained recognition into distinct pathways, GUIDED aligns each subtask with the module best suited for its respective roles.
Specifically, given a fine-grained class name, we first use a language model to extract a coarse-grained subject and its descriptive attributes.
Then the detector is guided solely by the subject embedding, ensuring stable localization unaffected by irrelevant or overrepresented attributes. To selectively retain helpful attributes, we introduce an attribute embedding fusion module that incorporates attribute information into detection queries in an attention-based manner. This mitigates over-representation while preserving discriminative power.
Finally, a region-level attribute discrimination module compares each detected region against full fine-grained class names using a refined vision-language model with a projection head for improved alignment.
Extensive experiments on FG-OVD and 3F-OVD benchmarks show that GUIDED achieves new state-of-the-art results, demonstrating the benefits of disentangled modeling and modular optimization. Jiaming Li 0010, Zhijia Liang, Weikai Chen 0001, Lin Ma 0002, Guanbin Li |
NeurIPS | 4 |
| 2025 | Towards Better & Faster Autoregressive Image Generation: From the Perspective of EntropyabstractIn this work, we first revisit the sampling issues in current autoregressive (AR) image generation models and identify that image tokens, unlike text tokens, exhibit lower information density and non-uniform spatial distribution. Accordingly, we present an entropy-informed decoding strategy that facilitates higher autoregressive generation quality with faster synthesis speed. Specifically, the proposed method introduces two main innovations: 1) dynamic temperature control guided by spatial entropy of token distributions, enhancing the balance between content diversity, alignment accuracy, and structural coherence in both mask-based and scale-wise models, without extra computational overhead, and 2) entropy-aware acceptance rules in speculative decoding, achieving near-lossless generation at about 85% of the inference cost of conventional acceleration methods. Extensive experiments across multiple benchmarks using diverse AR image generation models demonstrate the effectiveness and generalizability of our approach in enhancing both generation quality and sampling speed. Feng Zhao 0004, Pengyang Ling, Haibo Qiu, Zhixiang Wei, Hu Yu 0001, Jie Huang 0017, Zhixiong Zeng, Lin Ma 0002 |
NeurIPS | 9 |
| 2025 | VITRIX-UniViTAR: Unified Vision Transformer with Native ResolutionabstractConventional Vision Transformer streamlines visual modeling by employing a uniform input resolution, which underestimates the inherent variability of natural visual data and incurs a cost in spatial-contextual fidelity. While preliminary explorations have superficially investigated native resolution modeling, existing works still lack systematic training recipe from the visual representation perspective. To bridge this gap, we introduce Unified Vision Transformer with Native Resolution, i.e. UniViTAR, a family of homogeneous vision foundation models tailored for unified visual modality and native resolution scenario in the era of multimodal. Our framework first conducts architectural upgrades to the vanilla paradigm by integrating multiple advanced components. Building upon these improvements, a progressive training paradigm is introduced, which strategically combines two core mechanisms: (1) resolution curriculum learning, transitioning from fixed-resolution pretraining to native resolution tuning, thereby leveraging ViT’s inherent adaptability to variable-length sequences, and (2) visual modality adaptation via inter-batch image-video switching, which balances computational efficiency with enhanced temporal reasoning. In parallel, a hybrid training framework further synergizes sigmoid-based contrastive loss with feature distillation from a frozen teacher model, thereby accelerating early-stage convergence. Finally, trained exclusively on public accessible image-caption data, our UniViTAR family across multiple model scales from 0.3B to 1B achieves state-of-the-art performance on a wide variety of visual-related tasks. The code and models are available here. Limeng Qiao, Yiyang Gan, Bairui Wang, Lin Ma 0002 |
NeurIPS | 7 |
| 2025 | VITRIX-CLIPIN: Enhancing Fine-Grained Visual Understanding in CLIP via Instruction-Editing Data and Long CaptionsabstractDespite the success of Vision-Language Models (VLMs) like CLIP in aligning vision and language, their proficiency in detailed, fine-grained visual comprehension remains a key challenge. We present CLIP-IN, a novel framework that bolsters CLIP's fine-grained perception through two core innovations. Firstly, we leverage instruction-editing datasets, originally designed for image manipulation, as a unique source of hard negative image-text pairs. Coupled with a symmetric hard negative contrastive loss, this enables the model to effectively distinguish subtle visual-semantic differences. Secondly, CLIP-IN incorporates long descriptive captions, utilizing rotary positional encodings to capture rich semantic context often missed by standard CLIP. Our experiments demonstrate that CLIP-IN achieves substantial gains on the MMVP benchmark and various fine-grained visual recognition tasks, without compromising robust zero-shot performance on broader classification and retrieval tasks. Critically, integrating CLIP-IN's visual representations into Multimodal Large Language Models significantly reduces visual hallucinations and enhances reasoning abilities. This work underscores the considerable potential of synergizing targeted, instruction-based contrastive learning with comprehensive descriptive information to elevate the fine-grained understanding of VLMs. Limeng Qiao, Lin Ma 0002 |
NeurIPS | 4 |
| 2025 | MLLM-Tool: A Multimodal Large Language Model for Tool Agent LearningabstractRecently, the astonishing performance of large language models (LLMs) in natural language comprehension and generation tasks triggered lots of exploration of using them as central controllers to build agent systems. Multiple studies focus on bridging the LLMs to external tools to extend the application scenarios. However, the current LLMs' ability to perceive tool use is limited to a single text query, which may result in ambiguity in understanding the users' real intentions. LLMs are expected to eliminate that by perceiving the information in the visual-or auditory-grounded instructions. Therefore, in this paper, we propose MLLM-Tool, a system incorporating open-source LLMs and multi-modal encoders so that the learned LLMs can be conscious of multi-modal input instruction and then select the function-matched tool correctly. To facilitate the evaluation of the model's capability, we collect a dataset featuring multi-modal input tools from HuggingFace. Another essential feature of our dataset is that it also contains multiple potential choices for the same instruction due to the existence of identical functions and synonymous functions, which provides more potential solutions for the same query. The experiments reveal that our MLLM-Tool is capable of recommending appropriate tools for multi-modal instructions. Codes and data are available at github.com/MLLM-Tool/MLLM-Tool. Weixin Luo, Sixun Dong, Xiaohua Xuan, Lin Ma 0002, Shenghua Gao |
WACV | 6 |
| 2025 | Uni-MoE: Scaling Unified Multimodal LLMs With Mixture of ExpertsabstractRecent advancements in Multimodal Large Language Models (MLLMs) underscore the significance of scalable models and data to boost performance, yet this often incurs substantial computational costs. Although the Mixture of Experts (MoE) architecture has been employed to scale large language or visual-language models efficiently, these efforts typically involve fewer experts and limited modalities. To address this, our work presents the pioneering attempt to develop a unified MLLM with the MoE architecture, named Uni-MoE that can handle a wide array of modalities. Specifically, it features modality-specific encoders with connectors for a unified multimodal representation. We also implement a sparse MoE architecture within the LLMs to enable efficient training and inference through modality-level data parallelism and expert-level model parallelism. To enhance the multi-expert collaboration and generalization, we present a progressive training strategy: 1) Cross-modality alignment using various connectors with different cross-modality data, 2) Training modality-specific experts with cross-modality instruction data to activate experts' preferences, and 3) Tuning the whole Uni-MoE framework utilizing Low-Rank Adaptation (LoRA) on mixed multimodal instruction data. We evaluate the instruction-tuned Uni-MoE on a comprehensive set of multimodal datasets. The extensive experimental results demonstrate Uni-MoE's principal advantage of significantly reducing performance bias in handling mixed multimodal datasets, alongside improved multi-expert collaboration and generalization. Yunxin Li, Shenyuan Jiang, Baotian Hu, Longyue Wang, Wanqi Zhong, Wenhan Luo, Lin Ma 0002, Min Zhang 0005 |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2025 | A Pyramid Fusion MLP for Dense PredictionabstractRecently, MLP-based architectures have achieved competitive performance with convolutional neural networks (CNNs) and vision transformers (ViTs) across various vision tasks. However, most MLP-based methods introduce local feature interactions to facilitate direct adaptation to downstream tasks, thereby lacking the ability to capture global visual dependencies and multi-scale context, ultimately resulting in unsatisfactory performance on dense prediction. This paper proposes a competitive and effective MLP-based architecture called Pyramid Fusion MLP (PFMLP) to address the above limitation. Specifically, each block in PFMLP introduces multi-scale pooling and fully connected layers to generate feature pyramids, which are subsequently fused using up-sample layers and an additional fully connected layer. Employing different down-sample rates allows us to obtain diverse receptive fields, enabling the model to simultaneously capture long-range dependencies and fine-grained cues, thereby exploiting the potential of global context information and enhancing the spatial representation power of the model. Our PFMLP is the first lightweight MLP to obtain comparable results with state-of-the-art CNNs and ViTs on the ImageNet-1K benchmark.With larger FLOPs, it exceeds state-of-the-art CNNs, ViTs, and MLPs under similar computational complexity. Furthermore, experiments in object detection, instance segmentation, and semantic segmentation demonstrate that the visual representation acquired from PFMLP can be seamlessly transferred to downstream tasks, producing competitive results. All materials contain the training codes and logs are released at https://github.com/huangqiuyu/PFMLP. Qiuyu Huang, Zequn Jie, Lin Ma 0002, Li Shen 0008, Shenqi Lai |
IEEE Trans. Image Process. | 3 |
| 2025 | Weakly-Supervised 3D Visual Grounding Based on Visual Language AlignmentabstractLearning to ground natural language queries to target objects or regions in 3D point clouds is quite essential for 3D scene understanding. Nevertheless, existing 3D visual grounding approaches require a substantial number of bounding box annotations for text queries, which is time-consuming and labor-intensive to obtain. In this paper, we propose3D-VLA, a weakly supervised approach for3Dvisual grounding based onVisualLanguageAlignment. Our 3D-VLA exploits the superior ability of current large-scale vision-language models (VLMs) on aligning the semantics between texts and 2D images, as well as the naturally existing correspondences between 2D images and 3D point clouds, and thus implicitly constructs correspondences between texts and 3D point clouds with no need for fine-grained box annotations in the training procedure. During the inference stage, the learned text-3D correspondence will help us ground the text queries to the 3D target objects even without 2D images. To the best of our knowledge, this is the first work to investigate 3D visual grounding in a weakly supervised manner by involving large scale vision-language models, and extensive experiments on ReferIt3D and ScanRefer datasets demonstrate that our 3D-VLA achieves comparable and even superior results over the fully supervised methods. Xiaoxu Xu, Yitian Yuan, Qiudan Zhang, Wenhui Wu 0001, Zequn Jie, Lin Ma 0002, Xu Wang 0006 |
IEEE Trans. Multim. | 6 |
| 2024 | Instance-Aware Multi-Camera 3D Object Detection with Structural Priors Mining and Self-Boosting LearningabstractCamera-based bird-eye-view (BEV) perception paradigm has made significant progress in the autonomous driving field. Under such a paradigm, accurate BEV representation construction relies on reliable depth estimation for multi-camera images. However, existing approaches exhaustively predict depths for every pixel without prioritizing objects, which are precisely the entities requiring detection in the 3D space. To this end, we propose IA-BEV, which integrates image-plane instance awareness into the depth estimation process within a BEV-based detector. First, a category-specific structural priors mining approach is proposed for enhancing the efficacy of monocular depth generation. Besides, a self-boosting learning strategy is further proposed to encourage the model to place more emphasis on challenging objects in computation-expensive temporal stereo matching. Together they provide advanced depth estimation results for high-quality BEV features construction, benefiting the ultimate 3D detection. The proposed method achieves state-of-the-art performances on the challenging nuScenes benchmark, and extensive experimental results demonstrate the effectiveness of our designs. Zequn Jie, Shaoxiang Chen 0001, Lechao Cheng, Jingjing Chen 0001, Lin Ma 0002, Yu-Gang Jiang 0001 |
AAAI | 6 |
| 2024 | ColNeRF: Collaboration for Generalizable Sparse Input Neural Radiance FieldabstractNeural Radiance Fields (NeRF) have demonstrated impressive potential in synthesizing novel views from dense input, however, their effectiveness is challenged when dealing with sparse input. Existing approaches that incorporate additional depth or semantic supervision can alleviate this issue to an extent. However, the process of supervision collection is not only costly but also potentially inaccurate. In our work, we introduce a novel model: the Collaborative Neural Radiance Fields (ColNeRF) designed to work with sparse input. The collaboration in ColNeRF includes the cooperation among sparse input source images and the cooperation among the output of the NeRF. Through this, we construct a novel collaborative module that aligns information from various views and meanwhile imposes self-supervised constraints to ensure multi-view consistency in both geometry and appearance. A Collaborative Cross-View Volume Integration module (CCVI) is proposed to capture complex occlusions and implicitly infer the spatial location of objects. Moreover, we introduce self-supervision of target rays projected in multiple directions to ensure geometric and color consistency in adjacent regions. Benefiting from the collaboration at the input and output ends, ColNeRF is capable of capturing richer and more generalized scene representation, thereby facilitating higher-quality results of the novel view synthesis. Our extensive experimental results demonstrate that ColNeRF outperforms state-of-the-art sparse input generalizable NeRF methods. Furthermore, our approach exhibits superiority in fine-tuning towards adapting to new scenes, achieving competitive performance compared to per-scene optimized NeRF-based methods while significantly reducing computational costs. Our code is available at: https://github.com/eezkni/ColNeRF. Zhangkai Ni, Peiqi Yang, Wenhan Yang, Hanli Wang, Lin Ma 0002, Sam Kwong |
AAAI | 5 |
| 2024 | A Multimodal In-Context Tuning Approach for E-Commerce Product Description GenerationabstractIn this paper, we propose a new setting for generating product descriptions from images, augmented by marketing keywords. It leverages the combined power of visual and textual information to create descriptions that are more tailored to the unique features of products. For this setting, previous methods utilize visual and textual encoders to encode the image and keywords and employ a language model-based decoder to generate the product description. However, the generated description is often inaccurate and generic since same-category products have similar copy-writings, and optimizing the overall framework on large-scale samples makes models concentrate on common words yet ignore the product features. To alleviate the issue, we present a simple and effective Multimodal In-Context Tuning approach, named ModICT, which introduces a similar product sample as the reference and utilizes the in-context learning capability of language models to produce the description. During training, we keep the visual encoder and language model frozen, focusing on optimizing the modules responsible for creating multimodal in-context references and dynamic prompts. This approach preserves the language generation prowess of large language models (LLMs), facilitating a substantial increase in description diversity. To assess the effectiveness of ModICT across various language model scales and types, we collect data from three distinct product categories within the E-commerce domain. Extensive experiments demonstrate that ModICT significantly improves the accuracy (by up to 3.3% on Rouge-L) and diversity (by up to 9.4% on D-5) of generated results compared to conventional methods. Our findings underscore the potential of ModICT as a valuable tool for enhancing the automatic generation of product descriptions in a wide range of applications. Data and code are at https://github.com/HITsz-TMG/Multimodal-In-Context-Tuning Yunxin Li, Baotian Hu, Wenhan Luo, Lin Ma 0002, Min Zhang 0005 |
LREC/COLING | 4 |
| 2024 | InstaGen: Enhancing Object Detection by Training on Synthetic DatasetabstractIn this paper, we present a novel paradigm to enhance the ability of object detector, e.g., expanding categories or improving detection performance, by training on syn-thetic dataset generated from diffusion models. Specifically, we integrate an instance-level grounding head into a pre-trained, generative diffusion model, to augment it with the ability of localising instances in the generated images. The grounding head is trained to align the text embedding of category names with the regional visual feature of the diffusion model, using supervision from an off-the-shelf object detector, and a novel self-training scheme on (novel) categories not covered by the detector. We conduct thorough experiments to show that, this enhanced version of diffusion model, termed as InstaGen, can serve as a data synthe-sizer, to enhance object detectors by training on its generated samples, demonstrating superior performance over existing state-of-the-art methods in open-vocabulary (+4.5 AP) and data-sparse (+ 1. 2 ~ 5.2 AP) scenarios. Chengjian Feng, Zequn Jie, Weidi Xie, Lin Ma 0002 |
CVPR | 5 |
| 2024 | AlignSAM: Aligning Segment Anything Model to Open Context via Reinforcement LearningabstractPowered by massive curated training data, Segment Any-thing Model (SAM) has demonstrated its impressive generalization capabilities in open-world scenarios with the guidance of prompts. However, the vanilla SAM is class-agnostic and heavily relies on user-provided prompts to segment objects of interest. Adapting this method to diverse tasks is crucial for accurate target identification and to avoid suboptimal segmentation results. In this paper, we propose a novel framework, termed AlignSAM, designed for automatic prompting for aligning SAM to an open context through reinforcement learning. Anchored by an agent, AlignSAM enables the generality of the SAM model across diverse downstream tasks while keeping its parameters frozen. Specifically, AlignSAM initiates a prompting agent to iteratively refine segmentation predictions by interacting with the foundational model. It integrates a reinforcement learning policy network to provide informative prompts to the foundational models. Additionally, a semantic recal-ibration module is introduced to provide fine-grained labels of prompts, enhancing the model's proficiency in handling tasks encompassing explicit and implicit semantics. Experiments conducted on various challenging segmentation tasks among existing foundation models demonstrate the superiority of the proposed AlignSAM over state-of-the-art approaches. Project page: https://github.com/Duojun-Huang/AIignSAM-CVPR2024. Duojun Huang, Xinyu Xiong, Jichang Li, Zequn Jie, Lin Ma 0002, Guanbin Li |
CVPR | 6 |
| 2024 | Misalignment-Robust Frequency Distribution Loss for Image TransformationabstractThis paper aims to address a common challenge in deep learning-based image transformation methods, such as im-age enhancement and super-resolution, which heavily rely on precisely aligned paired datasets with pixel-level align-ments. However, creating precisely aligned paired images presents significant challenges and hinders the advance-ment of methods trained on such data. To overcome this challenge, this paper introduces a novel and simple frequency Distribution Loss (FDL) for computing distribution distance within the frequency domain. Specifically, we transform image features into the frequency domain using Discrete Fourier Transformation (DFT). Subsequently, frequency components (amplitude and phase) are processed separately to form the FDL loss function. Our method is empirically proven effective as a training constraint due to the thoughtful utilization of global information in the frequency domain. Extensive experimental evaluations, fo-cusing on image enhancement and super-resolution tasks, demonstrate that FDL outperforms existing misalignment-robust loss functions. Furthermore, we explore the poten-tial of our FDL for image style transfer that relies solely on completely misaligned data. Our code is available at: https://github.com/eezkni/FDL Zhangkai Ni, Juncheng Wu, Wenhan Yang, Hanli Wang, Lin Ma 0002 |
CVPR | 6 |
| 2024 | Making Large Language Models Better Planners with Reasoning-Decision Alignment
Shaoxiang Chen 0001, Sihao Lin, Zequn Jie, Lin Ma 0002, Guangrun Wang, Xiaodan Liang |
ECCV (36) | 6 |
| 2024 | 3D Weakly Supervised Semantic Segmentation with 2D Vision-Language Guidance
Xiaoxu Xu, Yitian Yuan, Jinlong Li 0003, Qiudan Zhang, Zequn Jie, Lin Ma 0002, Hao Tang 0005, Nicu Sebe, Xu Wang 0006 |
ECCV (73) | 6 |
| 2024 | UniMD: Towards Unifying Moment Retrieval and Temporal Action Detection
Yingsen Zeng, Chengjian Feng, Lin Ma 0002 |
ECCV (46) | 4 |
| 2024 | Aux-NAS: Exploiting Auxiliary Labels with Negligibly Extra Inference CostabstractWe aim at exploiting additional auxiliary labels from an independent (auxiliary) task to boost the primary task performance which we focus on, while preserving a single task inference cost of the primary task. While most existing auxiliary learning methods are optimization-based relying on loss weights/gradients manipulation, our method is architecture-based with a flexible asymmetric structure for the primary and auxiliary tasks, which produces different networks for training and inference. Specifically, starting from two single task networks/branches (each representing a task), we propose a novel method with evolving networks where only primary-to-auxiliary links exist as the cross-task connections after convergence. These connections can be removed during the primary task inference, resulting in a single-task inference cost. We achieve this by formulating a Neural Architecture Search (NAS) problem, where we initialize bi-directional connections in the search space and guide the NAS optimization converging to an architecture with only the single-side primary-to-auxiliary connections. Moreover, our method can be incorporated with optimization-based auxiliary learning approaches. Extensive experiments with six tasks on NYU v2, CityScapes, and Taskonomy datasets using VGG, ResNet, and ViT backbones validate the promising performance. The codes are available at https://github.com/ethanygao/Aux-NAS. Yuan Gao 0015, Wenhan Luo, Lin Ma 0002, Jin-Gang Yu, Gui-Song Xia, Jiayi Ma 0001 |
ICLR | 4 |
| 2024 | Lumen: Unleashing Versatile Vision-Centric Capabilities of Large Multimodal ModelsabstractLarge Multimodal Model (LMM) is a hot research topic in the computer vision area and has also demonstrated remarkable potential across multiple disciplinary fields. A recent trend is to further extend and enhance the perception capabilities of LMMs. The current methods follow the paradigm of adapting the visual task outputs to the format of the language model, which is the main component of a LMM. This adaptation leads to convenient development of such LMMs with minimal modifications, however, it overlooks the intrinsic characteristics of diverse visual tasks and hinders the learning of perception capabilities. To address this issue, we propose a novel LMM architecture named Lumen, a Large multimodal model with versatile vision-centric capability enhancement. We decouple the LMM's learning of perception capabilities into task-agnostic and task-specific stages. Lumen first promotes fine-grained vision-language concept alignment, which is the fundamental capability for various visual tasks. Thus the output of the task-agnostic stage is a shared representation for all the tasks we address in this paper. Then the task-specific decoding is carried out by flexibly routing the shared representation to lightweight task decoders with negligible training efforts. Comprehensive experimental results on a series of vision-centric and VQA benchmarks indicate that our Lumen model not only achieves or surpasses the performance of existing LMM-based approaches in a range of vision-centric tasks while maintaining general visual understanding and instruction following capabilities. Shaoxiang Chen 0001, Zequn Jie, Jingjing Chen 0001, Lin Ma 0002, Yu-Gang Jiang 0001 |
NeurIPS | 5 |
| 2024 | LESS: Label-Efficient and Single-Stage Referring 3D SegmentationabstractReferring 3D Segmentation is a visual-language task that segments all points of the specified object from a 3D point cloud described by a sentence of query. Previous works perform a two-stage paradigm, first conducting language-agnostic instance segmentation then matching with given text query. However, the semantic concepts from text query and visual cues are separately interacted during the training, and both instance and semantic labels for each object are required, which is time consuming and human-labor intensive. To mitigate these issues, we propose a novel Referring 3D Segmentation pipeline, Label-Efficient and Single-Stage, dubbed LESS, which is only under the supervision of efficient binary mask. Specifically, we design a Point-Word Cross-Modal Alignment module for aligning the fine-grained features of points and textual embedding. Query Mask Predictor module and Query-Sentence Alignment module are introduced for coarse-grained alignment between masks and query. Furthermore, we propose an area regularization loss, which coarsely reduces irrelevant background predictions on a large scale. Besides, a point-to-point contrastive loss is proposed concentrating on distinguishing points with subtly similar features. Through extensive experiments, we achieve state-of-the-art performance on ScanRefer dataset by surpassing the previous methods about 3.7% mIoU using only binary labels. Code is available at https://github.com/mellody11/LESS. Xuexun Liu, Xiaoxu Xu, Jinlong Li 0003, Qiudan Zhang, Xu Wang 0006, Nicu Sebe, Lin Ma 0002 |
NeurIPS | 7 |
| 2024 | DeTAL: Open-Vocabulary Temporal Action Localization With Decoupled NetworksabstractPre-trained visual-language (ViL) models have demonstrated good zero-shot capability in video understanding tasks, where they were usually adapted through fine-tuning or temporal modeling. However, in the task of open-vocabulary temporal action localization (OV-TAL), such adaption reduces the robustness of ViL models against different data distributions, leading to a misalignment between visual representations and text descriptions of unseen action categories. As a result, existing methods often strike a trade-off between action detection and classification. Aiming at this issue, this paper proposes DeTAL, a simple but effective two-stage approach for OV-TAL. DeTAL decouples action detection from action classification to avoid the compromise between them, and the state-of-the-art methods for close-set action localization can be handily adapted to OV-TAL, which significantly improves the performance. Meanwhile, DeTAL can easily tackle the scenario where action category annotations are unavailable in the training dataset. In the experiments, we propose a new cross-dataset setting to evaluate the zero-shot capability of different methods. And the results demonstrate that DeTAL outperforms the state-of-the-art methods for OV-TAL on both THUMOS14 and ActivityNet1.3. Zhiheng Li 0005, Ran Song 0001, Lin Ma 0002, Wei Zhang 0021 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2024 | IGCN: A Provably Informative GCN Embedding for Semi-Supervised Learning With Extremely Limited LabelsabstractGraph Neural Networks (GNNs) have gained much more attention in the representation learning for the graph-structured data. However, the labels are always limited in the graph, which easily leads to the overfitting problem and causes the poor performance. To solve this problem, we propose a new framework called IGCN, short for Informative Graph Convolutional Network, where the objective of IGCN is designed to obtain the informative embeddings via discarding the task-irrelevant information of the graph data based on the mutual information. As the mutual information for irregular data is intractable to compute, our framework is optimized via a surrogate objective, where two terms are derived to approximate the original objective. For the former term, it demonstrates that the mutual information between the learned embeddings and the ground truth should be high, where we utilize the semi-supervised classification loss and the prototype based supervised contrastive learning loss for optimizing it. For the latter term, it requires that the mutual information between the learned node embeddings and the initial embeddings should be high and we propose to minimize the reconstruction loss between them to achieve the goal of maximizing the latter term from the feature level and the layer level, which contains the graph encoder-decoder module and a novel architecture GCN$_{Info}$. Moreover, we provably show that the designed GCN$_{Info}$can better alleviate the information loss and preserve as much useful information of the initial embeddings as possible. Experimental results show that the IGCN outperforms the state-of-the-art methods on 7 popular datasets. Lin Zhang 0041, Ran Song 0001, Wenhao Tan, Lin Ma 0002, Wei Zhang 0021 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | LMEye: An Interactive Perception Network for Large Language ModelsabstractCurrent efficient approaches to building Multimodal Large Language Models (MLLMs) mainly incorporate visual information into LLMs with a simple visual mapping network such as a linear projection layer, a multilayer perceptron (MLP), or Q-former from BLIP-2. Such networks project the image feature once and do not consider the interaction between the image and the human inputs. Hence, the obtained visual information without being connected to human intention may be inadequate for LLMs to generate intention-following responses, which we refer to as static visual information. To alleviate this issue, our paper introduces LMEye, a human-like eye with a play-and-plug interactive perception network, designed to enable dynamic interaction between LLMs and external visual information. It can allow the LLM to request the desired visual information aligned with various human instructions, which we term dynamic visual information acquisition. Specifically, LMEye consists of a simple visual mapping network to provide the basic perception of an image for LLMs. It also contains additional modules responsible for acquiring requests from LLMs, performing request-based visual information seeking, and transmitting the resulting interacted visual information to LLMs, respectively. In this way, LLMs act to understand the human query, deliver the corresponding request to the request-based visual information interaction module, and generate the response based on the interleaved multimodal information. We evaluate LMEye through extensive experiments on multimodal benchmarks, demonstrating that it significantly improves zero-shot performances on various multimodal tasks compared to previous methods, with fewer parameters. Moreover, we also verify its effectiveness and scalability on various language models and video understanding, respectively. Yunxin Li, Baotian Hu, Xinyu Chen 0003, Lin Ma 0002, Yong Xu 0001, Min Zhang 0005 |
IEEE Trans. Multim. | 4 |
| 2024 | CroMIC-QA: The Cross-Modal Information Complementation Based Question AnsweringabstractThis paper proposes a new multi-modal question-answering task, named as Cross-Modal Information Complementation based Question Answering (CroMIC-QA), to promote the exploration on bridging the semantic gap between visual and linguistic signals. The proposed task is inspired by the common phenomenon that, in most user-generated QA scenarios, the information of the given textual question is incomplete, and thus it is required to merge the semantics of both the text and the accompanying image to infer the complete real question. In this work, the CroMIC-QA task is first formally defined and compared with the classic Visual Question Answering (VQA) task. On this basis, a specified dataset, CroMIC-QA-Agri, is collected from an online QA community in the agriculture domain for the proposed task. A group of experiments is conducted on this dataset, with the typical multi-modal deep architectures implemented and compared. The experimental results show that the appropriate text/image presentations and text-image semantic interaction methods are effective to improve the performance of the framework. Shun Qian, Bingquan Liu, Chengjie Sun, Zhen Xu 0003, Lin Ma 0002, Baoxun Wang |
IEEE Trans. Multim. | 5 |
| 2024 | Weakly-Supervised 3D Scene Graph Generation via Visual-Linguistic Assisted Pseudo-LabelingabstractLearning to build 3D scene graphs is essential for real-world perception in a structured and rich fashion. However, previous 3D scene graph generation methods utilize a fully supervised learning manner and require a large amount of entity-level annotation data of objects and relations, which is extremely resource-consuming and tedious to obtain. To tackle this problem, we propose 3D-VLAP, a weakly-supervised 3D scene graph generation method via Visual-Linguistic Assisted Pseudo-labeling. Specifically, our 3D-VLAP exploits the superior ability of current large-scale visual-linguistic models to align the semantics between texts and 2D images, as well as the naturally existing correspondences between 2D images and 3D point clouds, and thus implicitly constructs correspondences between texts and 3D point clouds. First, we establish the positional correspondence from 3D point clouds to 2D images via camera intrinsic and extrinsic parameters, thereby achieving alignment of 3D point clouds and 2D images. Subsequently, a large-scale cross-modal visual-linguistic model is employed to indirectly align 3D instances with the textual category labels of objects by matching 2D images with object category labels. The pseudo labels for objects and relations are then produced for 3D-VLAP model training by calculating the similarity between visual embeddings and textual category embeddings of objects and relations encoded by the visual-linguistic model, respectively. Ultimately, we design an edge self-attention based graph neural network to generate scene graphs of 3D point clouds. Experiments demonstrate that our 3D-VLAP achieves comparable results with current fully supervised methods, meanwhile alleviating the data annotation pressure. Xu Wang 0006, Qiudan Zhang, Wenhui Wu 0001, Mark Junjie Li, Lin Ma 0002, Jianmin Jiang |
IEEE Trans. Multim. | 6 |
| 2023 | Curriculum Multi-Negative Augmentation for Debiased Video GroundingabstractVideo Grounding (VG) aims to locate the desired segment from a video given a sentence query. Recent studies have found that current VG models are prone to over-rely the groundtruth moment annotation distribution biases in the training set. To discourage the standard VG model's behavior of exploiting such temporal annotation biases and improve the model generalization ability, we propose multiple negative augmentations in a hierarchical way, including cross-video augmentations from clip-/video-level, and self-shuffled augmentations with masks. These augmentations can effectively diversify the data distribution so that the model can make more reasonable predictions instead of merely fitting the temporal biases. However, directly adopting such data augmentation strategy may inevitably carry some noise shown in our cases, since not all of the handcrafted augmentations are semantically irrelevant to the groundtruth video. To further denoise and improve the grounding accuracy, we design a multi-stage curriculum strategy to adaptively train the standard VG model from easy to hard negative augmentations. Experiments on newly collected Charades-CD and ActivityNet-CD datasets demonstrate our proposed strategy can improve the performance of the base model on both i.i.d and o.o.d scenarios. Xiaohan Lan, Yitian Yuan, Hong Chen 0011, Xin Wang 0019, Zequn Jie, Lin Ma 0002, Zhi Wang 0001, Wenwu Zhu 0001 |
AAAI | 6 |
| 2023 | A Multi-Modal Context Reasoning Approach for Conditional Inference on Joint Textual and Visual CluesabstractConditional inference on joint textual and visual clues is a multi-modal reasoning task that textual clues provide prior permutation or external knowledge, which are complementary with visual content and pivotal to deducing the correct option.Previous methods utilizing pretrained vision-language models (VLMs) have achieved impressive performances, yet they show a lack of multimodal context reasoning capability, especially for text-modal information.To address this issue, we propose a Multi-modal Context Reasoning approach, named ModCR.Compared to VLMs performing reasoning via cross modal semantic alignment, it regards the given textual abstract semantic and objective image information as the pre-context information and embeds them into the language model to perform context reasoning.Different from recent vision-aided language models used in natural language processing, ModCR incorporates the multi-view semantic alignment information between language and vision by introducing the learnable alignment prefix between image and text in the pretrained language model.This makes the language model well-suitable for such multi-modal reasoning scenario on joint textual and visual clues.We conduct extensive experiments on two corresponding data sets and experimental results show significantly improved performance (exact gain by 4.8% on PMR test set) compared to previous strong baselines.Code Yunxin Li, Baotian Hu, Xinyu Chen 0003, Lin Ma 0002, Min Zhang 0005 |
ACL (1) | 5 |
| 2023 | A Neural Divide-and-Conquer Reasoning Framework for Image Retrieval from Linguistically Complex TextabstractPretrained Vision-Language Models (VLMs) have achieved remarkable performance in image retrieval from text.However, their performance drops drastically when confronted with linguistically complex texts that they struggle to comprehend.Inspired by the Divide-and-Conquer (Smith, 1985) algorithm and dualprocess theory (Groves and Thompson, 1970), in this paper, we regard linguistically complex texts as compound proposition texts composed of multiple simple proposition sentences and propose an end-to-end Neural Divide-and-Conquer Reasoning framework, dubbed NDCR.It contains three main components: 1) Divide: a proposition generator divides the compound proposition text into simple proposition sentences and produces their corresponding representations, 2) Conquer: a pretrained VLMsbased visual-linguistic interactor achieves the interaction between decomposed proposition sentences and images, 3) Combine: a neuralsymbolic reasoner combines the above reasoning states to obtain the final solution via a neural logic reasoning approach.According to the dual-process theory, the visual-linguistic interactor and neural-symbolic reasoner could be regarded as analogical reasoning System 1 and logical reasoning System 2. We conduct extensive experiments on a challenging image retrieval from contextual descriptions data set.Experimental results and analyses indicate NDCR significantly improves performance in the complex image-text reasoning problem. Yunxin Li, Baotian Hu, Lin Ma 0002, Min Zhang 0005 |
ACL (1) | 4 |
| 2023 | AeDet: Azimuth-Invariant Multi-View 3D Object DetectionabstractRecent LSS-based multi-view 3D object detection has made tremendous progress, by processing the features in Brid-Eye-View (BEV) via the convolutional detector. However, the typical convolution ignores the radial symmetry of the BEV features and increases the difficulty of the detector optimization. To preserve the inherent property of the BEV features and ease the optimization, we propose an azimuth-equivariant convolution (AeConv) and an azimuth-equivariant anchor. The sampling grid of AeConv is always in the radial direction, thus it can learn azimuth-invariant BEV features. The proposed anchor enables the detection head to learn predicting azimuth-irrelevant targets. In addition, we introduce a camera-decoupled virtual depth to unify the depth prediction for the images with different camera intrinsic parameters. The resultant detector is dubbed Azimith-equivariant Detector (AeDet). Extensive experiments are conducted on nuScenes, and AeDet achieves a 62.0% NDS, surpassing the recent multi-view 3D object detectors such as PETRv2 and BEVDepth by a large margin. Project page: https://fcjian.github.io/aedet. Chengjian Feng, Zequn Jie, Xiangxiang Chu, Lin Ma 0002 |
CVPR | 5 |
| 2023 | MSMDFusion: Fusing LiDAR and Camera at Multiple Scales with Multi-Depth Seeds for 3D Object DetectionabstractFusing LiDAR and camera information is essential for accurate and reliable 3D object detection in autonomous driving systems. This is challenging due to the difficulty of combining multi-granularity geometric and semantic features from two drastically different modalities. Recent approaches aim at exploring the semantic densities of camera features through lifting points in 2D camera images (referred to as “seeds”) into 3D space, and then incorporate 2D semantics via cross-modal interaction or fusion techniques. However, depth information is under-investigated in these approaches when lifting points into 3D space, thus 2D semantics can not be reliably fused with 3D points. Moreover, their multi-modal fusion strategy, which is implemented as concatenation or attention, either can not effectively fuse 2D and 3D information or is unable to perform fine-grained interactions in the voxel space. To this end, we propose a novel framework with better utilization of the depth information and fine-grained cross-modal interaction between LiDAR and camera, which consists of two important components. First, a Multi-Depth Unprojection (MDU) method is used to enhance the depth quality of the lifted points at each interaction level. Second, a Gated Modality-Aware Convolution (GMA-Conv) block is applied to modulate voxels involved with the camera modality in a fine-grained manner and then aggregate multi-modal features into a unified space. Together they provide the detection head with more comprehensive features from LiDAR and camera. On the nuScenes test benchmark, our proposed method, abbreviated as MSMD-Fusion, achieves state-of-the-art results on both 3D object detection and tracking tasks without using test-time-augmentation and ensemble techniques. The code is available at https://github.com/SxJyJay/MSMDFusion. Zequn Jie, Shaoxiang Chen 0001, Jingjing Chen 0001, Lin Ma 0002, Yu-Gang Jiang 0001 |
CVPR | 5 |
| 2023 | TriDet: Temporal Action Detection with Relative Boundary ModelingabstractIn this paper, we present a one-stage framework TriDet for temporal action detection. Existing methods often suffer from imprecise boundary predictions due to the ambiguous action boundaries in videos. To alleviate this problem, we propose a novel Trident-head to model the action boundary via an estimated relative probability distribution around the boundary. In the feature pyramid of TriDet, we propose an efficient Scalable-Granularity Perception (SGP) layer to mitigate the rank loss problem of self-attention that takes place in the video features and aggregate information across different temporal granularities. Benefiting from the Trident-head and the SGP-based feature pyramid, TriDet achieves state-of-the-art performance on three challenging benchmarks: THUMOS14, HACS and EPIC-KITCHEN 100, with lower computational costs, compared to previous methods. For example, TriDet hits an average mAP of 69.3% on THUMOS14, outperforming the previous best by 2.5%, but with only 74.6% of its latency. The code is released to https://github.com/dingfengshi/TriDet. Dingfeng Shi, Qiong Cao, Lin Ma 0002, Jia Li 0003, Dacheng Tao |
CVPR | 4 |
| 2023 | Adaptive Sparse Pairwise Loss for Object Re-IdentificationabstractObject re-identification (ReID) aims to find instances with the same identity as the given probe from a large gallery. Pairwise losses play an important role in training a strong ReID network. Existing pairwise losses densely exploit each instance as an anchor and sample its triplets in a mini-batch. This dense sampling mechanism inevitably introduces positive pairs that share few visual similarities, which can be harmful to the training. To address this problem, we propose a novel loss paradigm termed Sparse Pair-wise (SP) loss that only leverages few appropriate pairs for each class in a mini-batch, and empirically demonstrate that it is sufficient for the ReID tasks. Based on the proposed loss framework, we propose an adaptive positive mining strategy that can dynamically adapt to diverse intra-class variations. Extensive experiments show that SP loss and its adaptive variant AdaSP loss outperform other pairwise losses, and achieve state-of-the-art performance across several ReID benchmarks. Code is available at https://github.com/Astaxanthin/AdaSP. Zhen Cheng 0004, Lin Ma 0002 |
CVPR | 5 |
| 2023 | E2E-LOAD: End-to-End Long-form Online Action DetectionabstractRecently, feature-based methods for Online Action Detection (OAD) have been gaining traction. However, these methods are constrained by their fixed backbone design, which fails to leverage the potential benefits of a trainable backbone. This paper introduces an end-to-end learning network that revises these approaches, incorporating a backbone network design that improves effectiveness and efficiency. Our proposed model utilizes a shared initial spatial model for all frames and maintains an extended sequence cache, which enables low-cost inference. We promote an asymmetric spatiotemporal model that caters to long-form and short-form modeling. Additionally, we propose an innovative and efficient inference mechanism that accelerates extensive spatiotemporal exploration. Through comprehensive ablation studies and experiments, we validate the performance and efficiency of our proposed method. Remarkably, we achieve an end-to-end learning OAD of 17.3 (+12.6) FPS with 72.4% (+1.2%), 90.3% (+0.7%), and 48.1% (+26.0%) mAP on THMOUS’14, TVSeries, and HDD, respectively. The source code is available at https://github.com/sqiangcao99/E2E-LOAD. Shuqiang Cao, Weixin Luo, Bairui Wang, Wei Zhang 0021, Lin Ma 0002 |
ICCV | 5 |
| 2023 | Open-Vocabulary Semantic Segmentation with Decoupled One-Pass NetworkabstractRecently, the open-vocabulary semantic segmentation problem has attracted increasing attention and the best performing methods are based on two-stream networks: one stream for proposal mask generation and the other for segment classification using a pre-trained visual-language model. However, existing two-stream methods require passing a great number of (up to a hundred) image crops into the visual-language model, which is highly inefficient. To address the problem, we propose a network that only needs a single pass through the visual-language model for each input image. Specifically, we first propose a novel network adaptation approach, termed patch severance, to restrict the harmful interference between the patch embeddings in the pre-trained visual encoder. We then propose classification anchor learning to encourage the network to spatially focus on more discriminative features for classification. Extensive experiments demonstrate that the proposed method achieves outstanding performance, surpassing state-of-the-art methods while being 4 to 7 times faster at inference. Code: https://github.com/CongHan0808/DeOP.git Dengjie Li, Lin Ma 0002 |
ICCV | 5 |
| 2023 | Planning Assembly Sequence with Graph TransformerabstractAssembly Sequence Planning (ASP) is the essential process for modern manufacturing, proven to be NP-complete thus its effective and efficient solution has been a challenge for researchers in the field. In this paper, we present a graph-transformer based framework for the ASP problem which is trained and demonstrated on a self-collected ASP database. The ASP database contains a self-collected set of LEGO models. The LEGO model is abstracted to a heterogeneous graph structure after a thorough analysis of the original structure and feature extraction. The ground truth assembly sequence is first generated by brute-force search and then adjusted manually to be in line with human rational habits. Based on this self-collected ASP dataset, we propose a heterogeneous graph-transformer framework to learn the latent rules for assembly planning. We evaluated the proposed framework in a series of experiments. The results show that the similarity of the predicted and ground truth sequences can reach 0.44, a medium correlation measured by Kendall's τ. Meanwhile, we compared the different effects of node features and edge features and generated a feasible and reasonable assembly sequence as a benchmark for further research. Our dataset and code are available on: htps://github.com/AIR-DISCOVER/ICRA_ASP. Lin Ma 0002, Jiangtao Gong, Hao Chen 0062, Hao Zhao 0002, Wenbing Huang 0001, Guyue Zhou |
ICRA | 1 |
| 2023 | Suspected Objects Matter: Rethinking Model's Prediction for One-stage Visual GroundingabstractRecently, one-stage visual grounders attract high attention due to their comparable accuracy but significantly higher efficiency than two-stage grounders. However, inter-object relation modeling has not been well studied for one-stage grounders. Inter-object relationship modeling, though important, is not necessarily performed among all objects, as only part of them are related to the text query and may confuse the model. We call these objects "suspected objects". However, exploring their relationships in the one-stage paradigm is non-trivial because: (1) no object proposals are available as the basis on which to select suspected objects and perform relationship modeling; (2) suspected objects are more confusing than others, as they may share similar semantics, be entangled with certain relationships, etc, and thereby more easily mislead the model's prediction. Toward this end, we propose a Suspected Object Transformation mechanism (SOT), which can be seamlessly integrated into existing CNN and Transformer-based one-stage visual grounders to encourage the target object selection among the suspected ones. Suspected objects are dynamically discovered from a learned activation map adapted to the model's current discrimination ability during training. Afterward, on top of suspected objects, a Keyword-Aware Discrimination module (KAD) and an Exploration by Random Connection strategy (ERC) are concurrently proposed to help the model rethink its initial prediction. On the one hand, KAD leverages keywords contributing high to suspected object discrimination. On the other hand, ERC allows the model to seek the correct object instead of being trapped in a situation that always exploits the current false prediction. Extensive experiments demonstrate the effectiveness of our proposed method. Zequn Jie, Jingjing Chen 0001, Lin Ma 0002, Yu-Gang Jiang 0001 |
ACM Multimedia | 4 |
| 2023 | Punctuation-level Attack: Single-shot and Single Punctuation Can Fool Text ModelsabstractThe adversarial attacks have attracted increasing attention in various fields including natural language processing. The current textual attacking models primarily focus on fooling models by adding character-/word-/sentence-level perturbations, ignoring their influence on human perception. In this paper, for the first time in the community, we propose a novel mode of textual attack, punctuation-level attack. With various types of perturbations, including insertion, displacement, deletion, and replacement, the punctuation-level attack achieves promising fooling rates against SOTA models on typical textual tasks and maintains minimal influence on human perception and understanding of the text by mere perturbation of single-shot single punctuation. Furthermore, we propose a search method named Text Position Punctuation Embedding and Paraphrase (TPPEP) to accelerate the pursuit of optimal position to deploy the attack, without exhaustive search, and we present a mathematical interpretation of TPPEP. Thanks to the integrated Text Position Punctuation Embedding (TPPE), the punctuation attack can be applied at a constant cost of time. Experimental results on public datasets and SOTA models demonstrate the effectiveness of the punctuation attack and the proposed TPPE. We additionally apply the single punctuation attack to summarization, semantic-similarity-scoring, and text-to-image tasks, and achieve encouraging results. Chongyang Du, Tao Wang 0052, Kaihao Zhang, Wenhan Luo, Lin Ma 0002, Wei Liu 0005, Xiaochun Cao |
NeurIPS | 6 |
| 2023 | RecFormer: Recurrent Multi-modal Transformer with History-Aware Contrastive Learning for Visual Dialog
Liucun Lu, Jinghui Qin, Zequn Jie, Lin Ma 0002, Liang Lin 0004, Xiaodan Liang |
PRCV (1) | 4 |
| 2023 | Weakly supervised semantic segmentation via self-supervised destruction learning
Jinlong Li 0003, Zequn Jie, Xu Wang 0006, Yu Zhou 0027, Lin Ma 0002, Jianmin Jiang |
Neurocomputing | 5 |
| 2023 | MARN: Multi-level Attentional Reconstruction Networks for Weakly Supervised Video Temporal Grounding
Yijun Song, Jingwen Wang 0003, Lin Ma 0002, Jun Yu 0002, Jinxiu Liang, Zhou Yu 0001 |
Neurocomputing | 3 |
| 2023 | Weakly Supervised Semantic Segmentation Via Progressive Patch LearningabstractMost of the existing semantic segmentation approaches with image-level class labels as supervision, highly rely on the initial class activation map (CAM) generated from the standard classification network. In this paper, a novel “Progressive Patch Learning” approach is proposed to improve the local details extraction of the classification, producing the CAM better covering the whole object rather than only the most discriminative regions as in CAMs obtained in conventional classification models. “Patch Learning” destructs the feature maps into patches and independently processes each local patch in parallel before the final aggregation. Such a mechanism enforces the network to find weak information from the scattered discriminative local parts, achieving enhanced local details sensitivity. “Progressive Patch Learning” further extends the feature destruction and patch learning to multi-level granularities in a progressive manner. Cooperating with a multi-stage optimization strategy, such a “Progressive Patch Learning” mechanism implicitly provides the model with the feature extraction ability across different locality-granularities. As an alternative to the implicit multi-granularity progressive fusion approach, we additionally propose an explicit method to simultaneously fuse features from different granularities in a single model, further enhancing the CAM quality on the full object coverage. Our proposed method achieves outstanding performance on the PASCAL VOC 2012 dataset (e.g., with 69.6$\%$mIoU on thetestset), which surpasses most existing weakly supervised semantic segmentation methods. Jinlong Li 0003, Zequn Jie, Xu Wang 0006, Yu Zhou 0027, Xiaolin Wei, Lin Ma 0002 |
IEEE Trans. Multim. | 6 |
| 2023 | Fast and Robust Online Handwritten Chinese Character Recognition With Deep Spatial and Contextual Information Fusion NetworkabstractDeep convolutional neuralnetworks have achieved fairly high accuracy for single online handwritten Chinese character recognition (SOLHCCR). However, in real application scenarios, users always write multiple characters to form a complete sentence, and previous contextual information holds significant potential for improving the accuracy, robustness and efficiency of recognition. In this work, we first propose a simple and straightforward model named the vanilla compositional network (VCN) by coupling convolutional neural network with a sequence modeling architecture (i.e., a recurrent neural network or Transformer), which exploits the handwritten character’s previous contextual information. Although VCN performs much better than the previous state-of-the-art SOLHCCR models, it is a two-stage architecture in nature. It suffers from high fragility when confronting with poorly written characters such as sloppy writing, and missing or broken strokes, due to relying heavily on contextual information. To improve the robustness of the OLHCCR model, we further propose a novel deep spatial & contextual information fusion network (DSCIFN). It utilizes an autoregresssive framework pre-trained on a large-scale sentence corpora as the backbone component, and highly integrates the spatial features of handwritten characters and their previous contextual information in a multi-layer fusion module. To verify the effectiveness of models, we reorganize a new form of online Chinese handwritten character with its previous context dataset, named OHCCC. Extensive experimental results demonstrate that DSCIFN achieves state-of-the-art performance and has increased strong robustness compared to VCN and previous SOLHCCR models. The in-depth empirical analysis and case study indicate that DSCIFN can significantly improve the efficiency of handwriting input because it does not need complete strokes to recognize a handwritten Chinese character precisely. Yunxin Li, Qian Yang 0007, Qingcai Chen, Baotian Hu, Xiaolong Wang 0001, Lin Ma 0002 |
IEEE Trans. Multim. | 7 |
| 2023 | A Closer Look at Debiased Temporal Sentence Grounding in Videos: Dataset, Metric, and ApproachabstractTemporal Sentence Grounding in Videos (TSGV) , which aims to ground a natural language sentence that indicates complex human activities in an untrimmed video, has drawn widespread attention over the past few years. However, recent studies have found that current benchmark datasets may have obvious moment annotation biases, enabling several simple baselines even without training to achieve state-of-the-art (SOTA) performance. In this paper, we take a closer look at existing evaluation protocols for TSGV, and find that both the prevailing dataset splits and evaluation metrics are the devils that lead to untrustworthy benchmarking. Therefore, we propose to re-organize the two widely-used datasets, making the ground-truth moment distributions different in the training and test splits, i.e., out-of-distribution (OOD) test. Meanwhile, we introduce a new evaluation metric “dR@ n ,IoU= m ” that discounts the basic recall scores especially with small IoU thresholds, so as to alleviate the inflating evaluation caused by biased datasets with a large proportion of long ground-truth moments. New benchmarking results indicate that our proposed evaluation protocols can better monitor the research progress in TSGV. Furthermore, we propose a novel causality-based Multi-branch Deconfounding Debiasing (MDD) framework for unbiased moment prediction. Specifically, we design a multi-branch deconfounder to eliminate the effects caused by multiple confounders with causal intervention. In order to help the model better align the semantics between sentence queries and video moments, we enhance the representations during feature encoding. Specifically, for textual information, the query is parsed into several verb-centered phrases to obtain a more fine-grained textual feature. For visual information, the positional information has been decomposed from the moment features to enhance the representations of moments with diverse locations. Extensive experiments demonstrate that our proposed approach can achieve competitive results among existing SOTA approaches and outperform the base model with great gains. Xiaohan Lan, Yitian Yuan, Xin Wang 0019, Long Chen 0016, Zhi Wang 0001, Lin Ma 0002, Wenwu Zhu 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2022 | Visual Consensus Modeling for Video-Text RetrievalabstractIn this paper, we propose a novel method to mine the commonsense knowledge shared between the video and text modalities for video-text retrieval, namely visual consensus modeling. Different from the existing works, which learn the video and text representations and their complicated relationships solely based on the pairwise video-text data, we make the first attempt to model the visual consensus by mining the visual concepts from videos and exploiting their co-occurrence patterns within the video and text modalities with no reliance on any additional concept annotations. Specifically, we build a shareable and learnable graph as the visual consensus, where the nodes denoting the mined visual concepts and the edges connecting the nodes representing the co-occurrence relationships between the visual concepts. Extensive experimental results on the public benchmark datasets demonstrate that our proposed method, with the ability to effectively model the visual consensus, achieves state-of-the-art performances on the bidirectional video-text retrieval task. Our code is available at https://github.com/sqiangcao99/VCM. Shuqiang Cao, Bairui Wang, Wei Zhang 0021, Lin Ma 0002 |
AAAI | 4 |
| 2022 | Explore Inter-contrast between Videos via Composition for Weakly Supervised Temporal Sentence GroundingabstractWeakly supervised temporal sentence grounding aims to temporally localize the target segment corresponding to a given natural language query, where it provides video-query pairs without temporal annotations during training. Most existing methods use the fused visual-linguistic feature to reconstruct the query, where the least reconstruction error determines the target segment. This work introduces a novel approach that explores the inter-contrast between videos in a composed video by selecting components from two different videos and fusing them into a single video. Such a straightforward yet effective composition strategy provides the temporal annotations at multiple composed positions, resulting in numerous videos with temporal ground-truths for training the temporal sentence grounding task. A transformer framework is introduced with multi-tasks training to learn a compact but efficient visual-linguistic space. The experimental results on the public Charades-STA and ActivityNet-Caption dataset demonstrate the effectiveness of the proposed method, where our approach achieves comparable performance over the state-of-the-art weakly-supervised baselines. The code is available at https://github.com/PPjmchen/Composition_WSTG. Jiaming Chen 0001, Weixin Luo, Wei Zhang 0021, Lin Ma 0002 |
AAAI | 4 |
| 2022 | Task Generalizable Spatial and Texture Aware Image Downsizing Network
Lin Ma 0002, Hongsheng Li 0001, Qiang Wang 0023 |
BMVC | 1 |
| 2022 | PromptDet: Towards Open-Vocabulary Detection Using Uncurated Images
Chengjian Feng, Zequn Jie, Xiangxiang Chu, Haibing Ren, Xiaolin Wei, Weidi Xie, Lin Ma 0002 |
ECCV (9) | 8 |
| 2022 | MORE: Multi-Order RElation Mining for Dense Captioning in 3D Scenes
Shaoxiang Chen 0001, Zequn Jie, Jingjing Chen 0001, Lin Ma 0002, Yu-Gang Jiang 0001 |
ECCV (35) | 5 |
| 2022 | ReAct: Temporal Action Detection with Relational Queries
Dingfeng Shi, Qiong Cao, Jing Zhang 0037, Lin Ma 0002, Jia Li 0003, Dacheng Tao |
ECCV (10) | 5 |
| 2022 | Cycle-Interactive Generative Adversarial Network for Robust Unsupervised Low-Light EnhancementabstractGetting rid of the fundamental limitations in fitting to the paired training data, recent unsupervised low-light enhancement methods excel in adjusting illumination and contrast of images. However, for unsupervised low light enhancement, the remaining noise suppression issue due to the lacking of supervision of detailed signal largely impedes the wide deployment of these methods in real-world applications. Herein, we propose a novel Cycle-Interactive Generative Adversarial Network (CIGAN) for unsupervised low-light image enhancement, which is capable of not only better transferring illumination distributions between low/normal-light images but also manipulating detailed signals between two domains, e.g., suppressing/synthesizing realistic noise in the cyclic enhancement/degradation process. In particular, the proposed low-light guided transformation feed-forwards the features of low-light images from the generator of enhancement GAN (eGAN) into the generator of degradation GAN (dGAN). With the learned information of real low-light images, dGAN can synthesize more realistic diverse illumination and contrast in low-light images. Moreover, the feature randomized perturbation module in dGAN learns to increase the feature randomness to produce diverse feature distributions, persuading the synthesized low-light images to contain realistic noise. Extensive experiments demonstrate both the superiority of the proposed method and the effectiveness of each module in CIGAN. Zhangkai Ni, Wenhan Yang, Hanli Wang, Shiqi Wang 0001, Lin Ma 0002, Sam Kwong |
ACM Multimedia | 5 |
| 2022 | Chunk-aware Alignment and Lexical Constraint for Visual Entailment with Natural Language ExplanationsabstractVisual Entailment with natural language explanations aims to infer the relationship between a text-image pair and generate a sentence to explain the decision-making process. Previous methods rely mainly on a pre-trained vision-language model to perform the relation inference and a language model to generate the corresponding explanation. However, the pre-trained vision-language models mainly build token-level alignment between text and image yet ignore the high-level semantic alignment between the phrases (chunks) and visual contents, which is critical for vision-language reasoning. Moreover, the explanation generator based only on the encoded joint representation does not explicitly consider the critical decision-making points of relation inference. Thus the generated explanations are less faithful to visual-language reasoning. To mitigate these problems, we propose a unified Chunk-aware Alignment and Lexical Constraint based method, dubbed as CALeC. It contains a Chunk-aware Semantic Interactor (arr. CSI), a relation inferrer, and a Lexical Constraint-aware Generator (arr. LeCG). Specifically, CSI exploits the sentence structure inherent in language and various image regions to build chunk-aware semantic alignment. Relation inferrer uses an attention-based reasoning network to incorporate the token-level and chunk-level vision-language representations. LeCG utilizes lexical constraints to expressly incorporate the words or chunks focused by the relation inferrer into explanation generation, improving the faithfulness and informativeness of the explanations. We conduct extensive experiments on three datasets, and experimental results indicate that CALeC significantly outperforms other competitor models on inference accuracy and quality of generated explanations. Qian Yang 0007, Yunxin Li, Baotian Hu, Lin Ma 0002, Min Zhang 0005 |
ACM Multimedia | 4 |
| 2022 | Expansion and Shrinkage of Localization for Weakly-Supervised Semantic SegmentationabstractGenerating precise class-aware pseudo ground-truths, a.k.a, class activation maps (CAMs), is essential for Weakly-Supervised Semantic Segmentation. The original CAM method usually produces incomplete and inaccurate localization maps. To tackle with this issue, this paper proposes an Expansion and Shrinkage scheme based on the offset learning in the deformable convolution, to sequentially improve the recall and precision of the located object in the two respective stages. In the Expansion stage, an offset learning branch in a deformable convolution layer, referred to as expansion sampler'', seeks to sample increasingly less discriminative object regions, driven by an inverse supervision signal that maximizes image-level classification loss. The located more complete object region in the Expansion stage is then gradually narrowed down to the final object region during the Shrinkage stage. In the Shrinkage stage, the offset learning branch of another deformable convolution layer referred to as theshrinkage sampler'', is introduced to exclude the false positive background regions attended in the Expansion stage to improve the precision of the localization maps. We conduct various experiments on PASCAL VOC 2012 and MS COCO 2014 to well demonstrate the superiority of our method over other state-of-the-art methods for Weakly-Supervised Semantic Segmentation. The code is available at https://github.com/TyroneLi/ESOL_WSSS. Jinlong Li 0003, Zequn Jie, Xu Wang 0006, Xiaolin Wei, Lin Ma 0002 |
NeurIPS | 5 |
| 2022 | Beyond Monocular Deraining: Parallel Stereo Deraining Network Via Semantic Prior
Kaihao Zhang, Wenhan Luo, Yanjiang Yu, Wenqi Ren, Fang Zhao 0006, Lin Ma 0002, Wei Liu 0005, Hongdong Li |
Int. J. Comput. Vis. | 7 |
| 2022 | Content-aware Recommendation via Dynamic Heterogeneous Graph Convolutional Network
Tingting Liang, Lin Ma 0002, Congying Xia, Yuyu Yin |
Knowl. Based Syst. | 2 |
| 2022 | Liquid Warping GAN With Attention: A Unified Framework for Human Image SynthesisabstractWe tackle human image synthesis, including human motion imitation, appearance transfer, and novel view synthesis, within a unified framework. It means that the model, once being trained, can be used to handle all these tasks. The existing task-specific methods mainly use 2D keypoints (pose) to estimate the human body structure. However, they only express the position information with no ability to characterize the personalized shape of the person and model the limb rotations. In this paper, we propose to use a 3D body mesh recovery module to disentangle the pose and shape. It can not only model the joint location and rotation but also characterize the personalized body shape. To preserve the source information, such as texture, style, color, and face identity, we propose an Attentional Liquid Warping GAN with Attentional Liquid Warping Block (AttLWB) that propagates the source information in both image and feature spaces to the synthesized reference. Specifically, the source features are extracted by a denoising convolutional auto-encoder for characterizing the source identity well. Furthermore, our proposed method can support a more flexible warping from multiple sources. To further improve the generalization ability of the unseen source images, a one/few-shot adversarial learning is applied. In detail, it first trains a model in an extensive training set. Then, it finetunes the model by one/few-shot unseen image(s) in a self-supervised way to generate high-resolution ( 512 ×512 and 1024 ×1024) results. Also, we build a new dataset, namely Impersonator (iPER) dataset, for the evaluation of human motion imitation, appearance transfer, and novel view synthesis. Extensive experiments demonstrate the effectiveness of our methods in terms of preserving face identity, shape consistency, and clothes details. All codes and dataset are available on https://impersonator.org/work/impersonator-plus-plus.html. Wen Liu 0003, Zhixin Piao, Zhi Tu, Wenhan Luo, Lin Ma 0002, Shenghua Gao |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2022 | Semantic Conditioned Dynamic Modulation for Temporal Sentence Grounding in VideosabstractTemporal sentence grounding in videos aims to localize one target video segment, which semantically corresponds to a given sentence. Unlike previous methods mainly focusing on matching semantics between the sentence and different video segments, in this paper, we propose a novel semantic conditioned dynamic modulation (SCDM) mechanism, which leverages the sentence semantics to modulate the temporal convolution operations for better correlating and composing the sentence-relevant video contents over time. The proposed SCDM also performs dynamically with respect to the diverse video contents so as to establish a precise semantic alignment between sentence and video. By coupling the proposed SCDM with a hierarchical temporal convolutional architecture, video segments with various temporal scales are composed and localized. Besides, more fine-grained clip-level actionness scores are also predicted with the SCDM-coupled temporal convolution on the bottom layer of the overall architecture, which are further used to adjust the temporal boundaries of the localized segments and thereby lead to more accurate grounding results. Experimental results on benchmark datasets demonstrate that the proposed model can improve the temporal grounding accuracy consistently, and further investigation experiments also illustrate the advantages of SCDM on stabilizing the model training and associating relevant video contents for temporal sentence grounding. Our code for this paper is available at https://github.com/yytzsy/SCDM-TPAMI. Yitian Yuan, Lin Ma 0002, Jingwen Wang 0003, Wei Liu 0005, Wenwu Zhu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | Syntax Customized Video Captioning by Imitating Exemplar SentencesabstractEnhancing the diversity of sentences to describe video contents is an important problem arising in recent video captioning research. In this paper, we explore this problem from a novel perspective of customizing video captions by imitating exemplar sentence syntaxes. Specifically, given a video and any syntax-valid exemplar sentence, we introduce a new task of Syntax Customized Video Captioning (SCVC) aiming to generate one caption which not only semantically describes the video contents but also syntactically imitates the given exemplar sentence. To tackle the SCVC task, we propose a novel video captioning model, where a hierarchical sentence syntax encoder is first designed to extract the syntactic structure of the exemplar sentence, then a syntax conditioned caption decoder is devised to generate the syntactically structured caption expressing video semantics. As there is no available syntax customized groundtruth video captions, we tackle such a challenge by proposing a new training strategy, which leverages the traditional pairwise video captioning data and our collected exemplar sentences to accomplish the model learning. Extensive experiments, in terms of semantic, syntactic, fluency, and diversity evaluations, clearly demonstrate our model capability to generate syntax-varied and semantics-coherent video captions that well imitate different exemplar sentences with enriched diversities. Code is available at https://github.com/yytzsy/Syntax-Customized-Video-Captioning. Yitian Yuan, Lin Ma 0002, Wenwu Zhu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | Disentangled Feature Networks for Facial Portrait and Caricature GenerationabstractFacial portrait is an artistic form which draws faces by emphasizing discriminative or prominent parts of faces via various kinds of drawing tools. However, the complex interplay between the different facial factors, such as facial parts, background, and drawing styles, and the significant domain gap between natural facial images and their portrait counterparts makes the task challenging. In this paper, a flexible four-stream Disentangled Feature Networks (DFN) is proposed to learn disentangled feature representation of different facial factors and generate plausible portraits with reasonable exaggerations and richness in style. Four factors are encoded as embedding features, and combined to reconstruct facial portraits. Meanwhile, to make the process fully automatic (without manually specifying either portrait style or exaggerating form), we propose a new Adversarial Portrait Mapping Module (APMM) to map noise to the embedding feature space, as proxies for portrait style and exaggerating. Thanks to the proposedDFNandAPMM, we are able to manipulate the portrait style and facial geometric structures to generate a large number of portraits. Extensive experiments on two public datasets show that our proposed methods can generate a diverse set of artistic portraits. Kaihao Zhang, Wenhan Luo, Lin Ma 0002, Wenqi Ren, Hongdong Li |
IEEE Trans. Multim. | 3 |
| 2021 | Similarity Reasoning and Filtration for Image-Text MatchingabstractImage-text matching plays a critical role in bridging the vision and language, and great progress has been made by exploiting the global alignment between image and sentence, or local alignments between regions and words. However, how to make the most of these alignments to infer more accurate matching scores is still underexplored. In this paper, we propose a novel Similarity Graph Reasoning and Attention Filtration (SGRAF) network for image-text matching. Specifically, the vector-based similarity representations are firstly learned to characterize the local and global alignments in a more comprehensive manner, and then the Similarity Graph Reasoning (SGR) module relying on one graph convolutional neural network is introduced to infer relation-aware similarities with both the local and global alignments. The Similarity Attention Filtration (SAF) module is further developed to integrate these alignments effectively by selectively attending on the significant and representative alignments and meanwhile casting aside the interferences of non-meaningful alignments. We demonstrate the superiority of the proposed method with achieving state-of-the-art performances on the Flickr30K and MSCOCO datasets, and the good interpretability of SGR and SAF with extensive qualitative experiments and analyses. Haiwen Diao, Ying Zhang 0021, Lin Ma 0002, Huchuan Lu |
AAAI | 3 |
| 2021 | Relation-aware Instance Refinement for Weakly Supervised Visual GroundingabstractVisual grounding, which aims to build a correspondence between visual objects and their language entities, plays a key role in cross-modal scene understanding. One promising and scalable strategy for learning visual grounding is to utilize weak supervision from only image-caption pairs. Previous methods typically rely on matching query phrases directly to a precomputed, fixed object candidate pool, which leads to inaccurate localization and ambiguous matching due to lack of semantic relation constraints. In our paper, we propose a novel context-aware weakly-supervised learning method that incorporates coarse-to-fine object refinement and entity relation modeling into a two-stage deep network, capable of producing more accurate object representation and matching. To effectively train our network, we introduce a self-taught regression loss for the proposal locations and a classification loss based on parsed entity relations. Extensive experiments on two public benchmarks Flickr30K Entities and ReferItGame demonstrate the efficacy of our weakly grounding framework. The results show that we outperform the previous methods by a considerable margin, achieving 59.27% top-1 accuracy in Flickr30K Entities and 37.68% in the ReferItGame dataset respectively1. Yongfei Liu, Lin Ma 0002, Xuming He 0001 |
CVPR | 3 |
| 2021 | Neural Symbolic Representation Learning for Image CaptioningabstractTraditional image captioning models mainly rely on one encoder-decoder architecture to generate one natural sentence for a given image. Such an architecture mostly uses deep neural networks to extract the neural representations of the image while ignoring the information of abstractive concepts as well as their intertwined relationships conveyed in the image. To this end, to comprehensively characterize the image content and bridge the gap between neural representations and high-level abstractive concepts, we make the first attempt to investigate the ability of neural symbolic representation of the image for the image captioning task. We first parse and convert a given image to neural symbolic representation in the form of an attributed relational graph, with the nodes denoting the abstractive concepts and the branches indicating the relationships between connected nodes, respectively. By performing computations over the attributed relational graph, the neural symbolic representation evolves step by step, with the node and branch representations as well as their corresponding importance weights transiting step by step. Empirically, extensive experiments validate the effectiveness of the proposed method. It enables a more comprehensive understanding of the given image by integrating the neural representation and neural symbolic representation, with the state-of-the-art results being achieved on both the MSCOCO and Flickr30k datasets. Besides, the proposed neural symbolic representation is demonstrated to better generalize to other domains with significant performance improvements compared with existing methods on the cross domain image captioning task. Lin Ma 0002, Yanwei Fu 0001, Xiangyang Xue 0001 |
ICMR | 2 |
| 2021 | Two-stage Visual Cues Enhancement Network for Referring Image SegmentationabstractReferring Image Segmentation (RIS) aims at segmenting the target object from an image referred by one given natural language expression. The diverse and flexible expressions and complex visual contents in the images raise the RIS model with higher demands for investigating fine-grained matching behaviors between words in expressions and objects presented in images. However, such matching behaviors are hard to be learned and captured when the visual cues of referents (i.e. referred objects) are insufficient, as the referents of weak visual cues tend to be easily confused by cluttered background at boundary or even overwhelmed by salient objects in the image. And the insufficient visual cues issue can not be handled by the cross-modal fusion mechanisms as done in previous work.In this paper, we tackle this problem from a novel perspective of enhancing the visual information for the referents by devising a Two-stage Visual cues enhancement Network (TV-Net), where a novel Retrieval and Enrichment Scheme (RES) and an Adaptive Multi-resolution feature Fusion (AMF) module are proposed. Specifically, RES retrieves the most relevant image from an external data pool with regard to both the visual and textual similarities, and then enriches the visual information of the referent with the retrieved image for better multimodal feature learning. AMF further enhances the visual detailed information by incorporating the high-resolution feature maps from lower convolution layers of the image. Through the two-stage enhancement, our proposed TV-Net enjoys better performances in learning fine-grained matching behaviors between the natural language expression and image, especially when the visual information of the referent is inadequate, thus produces better segmentation results. Extensive experiments are conducted to validate the effectiveness of the proposed method on the RIS task, with our proposed TV-Net surpassing the state-of-the-art approaches on four benchmark datasets. Zequn Jie, Weixin Luo, Jingjing Chen 0001, Yu-Gang Jiang 0001, Xiaolin Wei, Lin Ma 0002 |
ACM Multimedia | 7 |
| 2021 | Cross-modality Discrepant Interaction Network for RGB-D Salient Object DetectionabstractThe popularity and promotion of depth maps have brought new vigor and vitality into salient object detection (SOD), and a mass of RGB-D SOD algorithms have been proposed, mainly concentrating on how to better integrate cross-modality features from RGB image and depth map. For the cross-modality interaction in feature encoder, existing methods either indiscriminately treat RGB and depth modalities, or only habitually utilize depth cues as auxiliary information of the RGB branch. Different from them, we reconsider the status of two modalities and propose a novel Cross-modality Discrepant Interaction Network (CDINet) for RGB-D SOD, which differentially models the dependence of two modalities according to the feature representations of different layers. To this end, two components are designed to implement the effective cross-modality interaction: 1) the RGB-induced Detail Enhancement (RDE) module leverages RGB modality to enhance the details of the depth features in low-level encoder stage. 2) the Depth-induced Semantic Enhancement (DSE) module transfers the object positioning and internal consistency of depth features to the RGB branch in high-level encoder stage. Furthermore, we also design a Dense Decoding Reconstruction (DDR) structure, which constructs a semantic block by combining multi-level encoder features to upgrade the skip connection in the feature decoding. Extensive experiments on five benchmark datasets demonstrate that our network outperforms $15$ state-of-the-art methods both quantitatively and qualitatively. Our code is publicly available at:https://rmcong.github.io/proj_CDINet.html. Chen Zhang 0013, Runmin Cong, Qinwei Lin, Lin Ma 0002, Feng Li 0037, Yao Zhao 0001, Sam Kwong |
ACM Multimedia | 4 |
| 2021 | Unsupervised text-to-image synthesis
Yanlong Dong, Ying Zhang 0021, Lin Ma 0002, Zhi Wang 0001, Jiebo Luo 0001 |
Pattern Recognit. | 3 |
| 2021 | Progressive Point Cloud Upsampling via Differentiable RenderingabstractIn this paper, we propose one novel progressive point cloud upsampling framework to tackle the non-uniform distribution issue during the point cloud upsampling process. Specifically, we design an Up-UNet feature expansion module which is capable of learning the local and global point features via a down-feature operator and an up-feature operator, respectively, to alleviate the non-uniform distribution issue and remove the outliers. Moreover, we design a hybrid loss function considering both the multi-scale reconstruction loss and the rendering loss. The multi-scale reconstruction loss enables each upsampling module to generate a denser point cloud, while the rendering loss via point-based differentiable rendering ensures that the proposed model preserves the point cloud structures. Extensive experimental results demonstrate that our proposed model achieves state-of-the-art performance in terms of both qualitative and quantitative evaluations. Xu Wang 0006, Lin Ma 0002, Shiqi Wang 0001, Sam Kwong, Jianmin Jiang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2021 | Pyramid Global Context Network for Image DehazingabstractHaze caused by atmospheric scattering and absorption would severely affect scene visibility of an image. Thus, image dehazing for haze removal has been widely studied in the literature. Within a hazy image, haze is not confined in a small local patch/position, while widely diffusing in a whole image. Under this circumstance, global context is a crucial factor in the success of dehazing, which was seldom investigated in existing dehazing algorithms. In the literature, the global context (GC) block has been designed to learn point-wise long-range dependencies of an image for global context modeling; however, patch-wise long-range dependencies were ignored. To image dehazing, patch-wise long-range dependencies should be highlighted to cooperate with patch-wise operations of image dehazing. In this paper, we first extend the point-wise GC into a Pyramid Global Context (PGC), which is a multi-scale GC, after undergoing the pyramid pooling. Thus, patch-wise long-range dependencies can be explored by the PGC. Then, the proposed PGC is plugged into a U-Net, getting an attentive U-Net. Further, the attentive U-Net is optimized by importing ResNet's shortcut connection and dilated convolution. Thus, the finalized dehazing model can explore both long-range and patch-wise context dependencies for global context modeling, which is crucial for image dehazing. The extensive experiments on synthetic databases and real-world hazy images demonstrate the superiority of our model over other representative state-of-the-art models from both quantitative and qualitative comparisons. Dong Zhao 0016, Long Xu 0001, Lin Ma 0002, Jia Li 0003, Yihua Yan |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2021 | Coupled Network for Robust Pedestrian Detection With Gated Multi-Layer Feature Extraction and Deformable Occlusion HandlingabstractPedestrian detection methods have been significantly improved with the development of deep convolutional neural networks. Nevertheless, detecting ismall-scaled pedestrians and occluded pedestrians remains a challenging problem. In this paper, we propose a pedestrian detection method with a couple-network to simultaneously address these two issues. One of the sub-networks, the gated multi-layer feature extraction sub-network, aims to adaptively generate discriminative features for pedestrian candidates in order to robustly detect pedestrians with large variations on scale. The second sub-network targets on handling the occlusion problem of pedestrian detection by using deformable regional region of interest (RoI)-pooling. We investigate two different gate units for the gated sub-network, namely, the channel-wise gate unit and the spatio-wise gate unit, which can enhance the representation ability of the regional convolutional features among the channel dimensions or across the spatial domain, repetitively. Ablation studies have validated the effectiveness of both the proposed gated multi-layer feature extraction sub-network and the deformable occlusion handling sub-network. With the coupled framework, our proposed pedestrian detector achieves promising results on both two pedestrian datasets, especially on detecting small or occluded pedestrians. On the CityPersons dataset, the proposed detector achieves the lowest missing rates (i.e. 40.78% and 34.60%) on detecting small and occluded pedestrians, surpassing the second best comparison method by 6.0% and 5.87%, respectively. Tianrui Liu 0001, Wenhan Luo, Lin Ma 0002, Junjie Huang 0001, Tania Stathaki, Tianhong Dai |
IEEE Trans. Image Process. | 3 |
| 2021 | Quality Evaluation for Image Retargeting With Instance SemanticsabstractTo meet the ever-increasing demand for devices with diversified displays, image retargeting has become a prevalent technique for adaptive image resizing. In practice, the retargeting operation inevitably causes impairments in the images; thus, image retargeting quality assessment (IRQA) is urgently needed and, can be used to guide algorithm optimization, selection and design. Unlike traditional image quality assessment, image retargeting introduces geometric distortions, which typically affect high-level image semantics. With this motivation, this paper presents a quality evaluation model for image retargeting based on INstance SEMantics (INSEM). Considering that the human visual system (HVS) perceives images highly dependent on apprehensible areas and that impairments in image retargeting mainly degrade the salient instances, an image instance is utilized as the basic semantic unit, and a top-down method is devised to extract instance-level semantic features for IRQA. In addition, taking into account the influence of semantic categories on the perception of retargeting quality, we further propose Semantic-based self-adaptive pooling (SSAP) to integrate instance-based semantic features. Finally, global features are incorporated to generate quality scores that are more consistent with people's perceptions. Extensive experiments and comparisons of three public databases, in terms of both intradatabase and cross-database settings, demonstrate the superiority of the proposed metric over state-of-the-art methods. Leida Li, Jinjian Wu, Lin Ma 0002, Yuming Fang 0001 |
IEEE Trans. Multim. | 4 |
| 2021 | PFAN++: Bi-Directional Image-Text Retrieval With Position Focused Attention NetworkabstractBi-directional image-text retrieval and matching attract much attention recently. This cross-domain task demands a fine understanding of both modalities for learning a measure of different modality data. In this paper, we propose a novel position focused attention network to investigate the relation between the visual and the textual views. This work integrates the prior object position to enhance the visual-text joint-embedding learning. The image is first split into blocks, which are treated as the basic position cells, and the position of an image region is inferred. Then, we propose a position attention to model the relations between the image region and position cells. Finally, we generate a valuable position feature to further enhance the region expression and model a more reliable relationship between the visual image and the textual sentence. Experiments on the popular datasets Flickr30K and MS-COCO show the effectiveness of the proposed method. Besides the public datasets, we also conduct experiments on our collected practical large-scale news dataset (Tencent-News) to validate the practical application value of the proposed method. As far as we know, this is the first attempt to test the performance on the practical application. Our method achieves the competitive performance on all of these three datasets. Yaxiong Wang, Xiuxiu Bai, Xueming Qian, Lin Ma 0002 |
IEEE Trans. Multim. | 5 |
| 2021 | CASNet: A Cross-Attention Siamese Network for Video Salient Object DetectionabstractRecent works on video salient object detection have demonstrated that directly transferring the generalization ability of image-based models to video data without modeling spatial-temporal information remains nontrivial and challenging. Considering both intraframe accuracy and interframe consistency of saliency detection, this article presents a novel cross-attention based encoder-decoder model under the Siamese framework (CASNet) for video salient object detection. A baseline encoder-decoder model trained with Lovász softmax loss function is adopted as a backbone network to guarantee the accuracy of intraframe salient object detection. Self- and cross-attention modules are incorporated into our model in order to preserve the saliency correlation and improve intraframe salient detection consistency. Extensive experimental results obtained by ablation analysis and cross-data set validation demonstrate the effectiveness of our proposed method. Quantitative results indicate that our CASNet model outperforms 19 state-of-the-art image- and video-based methods on six benchmark data sets. Yuzhu Ji, Haijun Zhang 0002, Zequn Jie, Lin Ma 0002, Q. M. Jonathan Wu |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2020 | Feature Deformation Meta-Networks in Image Captioning of Novel ObjectsabstractThis paper studies the task of image captioning with novel objects, which only exist in testing images. Intrinsically, this task can reflect the generalization ability of models in understanding and captioning the semantic meanings of visual concepts and objects unseen in training set, sharing the similarity to one/zero-shot learning. The critical difficulty thus comes from that no paired images and sentences of the novel objects can be used to help train the captioning model. Inspired by recent work (Chen et al. 2019b) that boosts one-shot learning by learning to generate various image deformations, we propose learning meta-networks for deforming features for novel object captioning. To this end, we introduce the feature deformation meta-networks (FDM-net), which is trained on source data, and learn to adapt to the novel object features detected by the auxiliary detection model. FDM-net includes two sub-nets: feature deformation, and scene graph sentence reconstruction, which produce the augmented image features and corresponding sentences, respectively. Thus, rather than directly deforming images, FDM-net can efficiently and dynamically enlarge the paired images and texts by learning to deform image features. Extensive experiments are conducted on the widely used novel object captioning dataset, and the results show the effectiveness of our FDM-net. Ablation study and qualitative visualization further give insights of our model. Tingjia Cao, Lin Ma 0002, Yanwei Fu 0001, Yu-Gang Jiang 0001, Xiangyang Xue 0001 |
AAAI | 4 |
| 2020 | Recurrent Nested Model for Sequence Generation
Lin Ma 0002 |
AAAI | 2 |
| 2020 | Temporally Grounding Language Queries in Videos by Contextual Boundary-Aware PredictionabstractThe task of temporally grounding language queries in videos is to temporally localize the best matched video segment corresponding to a given language (sentence). It requires certain models to simultaneously perform visual and linguistic understandings. Previous work predominantly ignores the precision of segment localization. Sliding window based methods use predefined search window sizes, which suffer from redundant computation, while existing anchor-based approaches fail to yield precise localization. We address this issue by proposing an end-to-end boundary-aware model, which uses a lightweight branch to predict semantic boundaries corresponding to the given linguistic information. To better detect semantic boundaries, we propose to aggregate contextual information by explicitly modeling the relationship between the current element and its neighbors. The most confident segments are subsequently selected based on both anchor and boundary predictions at the testing stage. The proposed model, dubbed Contextual Boundary-aware Prediction (CBP), outperforms its competitors with a clear margin on three public datasets. Jingwen Wang 0003, Lin Ma 0002 |
AAAI | 2 |
| 2020 | Fine-Grained Image-to-Image Transformation Towards Visual RecognitionabstractExisting image-to-image transformation approaches primarily focus on synthesizing visually pleasing data. Generating images with correct identity labels is challenging yet much less explored. It is even more challenging to deal with image transformation tasks with large deformation in poses, viewpoints, or scales while preserving the identity, such as face rotation and object viewpoint morphing. In this paper, we aim at transforming an image with a fine-grained category to synthesize new images that preserve the identity of the input image, which can thereby benefit the subsequent fine-grained image recognition and few-shot learning tasks. The generated images, transformed with large geometric deformation, do not necessarily need to be of high visual quality but are required to maintain as much identity information as possible. To this end, we adopt a model based on generative adversarial networks to disentangle the identity related and unrelated factors of an image. In order to preserve the fine-grained contextual details of the input image during the deformable transformation, a constrained nonalignment connection method is proposed to construct learnable highways between intermediate convolution blocks in the generator. Moreover, an adaptive identity modulation mechanism is proposed to transfer the identity information into the output image effectively. Extensive experiments on the CompCars and Multi-PIE datasets demonstrate that our model preserves the identity of the generated images much better than the state-of-the-art image-to-image transformation models, and as a result significantly boosts the visual recognition performance in fine-grained few-shot learning. Wei Xiong 0008, Yixuan Zhang 0008, Wenhan Luo, Lin Ma 0002, Jiebo Luo 0001 |
CVPR | 5 |
| 2020 | Cops-Ref: A New Dataset and Task on Compositional Referring Expression ComprehensionabstractReferring expression comprehension (REF) aims at identifying a particular object in a scene by a natural language expression. It requires joint reasoning over the textual and visual domains to solve the problem. Some popular referring expression datasets, however, fail to provide an ideal test bed for evaluating the reasoning ability of the models, mainly because 1) their expressions typically describe only some simple distinctive properties of the object and 2) their images contain limited distracting information. To bridge the gap, we propose a new dataset for visual reasoning in context of referring expression comprehension with two main features. First, we design a novel expression engine rendering various reasoning logics that can be flexibly combined with rich visual properties to generate expressions with varying compositionality. Second, to better exploit the full reasoning chain embodied in an expression, we propose a new test setting by adding additional distracting images containing objects sharing similar properties with the referent, thus minimising the success rate of reasoning-free cross-domain alignment. We evaluate several state-of-the-art REF models, but find none of them can achieve promising performance. A proposed modular hard mining strategy performs the best but still leaves substantial room for improvement. Zhenfang Chen, Peng Wang 0023, Lin Ma 0002, Kwan-Yee Kenneth Wong, Qi Wu 0001 |
CVPR | 3 |
| 2020 | Deblurring by Realistic BlurringabstractExisting deep learning methods for image deblurring typically train models using pairs of sharp images and their blurred counterparts. However, synthetically blurring images does not necessarily model the blurring process in real-world scenarios with sufficient accuracy. To address this problem, we propose a new method which combines two GAN models, i.e., a learning-to-Blur GAN (BGAN) and learning-to-DeBlur GAN (DBGAN), in order to learn a better model for image deblurring by primarily learning how to blur images. The first model, BGAN, learns how to blur sharp images with unpaired sharp and blurry image sets, and then guides the second model, DBGAN, to learn how to correctly deblur such images. In order to reduce the discrepancy between real blur and synthesized blur, a relativistic blur loss is leveraged. As an additional contribution, this paper also introduces a Real-World Blurred Image (RWBI) dataset including diverse blurry images. Our experiments show that the proposed method achieves consistently superior quantitative performance as well as higher perceptual quality on both the newly proposed dataset and the public GOPRO dataset. Kaihao Zhang, Wenhan Luo, Yiran Zhong, Lin Ma 0002, Björn Stenger, Wei Liu 0005, Hongdong Li |
CVPR | 4 |
| 2020 | Context-Gated Convolution
Xudong Lin 0003, Lin Ma 0002, Wei Liu 0005, Shih-Fu Chang |
ECCV (18) | 2 |
| 2020 | Beyond Monocular Deraining: Stereo Image Deraining via Semantic Understanding
Kaihao Zhang, Wenhan Luo, Wenqi Ren, Jingwen Wang 0003, Fang Zhao 0006, Lin Ma 0002, Hongdong Li |
ECCV (27) | 6 |
| 2020 | Lossy Geometry Compression Of 3d Point Cloud Data Via An Adaptive Octree-Guided NetworkabstractIn this paper, we propose a deep learning based framework for point cloud geometry lossy compression via hybrid representation of point cloud. First, the input raw 3D point cloud data is adaptively decomposed into non-overlapping local patches through adaptive Octree decomposition and clustering. Second, a framework of point cloud auto-encoder network with quantization layer is proposed for learning compact latent feature representation from each patch. Specifically, the proposed point cloud auto-encoder networks with different input size are trained for achieving optimal rate-distortion (RD) performance. Final, bitstream specifications of proposed compression systems with additional signaled meta-data and header information are designed to support parallel decoding and successive reconstruction. Experimental results shows that our proposed method can achieve 40.20% bitrate saving in average than the existing standard Geometry based Point Cloud Compression (G-PCC) codec. Xuanzheng Wen, Xu Wang 0006, Junhui Hou, Lin Ma 0002, Yu Zhou 0027, Jianmin Jiang |
ICME | 4 |
| 2020 | Grasp for Stacking via Deep Reinforcement LearningabstractIntegrated robotic arm system should contain both grasp and place actions. However, most grasping methods focus more on how to grasp objects, while ignoring the placement of the grasped objects, which limits their applications in various industrial environments. In this research, we propose a model-free deep Q-learning method to learn the grasping-stacking strategy end-to-end from scratch. Our method maps the images to the actions of the robotic arm through two deep networks: the grasping network (GNet) using the observation of the desk and the pile to infer the gripper's position and orientation for grasping, and the stacking network (SNet) using the observation of the platform to infer the optimal location when placing the grasped object. To make a long-range planning, the two observations are integrated in the grasping for stacking network (GSN). We evaluate the proposed GSN on a grasping-stacking task in both simulated and real-world scenarios. Wei Zhang 0021, Ran Song 0001, Lin Ma 0002, Yibin Li 0001 |
ICRA | 4 |
| 2020 | Unpaired Image Enhancement with Quality-Attention Generative Adversarial NetworkabstractIn this work, we aim to learn an unpaired image enhancement model, which can enrich low-quality images with the characteristics of high-quality images provided by users. We propose a quality attention generative adversarial network (QAGAN) trained on unpaired data based on the bidirectional Generative Adversarial Network (GAN) embedded with a quality attention module (QAM). The key novelty of the proposed QAGAN lies in the injected QAM for the generator such that it learns domain-relevant quality attention directly from the two domains. More specifically, the proposed QAM allows the generator to effectively select semantic-related characteristics from the spatial-wise and adaptively incorporate style-related attributes from the channel-wise, respectively. Therefore, in our proposed QAGAN, not only discriminators but also the generator can directly access both domains which significantly facilitate the generator to learn the mapping function. Extensive experimental results show that, compared with the state-of-the-art methods based on unpaired learning, our proposed method achieves better performance in both objective and subjective evaluations. Zhangkai Ni, Wenhan Yang, Shiqi Wang 0001, Lin Ma 0002, Sam Kwong |
ACM Multimedia | 4 |
| 2020 | Controllable Video Captioning with an Exemplar SentenceabstractIn this paper, we investigate a novel and challenging task, namely controllable video captioning with an exemplar sentence. Formally, given a video and a syntactically valid exemplar sentence, the task aims to generate one caption which not only describes the semantic contents of the video, but also follows the syntactic form of the given exemplar sentence. In order to tackle such an exemplar-based video captioning task, we propose a novel Syntax Modulated Caption Generator (SMCG) incorporated in an encoder-decoder-reconstructor architecture. The proposed SMCG takes video semantic representation as an input, and conditionally modulates the gates and cells of long short-term memory network with respect to the encoded syntactic information of the given exemplar sentence. Therefore, SMCG is able to control the states for word prediction and achieve the syntax customized caption generation. We conduct experiments by collecting auxiliary exemplar sentences for two public video captioning datasets. Extensive experimental results demonstrate the effectiveness of our approach on generating syntax controllable and semantic preserved video captions. By providing different exemplar sentences, our approach is capable of producing different captions with various syntactic structures, thus indicating a promising way to strengthen the diversity of video captioning. Code for this paper is available at https://github.com/yytzsy/SMCG. Yitian Yuan, Lin Ma 0002, Jingwen Wang 0003, Wenwu Zhu 0001 |
ACM Multimedia | 2 |
| 2020 | Every Moment Matters: Detail-Aware Networks to Bring a Blurry Image AliveabstractMotion-blurred images are the result of light accumulation over the period of camera exposure time, during which the camera and objects in the scene are in relative motion to each other. The inverse process of extracting an image sequence from a single motion-blurred image is an ill-posed vision problem. One key challenge is that the motions across frames are subtle, which makes the generating networks difficult to capture them and thus the recovery sequences lack motion details. In order to alleviate this problem, we propose a detail-aware network with three consecutive stages to improve the reconstruction quality by addressing specific aspects in the recovery process. The detail-aware network firstly models the dynamics using a cycle flow loss, resolving the temporal ambiguity of the reconstruction in the first stage. Then, a GramNet is proposed in the second stage to refine subtle motion between continuous frames using Gram matrices as motion representation. Finally, we introduce a HeptaGAN in the third stage to bridge the continuous and discrete nature of exposure time and recovered frames, respectively, in order to maintain rich detail. Experiments show that the proposed detail-aware networks produce sharp image sequences with rich details and subtle motion, outperforming the state-of-the-art methods. Kaihao Zhang, Wenhan Luo, Björn Stenger, Wenqi Ren, Lin Ma 0002, Hongdong Li |
ACM Multimedia | 5 |
| 2020 | Reconstruct and Represent Video Contents for Captioning via Reinforcement LearningabstractIn this paper, the problem of describing visual contents of a video sequence with natural language is addressed. Unlike previous video captioning work mainly exploiting the cues of video contents to make a language description, we propose a reconstruction network (RecNet) in a novel encoder-decoder-reconstructor architecture, which leverages both forward (video to sentence) and backward (sentence to video) flows for video captioning. Specifically, the encoder-decoder component makes use of the forward flow to produce a sentence based on the encoded video semantic features. Two types of reconstructors are subsequently proposed to employ the backward flow and reproduce the video features from local and global perspectives, respectively, capitalizing on the hidden state sequence generated by the decoder. Moreover, in order to make a comprehensive reconstruction of the video features, we propose to fuse the two types of reconstructors together. The generation loss yielded by the encoder-decoder component and the reconstruction loss introduced by the reconstructor are jointly cast into training the proposed RecNet in an end-to-end fashion. Furthermore, the RecNet is fine-tuned by CIDEr optimization via reinforcement learning, which significantly boosts the captioning performance. Experimental results on benchmark datasets demonstrate that the proposed reconstructor can boost the performance of video captioning consistently. Wei Zhang 0021, Bairui Wang, Lin Ma 0002, Wei Liu 0005 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2020 | Matching Image and Sentence With Multi-Faceted RepresentationsabstractIn this paper, we propose a novel multimodal matching model for the image and sentence based on their multiple representations. Each representation of the image or sentence undergoes an independent neural network, consisting of multiple layers of nonlinear mappings to yield the corresponding embedding. Besides exploiting the image and sentence relationship based on their embeddings, we propose one novel loss to further exploit the relationship within each single modality, namely, image and sentence based on the yielded multiple embeddings, which is used to train the neural networks simultaneously. The experimental results demonstrate that multiple representations can help to capture the image contents and the sentence semantic meaning more precisely, thus making comprehensive exploitations of the complicated image and sentence matching relationship. More concretely, the proposed matching model significantly outperforms the state-of-the-art approaches in bidirectional image-sentence retrieval on the Flickr8K, Flickr30K, and Microsoft COCO datasets. Lin Ma 0002, Zequn Jie, Yu-Gang Jiang 0001, Wei Liu 0005 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2020 | Multi-Exposure Decomposition-Fusion Model for High Dynamic Range Image Saliency DetectionabstractHigh dynamic range (HDR) imaging techniques have witnessed a great improvement in the past few decades. However, saliency detection task on HDR content is still far from well explored. In this paper, we introduce a multi-exposure decomposition-fusion model for HDR image saliency detection inspired by the brightness adaption mechanism. The proposed model is composed of three modules. Firstly, a decomposition module converts the input raw HDR image into a stack of LDR images by uniformly sampling the exposure time range. Secondly, a saliency region proposal network is employed to generate the candidate saliency maps for each LDR image in the exposure stack. Finally, an uncertainty weighting based fusion algorithm is applied to generate the overall saliency map for the input HDR image by merging the obtained LDR saliency maps. Extensive experiments show that our proposed model achieves superior performance compared with the state-of-the-art methods on the existing HDR eye fixation databases. The source code of the proposed model are made publicly available at https://github.com/sunnycia/DFHSal. Xu Wang 0006, Zhenhao Sun, Qiudan Zhang, Yuming Fang 0001, Lin Ma 0002, Shiqi Wang 0001, Sam Kwong |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2020 | Towards Unsupervised Deep Image Enhancement With Generative Adversarial NetworkabstractImproving the aesthetic quality of images is challenging and eager for the public. To address this problem, most existing algorithms are based on supervised learning methods to learn an automatic photo enhancer for paired data, which consists of low-quality photos and corresponding expert-retouched versions. However, the style and characteristics of photos retouched by experts may not meet the needs or preferences of general users. In this paper, we present an unsupervised image enhancement generative adversarial network (UEGAN), which learns the corresponding image-to-image mapping from a set of images with desired characteristics in an unsupervised manner, rather than learning on a large number of paired images. The proposed model is based on single deep GAN which embeds the modulation and attention mechanisms to capture richer global and local features. Based on the proposed model, we introduce two losses to deal with the unsupervised image enhancement: (1) fidelity loss, which is defined as a l2 regularization in the feature domain of a pre-trained VGG network to ensure the content between the enhanced image and the input image is the same, and (2) quality loss that is formulated as a relativistic hinge adversarial loss to endow the input image the desired characteristics. Both quantitative and qualitative results show that the proposed model effectively improves the aesthetic quality of images. Zhangkai Ni, Wenhan Yang, Shiqi Wang 0001, Lin Ma 0002, Sam Kwong |
IEEE Trans. Image Process. | 4 |
| 2019 | Localizing Natural Language in VideosabstractIn this paper, we consider the task of natural language video localization (NLVL): given an untrimmed video and a natural language description, the goal is to localize a segment in the video which semantically corresponds to the given natural language description. We propose a localizing network (LNet), working in an end-to-end fashion, to tackle the NLVL task. We first match the natural sentence and video sequence by cross-gated attended recurrent networks to exploit their fine-grained interactions and generate a sentence-aware video representation. A self interactor is proposed to perform crossframe matching, which dynamically encodes and aggregates the matching evidences. Finally, a boundary model is proposed to locate the positions of video segments corresponding to the natural sentence description by predicting the starting and ending points of the segment. Extensive experiments conducted on the public TACoS and DiDeMo datasets demonstrate that our proposed model performs effectively and efficiently against the state-of-the-art approaches. Jingyuan Chen 0003, Lin Ma 0002, Zequn Jie, Jiebo Luo 0001 |
AAAI | 2 |
| 2019 | Hierarchical Photo-Scene Encoder for Album StorytellingabstractIn this paper, we propose a novel model with a hierarchical photo-scene encoder and a reconstructor for the task of album storytelling. The photo-scene encoder contains two subencoders, namely the photo and scene encoders, which are stacked together and behave hierarchically to fully exploit the structure information of the photos within an album. Specifically, the photo encoder generates semantic representation for each photo while exploiting temporal relationships among them. The scene encoder, relying on the obtained photo representations, is responsible for detecting the scene changes and generating scene representations. Subsequently, the decoder dynamically and attentively summarizes the encoded photo and scene representations to generate a sequence of album representations, based on which a story consisting of multiple coherent sentences is generated. In order to fully extract the useful semantic information from an album, a reconstructor is employed to reproduce the summarized album representations based on the hidden states of the decoder. The proposed model can be trained in an end-to-end manner, which results in an improved performance over the state-of-the-arts on the public visual storytelling (VIST) dataset. Ablation studies further demonstrate the effectiveness of the proposed hierarchical photo-scene encoder and reconstructor. Bairui Wang, Lin Ma 0002, Wei Zhang 0021 |
AAAI | 2 |
| 2019 | Cousin Network Guided Sketch Recognition via Latent Attribute WarehouseabstractWe study the problem of sketch image recognition. This problem is plagued with two major challenges: 1) sketch images are often scarce in contrast to the abundance of natural images, rendering the training task difficult, and 2) the significant domain gap between sketch image and its natural image counterpart makes the task of bridging the two domains challenging. In order to overcome these challenges, in this paper we propose to transfer the knowledge of a network learned from natural images to a sketch network - a new deep net architecture which we term as cousin network. This network guides a sketch-recognition network to extract more relevant features that are close to those of natural images, via adversarial training. Moreover, to enhance the transfer ability of the classification model, a sketch-to-image attribute warehouse is constructed to approximate the transformation between the sketch domain and the real image domain. Extensive experiments conducted on the TU-Berlin dataset show that the proposed model is able to efficiently distill knowledge from natural images and achieves superior performance than the current state of the art. Kaihao Zhang, Wenhan Luo, Lin Ma 0002, Hongdong Li |
AAAI | 3 |
| 2019 | Weakly-Supervised Spatio-Temporally Grounding Natural Sentence in VideoabstractIn this paper, we address a novel task, namely weakly-supervised spatio-temporally grounding natural sentence in video.Specifically, given a natural sentence and a video, we localize a spatio-temporal tube in the video that semantically corresponds to the given sentence, with no reliance on any spatio-temporal annotations during training.First, a set of spatiotemporal tubes, referred to as instances, are extracted from the video.We then encode these instances and the sentence using our proposed attentive interactor which can exploit their fine-grained relationships to characterize their matching behaviors.Besides a ranking loss, a novel diversity loss is introduced to train the proposed attentive interactor to strengthen the matching behaviors of reliable instance-sentence pairs and penalize the unreliable ones.Moreover, we also contribute a dataset, called VID-sentence, based on the Im-ageNet video object detection dataset, to serve as a benchmark for our task.Extensive experimental results demonstrate the superiority of our model over the baseline approaches.Our code and the constructed VID-sentence dataset are available at: https://github.com/ JeffCHEN2017/WSSTG.git. Zhenfang Chen, Lin Ma 0002, Wenhan Luo, Kwan-Yee Kenneth Wong |
ACL (1) | 2 |
| 2019 | Image Deformation Meta-Networks for One-Shot LearningabstractHumans can robustly learn novel visual concepts even when images undergo various deformations and loose certain information. Mimicking the same behavior and synthesizing deformed instances of new concepts may help visual recognition systems perform better one-shot learning, i.e., learning concepts from one or few examples. Our key insight is that, while the deformed images may not be visually realistic, they still maintain critical semantic information and contribute significantly to formulating classifier decision boundaries. Inspired by the recent progress of meta-learning, we combine a meta-learner with an image deformation sub-network that produces additional training examples, and optimize both models in an end-to-end manner. The deformation sub-network learns to deform images by fusing a pair of images --- a probe image that keeps the visual content and a gallery image that diversifies the deformations. We demonstrate results on the widely used one-shot learning benchmarks (miniImageNet and ImageNet 1K Challenge datasets), which significantly outperform state-of-the-art approaches. Zitian Chen, Yanwei Fu 0001, Yu-Xiong Wang, Lin Ma 0002, Wei Liu 0005, Martial Hebert |
CVPR | 4 |
| 2019 | Spatio-Temporal Video Re-Localization by Warp LSTMabstractThe need for efficiently finding the video content a user wants is increasing because of the erupting of user-generated videos on the Web. Existing keyword-based or content-based video retrieval methods usually determine what occurs in a video but not when and where. In this paper, we make an answer to the question of when and where by formulating a new task, namely spatio-temporal video re-localization. Specifically, given a query video and a reference video, spatio-temporal video re-localization aims to localize tubelets in the reference video such that the tubelets semantically correspond to the query. To accurately localize the desired tubelets in the reference video, we propose a novel warp LSTM network, which propagates the spatio-temporal information for a long period and thereby captures the corresponding long-term dependencies. Another issue for spatio-temporal video re-localization is the lack of properly labeled video datasets. Therefore, we reorganize the videos in the AVA dataset to form a new dataset for spatio-temporal video re-localization research. Extensive experimental results show that the proposed model achieves superior performances over the designed baselines on the spatio-temporal video re-localization task. Yang Feng 0001, Lin Ma 0002, Wei Liu 0005, Jiebo Luo 0001 |
CVPR | 2 |
| 2019 | Unsupervised Image CaptioningabstractDeep neural networks have achieved great successes on the image captioning task. However, most of the existing models depend heavily on paired image-sentence datasets, which are very expensive to acquire. In this paper, we make the first attempt to train an image captioning model in an unsupervised manner. Instead of relying on manually labeled image-sentence pairs, our proposed model merely requires an image set, a sentence corpus, and an existing visual concept detector. The sentence corpus is used to teach the captioning model how to generate plausible sentences. Meanwhile, the knowledge in the visual concept detector is distilled into the captioning model to guide the model to recognize the visual concepts in an image. In order to further encourage the generated captions to be semantically consistent with the image, the image and caption are projected into a common latent space so that they can reconstruct each other. Given that the existing sentence corpora are mainly designed for linguistic research and are thus with little reference to image contents, we crawl a large-scale image description corpus of two million natural sentences to facilitate the unsupervised image captioning scenario. Experimental results show that our proposed model is able to produce quite promising results without any caption annotations. Yang Feng 0001, Lin Ma 0002, Wei Liu 0005, Jiebo Luo 0001 |
CVPR | 2 |
| 2019 | Multi-Granularity Generator for Temporal Action ProposalabstractTemporal action proposal generation is an important task, aiming to localize the video segments containing human actions in an untrimmed video. In this paper, we propose a multi-granularity generator (MGG) to perform the temporal action proposal from different granularity perspectives, relying on the video visual features equipped with the position embedding information. First, we propose to use a bilinear matching model to exploit the rich local information within the video sequence. Afterwards, two components, namely segment proposal producer (SPP) and frame actionness producer (FAP), are combined to perform the task of temporal action proposal at two distinct granularities. SPP considers the whole video in the form of feature pyramid and generates segment proposals from one coarse perspective, while FAP carries out a finer actionness evaluation for each video frame. Our proposed MGG can be trained in an end-to-end fashion. Through temporally adjusting the segment proposals with fine-grained information based on frame actionness, MGG achieves the superior performance over state-of-the-art methods on the public THUMOS-14 and ActivityNet-1.3 datasets. Moreover, we employ existing action classifiers to perform the classification of the proposals generated by MGG, leading to significant improvements compared against the competing methods for the video detection task. Lin Ma 0002, Yifeng Zhang 0001, Wei Liu 0005, Shih-Fu Chang |
CVPR | 2 |
| 2019 | Learning Joint Gait Representation via Quintuplet Loss MinimizationabstractGait recognition is an important biometric method popularly used in video surveillance, where the task is to identify people at a distance by their walking patterns from video sequences. Most of the current successful approaches for gait recognition either use a pair of gait images to form a cross-gait representation or rely on a single gait image for unique-gait representation. These two types of representations emperically complement one another. In this paper, we propose a new Joint Unique-gait and Cross-gait Network (JUCNet), to combine the advantages of unique-gait representation with that of cross-gait representation, leading to an significantly improved performance. Another key contribution of this paper is a novel quintuplet loss function, which simultaneously increases the inter-class differences by pushing representations extracted from different subjects apart and decreases the intra-class variations by pulling representations extracted from the same subject together. Experiments show that our method achieves the state-of-the-art performance tested on standard benchmark datasets, demonstrating its superiority over existing methods. Kaihao Zhang, Wenhan Luo, Lin Ma 0002, Wei Liu 0005, Hongdong Li |
CVPR | 3 |
| 2019 | Liquid Warping GAN: A Unified Framework for Human Motion Imitation, Appearance Transfer and Novel View SynthesisabstractWe tackle the human motion imitation, appearance transfer, and novel view synthesis within a unified framework, which means that the model once being trained can be used to handle all these tasks. The existing task-specific methods mainly use 2D keypoints (pose) to estimate the human body structure. However, they only expresses the position information with no abilities to characterize the personalized shape of the individual person and model the limbs rotations. In this paper, we propose to use a 3D body mesh recovery module to disentangle the pose and shape, which can not only model the joint location and rotation but also characterize the personalized body shape. To preserve the source information, such as texture, style, color, and face identity, we propose a Liquid Warping GAN with Liquid Warping Block (LWB) that propagates the source information in both image and feature spaces, and synthesizes an image with respect to the reference. Specifically, the source features are extracted by a denoising convolutional auto-encoder for characterizing the source identity well. Furthermore, our proposed method is able to support a more flexible warping from multiple sources. In addition, we build a new dataset, namely Impersonator (iPER) dataset, for the evaluation of human motion imitation, appearance transfer, and novel view synthesis. Extensive experiments demonstrate the effectiveness of our method in several aspects, such as robustness in occlusion case and preserving face identity, shape consistency and clothes details. All codes and datasets are available on https://svip-lab.github.io/project/impersonator.html. Wen Liu 0003, Zhixin Piao, Jie Min, Wenhan Luo, Lin Ma 0002, Shenghua Gao |
ICCV | 5 |
| 2019 | Controllable Video Captioning With POS Sequence Guidance Based on Gated Fusion NetworkabstractIn this paper, we propose to guide the video caption generation with Part-of-Speech (POS) information, based on a gated fusion of multiple representations of input videos. We construct a novel gated fusion network, with one particularly designed cross-gating (CG) block, to effectively encode and fuse different types of representations, e.g., the motion and content features of an input video. One POS sequence generator relies on this fused representation to predict the global syntactic structure, which is thereafter leveraged to guide the video captioning generation and control the syntax of the generated sentence. Specifically, a gating strategy is proposed to dynamically and adaptively incorporate the global syntactic POS information into the decoder for generating each word. Experimental results on two benchmark datasets, namely MSR-VTT and MSVD, demonstrate that the proposed model can well exploit complementary information from multiple representations, resulting in improved performances. Moreover, the generated global POS information can well capture the global syntactic structure of the sentence, and thus be exploited to control the syntactic structure of the description. Such POS information not only boosts the video captioning performance but also improves the diversity of the generated captions. Our code is at: https://github.com/vsislab/Controllable_XGating. Bairui Wang, Lin Ma 0002, Wei Zhang 0021, Jingwen Wang 0003, Wei Liu 0005 |
ICCV | 2 |
| 2019 | Hallucinating Optical Flow Features for Video ClassificationabstractAppearance and motion are two key components to depict and characterize the video content. Currently, the two-stream models have achieved state-of-the-art performances on video classification. However, extracting motion information, specifically in the form of optical flow features, is extremely computationally expensive, especially for large-scale video classification. In this paper, we propose a motion hallucination network, namely MoNet, to imagine the optical flow features from the appearance features, with no reliance on the optical flow computation. Specifically, MoNet models the temporal relationships of the appearance features and exploits the contextual relationships of the optical flow features with concurrent connections. Extensive experimental results demonstrate that the proposed MoNet can effectively and efficiently hallucinate the optical flow features, which together with the appearance features consistently improve the video classification performances. Moreover, MoNet can help cutting down almost a half of computational and data-storage burdens for the two-stream video classification. Our code is available at: https://github.com/YongyiTang92/MoNet-Features Yongyi Tang, Lin Ma 0002, Lianqiang Zhou |
IJCAI | 2 |
| 2019 | Position Focused Attention Network for Image-Text MatchingabstractImage-text matching tasks have recently attracted a lot of attention in the computer vision field. The key point of this cross-domain problem is how to accurately measure the similarity between the visual and the textual contents, which demands a fine understanding of both modalities. In this paper, we propose a novel position focused attention network (PFAN) to investigate the relation between the visual and the textual views. In this work, we integrate the object position clue to enhance the visual-text joint-embedding learning. We first split the images into blocks, by which we infer the relative position of region in the image. Then, an attention mechanism is proposed to model the relations between the image region and blocks and generate the valuable position feature, which will be further utilized to enhance the region expression and model a more reliable relationship between the visual image and the textual sentence. Experiments on the popular datasets Flickr30K and MS-COCO show the effectiveness of the proposed method. Besides the public datasets, we also conduct experiments on our collected practical news dataset (Tencent-News) to validate the practical application value of proposed method. As far as we know, this is the first attempt to test the performance on the practical application. Our method can achieve the state-of-art performance on all of these three datasets. Yaxiong Wang, Xueming Qian, Lin Ma 0002 |
IJCAI | 4 |
| 2019 | Sentence Specified Dynamic Video Thumbnail GenerationabstractWith the tremendous growth of videos over the Internet, video thumbnails, providing video content previews, are becoming increasingly crucial to influencing users' online searching experiences. Conventional video thumbnails are generated once purely based on the visual characteristics of videos, and then displayed as requested. Hence, such video thumbnails, without considering the users' searching intentions, cannot provide a meaningful snapshot of the video contents that users concern. In this paper, we define a distinctively new task, namely sentence specified dynamic video thumbnail generation, where the generated thumbnails not only provide a concise preview of the original video contents but also dynamically relate to the users' searching intentions with semantic correspondences to the users' query sentences. To tackle such a challenging task, we propose a novel graph convolved video thumbnail pointer (GTP). Specifically, GTP leverages a sentence specified video graph convolutional network to model both the sentence-video semantic interaction and the internal video relationships incorporated with the sentence information, based on which a temporal conditioned pointer network is then introduced to sequentially generate the sentence specified video thumbnails. Moreover, we annotate a new dataset based on ActivityNet Captions for the proposed new task, which consists of 10,000+ video-sentence pairs with each accompanied by an annotated sentence specified video thumbnail. We demonstrate that our proposed GTP outperforms several baseline methods on the created dataset, and thus believe that our initial results along with the release of the new dataset will inspire further research on sentence specified dynamic video thumbnail generation. Dataset and code are available at https://github.com/yytzsy/GTP Yitian Yuan, Lin Ma 0002, Wenwu Zhu 0001 |
ACM Multimedia | 2 |
| 2019 | Exploiting Local and Global Structure for Point Cloud Semantic Segmentation with Contextual Point RepresentationsabstractIn this paper, we propose one novel model for point cloud semantic segmentation,which exploits both the local and global structures within the point cloud based onthe contextual point representations. Specifically, we enrich each point represen-tation by performing one novel gated fusion on the point itself and its contextualpoints. Afterwards, based on the enriched representation, we propose one novelgraph pointnet module, relying on the graph attention block to dynamically com-pose and update each point representation within the local point cloud structure.Finally, we resort to the spatial-wise and channel-wise attention strategies to exploitthe point cloud global structure and thereby yield the resulting semantic label foreach point. Extensive results on the public point cloud databases, namely theS3DIS and ScanNet datasets, demonstrate the effectiveness of our proposed model,outperforming the state-of-the-art approaches. Our code for this paper is available at https://github.com/fly519/ELGS. Xu Wang 0006, Jingming He, Lin Ma 0002 |
NeurIPS | 3 |
| 2019 | Semantic Conditioned Dynamic Modulation for Temporal Sentence Grounding in VideosabstractTemporal sentence grounding in videos aims to detect and localize one target video segment, which semantically corresponds to a given sentence. Existing methods mainly tackle this task via matching and aligning semantics between a sentence and candidate video segments, while neglect the fact that the sentence information plays an important role in temporally correlating and composing the described contents in videos. In this paper, we propose a novel semantic conditioned dynamic modulation (SCDM) mechanism, which relies on the sentence semantics to modulate the temporal convolution operations for better correlating and composing the sentence related video contents over time. More importantly, the proposed SCDM performs dynamically with respect to the diverse video contents so as to establish a more precise matching relationship between sentence and video, thereby improving the temporal grounding accuracy. Extensive experiments on three public datasets demonstrate that our proposed model outperforms the state-of-the-arts with clear margins, illustrating the ability of SCDM to better associate and localize relevant video contents for temporal sentence grounding. Our code for this paper is available at https://github.com/yytzsy/SCDM. Yitian Yuan, Lin Ma 0002, Jingwen Wang 0003, Wei Liu 0005, Wenwu Zhu 0001 |
NeurIPS | 2 |
| 2019 | Bidirectional image-sentence retrieval by local and global deep matching
Lin Ma 0002, Zequn Jie, Xu Wang 0006 |
Neurocomputing | 1 |
| 2019 | Reversible data hiding for high dynamic range images using edge information
Xuanyu He, Wei Zhang 0021, Lin Ma 0002, Yibin Li 0001 |
Multim. Tools Appl. | 4 |
| 2019 | Toward Efficient Action Recognition: Principal Backpropagation for Training Two-Stream NetworksabstractIn this paper, we propose the novel principal backpropagation networks (PBNets) to revisit the backpropagation algorithms commonly used in training two-stream networks for video action recognition. We content that existing approaches always take all the frames/snippets for the backpropagation not optimal for video recognition since the desired actions only occur in a short period within a video. To remedy these drawbacks, we design a watch-and-choose mechanism. In particular, the watching stage exploits a dense snippet-wise temporal pooling strategy to discover the global characteristic for each input video, while the choosing phase only backpropagates a small number of representative snippets that are selected with two novel strategies, i.e., Max-rule and KL-rule. We prove that with the proposed selection strategies, performing the backpropagation on the selected subset is capable of decreasing the loss of the whole snippets as well. The proposed PBNets are evaluated on two standard video action recognition benchmarks UCF101 and HMDB51, where it surpasses the state of the arts consistently, but requiring less memory and computation to achieve high performance. Wenbing Huang 0001, Lijie Fan, Mehrtash Harandi, Lin Ma 0002, Huaping Liu 0001, Wei Liu 0005, Chuang Gan 0001 |
IEEE Trans. Image Process. | 4 |
| 2019 | Low-Light Image Enhancement via a Deep Hybrid NetworkabstractCamera sensors often fail to capture clear images or videos in a poorly lit environment. In this paper, we propose a trainable hybrid network to enhance the visibility of such degraded images. The proposed network consists of two distinct streams to simultaneously learn the global content and the salient structures of the clear image in a unified network. More specifically, the content stream estimates the global content of the low-light input through an encoder-decoder network. However, the encoder in the content stream tends to lose some structure details. To remedy this, we propose a novel spatially variant recurrent neural network (RNN) as an edge stream to model edge details, with the guidance of another auto-encoder. The experimental results show that the proposed network favorably performs against the state-of-the-art low-light image enhancement algorithms. Wenqi Ren, Sifei Liu, Lin Ma 0002, Qianqian Xu 0001, Xiangyu Xu 0002, Xiaochun Cao, Junping Du 0001, Ming-Hsuan Yang 0001 |
IEEE Trans. Image Process. | 3 |
| 2019 | Deep Video Dehazing With Semantic SegmentationabstractRecent research have shown the potential of using convolutional neural networks (CNNs) to accomplish single image dehazing. In this work, we take one step further to explore the possibility of exploiting a network to perform haze removal for videos. Unlike single image dehazing, video based approaches can take advantage of the abundant information that exists across neighboring frames. In this work, assuming that a scene point yields highly correlated transmission values between adjacent video frames, we develop a deep learning solution for video dehazing, where a CNN is trained end-to-end to learn how to accumulate information across frames for transmission estimation. The estimated transmission map is subsequently used to recover a haze-free frame via atmospheric scattering model. In addition, as the semantic information of a scene provides a strong prior for image restoration, we propose to incorporate global semantic priors as input to regularize the transmission maps so that the estimated maps can be smooth in the regions of the same object and only discontinuous across the boundaries of different objects. To train this network, we generate a dataset consisted of synthetic hazy and haze-free videos for supervision based on the NYU depth dataset. We show that the features learned from this dataset are capable of removing haze that arises in outdoor scenes in a wide range of videos. Extensive experiments demonstrate that the proposed algorithm performs favorably against the state-of-the-art methods on both synthetic and real-world videos. Wenqi Ren, Jingang Zhang, Xiangyu Xu 0002, Lin Ma 0002, Xiaochun Cao, Gaofeng Meng, Wei Liu 0005 |
IEEE Trans. Image Process. | 4 |
| 2019 | Adversarial Spatio-Temporal Learning for Video DeblurringabstractCamera shake or target movement often leads to undesired blur effects in videos captured by a hand-held camera. Despite significant efforts having been devoted to video-deblur research, two major challenges remain: 1) how to model the spatio-temporal characteristics across both the spatial domain (i.e., image plane) and the temporal domain (i.e., neighboring frames) and 2) how to restore sharp image details with respect to the conventionally adopted metric of pixel-wise errors. In this paper, to address the first challenge, we propose a deblurring network (DBLRNet) for spatial-temporal learning by applying a 3D convolution to both the spatial and temporal domains. Our DBLRNet is able to capture jointly spatial and temporal information encoded in neighboring frames, which directly contributes to the improved video deblur performance. To tackle the second challenge, we leverage the developed DBLRNet as a generator in the generative adversarial network (GAN) architecture and employ a content loss in addition to an adversarial loss for efficient adversarial training. The developed network, which we name as deblurring GAN, is tested on two standard benchmarks and achieves the state-of-the-art performance. Kaihao Zhang, Wenhan Luo, Yiran Zhong, Lin Ma 0002, Wei Liu 0005, Hongdong Li |
IEEE Trans. Image Process. | 4 |
| 2018 | Learning to Guide Decoding for Image CaptioningabstractRecently, much advance has been made in image captioning, and an encoder-decoder framework has achieved outstanding performance for this task. In this paper, we propose an extension of the encoder-decoder framework by adding a component called guiding network. The guiding network models the attribute properties of input images, and its output is leveraged to compose the input of the decoder at each time step. The guiding network can be plugged into the current encoder-decoder framework and trained in an end-to-end manner. Hence, the guiding vector can be adaptively learned according to the signal from the decoder, making itself to embed information from both image and language. Additionally, discriminative supervision can be employed to further improve the quality of guidance. The advantages of our proposed approach are verified by experiments carried out on the MS COCO dataset. Lin Ma 0002, Hanwang Zhang, Wei Liu 0005 |
AAAI | 2 |
| 2018 | Regularizing RNNs for Caption Generation by Reconstructing the Past With the PresentabstractRecently, caption generation with an encoder-decoder framework has been extensively studied and applied in different domains, such as image captioning, code captioning, and so on. In this paper, we propose a novel architecture, namely Auto-Reconstructor Network (ARNet), which, coupling with the conventional encoder-decoder framework, works in an end-to-end fashion to generate captions. ARNet aims at reconstructing the previous hidden state with the present one, besides behaving as the input-dependent transition operator. Therefore, ARNet encourages the current hidden state to embed more information from the previous one, which can help regularize the transition dynamics of recurrent neural networks (RNNs). Extensive experimental results show that our proposed ARNet boosts the performance over the existing encoder-decoder models on both image captioning and source code captioning tasks. Additionally, ARNet remarkably reduces the discrepancy between training and inference processes for caption generation. Furthermore, the performance on permuted sequential MNIST demonstrates that ARNet can effectively regularize RNN, especially on modeling long-term dependencies. Our code is available at: https://github.com/chenxinpeng/ARNet. Lin Ma 0002, Jian Yao 0002, Wei Liu 0005 |
CVPR | 2 |
| 2018 | Gated Fusion Network for Single Image DehazingabstractIn this paper, we propose an efficient algorithm to directly restore a clear image from a hazy input. The proposed algorithm hinges on an end-to-end trainable neural network that consists of an encoder and a decoder. The encoder is exploited to capture the context of the derived input images, while the decoder is employed to estimate the contribution of each input to the final dehazed result using the learned representations attributed to the encoder. The constructed network adopts a novel fusion-based strategy which derives three inputs from an original hazy image by applying White Balance (WB), Contrast Enhancing (CE), and Gamma Correction (GC). We compute pixel-wise confidence maps based on the appearance differences between these different inputs to blend the information of the derived inputs and preserve the regions with pleasant visibility. The final dehazed image is yielded by gating the important features of the derived inputs. To train the network, we introduce a multi-scale approach such that the halo artifacts can be avoided. Extensive experimental results on both synthetic and real-world images demonstrate that the proposed algorithm performs favorably against the state-of-the-art algorithms. Wenqi Ren, Lin Ma 0002, Jiawei Zhang 0002, Jinshan Pan, Xiaochun Cao, Wei Liu 0005, Ming-Hsuan Yang 0001 |
CVPR | 2 |
| 2018 | Reconstruction Network for Video CaptioningabstractIn this paper, the problem of describing visual contents of a video sequence with natural language is addressed. Unlike previous video captioning work mainly exploiting the cues of video contents to make a language description, we propose a reconstruction network (RecNet) with a novel encoder-decoder-reconstructor architecture, which leverages both the forward (video to sentence) and backward (sentence to video) flows for video captioning. Specifically, the encoder-decoder makes use of the forward flow to produce the sentence description based on the encoded video semantic features. Two types of reconstructors are customized to employ the backward flow and reproduce the video features based on the hidden state sequence generated by the decoder. The generation loss yielded by the encoder-decoder and the reconstruction loss introduced by the reconstructor are jointly drawn into training the proposed RecNet in an end-to-end fashion. Experimental results on benchmark datasets demonstrate that the proposed reconstructor can boost the encoder-decoder models and leads to significant gains in video caption accuracy. Bairui Wang, Lin Ma 0002, Wei Zhang 0021, Wei Liu 0005 |
CVPR | 2 |
| 2018 | Bidirectional Attentive Fusion With Context Gating for Dense Video CaptioningabstractDense video captioning is a newly emerging task that aims at both localizing and describing all events in a video. We identify and tackle two challenges on this task, namely, (1) how to utilize both past and future contexts for accurate event proposal predictions, and (2) how to construct informative input to the decoder for generating natural event descriptions. First, previous works predominantly generate temporal event proposals in the forward direction, which neglects future video context. We propose a bidirectional proposal method that effectively exploits both past and future contexts to make proposal predictions. Second, different events ending at (nearly) the same time are indistinguishable in the previous works, resulting in the same captions. We solve this problem by representing each event with an attentive fusion of hidden states from the proposal module and video contents (e.g., C3D features). We further propose a novel context gating mechanism to balance the contributions from the current event and its surrounding contexts dynamically. We empirically show that our attentively fused event representation is superior to the proposal hidden states or video contents alone. By coupling proposal and captioning modules into one unified framework, our model outperforms the state-of-the-arts on the ActivityNet Captions dataset with a relative gain of over 100% (Meteor score increases from 4.82 to 9.65). Jingwen Wang 0003, Lin Ma 0002, Wei Liu 0005, Yong Xu 0001 |
CVPR | 3 |
| 2018 | Learning to Generate Time-Lapse Videos Using Multi-Stage Dynamic Generative Adversarial NetworksabstractTaking a photo outside, can we predict the immediate future, e.g., how would the cloud move in the sky? We address this problem by presenting a generative adversarial network (GAN) based two-stage approach to generating realistic time-lapse videos of high resolution. Given the first frame, our model learns to generate long-term future frames. The first stage generates videos of realistic contents for each frame. The second stage refines the generated video from the first stage by enforcing it to be closer to real videos with regard to motion dynamics. To further encourage vivid motion in the final generated video, Gram matrix is employed to model the motion more precisely. We build a large scale time-lapse dataset, and test our approach on this new dataset. Using our model, we are able to generate realistic videos of up to 128 Ã- 128 resolution for 32 frames. Quantitative and qualitative experiment results demonstrate the superiority of our model over the state-of-the-art models. Wei Xiong 0008, Wenhan Luo, Lin Ma 0002, Wei Liu 0005, Jiebo Luo 0001 |
CVPR | 3 |
| 2018 | Video Re-localization
Yang Feng 0001, Lin Ma 0002, Wei Liu 0005, Tong Zhang 0001, Jiebo Luo 0001 |
ECCV (14) | 2 |
| 2018 | Neural Stereoscopic Image Style Transfer
Xinyu Gong, Hao-Zhi Huang 0001, Lin Ma 0002, Fumin Shen, Wei Liu 0005, Tong Zhang 0001 |
ECCV (5) | 3 |
| 2018 | Recurrent Fusion Network for Image Captioning
Lin Ma 0002, Yu-Gang Jiang 0001, Wei Liu 0005, Tong Zhang 0001 |
ECCV (2) | 2 |
| 2018 | Unsupervised Image-to-Image Translation with Stacked Cycle-Consistent Adversarial Networks
Minjun Li, Hao-Zhi Huang 0001, Lin Ma 0002, Wei Liu 0005, Tong Zhang 0001, Yu-Gang Jiang 0001 |
ECCV (9) | 3 |
| 2018 | Temporally Grounding Natural Sentence in VideoabstractWe introduce an effective and efficient method that grounds (i.e., localizes) natural sentences in long, untrimmed video sequences.Specifically, a novel Temporal GroundNet (TGN) 1 is proposed to temporally capture the evolving fine-grained frame-by-word interactions between video and sentence.TGN sequentially scores a set of temporal candidates ended at each frame based on the exploited frameby-word interactions, and finally grounds the segment corresponding to the sentence.Unlike traditional methods treating the overlapping segments separately in a sliding window fashion, TGN aggregates the historical information and generates the final grounding result in one single pass.We extensively evaluate our proposed TGN on three public datasets with significant improvements over the stateof-the-arts.We further show the consistent effectiveness and efficiency of TGN through an ablation study and a runtime test.* Work done while Jingyuan Chen and Xinpeng Chen were Research Interns with Tencent AI Lab. Jingyuan Chen 0003, Lin Ma 0002, Zequn Jie, Tat-Seng Chua |
EMNLP | 3 |
| 2018 | Safe Element Screening for Submodular Function MinimizationabstractSubmodular functions are discrete analogs of convex functions, which have applications in various fields, including machine learning and computer vision. However, in large-scale applications, solving Submodular Function Minimization (SFM) problems remains challenging. In this paper, we make the first attempt to extend the emerging technique named screening in large-scale sparse learning to SFM for accelerating its optimization process. We first conduct a careful studying of the relationships between SFM and the corresponding convex proximal problems, as well as the accurate primal optimum estimation of the proximal problems. Relying on this study, we subsequently propose a novel safe screening method to quickly identify the elements guaranteed to be included (we refer to them as active) or excluded (inactive) in the final optimal solution of SFM during the optimization process. By removing the inactive elements and fixing the active ones, the problem size can be dramatically reduced, leading to great savings in the computational cost without sacrificing any accuracy. To the best of our knowledge, the proposed method is the first screening method in the fields of SFM and even combinatorial optimization, thus pointing out a new direction for accelerating SFM algorithms. Experiment results on both synthetic and real datasets demonstrate the significant speedups gained by our approach. Lin Ma 0002, Wei Liu 0005, Tong Zhang 0001 |
ICML | 3 |
| 2018 | Image-level to Pixel-wise Labeling: From Theory to PracticeabstractConventional convolutional neural networks (CNNs) have achieved great success in image semantic segmentation. Existing methods mainly focus on learning pixel-wise labels from an image directly. In this paper, we advocate tackling the pixel-wise segmentation problem by considering the image-level classification labels. Theoretically, we analyze and discuss the effects of image-level labels on pixel-wise segmentation from the perspective of information theory. In practice, an end-to-end segmentation model is built by fusing the image-level and pixel-wise labeling networks. A generative network is included to reconstruct the input image and further boost the segmentation model training with an auxiliary loss. Extensive experimental results on benchmark dataset demonstrate the effectiveness of the proposed method, where good image-level labels can significantly improve the pixel-wise segmentation accuracy. Tiezhu Sun, Wei Zhang 0021, Zhijie Wang 0010, Lin Ma 0002, Zequn Jie |
IJCAI | 4 |
| 2018 | Long-Term Human Motion Prediction by Modeling Motion Context and Enhancing Motion DynamicsabstractHuman motion prediction aims at generating future frames of human motion based on an observed sequence of skeletons. Recent methods employ the latest hidden states of a recurrent neural network (RNN) to encode the historical skeletons, which can only address short-term prediction. In this work, we propose a motion context modeling by summarizing the historical human motion with respect to the current prediction. A modified highway unit (MHU) is proposed for efficiently eliminating motionless joints and estimating next pose given the motion context. Furthermore, we enhance the motion dynamic by minimizing the gram matrix loss for long-term motion prediction. Experimental results show that the proposed model can promisingly forecast the human future movements, which yields superior performances over related state-of-the-art approaches. Moreover, specifying the motion context with the activity labels enables our model to perform human motion transfer. Yongyi Tang, Lin Ma 0002, Wei Liu 0005, Wei-Shi Zheng 0001 |
IJCAI | 2 |
| 2018 | Deep Non-Blind Deconvolution via Generalized Low-Rank ApproximationabstractIn this paper, we present a deep convolutional neural network to capture the inherent properties of image degradation, which can handle different kernels and saturated pixels in a unified framework. The proposed neural network is motivated by the low-rank property of pseudo-inverse kernels. We first compute a generalized low-rank approximation for a large number of blur kernels, and then use separable filters to initialize the convolutional parameters in the network. Our analysis shows that the estimated decomposed matrices contain the most essential information of the input kernel, which ensures the proposed network to handle various blurs in a unified framework and generate high-quality deblurring results. Experimental results on benchmark datasets with noise and saturated pixels demonstrate that the proposed algorithm performs favorably against state-of-the-art methods. Wenqi Ren, Jiawei Zhang 0002, Lin Ma 0002, Jinshan Pan, Xiaochun Cao, Wangmeng Zuo, Wei Liu 0005, Ming-Hsuan Yang 0001 |
NeurIPS | 3 |
| 2018 | Parsimonious Quantile Regression of Financial Asset Tail Dynamics via Sequential LearningabstractWe propose a parsimonious quantile regression framework to learn the dynamic tail behaviors of financial asset returns. Our model captures well both the time-varying characteristic and the asymmetrical heavy-tail property of financial time series. It combines the merits of a popular sequential neural network model, i.e., LSTM, with a novel parametric quantile function that we construct to represent the conditional distribution of asset returns. Our model also captures individually the serial dependences of higher moments, rather than just the volatility. Across a wide range of asset classes, the out-of-sample forecasts of conditional quantiles or VaR of our model outperform the GARCH family. Further, the proposed approach does not suffer from the issue of quantile crossing, nor does it expose to the ill-posedness comparing to the parametric probability density function approach. Lin Ma 0002, Wei Liu 0005, Qi Wu 0009 |
NeurIPS | 3 |
| 2018 | Deep intensity guidance based compression artifacts reduction for depth map
Xu Wang 0006, Yun Zhang 0002, Lin Ma 0002, Sam Kwong, Jianmin Jiang |
J. Vis. Commun. Image Represent. | 4 |
| 2018 | Image processing for synthesis imaging of mingantu spectral radioheliograph (MUSER)
Long Xu 0001, Yihua Yan, Lin Ma 0002, Yun Zhang 0002 |
Multim. Tools Appl. | 3 |
| 2018 | Quaternion representation based visual saliency for stereoscopic image quality assessment
Xu Wang 0006, Lin Ma 0002, Sam Kwong, Yu Zhou 0027 |
Signal Process. | 2 |
| 2018 | Screen Content Image Quality Assessment Using Multi-Scale Difference of GaussianabstractIn this paper, a novel image quality assessment (IQA) model for the screen content images (SCIs) is proposed by using multi-scale difference of Gaussian (MDOG). Motivated by the observation that the human visual system (HVS) is sensitive to the edges while the image details can be better explored in different scales, the proposed model exploits MDOG to effectively characterize the edge information of the reference and distorted SCIs at two different scales, respectively. Then, the degree of edge similarity is measured in terms of the smaller-scale edge map. Finally, the edge strength computed based on the larger-scale edge map is used as the weighting factor to generate the final SCI quality score. Experimental results have shown that the proposed IQA model for the SCIs produces high consistency with human perception of the SCI quality and outperforms the state-of-the-art quality models. Ying Fu 0004, Huanqiang Zeng, Lin Ma 0002, Zhangkai Ni, Jianqing Zhu, Kai-Kuang Ma |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2018 | A Gabor Feature-Based Quality Assessment Model for the Screen Content ImagesabstractIn this paper, an accurate and efficient full-reference image quality assessment (IQA) model using the extracted Gabor features, called Gabor feature-based model (GFM), is proposed for conducting objective evaluation of screen content images (SCIs). It is well-known that the Gabor filters are highly consistent with the response of the human visual system (HVS), and the HVS is highly sensitive to the edge information. Based on these facts, the imaginary part of the Gabor filter that has odd symmetry and yields edge detection is exploited to the luminance of the reference and distorted SCI for extracting their Gabor features, respectively. The local similarities of the extracted Gabor features and two chrominance components, recorded in the LMN color space, are then measured independently. Finally, the Gabor-feature pooling strategy is employed to combine these measurements and generate the final evaluation score. Experimental simulation results obtained from two large SCI databases have shown that the proposed GFM model not only yields a higher consistency with the human perception on the assessment of SCIs but also requires a lower computational complexity, compared with that of classical and state-of-the-art IQA models. The source code for the proposed GFM will be available at http://smartviplab.org/pubilcations/GFM.html. Zhangkai Ni, Huanqiang Zeng, Lin Ma 0002, Junhui Hou, Jing Chen 0001, Kai-Kuang Ma |
IEEE Trans. Image Process. | 3 |
| 2017 | Real-Time Neural Style Transfer for VideosabstractRecent research endeavors have shown the potential of using feed-forward convolutional neural networks to accomplish fast style transfer for images. In this work, we take one step further to explore the possibility of exploiting a feed-forward network to perform style transfer for videos and simultaneously maintain temporal consistency among stylized video frames. Our feed-forward network is trained by enforcing the outputs of consecutive frames to be both well stylized and temporally consistent. More specifically, a hybrid loss is proposed to capitalize on the content information of input frames, the style information of a given style image, and the temporal information of consecutive frames. To calculate the temporal loss during the training stage, a novel two-frame synergic training mechanism is proposed. Compared with directly applying an existing image style transfer method to videos, our proposed method employs the trained network to yield temporally consistent stylized videos which are much more visually pleasant. In contrast to the prior video style transfer method which relies on time-consuming optimization on the fly, our method runs in real time while generating competitive visual results. Hao-Zhi Huang 0001, Hao Wang 0050, Wenhan Luo, Lin Ma 0002, Zhifeng Li 0001, Wei Liu 0005 |
CVPR | 4 |
| 2017 | Multimodal deep learning for solar radio burst classification
Lin Ma 0002, Zhuo Chen 0006, Long Xu 0001, Yihua Yan |
Pattern Recognit. | 1 |
| 2017 | Multi-Task Rank Learning for Image Quality AssessmentabstractIn practice, images are distorted by more than one distortion. For image quality assessment (IQA), existing machine learning (ML)-based methods generally establish a unified model for all the distortion types, or each model is trained independently for each distortion type, which is therefore distortion aware. In distortion-aware methods, the common features among different distortions are not exploited. In addition, there are fewer training samples for each model training task, which may result in overfitting. To address these problems, we propose a multi-task learning framework to train multiple IQA models together, where each model is for each distortion type; however, all the training samples are associated with each model training task. Thus, the common features among different distortion types and the said underlying relatedness among all the learning tasks are exploited, which would benefit the generalization ability of trained models and prevent overfitting possibly. In addition, pairwise image quality ranking instead of image quality rating is optimized in our learning task, which is fundamentally departed from traditional ML-based IQA methods toward better performance. The experimental results confirm that the proposed multi-task rank-learning-based IQA metric is prominent against all state-of-the-art nonreference IQA approaches. Long Xu 0001, Jia Li 0003, Weisi Lin, Yongbing Zhang 0002, Lin Ma 0002, Yuming Fang 0001, Yihua Yan |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2017 | ESIM: Edge Similarity for Screen Content Image Quality AssessmentabstractIn this paper, an accurate full-reference image quality assessment (IQA) model developed for assessing screen content images (SCIs), called the edge similarity (ESIM), is proposed. It is inspired by the fact that the human visual system (HVS) is highly sensitive to edges that are often encountered in SCIs; therefore, essential edge features are extracted and exploited for conducting IQA for the SCIs. The key novelty of the proposed ESIM lies in the extraction and use of three salient edge features-i.e., edge contrast, edge width, and edge direction. The first two attributes are simultaneously generated from the input SCI based on a parametric edge model, while the last one is derived directly from the input SCI. The extraction of these three features will be performed for the reference SCI and the distorted SCI, individually. The degree of similarity measured for each above-mentioned edge attribute is then computed independently, followed by combining them together using our proposed edge-width pooling strategy to generate the final ESIM score. To conduct the performance evaluation of our proposed ESIM model, a new and the largest SCI database (denoted as SCID) is established in our work and made to the public for download. Our database contains 1800 distorted SCIs that are generated from 40 reference SCIs. For each SCI, nine distortion types are investigated, and five degradation levels are produced for each distortion type. Extensive simulation results have clearly shown that the proposed ESIM model is more consistent with the perception of the HVS on the evaluation of distorted SCIs than the multiple state-of-the-art IQA methods. Zhangkai Ni, Lin Ma 0002, Huanqiang Zeng, Jing Chen 0001, Canhui Cai, Kai-Kuang Ma |
IEEE Trans. Image Process. | 2 |
| 2017 | Objective Quality Assessment of Image Retargeting by Incorporating Fidelity Measures and Inconsistency DetectionabstractThe tremendous growth in mobile devices has resulted in huge generation and usage of digital images. Image quality assessment is thus an important issue for mobile media applications. In this paper, we focus on the quality evaluation of images generated by content-aware image retargeting, in which the reference and the distorted images are of different sizes. Through retargeting, many types of deformation inconsistency lead to shape distortion, deformation artifacts, and content information loss, worsening its perceptual quality. The deformation inconsistency occurs on different levels of the retargeted images. Limited by the accuracy of the alignment between the original and retargeted images, previous methods only focus on pixel-level and patch-level fidelity analyses and fail to detect deformation inconsistency. In this paper, we improve the alignment algorithm and propose a three-level representation of the retargeting process. Based on the analysis of this three-level representation, both fidelity measures and inconsistency detection are combined to determine the final retargeting quality. The proposed algorithm is validated on the public data sets RetargetMe and CUHK. Experimental results demonstrate that inconsistency detection contributes to accurately assessing the image retargeting perceptual quality. This inspires us to investigate more about deformation inconsistency to formulate the objective quality of image retargeting. Yichi Zhang 0014, King Ngi Ngan, Lin Ma 0002, Hongliang Li 0001 |
IEEE Trans. Image Process. | 3 |
| 2016 | Screen content image quality assessment using edge modelabstractSince the human visual system (HVS) is highly sensitive to edges, a novel image quality assessment (IQA) metric for assessing screen content images (SCIs) is proposed in this paper. The turnkey novelty lies in the use of an existing parametric edge model to extract two types of salient attributes - namely, edge contrast and edge width, for the distorted SCI under assessment and its original SCI, respectively. The extracted information is subject to conduct similarity measurements on each attribute, independently. The obtained similarity scores are then combined using our proposed edge-width pooling strategy to generate the final IQA score. Hopefully, this score is consistent with the judgment made by the HVS. Experimental results have shown that the proposed IQA metric produces higher consistency with that of the HVS on the evaluation of the image quality of the distorted SCI than that of other state-of-the-art IQA metrics. Zhangkai Ni, Lin Ma 0002, Huanqiang Zeng, Canhui Cai, Kai-Kuang Ma |
ICIP | 2 |
| 2016 | Perceptual image quality enhancement for solar radio imageabstractIn solar radio observation, the visualization of data is very important since it can more intuitively and clearly deliver interest information of solar radio activities to astronomers. As to visualization, we highly expect good visual quality of images/videos in favor of the discovery of solar radio events recorded by observation data. The existing imaging system cannot guarantee good visual quality of solar radio data visualization. In this paper, an image quality enhancement algorithm is developed to improve solar radio extreme ultraviolet (EUV) images from Solar Dynamics Observatory (SDO). Firstly, the guided filter is employed to smooth image, which outputs an image with good skeleton and edges. Since the fine structures of solar radio activities are embedded in high frequency components of a solar radio image, we propose a novel structure preserving filtering to amplify the different signal of original input image subtracting smoothed one. Afterwards, fusing the amplified details and smoothed one together, the final enhanced image is generated. The experimental results prove that the image quality is significantly improved by using the proposed image quality enhancement algorithm. Long Xu 0001, Lin Ma 0002, Zhuo Chen 0006, Xianyou Zeng, Yihua Yan |
QoMEX | 2 |
| 2016 | Complex singular value decomposition based stereoscopic image quality assessmentabstractDesigning a reliable and generic perceptual quality metric is a challenging issue in three-dimensional (3D) visual signal processing. Due to the limited knowledge on 3D perceptual, it is difficult to fuse the visual information of left and right views in an effective way. In this paper, we propose a complex singular value decomposition (CSVD) based stereoscopic image quality assessment (SIQA) metric. First, the corresponding blocks of the left/right view are grouped into complex representation (CR) block through the scale-invariant feature transform (SIFT) view matching process. Then we compute the CSVD coefficients of each CR block. Final, a CSVD based quality pooling stage is employed to predict the final visual quality of the distorted 3D image. Experimental results demonstrate that the proposed metric has good consistency with 3D perception of human. Xu Wang 0006, Lin Ma 0002, Yu Zhou 0027, Sam Kwong |
VCIP | 3 |
| 2016 | Content adaptive directional transform for high efficiency video codingabstractHEVC is an emerging new standard for digital video compression, which is regarded as a successor to H.264/AVC standard. It still belongs to block-based hybrid video coding framework. The block patterns range from 4×4 to 64×64 blocks, and DCT is extended from 4×4 to 32×32. 2D-DCT for image is performed along the vertical and horizontal directions, so it is good at the energy compaction of residual block with vertical or horizontal edges. However, the edges are usually neither horizontal nor vertical for most cases, such as neither vertical nor horizontal intra prediction, so directional transform was explored in the past several years. In this paper, a directional transform adaptive to image content is proposed. Firstly, the prediction residual blocks are collected from coding a number of video sequences with plenty of image content. Secondly, for each kind of image content, the residual blocks are clustered to form the given number of clusters. Thirdly, each cluster contributes a transform basis after Singular Value Decomposition (SVD). The experimental results in terms of PSNR gains demonstrate the efficiency of the proposed algorithm with the comparison with the standard HM software. Long Xu 0001, Lin Ma 0002, Yun Zhang 0002, Yihua Yan |
VCIP | 2 |
| 2016 | Interest points based collaborative trackingabstractIn this paper, we propose a robust collaborative tracking algorithm based on interest points detection and template matching in sparse representation framework. In the proposed tracker, the target dictionary and the candidate dictionary are constructed with the patches around interest points of the previous frame and the current frame, respectively. The correspondence between target points and candidate points is computed by solving an Li minimization problem. Only the mutually matched target and candidate point pairs are selected, and the displacements of them are measured to generate the candidate targets in current frame. To find the best one among all candidate targets, each of them is sparsely represented by target templates given as benchmarks in initial frame. The candidate target with the smallest projection error produces the final tracking result. The experimental results show that the proposed tracker is superior to the state-of-the-art methods remarkably with respect to the tracking accuracy. Xianyou Zeng, Long Xu 0001, Lin Ma 0002, Ruizhen Zhao |
VCIP | 3 |
| 2016 | Reorganized DCT-based image representation for reduced reference stereoscopic image quality assessment
Lin Ma 0002, Xu Wang 0006, Qiong Liu 0001, King Ngi Ngan |
Neurocomputing | 1 |
| 2016 | Imaging and representation learning of solar radio spectrums for classification
Zhuo Chen 0006, Lin Ma 0002, Long Xu 0001, Chengming Tan, Yihua Yan |
Multim. Tools Appl. | 2 |
| 2016 | Learning structure of stereoscopic image for no-reference quality assessment with convolutional neural network
Wei Zhang 0021, Chenfei Qu, Lin Ma 0002, Jingwei Guan, Rui Huang 0001 |
Pattern Recognit. | 3 |
| 2016 | Gradient Direction for Screen Content Image Quality AssessmentabstractIn this letter, we make the first attempt to explore the usage of the gradient direction to conduct the perceptual quality assessment of the screen content images (SCIs). Specifically, the proposed approach first extracts the gradient direction based on the local information of the image gradient magnitude, which not only preserves gradient direction consistency in local regions, but also demonstrates sensitivities to the distortions introduced to the SCI. A deviation-based pooling strategy is subsequently utilized to generate the corresponding image quality index. Moreover, we investigate and demonstrate the complementary behaviors of the gradient direction and magnitude for SCI quality assessment. By jointly considering them together, our proposed SCI quality metric outperforms the state-of-the-art quality metrics in terms of correlation with human visual system perception. Zhangkai Ni, Lin Ma 0002, Huanqiang Zeng, Canhui Cai, Kai-Kuang Ma |
IEEE Signal Process. Lett. | 2 |
| 2016 | Just Noticeable Difference Estimation for Screen Content ImagesabstractWe propose a novel just noticeable difference (JND) model for a screen content image (SCI). The distinct properties of the SCI result in different behaviors of the human visual system when viewing the textual content, which motivate us to employ a local parametric edge model with an adaptive representation of the edge profile in JND modeling. In particular, we decompose each edge profile into its luminance, contrast, and structure, and then evaluate the visibility threshold in different ways. The edge luminance adaptation, contrast masking, and structural distortion sensitivity are studied in subjective experiments, and the final JND model is established based on the edge profile reconstruction with tolerable variations. Extensive experiments are conducted to verify the proposed JND model, which confirm that it is accurate in predicting the JND profile, and outperforms the state-of-the-art schemes in terms of the distortion masking ability. Furthermore, we explore the applicability of the proposed JND model in the scenario of perceptually lossless SCI compression, and experimental results show that the proposed scheme can outperform the conventional JND guided compression schemes by providing better visual quality at the same coding bits. Shiqi Wang 0001, Lin Ma 0002, Yuming Fang 0001, Weisi Lin, Siwei Ma 0001, Wen Gao 0001 |
IEEE Trans. Image Process. | 2 |
| 2016 | No-Reference Retargeted Image Quality Assessment Based on Pairwise Rank LearningabstractIn this paper, we propose a novel no-reference image quality assessment method for the retargeted image based on the pairwise rank learning approach. Each retargeted image needs to be first represented as a feature vector, which not only captures the image characteristics but also is sensitive to distortions during the retargeting process. As such, we investigate and examine different image representations for their abilities depicting the perceptual quality of retargeted image. Based on the image representations, we resort to the pairwise rank learning approach to discriminate the perceptual quality between the retargeted image pairs. Experimental results demonstrate that the proposed method can effectively depict the perceptual quality of the retargeted image, which can even perform comparably with the full-reference quality assessment methods. Lin Ma 0002, Long Xu 0001, Yichi Zhang 0014, Yihua Yan, King Ngi Ngan |
IEEE Trans. Multim. | 1 |
| 2016 | Free-Energy Principle Inspired Video Quality Metric and Its Use in Video CodingabstractIn this paper, we extend the free-energy principle to video quality assessment (VQA) by incorporating with the recent psychophysical study on human visual speed perception (HVSP). A novel video quality metric, namely the free-energy principle inspired video quality metric (FePVQ), is therefore developed and applied to perceptual video coding optimization. The free-energy principle suggests that the human visual system (HVS) can actively predict “orderly” information and avoid “disorderly” information for image perception. Basically, “orderly” is associated with the skeletons and edges of objects, and “disorderly” mostly concerns textures in images. Based on this principle, an image is separated into orderly and disorderly regions, and processed differently in image quality assessment. For videos, visual attention, or fixation, is associated with the objects with significant motion according to HVSP, resulting in a motion strength factor in the FePVQ so that the free-energy principle is extended into spatio-temporal domain for VQA. In addition, we investigate the application of the FePVQ in perceptual rate distortion optimization (RDO). For this purpose, the FePVQ is realized with low computational cost by using the relative total variation model and the block-wise motion vectors of video coding to simulate the free-energy principle and the HVSP, respectively. The experimental results indicate that the proposed FePVQ is highly consistent with the HVS perception. The linear correlation coefficient and Spearman's rank-order correlation coefficient are up to 0.8324 and 0.8281 on the LIVE video database. Better perceptual quality of encoded video sequences is achieved by FePVQ-motivated RDO in video coding. Long Xu 0001, Weisi Lin, Lin Ma 0002, Yongbing Zhang 0002, Yuming Fang 0001, King Ngi Ngan, Songnan Li, Yihua Yan |
IEEE Trans. Multim. | 3 |
| 2015 | Multi-task rank learning for image quality assessmentabstractIn practice, multiple types of distortions are associated with an image quality degradation process. The existing machine learning (ML) based image quality assessment (IQA) approaches generally established a unified model for all distortion types, or each model is trained independently for each distortion type by using single-task learning, which lead to the poor generalization ability of the models as applied to practical image processing. There are often the underlying cross relatedness amongst these single-task learnings in IQA, which is ignored by the previous approaches. To solve this problem, we propose a multi-task learning framework to train IQA models simultaneously across individual tasks each of which concerns one distortion type. These relatedness can be therefore exploited to improve the generalization ability of IQA models from single-task learning. In addition, pairwise image quality rank instead of image quality rating is optimized in learning task. By mapping image quality rank to image quality rating, a novel no-reference (NR) IQA approach can be derived. The experimental results confirm that the proposed Multi-task Rank Learning based IQA (MRLIQ) approach is prominent among all state-of-the-art NR-IQA approaches. Long Xu 0001, Jia Li 0003, Weisi Lin, Yongbing Zhang 0002, Lin Ma 0002, Yuming Fang 0001, Yun Zhang 0002, Yihua Yan |
ICASSP | 5 |
| 2015 | Multimodal Learning for Classification of Solar Radio SpectrumabstractThis paper proposes the first attempt to utilize multi-modal learning method for the representation learning of the solar radio spectrums. The solar radio signals sensed from differ-ent frequency channels, which present different characteristics, are regarded as different modalities. We employ a multimodal neural network to learn the representations of the solar radio spectrum, which can distinguish the differences and learn the interactions between different modalities. The original solar ra-dio spectrums are firstly pre-processed, including normalization, denoising, channel competition and etc., before being fed into the multimodal learning network. Experimental results have demon-strated that the proposed multimodal learning network can learn the representation of the solar radio spectrum more effectively, and improve the classification accuracy. Zhuo Chen 0006, Lin Ma 0002, Long Xu 0001, Ying Weng, Yihua Yan |
SMC | 2 |
| 2015 | Rank Learning Based No-Reference Quality Assessment of Retargeted ImagesabstractIn this paper, we first propose a novel no-reference (NR) image quality assessment (IQA) method for retargeted image based on the rank learning approach. Firstly, image features for each retargeted image are extracted, which should not only represent the image characteristics but also be sensitive to the retargeted distortions. Specifically, the image feature should be able to capture the shape distortions, which are the commonly encountered distortions of the retargeted image. Based on the extracted image features, the rank learning method is employed to train a model to discriminate the perceptual quality of the retargeted image. Experimental results demonstrate that the proposed method can effectively depict the perceptual quality of the retargeted image, which can even perform comparably with the full-reference (FR) quality assessment methods. Lin Ma 0002, Long Xu 0001, Yichi Zhang 0014, King Ngi Ngan, Yihua Yan |
SMC | 1 |
| 2015 | Multimodal learning for facial expression recognition
Wei Zhang 0021, Youmei Zhang, Lin Ma 0002, Jingwei Guan, Shijie Gong |
Pattern Recognit. | 3 |
| 2013 | Reduced reference video quality assessment based on spatial HVS mutual masking and temporal motion estimationabstractIn this paper, an effective reduced reference (RR) video quality assessment (VQA) is proposed by depicting both the spatial and temporal statistical characteristics of the video signals. For each video frame, spatial information change (SIC) is employed to depict the energy variation. A novel mutual masking strategy based on the extracted SIC is proposed to accurately simulate the human visual system (HVS) texture masking property. For adjacent video frames, the temporal relationship is depicted by block-based motion estimation (BME). The generalized Gaussian density (GGD) function is employed to depict the histogram natural statistic of the residual frame after BME. The city-block distance (CBD) is used to measure the distance between histograms of the original and distorted video sequence. By pooling the measurements from both spatial and temporal perspectives, an efficient RR VQA is constructed. With the evaluations on the public video quality database, the proposed RR VQA demonstrated to be more effective than the representative RR VQAs and even the full-reference (FR) VQAs, such as peak signal-to-noise ratio (PSNR) and structure similarity index (SSIM) in matching the subjective ratings. Furthermore, the proposed RR VQA demonstrated to be much more effective and efficient, requiring only a very small number of bits for the RR feature representation. Lin Ma 0002, King Ngi Ngan, Long Xu 0001 |
ICME | 1 |
| 2013 | Overview of quality assessment for visual signals and newly emerged trendsabstractQuality assessment is not only essential on its own for testing, optimizing, benchmarking, monitoring and inspecting related systems and services, but also plays an essential role in the design of virtually all visual signal processing and communication algorithms, as well as various related decision making processes. In the paper, we provide an overview of the quality assessment approaches for traditional visual signals, as well as the newly emerged ones, which covers the subjective quality evaluation and objective quality metrics of scalable and mobile videos, high dynamic range (HDR) images, image segmentation results, 3D images/videos, and retargeted images. Also the challenges for designing effective quality metrics and corresponding applications are discussed. Lin Ma 0002, Chenwei Deng, King Ngi Ngan, Weisi Lin |
ISCAS | 1 |
| 2013 | High quality image construction from multiple low quality copiesabstractIn this paper, the authors proposed to construct a high quality image based on multiple low quality input images. The relationship between one pixel and its neighbourhood should be consistent between different degraded images. Therefore, the reconstruction method is proposed by enforcing the pixel consistency property, which is ensured by estimating the parameters of piecewise image model for each pixel. Subsequently, the reconstructed coefficients are regularized within a reasonable range. Experimental results on multiple images (with different distortions of different levels) have demonstrated that the proposed method can effectively alleviate the noises meanwhile preserve the detailed information. Better quality images in terms of both objective and subjective measurements can be generated. Lin Ma 0002, Long Xu 0001, Qian Zhang 0001, King Ngi Ngan |
MMSP | 1 |
| 2013 | Visual quality metric for perceptual video codingabstractThe visual quality assessment (VQA) becomes prevailing in the studies of image and video coding. It assesses the quality of image or video more accurately than mean square error (MSE) with respect to the human visual system (HVS). Toward perceptual video coding, MSE is weighted spatially and temporally to simulate the HVS response to visual signal in this paper. Firstly, the image content is depicted by edge strength to compose spatial weighting factors. Secondly, the motion strength calculated from motion vector of each block gives temporal weighting factors. Thirdly, the motion trajectory based saliency map for video signal is integrated as another weighting factor of MSE. The proposed VQM not only efficiently model HVS but also relate to quantization parameter (QP) capable of guiding perceptual video coding. A perceptual rate distortion optimization (RDO) is established on the proposed VQM. The experimental results indicate that the proposed VQM is consistent well with HVS. In addition, the better rate-distortion efficiency and accurate bit rate control can be achieved by the proposed visual quality control algorithm. Long Xu 0001, Lin Ma 0002, King Ngi Ngan, Weisi Lin, Ying Weng |
VCIP | 2 |
| 2013 | Anaglyph image generation by matching color appearance attributes
Songnan Li, Lin Ma 0002, King Ngi Ngan |
Signal Process. Image Commun. | 2 |
| 2013 | Reduced-reference image quality assessment in reorganized DCT domain
Lin Ma 0002, Songnan Li, King Ngi Ngan |
Signal Process. Image Commun. | 1 |
| 2013 | Consistent Visual Quality Control in Video CodingabstractVisual quality consistency is one of the most important issues in video quality assessment. When people view a sequential video, they may have an unpleasant perceptual experience if the video has an inconsistent visual quality even though the average visual quality of the video is not compromised. Thus, consistent visual quality control is mostly expected in general video encoding with limited channel bandwidth and buffer resources. However, there still has not been enough study on such an issue. In this paper, a new objective visual quality metric (VQM) is proposed first, which can easily be incorporated into video coding for guiding video coding. Second, a VQM-based window model is proposed to handle the tradeoff between visual quality consistency and buffer constraint in video coding. Third, a window-level rate control algorithm is developed to accomplish visual quality control based on the above two proposals. Finally, experimental results prove that consistent visual quality, high rate-distortion efficiency, accurate bit control, and compliant buffer constraint can be achieved by the proposed rate control algorithm. Long Xu 0001, Songnan Li, King Ngi Ngan, Lin Ma 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2012 | Study of subjective and objective quality assessment of retargeted imagesabstractThis paper presents the result of a recent large-scale subjective study of image retargeting quality on a collection of images generated by several representative image retargeting methods. Owning to many approaches to image retargeting that are developed, there is a need for a diverse independent public database of the retargeted images and the corresponding subjective scores that is freely available. We build an image retargeting quality database, in which 171 retargeted images (obtained from 57 natural source images of different contents) were generated by several representative image retargeting methods. The perceptual quality of each image is evaluated by at least 30 human subjects and the mean opinion scores (MOS) were recorded. Furthermore, several publicly available quality metrics for the retargeted images are evaluated on the built database. The database is made available [1] to the research community in order to further research on the perceptual quality assessment of the retargeted images. Lin Ma 0002, Weisi Lin, Chenwei Deng, King Ngi Ngan |
ISCAS | 1 |
| 2012 | Learning-based image restoration for compressed images
Lin Ma 0002, Debin Zhao, Wen Gao 0001 |
Signal Process. Image Commun. | 1 |
| 2012 | Full-Reference Video Quality Assessment by Decoupling Detail Losses and Additive ImpairmentsabstractVideo quality assessment plays a fundamental role in video processing and communication applications. In this paper, we study the use of motion information and temporal human visual system (HVS) characteristics for objective video quality assessment. In our previous work, two types of spatial distortions, i.e., detail losses and additive impairments, are decoupled and evaluated separately for spatial quality assessment. The detail losses refer to the loss of useful visual information that will affect the content visibility, and the additive impairments represent the redundant visual information in the test image, such as the blocking or ringing artifacts caused by data compression and so on. In this paper, a novel full-reference video quality metric is developed, which conceptually comprises the following processing steps: 1) decoupling detail losses and additive impairments within each frame for spatial distortion measure; 2) analyzing the video motion and using the HVS characteristics to simulate the human perception of the spatial distortions; and 3) taking into account cognitive human behaviors to integrate frame-level quality scores into sequence-level quality score. Distinguished from most studies in the literature, the proposed method comprehensively investigates the use of motion information in the simulation of HVS processing, e.g., to model the eye movement, to predict the spatio-temporal HVS contrast sensitivity, to implement the temporal masking effect, and so on. Furthermore, we also prove the effectiveness of decoupling detail losses and additive impairments for video quality assessment. The proposed method is tested on two subjective quality video databases, LIVE and IVP, and demonstrates the state-of-the-art performance in matching subjective ratings. Songnan Li, Lin Ma 0002, King Ngi Ngan |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2012 | Reduced-Reference Video Quality Assessment of Compressed Video SequencesabstractIn this paper, a novel reduced-reference (RR) video quality assessment (VQA) is proposed by exploiting the spatial information loss and the temporal statistical characteristics of the interframe histogram. From the spatial perspective, an energy variation descriptor (EVD) is proposed to measure the energy change of each individual encoded frame, which results from the quantization process. Besides depicting the energy change, EVD can further simulate the texture masking property of the human visual system (HVS). From the temporal perspective, the generalized Gaussian density (GGD) function is employed to capture the natural statistics of the interframe histogram distribution. The city-block distance (CBD) is used to calculate the histogram distance between the original video sequence and the encoded one. For simplicity, the difference image between adjacent frames is employed to characterize the temporal interframe relationship. By combining the spatial EVD together with the temporal CBD, an efficient RR VQA is developed. Evaluation on the subjective quality video database demonstrates that the proposed method outperforms the representative RR video quality metric and the full-reference VQAs, such as peak signal-to-noise ratio and structure similarity index in matching subjective ratings. This means that the proposed metric is more consistent with the HVS perception. Furthermore, as only a small number of RR features are extracted for representing the original video sequence (each frame requires only one parameter for describing EVD and three parameters for recording GGD), the RR features can be embedded into the video sequences or transmitted through the ancillary data channel, which can be used in the video quality monitoring system. Lin Ma 0002, Songnan Li, King Ngi Ngan |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2011 | Motion trajectory based visual saliency for video quality assessmentabstractIn this paper, we propose a novel visual saliency detection method for video sequences by considering the object motion trajectories. Firstly, each frame of the video sequence is described in a new Quaternion Representation (QR), which comprises the spatial image content and the temporal motion characteristics. Based on the QR, Quaternion Fourier Transform (QFT) is employed to construct the visual salien-cy of the video sequence. Finally, the detected visual salien-cy map is incorporated with several video quality metrics. Compared with other visual saliency models, the proposed method can improve the performances of video quality metrics. It further confirms that the proposed visual saliency model can accurately depict the Human Vision System (HVS) properties. Lin Ma 0002, Songnan Li, King Ngi Ngan |
ICIP | 1 |
| 2011 | Perceptual image compression via adaptive block- based super-resolution directed down-samplingabstractIn this paper, we propose a novel perceptual image coding scheme via adaptive block-based super-resolution directed down-sampling. At the encoder side, for each macroblock of a given image, Rate Distortion Optimization (RDO) determines whether it is encoded at the original or down-sampled resolution. The down-sampling process is directed by super-resolution, which generates the down-sampled block by minimizing the reconstruction errors between the original macroblock and the one restored by the corresponding super-resolution method. At the decoder side, in order to reduce the complexity, the super-resolution method reconstructs the full-resolution macroblock in the DCT domain together with the inverse DCT. Experimental results have demonstrated that the proposed method can produce higher quality images in terms of both PSNR and visual quality compared with the existing methods. Lin Ma 0002, Songnan Li, King Ngi Ngan |
ISCAS | 1 |
| 2011 | Adaptive Block-size Transform based Just-Noticeable Difference model for images/videos
Lin Ma 0002, King Ngi Ngan, Fan Zhang 0093, Songnan Li |
Signal Process. Image Commun. | 1 |
| 2011 | Image Quality Assessment by Separately Evaluating Detail Losses and Additive ImpairmentsabstractIn the research field of image processing, mean squared error (MSE) and peak signal-to-noise ratio (PSNR) are extensively adopted as the objective visual quality metrics, mainly because of their simplicity for calculation and optimization. However, it has been well recognized that these pixel-based difference measures correlate poorly with the human perception. Inspired by existing works, in this paper we propose a novel algorithm which separately evaluates detail losses and additive impairments for image quality assessment. The detail loss refers to the loss of useful visual information which affects the content visibility, and the additive impairment represents the redundant visual information whose appearance in the test image will distract viewer's attention from the useful contents causing unpleasant viewing experience. To separate detail losses and additive impairments, a wavelet-domain decoupling algorithm is developed which can be used for a host of distortion types. Two HVS characteristics, i.e., the contrast sensitivity function and the contrast masking effect, are taken into account to approximate the HVS sensitivities. We propose two simple quality measures to correlate detail losses and additive impairments with visual quality, respectively. Based on the findings inthat observers judge low-quality images in terms of the ability to interpret the content, the outputs of the two quality measures are adaptively combined to yield the overall quality index. By conducting experiments based on five subjectively-rated image databases, we demonstrate that the proposed metric has a better or similar performance in matching subjective ratings when compared with the state-of-the-art image quality metrics. Songnan Li, Fan Zhang 0093, Lin Ma 0002, King Ngi Ngan |
IEEE Trans. Multim. | 3 |
| 2011 | Reduced-Reference Image Quality Assessment Using Reorganized DCT-Based Image RepresentationabstractIn this paper, a novel reduced-reference (RR) image quality assessment (IQA) is proposed by statistical modeling of the discrete cosine transform (DCT) coefficient distributions. In order to reduce the RR data rates and further exploit the identical nature of the coefficient distributions between adjacent DCT subbands, the DCT coefficients are reorganized into a three-level coefficient tree. Subsequently, generalized Gaussian density (GGD) is employed to model the coefficient distribution of each reorganized DCT subband. The city-block distance is employed to measure the difference between the two images. Experimental results demonstrate that only a small number of RR features is sufficient for representing the image perceptual quality. The proposed method outperforms the RR WNISM and even the full-reference (FR) quality metric PSNR. Lin Ma 0002, Songnan Li, Fan Zhang 0093, King Ngi Ngan |
IEEE Trans. Multim. | 1 |
| 2011 | Practical Image Quality Metric Applied to Image CodingabstractPerceptual image coding requires an effective image quality metric, yet most of the existing metrics are complex and can hardly guide the compression effectively. This paper proposes a practical full-reference metric with consideration of the texture masking effect and contrast sensitivity function. The metric is capable of evaluating typical image impairments in real-world applications and can achieve the comparable performance as the state-of-the-art metrics on the publicly available subjectively-rated image databases. Due to its simplicity, the metric is embedded into JPEG image coding to ensure a better perceptual rate-distortion performance. Fan Zhang 0093, Lin Ma 0002, Songnan Li, King Ngi Ngan |
IEEE Trans. Multim. | 2 |
| 2010 | Video Quality Assessment based on Adaptive Block-size Transform Just-Noticeable Difference modelabstractIn this paper, we propose a full reference Video Quality Assessment (VQA) algorithm based on the Adaptive Block-size Transform Just-Noticeable Difference (ABT-JND) model. Firstly, ABT-JND is introduced for its efficiency of modeling the Human Vision System (HVS) characteristics. Based on the ABT-JND model, the full reference VQA is developed, by capturing HVS responses of spatio-temporal distortions over different block-size transforms. Experimental results have demonstrated that the proposed VQA outperforms other VQA methods, while slightly poorer than MOVIE. However, it maintains a very simple formulation. Since the proposed VQA performs on transform domain, it could be easily applied on many related applications, such as video compression, watermarking, and so on. Lin Ma 0002, Fan Zhang 0093, Songnan Li, King Ngi Ngan |
ICIP | 1 |
| 2010 | Adaptive block-size transform based just-noticeable difference profile for videosabstractIn this paper, we propose a novel adaptive block-size transform (ABT) based just-noticeable difference (JND) model for videos. Firstly, the ABT-based spatial JND profile is extended to spatial-temporal JND model for videos by considering temporal contrast sensitivity function (TCSF), eye movement, and the motion information of the objects in video sequence. Furthermore, a metric named motion characteristics distance (MCD) is proposed to depict the motion characteristics similarity between a macroblock and its corresponding sub-blocks. Based on the proposed MCD and the obtained spatial image content information, a novel balanced strategy is proposed to determine which transform size is employed to generate the resulting JND model. Experimental results have demonstrated that our proposed scheme could tolerate more distortions while preserving better perceptual quality than other JND profiles, which means that the proposed model consists well with human vision system (HVS). Moreover, for the balanced strategy, experiments have shown that temporal motion characteristics accord very well with the spatial image content information, which has demonstrated the efficiency of our proposed balanced strategy. Lin Ma 0002, King Ngi Ngan |
ISCAS | 1 |
| 2010 | Temporal inconsistency measure for video quality assessmentabstractVisual quality assessment plays a crucial role in many vision-related signal processing applications. In the literature, more efforts have been spent on spatial visual quality measure. Although a large number of video quality metrics have been proposed, the methods to use temporal information for quality assessment are less diversified. In this paper, we propose a novel method to measure the temporal impairments. The proposed method can be incorporated into any image quality metric to extend it into a video quality metric. Moreover, it is easy to apply the proposed method in video coding system to incorporate with MSE for rate-distortion optimization. Songnan Li, Lin Ma 0002, Fan Zhang 0093, King Ngi Ngan |
PCS | 2 |
| 2010 | Limitation and challenges of image quality measurementabstractSubjectively-rated image databases have become increasingly popular in the evaluation of image quality measurement algorithms. Several groups recently have improved their metrics' performance in matching these databases, using particular HVS (human visual system) properties or image statistical models. However, it is difficult to know whether these improvements are due to progress towards mimicking the perceptual properties, or are due to matching some characteristics of the databases. This paper demonstrates an inherent limitation in using such databases, showing that our very simple metric, built on the contrast masking effect, is able to perform as good as many state-of-the-art metrics. It is also argued that existent databases neither contain enough images with particularly biased distortions to test the significance of single HVS property, nor cover diverse distortion types to reflect the requirement of emerging applications. Fan Zhang 0093, Songnan Li, Lin Ma 0002, King Ngi Ngan |
VCIP | 3 |
| 2010 | Visual Horizontal Effect for Image Quality AssessmentabstractIn this paper, an image quality metric is proposed by modeling the visual Horizontal Effect (HE) and saliency property over structural distortions. Specifically, Structrue SIMilarity (SSIM) is firstly performed to obtain the structural distortion map. Subsequently, the obtained distortion map is refined by the visual HE model, which depicts visual sensitivities of oriented stimuli over different oriented contents. Finally, in order to describe the local Human Visual System (HVS) conspicuities, a saliency pooling strategy is proposed to generate the resulting image quality index. The experimental results have demonstrated that the proposed method outperforms SSIM and Visual Information Fidelity (VIF), which indicates that the obtained similarity index is more consistent with the perceptual evaluation of image quality. Lin Ma 0002, Songnan Li, King Ngi Ngan |
IEEE Signal Process. Lett. | 1 |
| 2008 | Three-tiered network model for image hallucinationabstractIn this paper, we propose a novel three-tiered network model for image hallucination based on the learnt knowledge composed of image patches relating low and high resolution. A common problem of previous hallucination methods is that irregularities are usually introduced into the constructed high-resolution images. We remove the irregularities in three steps. First, the hallucination with primal sketch priors is performed to construct a coarse high-frequency component. Second, enhancement is implemented to enforce local compatibility between the patches in the constructed component. Third, a Markov network is utilized to refine the enhanced high-frequency component. Experiments demonstrate that our model can hallucinate higher-quality images than existing methods. Lin Ma 0002, Yan Lu 0001, Feng Wu 0001, Debin Zhao |
ICIP | 1 |