EDBT 2026 Demo / reviewers in the wild / expert
Di Lin 0002
dblp:20/3191-2
· DBLP profile ↗
62ranked-venue papers
12as first author
45since 2021 · last 2026
0000-0002-9324-800XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 48 · 11 first-author · 36 since 2021Graphics, computer vision, multimedia, augmented reality and games · 29 · 7 first-author · 17 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Systems, architecture and hardware · 3 · 3 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CorrectAD: A Self-Correcting Agentic System to Improve End-to-end Planning in Autonomous DrivingabstractEnd-to-end planning methods are the de-facto standard of the current autonomous driving system, while the robustness of the data-driven approaches suffers due to the notorious long-tail problem (i.e., rare but safety-critical failure cases). In this work, we explore whether recent diffusion-based video generation methods (a.k.a. world models), paired with structured 3D layouts, can enable a fully automated pipeline to self-correct such failure cases. We first introduce an agent to simulate the role of product manager, dubbed PM-Agent, which formulates data requirements to collect data similar to the failure cases. Then, we use a generative model that can simulate both data collection and annotation. However, existing generative models struggle to generate high-fidelity data conditioned on 3D layouts. To address this, we propose DriveSora, which can generate spatiotemporally consistent videos aligned with the 3D annotations requested by PM-Agent. We integrate these components into our self-correcting agentic system, CorrectAD. Importantly, our pipeline is end-to-end model agnostic and can be applied to improve any end-to-end planner. Evaluated on both nuScenes and a more challenging in-house dataset across multiple end-to-end planners, CorrectAD corrects 62.5% and 49.8% of failure cases, reducing collision rates by 39% and 27%, respectively. Enhui Ma, Junpeng Jiang, Kun Zhan, Xueyang Zhang, Xianpeng Lang, Di Lin 0002, Kaicheng Yu |
AAAI | 13 |
| 2026 | FreeMem: Enhancing Consistency in Long Video Generation via Tuning-Free MemoryabstractText-to-Video (T2V) generation has advanced greatly, yet maintaining consistency remains challenging, especially for tuning-free long video generation. We attribute the consistency problem to cumulative deviations for long video generation at three levels: the random noise lacking correlation results initial deviation between frames; discrepancy in semantic feature tokens between denoising network blocks gradually accumulates as the frame count grows, leading to greater deviations; attention mechanisms struggle to capture global relationships across distant frames in long videos. To address these, we propose FreeMem, a tuning-free framework leveraging hierarchical memory update and injection: the noise memory stabilizes consistency by manipulating low and high frequency components in the initial noise space; the token memory combats inconsistency through adaptive fusion of historical and current semantic feature tokens between denoising network blocks; and the attention memory establishes persistent cache to model long-range relationships within self attention layers. Evaluated on VBench, FreeMem improves subject and background consistency matrics across various methods, offering a practical solution for low-cost, high-consistency long video generation. Jibin Peng, Di Lin 0002, Zhecheng Xu, Wuyuan Xie, Miaohui Wang, Lingyu Liang, Qing Guo 0005 |
AAAI | 2 |
| 2026 | Exploiting Geometric Structures for Modeling Multi-Agent Behaviors: A New ThinkingabstractIn this paper, we rethink model agent behaviors from a geometric structure perspective in multi-agent reinforcement learning. Modeling agent behaviors is essential for understanding how agents interact and facilitating effective decisions. The key lies in capturing the dependencies and sequential relationships among agent decisions. Since each decision influences the subsequent choices, this forms a hierarchical and nested tree-like structure of interdependencies. While modeling tree-like data in Euclidean spaces could cause distortion, which results in a loss of agent decision structure information. Motivated by this, we reconsider model agent behaviors in hyperbolic space and propose the Hyperbolic Multi-Agent Representations (HMAR) method, which projects the agent behaviors into a Poincaré ball and leverages hyperbolic neural networks to learn agent policy representations. Additionally, we designed a contrastive loss function to train this network, minimizing the distance in feature space between different representations of the same agent while maximizing the distance between representations of distinct agents. Experimental results provide empirical evidence for the effectiveness of the HMAR method in cooperative and competitive environments, demonstrating the potential of hyperbolic agent representations for effective decision-making in multi-agent environments. Bohao Qu, Xiaofeng Cao 0002, Bing Li 0001, Menglin Zhang, Tuan-Anh Vu, Di Lin 0002, Qing Guo 0005 |
AAAI | 6 |
| 2026 | Firing Bits Where It Matters: Spiking-Guided Just Recognizable Distortion Modeling for Machine-Centric Video CodingabstractJust recognizable distortion (JRD) has emerged as a promising paradigm for machine-centric video coding. However, existing JRD-guided coding methods are limited by coarse annotation granularity and high computational cost, which hinder their deployment. In this paper, we first investigate the impact of different JRD annotation strategies on downstream task performance. By incorporating both instance-level and contextual information, we construct a new JRD dataset with fine-grained annotations compatible with object detection and instance segmentation tasks. To enhance quantization parameter (QP) map prediction while maintaining computational efficiency, we propose a novel spiking neural network (SNN)-based framework that decomposes video frames into spatial structures, channel interactions, and temporal patterns. Furthermore, we introduce a spiking attention mechanism to aggregate task-relevant features and employ adaptive scaling vectors to suppress machine-perceived redundancy, enabling targeted bitrate allocation aligned with task-critical content. Extensive experiments on multiple datasets and backbones demonstrate that our approach consistently outperforms state-of-the-art codec-based and JRD-guided methods in maintaining task performance at ultra-low bitrates, while significantly reducing computational overhead. Wuyuan Xie, Zhenming Li, Yuwu Lu, Di Lin 0002, Yun Song, Miaohui Wang |
AAAI | 4 |
| 2026 | Causal Inference via Style Bias Deconfounding for Domain GeneralizationabstractDeep neural networks (DNNs) often struggle with out-of-distribution data, limiting their reliability in real-world visual applications. To address this issue, domain generalization methods have been developed to learn domain-invariant features from single or multiple training domains, enabling generalization to unseen testing domains. However, existing approaches usually overlook the impact of style frequency within the training set. This oversight predisposes models to capture spurious visual correlations caused by style confounding factors, rather than learning truly causal representations, thereby undermining inference reliability. In this work, we introduce Style Deconfounding Causal Learning (SDCL), a novel causal inference-based framework that explicitly addresses style as a confounding factor to enhance domain generalization in image modalities. Our approaches begins with constructing a structural causal model (SCM) tailored to the domain generalization problem and applies a backdoor adjustment strategy to account for style influence. Building on this foundation, we design a style-guided expert module (SGEM) to adaptively clusters style distributions during training, capturing the global confounding style. Additionally, a backdoor causal learning module (BDCL) performs causal interventions during feature extraction, ensuring fair integration of global confounding styles into sample predictions, effectively reducing style bias. The SDCL framework is highly versatile and can be seamlessly integrated with state-of-the-art data augmentation techniques. Extensive experiments across diverse natural and medical image recognition tasks validate its efficacy, demonstrating superior performance in both multi-domain and the more challenging single-domain generalization scenarios. Di Lin 0002, Hao Chen 0011, Hongying Liu 0001, Wei Feng 0005 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2026 | Time-Variant Image Inpainting via Interactive Distribution Transition EstimationabstractIn this work, we focus on a novel and practical task, i.e., Time-vAriant iMage inPainting (TAMP). The aim of TAMP is to restore a damaged target image by leveraging the complementary information from a reference image, where both images capture the same scene but with a significant time gap in between, i.e., time-variant images. Different from conventional reference-guided image inpainting, the reference image under TAMP setup presents significant content distinction to the target image and potentially also suffers from damages. Such an application frequently happens in our daily life to restore a damaged image by referring to another reference image, where there is no guarantee of the reference image's source and quality. In particular, our study finds that even SOTA reference-guided image inpainting methods fail to achieve plausible results due to the chaotic image complementation. To address such an ill-posed problem, we propose a novel Interactive Distribution Transition Estimation (InDiTE) module which interactively complements the time-variant images with appropriate semantics thus facilitate the restoration of damaged regions. To further boost the performance, we propose our TAMP solution, namely Interactive Distribution Transition Estimation-driven Diffusion (InDiTE-Diff), which integrates InDiTE with SOTA diffusion model and conducts latent cross-reference during sampling. Moreover, considering the lack of benchmarks for TAMP task, we newly assembled a dataset, i.e., TAMP-Street, based on existing image and mask datasets. We conduct experiments on the TAMP-Street datasets under two different time-variant image inpainting settings, which show our method consistently outperform SOTA reference-guided image inpainting methods for solving TAMP. Yun Xing 0001, Qing Guo 0005, Yihao Huang 0001, Xiaofeng Cao 0002, Luqi Gong, Di Lin 0002, Ivor W. Tsang, Lei Ma 0003 |
IEEE Trans. Image Process. | 7 |
| 2025 | mmFAS: Multimodal Face Anti-Spoofing Using Multi-Level Alignment and Switch-Attention FusionabstractThe increasing number of presentation attacks on reliable face matching has raised concerns and garnered attention towards face anti-spoofing (FAS). However, existing methods for FAS modeling commonly fuse multiple visual modalities (e.g., RGB, Depth, and Infrared) in a straightforward manner, disregarding latent feature gaps that can hinder representation learning. To address this challenge, we propose a novel multimodal FAS framework (mmFAS) that focuses on explicit alignment and fusion of latent features across different modalities. Specifically, we develop a multimodal alignment module to alleviate the latent feature gap by using instance-level contrastive learning and class-level matching simultaneously. Further, we explore a new switch-attention based fusion module to automatically aggregate complementary information and control model complexity. To evaluate the anti-spoofing performance more effectively, we adopt a challenging yet meaningful cross-database protocol involving four benchmark multimodal FAS datasets to simulate realworld scenarios. Extensive experimental results demonstrate the effectiveness of mmFAS in improving the accuracy of FAS systems, outperforming 10 representative methods. Geng Chen 0006, Wuyuan Xie, Di Lin 0002, Ye Liu 0005, Miaohui Wang |
AAAI | 3 |
| 2025 | SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World EnvironmentsabstractLarge vision-language models (LVLMs) have shown remarkable capabilities in interpreting visual content. While existing works demonstrate these models’ vulnerability to deliberately placed adversarial texts, such texts are often easily identifiable as anomalous. In this paper, we present the first approach to generate scene-coherent typographic adversarial attacks that mislead advanced LVLMs while maintaining visual naturalness through the capability of the LLM-based agent. Our approach addresses three critical questions: what adversarial text to generate, where to place it within the scene, and how to integrate it seamlessly. We propose a training-free, multi-modal LLM-driven scene-coherent typographic adversarial planning (SceneTAP) that employs a three-stage process: scene understanding, adversarial planning, and seamless integration. The SceneTAP utilizes chain-of-thought reasoning to comprehend the scene, formulate effective adversarial text, strategically plan its placement, and provide detailed instructions for natural integration within the image. This is followed by a scene-coherent TextDiffuser that executes the attack using a local diffusion mechanism. We extend our method to real-world scenarios by printing and placing generated patches in physical environments, demonstrating its practical implications. Extensive experiments show that our scene-coherent adversarial text successfully misleads state-of-the-art LVLMs, including ChatGPT-4o, even after capturing new images of physical setups. Our evaluations demonstrate a significant increase in attack success rates while maintaining visual naturalness and contextual appropriateness. This work highlights vulnerabilities in current vision-language models to sophisticated, scene-coherent adversarial attacks and provides insights into potential defense mechanisms. We release our code at https://github.com/tsingqguo/scenetap. Jie Zhang 0002, Di Lin 0002, Tianwei Zhang 0004, Ivor W. Tsang, Yang Liu 0003, Qing Guo 0005 |
CVPR | 4 |
| 2025 | Generative Hard Example Augmentation for Semantic Point Cloud SegmentationabstractThe recent progress in semantic point cloud segmentation is attributed to deep networks, which require a large amount of point cloud data for training. However, how to collect substantial point-wise annotations of the point clouds at affordable cost for the end-to-end network training still needs to be solved. In this paper, we propose Generative Hard Example Augmentation (GHEA) to achieve novel examples of point clouds, which enrich the data for training the segmentation network. Firstly, GHEA employs the generative network to embed the discrepancy between the point clouds into the latent space. From the latent space, we sample multiple discrepancies for reshaping a point cloud to various examples, contributing to the richness of the training data. Secondly, GHEA mixes the reshaped point clouds by respecting their segmentation errors. This mixup allows the reshaped point clouds, which are difficult to segment, to join as the challenging example for network training. We evaluate the effectiveness of GHEA, which helps the popular segmentation networks to improve the performances. Qi Zhang 0071, Jibin Peng, Wei Feng 0005, Di Lin 0002 |
CVPR | 5 |
| 2025 | NoiseController: Towards Consistent Multi-View Video Generation via Noise Decomposition and Collaboration
Haotian Dong, Xin Wang 0118, Di Lin 0002, Yipeng Wu, Kairui Yang, Ping Li 0016, Qing Guo 0005 |
ICCV | 3 |
| 2025 | Trajectory-LLM: A Language-based Data Generator for Trajectory Prediction in Autonomous DrivingabstractVehicle trajectory prediction is a crucial aspect of autonomous driving, which requires extensive trajectory data to train prediction models to understand the complex, varied, and unpredictable patterns of vehicular interactions. However, acquiring real-world data is expensive, so we advocate using Large Language Models (LLMs) to generate abundant and realistic trajectories of interacting vehicles efficiently. These models rely on textual descriptions of vehicle-to-vehicle interactions on a map to produce the trajectories. We introduce Trajectory-LLM (Traj-LLM), a new approach that takes brief descriptions of vehicular interactions as input and generates corresponding trajectories. Unlike language-based approaches that translate text directly to trajectories, Traj-LLM uses reasonable driving behaviors to align the vehicle trajectories with the text. This results in an "interaction-behavior-trajectory" translation process. We have also created a new dataset, Language-to-Trajectory (L2T), which includes 240K textual descriptions of vehicle interactions and behaviors, each paired with corresponding map topologies and vehicle trajectory segments. By leveraging the L2T dataset, Traj-LLM can adapt interactive trajectories to diverse map topologies. Furthermore, Traj-LLM generates additional data that enhances downstream prediction models, leading to consistent performance improvements across public benchmarks. The source code is released at https://github.com/TJU-IDVLab/Traj-LLM. Kairui Yang, Gengjie Lin, Haotian Dong, Yipeng Wu, Die Zuo, Jibin Peng, Ziyuan Zhong, Xin Wang 0118, Qing Guo 0005, Xiaosong Jia, Junchi Yan, Di Lin 0002 |
ICLR | 14 |
| 2025 | Causal Regularization Graph Attention Network for Stable Variables Decoupling and Fault Diagnosis of Complex Industrial ProcessesabstractGraph neural networks (GNNs) have emerged as a powerful tool for fault diagnosis in industrial processes, where components and their relationships can be modeled as nodes and edges in a graph. However, most GNN models are based on the I.I.D. hypothesis and learn both causal and spurious correlations. In real-world applications, test data distributions often differ from training data, leading to changes in spurious correlations and inaccurate diagnosis results. To address this issue, this paper proposes a causal regularization graph attention network (CR-GAT) that decouples causal and confounding variables and eliminates spurious correlations. The method involves extracting high-level graph variables using differentiable pooling, measuring nonlinear dependence between variables with the Hilbert-Schmidt Independence Criterion (HSIC), and iteratively optimizing GAT loss and variable weights to learn true causal relationships. Comparative experiments on three-phase flow and power system datasets demonstrate the effectiveness of the proposed method. Puyuan Hu, Siheng Zhao, Shifei Ma, Di Lin 0002, Weidong Zhang 0004 |
INDIN | 5 |
| 2025 | PhysGCN-DL: Physics-Informed Graph Convolutional Networks with Diversity-Aware Loss Optimization for Multimodal Pedestrian Trajectory PredictionabstractPedestrian trajectory prediction ensures safe navigation in autonomous driving and intelligent robots. Existing methods have shown promising results but still face challenges in handling dynamic environments, social interactions, and high-dimensional data. In this paper, we propose a novel PhysGCN-DL within the itransformer framework to address these challenges. Our model incorporates physically-inspired dynamic interaction modeling by representing physical interactions between pedestrians as edge weights in graph convolution. This approach captures the heterogeneity of pedestrian movement and improves the interpretability of social interactions. Moreover, we design a novel loss function to jointly enhance prediction diversity and accuracy, thereby improving the model’s robustness across both dense and sparse scenarios. Empirical evaluations confirm that our approach outperforms existing methods in generating accurate and diverse pedestrian trajectories. Zihan Jiang 0005, Haibo Lu, BoYuan Yang, Di Lin 0002, Weidong Zhang 0004 |
IROS | 6 |
| 2025 | CVLN-Think: Causal Inference with Counterfactual Style Adaptation for Continuous Vision-and-Language NavigationabstractVision-and-Language Navigation in Continuous Environments (VLN-CE) presents challenges due to environmental variations and domain shifts, making it difficult for agents to generalize beyond seen environments. Most existing methods rely on learning correlations between observations and actions from training data, which leads to spurious dependencies on environmental biases. To address this, we propose CVLN-Think (CVT), a novel navigation model that incorporates causal inference to enhance robustness and adaptability. Specifically, Style Causal Adjuster (SCA) generates counterfactual style observations, enabling agents to learn invariant spatial structures rather than overfitting to dataset-specific visual patterns. Furthermore, Thinking Cause Navigation Engine (TCNE) applies causal intervention to adjust navigation decisions by identifying and mitigating biases from prior experience. Unlike conventional approaches that passively learn from data distributions, our model actively thinks along the "observation-action" chain to make more reliable navigation predictions. Experimental results demonstrate that our approach achieves satisfactory performance on VLN-CE tasks. Further analysis indicates that our method possesses stronger generalization capabilities, highlighting the superiority of our proposed approach. Di Lin 0002, Weidong Zhang 0004 |
IROS | 3 |
| 2025 | Open-Vocabulary Part Segmentation via Progressive and Boundary-Aware StrategyabstractOpen-vocabulary part segmentation (OVPS) struggles with structurally connected boundaries due to the inherent conflict between continuous image features and discrete classification mechanism. To address this, we propose PBAPS, a novel training-free framework specifically designed for OVPS. PBAPS leverages structural knowledge of object-part relationships to guide a progressive segmentation from objects to fine-grained parts. To further improve accuracy at challenging boundaries, we introduce a Boundary-Aware Refinement (BAR) module that identifies ambiguous boundary regions by quantifying classification uncertainty, enhances the discriminative features of these ambiguous regions using high-confidence context, and adaptively refines part prototypes to better align with the specific image. Experiments on Pascal-Part-116, ADE20K-Part-234, PartImageNet demonstrate that PBAPS significantly outperforms state-of-the-art methods, achieving 46.35\% mIoU and 34.46\% bIoU on Pascal-Part-116. Our code is available at https://github.com/TJU-IDVLab/PBAPS. Xinlong Li, Di Lin 0002, Shaoyiyi Gao, Qing Guo 0005 |
NeurIPS | 2 |
| 2025 | EfficientDeRain+: Learning Uncertainty-Aware Filtering via RainMix Augmentation for High-Efficiency Deraining
Qing Guo 0005, Hua Qi, Jingyang Sun, Felix Juefei-Xu, Lei Ma 0003, Di Lin 0002, Wei Feng 0005, Song Wang 0002 |
Int. J. Comput. Vis. | 6 |
| 2025 | Causal Counterfactual Faithfulness Generation for Open-Set Fault Diagnosis of Complex Industrial ProcessesabstractTraditional intelligent fault diagnosis models are usually capable of diagnosing known types of faults. However, in the field of industrial fault diagnosis in open environments, it is almost impossible to collect training samples that cover all fault categories. Therefore, when encountering unknown types of fault, traditional methods tend to misclassify them as known categories. To address this issue, a causal counterfactual faithfulness generation method is proposed for open-set fault diagnosis of complex industrial processes. Initially, the signal data from fault sensors are processed into graph data composed of nodes and edges. Then, the features of nodes and their adjacent nodes are learned and integrated into graph architecture to generate new fault sample attributes. Subsequently, the causal generative model infers the category features and combines known fault categories to generate counterfactual samples. Finally, the sample’s classification as an unknown category is ultimately determined by testing the principle of consistency. The proposed method can significantly improve the accuracy of open-set diagnosis without affecting the accuracy of closed-set classification. Comparison experiments with multiple baseline models in two fault datasets illustrated that the proposed method shows an improvement in almost all indicators, which ultimately verified the effectiveness of the proposed method in the task of fault diagnosis in open environments. Puyuan Hu, Siheng Zhao, Di Lin 0002, Weidong Zhang 0004, Steven X. Ding |
IEEE Trans. Ind. Informatics | 4 |
| 2025 | Causal Disentangled Graph Neural Network for Fault Diagnosis of Complex Industrial ProcessabstractGraph neural networks (GNNs) are good at capturing the intricate topologies and dependencies among components and are outstanding in fault diagnosis tasks of complex industrial process. Bias substructures consisting of irrelevant sensor signals and noise data are simpler compared to causal substructures consisting of fault signals, and GNNs tend to utilize the letter to quickly achieve low loss. However, spurious correlations in the bias substructures will mislead predictions. To address this issue, this study takes the disentanglement of causal and bias substructures as the key to improve model stability. A causal disentangled GNN (CDGNN) is proposed. First, sensor signals are transformed into graph data employing an attention mechanism to capture the interactions between them. Then, a causal disentanglement learning module is designed to extract causal subgraphs from input graphs. Finally, causal subgraph features from different source machines are aggregated to form a complete graph representation. Experimental results on two complex industrial datasets indicate that CDGNN is an effective and stable method for fault diagnosis. Quanhu Zhang, Di Lin 0002, Weidong Zhang 0004, Steven X. Ding |
IEEE Trans. Ind. Informatics | 3 |
| 2025 | CarveNet: Carving Point-Block for Complex 3D Shape Completionabstract3D point cloud completion is very challenging because it relies on accurately understanding the complex 3D shapes (e.g., high-curvature, concave/convex, and hollowed-out 3D shapes) and the unknown & diverse patterns of the partially available point clouds. In this paper, we propose a novel solution, i.e.,Point-block Carving(PC), for completing the complex 3D point cloud completion. Given the partial point cloud as the guidance, we carve a 3D block that contains the uniformly distributed 3D points, yielding the entire point cloud. We propose a new network architecture to achieve PC, i.e.,CarveNet. This network conducts the exclusive convolution on each block point, where the convolutional kernels are trained on the 3D shape data. CarveNet determines which point should be carved to recover the complete shapes' details effectively. Furthermore, we propose a sensor-aware method for data augmentation, i.e.,SensorAug, for training CarveNet on richer patterns of partial point clouds, thus enhancing the completion power of the network. The extensive evaluations on the ShapeNet, ShapNet-55/34 and KITTI datasets demonstrate the generality of our approach on the partial point clouds with diverse patterns. On these datasets, CarveNet successfully outperforms the state-of-the-art methods. Qing Guo 0005, Zhijie Wang 0014, Lubo Wang, Haotian Dong, Felix Juefei-Xu, Di Lin 0002, Lei Ma 0003, Wei Feng 0005, Yang Liu 0003 |
IEEE Trans. Multim. | 6 |
| 2025 | Accurate-PGNet: Learning to Assemble Perceptual Body Parts for Accurate Human Skeleton EstablishmentabstractThe human skeleton establishment aims to provide accurate localization information of the human body from RGB images and establish a complete human skeleton for many applications, such as action recognition, video surveillance, and human-computer interaction. Considering the inherent human body structure, many recent methods group the relevant body parts and utilize the deep convolutional network to learn the visual context from the part groups. However, the grouping approaches used in these methods heavily rely on prior knowledge of the human body shape but lose important relationships between parts. In this paper, we introduce the Accurate Part Grouping Network (Accurate-PGNet), a novel network for hierarchically grouping body parts in a data-driven manner. In contrast to the previous methods, we use neural architecture search (NAS) to optimize the architecture of Accurate-PGNet and properly group the body parts. The part grouping respects the diverse visual patterns of parts, producing groups containing different body parts. From each group, we learn the visual feature map. It helps to capture the correlation between parts and predict their locations. The feature maps of the part groups are merged hierarchically to capture the higher-order context of parts in larger groups. We extensively evaluated our method on the challenging benchmarks, demonstrating that Accurate-PGNet effectively helps to achieve state-of-the-art results. Di Lin 0002, Xin Wang 0118, George Baciu, C. L. Philip Chen, Ping Li 0016 |
IEEE Trans. Multim. | 2 |
| 2025 | Point-to-Set Metric-Gated Mixture of Experts for Multisource Domain Adaptation Fault DiagnosisabstractThe multisource unsupervised domain adaptation (MUDA) scenario poses a significant challenge in the field of intelligent fault diagnosis (IFD), where the goal is to transfer the knowledge learned from multiple labeled source domains to an unlabeled target domain. Existing IFD-oriented MUDA approaches frequently fail to recognize the distinct importance of each source domain relative to specific target samples, or lack flexibility in integrating diagnostic insights from multiple sources. In response, a novel MUDA approach is proposed for IFD, termed point-to-set metric-gated mixture of experts (PSMMoEs). This method leverages a mixture-of-experts (MoEs) framework to automatically integrate the complementary information from multiple source domains. It develops a deep point-to-set distance (PSD) metric learning technique within the MoE's gating mechanism, effectively fusing domain-specific features by assessing the similarity between individual target samples and each source domain. The method ensures balanced training across progressive stages, harmonizing multitask learning with joint training for the MoE framework. Furthermore, a multilayer maximum mean discrepancy (MMD) measurement is employed for domain alignment, ensuring feature alignment across different domains at multiple levels. In order to assess the efficacy of the proposed method, it is compared with several leading domain adaptation methods on publicly available and laboratory-based rotating machinery fault datasets. The experimental results demonstrate superior classification and adaptation capabilities of the proposed fault diagnosis method. Boyuan Yang 0002, Di Lin 0002, Ping Li 0016, C. L. Philip Chen |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2025 | Temporal-Interim Pose Synthesis and Distillation for Dynamic Human Pose EstimationabstractIn the task of dynamic human pose estimation (dynamic HPE), the temporal relationships between human body parts should be captured comprehensively to understand the dynamic human motions, where the correlated motion information eventually helps to recognize body parts. The popular methods are successful in terms of utilizing long-term motion information captured by low-speed cameras. Yet they neglect the underlying intermediate motions between captured frames, which comprise the temporal-interim poses lost in the video. In this article, we introduce a novel framework, temporal-interim pose synthesis and distillation, to produce and leverage the intermediate motion information for dynamic motion establishment. The pose synthesis yields the visual feature maps of the intermediate poses, which appear between the existing video frames. It allows the synthesized and current poses to form richer motion patterns. Next, the pose distillation divides the body parts into several groups, where it learns the specific part-wise relationship within each group. It degrades the complexity of learning useful part-wise relationships from rich motion patterns and extracts more detailed motion information for fine-grained part groups. We extensively evaluate our method on challenging datasets for dynamic pose estimation, achieving state-of-the-artresults. Di Lin 0002, Xin Wang 0118, Bin Sheng 0001, George Baciu, C. L. Philip Chen, Ping Li 0016 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2025 | HRC-Net: Learning Visual Hypothesis, Representative, and Collaboration for Multi-Domain Image InpaintingabstractMulti-domain image inpainting utilizes complementary contextual information from auxiliary domain images to restore corrupted regions. While existing methods reconstruct auxiliary images to provide additional guidance, they face fundamental limitations: recovered pixels with complex patterns often lack representative details, while oversimplified patterns offer insufficient contextual information. To address these challenges, we propose HRC-Net, a novel framework incorporating three generative sub-networks for the comprehensive image inpainting task. Our architecture consists of: (1) A Hypothesis Sub-network that enables robust samplings of pixel-wise hypotheses from multi-domain inputs; (2) A Representative Sub-network that learns to score hypothesis quality based on contextual relevance; and (3) a Collaboration Sub-network that optimizes adaptive fusion kernels to integrate the most pertinent details. Together, these components model the joint distribution of representative scores and convolutional kernels, fostering a precise interaction between auxiliary hypotheses and target image corruption to meticulously repair the target image. Extensive evaluations across multiple benchmark datasets demonstrate HRC-Net's superior performance, significantly outperforming state-of-the-art methods in both quantitative metrics and visual quality. Xin Wang 0118, Di Lin 0002, Wanchao Su, Ji Du, Jie Zhang 0090, Haotian Dong, Ke Xu 0010, Qing Guo 0005, Ping Li 0016 |
ACM Trans. Graph. | 2 |
| 2024 | msLPCC: A Multimodal-Driven Scalable Framework for Deep LiDAR Point Cloud CompressionabstractLiDAR sensors are widely used in autonomous driving, and the growing storage and transmission demands have made LiDAR point cloud compression (LPCC) a hot research topic. To address the challenges posed by the large-scale and uneven-distribution (spatial and categorical) of LiDAR point data, this paper presents a new multimodal-driven scalable LPCC framework. For the large-scale challenge, we decouple the original LiDAR data into multi-layer point subsets, compress and transmit each layer separately, so as to ensure the reconstruction quality requirement under different scenarios. For the uneven-distribution challenge, we extract, align, and fuse heterologous feature representations, including point modality with position information, depth modality with spatial distance information, and segmentation modality with category information. Extensive experimental results on the benchmark SemanticKITTI database validate that our method outperforms 14 recent representative LPCC methods. Miaohui Wang, Runnan Huang, Hengjin Dong, Di Lin 0002, Yun Song, Wuyuan Xie |
AAAI | 4 |
| 2024 | LRR: Language-Driven Resamplable Continuous Representation against Adversarial Tracking AttacksabstractVisual object tracking plays a critical role in visual-based autonomous systems, as it aims to estimate the position and size of the object of interest within a live video. Despite significant progress made in this field, state-of-the-art (SOTA) trackers often fail when faced with adversarial perturbations in the incoming frames. This can lead to significant robustness and security issues when these trackers are deployed in the real world. To achieve high accuracy on both clean and adversarial data, we propose building a spatial-temporal continuous representation using the semantic text guidance of the object of interest. This novel continuous representation enables us to reconstruct incoming frames to maintain semantic and appearance consistency with the object of interest and its clean counterparts. As a result, our proposed method successfully defends against different SOTA adversarial tracking attacks while maintaining high accuracy on clean data. In particular, our method significantly increases tracking accuracy under adversarial attacks with around 90% relative improvement on UAV123, which is even higher than the accuracy on clean data. Jianlang Chen, Xuhong Ren, Qing Guo 0005, Felix Juefei-Xu, Di Lin 0002, Wei Feng 0005, Lei Ma 0003, Jianjun Zhao 0001 |
ICLR | 5 |
| 2024 | Sim2Real-Fire: A Multi-modal Simulation Dataset for Forecast and Backtracking of Real-world Forest FireabstractThe latest research on wildfire forecast and backtracking has adopted AI models, which require a large amount of data from wildfire scenarios to capture fire spread patterns. This paper explores using cost-effective simulated wildfire scenarios to train AI models and apply them to the analysis of real-world wildfire. This solution requires AI models to minimize the Sim2Real gap, a brand-new topic in the fire spread analysis research community. To investigate the possibility of minimizing the Sim2Real gap, we collect the Sim2Real-Fire dataset that contains 1M simulated scenarios with multi-modal environmental information for training AI models. We prepare 1K real-world wildfire scenarios for testing the AI models. We also propose a deep transformer, S2R-FireTr, which excels in considering the multi-modal environmental information for forecasting and backtracking the wildfire. S2R-FireTr surpasses state-of-the-art methods in real-world wildfire scenarios. Keqiu Li, Li Guohui, Changqing Ji, Lubo Wang, Die Zuo, Qing Guo 0005, Manyu Wang 0001, Di Lin 0002 |
NeurIPS | 11 |
| 2024 | Voxel Proposal Network via Multi-Frame Knowledge Distillation for Semantic Scene CompletionabstractSemantic scene completion is a difficult task that involves completing the geometry and semantics of a scene from point clouds in a large-scale environment. Many current methods use 3D/2D convolutions or attention mechanisms, but these have limitations in directly constructing geometry and accurately propagating features from related voxels, the completion likely fails while propagating features in a single pass without considering multiple potential pathways. And they are generally only suitable for static scenes and struggle to handle dynamic aspects. This paper introduces Voxel Proposal Network (VPNet) that completes scenes from 3D and Bird's-Eye-View (BEV) perspectives. It includes Confident Voxel Proposal based on voxel-wise coordinates to propose confident voxels with high reliability for completion. This method reconstructs the scene geometry and implicitly models the uncertainty of voxel-wise semantic labels by presenting multiple possibilities for voxels. VPNet employs Multi-Frame Knowledge Distillation based on the point clouds of multiple adjacent frames to accurately predict the voxel-wise labels by condensing various possibilities of voxel relationships. VPNet has shown superior performance and achieved state-of-the-art results on the SemanticKITTI and SemanticPOSS datasets. Lubo Wang, Di Lin 0002, Kairui Yang, Qing Guo 0005, Wuyuan Xie, Miaohui Wang, Lingyu Liang, Ping Li 0016 |
NeurIPS | 2 |
| 2024 | Non-iterative scribble-supervised learning with pacing pseudo-masks for medical image segmentation
Zefan Yang, Di Lin 0002, Dong Ni 0001, Yi Wang 0031 |
Expert Syst. Appl. | 2 |
| 2024 | Recurrent feature propagation and edge skip-connections for automatic abdominal organ segmentation
Zefan Yang, Di Lin 0002, Dong Ni 0001, Yi Wang 0031 |
Expert Syst. Appl. | 2 |
| 2024 | Intelligent Bearing Anomaly Detection for Industrial Internet of Things Based on Auto-Encoder Wasserstein Generative Adversarial NetworkabstractBearing anomaly detection plays a crucial role in modern industries as most rotating machinery faults are attributed to faulty bearings. However, acquiring fault samples in industry is a time-consuming and expensive process. To address this issue, this paper presents an integrated unsupervised learning method named AE-AnoWGAN (Autoencoder Wasserstein Generative Adversarial Network). AE-AnoWGAN is capable of detecting abnormal bearings and performing anomaly localization without the need for labeled data. In this approach, industrial data is initially processed using continuous wavelet transform to convert it into time-frequency representations (TFRs). These TFRs are then fed into the integrated AE-AnoWGAN for training. AE-AnoWGAN consists of multiple encoder-decoder and discriminator pairs, which are randomly paired and trained using adversarial training. The encoder maps the TFRs to a latent space, and the pre-trained generator acts as the decoder to generate reconstructed TFRs. During the testing phase, the model calculates anomaly scores for the input TFRs. Experimental evaluations were conducted using the PU bearing dataset and IMS bearing dataset. Comparative results demonstrate that the proposed AE-AnoWGAN method outperforms existing approaches in terms of anomaly detection accuracy. Moreover, the method exhibits high anomaly detection efficiency, making it suitable for real-time monitoring applications. Furthermore, this method provides practical value by enabling anomaly localization and bearing degradation estimation of TFRs. Di Lin 0002, Weidong Zhang 0004 |
IEEE Internet Things J. | 3 |
| 2024 | Sign language translation with hierarchical memorized context in question answering scenarios
Liqing Gao, Wei Feng 0005, Rui-Ze Han, Di Lin 0002, Liang Wang 0001 |
Neural Comput. Appl. | 5 |
| 2023 | CVSformer: Cross-View Synthesis Transformer for Semantic Scene CompletionabstractSemantic scene completion (SSC) requires an accurate understanding of the geometric and semantic relationships between the objects in the 3D scene for reasoning the occluded objects. The popular SSC methods voxelize the 3D objects, allowing the deep 3D convolutional network (3D CNN) to learn the object relationships from the complex scenes. However, the current networks lack the controllable kernels to model the object relationship across multiple views, where appropriate views provide the relevant information for suggesting the existence of the occluded objects. In this paper, we propose Cross-View Synthesis Transformer (CVSformer), which consists of Multi-View Feature Synthesis and Cross-View Transformer for learning cross-view object relationships. In the multi-view feature synthesis, we use a set of 3D convolutional kernels rotated differently to compute the multi-view features for each voxel. In the cross-view transformer, we employ the cross-view fusion to comprehensively learn the cross-view relationships, which form useful information for enhancing the features of individual views. We use the enhanced features to predict the geometric occupancies and semantic labels of all voxels. We evaluate CVSformer on public datasets, where CVS-former yields state-of-the-art results. Our code is available at https://github.com/donghaotian123/CVSformer. Haotian Dong, Enhui Ma, Lubo Wang, Miaohui Wang, Wuyuan Xie, Qing Guo 0005, Ping Li 0016, Lingyu Liang, Kairui Yang, Di Lin 0002 |
ICCV | 10 |
| 2023 | Leveraging Inpainting for Single-Image Shadow RemovalabstractFully-supervised shadow removal methods achieve the best restoration qualities on public datasets but still generate some shadow remnants. One of the reasons is the lack of large-scale shadow & shadow-free image pairs. Unsupervised methods can alleviate the issue but their restoration qualities are much lower than those of fully-supervised methods. In this work, we find that pretraining shadow removal networks on the image inpainting dataset can reduce the shadow remnants significantly: a naive encoder-decoder network gets competitive restoration quality w.r.t. the state-of-the-art methods via only 10% shadow & shadow-free image pairs. After analyzing networks with/without inpainting pretraining via the information stored in the weight (IIW), we find that inpainting pretraining improves restoration quality in non-shadow regions and enhances the generalization ability of networks significantly. Additionally, shadow removal fine-tuning enables networks to fill in the details of shadow regions. Inspired by these observations we formulate shadow removal as an adaptive fusion task that takes advantage of both shadow removal and image inpainting. Specifically, we develop an adaptive fusion network consisting of two encoders, an adaptive fusion block, and a decoder. The two encoders are responsible for extracting the features from the shadow image and the shadow-masked image respectively. The adaptive fusion block is responsible for combining these features in an adaptive manner. Finally, the decoder converts the adaptive fused features to the desired shadow-free result. The extensive experiments show that our method empowered with inpainting outperforms all state-of-the-art methods. We have realized codes and models in https://github.com/tsingqguo/inpaint4shadow Qing Guo 0005, Rabab Abdelfattah, Di Lin 0002, Wei Feng 0005, Ivor W. Tsang, Song Wang 0002 |
ICCV | 4 |
| 2023 | Open Compound Domain Adaptation with Object Style Compensation for Semantic SegmentationabstractMany methods of semantic image segmentation have borrowed the success of open compound domain adaptation. They minimize the style gap between the images of source and target domains, more easily predicting the accurate pseudo annotations for target domain's images that train segmentation network. The existing methods globally adapt the scene style of the images, whereas the object styles of different categories or instances are adapted improperly. This paper proposes the Object Style Compensation, where we construct the Object-Level Discrepancy Memory with multiple sets of discrepancy features. The discrepancy features in a set capture the style changes of the same category's object instances adapted from target to source domains. We learn the discrepancy features from the images of source and target domains, storing the discrepancy features in memory. With this memory, we select appropriate discrepancy features for compensating the style information of the object instances of various categories, adapting the object styles to a unified style of source domain. Our method enables a more accurate computation of the pseudo annotations for target domain's images, thus yielding state-of-the-art results on different datasets. Tingliang Feng, Xueyang Liu, Wei Feng 0005, Di Lin 0002 |
NeurIPS | 7 |
| 2023 | Phase-based fine-grained change detection
Xuzhi Wang, Di Lin 0002, Wei Feng 0005 |
Expert Syst. Appl. | 3 |
| 2023 | Tampering localization and self-recovery using block labeling and adaptive significance
Xiaochen Yuan, Tong Liu 0021, Chan-Tong Lam, Guoheng Huang, Di Lin 0002, Ping Li 0016 |
Expert Syst. Appl. | 6 |
| 2023 | TAGNet: Learning Configurable Context Pathways for Semantic SegmentationabstractState-of-the-art semantic segmentation methods capture the relationship between pixels to facilitate contextual information exchange. Advanced methods utilize fixed pathways for context exchange, lacking the flexibility to harness the most relevant context for each pixel. In this paper, we present Configurable Context Pathways (CCPs), a novel model for establishing pathways for augmenting contextual information. In contrast to previous pathway models, CCPs are learned, leveraging configurable regions to form information flows between pairs of pixels. We propose TAGNet to adaptively configure the regions, which span over the entire image space, driven by the relationships between the remote pixels. Subsequently, the information flows along the pathways are updated gradually by the information provided by sequences of configurable regions, forming more powerful contextual information. We extensively evaluate the traveling, adaption, and gathering (TAG) stages of our network on the public benchmarks, demonstrating that all of the stages successfully improve the segmentation accuracy and help to surpass the state-of-the-art results. The code package is available at: https://github.com/dilincv/TAGNet. Di Lin 0002, Dingguo Shen, Yuanfeng Ji, Siting Shen, Mingrui Xie, Wei Feng 0005, Hui Huang 0004 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | Surface Geometry Processing: An Efficient Normal-Based Detail RepresentationabstractWith the rapid development of high-resolution 3D vision applications, the traditional way of manipulating surface detail requires considerable memory and computing time. To address these problems, we introduce an efficient surface detail processing framework in 2D normal domain, which extracts new normal feature representations as the carrier of micro geometry structures that are illustrated both theoretically and empirically in this article. Compared with the existing state of the arts, we verify and demonstrate that the proposed normal-based representation has three important properties, including detail separability, detail transferability and detail idempotence. Finally, three new schemes are further designed for geometric surface detail processing applications, including geometric texture synthesis, geometry detail transfer, and 3D surface super-resolution. Theoretical analysis and experimental results on the latest benchmark dataset verify the effectiveness and versatility of our normal-based representation, which accepts 30 times of the input surface vertices but at the same time only takes 6.5% memory cost and 14.0% running time in comparison with existing competing algorithms. Wuyuan Xie, Miaohui Wang, Di Lin 0002, Boxin Shi, Jianmin Jiang |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Coarse-to-Fine: Progressive Knowledge Transfer-Based Multitask Convolutional Neural Network for Intelligent Large-Scale Fault DiagnosisabstractIn modern industry, large-scale fault diagnosis of complex systems is emerging and becoming increasingly important. Most deep learning-based methods perform well on small number of fault diagnosis, but cannot converge to satisfactory results when handling large-scale fault diagnosis because the huge number of fault types will lead to the problems of intra/inter-class distance unbalance and poor local minima in neural networks. To address the above problems, a progressive knowledge transfer-based multitask convolutional neural network (PKT-MCNN) is proposed. First, to construct the coarse-to-fine knowledge structure intelligently, a structure learning algorithm is proposed via clustering fault types in different coarse-grained nodes. Thus, the intra/inter-class distance unbalance problem can be mitigated by spreading similar tasks into different nodes. Then, an MCNN architecture is designed to learn the coarse and fine-grained task simultaneously and extract more general fault information, thereby pushing the algorithm away from poor local minima. Last but not least, a PKT algorithm is proposed, which can not only transfer the coarse-grained knowledge to the fine-grained task and further alleviate the intra/inter-class distance unbalance in feature space, but also regulate different learning stages by adjusting the attention weight to each task progressively. To verify the effectiveness of the proposed method, a dataset of a nuclear power system with 66 fault types was collected and analyzed. The results demonstrate that the proposed method can be a promising tool for large-scale fault diagnosis. Yu Wang 0106, Di Lin 0002, Ping Li 0016, Qinghua Hu, C. L. Philip Chen |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2022 | MISF: Multi-level Interactive Siamese Filtering for High-Fidelity Image InpaintingabstractAlthough achieving significant progress, existing deep generative inpainting methods still show low generalization across different scenes. As a result, the generated images usually contain artifacts or the filled pixels differ greatly from the ground truth, making them far from real-world applications. Image-level predictive filtering is a widely used restoration technique by predicting suitable kernels adaptively according to different input scenes. Inspired by this inherent advantage, we explore the possibility of addressing image inpainting as a filtering task. To this end, we first study the advantages and challenges of the image-level predictive filtering for inpainting: the method can preserve local structures and avoid artifacts but fails to fill large missing areas. Then, we propose the semantic filtering by conducting filtering on deep feature level, which fills the missing semantic information but fails to recover the details. To address the issues while adopting the respective advantages, we propose a novel filtering technique, i.e., Multi-level Interactive Siamese Filtering (MISF) containing two branches: kernel prediction branch (KPB) and semantic & image filtering branch (SIFB). These two branches are interactively linked: SIFB provides multi-level features for KPB while KPB predicts dynamic kernels for SIFB. As a result, the final method takes the advantage of effective semantic & image-level filling for high-fidelity inpainting. Moreover, we discuss the relationship between MISF and the naive encoder-decoder-based inpainting, inferring that MISF provides novel dynamic convolutional operations to enhance the high generalization capability across scenes. We validate our method on three challenging datasets, i.e., Dunhuang, Places2, and CelebA. Our method outperforms state-of-the-art baselines on four metrics, i.e.,$L_{1}$, PSNR, SSIM, and LPIPS. Qing Guo 0005, Di Lin 0002, Ping Li 0016, Wei Feng 0005, Song Wang 0002 |
CVPR | 3 |
| 2022 | Cross-Image Context for Single Image InpaintingabstractVisual context is of crucial importance for image inpainting. The contextual information captures the appearance and semantic correlation between the image regions, helping to propagate the information of the complete regions for reasoning the content of the corrupted regions. Many inpainting methods compute the visual context based on the regions within the single image. In this paper, we propose the Cross-Image Context Memory (CICM) for learning and using the cross-image context to recover the corrupted regions. CICM consists of multiple sets of the cross-image representations learned from the image regions with different visual patterns. The regional representations are learned across different images, thus providing richer context that benefit the inpainting task. The experimental results demonstrate the effectiveness and generalization of CICM, which achieves state-of-the-art performances on various datasets for single image inpainting. Tingliang Feng, Wei Feng 0005, Di Lin 0002 |
NeurIPS | 4 |
| 2022 | Generative Status Estimation and Information Decoupling for Image Rain RemovalabstractImage rain removal requires the accurate separation between the pixels of the rain streaks and object textures. But the confusing appearances of rains and objects lead to the misunderstanding of pixels, thus remaining the rain streaks or missing the object details in the result. In this paper, we propose SEIDNet equipped with the generative Status Estimation and Information Decoupling for rain removal. In the status estimation, we embed the pixel-wise statuses into the status space, where each status indicates a pixel of the rain or object. The status space allows sampling multiple statuses for a pixel, thus capturing the confusing rain or object. In the information decoupling, we respect the pixel-wise statuses, decoupling the appearance information of rain and object from the pixel. Based on the decoupled information, we construct the kernel space, where multiple kernels are sampled for the pixel to remove the rain and recover the object appearance. We evaluate SEIDNet on the public datasets, achieving state-of-the-art performances of image rain removal. The experimental results also demonstrate the generalization of SEIDNet, which can be easily extended to achieve state-of-the-art performances on other image restoration tasks (e.g., snow, haze, and shadow removal). Di Lin 0002, Xin Wang 0118, Miaohui Wang, Wuyuan Xie, Qing Guo 0005, Ping Li 0016 |
NeurIPS | 1 |
| 2022 | Fast and robust active camera relocalization in the wild for fine-grained change detection
Qian Zhang 0051, Wei Feng 0005, Yi-Bo Shi, Di Lin 0002 |
Neurocomputing | 4 |
| 2022 | Deep LSAC for Fine-Grained RecognitionabstractFine-grained recognition emphasizes the identification of subtle differences among object categories given objects that appear in different shapes and poses. These variances should be reduced for reliable recognition. We propose a fine-grained recognition system that incorporates localization, segmentation, alignment, and classification in a unified deep neural network. The input to the classification module includes functions that enable backward-propagation (BP) in constructing the solver. Our major contribution is to propose a valve linkage function (VLF) for BP chaining and form our deep localization, segmentation, alignment, and classification (LSAC) system. The VLF can adaptively compromise errors of classification and alignment when training the LSAC model. It in turn helps to update the localization and segmentation. We evaluate our framework on two widely used fine-grained object data sets. The performance confirms the effectiveness of our LSAC system. Di Lin 0002, Yi Wang 0031, Lingyu Liang, Ping Li 0016, C. L. Philip Chen |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2021 | Reciprocal Learning for Semi-supervised Segmentation
Xiangyun Zeng, Rian Huang, Yuming Zhong, Chu Han, Di Lin 0002, Dong Ni 0001, Yi Wang 0031 |
MICCAI (2) | 6 |
| 2020 | RANet: Region Attention Network for Semantic SegmentationabstractRecent semantic segmentation methods model the relationship between pixels to construct the contextual representations. In this paper, we introduce the \emph{Region Attention Network} (RANet), a novel attention network for modeling the relationship between object regions. RANet divides the image into object regions, where we select representative information. In contrast to the previous methods, RANet configures the information pathways between the pixels in different regions, enabling the region interaction to exchange the regional context for enhancing all of the pixels in the image. We train the construction of object regions, the selection of the representative regional contents, the configuration of information pathways and the context exchange between pixels, jointly, to improve the segmentation accuracy. We extensively evaluate our method on the challenging segmentation benchmarks, demonstrating that RANet effectively helps to achieve the state-of-the-art results. Dingguo Shen, Yuanfeng Ji, Ping Li 0016, Yi Wang 0031, Di Lin 0002 |
NeurIPS | 5 |
| 2020 | Zig-Zag Network for Semantic Segmentation of RGB-D ImagesabstractSemantic segmentation of images requires an understanding of appearances of objects and their spatial relationships in scenes. The fully convolutional network (FCN) has been successfully applied to recognize objects' appearances, which are represented with RGB channels. Images augmented with depth channels provide more understanding of the geometric information of the scene in an image. In this paper, we present a multiple-branch neural network to utilize depth information to assist in the semantic segmentation of images. Our approach splits the image into layers according to the "scene-scale". We introduce the context-aware receptive field (CARF), which provides better control of the relevant context information of learned features. Each branch of the network is equipped with CARF to adaptively aggregate the context information of image regions, leading to a more focused domain that is easier to learn. Furthermore, we propose a new zig-zag architecture to exchange information between the feature maps at different levels, augmented by the CARFs of the backbone network and decoder network. With the flexible information propagation allowed by our zig-zag network, we enrich the context information of feature maps for the segmentation. We show that the zig-zag network achieves state-of-the-art performances on several public datasets. Di Lin 0002, Hui Huang 0004 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2020 | SCN: Switchable Context Network for Semantic Segmentation of RGB-D ImagesabstractContext representations have been widely used to profit semantic image segmentation. The emergence of depth data provides additional information to construct more discriminating context representations. Depth data preserves the geometric relationship of objects in a scene, which is generally hard to be inferred from RGB images. While deep convolutional neural networks (CNNs) have been successful in solving semantic segmentation, we encounter the problem of optimizing CNN training for the informative context using depth data to enhance the segmentation accuracy. In this paper, we present a novel switchable context network (SCN) to facilitate semantic segmentation of RGB-D images. Depth data is used to identify objects existing in multiple image regions. The network analyzes the information in the image regions to identify different characteristics, which are then used selectively through switching network branches. With the content extracted from the inherent image structure, we are able to generate effective context representations that are aware of both image structures and object relationships, leading to a more coherent learning of semantic segmentation network. We demonstrate that our SCN outperforms state-of-the-art methods on two public datasets. Di Lin 0002, Ruimao Zhang, Yuanfeng Ji, Ping Li 0016, Hui Huang 0004 |
IEEE Trans. Cybern. | 1 |
| 2019 | ZigZagNet: Fusing Top-Down and Bottom-Up Context for Object SegmentationabstractMulti-scale context information has proven to be essential for object segmentation tasks. Recent works construct the multi-scale context by aggregating convolutional feature maps extracted by different levels of a deep neural network. This is typically done by propagating and fusing features in a one-directional, top-down and bottom-up, manner. In this work, we introduce ZigZagNet, which aggregates a richer multi-context feature map by using not only dense top-down and bottom-up propagation, but also by introducing pathways crossing between different levels of the top-down and the bottom-up hierarchies, in a zig-zag fashion. Furthermore, the context information is exchanged and aggregated over multiple stages, where the fused feature maps from one stage are fed into the next one, yielding a more comprehensive context for improved segmentation performance. Our extensive evaluation on the public benchmarks demonstrates that ZigZagNet surpasses the state-of-the-art accuracy for both semantic segmentation and instance segmentation tasks. Di Lin 0002, Dingguo Shen, Siting Shen, Yuanfeng Ji, Dani Lischinski, Daniel Cohen-Or, Hui Huang 0004 |
CVPR | 1 |
| 2019 | PRSNet: Part Relation and Selection Network for Bone Age Assessment
Yuanfeng Ji, Hao Chen 0011, Dan Lin 0009, Di Lin 0002 |
MICCAI (6) | 5 |
| 2019 | Multiview-coherent disocclusion synthesis using connected regions optimizationabstractAbstract Handling of missing areas is a key step for depth‐based rendering to synthesize virtual views. Existing methods usually consider finding candidate pixels from only one reference view to fill missing areas. However, the information provided by one reference view is restricted by the position of the view. By utilizing two reference views located on both the left and right sides of the virtual view, we propose to synthesize the missing areas at hole level with connected regions optimization. To avoid the appearance of ghost boundary, we apply morphological operations to generate a boundary band map for the depth map, which restricts the warping of the virtual view. We use a binary map to mark unknown pixels in a hole, label the connected unknown regions, and count the area of each connected region, which decide the order in our enhanced inpainting synthesis. Besides, we separate the foreground and background regions of the depth map to constrain the searching of candidate pixels. Multiple experiments on virtual view synthesis have shown the effectiveness and high quality of our multiview‐coherent disocclusion synthesis. Ping Li 0016, Yuxi Jin, Bin Sheng 0001, Di Lin 0002, Yongwei Nie, Enhua Wu |
Comput. Animat. Virtual Worlds | 4 |
| 2019 | SAGNet: structure-aware generative network for 3D-shape modelingabstractWe present SAGNet, a structure-aware generative model for 3D shapes. Given a set of segmented objects of a certain class, the geometry of their parts and the pairwise relationships between them (the structure) are jointly learned and embedded in a latent space by an autoencoder. The encoder intertwines the geometry and structure features into a single latent code, while the decoder disentangles the features and reconstructs the geometry and structure of the 3D model. Our autoencoder consists of two branches, one for the structure and one for the geometry. The key idea is that during the analysis, the two branches exchange information between them, thereby learning the dependencies between structure and geometry and encoding two augmented features, which are then fused into a single latent code. This explicit intertwining of information enables separately controlling the geometry and the structure of the generated models. We evaluate the performance of our method and conduct an ablation study. We explicitly show that encoding of shapes accounts for both similarities in structure and geometry. A variety of quality results generated by SAGNet are presented. Di Lin 0002, Dani Lischinski, Daniel Cohen-Or, Hui Huang 0004 |
ACM Trans. Graph. | 3 |
| 2018 | Multi-scale Context Intertwining for Semantic Segmentation
Di Lin 0002, Yuanfeng Ji, Dani Lischinski, Daniel Cohen-Or, Hui Huang 0004 |
ECCV (3) | 1 |
| 2018 | Semantic object reconstruction via casual handheld scanningabstractWe introduce a learning-based method to reconstruct objects acquired in a casual handheld scanning setting with a depth camera. Our method is based on two core components. First, a deep network that provides a semantic segmentation and labeling of the frames of an input RGBD sequence. Second, an alignment and reconstruction method that employs the semantic labeling to reconstruct the acquired object from the frames. We demonstrate that the use of a semantic labeling improves the reconstructions of the objects, when compared to methods that use only the depth information of the frames. Moreover, since training a deep network requires a large amount of labeled data, a key contribution of our work is an active self-learning framework to simplify the creation of the training data. Specifically, we iteratively predict the labeling of frames with the neural network, reconstruct the object from the labeled frames, and evaluate the confidence of the labeling, to incrementally train the neural network while requiring only a small amount of user-provided annotations. We show that this method enables the creation of data for training a neural network with high accuracy, while requiring only little manual effort. Ruizhen Hu, Oliver van Kaick, Luanmin Chen, Di Lin 0002, Daniel Cohen-Or, Hui Huang 0004 |
ACM Trans. Graph. | 5 |
| 2017 | Cascaded Feature Network for Semantic Segmentation of RGB-D ImagesabstractFully convolutional network (FCN) has been successfully applied in semantic segmentation of scenes represented with RGB images. Images augmented with depth channel provide more understanding of the geometric information of the scene in the image. The question is how to best exploit this additional information to improve the segmentation performance. In this paper, we present a neural network with multiple branches for segmenting RGB-D images. Our approach is to use the available depth to split the image into layers with common visual characteristic of objects/scenes, or common “scene-resolution”. We introduce context-aware receptive field (CaRF) which provides a better control on the relevant contextual information of the learned features. Equipped with CaRF, each branch of the network semantically segments relevant similar scene-resolution, leading to a more focused domain which is easier to learn. Furthermore, our network is cascaded with features from one branch augmenting the features of adjacent branch. We show that such cascading of features enriches the contextual information of each branch and enhances the overall performance. The accuracy that our network achieves outperforms the state-of-the-art methods on two public datasets. Di Lin 0002, Guangyong Chen, Daniel Cohen-Or, Pheng-Ann Heng, Hui Huang 0004 |
ICCV | 1 |
| 2017 | Learning to Aggregate Ordinal Labels by Maximizing Separating WidthabstractWhile crowdsourcing has been a cost and time efficient method to label massive samples, one critical issue is quality control, for which the key challenge is to infer the ground truth from noisy or even adversarial data by various users. A large class of crowdsourcing problems, such as those involving age, grade, level, or stage, have an ordinal structure in their labels. Based on a technique of sampling estimated label from the posterior distribution, we define a novel separating width among the labeled observations to characterize the quality of sampled labels, and develop an efficient algorithm to optimize it through solving multiple linear decision boundaries and adjusting prior distributions. Our algorithm is empirically evaluated on several real world datasets, and demonstrates its supremacy over state-of-the-art methods. Guangyong Chen, Shengyu Zhang 0002, Di Lin 0002, Hui Huang 0004, Pheng-Ann Heng |
ICML | 3 |
| 2017 | Two-Class Weather ClassificationabstractGiven a single outdoor image, we propose a collaborative learning approach using novel weather features to label the image as either sunny or cloudy. Though limited, this two-class classification problem is by no means trivial given the great variety of outdoor images captured by different cameras where the images may have been edited after capture. Our overall weather feature combines the data-driven convolutional neural network (CNN) feature and well-chosen weather-specific features. They work collaboratively within a unified optimization framework that is aware of the presence (or absence) of a given weather cue during learning and classification. In this paper we propose a new data augmentation scheme to substantially enrich the training data, which is used to train a latent SVM framework to make our solution insensitive to global intensity transfer. Extensive experiments are performed to verify our method. Compared with our previous work and the sole use of a CNN classifier, this paper improves the accuracy up to 7-8 percent. Our weather image dataset is available together with the executable of our classifier. Cewu Lu, Di Lin 0002, Jiaya Jia, Chi-Keung Tang |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2017 | RSCM: Region Selection and Concurrency Model for Multi-Class Weather RecognitionabstractToward weather condition recognition, we emphasize the importance of regional cues in this paper and address a few important problems regarding appropriate representation, its differentiation among regions, and weather-condition feature construction. Our major contribution is, first, to construct a multi-class benchmark data set containing 65 000 images from six common categories for sunny, cloudy, rainy, snowy, haze, and thunder weather. This data set also benefits weather classification and attribute recognition. Second, we propose a deep learning framework named region selection and concurrency model (RSCM) to help discover regional properties and concurrency. We evaluate RSCM on our multi-class benchmark data and another public data set for weather recognition. Di Lin 0002, Cewu Lu, Hui Huang 0004, Jiaya Jia |
IEEE Trans. Image Process. | 1 |
| 2016 | ScribbleSup: Scribble-Supervised Convolutional Networks for Semantic SegmentationabstractLarge-scale data is of crucial importance for learning semantic segmentation models, but annotating per-pixel masks is a tedious and inefficient procedure. We note that for the topic of interactive image segmentation, scribbles are very widely used in academic research and commercial software, and are recognized as one of the most userfriendly ways of interacting. In this paper, we propose to use scribbles to annotate images, and develop an algorithm to train convolutional networks for semantic segmentation supervised by scribbles. Our algorithm is based on a graphical model that jointly propagates information from scribbles to unmarked pixels and learns network parameters. We present competitive object semantic segmentation results on the PASCAL VOC dataset by using scribbles as annotations. Scribbles are also favored for annotating stuff (e.g., water, sky, grass) that has no well-defined shape, and our method shows excellent results on the PASCALCONTEXT dataset thanks to extra inexpensive scribble annotations. Our scribble annotations on PASCAL VOC are available at http://research.microsoft.com/en-us/um/ people/jifdai/downloads/scribble_sup. Di Lin 0002, Jifeng Dai, Jiaya Jia, Kaiming He, Jian Sun 0001 |
CVPR | 1 |
| 2015 | Deep LAC: Deep localization, alignment and classification for fine-grained recognitionabstractWe propose a fine-grained recognition system that incorporates part localization, alignment, and classification in one deep neural network. This is a nontrivial process, as the input to the classification module should be functions that enable back-propagation in constructing the solver. Our major contribution is to propose a valve linkage function (VLF) for back-propagation chaining and form our deep localization, alignment and classification (LAC) system. The VLF can adaptively compromise the errors of classification and alignment when training the LAC model. It in turn helps update localization. The performance on fine-grained object data bears out the effectiveness of our LAC system. Di Lin 0002, Xiaoyong Shen, Cewu Lu, Jiaya Jia |
CVPR | 1 |
| 2014 | Learning Important Spatial Pooling Regions for Scene ClassificationabstractWe address the false response influence problem when learning and applying discriminative parts to construct the mid-level representation in scene classification. It is often caused by the complexity of latent image structure when convolving part filters with input images. This problem makes mid-level representation, even after pooling, not distinct enough to classify input data correctly to categories. Our solution is to learn important spatial pooling regions along with their appearance. The experiments show that this new framework suppresses false response and produces improved results on several datasets, including MIT-Indoor, 15-Scene, and UIUC 8-Sport. When combined with global image features, our method achieves state-of-the-art performance on these datasets. Di Lin 0002, Cewu Lu, Renjie Liao 0001, Jiaya Jia |
CVPR | 1 |
| 2014 | Two-Class Weather ClassificationabstractGiven a single outdoor image, this paper proposes a collaborative learning approach for labeling it as either sunny or cloudy. Never adequately addressed, this twoclass classification problem is by no means trivial given the great variety of outdoor images. Our weather feature combines special cues after properly encoding them into feature vectors. They then work collaboratively in synergy under a unified optimization framework that is aware of the presence (or absence) of a given weather cue during learning and classification. Extensive experiments and comparisons are performed to verify our method. We build a new weather image dataset consisting of 10K sunny and cloudy images, which is available online together with the executable. Cewu Lu, Di Lin 0002, Jiaya Jia, Chi-Keung Tang |
CVPR | 2 |