Si Li 0001

dblp:54/6603-1 · DBLP profile ↗
← Back
67ranked-venue papers
3as first author
48since 2021 · last 2026
0000-0001-9823-3870ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 50 · 1 first-author · 37 since 2021Graphics, computer vision, multimedia, augmented reality and games · 28 · 1 first-author · 23 since 2021Databases, data management, data science and information retrieval · 7 · 2 first-author · 3 since 2021
YearPublicationVenuePosition
2026 ReContraster: Making Your Posters Stand Out with Regional Contrast
abstract
Effective poster design requires rapidly capturing attention and clearly conveying messages.Inspired by the "contrast effects" principle, we propose ReContraster, the first training-free model to leverage regional contrast to make posters stand out.By emulating the cognitive behaviors of a poster designer, ReContraster introduces the compositional multi-agent system to identify elements, organize layout, and evaluate generated poster candidates.To further ensure harmonious transitions across region boundaries, ReContraster integrates the hybrid denoising strategy during the diffusion process.We additionally contribute a new benchmark dataset for comprehensive evaluation.Seven quantitative metrics and four user studies confirm its superiority over relevant state-of-the-art methods, producing visually striking and aesthetically appealing posters.
Peixuan Zhang, Zijian Jia, Ziqi Cai, Shuchen Weng, Si Li 0001, Boxin Shi
ACL (1)5
2026 L-VOCAL: Language-based Video Colorization with Audio Alignment
Shuchen Weng, Huan Ouyang, Yuchen Hong, Lihan Lin, Si Li 0001, Boxin Shi
Int. J. Comput. Vis.6
2026 Affective Image Editing: Shaping Emotional Factors via Text Descriptions
Peixuan Zhang, Shuchen Weng, Chengxuan Zhu, Binghao Tang, Zijian Jia, Si Li 0001, Boxin Shi
Int. J. Comput. Vis.6
2026 L-C4: Language-based video colorization for creative and consistent color
Shuchen Weng, Huan Ouyang, Lihan Lin, Yu Li 0003, Si Li 0001, Boxin Shi
Neurocomputing6
2026 Toward Deeper Emotional Reflection: Crafting Affective Image Filters With Generative Priors
abstract
Social media platforms enable users to express emotions by posting text with accompanying images. In this paper, we propose the Affective Image Filter (AIF) task, which aims to reflect visually-abstract emotionsfrom text into visually-concrete images, thereby creating emotionally compelling results. We first introduce the AIF dataset and the formulation of the AIF models. Then, we present AIF-B as an initial attempt based on a multi-modal transformer architecture. After that, we propose AIF-D as an extension of AIF-B towards deeper emotional reflection, effectively leveraging generative priors from pre-trained large-scale diffusion models. Quantitative and qualitative experiments demonstrate that AIF models achieve superior performance for both content consistency and emotional fidelity compared to state-of-the-art methods. Extensive user study experiments demonstrate that AIF models are significantly more effective at evoking specific emotions. Based on the presented results, we comprehensively discuss the value and potential of AIF models.
Peixuan Zhang, Shuchen Weng, Jiajun Tang 0001, Si Li 0001, Boxin Shi
IEEE Trans. Pattern Anal. Mach. Intell.4
2025 VIRES: Video Instance Repainting via Sketch and Text Guided Generation
abstract
We introduce VIRES, a video instance repainting method with sketch and text guidance, enabling video instance repainting, replacement, generation, and removal. Existing approaches struggle with temporal consistency and accurate alignment with the provided sketch sequence. VIRES leverages the generative priors of text-to-video models to maintain temporal consistency and produce visually pleasing results. We propose the Sequential ControlNet with the standardized self-scaling, which effectively extracts structure layouts and adaptively captures high-contrast sketch details. We further augment the diffusion transformer backbone with the sketch attention to interpret and inject fine-grained sketch semantics. A sketch-aware encoder ensures that repainted results are aligned with the provided sketch sequence. Additionally, we contribute the VIRESET, a dataset with detailed annotations tailored for training and evaluating video instance editing methods. Experimental results demonstrate the effectiveness of VIRES, which outperforms state-of-the-art methods in visual quality, temporal consistency, condition alignment, and human ratings. The code, dataset and pretrained models are available at: https://hjzheng.net/projects/VIRES.
Shuchen Weng, Haojie Zheng, Peixuan Zhang, Yuchen Hong, Si Li 0001, Boxin Shi
CVPR6
2025 PolGS: Polarimetric Gaussian Splatting for Fast Reflective Surface Reconstruction
Yufei Han 0002, Bowen Tie, Heng Guo 0003, Youwei Lyu, Si Li 0001, Boxin Shi, Zhanyu Ma
ICCV5
2025 PolarAnything: Diffusion-based Polarimetric Image Synthesis
abstract
Polarization images facilitate image enhancement and 3D reconstruction tasks, but the limited accessibility of polarization cameras hinders their broader application. This gap drives the need for synthesizing photorealistic polarization images. The existing polarization simulator Mitsuba relies on a parametric polarization image formation model and requires extensive 3D assets covering shape and PBR materials, preventing it from generating large-scale photorealistic images. To address this problem, we propose PolarAnything, capable of synthesizing polarization images from a single RGB input with both photorealism and physical accuracy, eliminating the dependency on 3D asset collections. Drawing inspiration from the zero-shot performance of pretrained diffusion models, we introduce a diffusion-based generative framework with an effective representation strategy that preserves the fidelity of polarization properties. Experiments show that our model generates high-quality polarization images and supports downstream tasks like shape from polarization.
Kailong Zhang, Youwei Lyu, Heng Guo 0003, Si Li 0001, Zhanyu Ma, Boxin Shi
ICCV4
2025 DF-Net: A Dual Fusion Network for Accurate Video Temporal Grounding
abstract
Video Temporal Grounding (VTG) involves locating the start and end times of a video clip based on a given textual query. Existing methods face challenges including temporal boundary localization bias and class imbalance in points classification. To alleviate these, we propose a Dual Fusion Network (DF-Net) in which cross-modal fusion processes occur in the joint encoder and decoder. In the encoder, we design a multi-step sampling strategy and knowledge aggregation module to extract and aggregate features from central frames and neighbors and undergo the first cross-model fusion. In decoding, we employ a transformer decoder to design the second cross-modal fusion and propose a novel Kullback-Leibler (KL) divergence-based objective function combined with focal and distance-based regression losses. Experiments on short and long video benchmarks demonstrate significant performance gains.
Haolong Yan, Binghao Tang, Boda Lin, Si Li 0001
ICME5
2025 Audio-Sync Video Generation with Multi-Stream Temporal Control
abstract
Audio is inherently temporal and closely synchronized with the visual world, making it a naturally aligned and expressive control signal for controllable video generation (e.g., movies). Beyond control, directly translating audio into video is essential for understanding and visualizing rich audio narratives (e.g., Podcasts or historical recordings). However, existing approaches fall short in generating high-quality videos with precise audio-visual synchronization, especially across diverse and complex audio types. In this work, we introduce MTV, a versatile framework for audio-sync video generation. MTV explicitly separates audios into speech, effects, and music tracks, enabling disentangled control over lip motion, event timing, and visual mood, respectively—resulting in fine-grained and semantically aligned video generation. To support the framework, we additionally present DEMIX, a dataset comprising high-quality cinematic videos and demixed audio tracks. DEMIX is structured into five overlapped subsets, enabling scalable multi-stage training for diverse generation scenarios. Extensive experiments demonstrate that MTV achieves state-of-the-art performance across six standard metrics spanning video quality, text-video consistency, and audio-video alignment.
Shuchen Weng, Haojie Zheng, Si Li 0001, Boxin Shi
NeurIPS4
2025 Bi-directional dual contrastive adapting method for alleviating hallucination in visual question answering
Haolong Yan, Binghao Tang, Boda Lin, Yanxian Bi, Si Li 0001
Expert Syst. Appl.7
2025 DMR2G: diffusion model for radiology report generation
Huan Ouyang, Binghao Tang, Si Li 0001
Multim. Tools Appl.4
2025 OpenCIR: Conditional Image Repainting With Open Condition Mixture
abstract
In this paper, we introduce OpenCIR, a fully-functional Conditional Image Repainting (CIR) model designed for local image editing. Given an image and a combination of conditions related to geometry, texture, and color, CIR models are required to repaint instances and seamlessly composite them with the original images. Previous CIR models suffer from limited object categories, restricted condition modalities, and demanded geometry precision. In contrast, leveraging the generative priors from pre-trained models, OpenCIR could repaint open object categories. Equipped with redesigned condition injection modules and the condition extension strategy, OpenCIR is able to understand open condition modalities. Adopting the contour refinement strategy, OpenCIR allows users to specify instances with open geometry precision. In addition, we contribute the Open-CIR dataset, which includes detailed annotations, tailored for the comprehensive training and evaluation of the OpenCIR model. Extensive experiments demonstrate that OpenCIR outperforms relevant state-of-the-art methods, achieving superior visual quality, and more favorable results by human evaluators.
Shuchen Weng, Xiaocheng Gong, Haojie Zheng, Si Li 0001, Boxin Shi
IEEE Trans. Pattern Anal. Mach. Intell.5
2024 DP-FedFace: Privacy-Preserving Facial Recognition in Real Federated Scenarios
abstract
Advanced deep learning-based face recognition models require extensive datasets for optimal performance. However, increasing privacy concerns drive the limitation of face image access on devices to prevent personal information leaks. To address this, federated learning, which allows decentralized data collaboration, has gained popularity. However, traditional federated learning methods risk privacy by transmitting identity proxies to servers. We propose DP-FedFace, a privacy framework specifically designed for a realistic scenario where each client contains only the owner's face images (one identity per client). It uses the difference between human and model perception to eliminate visualization-critical low-frequency components, thus protecting user privacy. We also introduce a novel, learnable privacy cost allocation mechanism that optimizes allocation strategies and adds noise to frequency domain features. Extensive experiments demonstrate that DP-FedFace maintains high recognition accuracy while offers robust privacy protection.
Si Li 0001
CIKM2
2024 CAG: A Consistency-Adaptive Text-Image Alignment Generation for Joint Multimodal Entity-Relation Extraction
abstract
Joint Multimodal Entity-Relation Extraction (JMERE) aims to extract entity-relationship triples in texts from given image-text pairs. As a joint multimodal information extraction task, it has attracted increasing research interest. Previous works of JMERE typically utilize graph networks to align textual entities and visual objects and achieve promising performance. However, these methods do not pay attention to the inconsistency between text and image and the straight alignment could limit the performance of JMERE models. In this paper, we propose a Consistency-adaptive text-image Alignment Generation (CAG) framework for various text-image consistency scenarios. Specifically, we propose a Consistency Factor (CF) to measure the consistency between images and texts. We also design consistency-adaptive contrastive learning based on CF, which can reduce the impact of inconsistent visual and textual information. Additionally, we adopt JMERE-specifical instruction tuning for better entity-relationship triplet generation. Experimental results on the JMERE dataset demonstrate that our proposed CAG is effective and achieves state-of-the-art performance.
Xiaocheng Gong, Binghao Tang, Yayue Deng, Huan Ouyang, Lei Luo 0008, Yunling Feng, Bin Duan 0005, Si Li 0001
CIKM11
2024 Type-Aware Decoding Via Explicitly Aggregating Event Information for Document-Level Event Extraction
abstract
Document-level event extraction (DEE) faces two main challenges: arguments-scattering and multi-event. Although previous methods attempt to address these challenges, they overlook the interference of event-unrelated sentences during event detection and neglect the mutual interference of different event roles during argument extraction. Therefore, this paper proposes a novel Schema-based Explicitly Aggregating (SEA) model to address these limitations. SEA aggregates event information into event type and role representations, enabling the decoding of event records based on specific type-aware representations. By detecting each event based on its event type representation, SEA mitigates the interference caused by event-unrelated information. Furthermore, SEA extracts arguments for each role based on its role-aware representations, reducing mutual interference between different roles. Experimental results on the ChFinAnn and DuEE-fin datasets show that SEA outperforms the SOTA methods.
Yidong Shi, Shudong Lu, Guanting Dong 0001, Xiaocheng Gong, Si Li 0001
ICASSP8
2024 Consensus Co-teaching for Dynamically Learning with Noisy Labels
abstract
Noisy labels in datasets may cause deep neural networks (DNNs) to memorize misleading information, thereby impacting generalization performance. Therefore, learning with noisy labels holds significant practical importance. Small-loss sample selection methods often underutilize available data and constrain the model’s potential, especially in the presence of complex or high-ratio noisy labels. In this paper, we propose Consensus Co-teaching (CoCo-teaching), introducing a consensus loss that operates on all data samples for additional supervision. This promotes the convergence of the two models towards greater similarity. Furthermore, a dynamic learning scheme is implemented to resist the memorization effects of DNNs, where models progressively rely on their consensus instead of the noisy labels during training. Meanwhile, we leverage the flip semantic consistency of images to enhance model divergence. Extensive experiments demonstrate the superiority of CoCo-teaching on synthetic and real-world noisy datasets. The code is available at https://github.com/Wangwenjing520/CoCo-teaching.git.
Si Li 0001
ICME2
2024 Focusing on All Refined Attention Regions for Noisy Label Facial Expression Recognition
abstract
Noisy label Facial Expression Recognition (FER) poses greater challenges than traditional noisy label classification tasks, primarily due to inter-class similarity and annotation ambiguity. In this paper, we observe that FER models memorize noisy samples by emphasizing local features considered relevant to the noisy labels. Motivated by this insight, we propose a novel Focus All Network (FAN) to suppress the learning of noisy samples. Similar to human vision, our network repetitively directs attention to various parts of the image, achieving full attention to faces and focusing on subtle distinguishing elements within each expression category. A simple attention model effectively aggregates these finer details, honing in on the most influential discriminative components of the image and correcting potential attention bias. Moreover, the simplicity of our network makes it a plug-and-play module. Extensive experiments on multiple expression benchmarks demonstrate that our method achieves competitive results.
Si Li 0001
ICME2
2024 Leveraging Generative Large Language Models with Visual Instruction and Demonstration Retrieval for Multimodal Sarcasm Detection
abstract
Binghao Tang, Boda Lin, Haolong Yan, Si Li. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Binghao Tang, Boda Lin, Haolong Yan, Si Li 0001
NAACL-HLT4
2024 SfPUEL: Shape from Polarization under Unknown Environment Light
abstract
Shape from polarization (SfP) benefits from advancements like polarization cameras for single-shot normal estimation, but its performance heavily relies on light conditions. This paper proposes SfPUEL, an end-to-end SfP method to jointly estimate surface normal and material under unknown environment light. To handle this challenging light condition, we design a transformer-based framework for enhancing the perception of global context features. We further propose to integrate photometric stereo (PS) priors from pretrained models to enrich extracted features for high-quality normal predictions. As metallic and dielectric materials exhibit different BRDFs, SfPUEL additionally predicts dielectric and metallic material segmentation to further boost performance. Experimental results on synthetic and our collected real-world dataset demonstrate that SfPUEL significantly outperforms existing SfP and single-shot normal estimation methods. The code and dataset is available at https://github.com/YouweiLyu/SfPUEL.
Youwei Lyu, Heng Guo 0003, Kailong Zhang, Si Li 0001, Boxin Shi
NeurIPS4
2024 ELEMO: Elements Focused Emotion Recognition for Sticker Images
Boda Lin, Binghao Tang, Haolong Yan, Si Li 0001
PRCV (5)5
2024 Scalable, explainable, adaptive information extraction from structure-aware nearest neighbor
Shudong Lu, Si Li 0001, Jun Guo 0002
Neurocomputing2
2024 Multimodal summarization with modality features alignment and features filtering
Binghao Tang, Boda Lin, Si Li 0001
Neurocomputing4
2024 SPLiT: Single Portrait Lighting Estimation via a Tetrad of Face Intrinsics
abstract
This paper proposes a novel pipeline to estimate a non-parametric environment map with high dynamic range from a single human face image. Lighting-independent and -dependent intrinsic images of the face are first estimated separately in a cascaded network. The influence of face geometry on the two lighting-dependent intrinsics, diffuse shading and specular reflection, are further eliminated by distributing the intrinsics pixel-wise onto spherical representations using the surface normal as indices. This results in two representations simulating images of a diffuse sphere and a glossy sphere under the input scene lighting. Taking into account the distinctive nature of light sources and ambient terms, we further introduce a two-stage lighting estimator to predict both accurate and realistic lighting from these two representations. Our model is trained supervisedly on a large-scale and high-quality synthetic face image dataset. We demonstrate that our method allows accurate and detailed lighting estimation and intrinsic decomposition, outperforming state-of-the-art methods both qualitatively and quantitatively on real face images.
Yean Cheng, Yongjie Zhu, Si Li 0001, Gang Pan 0001, Boxin Shi
IEEE Trans. Pattern Anal. Mach. Intell.5
2023 Polarization-Aware Low-Light Image Enhancement
abstract
Polarization-based vision algorithms have found uses in various applications since polarization provides additional physical constraints. However, in low-light conditions, their performance would be severely degenerated since the captured polarized images could be noisy, leading to noticeable degradation in the degree of polarization (DoP) and the angle of polarization (AoP). Existing low-light image enhancement methods cannot handle the polarized images well since they operate in the intensity domain, without effectively exploiting the information provided by polarization. In this paper, we propose a Stokes-domain enhancement pipeline along with a dual-branch neural network to handle the problem in a polarization-aware manner. Two application scenarios (reflection removal and shape from polarization) are presented to show how our enhancement can improve their results.
Chu Zhou, Minggui Teng, Youwei Lyu, Si Li 0001, Chao Xu 0006, Boxin Shi
AAAI4
2023 L-CoIns: Language-based Colorization With Instance Awareness
abstract
Language-based colorization produces plausible colors consistent with the language description provided by the user. Recent studies introduce additional annotation to prevent color-object coupling and mismatch issues, but they still have difficulty in distinguishing instances corresponding to the same object words. In this paper, we propose a transformer-based framework to automatically aggregate similar image patches and achieve instance awareness without any additional knowledge. By applying our presented luminance augmentation and counter-color loss to break down the statistical correlation between luminance and color words, our model is driven to synthesize colors with better descriptive consistency. We further collect a dataset to provide distinctive visual characteristics and detailed language descriptions for multiple instances in the same image. Extensive experiments demonstrate our advantages of synthesizing visually pleasing and description-consistent results of instance-aware colorization.
Shuchen Weng, Peixuan Zhang, Yu Li 0003, Si Li 0001, Boxin Shi
CVPR5
2023 Complementary Intrinsics from Neural Radiance Fields and CNNs for Outdoor Scene Relighting
abstract
Relighting an outdoor scene is challenging due to the diverse illuminations and salient cast shadows. Intrinsic image decomposition on outdoor photo collections could partly solve this problem by weakly supervised labels with albedo and normal consistency from multiview stereo. With neural radiance fields (NeRF), editing the appearance code could produce more realistic results without interpreting the outdoor scene image formation explicitly. This paper proposes to complement the intrinsic estimation from volume rendering using NeRF and from inversing the photometric image formation model using convolutional neural networks (CNNs). The former produces richer and more reliable pseudo labels (cast shadows and sky appearances in addition to albedo and normal) for training the latter to predict interpretable and editable lighting parameters via a single-image prediction pipeline. We demonstrate the advantages of our method for both intrinsic image decomposition and relighting for various real outdoor scenes.
Xuanning Cui, Yongjie Zhu, Jiajun Tang 0001, Si Li 0001, Zhaofei Yu, Boxin Shi
CVPR5
2023 Affective Image Filter: Reflecting Emotions from Text to Images
abstract
Understanding the emotions in text and presenting them visually is a very challenging problem that requires a deep understanding of natural language and high-quality image synthesis simultaneously. In this work, we propose Affective Image Filter (AIF), a novel model that is able to understand the visually-abstract emotions from the text and reflect them to visually-concrete images with appropriate colors and textures. We build our model based on the multi-modal transformer architecture, which unifies both images and texts into tokens and encodes the emotional prior knowledge. Various loss functions are proposed to understand complex emotions and produce appropriate visualization. In addition, we collect and contribute a new dataset with abundant aesthetic images and emotional texts for training and evaluating the AIF model. We carefully design four quantitative metrics and conduct a user study to comprehensively evaluate the performance, which demonstrates our AIF model outperforms state-of-the-art methods and could evoke specific emotional responses from human observers.
Shuchen Weng, Peixuan Zhang, Si Li 0001, Boxin Shi
ICCV5
2023 L-CAD: Language-based Colorization with Any-level Descriptions using Diffusion Priors
abstract
Language-based colorization produces plausible and visually pleasing colors under the guidance of user-friendly natural language descriptions. Previous methods implicitly assume that users provide comprehensive color descriptions for most of the objects in the image, which leads to suboptimal performance. In this paper, we propose a unified model to perform language-based colorization with any-level descriptions. We leverage the pretrained cross-modality generative model for its robust language understanding and rich color priors to handle the inherent ambiguity of any-level descriptions. We further design modules to align with input conditions to preserve local spatial structures and prevent the ghosting effect. With the proposed novel sampling strategy, our model achieves instance-aware colorization in diverse and complex scenarios. Extensive experimental results demonstrate our advantages of effectively handling any-level descriptions and outperforming both language-based and automatic colorization methods. The code and pretrained models are available at: https://github.com/changzheng123/L-CAD.
Shuchen Weng, Peixuan Zhang, Yu Li 0003, Si Li 0001, Boxin Shi
NeurIPS5
2023 Physics-Guided Reflection Separation From a Pair of Unpolarized and Polarized Images
abstract
Undesirable reflections contained in photos taken in front of glass windows or doors often degrade visual quality of the image. Separating two layers apart benefits both human and machine perception. The polarization status of the light changes after refraction or reflection, providing more observations of the scene, which can benefit the reflection separation. Different from previous works that take three or more polarization images as input, we propose to exploit physical constraints from a pair of unpolarized and polarized images to separate reflection and transmission layers in this paper. Due to the simplified capturing setup, the system is more under-determined compared to the existing polarization-based works. In order to solve this problem, we propose to estimate the semi-reflector orientation first to make the physical image formation well-posed, and then learn to reliably separate two layers using additional networks based on both physical and numerical analysis. In addition, a motion estimation network is introduced to handle the misalignment of paired input. Quantitative and qualitative experimental results show our approach performs favorably over existing polarization and single image based solutions.
Youwei Lyu, Zhaopeng Cui, Si Li 0001, Marc Pollefeys, Boxin Shi
IEEE Trans. Pattern Anal. Mach. Intell.3
2023 Shape From Polarization With Distant Lighting Estimation
abstract
This article presents a new approach for surface normal recovery from polarization images under an unknown distant light. Polarization provides rich cues of object geometry and material, but it is also influenced by different lighting conditions. Different from previous Shape-from-Polarization (SfP) methods, which rely on handcrafted or data-driven priors, we analytically investigate the benefits of estimating distant lighting for resolving the ambiguity in normal estimation from SfP using the polarimetric Bidirectional Reflectance Distribution Function (pBRDF) based image formation model. We then propose a two-stage learning framework that first effectively exploits polarization and shading cues to estimate the reflectance and lighting information and then optimizes the initial normal as the geometric prior. Leveraging the normal prior with the polarization cues from the input images, our network further generates the surface normal with more details in the second stage. We also present a data generation pipeline derived from the pBRDF model enabling model training and create a real dataset for evaluation of SfP approaches. Extensive ablation studies show the effectiveness of our designed architecture, and our approach outperforms existing methods in quantitative and qualitative experiments on real data.
Youwei Lyu, Lingran Zhao, Si Li 0001, Boxin Shi
IEEE Trans. Pattern Anal. Mach. Intell.3
2023 Polarization Guided HDR Reconstruction via Pixel-Wise Depolarization
abstract
Taking photos with digital cameras often accompanies saturated pixels due to their limited dynamic range, and it is far too ill-posed to restore them. Capturing multiple low dynamic range images with bracketed exposures can make the problem less ill-posed, however, it is prone to ghosting artifacts caused by spatial misalignment among images. A polarization camera can capture four spatially-aligned and temporally-synchronized polarized images with different polarizer angles in a single shot, which can be used for ghost-free high dynamic range (HDR) reconstruction. However, real-world scenarios are still challenging since existing polarization-based HDR reconstruction methods treat all pixels in the same manner and only utilize the spatially-variant exposures of the polarized images (without fully exploiting the degree of polarization (DoP) and the angle of polarization (AoP) of the incoming light to the sensor, which encode abundant structural and contextual information of the scene) to handle the problem still in an ill-posed manner. In this paper, we propose a pixel-wise depolarization strategy to solve the polarization guided HDR reconstruction problem, by classifying the pixels based on their levels of ill-posedness in HDR reconstruction procedure and applying different solutions to different classes. To utilize the strategy with better generalization ability and higher robustness, we propose a network-physics-hybrid polarization-based HDR reconstruction pipeline along with a neural network tailored to it, fully exploiting the DoP and AoP. Experimental results show that our approach achieves state-of-the-art performance on both synthetic and real-world images.
Chu Zhou, Yufei Han 0002, Minggui Teng, Jin Han 0001, Si Li 0001, Chao Xu 0006, Boxin Shi
IEEE Trans. Image Process.5
2023 Reflection Removal With NIR and RGB Image Feature Fusion
abstract
Removing undesirable reflections in photographs benefits both human perceptions and downstream computer vision tasks, but it is a highly ill-posed problem based on a single RGB image. Different from RGB images, near-infrared (NIR) images captured by an active NIR camera are less likely to be affected by reflections when glass and camera planes form certain angles, while textures on objects could “vanish” in some situations. Based on this observation, we propose a cascaded reflection removal network with an image feature fusion strategy to utilize auxiliary information in active NIR images. To tackle the insufficiency of training data, we propose a data generation pipeline to approximate perceptual properties and the reflection-suppressing nature of active NIR images. We further build a dataset with synthetic and real images to facilitate the research. Experimental results show that the proposed method outperforms state-of-the-art reflection removal methods in both quantitative metrics and visual quality.
Yuchen Hong, Youwei Lyu, Si Li 0001, Boxin Shi
IEEE Trans. Multim.3
2022 L-CoDe: Language-Based Colorization Using Color-Object Decoupled Conditions
abstract
Colorizing a grayscale image is inherently an ill-posed problem with multi-modal uncertainty. Language-based colorization offers a natural way of interaction to reduce such uncertainty via a user-provided caption. However, the color-object coupling and mismatch issues make the mapping from word to color difficult. In this paper, we propose L-CoDe, a Language-based Colorization network using color-object Decoupled conditions. A predictor for object-color corresponding matrix (OCCM) and a novel attention transfer module (ATM) are introduced to solve the color-object coupling problem. To deal with color-object mismatch that results in incorrect color-object correspondence, we adopt a soft-gated injection module (SIM). We further present a new dataset containing annotated color-object pairs to provide supervisory signals for resolving the coupling problem. Experimental results show that our approach outperforms state-of-the-art methods conditioned on captions.
Shuchen Weng, Jiajun Tang 0001, Si Li 0001, Boxin Shi
AAAI5
2022 UniCoRN: A Unified Conditional Image Repainting Network
abstract
Conditional image repainting (CIR) is an advanced image editing task, which requires the model to generate visual content in user-specified regions conditioned on multiple cross-modality constraints, and composite the visual content with the provided background seamlessly. Existing methods based on two-phase architecture design assume dependency between phases and cause color-image incongruity. To solve these problems, we propose a novel Unified Conditional image Repainting Network (UniCoRN). We break the two-phase assumption in the CIR task by constructing the interaction and dependency relationship between background and other conditions. We further introduce the hierarchical structure into cross-modality similarity model to capture feature patterns at different levels and bridge the gap between visual content and color condition. A new Landscape-CIR dataset is collected and annotated to expand the application scenarios of the CIR task. Experiments show that UniCoRN achieves higher synthetic quality, better condition consistency, and more realistic compositing effect.
Jimeng Sun 0002, Shuchen Weng, Si Li 0001, Boxin Shi
CVPR4
2022 L-CoDer: Language-Based Colorization with Color-Object Decoupling Transformer
Shuchen Weng, Yu Li 0003, Si Li 0001, Boxin Shi
ECCV (18)4
2022 Estimating Spatially-Varying Lighting in Urban Scenes with Disentangled Representation
Jiajun Tang 0001, Yongjie Zhu, Jun Hoong Chan, Si Li 0001, Boxin Shi
ECCV (6)5
2022 CT2: Colorization Transformer via Color Tokens
Shuchen Weng, Jimeng Sun 0002, Yu Li 0003, Si Li 0001, Boxin Shi
ECCV (7)4
2022 AdsCVLR: Commercial Visual-Linguistic Representation Modeling in Sponsored Search
abstract
Sponsored search advertisements (ads) appear next to search results when consumers look for products and services on search engines. As the fundamental basis of search ads, relevance modeling has attracted increasing attention due to the significant research challenges and tremendous practical value. In this paper, we address the problem of multi-modal modeling in sponsored search, which models the relevance between user query and commercial ads with multi-modal structured information. To solve this problem, we propose a transformer architecture with Ads data on Commercial Visual-Linguistic Representation (AdsCVLR) with contrastive learning that naturally extends the transformer encoder with the complementary multi-modal inputs, serving as a strong aggregator of image-text features. We also make a public advertising dataset, which includes 480K labeled query-ad pairwise data with structured information of image, title, seller, description, and so on. Empirically, we evaluate the AdsCVLR model over the large industry dataset, and the experimental results of online/offline tests show the superiority of our method.
Yongjie Zhu, Chunhui Han, Yuefeng Zhan, Bochen Pang, Zhaoju Li, Hao Sun 0015, Si Li 0001, Boxin Shi, Nan Duan 0001, Ruofei Zhang, Liangjie Zhang, Qi Zhang 0066
ACM Multimedia7
2022 Event detection from text using path-aware graph convolutional network
Shudong Lu, Si Li 0001, Haibo Lan, Jun Guo 0002
Appl. Intell.2
2022 Explainable document-level event extraction via back-tracing to sentence-level event clues
Shudong Lu, Si Li 0001, Jun Guo 0002
Knowl. Based Syst.3
2022 Hybrid Face Reflectance, Illumination, and Shape From a Single Image
abstract
We propose HyFRIS-Net to jointly estimate the hybrid reflectance and illumination models, as well as the refined face shape from a single unconstrained face image in a pre-defined texture space. The proposed hybrid reflectance and illumination representation ensure photometric face appearance modeling in both parametric and non-parametric spaces for efficient learning. While forcing the reflectance consistency constraint for the same person and face identity constraint for different persons, our approach recovers an occlusion-free face albedo with disambiguated color from the illumination color. Our network is trained in a self-evolving manner to achieve general applicability on real-world data. We conduct comprehensive qualitative and quantitative evaluations with state-of-the-art methods to demonstrate the advantages of HyFRIS-Net in modeling photo-realistic face albedo, illumination, and shape.
Yongjie Zhu, Chen Li 0031, Si Li 0001, Boxin Shi, Yu-Wing Tai
IEEE Trans. Pattern Anal. Mach. Intell.3
2022 CorefDPR: A Joint Model for Coreference Resolution and Dropped Pronoun Recovery in Chinese Conversations
abstract
In this work, we present that coreference resolution and dropped pronoun recovery are two strongly related tasks in Chinese conversations, as recovering the dropped pronoun needs to explore the referent of the pronoun at first. Meanwhile, the omitted entity mention should be recovered before its coreferences are resolved. This motivates us to propose CorefDPR, a novel model to jointly resolve these two tasks and make them enhance each other. CorefDPR firstly utilizes a pre-trained language model to encode tokens in the conversation snippet. Then, the coreference resolution layer detects all entity mentions from the candidate text spans and groups them as coreferent mention clusters based on the contextualized token states. Furthermore, the pronoun recovery layer explores the referent of each dropped pronoun from the coreferent mention clusters and predicts the probability distribution over pronoun category for each token. Finally, a general conditional random fields (GCRF) is employed to globally optimize the pronoun recovery sequence of the snippet by modeling both intra-utterance and cross-utterance pronoun dependencies, and the recovered pronouns are further linked back to corresponding mention clusters to complete them. Experimental results on the benchmark demonstrate that our proposed model outperformed the state-of-the-art baselines of both these two tasks, and the exploratory experiments also demonstrate that these two tasks mutually benefit each other.
Si Li 0001, Sheng Gao 0001, Jun Guo 0002
IEEE ACM Trans. Audio Speech Lang. Process.2
2021 A Joint Model for Dropped Pronoun Recovery and Conversational Discourse Parsing in Chinese Conversational Speech
abstract
Jingxuan Yang, Kerui Xu, Jun Xu, Si Li, Sheng Gao, Jun Guo, Nianwen Xue, Ji-Rong Wen. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Kerui Xu, Jun Xu 0001, Si Li 0001, Sheng Gao 0001, Jun Guo 0002, Nianwen Xue, Ji-Rong Wen
ACL/IJCNLP (1)4
2021 Spatially-Varying Outdoor Lighting Estimation From Intrinsics
abstract
We present SOLID-Net, a neural network for spatially- varying outdoor lighting estimation from a single outdoor image for any 2D pixel location. Previous work has used a unified sky environment map to represent outdoor lighting. Instead, we generate spatially-varying local lighting environment maps by combining global sky environment map with warped image information according to geometric information estimated from intrinsics. As no outdoor dataset with image and local lighting ground truth is readily available, we introduce the SOLID-Img dataset with physically- based rendered images and their corresponding intrinsic and lighting information. We train a deep neural network to regress intrinsic cues with physically-based constraints and use them to conduct global and local lightings estimation. Experiments on both synthetic and real datasets show that SOLID-Net significantly outperforms previous methods.
Yongjie Zhu, Yinda Zhang 0001, Si Li 0001, Boxin Shi
CVPR3
2021 DeRenderNet: Intrinsic Image Decomposition of Urban Scenes with Shape-(In)dependent Shading Rendering
abstract
We propose DeRenderNet, a deep neural network to decompose the albedo and latent lighting, and render shape-(in)dependent shadings, given a single image of an outdoor urban scene, trained in a self-supervised manner. To achieve this goal, we propose to use the albedo maps extracted from scenes in videogames as direct supervision and pre-compute the normal and shadow prior maps based on the depth maps provided as indirect supervision. Compared with state-of-the-art intrinsic image decomposition methods, DeRenderNet produces shadow-free albedo maps with clean details and an accurate prediction of shadows in the shape-independent shading, which is shown to be effective in re-rendering and improving the accuracy of high-level vision tasks for urban scenes.
Yongjie Zhu, Jiajun Tang 0001, Si Li 0001, Boxin Shi
ICCP3
2021 Memetic Federated Learning for Biomedical Natural Language Processing
Xinya Zhou, Conghui Tan, Di Jiang 0004, Bosen Zhang, Si Li 0001, Qian Xu 0005, Sheng Gao 0001
NLPCC (2)5
2021 Meta-Learned Specific Scenario Interest Network for User Preference Prediction
abstract
User preference prediction is a task of learning user interests through user-item interactions. Most existing studies capture user interests based on historical behaviors without considering specific scenario information. However, the users may have special interests in these specific scenarios and sometimes user historical behaviors are limited. In this paper, we propose a Meta-Learned Specific Scenario Interest Network (Meta-SSIN) to predict user preference of target item by capturing specific scenario interests. Meta-SSIN uses multiple independent meta-learning modules to model historical behaviors in each scenario. The independent module can capture special interests based on limited behaviors. Experimental results on three datasets show that Meta-SSIN outperforms compared state-of-the-art methods.
Hehuan Liu, Si Li 0001, Jun Guo 0002
SIGIR4
2020 Near-Infrared Image Guided Reflection Removal
abstract
Removing reflections from a single RGB image is a highly ill-posed problem. Unlike RGB images, near-infrared (NIR) images obtained through an active NIR camera are less likely to be affected by reflections when glass and camera planes form certain angles, while textures on objects could “vanish” under certain circumstances. Based on this observation, we propose a two-stream neural network to remove undesired reflections in an RGB image with the guidance of an NIR image. To tackle the insufficiency of training data, we propose a synthetic data generation pipeline that simulates the reflection-suppressing nature of the active NIR imaging and build a dataset mixed with synthetic and real data. Experimental results show that the proposed method outperforms state-of-the-art reflection removal methods in both quantitative metrics and visual quality.
Yuchen Hong, Youwei Lyu, Si Li 0001, Boxin Shi
ICME3
2020 Knowledge-based Context-aware Multi-turn Conversational Model with Hierarchical Attention
abstract
We study response generation in multi-turn open- domain dialogue systems. Background knowledge based response generation has been developed to make dialogue models generate more informative and appropriate responses. However, these knowledge-based dialogue models are limited to the domain of single round conversation, and fail to consider the role of dialogue context in the selection of relevant knowledge and response generation. As a result, these models might lose some useful information in the dialogue context and generate irrelevant responses. We argue that both dialogue context and relevant knowledge play important roles in the response generation of multiturn open-domain dialogue systems. We propose a Knowledge- based Context-aware Multi-turn Conversational (KCMC) model to consider both dialogue context and relevant knowledge in a unified framework. The Knowledge Fusion module is designed to augment the semantic representation of dialogue context with associated knowledge triples. And we introduce hierarchical encoders to model the hierarchy of dialogue context and to capture important information in the dialogue context. Furthermore, a hierarchical attention mechanism attends to important parts of knowledge triples, which facilitates better knowledge selection and response generation. Through extensive experiments on two datasets, we demonstrate that the proposed model is capable of generating more informative and appropriate responses than baseline models.
Chunquan Chen, Si Li 0001
IJCNN2
2020 Weaken Grammatical Error Influence in Chinese Grammatical Error Correction
Jinggui Liang, Si Li 0001
NLPCC (2)2
2020 Outline Extraction with Question-Specific Memory Cells
abstract
Outline extraction has been widely applied in online consultation to help experts quickly understand individual cases. Given a specific case described as unstructured plain text, outline extraction aims to make a summary for this case by answering a set of questions, which in fact is a new type of machine reading comprehension task. Inspired by a recently popular memory network, we propose a novel question-specific memory cell network (QSMCN) to extract information related to multiple questions on-the-fly as it reads texts. QSMCN constructs a specific memory cell for each question, which is sequentially expanded in recurrent neural network style. Each cell contains three specific vectors to first identify whether current input is related to corresponding question and then update question-specific case representation. We add a penalization term in the loss function to make extracted knowledge more reasonable and interpretable. To support this study, we construct a new outline extraction corpus, InjuryCase, 1 which is composed of 3,995 real Chinese occupational injury cases. Experimental results show that our method makes a significant improvement. We further apply the proposed framework on two multi-aspect extraction tasks and find that the proposed model also remarkably outperforms existing state-of-the-art methods of the aspect extraction task.
Haotian Cui, Si Li 0001, Sheng Gao 0001, Jun Guo 0002, Zhengdong Lu
ACM Trans. Asian Low Resour. Lang. Inf. Process.3
2019 A Prism Module for Semantic Disentanglement in Name Entity Recognition
abstract
Natural Language Processing has been perplexed for many years by the problem that multiple semantics are mixed inside a word, even with the help of context.To solve this problem, we propose a prism module to disentangle the semantic aspects of words and reduce noise at the input layer of a model.In the prism module, some words are selectively replaced with task-related semantic aspects, then these denoised word representations can be fed into downstream tasks to make them easier.Besides, we also introduce a structure to train this module jointly with the downstream model without additional data.This module can be easily integrated into the downstream model and significantly improve the performance of baselines on named entity recognition (NER) task.The ablation analysis demonstrates the rationality of the method.As a side effect, the proposed method also provides a way to visualize the contribution of each word.1
Daqi Zheng, Zhengdong Lu, Sheng Gao 0001, Si Li 0001
ACL (1)6
2019 Unsupervised Image Retrieval With Mask-Based Prominent Feature Accumulation
abstract
For unsupervised image retrieval, which features are chosen for final representation determines its performance. Nowadays, unsupervised methods can deal with most image retrieval tasks. However, it is still a challenging task to retrieve images with complex background. In this paper, we propose a new approach of mask-based prominent feature accumulation (MPFA), which utilizes MAX-Mask and SUM-Mask to retain significant features in each channel for all database images. Channels are then sorted by MPFA to select representative channels of feature maps extracted from pre-trained CNN. After that, the final image representation is generated by aggregating the selected channels. Experiments on the public datasets show improvement of our proposed approach compared to state-of-the-art methods, especially for images with complex background.
Haitao Yang 0008, Si Li 0001
ICIP4
2019 Cyclone Intensity Estimate with Context-Aware Cyclegan
abstract
Deep learning approaches to cyclone intensity estimation have recently shown promising results. However, suffering from the extreme scarcity of cyclone data on specific intensity, most existing deep learning methods fail to achieve satisfactory performance on cyclone intensity estimation, especially on classes with few instances. To avoid the degradation of recognition performance caused by scarce samples, we propose a context-aware CycleGAN which learns the latent evolution features from adjacent cyclone intensity and synthesizes CNN features of classes lacking samples from unpaired source classes. Specifically, our approach synthesizes features conditioned on the learned evolution features, while the extra information is not required. Experimental results of several evaluation methods show the effectiveness of our approach, even can predicting unseen classes.
Haitao Yang 0008, Mingfei Cheng, Si Li 0001
ICIP4
2019 From Market to Dish: Multi-ingredient Image Recognition for Personalized Recipe Recommendation
abstract
Recognition of food ingredients enables applications on recipe recommendation for developing a healthier eating habit. Existing ingredients recognition methods largely rely on ideal images captured in a controlled environment, while ingredients are usually displayed unorderly in a complex environment in the market. We propose the multi-ingredient recognition problem in the market and develop a Spatial Regularization Network (SRN) based method to solve it by using a newly collected multiple vegetable image dataset captured in the market. We further use the recognition result to develop a recipe recommendation system to satisfy the daily nutrition requirements and individual preference of each user. Experiments show that our multi-ingredient recognition outperforms previous methods over 14% in mAP and recommendation model shows an improvement of over 23% in HR@10.
Lin Zhang 0014, Jianbo Zhao 0002, Si Li 0001, Boxin Shi, Ling-Yu Duan
ICME3
2019 Reflection Separation using a Pair of Unpolarized and Polarized Images
abstract
When we take photos through glass windows or doors, the transmitted background scene is often blended with undesirable reflection. Separating two layers apart to enhance the image quality is of vital importance for both human and machine perception. In this paper, we propose to exploit physical constraints from a pair of unpolarized and polarized images to separate reflection and transmission layers. Due to the simplified capturing setup, the system becomes more underdetermined compared with existing polarization based solutions that take three or more images as input. We propose to solve semireflector orientation estimation first to make the physical image formation well-posed and then learn to reliably separate two layers using a refinement network with gradient loss. Quantitative and qualitative experimental results show our approach performs favorably over existing polarization and single image based solutions.
Youwei Lyu, Zhaopeng Cui, Si Li 0001, Marc Pollefeys, Boxin Shi
NeurIPS3
2018 Word-Driven and Context-Aware Review Modeling for Recommendation
abstract
Recently, convolutional neural networks(CNNs) has been demonstrated to effectively model reviews in recommender systems, due to the learning of contextual features such as surrounding words and word order for reviews. However, CNNs with max-pooling fails to capture the count information of contextual features, since the feature map generated by a convolution filter can only get a max feature value with max-pooling. If the max feature value appears more than once in the feature map, CNNs will lose the count information of the contextual feature. The count information is quite critical for modeling reviews, for example, ten "a good quality" is more credible than one "a good quality" for representing item properties, and five "the scenery is" shows that a user may pay more attention to the scenery than one "the scenery is" does. Our model, called WCN-MF, extends CNNs by introducing a new module named Deep Latent Dirichlet Allocation (DLDA) to capture the count information of contextual features. DLDA is inspired by the fact that contextual features consist of words, hence, we can first capture the count information of words, and second generate contextual features with words. By combining DLDA with CNNs, we can get a word-driven and context-aware review representation. Further, we incorporate the review representation with Matrix Factorization for recommendation. Our evaluations on three real- world datasets reveal that our model can significantly outperform the state-of-the-art recommendation models.
Si Li 0001, Guang Chen 0003
CIKM2
2018 From Random to Supervised: A Novel Dropout Mechanism Integrated with Global Information
abstract
Dropout is used to avoid overfitting by randomly dropping units from the neural networks during training.Inspired by dropout, this paper presents GI-Dropout, a novel dropout method integrating with global information to improve neural networks for text classification.Unlike the traditional dropout method in which the units are dropped randomly according to the same probability, we aim to use explicit instructions based on global information of the dataset to guide the training process.With GI-Dropout, the model is supposed to pay more attention to inapparent features or patterns.Experiments demonstrate the effectiveness of the dropout with global information on seven text classification tasks, including sentiment analysis and topic classification.
Hengru Xu, Renfen Hu, Si Li 0001, Sheng Gao 0001
CoNLL4
2017 A Compare-Aggregate Model with Dynamic-Clip Attention for Answer Selection
abstract
Answer selection for question answering is a challenging task, since it requires effective capture of the complex semantic relations between questions and answers. Previous remarkable approaches mainly adopt general Compare-Aggregate framework that performs word-level comparison and aggregation. In this paper, unlike previous Compare-Aggregate models which utilize the traditional attention mechanism to generate corresponding word-level vector before comparison, we propose a novel attention mechanism named Dynamic-Clip Attention which is directly integrated into the Compare-Aggregate framework. Dynamic-Clip Attention focuses on filtering out noise in attention matrix, in order to better mine the semantic relevance of word-level vectors. At the same time, different from previous Compare-Aggregate works which treat answer selection task as a pointwise classification problem, we propose a listwise ranking approach to model this task to learn the relative order of candidate answers. Experiments on TrecQA and WikiQA datasets show that our proposed model achieves the state-of-the-art performance.
Weijie Bian, Si Li 0001, Guang Chen 0003, Zhiqing Lin
CIKM2
2017 Aggregating Class Interactions for Hierarchical Attention Relation Extraction
Si Li 0001, Guang Chen 0003
ICONIP (2)2
2017 Neural Domain Adaptation with Contextualized Character Embedding for Chinese Word Segmentation
Zuyi Bao, Si Li 0001, Sheng Gao 0001, Weiran Xu
NLPCC2
2017 Improved Compare-Aggregate Model for Chinese Document-Based Question Answering
Weijie Bian, Si Li 0001, Guang Chen 0003, Zhiqing Lin
NLPCC3
2017 Product ranking using hierarchical aspect structures
Si Li 0001, Zhaoyan Ming, Yan Leng, Jun Guo 0002
J. Intell. Inf. Syst.1
2012 SSHLDA: A Semi-Supervised Hierarchical Topic Model
Xianling Mao, Zhaoyan Ming, Tat-Seng Chua, Si Li 0001, Hongfei Yan, Xiaoming Li 0001
EMNLP-CoNLL4
2011 Product comparison using comparative relations
abstract
This paper proposes a novel Product Comparison approach. The comparative relations between products are first mined from both user reviews on multiple review websites and community-based question answering pairs containing product comparison information. A unified graph model is then developed to integrate the resultant comparative relations for product comparison. Experiments on popular electronic products show that the proposed approach outperforms the state-of-the-art methods.
Si Li 0001, Zhengjun Zha, Zhaoyan Ming, Meng Wang 0001, Tat-Seng Chua, Jun Guo 0002, Weiran Xu
SIGIR1
2010 Exploiting Combined Multi-level Model for Document Sentiment Analysis
abstract
This paper focuses on the task of text sentiment analysis in hybrid online articles and web pages. Traditional approaches of text sentiment analysis typically work at a particular level, such as phrase, sentence or document level, which might not be suitable for the documents with too few or too many words. Considering every level analysis has its own advantages, we expect that a combination model may achieve better performance. In this paper, a novel combined model based on phrase and sentence level's analyses and a discussion on the complementation of different levels' analyses are presented. For the phrase-level sentiment analysis, a newly defined Left-Middle-Right template and the Conditional Random Fields are used to extract the sentiment words. The Maximum Entropy model is used in the sentence-level sentiment analysis. The experiment results verify that the combination model with specific combination of features is better than single level model.
Si Li 0001, Hao Zhang 0022, Weiran Xu, Guang Chen 0003, Jun Guo 0002
ICPR1