EDBT 2026 Demo / reviewers in the wild / expert
Lingzhi Zhang
dblp:70/1413
· DBLP profile ↗
31ranked-venue papers
14as first author
22since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 18 · 10 first-author · 13 since 2021Artificial intelligence and machine learning · 17 · 8 first-author · 13 since 2021Databases, data management, data science and information retrieval · 2Systems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | From Words to Pixels: A Comprehensive Survey on Large Language Models in Visual SegmentationabstractVisual segmentation, the task of segmenting an image into semantically meaningful regions, is a cornerstone in machine learning and has widespread applications in industry. Nevertheless, visual segmentation with instruction has been a challenging task for many years. This largely stems from the cross-modal discrepancy between language and image domains, resulting in difficulty in relating the instruction semantics and the pixel-level predictions. In recent years, the remarkable reasoning capabilities of Large Language Models (LLMs) and Large Multimodal Models (LMMs) have spurred a new wave of research aiming to bridge the disparity between natural language instructions and pixel-level understanding. This survey offers the first comprehensive overview of the rapidly evolving field of LLM-driven visual segmentation. We categorize existing approaches based on their core objectives and methodologies, including reasoning-based segmentation, open-vocabulary segmentation, grounding techniques connecting language to pixels, and extensions to video domains. We review recent seminal works in LLM-based visual segmentation, analyzing their architectural innovations, training strategies, and benchmark performance. Furthermore, we discuss the common datasets, evaluation metrics, and identify key challenges and promising future directions at the intersection of language and visual segmentation. We hope this survey serves as a valuable resource for researchers and practitioners seeking to understand the current landscape and future directions of leveraging LLMs for sophisticated visual segmentation tasks and applications. The resource summary is available at https://github.com/wyzjack/Awesome-LLM-Visual-Segmentation. Yizhou Wang 0006, Mang Tik Chiu, Lingzhi Zhang, Xuan Shen, Sohrab Amirghodsi, Yun Fu 0001 |
ACL (1) | 3 |
| 2026 | Fine-grained Defocus Blur Control for Generative Image ModelsabstractCurrent text-to-image diffusion models excel at generating diverse, high-quality images, yet they struggle to incorporate fine-grained camera metadata such as precise aperture settings. In this work, we introduce a novel text-to-image diffusion framework that leverages camera metadata, or EXIF data, which is often embedded in image files, with an emphasis on generating controllable lens blur. Our method mimics the physical image formation process by first generating an all-in-focus image, estimating its monocular depth, predicting a plausible focus distance with a novel focus distance transformer, and then forming a defocused image with an existing differentiable lens blur model [32]. Gradients flow backwards through this whole process, allowing us to learn without explicit supervision to generate defocus effects based on content elements and the provided EXIF data. At inference time, this enables precise interactive user control over defocus effects while preserving scene contents, which is not achievable with existing diffusion models. Experimental results demonstrate that our model enables superior fine-grained control without altering the depicted scene. Ayush Shrivastava, Connelly Barnes, Xuaner Cecilia Zhang, Lingzhi Zhang, Andrew Owens, Sohrab Amirghodsi, Eli Shechtman |
WACV | 4 |
| 2025 | FINECAPTION: Compositional Image Captioning Focusing on Wherever You Want at Any GranularityabstractThe advent of large Vision-Language Models (VLMs) has significantly advanced multimodal tasks, enabling more sophisticated and accurate reasoning across various applications, including image and video captioning, visual question answering, and cross-modal retrieval. Despite their superior capabilities, VLMs struggle with fine-grained image regional composition information perception. Specifically, they have difficulty accurately aligning the segmentation masks with the corresponding semantics and precisely describing the compositional aspects of the referred regions. However, compositionality – the ability to understand and generate novel combinations of known visual and textual components – is critical for facilitating coherent reasoning and understanding across modalities by VLMs. To address this issue, we propose FineCaption, a novel VLM that can recognize arbitrary masks as referential inputs and process high-resolution images for compositional image captioning at different granularity levels. To support this endeavor, we introduce CompositionCap, a new dataset for multi-grained region compositional image captioning, which introduces the task of compositional attribute-aware regional image captioning. Empirical results demonstrate the effectiveness of our proposed model compared to other state-of-the-art VLMs. Additionally, we analyze the capabilities of current VLMs in recognizing various visual prompts for compositional region image captioning, highlighting areas for improvement in VLM design and training. https://hanghuacs.github.io/FineCaption/ Hang Hua, Qing Liu 0017, Lingzhi Zhang, Jing Shi 0005, Soo Ye Kim, Yilin Wang 0002, Jianming Zhang 0001, Zhe Lin 0001, Jiebo Luo 0001 |
CVPR | 3 |
| 2025 | Layer- and Timestep-Adaptive Differentiable Token Compression Ratios for Efficient Diffusion TransformersabstractDiffusion Transformers (DiTs) have achieved state-of-the-art (SOTA) image generation quality but suffer from high latency and memory inefficiency, making them difficult to deploy on resource-constrained devices. One major efficiency bottleneck is that existing DiTs apply equal computation across all regions of an image. However, not all image tokens are equally important, and certain localized areas require more computation, such as objects. To address this, we propose DiffCR, a dynamic DiT inference framework with differentiable compression ratios, which automatically learns to dynamically route computation across layers and timesteps for each image token, resulting in efficient DiTs. Specifically, DiffCR integrates three features: (1) A token-level routing scheme where each DiT layer includes a router that is fine-tuned jointly with model weights to predict token importance scores. In this way, unimportant tokens bypass the entire layer’s computation; (2) A layer-wise differentiable ratio mechanism where different DiT layers automatically learn varying compression ratios from a zero initialization, resulting in large compression ratios in redundant layers while others remain less compressed or even uncompressed; (3) A timestep-wise differentiable ratio mechanism where each denoising timestep learns its own compression ratio. The resulting pattern shows higher ratios for noisier timesteps and lower ratios as the image becomes clearer. Extensive experiments on text-to-image and inpainting tasks show that DiffCR effectively captures dynamism across token, layer, and timestep axes, achieving superior tradeoffs between generation quality and efficiency compared to prior works. The project website is available here. Haoran You, Connelly Barnes, Yuqian Zhou, Zhenbang Du, Lingzhi Zhang, Yotam Nitzan, Zhe Lin 0001, Eli Shechtman, Sohrab Amirghodsi, Yingyan (Celine) Lin |
CVPR | 7 |
| 2025 | Refer to Any Segmentation Mask Group with Vision-Language PromptsabstractRecent image segmentation models have advanced to segment images into high-quality masks for visual entities, and yet they cannot provide comprehensive semantic understanding for complex queries based on both language and vision. This limitation reduces their effectiveness in applications that require user-friendly interactions driven by vision-language prompts. To bridge this gap, we introduce a novel task of omnimodal referring expression segmentation (ORES). In this task, a model produces a group of masks based on arbitrary prompts specified by text only or text plus reference visual entities. To address this new challenge, we propose a novel framework to "Refer to Any Segmentation Mask Group" (RAS), which augments segmentation models with complex multimodal interactions and comprehension via a mask-centric large multimodal model. For training and benchmarking ORES models, we create datasets MaskGroups-2M and MaskGroups-HQ to include diverse mask groups specified by text and reference entities. Through extensive evaluation, we demonstrate superior performance of RAS on our new ORES task, as well as classic referring expression segmentation (RES) and generalized referring expression segmentation (GRES) tasks. Project page: https://Ref2Any.github.io. Shengcao Cao, Zijun Wei, Jason Kuen, Kangning Liu, Lingzhi Zhang, Jiuxiang Gu, Hyunjoon Jung, Liangyan Gui, Yu-Xiong Wang |
ICCV | 5 |
| 2025 | LogAD: A Multi-Feature Fusion Approach for Log Anomaly DetectionabstractWith the increasing complexity of software systems, log-based anomaly detection has become critical for ensuring system reliability. However, existing methods often suffer from limited feature integration and insufficient semantic representation, leading to unstable detection performance. To address these challenges, this paper proposes a multi-feature fusion framework for log anomaly detection, leveraging heterogeneous graph neural networks (HGNNs) to capture rich semantic relationships. First, we design a hybrid preprocessing pipeline that combines log parsing (via Drain), session-fixed window grouping, and hybrid label estimation using HDBSCAN clustering and HNSW-based similarity search. This step mitigates label scarcity while enhancing feature representation robustness. Second, we construct a heterogeneous graph with three node types-log sequences, templates, and parameters-to model interdependencies between log events through meta-paths, enabling comprehensive feature fusion. Third, a heterogeneous graph attention network (HGAT) with multi-head attention is developed to prioritize critical patterns across meta-paths, improving anomaly discrimination. Experimental results on benchmark datasets demonstrate that our model outperforms state-of-the-art baselines in accuracy and F1-score. Furthermore, we implement LogAD, an automated detection tool integrating ELK-stack-based log management, multi-feature anomaly detection, and security-focused operational support. The system's visualization interface and efficient processing pipeline provide a practical solution for real-world deployment. This work advances log analysis by bridging feature isolation and semantic sparsity, offering both algorithmic innovation and engineering applicability. Guangzu Wang, Lingzhi Zhang, Tianyu Wo, Xu Wang 0007, Chunming Hu |
JCC | 2 |
| 2025 | Good Seed Makes a Good Crop: Discovering Secret Seeds in Text-to-Image Diffusion Models
Katherine Xu, Lingzhi Zhang, Jianbo Shi |
WACV | 2 |
| 2025 | Detecting Origin Attribution for Text-to-Image Diffusion ModelsabstractModern text-to-image (T2I) diffusion models can generate images with remarkable realism and creativity. These advancements have sparked research in fake image detection and attribution, yet prior studies have not fully explored the practical and scientific dimensions of this task. In addition to attributing images to 12 state-of-the-art T2I generators, we provide extensive analyses on what inference stage hyperparameters and image modifications are discernible. Our experiments reveal that initialization seeds are highly detectable, along with other subtle variations in the image generation process to some extent. We further investigate what visual traces are leveraged in image attribution by perturbing high-frequency details and employing midlevel representations of image style and structure. Notably, altering high-frequency information causes only slight reductions in accuracy, and training an attributor on style representations outperforms training on RGB images. Our analyses underscore that fake images are detectable and attributable at various levels of visual granularity. Katherine Xu, Lingzhi Zhang, Jianbo Shi |
WACV | 2 |
| 2025 | Optimal Control for Constrained Discrete-Time Nonlinear Systems Based on Safe Reinforcement LearningabstractThe state and input constraints of nonlinear systems could greatly impede the realization of their optimal control when using reinforcement learning (RL)-based approaches since the commonly used quadratic utility functions cannot meet the requirements of solving constrained optimization problems. This article develops a novel optimal control approach for constrained discrete-time (DT) nonlinear systems based on safe RL. Specifically, a barrier function (BF) is introduced and incorporated with the value function to help transform a constrained optimization problem into an unconstrained one. Meanwhile, the minimum of such an optimization problem can be guaranteed to occur at the origin. Then a constrained policy iteration (PI) algorithm is developed to realize the optimal control of the nonlinear system and to enable the state and input constraints to be satisfied. The constrained optimal control policy and its corresponding value function are derived through the implementation of two neural networks (NNs). Performance analysis shows that the proposed control approach still retains the convergence and optimality properties of the traditional PI algorithm. Simulation results of three examples reveal its effectiveness. Lingzhi Zhang, Lei Xie 0007, Yi Jiang 0007, Zhishan Li, Xueqin Amy Liu |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | Brush2Prompt: Contextual Prompt Generator for Object InpaintingabstractObject inpainting is a task that involves adding objects to real images and seamlessly compositing them. With the recent commercialization of products like Stable Diffusion and Generative Fill, inserting objects into images by using prompts has achieved impressive visual results. In this paper, we propose a prompt suggestion model to simplify the process of prompt input. When the user provides an image and a mask, our model predicts suitable prompts based on the partial contextual information in the masked image, and the shape and location of the mask. Specifically, we introduce a concept-diffusion in the CLIP space that predicts CLIP-text embeddings from a masked image. These diffused embeddings can be directly injected into open-source in-painting models like Stable Diffusion and its variants. Alternatively, they can be decoded into natural language for use in other publicly available applications such as Generative Fill. Our prompt suggestion model demonstrates a balanced accuracy and diversity, showing its capability to be both contextually aware and creatively adaptive. Mang Tik Chiu, Yuqian Zhou, Lingzhi Zhang, Zhe Lin 0001, Connelly Barnes, Sohrab Amirghodsi, Eli Shechtman, Humphrey Shi |
CVPR | 3 |
| 2024 | Amodal Completion via Progressive Mixed Context DiffusionabstractOur brain can effortlessly recognize objects even when partially hidden from view. Seeing the visible of the hidden is called amodal completion; however, this task remains a challenge for generative AI despite rapid progress. We propose to sidestep many of the difficulties of existing approaches, which typically involve a two-step process of predicting amodal masks and then generating pixels. Our method involves thinking outside the box, literally! We go outside the object bounding box to use its context to guide a pretrained diffusion inpainting model, and then progressively grow the occluded object and trim the extra background. We overcome two technical challenges: 1) how to be free of unwanted co-occurrence bias, which tends to regenerate similar occluders, and 2) how to judge if an amodal completion has succeeded. Our amodal completion method exhibits improved photorealistic completion results compared to existing approaches in numerous successful completion cases. And the best part? It doesn't require any special training or fine-tuning of models. Katherine Xu, Lingzhi Zhang, Jianbo Shi |
CVPR | 2 |
| 2024 | Event-Triggered Constrained Optimal Control for Organic Rankine Cycle Systems via Safe Reinforcement LearningabstractThe organic Rankine cycle (ORC) is an effective application for converting low-grade heat sources into power and is crucial for environmentally friendly production and energy recovery. However, the inherent complexity of the mechanism, its strong and unidentified nonlinearity, and the presence of control constraints severely impair the design of its optimal controller. To solve these issues, this study provides a novel event-triggered (ET) constrained optimal control approach for the ORC systems based on a safe reinforcement learning technique to find the optimal control law. Instead of employing the usual non-quadratic integral form to solve the control-limited optimal control problems, a constraint handling strategy based on a relaxed weighted barrier function (BF) technique is proposed. By adding the BF terms to the original value function, a modified value iteration algorithm is developed to make the control input solutions that tend to violate the constraints be pushed back and maintained in their safe sets. In addition, the ET mechanism proposed in this article is critically required for the ORC systems, and it can significantly reduce the computational load. The combination of these two techniques allows the ORC systems to achieve set-point tracking control and satisfy the control restrictions. The proposed approach is conducted based on a heuristic dynamic programming framework with three neural networks (NNs) involved. The safety and convergence of the proposed approach and the stability of the closed-loop system are analyzed. Simulation results and comparisons are presented to demonstrate its effectiveness. Lingzhi Zhang, Runze Lin, Lei Xie 0007, Wei Dai 0004 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | Operational Optimal Tracking Control for Industrial Multirate Systems Subject to Unknown DisturbancesabstractIt is well common for industrial processes to employ a hierarchical control structure involving a basic loop process and an operation loop process with two timescales. However, the control system suffers from another multirate challenge where control and sampling rates may differ even within a single loop. Additionally, the underlying complex mechanism of the operation loop further complicates the accurate modeling of its dynamics, especially in the presence of external unknown disturbances. This gives rise to the difficulty in obtaining desired control performance. To overcome these problems, this article develops a novel operational optimal tracking control method for a class of multirate systems subject to unknown disturbances. To this end, a lifting technique is integrated with a general model predictive controller for the basic loop process, aimed at handling the asynchronism phenomenon and achieving loop setpoint tracking control. Furthermore, a nonlinear disturbance observer is used for estimating the unknown external disturbance of the operation loop process. In this way, offset-free tracking control of the system, along with loop setpoints optimization, can be achieved using the policy iteration reinforcement learning algorithm. The convergence of the proposed method is analyzed and tangible improvements are verified by simulations. Lingzhi Zhang, Lei Xie 0007, Wei Dai 0004, Shan Lu 0009 |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2023 | Perceptual Artifacts Localization for Image Synthesis TasksabstractRecent advancements in deep generative models have facilitated the creation of photo-realistic images across various tasks. However, these generated images often exhibit perceptual artifacts in specific regions, necessitating manual correction. In this study, we present a comprehensive empirical examination of Perceptual Artifacts Localization (PAL) spanning diverse image synthesis endeavors. We introduce a novel dataset comprising 10, 168 generated images, each annotated with per-pixel perceptual artifact labels across ten synthesis tasks. A segmentation model, trained on our proposed dataset, effectively localizes artifacts across a range of tasks. Additionally, we illustrate its proficiency in adapting to previously unseen models using minimal training samples. We further propose an innovative zoom-in inpainting pipeline that seamlessly rectifies perceptual artifacts in the generated images. Through our experimental analyses, we elucidate several invaluable downstream applications, such as automated artifact rectification, non-referential image quality evaluation, and abnormal region detection in images. The dataset and code are released here: https://owenzlz.github.io/PAL4VST Lingzhi Zhang, Zhengjie Xu, Connelly Barnes, Yuqian Zhou, Qing Liu 0017, He Zhang 0004, Sohrab Amirghodsi, Zhe Lin 0001, Eli Shechtman, Jianbo Shi |
ICCV | 1 |
| 2022 | Inpainting at Modern Camera Resolution by Guided PatchMatch with Auto-curation
Lingzhi Zhang, Connelly Barnes, Kevin Wampler, Sohrab Amirghodsi, Eli Shechtman, Zhe Lin 0001, Jianbo Shi |
ECCV (17) | 1 |
| 2022 | Perceptual Artifacts Localization for Inpainting
Lingzhi Zhang, Yuqian Zhou, Connelly Barnes, Sohrab Amirghodsi, Zhe Lin 0001, Eli Shechtman, Jianbo Shi |
ECCV (29) | 1 |
| 2022 | Fine-Grained Egocentric Hand-Object Segmentation: Dataset, Model, and Applications
Lingzhi Zhang, Shenghao Zhou, Simon Stent, Jianbo Shi |
ECCV (29) | 1 |
| 2022 | Inpaint2Learn: A Self-Supervised Framework for Affordance LearningabstractPerceiving affordances –the opportunities of interaction in a scene, is a fundamental ability of humans. It is an equally important skill for AI agents and robots to better understand and interact with the world. However, labeling affordances in the environment is not a trivial task. To address this issue, we propose a task-agnostic framework, named Inpaint2Learn, that generates affordance labels in a fully automatic manner and opens the door for affordance learning in the wild. To demonstrate its effectiveness, we apply it to three different tasks: human affordance prediction, Location2Object and 6D object pose hallucination. Our experiments and user studies show that our models, trained with the Inpaint2Learn scaffold, are able to generate diverse and visually plausible results in all three scenarios. Lingzhi Zhang, Weiyu Du, Shenghao Zhou, Jiancong Wang, Jianbo Shi |
WACV | 1 |
| 2022 | Multi-Rate Layered Operational Optimal Control for Large-Scale Industrial ProcessesabstractIn large-scale process industries, one of the great challenges is to achieve optimum operation of systems with multi-time-scale property and partially unknown models. To this end, this article proposes a novel multi-rate layered operational optimal control (OOC) method, which employs lifting technique to unify the relatively fast dual-rate of basic loop layer and relatively slow single-rate of operational layer. Besides, by integrating model-based predictive control of basic loop layer with data-based actor-critic reinforcement learning (RL) of operational layer, it overcomes the difficulty of building the operational process dynamic model. The convergence of the proposed method is proved, and dense medium separation (DMS) process is taken as an application case to illustrate the effectiveness of our proposed method via a self-developed simulation platform. Wei Dai 0004, Tongyun Li, Lingzhi Zhang, Yao Jia 0001, Huaicheng Yan 0001 |
IEEE Trans. Ind. Informatics | 3 |
| 2021 | Speech Emotion Recognition Model with Time-Scale-Invariance MFCCs as InputabstractSpeech Emotion Recognition (SER) is a significant task for human communication. In the recent years, Mel-frequency Cepstrum Coefficient (MFCC) feature can be usually utilized in the related tasks of speech emotion recognition. In this study, we developed a multi-head-attention CNN model with auxiliary task of gender task. Base on proposed model, we explore the effect of different time-scale MFCCs and different combination of them as input on the performance of proposed model. Experimental results show that MFCC having higher resolution in time-scale as input can help model achieving better performance of speech emotion recognition with a moderate range. Also, it can help model achieving better performance to combine different time-scale MFCCs appropriately. Xiaohan Xie, Jiaqi Lou, Lingzhi Zhang |
IWCMC | 3 |
| 2021 | Merged Biogeography-Based Optimization Algorithm for Color Image SegmentationabstractImage segmentation is an important step in image processing. Segmentation based on threshold is a common method. Searching the suitable threshold vector is essentially an optimization problem, especially for color image which have higher dimensions and complexity. This paper proposes a merged Biogeography-Based Optimization algorithm (MBBO) for color image segmentation based on threshold. We merge a mutation operation into the migration operator of BBO to enhance the global search ability. Then we merge a chemotaxis operation into the mutation operator of BBO to enhance the local search ability. A greedy selection method is also used to further improve the performance and reduce computation complexity. Experimental results show that MBBO obtains better optimization performance, stronger stability and faster running speed compared with other existing algorithms. Lingzhi Zhang, Xiaohan Xie |
IWCMC | 1 |
| 2021 | Dual-Rate Adaptive Optimal Tracking Control for Dense Medium Separation Process Using Neural NetworksabstractDense medium separation (DMS) is of great significance for coal cleaning. The DMS control system always involves dense medium density adjustment and ash content control that are operating on fast and slow time scales, respectively. The inherent time-varying and strongly nonlinear characteristics of the DMS process give rise to challenges for the design of this multitime scale control system. To address this issue, this article proposes a dual-rate adaptive optimal tracking control approach for the DMS system. For the basic loop process, a nonlinear adaptive PI controller containing a neural network (NN)-based unmodeled dynamics compensator is proposed. Then, a lifting technique is used to unify the time scales of the two loops accompanied by formulating a generalized controlled object, whose dynamics is completely unknown. On this basis, a data-driven operation optimization control method that combines adaptive dynamic programming algorithm and reference control is developed, which is implemented using NNs. Finally, the stability of the proposed method is analyzed. The simulation results indicate its effectiveness. Wei Dai 0004, Lingzhi Zhang, Jun Fu 0001, Tianyou Chai |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2020 | Nested Scale-Editing for Conditional Image SynthesisabstractWe propose an image synthesis approach that provides stratified navigation in the latent code space. With a tiny amount of partial or very low-resolution image, our approach can consistently out-perform state-of-the-art counterparts in terms of generating the closest sampled image to the ground truth. We achieve this through scale-independent editing while expanding scale-specific diversity. Scale-independence is achieved with a nested scale disentanglement loss. Scale-specific diversity is created by incorporating a progressive diversification constraint. We introduce semantic persistency across the scales by sharing common latent codes. Together they provide better control of the image synthesis process. We evaluate the effectiveness of our proposed approach through various tasks, including image outpainting, image superresolution, and cross-domain image translation. Lingzhi Zhang, Jiancong Wang, Yinshuang Xu, Jie Min, Tarmily Wen, James C. Gee, Jianbo Shi |
CVPR | 1 |
| 2020 | Learning Object Placement by Inpainting for Compositional Data Augmentation
Lingzhi Zhang, Tarmily Wen, Jie Min, Jiancong Wang, Jianbo Shi |
ECCV (13) | 1 |
| 2020 | Image Compression with Laplacian Guided Scale Space InpaintingabstractWe present an image compression algorithm that preserves high-frequency details and information of rare occurrences. Our approach can be thought of as image inpainting in the frequency scale space. Given an image, we construct a Laplacian image pyramid, and store only the finest and coarsest levels, thereby removing the middle-frequency of the image. Using a network backbone borrowed from an image super-resolution algorithm, we train our network to hallucinate the missing middle-level Laplacian image. We introduce a novel training paradigm where we train our algorithm using only a face dataset where the faces are aligned and scaled correctly. We demonstrate that image compression learned on this restricted dataset leads to better GAN network [1] convergence and generalization to completely different image domains. We also show that Lapacian inpainting could be simplified further with a few selective pixels as seeds. Lingzhi Zhang, Pujika Kumar, Manuj Sabharwal, Andy Kuzma, Jianbo Shi |
ICIP | 1 |
| 2020 | Deep Image BlendingabstractImage composition is an important operation to create visual content. Among image composition tasks, image blending aims to seamlessly blend an object from a source image onto a target image with lightly mask adjustment. A popular approach is Poisson image blending [23], which enforces the gradient domain smoothness in the composite image. However, this approach only considers the boundary pixels of target image, and thus can not adapt to texture of target background image. In addition, the colors of the target image often seep through the original source object too much causing a significant loss of content of the source object. We propose a Poisson blending loss that achieves the same purpose of Poisson image blending. In addition, we jointly optimize the proposed Poisson blending loss as well as the style and content loss computed from a deep network, and reconstruct the blending region by iteratively updating the pixels using the L-BFGS solver. In the blending image, we not only smooth out gradient domain of the blending boundary but also add consistent texture into the blending region. User studies show that our method outperforms strong baselines as well as state-of-the-art approaches when placing objects onto both paintings and real-world images. Code is available at: https://github.com/owenzlz/DeepImageBlending. Lingzhi Zhang, Tarmily Wen, Jianbo Shi |
WACV | 1 |
| 2020 | Multimodal Image Outpainting with Regularized Normalized DiversificationabstractIn this paper, we study the problem of generating a set of realistic and diverse backgrounds when given only a small foreground region. We refer to this task as image outpainting. The technical challenge of this task is to synthesize not only plausible but also diverse image outputs. Traditional generative adversarial networks suffer from mode collapse. While recent approaches [32], [28] propose to maximize or preserve the pairwise distance between generated samples with respect to their latent distance, they do not explicitly prevent the diverse samples of different conditional inputs from collapsing. Therefore, we propose a new regularization method to encourage diverse sampling in conditional synthesis. In addition, we propose a feature pyramid discriminator to improve the image quality. Our experimental results show that our model can produce more diverse images without sacrificing visual quality compared to state-of-the-arts approaches in both the CelebA face dataset [29] and the Cityscape scene dataset [2]. Code is available at: https://github.com/owenzlz/DiverseOutpaint. Lingzhi Zhang, Jiancong Wang, Jianbo Shi |
WACV | 1 |
| 2019 | A Multiobjective Evolutionary Algorithm Based on Coordinate TransformationabstractIn this paper, a novel multiobjective evolutionary algorithm (MOEA/CT) is proposed to better manage convergence and distribution of solutions when MOEAs are used for solving multiobjective optimization problems. The coordinate transformation strategy, an external archive update strategy, and a diversity maintenance strategy are proposed in MOEA/CT. The coordinate transformation strategy in the objective space is designed to find more efficient solutions that can accelerate the convergence process. Based on the coordinate transformation strategy, a novel update strategy and diversity maintenance approach for selecting nondominated solutions from the external archive set are integrated in MOEA/CT for getting better distribution of the solutions. The proposed MOEA/CT is compared with eight state-of-art algorithms on six biobjective and seven tri-objective test problems. In terms of four performance metrics, the comparative experimental results demonstrate that MOEA/CT outperforms the other eight competitors and it can achieve solutions with better distribution and better convergence to the Pareto front. In addition, parameter sensitivity analysis is provided to investigate the effect of a key parameter in MOEA/CT; the proposed three strategies are also studied individually to investigate their contribution to MOEA/CT; the performance analysis along with the capacity of external archive is given to clearly make the influence in MOEA/CT; finally, the scalability performance of MOEA/CT is investigated and compared with five notable many-objective evolutionary algorithms on the DTLZ and WFG test suites with 5, 8, 10, and 15 objectives. Wei Fang 0001, Lingzhi Zhang, Shengxiang Yang, Jun Sun 0008, Xiaojun Wu 0001 |
IEEE Trans. Cybern. | 2 |
| 2017 | A novel quantum-behaved particle swarm optimization with random selection for large scale optimizationabstractLarge scale optimization has become a well-recognised field in many science and engineering applications and a variety of metaheuristic algorithms adopting cooperative coevolution (CC) framework with problem decomposition have been applied to solve them. In this paper, a novel decomposition strategy termed as random selection is proposed. In random selection strategy, only a small part of decision variables are randomly selected to form a group for evolving at every iteration and the maximum number of randomly selected decision variables are limited by the parameter RSSCALE. By random selection, the randomly selected searching subspace is explored sufficiently in each iteration and the whole search space can be fully covered after several iterations. We evaluate the random selection strategy by combining quantum-behaved particle swarm optimization (RSQPSO) and a comparative study is carried out on a set of benchmark functions between RSQPSO and four state-of-the-art algorithms, which were specially designed for large scale optimization. The comparative results show that the proposed approach performs well for solving large scale optimization problems. Wei Fang 0001, Lingzhi Zhang, Xiaojun Wu 0001, Jun Sun 0008 |
CEC | 2 |
| 2005 | IMAX: The Big Picture of Dynamic XML StatisticsabstractCurrent approaches for estimating the cardinality of XML queries are applicable to a static scenario wherein the underlying XML data does not change subsequent to the collection of statistics on the repository. However, in practice, many XML-based applications are dynamic and involve frequent updates to the data. In this paper, we investigate efficient strategies for incrementally maintaining statistical summaries as and when updates are applied to the data. Specifically, we propose algorithms that handle both the addition of new documents as well as random insertions in the existing document trees. We also show, through a detailed performance evaluation, that our incremental techniques are significantly faster than the naive recomputation approach; and that estimation accuracy can be maintained even with a fixed memory budget. Maya Ramanath, Lingzhi Zhang, Juliana Freire, Jayant R. Haritsa |
ICDE | 2 |
| 2004 | A Flexible Infrastructure for Gathering XML Statistics and Estimating Query CardinalityabstractA key component of XML data management systems is the result size estimator, which estimates the cardinalities of user queries. Estimated cardinalities are needed in a variety of tasks, including query optimization and cost-based storage design; and they can also be used to give users early feedback about the expected outcome of their queries. In contrast to previously proposed result estimators, which use specialized data structures and estimation algorithms, StatiX uses histograms to uniformly capture both the structural and value skew present in documents. The original version of StatiX was built as a proof of concept. With the goal of making the system publicly available, we have built StatiX++, a new and improved version of StatiX, which extends the original system in significant ways. In this demonstration, we show the key features of StatiX++. Juliana Freire, Maya Ramanath, Lingzhi Zhang |
ICDE | 3 |