EDBT 2026 Demo / reviewers in the wild / expert
Minshan Xie
dblp:167/6600
· DBLP profile ↗
15ranked-venue papers
6as first author
13since 2021 · last 2026
0000-0002-6288-1611ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 13 · 6 first-author · 11 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | VLM-Guard: Defending Jailbreaks by Monitoring Only Hundreds of Safety-Critical Neurons
Jinyin Hu, Jiawei Zhou 0013, Minshan Xie, Zhonghao Yang 0003, Jing Li 0034, Huadi Zheng, Jie Shi 0005, Daojing He, Yu Li 0007 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2025 | Advancing Manga Analysis: Comprehensive Segmentation Annotations for the Manga109 DatasetabstractManga, a popular form of multimodal artwork, has traditionally been overlooked in deep learning advancements due to the absence of a robust dataset and comprehensive annotation. Manga segmentation is the key to the digital migration of manga. There exists a significant domain gap between the manga and the natural images, that fails most existing learning-based methods. To address this gap, we introduce an augmented segmentation annotation for the Manga109 dataset, a collection of 109 manga volumes, that offers intricate artworks in a rich variety of styles. We introduce a detailed annotation that extends beyond the original simple bounding boxes to the segmentation masks with pixel-level precision. It provides object category, location, and instance information that can be used for semantic segmentation and instance segmentation. We also provide a comprehensive analysis of our annotation dataset from various aspects. We further measure the improvement of the state-of-the-art segmentation model after training it with our augmented dataset. The benefits of this augmented dataset are profound, with the potential to significantly enhance manga analysis algorithms and catalyze the novel development in digital art processing and cultural analytics. This annotation, named MangaSeg, is publicly available at https://huggingface.co/datasets/MS92/MangaSegmentation. Minshan Xie, Hanyuan Liu, Chengze Li, Tien-Tsin Wong |
CVPR | 1 |
| 2025 | BlueNeg: A 35MM Negative Film Dataset for Restoring Channel-Heterogeneous Deterioration
Hanyuan Liu, Chengze Li, Minshan Xie, Zhenni Wang, Jiawen Liang, Andrew Chi-Sing Leung, Tien-Tsin Wong |
ICCV | 3 |
| 2025 | ColorDiffuser: Video Colorization with Pretrained Text-to-Image Diffusion Models
Hanyuan Liu, Minshan Xie, Jinbo Xing, Chengze Li, Andrew Chi-Sing Leung, Tien-Tsin Wong |
ACM Multimedia | 2 |
| 2025 | Synchronized Multi-Frame Diffusion for Temporally Consistent Video StylizationabstractAbstract Text‐guided video‐to‐video stylization transforms the visual appearance of a source video to a different appearance guided on textual prompts. Existing text‐guided image diffusion models can be extended for stylized video synthesis. However, they struggle to generate videos with both highly detailed appearance and temporal consistency. In this paper, we propose a synchronized multi‐frame diffusion framework to maintain both the visual details and the temporal consistency. Frames are denoised in a synchronous fashion, and more importantly, information of different frames is shared since the beginning of the denoising process. Such information sharing ensures that a consensus, in terms of the overall structure and color distribution, among frames can be reached in the early stage of the denoising process before it is too late. The optical flow from the original video serves as the connection, and hence the venue for information sharing, among frames. We demonstrate the effectiveness of our method in generating high‐quality and diverse results in extensive experiments. Our method shows superior qualitative and quantitative results compared to state‐of‐the‐art video editing methods. Minshan Xie, Hanyuan Liu, Chengze Li, Tien-Tsin Wong |
Comput. Graph. Forum | 1 |
| 2025 | Screentone-Preserved Manga RetargetingabstractAbstract As a popular comic style, manga offers a unique impression by utilizing a rich set ofbitonal patterns, or screentones, for illustration. However, screentones can easily be degraded when manga is resized in terms of aspect ratio and resolution for manga re‐layout and e‐manga migration applications. To tackle this problem, we propose the first automatic manga retargeting method that synthesizes a retargeted manga image while preserving the prominent structure and fine screentone intended by the manga artist. While modern natural photo retargeting methods can achieve prominent structure preservation, preserving screentones within arbitrarily shaped regions is very challenging due to two properties of manga: (i) pattern constancy under translation, and (ii) non‐compatibility with interpolation. To circumvent this barrier, we propose learning a quantized representation of screentones that is translation‐invariant and pointwisely representable through a tailored manga reconstruction network with a screentone‐anchored codebook. Thanks to these merits, we can perform the re‐synthesis operation using existing photo retargeting methods and achieve the desired manga retargeting results. We conducted extensive qualitative and quantitative experiments to validate the effectiveness of our method, and we achieved notably compelling results compared to alternative methods. Minshan Xie, Menghan Xia, Chengze Li, Xueting Liu 0001, Tien-Tsin Wong |
Comput. Graph. Forum | 1 |
| 2024 | SKETCH2MANGA: Shaded Manga Screening from Sketch with Diffusion ModelsabstractWhile manga is a popular entertainment form, creating manga is tedious, especially adding screentones to the created sketch, namely manga screening. Unfortunately, there is no existing method that tailors for automatic manga screening, probably due to the difficulty in generating shaded high-frequency screentones of high-quality. Classic manga screening approaches generally require user input to provide screentone exemplars or a reference manga image. Recent deep learning models enable automatic generation by learning from a large-scale dataset. However, the state-of-the-art models still fail to generate high-quality shaded screentones due to the lack of a tailored model and high-quality manga training data. In this paper, we propose a novel sketch-to-manga framework that first generates a color illustration from the sketch and then generates a screentoned manga based on the intensity guidance. Our method significantly outperforms existing methods in generating high-quality manga with shaded high-frequency screentones. Xueting Liu 0001, Chengze Li, Minshan Xie, Tien-Tsin Wong |
ICIP | 4 |
| 2024 | Text-Guided Texturing by Synchronized Multi-View DiffusionabstractThis paper introduces a novel approach to synthesize texture to dress up a 3D object, given a text prompt. Based on the pre-trained text-to-image (T2I) diffusion model, existing methods usually employ a project-and-inpaint approach, in which a view of the given object is first generated and warped to another view for inpainting. But it tends to generate inconsistent texture due to the asynchronous diffusion of multiple views. We believe that such asynchronous diffusion and insufficient information sharing among views are the root causes of the inconsistent artifacts. In this paper, we propose a synchronized multi-view diffusion approach that allows the diffusion processes from different views to reach a consensus on the generated content early in the process, and hence ensures the texture consistency. To synchronize the diffusion, we share the denoised content among different views in each denoising step, specifically by blending the latent content in the texture domain from overlapping views. Our method demonstrates superior performance in generating consistent, seamless and highly detailed textures, comparing to state-of-the-art methods. © 2024 Copyright held by the owner/author(s). Minshan Xie, Hanyuan Liu, Tien-Tsin Wong |
SIGGRAPH Asia | 2 |
| 2024 | LF2MV: Learning an Editable Meta-View Towards Light Field RepresentationabstractLight fields are 4D scene representations that are typically structured as arrays of views or several directional samples per pixel in a single view. However, this highly correlated structure is not very efficient to transmit and manipulate, especially for editing. To tackle this issue, we propose a novel representation learning framework that can encode the light field into a single meta-view that is both compact and editable. Specifically, the meta-view composes of three visual channels and a complementary meta channel that is embedded with geometric and residual appearance information. The visual channels can be edited using existing 2D image editing tools, before reconstructing the whole edited light field. To facilitate edit propagation against occlusion, we design a special editing-aware decoding network that consistently propagates the visual edits to the whole light field upon reconstruction. Extensive experiments show that our proposed method achieves competitive representation accuracy and meanwhile enables consistent edit propagation. Menghan Xia, Jose Echevarria, Minshan Xie, Tien-Tsin Wong |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2021 | Exploiting Aliasing for Manga RestorationabstractAs a popular entertainment art form, manga enriches the line drawings details with bitonal screentones. However, manga resources over the Internet usually show screen-tone artifacts because of inappropriate scanning/rescaling resolution. In this paper, we propose an innovative two-stage method to restore quality bitonal manga from de-graded ones. Our key observation is that the aliasing induced by downsampling bitonal screentones can be utilized as informative clues to infer the original resolution and screentones. First, we predict the target resolution from the degraded manga via the Scale Estimation Network (SE-Net) with spatial voting scheme. Then, at the target resolution, we restore the region-wise bitonal screentones via the Manga Restoration Network (MR-Net) discriminatively, depending on the degradation degree. Specifically, the original screentones are directly restored in pattern-identifiable regions, and visually plausible screentones are synthesized in pattern-agnostic regions. Quantitative evaluation on synthetic data and visual assessment on real-world cases illustrate the effectiveness of our method. Minshan Xie, Menghan Xia, Tien-Tsin Wong |
CVPR | 1 |
| 2021 | A Real-Time Mobile Application for Cattle Tracking using Video Captured from a DroneabstractIn many countries, governments have issued laws on cattle traceability, which require farmers to keep track of their livestock. While there are many algorithms to monitor cattle in confined indoor environments, monitoring cattle of a large number in outdoor pastures is still an open research problem. In recent years, with the advancement of unmanned aerial vehicles, e.g., drones, cattle monitoring can be based on images captured by drones. However, the challenges include image analysis in real-time, and keeping track of a dynamic scene (cattle movements) based on a moving sensor device. In this work, we develop an iOS application that can send captured herd images to our deep-learning-based server for cow segmentation and counting. Furthermore, the app can guide the operator to control the drone flight route and viewing perspectives. Chuyang Liu, Zihao Jian, Minshan Xie, Irene Cheng 0001 |
ISNCC | 3 |
| 2021 | Seamless manga inpainting with semantics awarenessabstractManga inpainting fills up the disoccluded pixels due to the removal of dialogue balloons or "sound effect" text. This process is long needed by the industry for the language localization and the conversion to animated manga. It is mostly done manually, as existing methods (mostly for natural image inpainting) cannot produce satisfying results. Manga inpainting is more tricky than natural image inpainting because its highly abstract illustration using structural lines and screentone patterns, which confuses the semantic interpretation and visual content synthesis. In this paper, we present the first manga inpainting method, a deep learning model, that generates high-quality results. Instead of direct inpainting, we propose to separate the complicated inpainting into two major phases, semantic inpainting and appearance synthesis. This separation eases both the feature understanding and hence the training of the learning model. A key idea is to disentangle the structural line and screentone, that helps the network to better distinguish the structural line and the screentone features for semantic interpretation. Both the visual comparison and the quantitative experiments evidence the effectiveness of our method and justify its superiority over existing state-of-the-art methods in the application of manga inpainting. Minshan Xie, Menghan Xia, Xueting Liu 0001, Chengze Li, Tien-Tsin Wong |
ACM Trans. Graph. | 1 |
| 2021 | Perceptual-Aware Sketch Simplification Based on Integrated VGG LayersabstractDeep learning has been recently demonstrated as an effective tool for raster-based sketch simplification. Nevertheless, it remains challenging to simplify extremely rough sketches. We found that a simplification network trained with a simple loss, such as pixel loss or discriminator loss, may fail to retain the semantically meaningful details when simplifying a very sketchy and complicated drawing. In this paper, we show that, with a well-designed multi-layer perceptual loss, we are able to obtain aesthetic and neat simplification results preserving semantically important global structures as well as fine details without blurriness and excessive emphasis on local structures. To do so, we design a multi-layer discriminator by fusing all VGG feature layers to differentiate sketches and clean lines. The weights used in layer fusing are automatically learned via an intelligent adjustment mechanism. Furthermore, to evaluate our method, we compare our method to state-of-the-art methods through multiple experiments, including visual comparison and intensive user study. Xuemiao Xu, Minshan Xie, Peiqi Miao, Wenpeng Xiao, Huaidong Zhang, Xueting Liu 0001, Tien-Tsin Wong |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2020 | Manga filling style conversion with screentone variational autoencoderabstractWestern color comics and Japanese-style screened manga are two popular comic styles. They mainly differ in the style of region-filling. However, the conversion between the two region-filling styles is very challenging, and manually done currently. In this paper, we identify that the major obstacle in the conversion between the two filling styles stems from the difference between the fundamental properties of screened region-filling and colored region-filling. To resolve this obstacle, we propose a screentone variational autoencoder, ScreenVAE, to map the screened manga to an intermediate domain. This intermediate domain can summarize local texture characteristics and is interpolative. With this domain, we effectively unify the properties of screening and color-filling, and ease the learning for bidirectional translation between screened manga and color comics. To carry out the bidirectional translation, we further propose a network to learn the translation between the intermediate domain and color comics. Our model can generate quality screened manga given a color comic, and generate color comic that retains the original screening intention by the bitonal manga artist. Several results are shown to demonstrate the effectiveness and convenience of the proposed method. We also demonstrate how the intermediate domain can assist other applications such as manga inpainting and photo-to-comic conversion. Minshan Xie, Chengze Li, Xueting Liu 0001, Tien-Tsin Wong |
ACM Trans. Graph. | 1 |
| 2017 | ASCII Art Synthesis from Natural PhotographsabstractWhile ASCII art is a worldwide popular art form, automatic generating structure-based ASCII art from natural photographs remains challenging. The major challenge lies on extracting the perception-sensitive structure from the natural photographs so that a more concise ASCII art reproduction can be produced based on the structure. However, due to excessive amount of texture in natural photos, extracting perception-sensitive structure is not easy, especially when the structure may be weak and within the texture region. Besides, to fit different target text resolutions, the amount of the extracted structure should also be controllable. To tackle these challenges, we introduce a visual perception mechanism of non-classical receptive field modulation (non-CRF modulation) from physiological findings to this ASCII art application, and propose a new model of non-CRF modulation which can better separate the weak structure from the crowded texture, and also better control the scale of texture suppression. Thanks to our non-CRF model, more sensible ASCII art reproduction can be obtained. In addition, to produce more visually appealing ASCII arts, we propose a novel optimization scheme to obtain the optimal placement of proportional-font characters. We apply our method on a rich variety of images, and visually appealing ASCII art can be obtained in all cases. Xuemiao Xu, Linyuan Zhong, Minshan Xie, Xueting Liu 0001, Harry Qin, Tien-Tsin Wong |
IEEE Trans. Vis. Comput. Graph. | 3 |