VLDB 2026 Research / reviewers in the wild / expert
Shisong Lin
dblp:244/8190
· DBLP profile ↗
5ranked-venue papers
3as first author
4since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 3 since 2021Security and privacy · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | UTDesign: A Unified Framework for Stylized Text Editing and Generation in Graphic Design ImagesabstractAI-assisted graphic design has emerged as a powerful tool for automating the creation and editing of design elements such as posters, banners, and advertisements. While diffusion-based text-to-image models have demonstrated strong capabilities in visual content generation, their text rendering performance, particularly for small-scale typography and non-Latin scripts, remains limited. In this paper, we propose UTDesign, a unified framework for high-precision stylized text editing and conditional text generation in design images, supporting both English and Chinese scripts. Our framework introduces a novel DiT-based text style transfer model trained from scratch on a synthetic dataset, capable of generating transparent RGBA text foregrounds that preserve the style of reference glyphs. We further extend this model into a conditional text generation framework by training a multi-modal condition encoder on a curated dataset with detailed text annotations, enabling accurate, style-consistent text synthesis conditioned on background images, prompts, and layout specifications. Finally, we integrate our approach into a fully automated text-to-design (T2D) pipeline by incorporating pre-trained text-to-image (T2I) models and an MLLM-based layout planner. Extensive experiments demonstrate that UTDesign achieves state-of-the-art performance among open-source methods in terms of stylistic consistency and text accuracy, and also exhibits unique advantages compared to proprietary commercial approaches. Code and data for this paper are available at https://github.com/ZYM-PKU/UTDesign. Yuanpeng Gao, Jiwei Duan, Shisong Lin, Longfei Xiong, Zhouhui Lian |
SIGGRAPH Asia | 5 |
| 2021 | PointFace: Point Set Based Feature Learning for 3D Face RecognitionabstractThough 2D face recognition (FR) has achieved great success due to powerful 2D CNNs and large-scale training data, it is still challenged by extreme poses and illumination conditions. On the other hand, 3D FR has the potential to deal with aforementioned challenges in the 2D domain. However, most of available 3D FR works transform 3D surfaces to 2D maps and utilize 2D CNNs to extract features. The works directly processing point clouds for 3D FR is very limited in literature. To bridge this gap, in this paper, we propose a light-weight framework, named PointFace, to directly process point set data for 3D FR. Inspired by contrastive learning, our PointFace use two weight-shared encoders to directly extract features from a pair of 3D faces. A feature similarity loss is designed to guide the encoders to obtain discriminative face representations. We also present a pair selection strategy to generate positive and negative pairs to boost training. Extensive experiments on Lock3DFace and Bosphorus show that the proposed PointFace outperforms state-of-the-art 2D CNN based methods. Changyuan Jiang, Shisong Lin, Wei Chen 0092, Feng Liu 0013, LinLin Shen |
IJCB | 2 |
| 2021 | High Quality Facial Data Synthesis and Fusion for 3D Low-quality Face Recognitionabstract3D face recognition (FR) is a popular topic in computer vision, since 3D face data is invariant to pose and illumination condition changes which easily affect the performance of 2D FR. Though many 3D solutions have achieved impressive performances on public high-quality 3D face databases, few works concentrate on low-quality 3D FR. As the quality of 3D face acquired by widely used low-cost RGB-D sensors is really low, more robust methods are required to achieve satisfying performance on these 3D face data. To address this issue, we propose a novel two-stage pipeline to improve the performance of 3D FR. In the first stage, we utilize pix2pix network to restore the quality of low-quality face. In the second stage, we launch a multi-quality fusion network (MQFNet) to fuse the features from different qualities and enhance FR performance. Our proposed network achieves the state-of-the-art performance on the Lock3DFace database. Furthermore, extensive controlled experiments are conducted to demonstrate the effectiveness of each model of our network. Shisong Lin, Changyuan Jiang, Feng Liu 0013, LinLin Shen |
IJCB | 1 |
| 2021 | Orthogonalization-Guided Feature Fusion Network for Multimodal 2D+3D Facial Expression RecognitionabstractAs 2D and 3D data present different views of the same face, the features extracted from them can be both complementary and redundant. In this paper, we present a novel and efficient orthogonalization-guided feature fusion network, namely OGF$^2$Net, to fuse the features extracted from 2D and 3D faces for facial expression recognition. While 2D texture maps are fed into a 2D feature extraction pipeline (FE2DNet), the attribute maps generated from 3D data are concatenated as input of the 3D feature extraction pipeline (FE3DNet). The two networks are separately trained at the first stage and frozen in the second stage for late feature fusion, which can well address the unavailability of a large number of 3D+2D face pairs. To reduce the redundancies among features extracted from 2D and 3D streams, we design an orthogonal loss-guided feature fusion network to orthogonalize the features before fusing them. Experimental results show that the proposed method significantly outperforms the state-of-the-art algorithms on both the BU-3DFE and Bosphorus databases. While accuracies as high as 89.05% (P1 protocol) and 89.07% (P2 protocol) are achieved on the BU-3DFE database, an accuracy of 89.28% is achieved on the Bosphorus database. The complexity analysis also suggests that our approach achieves a higher processing speed while simultaneously requiring lower memory costs. Shisong Lin, Mengchao Bai, Feng Liu 0013, LinLin Shen, Yicong Zhou |
IEEE Trans. Multim. | 1 |
| 2019 | Local Feature Tensor Based Deep Learning for 3D Face RecognitionabstractA local feature tensor similarity based deep learning approach is proposed in this paper for 3D face recognition. Once a set of salient points on the 3D mesh are detected, three scale and rotation invariant features are extracted to represent local surface around each salient point. The local features of all the salient points are concatenated to produce a 3rdorder feature tensor to represent a 3D face. Similarity of two 3D faces can thus be measured by a similarity tensor calculated using the two feature tensors. To address the unavailability of large 3D face samples, a feature tensor based data augmentation approach is proposed to augment the number of feature tensors. Experimental results show that the ResNet model trained using the augmented feature tensors achieves the best performance among state of the art competitors, i.e. 99.71% and 96.2% accuracy are achieved for Bosphorus and BU3DFE database, respectively. Shisong Lin, Feng Liu 0013, LinLin Shen |
FG | 1 |