VLDB 2026 Research / reviewers in the wild / expert
Yan Kong
dblp:25/11202
· DBLP profile ↗
32ranked-venue papers
11as first author
14since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 5 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 first-author · 3 since 2021Systems, architecture and hardware · 4 · 2 first-authorSoftware engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | "Privacy across the boundary": Examining Perceived Privacy Risk Across Data Transmission and Sharing Ranges of Smart Home Personal AssistantsabstractAs Smart Home Personal Assistants (SPAs) evolve into social agents, understanding user privacy necessitates interpersonal communication frameworks, such as Privacy Boundary Theory (PBT). To ground our investigation, our three-phase preliminary study (1) identified transmission and sharing ranges as key boundary-related risk factors, (2) categorized relevant SPA functions and data types, and (3) analyzed commercial practices, revealing widespread data sharing and non-transparent safeguards. A subsequent mixed-methods study (N=412 survey, N=40 interviews among the survey participants) assessed users’ perceived privacy risks across data types, transmission ranges and sharing ranges. Results demonstrate a significant, non-linear escalation in perceived risk when data crosses two critical boundaries: the ‘public network’ (transmission) and ‘third parties’ (sharing). This boundary effect holds robustly across data types and demographics. Furthermore, risk perception is modulated by data attributes (e.g., social relational data), and contextual privacy calculus. Conversely, anonymization safeguards show limited efficacy especially for third-party sharing, a finding attributed to user distrust. These findings empirically ground PBT in the SPA context and inform design of boundary-aware privacy protection. Haobin Xing, Yan Kong, Xin Yi 0001, Kanye Ye Wang, Hewu Li |
CHI | 5 |
| 2026 | VisGuardian: A Lightweight Group-based Visual Privacy Control Technique For Smart Glasses in Home EnvironmentsabstractAlways-on sensing of AI applications on AR glasses makes traditional permission techniques inefficient for context-dependent private visual data within home environments. Home presents a challenging privacy context due to massive sensitive objects and the intimate nature of daily routines. We propose VisGuardian, a fine-grained content-based visual permission technique for AR glasses. VisGuardian features a group-based control mechanism that enables users to efficiently manage permissions for multiple private objects. VisGuardian detects objects using YOLO and adopts a pre-classified schema to group them. By selecting a single object, users can obscure groups of related objects based on criteria including privacy sensitivity, object category, or spatial proximity. A technical evaluation shows VisGuardian achieves mAP50 of 0.6704 with only 14.0 ms latency and a 1.7% increase in battery consumption per hour. Furthermore, a user study (N=24) comparing VisGuardian to slider-based and object-based baselines found it to be significantly faster for setting permissions and was preferred by users for its efficiency, effectiveness, and ease of use. Qucheng Zang, Yongquan Hu, Jiachen Du, Yan Kong, Xinyi Fu 0003, Suranga Nanayakkara, Xin Yi 0001, Hewu Li |
CHI | 6 |
| 2026 | DPLoRA: Stable Diffusion Model distillation via Dynamic Parallel LoRA branches
Weibin Zeng, Yan Kong, Fuzhang Wu, Youheng Ren, Sicheng Shen, Yuqing Fan, Weiming Dong |
Pattern Recognit. | 2 |
| 2025 | Algorithmic Inversion: A Learnable Algorithm Representation for Code GenerationabstractThe prevalent fine-tuning paradigm for large language models (LLMs) has demonstrated strong performance on various code generation tasks. However, these models still fall short when confronted with algorithmic programming problems, where precise algorithmic reasoning is required. Humans typically adopt diverse algorithmic techniques to tackle complex programming problems, enabling general analysis and accurate implementation. Building on this observation, we propose a method that learns compact, LLM-friendly representations of algorithmic knowledge, termed Algorithmic Inversion (AI), which aims to aid LLMs in understanding programming problems. Specifically, we apply a lightweight fine-tuning process on codeoriented models to automatically learn algorithm embeddings. When concatenated with the inputs, the algorithm embeddings act as instructive signals, guiding LLMs in generating correct code solutions by providing contextual algorithmic hints. We apply our approach to models of three different parameter sizes and evaluate them on three algorithmic programming benchmarks. Our extensive experiments show that applying AI to small (1.5B parameters) models results in absolute improvements of up to 1.8 on Pass@1, while large models (1 5 B parameters) achieve improvements of up to 1.4, compared to Prompt-Tuning. Additionally, our method outperforms traditional full fine-tuning approaches by a significant margin across all tested benchmarks. Furthermore, our analysis of the generated code reveals that AI effectively enhances the model's problem-solving process by providing clear algorithmic guidance. Codes and datasets are available11https://github.com/joeysbase/Algorithmic-Inversion . Zhongyi Shi, Fuzhang Wu, Weibin Zeng, Yan Kong, Sicheng Shen |
ICPC | 4 |
| 2025 | Query-Level Alignment for End-to-End Lesion Detection with Human Gaze
Yan Kong, Zhixiang Peng, Yonghao Li, Jiangdong Cai, Sheng Wang 0014, Qian Wang 0001, Yuqi Fang, Caifeng Shan |
MICCAI (13) | 1 |
| 2024 | Gaze-DETR: Using Expert Gaze to Reduce False Positives in Vulvovaginal Candidiasis Screening
Yan Kong, Sheng Wang 0014, Jiangdong Cai, Zihao Zhao 0002, Zhenrong Shen 0001, Yonghao Li, Manman Fei, Qian Wang 0001 |
MICCAI (4) | 1 |
| 2024 | Transformer with convolution and graph-node co-embedding: An accurate and interpretable vision backbone for predicting gene expressions from local histopathological imageabstractInferring gene expressions from histopathological images has long been a fascinating yet challenging task, primarily due to the substantial disparities between the two modality. Existing strategies using local or global features of histological images are suffering model complexity, GPU consumption, low interpretability, insufficient encoding of local features, and over-smooth prediction of gene expressions among neighboring sites. In this paper, we develop TCGN (Transformer with Convolution and Graph-Node co-embedding method) for gene expression estimation from H&E-stained pathological slide images. TCGN comprises a combination of convolutional layers, transformer encoders, and graph neural networks, and is the first to integrate these blocks in a general and interpretable computer vision backbone. Notably, TCGN uniquely operates with just a single spot image as input for histopathological image analysis, simplifying the process while maintaining interpretability. We validate TCGN on three publicly available spatial transcriptomic datasets. TCGN consistently exhibited the best performance (with median PCC 0.232). TCGN offers superior accuracy while keeping parameters to a minimum (just 86.241 million), and it consumes minimal memory, allowing it to run smoothly even on personal computers. Moreover, TCGN can be extended to handle bulk RNA-seq data while providing the interpretability. Enhancing the accuracy of omics information prediction from pathological images not only establishes a connection between genotype and phenotype, enabling the prediction of costly-to-measure biomarkers from affordable histopathological images, but also lays the groundwork for future multi-modal data modeling. Our results confirm that TCGN is a powerful tool for inferring gene expressions from histopathological images in precision health applications. Yan Kong, Ronghan Li, Zuoheng Wang, Hui Lu 0004 |
Medical Image Anal. | 2 |
| 2023 | Modeling the Trade-off of Privacy Preservation and Activity Recognition on Low-Resolution ImagesabstractA computer vision system using low-resolution image sensors can provide intelligent services (e.g., activity recognition) but preserve unnecessary visual privacy information from the hardware level. However, preserving visual privacy and enabling accurate machine recognition have adversarial needs on image resolution. Modeling the trade-off of privacy preservation and machine recognition performance can guide future privacy-preserving computer vision systems using low-resolution image sensors. In this paper, using the at-home activity of daily livings (ADLs) as the scenario, we first obtained the most important visual privacy features through a user survey. Then we quantified and analyzed the effects of image resolution on human and machine recognition performance in activity recognition and privacy awareness tasks. We also investigated how modern image super-resolution techniques influence these effects. Based on the results, we proposed a method for modeling the trade-off of privacy preservation and activity recognition on low-resolution images. Yuntao Wang 0001, Zirui Cheng, Xin Yi 0001, Yan Kong, Xuhai Xu, Yukang Yan, Chun Yu, Shwetak N. Patel, Yuanchun Shi |
CHI | 4 |
| 2023 | Squeez'In: Private Authentication on Smartphones based on Squeezing GesturesabstractIn this paper, we proposed Squeez’In, a technique on smartphones that enabled private authentication by holding and squeezing the phone with a unique pattern. We first explored the design space of practical squeezing gestures for authentication by analyzing the participants’ self-designed gestures and squeezing behavior. Results showed that varying-length gestures with two levels of touch pressure and duration were the most natural and unambiguous. We then implemented Squeez’In on an off-the-shelf capacitive sensing smartphone, and employed an SVM-GBDT model for recognizing gestures and user-specific behavioral patterns, achieving 99.3% accuracy and 0.93 F1-score when tested on 21 users. A following 14-day study validated the memorability and long-term stability of Squeez’In. During usability evaluation, compared with gesture and pin code, Squeez’In achieved significantly faster authentication speed and higher user preference in terms of privacy and security. Xin Yi 0001, Louisa Shi, Fengyan Han, Yan Kong, Hewu Li, Yuanchun Shi |
CHI | 6 |
| 2023 | Superresolved spatial transcriptomics transferred from a histological context
Xiaocheng Zhou, Yan Kong, Hui Lu 0004 |
Appl. Intell. | 3 |
| 2023 | Analysis of the influence of population distribution characteristics on swarm intelligence optimization algorithms
Rongxin Hu, Liyong Bao, Yan Kong |
Inf. Sci. | 5 |
| 2023 | Application of DQN-IRL Framework in Doudizhu's Sparse Reward
Yan Kong, Hongyuan Shi, Xiaocong Wu, Yefeng Rui |
Neural Process. Lett. | 1 |
| 2023 | Critical path-driven exploration on the MiniGrid environment
Yan Kong, Yu Dou |
Serv. Oriented Comput. Appl. | 1 |
| 2021 | Deep Gaussian Mixture Model on Multiple Interpretable Features of Fetal Heart Rate for Pregnancy Wellness
Yan Kong, Bin Xu 0001, Bowen Zhao 0004, Ji Qi 0003 |
PAKDD (1) | 1 |
| 2019 | Learning to Film From Professional Human Motion VideosabstractWe investigate the problem of 6 degrees of freedom (DOF) camera planning for filming professional human motion videos using a camera drone. Existing methods either plan motions for only a pan-tilt-zoom (PTZ) camera, or adopt ad-hoc solutions without carefully considering the impact of video contents and previous camera motions on the future camera motions. As a result, they can hardly achieve satisfactory results in our drone cinematography task. In this study, we propose a learning-based framework which incorporates the video contents and previous camera motions to predict the future camera motions that enable the capture of professional videos. Specifically, the inputs of our framework are video contents which are represented using subject-related feature based on 2D skeleton and scene-related features extracted from background RGB images, and camera motions which are represented using optical flows. The correlation between the inputs and output future camera motions are learned via a sequence-to-sequence convolutional long short-term memory (Seq2Seq ConvLSTM) network from a large set of video clips. We deploy our approach to a real drone cinematography system by first predicting the future camera motions, and then converting them to the drone's control commands via an odometer. Our experimental results on extensive datasets and showcases exhibit significant improvements in our approach over conventional baselines and our approach can successfully mimic the footage of a professional cameraman. Chong Huang 0005, David Chuan-En Lin, Yan Kong, Peng Chen 0008, Xin Yang 0008, Kwang-Ting Cheng |
CVPR | 4 |
| 2019 | Learning to Capture a Film-Look Video with a Camera DroneabstractThe development of intelligent drones has simplified aerial filming and provided smarter assistant tools for users to capture a film-look footage. Existing methods of autonomous aerial filming either specify predefined camera movements for a drone to capture a footage, or employ heuristic approaches for camera motion planning. However, both predefined movements and heuristically planned motions are hardly able to provide cinematic footages for various dynamic scenarios. In this paper, we propose a data-driven learning-based approach, which can imitate a professional cameraman's intention for capturing a film-look aerial footage of a single subject in real-time. We model the decision-making process of the cameraman with two steps: 1) we train a network to predict the future image composition and camera position, and 2) our system then generates control commands to achieve the desired shot framing. At the system level, we deploy our algorithm on the limited resources of a drone and demonstrate the feasibility of running automatic filming onboard in real-time. Our experiments show how our data-driven planning approach achieves film-look footages and successfully mimics the work of a professional cameraman. Chong Huang 0005, Yan Kong, Peng Chen 0008, Xin Yang 0008, Kwang-Ting Cheng |
ICRA | 3 |
| 2019 | Empirical Evaluation of Deep Learning-Based Travel Time Prediction
Mengyan Wang, Weihua Li 0007, Yan Kong, Quan Bai 0001 |
PKAW | 3 |
| 2019 | Gradient-aware blind face inpainting for deep face verification
Fuzhang Wu, Yan Kong, Weiming Dong |
Neurocomputing | 2 |
| 2018 | Through-the-Lens Drone FilmingabstractAerial filming in action scenes using a drone is difficult for inexperienced flyers because manipulating a remote controller and meeting the desired image composition are two independent, while concurrent, tasks. Existing systems attempt to utilize wearable GPS-based or infrared-based sensors to track the human movement and to assist in capturing footage. However, these sensors work only in either indoor (infrared-based) or outdoor environments (GPS-based), but not both. In this paper, we introduce a novel drone filming system which integrates monocular 3D human pose estimation and localization into a drone platform to remove the constraints imposed by wearable-sensor-based solutions. Meanwhile, given the estimated position, we propose a novel drone control system, called “through-the-lens drone filming”, to allow a cameraman to conveniently control the drone by manipulating a 3D model in the preview, which closes the gap between the flight control and the viewpoint design. Our system includes two key enabling techniques: 1) subject localization based on visual-inertial fusion, and 2) through-the-lens camera planning. This is the first drone camera system which allows users to capture human actions by manipulating the camera in a virtual environment. From the drone hardware, we integrate a gimbal camera and two GPUs into the limited space of a drone and demonstrate the feasibility of running the entire system onboard with insignificant delays, which are sufficient for filming in our real-time application. Experimental results, in both simulation and real-world scenarios, demonstrate that our techniques can greatly ease camera control and capture better videos. Chong Huang 0005, Yan Kong, Peng Chen 0008, Xin Yang 0008, Kwang-Ting Cheng |
IROS | 3 |
| 2018 | An intelligent agent-based method for task allocation in competitive cloud environmentsabstractSummary In market‐based cloud environments, both resource consumers and providers are self‐interested; additionally, they can come and leave the environment freely. Therefore, the environment is competitive and uncertain. Because of the competition, participants may cheat in making deals, and this represents that the environment is insecure to resource providers who intend to earn profits through renting their resources to the tasks of resource consumers. Against this, in this paper, intelligent agents are designed to strategically quote for the tasks that they are interested in, on behalf of resource providers. Agents could quote according to the messages it obtained and the information learnt and predicted from the messages, to minimize the influence of insecure factors, such as cheating, competition, and dynamism. The experimental evaluation shows that the proposed method outperforms both a well‐known multiresource negotiation‐based task allocation method and a max‐sum belief propagation–based method. Yan Kong, Minjie Zhang 0001, Dayong Ye, Jinxiu Zhu |
Concurr. Comput. Pract. Exp. | 1 |
| 2017 | A belief propagation-based method for task allocation in open and dynamic cloud environments
Yan Kong, Minjie Zhang 0001, Dayong Ye |
Knowl. Based Syst. | 1 |
| 2016 | An Auction-Based Approach for Group Task Allocation in an Open Network EnvironmentabstractTo solve the problem of group task allocation with time constraints in open and dynamic network environments, this paper proposes a decentralized combinatorial auction-based approach for group task allocation. In the proposed approach, both resource providers and consumers are modeled as intelligent agents. The proposed approach is decentralized, so all the agents are limited to communicating with their neighboring agents. The proposed approach also allows agents to enter and leave the network environments freely, and is robust for the dynamism and openness of the network environments. Tasks in the proposed approach have deadlines, and may need the collaboration of a group of self-interested providers. The experimental results demonstrate that the proposed approach outperforms two well-known task allocation approaches in terms of success rate of task allocation, the individual utility of the agents, the speed of task allocation and scalability. Yan Kong, Minjie Zhang 0001, Dayong Ye |
Comput. J. | 1 |
| 2016 | Category co-occurrence modeling for large scale scene recognition
Xinhang Song, Shuqiang Jiang, Luis Herranz, Yan Kong, Kai Zheng 0001 |
Pattern Recognit. | 4 |
| 2016 | Image Retargeting by Texture-Aware SynthesisabstractReal-world images usually contain vivid contents and rich textural details, which will complicate the manipulation on them. In this paper, we design a new framework based on exampled-based texture synthesis to enhance content-aware image retargeting. By detecting the textural regions in an image, the textural image content can be synthesized rather than simply distorted or cropped. This method enables the manipulation of textural & non-textural regions with different strategies since they have different natures. We propose to retarget the textural regions by example-based synthesis and non-textural regions by fast multi-operator. To achieve practical retargeting applications for general images, we develop an automatic and fast texture detection method that can detect multiple disjoint textural regions. We adjust the saliency of the image according to the features of the textural regions. To validate the proposed method, comparisons with state-of-the-art image retargeting techniques and a user study were conducted. Convincing visual results are shown to demonstrate the effectiveness of the proposed method. Weiming Dong, Fuzhang Wu, Yan Kong, Xing Mei, Tong-Yee Lee, Xiaopeng Zhang 0001 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2016 | Measuring and Predicting Visual Importance of Similar ObjectsabstractSimilar objects are ubiquitous and abundant in both natural and artificial scenes. Determining the visual importance of several similar objects in a complex photograph is a challenge for image understanding algorithms. This study aims to define the importance of similar objects in an image and to develop a method that can select the most important instances for an input image from multiple similar objects. This task is challenging because multiple objects must be compared without adequate semantic information. This challenge is addressed by building an image database and designing an interactive system to measure object importance from human observers. This ground truth is used to define a range of features related to the visual importance of similar objects. Then, these features are used in learning-to-rank and random forest to rank similar objects in an image. Importance predictions were validated on 5,922 objects. The most important objects can be identified automatically. The factors related to composition (e.g., size, location, and overlap) are particularly informative, although clarity and color contrast are also important. We demonstrate the usefulness of similar object importance on various applications, including image retargeting, image compression, image re-attentionizing, image admixture, and manipulation of blindness images. Yan Kong, Weiming Dong, Xing Mei, Chongyang Ma, Tong-Yee Lee, Siwei Lyu, Feiyue Huang, Xiaopeng Zhang 0001 |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2016 | Feature-aware natural texture synthesis
Fuzhang Wu, Weiming Dong, Yan Kong, Xing Mei, Dong-Ming Yan 0001, Xiaopeng Zhang 0001, Jean-Claude Paul |
Vis. Comput. | 3 |
| 2015 | Evaluating the Quality of Face Alignment without Ground TruthabstractThe study of face alignment has been an area of intense research in computer vision, with its achievements widely used in computer graphics applications. The performance of various face alignment methods is often image-dependent or somewhat random because of their own strategy. This study aims to develop a method that can select an input image with good face alignment results from many results produced by a single method or multiple ones. The task is challenging because different face alignment results need to be evaluated without any ground truth. This study addresses this problem by designing a feasible feature extraction scheme to measure the quality of face alignment results. The feature is then used in various machine learning algorithms to rank different face alignment results. Our experiments show that our method is promising for ranking face alignment results and is able to pick good face alignment results, which can enhance the overall performance of a face alignment method with a random strategy. We demonstrate the usefulness of our ranking-enhanced face alignment algorithm in two practical applications: face cartoon stylization and digital face makeup. Kekai Sheng, Weiming Dong, Yan Kong, Xing Mei, Chengjie Wang 0001, Feiyue Huang, Bao-Gang Hu |
Comput. Graph. Forum | 3 |
| 2015 | A negotiation-based method for task allocation with time constraints in open grid environmentsabstractSummary This paper addresses the task allocation problem in an open, dynamic grid environments and service‐oriented environments. In such environments, both grid/service providers and consumers can be modelled as intelligent agents. These agents can leave and enter the environment freely at any time. Task allocation under time constraints becomes a challenging issue in such environments because it is difficult to apply a central controller during the allocation process due to the openness and dynamism of the environments. This paper proposes a negotiation‐based method for task allocation under time constraints in an open, dynamic grid environment, where both consumer and provider agents can freely enter or leave the environment. In this method, there is no central controller available, and agents negotiate with each other for task allocation based only on local views. The experimental results show that the proposed method can outperform the current methods in terms of the success rate of task allocation and the total profit obtained from the allocated tasks by agents under different time constraints. Copyright © 2014 John Wiley & Sons, Ltd. Yan Kong, Minjie Zhang 0001, Dayong Ye |
Concurr. Comput. Pract. Exp. | 1 |
| 2014 | Summarization-Based Image Resizing by Intelligent Object CarvingabstractImage resizing can be more effectively achieved with a better understanding of image semantics. In this paper, similar patterns that exist in many real-world images are analyzed. By interactively detecting similar objects in an image, the image content can be summarized rather than simply distorted or cropped. This method enables the manipulation of image pixels or patches as well as semantic objects in the scene during image resizing process. Given the special nature of similar objects in a general image, the integration of a novel object carving (OC) operator with the multi-operator framework is proposed for summarizing similar objects. The object removal sequence in the summarization strategy directly affects resizing quality. The method by which to evaluate the visual importance of the object as well as to optimally select the candidates for object carving is demonstrated. To achieve practical resizing applications for general images, a template matching-based method is developed. This method can detect similar objects even when they are of various colors, transformed in terms of perspective, or partially occluded. To validate the proposed method, comparisons with state-of-the-art resizing techniques and a user study were conducted. Convincing visual results are shown to demonstrate the effectiveness of the proposed method. Weiming Dong, Tong-Yee Lee, Fuzhang Wu, Yan Kong, Xiaopeng Zhang 0001 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2013 | Content-Based Colour TransferabstractAbstract This paper presents a novel content‐based method for transferring the colour patterns between images. Unlike previous methods that rely on image colour statistics, our method puts an emphasis on high‐level scene content analysis. We first automatically extract the foreground subject areas and background scene layout from the scene. The semantic correspondences of the regions between source and target images are established. In the second step, the source image is re‐coloured in a novel optimization framework, which incorporates the extracted content information and the spatial distributions of the target colour styles. A new progressive transfer scheme is proposed to integrate the advantages of both global and local transfer algorithms, as well as avoid the over‐segmentation artefact in the result. Experiments show that with a better understanding of the scene contents, our method well preserves the spatial layout, the colour distribution and the visual coherence in the transfer process. As an interesting extension, our method can also be used to re‐colour video clips with spatially‐varied colour effects. Fuzhang Wu, Weiming Dong, Yan Kong, Xing Mei, Jean-Claude Paul, Xiaopeng Zhang 0001 |
Comput. Graph. Forum | 3 |
| 2013 | SimLocator: robust locator of similar objects in images
Yan Kong, Weiming Dong, Xing Mei, Xiaopeng Zhang 0001, Jean-Claude Paul |
Vis. Comput. | 1 |
| 1999 | An adaptive critic approach for self-learning stock tradingabstractThe paper describes a stock trading system with self learning capability using adaptive critic designs. The same approach can be formulated for other financial applications such as trading of bonds, options, futures, commodities, and the like. Derong Liu 0001, Yan Kong, Edward G. Luxford |
CIFEr | 2 |