EDBT 2026 Demo / reviewers in the wild / expert
Jiayan Qiu
dblp:174/1895
· DBLP profile ↗
17ranked-venue papers
4as first author
12since 2021 · last 2026
0000-0002-1807-6666ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 4 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 3 first-author · 8 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Remodeling Semantic Relationships in Vision-Language Fine-TuningabstractVision-language fine-tuning has emerged as an efficient paradigm for constructing multimodal foundation models. While textual context often highlights semantic relationships within an image, existing fine-tuning methods typically overlook this information when aligning vision and language, thus leading to suboptimal performance. Toward solving this problem, we propose a method that can improve multimodal alignment and fusion based on both semantics and relationships.Specifically, we first extract multilevel semantic features from different vision encoder to capture more visual cues of the relationships. Then, we learn to project the vision features to group related semantics, among which are more likely to have relationships. Finally, we fuse the visual features with the textual by using inheritable cross-attention, where we globally remove the redundant visual relationships by discarding visual-language feature pairs with low correlation. We evaluate our proposed method on eight foundation models and two downstream tasks, visual question answering and image captioning, and show that it outperforms all existing methods. Liu Liu 0014, Baosheng Yu, Jiayan Qiu |
AAAI | 4 |
| 2026 | Understanding Interaction as You Need: Intention-Driven Pedestrian Behavior PredictionabstractPrediction of pedestrian behavior is crucial for autonomous driving systems and intelligent transportation.Conventional methods predict the behavior based solely on either the pedestrian intention or the distance-related interactions between the pedestrian and its surroundings. However, these methods overlook the associations between intention and interaction for behavior prediction, in which they should be aligned with each other, thus leading to sub-optimal predictions. To solve this problem, we propose to predict the behavior by learning the association between intention and interaction, enabling them to mutually enhance each other during the prediction. Specifically, we first predict the short-term intention of all objects, including the target pedestrian and its surroundings.Then, instead of using the distance-related interactions, we predict the interactions by learning the correlated intentions. Finally, the intention-driven interactions refine the initial intention prediction, thus ensuring the alignment between intention and interaction for behavior prediction. We evaluate our method on two downstream tasks, the pedestrian trajectory prediction and pedestrian intention estimation, and show that it outperforms all the existing methods. Hang Yu 0006, Yansen Yu, Jiayan Qiu |
AAAI | 3 |
| 2025 | Pedestrian Trajectory Prediction Driven by Bidirectional Intention-InteractionabstractIn complex environments, humans estimate the surrounding interactions to plan efficient trajectories. Accurately assessing the relationships and intensity of these interactions is crucial to optimize trajectory prediction. Existing methods implicitly integrate interaction features into trajectory representations, making it difficult to disentangle their influence and leading to sensitivity to indirect factors such as distance. We note that interactions and intentions are strongly correlated, as dynamic interactions among pedestrians are guided by their respective intentions and, in turn, influence these intentions. Therefore, we propose a novel intention-interaction bidirectionally driven pedestrian trajectory prediction network (I2I-Net). The network decomposes future trajectories into short-term intention-guided segments and iteratively refines intentions based on interaction estimation at each time step. Our proposed method achieves state-of-the-art performance in experiments on the ETH & UCY and SDD datasets, validating the method’s ability to precisely model the influence of human interactions on pedestrian trajectories. Hang Yu 0006, Yansen Yu, Jiayan Qiu |
ICME | 3 |
| 2025 | Controllable Data Generation with Hierarchical Neural RepresentationsabstractImplicit Neural Representations (INRs) represent data as continuous functions using the parameters of a neural network, where data information is encoded in the parameter space. Therefore, modeling the distribution of such parameters is crucial for building generalizable INRs. Existing approaches learn a joint distribution of these parameters via a latent vector to generate new data, but such a flat latent often fails to capture the inherent hierarchical structure of the parameter space, leading to entangled data semantics and limited control over the generation process. Here, we propose a Controllable Hierarchical Implicit Neural Representation (CHINR) framework, which explicitly models conditional dependencies across layers in the parameter space. Our method consists of two stages: In Stage-1, we construct a Layers-of-Experts (LoE) network, where each layer modulates distinct semantics through a unique latent vector, enabling disentangled and expressive representations. In Stage-2, we introduce a Hierarchical Conditional Diffusion Model (HCDM) to capture conditional dependencies across layers, allowing for controllable and hierarchical data generation at various semantic granularities. Extensive experiments across different modalities demonstrate that CHINR improves generalizability and offers flexible hierarchical control over the generated content. Sheyang Tang, Jiayan Qiu, Zhou Wang 0001 |
ICML | 3 |
| 2025 | Dynamic Shadow Unveils Invisible Semantics for Video OutpaintingabstractConventional video outpainting methods primarily focus on maintaining coherent textures and visual consistency across frames.
However, they often fail at handling dynamic scenes due to the complex motion of objects or camera movement, leading to temporal incoherence and visible flickering artifacts across frames. This is primarily because they lack instance-aware modeling to accurately separate and track individual object motions throughout the video. In this paper, we propose a novel video outpainting framework that explicitly takes shadow-object pairs into consideration to enhance the temporal and spatial consistency of instances, even when they are temporarily invisible. Specifically, we first track the shadow-object pairs across frames and predict the instances in the scene to unveil the spatial regions of invisible instances. Then, these prediction results are fed to guide the instance-aware optical flow completion to unveil the temporal motion of invisible instances. Next, these spatiotemporal guidances of instances are used to guide the video outpainting process. Finally, a video-aware discriminator is implemented to enhance alignment among dynamic shadows and the extended semantics in the scene. Comprehensive experiments underscore the superiority of our approach, outperforming existing state-of-the-art methods in widely recognized benchmarks. Hang Yu 0006, Jiayan Qiu |
NeurIPS | 3 |
| 2025 | SIGMA: Refining Large Language Model Reasoning via Sibling-Guided Monte Carlo AugmentationabstractEnhancing large language models by simply scaling up datasets has begun to yield diminishing returns, shifting the spotlight to data quality. Monte Carlo Tree Search (MCTS) has emerged as a powerful technique for generating high-quality chain-of-thought data, yet conventional approaches typically retain only the top-scoring trajectory from the search tree, discarding sibling nodes that often contain valuable partial insights, recurrent error patterns, and alternative reasoning strategies. This unconditional rejection of non-optimal reasoning branches may waste vast amounts of informative data in the whole search tree. We propose SIGMA (Sibling Guided Monte Carlo Augmentation), a novel framework that reintegrates these discarded sibling nodes to refine LLM reasoning. SIGMA forges semantic links among sibling nodes along each search path and applies a two-stage refinement: a critique model identifies overlooked strengths and weaknesses across the sibling set, and a revision model conducts text-based backpropagation to refine the top-scoring trajectory in light of this comparative feedback. By recovering and amplifying the underutilized but valuable signals from non-optimal reasoning branches, SIGMA substantially improves reasoning trajectories. On the challenging MATH benchmark, our SIGMA-tuned 7B model achieves 54.92\% accuracy using only 30K samples, outperforming state-of-the-art models trained on 590K samples. This result highlights that our sibling-guided optimization not only significantly reduces data usage but also significantly boosts LLM reasoning. Yanwei Ren, Fuxiang Wu, Jiayan Qiu, Jiaxing Huang 0001, Baosheng Yu, Liu Liu 0014 |
NeurIPS | 4 |
| 2024 | Shadow-Enlightened Image OutpaintingabstractConventional image outpainting methods usually treat unobserved areas as unknown and extend the scene only in terms of semantic consistency, thus overlooking the hidden information in shadows cast by unobserved areas, such as the invisible shapes and semantics. In this paper, we propose to extract and utilize the hidden information of un-observed areas from their shadows to enhance image out-painting. To this end, we propose an end-to-end deep approach that explicitly looks into the shadows within the image. Specifically, we extract shadows from the input image and identify instance-level shadow regions cast by the un-observed areas. Then, the instance-level shadow representations are concatenated to predict the scene layout of each unobserved instance and outpaint the unobserved areas. Finally, two discriminators are implemented to enhance alignment between the extended semantics and their shadows. In the experiments, we show that our proposed approach provides complementary cues for outpainting and achieves considerable improvement on all datasets by adopting our approach as a plug-in module. Hang Yu 0006, Shaorong Xie, Jiayan Qiu |
CVPR | 4 |
| 2024 | Visual Relationship Transformation
Jiayan Qiu, Baosheng Yu |
ECCV (65) | 2 |
| 2022 | Relationship Spatialization for Depth Estimation
Jiayan Qiu, Xinchao Wang |
ECCV (37) | 2 |
| 2021 | Scene EssenceabstractWhat scene elements, if any, are indispensable for recognizing a scene? We strive to answer this question through the lens of an exotic learning scheme. Our goal is to identify a collection of such pivotal elements, which we term as Scene Essence, to be those that would alter scene recognition if taken out from the scene. To this end, we devise a novel approach that learns to partition the scene objects into two groups, essential ones and minor ones, under the supervision that if only the essential ones are kept while the minor ones are erased in the input image, a scene recognizer would preserve its original prediction. Specifically, we introduce a learnable graph neural network (GNN) for labelling scene objects, based on which the minor ones are wiped off by an off-the-shelf image inpainter. The features of the inpainted image derived in this way, together with those learned from the GNN with the minor-object nodes pruned, are expected to fool the scene discriminator. Both subjective and objective evaluations on Places365, SUN397, and MIT67 datasets demonstrate that, the learned Scene Essence yields a visually plausible image that convincingly retains the original scene category. Jiayan Qiu, Yiding Yang, Xinchao Wang, Dacheng Tao |
CVPR | 1 |
| 2021 | Part Uncertainty Estimation Convolutional Neural Network For Person Re-IdentificationabstractDue to the large amount of noisy data in person re-identification (ReID) task, the ReID models are usually affected by the data uncertainty. Therefore, the deep uncertainty estimation method is important for improving the model robustness and matching accuracy. To this end, we propose a part-based uncertainty convolutional neural network (PUCNN), which introduces the part-based uncertainty estimation into the baseline model. On the one hand, PUCNN improves the model robustness to noisy data by distributilizing the feature embedding and constraining the part-based uncertainty. On the other hand, PUCNN improves the cumulative matching characteristics (CMC) performance of the model by filtering out low-quality training samples according to the estimated uncertainty score. The experiments on both non-video datasets, the noised Market-1501 and DukeMTMC, and video datasets, PRID2011, iLiDS-VID and MARS, demonstrate that our proposed method achieves encouraging and promising performance. Wenyu Sun, Jiyang Xie 0001, Jiayan Qiu, Zhanyu Ma |
ICIP | 3 |
| 2021 | Matching Seqlets: An Unsupervised Approach for Locality Preserving Sequence MatchingabstractIn this paper, we propose a novel unsupervised approach for sequence matching by explicitly accounting for the locality properties in the sequences. In contrast to conventional approaches that rely on frame-to-frame matching, we conduct matching using sequencelet or seqlet, a sub-sequence wherein the frames share strong similarities and are thus grouped together. The optimal seqlets and matching between them are learned jointly, without any supervision from users. The learned seqlets preserve the locality information at the scale of interest and resolve the ambiguities during matching, which are omitted by frame-based matching methods. We show that our proposed approach outperforms the state-of-the-art ones on datasets of different domains including human actions, facial expressions, speech, and character strokes. Jiayan Qiu, Xinchao Wang, Pascal Fua, Dacheng Tao |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2020 | Distilling Knowledge From Graph Convolutional NetworksabstractExisting knowledge distillation methods focus on convolutional neural networks (CNNs), where the input samples like images lie in a grid domain, and have largely overlooked graph convolutional networks (GCN) that handle non-grid data. In this paper, we propose to our best knowledge the first dedicated approach to distilling knowledge from a pre-trained GCN model. To enable the knowledge transfer from the teacher GCN to the student, we propose a local structure preserving module that explicitly accounts for the topological semantics of the teacher. In this module, the local structure information from both the teacher and the student are extracted as distributions, and hence minimizing the distance between these distributions enables topology-aware knowledge transfer from the teacher, yielding a compact yet high-performance student model. Moreover, the proposed approach is readily extendable to dynamic graph models, where the input graphs for the teacher and the student may differ. We evaluate the proposed method on two different datasets using GCN models of different architectures, and demonstrate that our method achieves the state-of-the-art knowledge distillation performance for GCN models. Yiding Yang, Jiayan Qiu, Mingli Song, Dacheng Tao, Xinchao Wang |
CVPR | 2 |
| 2020 | Hallucinating Visual Instances in Total Absentia
Jiayan Qiu, Yiding Yang, Xinchao Wang, Dacheng Tao |
ECCV (5) | 1 |
| 2020 | Learning Propagation Rules for Attribution Map Generation
Yiding Yang, Jiayan Qiu, Mingli Song, Dacheng Tao, Xinchao Wang |
ECCV (20) | 2 |
| 2019 | World From BlurabstractWhat can we tell from a single motion-blurred image? We show in this paper that a 3D scene can be revealed. Unlike prior methods that focus on producing a deblurred image, we propose to estimate and take advantage of the hidden message of a blurred image, the relative motion trajectory, to restore the 3D scene collapsed during the exposure process. To this end, we train a deep network that jointly predicts the motion trajectory, the deblurred image, and the depth one, all of which in turn form a collaborative and self-supervised cycle that supervise one another to reproduce the input blurred image, enabling plausible 3D scene reconstruction from a single blurred image. We test the proposed model on several large-scale datasets we constructed based on benchmarks, as well as real-world blurred images, and show that it yields very encouraging quantitative and qualitative results. Jiayan Qiu, Xinchao Wang, Stephen J. Maybank, Dacheng Tao |
CVPR | 1 |
| 2018 | Towards Evolutionary CompressionabstractCompressing convolutional neural networks (CNNs) is essential for transferring the success of CNNs to a wide variety of applications to mobile devices. In contrast to directly recognizing subtle weights or filters as redundant in a given CNN, this paper presents an evolutionary method to automatically eliminate redundant convolution filters. We represent each compressed network as a binary individual of specific fitness. Then, the population is upgraded at each evolutionary iteration using genetic operations. As a result, an extremely compact CNN is generated using the fittest individual, which has the original network structure and can be directly deployed in any off-the-shelf deep learning libraries. In this approach, either large or small convolution filters can be redundant, and filters in the compressed network are more distinct. In addition, since the number of filters in each convolutional layer is reduced, the number of filter channels and the size of feature maps are also decreased, naturally improving both the compression and speed-up ratios. Experiments on benchmark deep CNN models suggest the superiority of the proposed algorithm over the state-of-the-art compression methods, e.g. combined with the parameter refining approach, we can reduce the storage requirement and the floating-point multiplications of ResNet-50 by a factor of 14.64x and 5.19x, respectively, without affecting its accuracy. Yunhe Wang 0001, Chang Xu 0002, Jiayan Qiu, Chao Xu 0006, Dacheng Tao |
KDD | 3 |