Volume 2,Issue 9
跨平台图像喂入对 AIGC 文生图生成稳定性的影响
—— 基于统一提示条件的实证研究
随着文生图工具进入创作流程,重复生成的不稳定性影响人物一致性与产出效率。本文在统一结构化Prompt(P2)条件下,以四个平台开展真实生成实验(K=15),对比Feed=0与Feed=5两种喂图条件下的人物主体一致性变化。结果显示,多数平台在喂入参考图后稳定性提升,但提升幅度存在平台差异。研究为跨平台创作中图像喂入策略的使用提供可审查的经验依据。
[1]Rombach R, Blattmann A, Lorenz D, et al. High-resolution image synthesis with latent diffusion models[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. New Orleans: IEEE, 2022: 10684-10695.
[2]Ho J, Jain A, Abbeel P. Denoising diffusion probabilistic models[C]//Advances in Neural Information Processing Systems. Vancouver: NeurIPS, 2020: 6840-6851.
[3]Saharia C, Chan W, Saxena S, et al. Photorealistic text-to-image diffusion models with deep language understanding[C]//Advances in Neural Information Processing Systems. New Orleans: NeurIPS, 2022: 36479-36494.
[4]Zhang L, Rao A, Agrawala M. Adding conditional control to text-to-image diffusion models[C]//Proceedings of the IEEE/CVF International Conference on Computer Vision.Paris: IEEE, 2023.(Open Access 版)获取路径:https://openaccess.thecvf.com/content/ICCV2023/papers/Zhang_Adding_Conditional_Control_to_Text-to-Image_Diffusion_Models_ICCV_2023_paper.pdf
[5]Ruiz N, Li Y, Ouyang P, et al. DreamBooth: Fine-tuning text-to-image diffusion models for subject-driven generation[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Vancouver: IEEE, 2023: 22500-22510.
[6]Gal R, Alaluf Y, Atzmon Y, et al. An image is worth one word: Personalizing text-to-image generation using textual inversion[C]//Proceedings of the International Conference on Learning Representations. Vienna: ICLR, 2023.
[7]Liu B, Zhang Y, Hu H. Consistency and controllability in text-to-image generation: A survey[J]. ACM Computing Surveys, 2023, 56(6): 1-36.
[8]Liu F, Ren Y, Huang J. Evaluating visual consistency in text-to-image generation[C]//Proceedings of the ACM International Conference on Multimedia. Lisbon: ACM,2022: 4152-4161.
[9]Hessel J, Holtzman A, Forbes M, et al. CLIPScore: A reference-free evaluation metric for image captioning[C]//Proceedings of the Conference on Empirical Methods in
Natural Language Processing. Punta Cana: ACL, 2021: 7514-7528.
[10]Radford A, Kim J W, Hallacy C, et al. Learning transferable visual models from natural language supervision[C]//Proceedings of the International Conference on Machine Learning. Virtual: PMLR, 2021: 8748-8763.
[11]Oppenlaender J. A taxonomy of prompt modifiers for text-to-image generation[EB/OL]. (2022-04-20) [2025-12-22]. 获取路径:https://arxiv.org/abs/2204.13988
[12]Reynolds L, McDonell K. Prompt programming for large language models: Beyond the few-shot paradigm[C]//Proceedings of the CHI Conference on Human Factors in Computing Systems. 2021. DOI:10.1145/3411763.3451760.(条目信息来源:dblp)
[13]Manovich L. AI aesthetics[M]. Moscow: Strelka Press, 2018.
[14]Elkins J, Chun A. Can AI make art?[J]. Arts, 2020, 9(4): 1-14.
[15]Boden M A. Creativity and artificial intelligence[J]. Artificial Intelligence, 1998, 103(1-2): 347-356.
[16]La tour B. Reassembling the social: An introduction to actor-network-theory[M]. Oxford: Oxford University Press, 2005.