This is a version adjusted based on the previous model. After testing multiple language models, the final choice was
JoyCaption2Split language model. Before the image enters the latent space, I first performed scaling, which is a very critical step, because when some images have insufficient pixels, a lot of details will be lost. However, after reducing the size and then enlarging it
can complete the pixels. These steps are very important) It can significantly improve the image quality without changing the original image size. The facial changes in portraits are still
relatively difficult to control, and the reversion amplitude can be reduced to around 0.2.