Photorealistic Lighting z-image-turbo-flow-dpo
Back
Photorealistic Lighting z-image-turbo-flow-dpo
Photorealistic Lighting z-image-turbo-flow-dpo
Photorealistic Lighting z-image-turbo-flow-dpo
Photorealistic Lighting z-image-turbo-flow-dpo
Photorealistic Lighting z-image-turbo-flow-dpo
Photorealistic Lighting z-image-turbo-flow-dpo

Photorealistic Lighting z-image-turbo-flow-dpo

35.4K Views113 Likes431 Favorites
PhotographyRealisticGirlBoySpace sceneLandscapeStylizationCharacterization

Dream2046

Dream2046

PhotographyRealisticGirlBoySpace sceneLandscapeStylizationCharacterization

Model Information

Active
Original author:
F16
Model Type:
LoRA
Basic Model:
Z-image-turbo
Resource Name:
models/loras/ZIT-flow-dpo-lora.safetensors
MD5:
41c63a50cf73fd94f632e1c9128e713d

Z-Image-Turbo Photorealistic Lighting & Shadows LoRA (Flow-DPO)

This is a LoRA adapter specifically designed for Alibaba-Tongyi/Z-Image-Turbo, fine-tuned using Flow-DPO (Direct Preference Optimization for Flow Matching) to significantly enhance photorealistic lighting, cinematic shadows, and overall image quality.

By applying Flow-DPO on spatially strictly aligned image pairs, this LoRA effectively resolves artifacts common in ultra-fast distillation models such as "flatness," "overexposure," or "plastic look," generating stunning and physically accurate lighting effects in just 8 inference steps.

Training Details & Methodology

The model was trained using a custom implementation of Flow-DPO (Improving Video Generation with Human Feedback, arXiv:2501.13918).

1. Dataset (Strictly Spatially Aligned)

To prevent the model from hallucinating or altering image structure (catastrophic forgetting), the preference dataset was constructed using strict spatial alignment:

Chosen: High-quality professional photography with perfect lighting, shadows, and textures.

Rejected: Programmatic degradation applied to the exact same images (Gaussian blur, reduced contrast, extreme exposure shifts, Gaussian noise, and severe JPEG compression artifacts).

Alignment: No cropping or warping operations were performed, ensuring that the flow matching trajectory only learns to correct lighting and textures.

2. Discrete Timestep Distillation Preservation

Unlike continuous sampling timesteps $t \in [0, 1]$ in standard diffusion models, Z-Image-Turbo is a distillation model specifically optimized for 8 fixed timesteps.

During Flow-DPO training, we dynamically extracted the precise discrete $t$ distribution from FlowMatchEulerDiscreteScheduler and strictly restricted random sampling to these 8 nodes. This ensures that LoRA maintains the extreme speed of the Turbo model without causing output blurriness.

3. Hyperparameters

Base Model: Alibaba-Tongyi/Z-Image-Turbo (6B Single-Stream DiT)

Learning Rate: 1e-4

KL Penalty ($\beta$): 1.0

Effective Batch Size: 1

Limitations

Not an Image-to-Image Restorer: This LoRA modifies the prior distribution of text-to-image generation. It is designed to generate better raw images from text prompts, rather than acting as an img2img filter to repair user-uploaded low-quality photos (unless combined with RF-Inversion technology, which is highly unstable in 8-step models).

Color Saturation

If the LoRA weight is too high (e.g., > 1.5), the DPO boundary maximization property may cause the image to become overly sharpened or saturated. For the best photorealistic results, keep the weight in the range of 0.6 - 1.0.

This model is sourced from an external transfer (transfer address: https://modelscope.cn/models/F16/z-image-turbo-flow-dpo ),if the original author has objections to this transfer, you can click,
Appeal
We will, within 24 hours, edit, delete, or transfer the model to the original author according to the original author's request

Z-Image-Turbo Photorealistic Lighting & Shadows LoRA (Flow-DPO)

This is a LoRA adapter specifically designed for Alibaba-Tongyi/Z-Image-Turbo, fine-tuned using Flow-DPO (Direct Preference Optimization for Flow Matching) to significantly enhance photorealistic lighting, cinematic shadows, and overall image quality.

By applying Flow-DPO on spatially strictly aligned image pairs, this LoRA effectively resolves artifacts common in ultra-fast distillation models such as "flatness," "overexposure," or "plastic look," generating stunning and physically accurate lighting effects in just 8 inference steps.

Training Details & Methodology

The model was trained using a custom implementation of Flow-DPO (Improving Video Generation with Human Feedback, arXiv:2501.13918).

1. Dataset (Strictly Spatially Aligned)

To prevent the model from hallucinating or altering image structure (catastrophic forgetting), the preference dataset was constructed using strict spatial alignment:

Chosen: High-quality professional photography with perfect lighting, shadows, and textures.

Rejected: Programmatic degradation applied to the exact same images (Gaussian blur, reduced contrast, extreme exposure shifts, Gaussian noise, and severe JPEG compression artifacts).

Alignment: No cropping or warping operations were performed, ensuring that the flow matching trajectory only learns to correct lighting and textures.

2. Discrete Timestep Distillation Preservation

Unlike continuous sampling timesteps $t \in [0, 1]$ in standard diffusion models, Z-Image-Turbo is a distillation model specifically optimized for 8 fixed timesteps.

During Flow-DPO training, we dynamically extracted the precise discrete $t$ distribution from FlowMatchEulerDiscreteScheduler and strictly restricted random sampling to these 8 nodes. This ensures that LoRA maintains the extreme speed of the Turbo model without causing output blurriness.

3. Hyperparameters

Base Model: Alibaba-Tongyi/Z-Image-Turbo (6B Single-Stream DiT)

Learning Rate: 1e-4

KL Penalty ($\beta$): 1.0

Effective Batch Size: 1

Limitations

Not an Image-to-Image Restorer: This LoRA modifies the prior distribution of text-to-image generation. It is designed to generate better raw images from text prompts, rather than acting as an img2img filter to repair user-uploaded low-quality photos (unless combined with RF-Inversion technology, which is highly unstable in 8-step models).

Color Saturation

If the LoRA weight is too high (e.g., > 1.5), the DPO boundary maximization property may cause the image to become overly sharpened or saturated. For the best photorealistic results, keep the weight in the range of 0.6 - 1.0.

This model is sourced from an external transfer (transfer address: https://modelscope.cn/models/F16/z-image-turbo-flow-dpo ),if the original author has objections to this transfer, you can click,
Appeal
We will, within 24 hours, edit, delete, or transfer the model to the original author according to the original author's request