This workflow is a professional martial arts action video generation workflow deeply customized based on the MiniMax-H3(FL2VA) base model, paired with the newly iterated Wushu Action LoRA V7 model. It adopts a dual-engine capture and training scheme, independently trained using both the musubi-tuner and AI-Toolkit frameworks to form dual-version model assets. This achieves high-quality, high-fidelity AI generation of Chinese martial arts action videos, serving as a plug-and-play cinematic martial arts animation workflow adapted for ComfyUI.
Addressing pain points in traditional AI martial arts videos such as stiff movements, chaotic forms, blurry visuals, and incoherent actions, this workflow has undergone comprehensive optimization and upgrading. The training dataset features a curated selection of 924 professional martial arts clips, discarding fuzzy generalized tagging in favor of refined labeling using professional martial arts terminology. This resolves core issues from older versions such as non-professional labeling, fragmented clip cutting, and motion distortion. Furthermore, it adapts to the partitioned base model characteristics of FL2VA, avoiding the adaptation flaws of PRUNED quantized base models and significantly enhancing the standardization, coherence, and realism of martial arts actions.
The workflow architecture is extremely minimalist and efficient, adapting to the full MiniMax H3 model ecosystem. It includes a standardized deployment scheme featuring an exclusive base model, text encoder, and audio-video dual VAE, allowing users to quickly set it up simply by placing model files into their respective categories. It comes with a built-in exclusive standardized sampling parameter system, locking in the Euler sampler + Simple scheduler, 1.0 exclusive CFG distillation parameters, and dedicated resolution and frame count rules adapted to the H3 model. This completely avoids image collapse and motion distortion caused by incorrect parameters, allowing beginners to directly copy and use the parameters with zero debugging barrier.
At the content generation level, the workflow is equipped with a three-stage storyboard prompt system and a complete martial arts terminology lexicon. It covers all categories of martial arts actions including kicking techniques, fists and palms, weaponry, footwork, body movements, and offense-defense transitions, supporting solo routines, two-person combat spars, and various weapon-based martial arts scenes involving swords, sabers, spears, and staves. It presets multiple mature general prompt templates while supporting the free combination of moves, camera angles, scene atmospheres, and combat sound effects, accurately distinguishing between dynamic martial arts actions and static visuals to eliminate issues like sluggish characters and motion piling.
To meet the demands of different usage scenarios, LoRA offers multiple optional versions adapted to different quantization specs of the FL2VA base model. It allows users to compare the effects of the dual training frameworks, flexibly adapting to scenarios such as personal creation, short-form video editing, and cinematic animation previz production. The workflow is compatible with conventional image quality restoration plugins, allowing low-parameter blur issues to be optimized by increasing sampling steps, with a subsequent iterative launch of the Ref2VA version to further expand multi-reference generation capabilities. Balancing professionalism, stability, and practicality, this overall workflow is currently the highest-fidelity and most adaptable dedicated martial arts action generation solution under the MiniMax H3 framework.


No creations yet

No creations available.