Your agent calls one API to train a character LoRA, generate images, and cut them into video — self-hosted on your own GPU, no cloud bill, no ComfyUI graph to wire by hand.

Every one of these came out of a self-hosted FlixML install on a single consumer GPU. The stills are SDXL, the clip is Wan 2.2. FlixML also drives lip-sync (InfiniteTalk), instruction edits (Qwen) and FLUX.2 — full breakdown, VRAM included, in the model directory.

SDXL
Cinematic scene from a text prompt.

SDXL
Landscape, no reference image.

SDXL
Architectural interior, hard daylight.

SDXL
Macro detail, backlit.

SDXL
Creature work, neon-lit night scene.
Wan 2.2
That still, animated — image in, video out.
FlixML is open source and free to run. We build and test every workflow on it ourselves, every day — that's where the guides and model breakdowns on this site come from. Clone it, point it at your own GPU, and it's yours.
[ "sdxl_base", "sdxl_lora", "wan22_i2v", "wan22_fun_camera", "infinitetalk_i2v", "infinitetalk_v2v", "flux2_base", "qwen_multiangle", ... ]
How to run the workflows behind image and video generation, model by model — plus what's moving in the wider model landscape.