ComfyUI Influencer Pipeline SetupComfyUI Workflow Setup on Vast.ai — AI Influencer Pipeline (Image, Video, Talking Head)
Budget: $30 – $250 AUD
PROJECT OVERVIEW
I need an experienced ComfyUI specialist to set up a complete AI influencer content creation pipeline on Vast.ai (RTX 4090 / 24GB VRAM). The goal is a fully working environment where I can generate consistent AI character images (SFW and NSFW), short video clips, dance/motion transfer videos, and long-form talking head / UGC-style videos — all from ComfyUI. I will be training multiple character LoRAs, so all workflows must allow me to easily select which character to use from a dropdown before generating. Switching between characters should be as simple as picking a different LoRA file — no reconfiguration needed.
I will handle LoRA training myself. I need you to build and configure the workflows, install all required models and custom nodes, and deliver everything documented so I can replicate the setup on a fresh Vast.ai instance.
SCOPE OF WORK
1) Vast.ai Environment Setup
Configure the official ComfyUI template on Vast.ai (RTX 4090). Install ComfyUI Manager and all required custom nodes. Create a setup/provisioning script (bash) that I can paste into Jupyter Terminal on a fresh instance to automatically install everything — models, nodes, dependencies. Organize model files into correct directories.
2) Image Generation Workflow
Flux Dev text-to-image workflow with LoRA loader so I can plug in my own trained LoRA. ControlNet / IP-Adapter integration for pose and composition control. ADetailer or face enhancement nodes for realistic skin and facial detail. Upscaling workflow (Real-ESRGAN or similar). Must support both SFW and NSFW generation without filters.
3) LoRA Training Workflow
Working Flux LoRA training workflow using ComfyUI Flux Trainer. Pre-configured with recommended settings for character training (network_dim, network_alpha, steps, optimizer). I will train my own LoRA — just need the workflow ready to go with clear instructions on where to place my dataset and how to adjust parameters. The training workflow should be reusable for multiple characters — I need to be able to swap in a new dataset folder and trigger word and train a new LoRA without reconfiguring anything else. All trained LoRA files should be stored in an organized folder structure so I can easily select which character to use from a dropdown in the image generation, video, and InfiniteTalk workflows. The provisioning script should also re-download all my trained LoRA files when setting up a fresh instance (or include instructions for uploading them).
4) Short Video Generation Workflow
Wan 2.2 (5B) image-to-video workflow. Image input to animated short clip (3-10 seconds). Character consistency between source image and video output. Frame interpolation for smooth output. Must include LoRA loader so I can select which character to generate with.
5) InfiniteTalk — Long-Form Talking Head Video Workflow
InfiniteTalk integration via Wan Video Wrapper (Kijai). Image plus audio to unlimited-length talking video with lip sync. Text-to-speech node integration (Chatterbox or similar) so I can type a script and generate audio plus video in one pipeline. Chunk-based generation with proper overlap settings for seamless long-form output. Frame interpolation for final output smoothness. Both single-speaker and multi-speaker setups.
6) Dance / Motion Transfer Workflow
I need to be able to take a TikTok dance video (or any reference video with body movement) and transfer that motion onto my AI influencer character. The workflow should take two inputs: a reference image of my character (generated from my LoRA) and a driving dance video. It should output a video of my character performing the same dance. Preferred approach is SteadyDancer (Wan 2.1 based, Kijai port) for best quality. If SteadyDancer has limitations, an AnimateDiff plus ControlNet (OpenPose/DWPose) pipeline with face swap (ReActor) is acceptable as a fallback. Frame interpolation and upscaling for final output. Should preserve character identity and face consistency throughout the video.
7) Documentation and Deliverables
All workflow JSON files (drag-and-drop ready). A single bash setup script that installs everything on a clean Vast.ai ComfyUI instance — models, custom nodes, dependencies. A list of all models used with Hugging Face / Civitai download links. Written guide or video walkthrough explaining each workflow and how to use it. The setup should be modular — I may want to add more workflows in the future, so the environment should be clean and well-organized.
REQUIREMENTS
Proven experience with ComfyUI workflows (show examples or portfolio). Experience with Flux, Wan 2.2, InfiniteTalk / MultiTalk, SteadyDancer or AnimateDiff dance transfer. Comfortable working with NSFW content (unrestricted generation). Experience deploying on Vast.ai or RunPod (cloud GPU). Must provide a working provisioning script — not just instructions. All workflows must support multiple trained LoRA characters with easy switching — no reconfiguration between characters.
NICE TO HAVE (Optional / Future Work)
Additional video models (HunyuanVideo, CogVideoX, etc.). Face swap workflow. Batch generation and automation workflows. API integration for automated content pipelines. Workflow for consistent character across different scenes using IP-Adapter. Any other useful workflows you recommend — I am open to suggestions.
I need an experienced ComfyUI specialist to set up a complete AI influencer content creation pipeline on Vast.ai (RTX 4090 / 24GB VRAM). The goal is a fully working environment where I can generate consistent AI character images (SFW and NSFW), short video clips, dance/motion transfer videos, and long-form talking head / UGC-style videos — all from ComfyUI. I will be training multiple character LoRAs, so all workflows must allow me to easily select which character to use from a dropdown before generating. Switching between characters should be as simple as picking a different LoRA file — no reconfiguration needed.
I will handle LoRA training myself. I need you to build and configure the workflows, install all required models and custom nodes, and deliver everything documented so I can replicate the setup on a fresh Vast.ai instance.
SCOPE OF WORK
1) Vast.ai Environment Setup
Configure the official ComfyUI template on Vast.ai (RTX 4090). Install ComfyUI Manager and all required custom nodes. Create a setup/provisioning script (bash) that I can paste into Jupyter Terminal on a fresh instance to automatically install everything — models, nodes, dependencies. Organize model files into correct directories.
2) Image Generation Workflow
Flux Dev text-to-image workflow with LoRA loader so I can plug in my own trained LoRA. ControlNet / IP-Adapter integration for pose and composition control. ADetailer or face enhancement nodes for realistic skin and facial detail. Upscaling workflow (Real-ESRGAN or similar). Must support both SFW and NSFW generation without filters.
3) LoRA Training Workflow
Working Flux LoRA training workflow using ComfyUI Flux Trainer. Pre-configured with recommended settings for character training (network_dim, network_alpha, steps, optimizer). I will train my own LoRA — just need the workflow ready to go with clear instructions on where to place my dataset and how to adjust parameters. The training workflow should be reusable for multiple characters — I need to be able to swap in a new dataset folder and trigger word and train a new LoRA without reconfiguring anything else. All trained LoRA files should be stored in an organized folder structure so I can easily select which character to use from a dropdown in the image generation, video, and InfiniteTalk workflows. The provisioning script should also re-download all my trained LoRA files when setting up a fresh instance (or include instructions for uploading them).
4) Short Video Generation Workflow
Wan 2.2 (5B) image-to-video workflow. Image input to animated short clip (3-10 seconds). Character consistency between source image and video output. Frame interpolation for smooth output. Must include LoRA loader so I can select which character to generate with.
5) InfiniteTalk — Long-Form Talking Head Video Workflow
InfiniteTalk integration via Wan Video Wrapper (Kijai). Image plus audio to unlimited-length talking video with lip sync. Text-to-speech node integration (Chatterbox or similar) so I can type a script and generate audio plus video in one pipeline. Chunk-based generation with proper overlap settings for seamless long-form output. Frame interpolation for final output smoothness. Both single-speaker and multi-speaker setups.
6) Dance / Motion Transfer Workflow
I need to be able to take a TikTok dance video (or any reference video with body movement) and transfer that motion onto my AI influencer character. The workflow should take two inputs: a reference image of my character (generated from my LoRA) and a driving dance video. It should output a video of my character performing the same dance. Preferred approach is SteadyDancer (Wan 2.1 based, Kijai port) for best quality. If SteadyDancer has limitations, an AnimateDiff plus ControlNet (OpenPose/DWPose) pipeline with face swap (ReActor) is acceptable as a fallback. Frame interpolation and upscaling for final output. Should preserve character identity and face consistency throughout the video.
7) Documentation and Deliverables
All workflow JSON files (drag-and-drop ready). A single bash setup script that installs everything on a clean Vast.ai ComfyUI instance — models, custom nodes, dependencies. A list of all models used with Hugging Face / Civitai download links. Written guide or video walkthrough explaining each workflow and how to use it. The setup should be modular — I may want to add more workflows in the future, so the environment should be clean and well-organized.
REQUIREMENTS
Proven experience with ComfyUI workflows (show examples or portfolio). Experience with Flux, Wan 2.2, InfiniteTalk / MultiTalk, SteadyDancer or AnimateDiff dance transfer. Comfortable working with NSFW content (unrestricted generation). Experience deploying on Vast.ai or RunPod (cloud GPU). Must provide a working provisioning script — not just instructions. All workflows must support multiple trained LoRA characters with easy switching — no reconfiguration between characters.
NICE TO HAVE (Optional / Future Work)
Additional video models (HunyuanVideo, CogVideoX, etc.). Face swap workflow. Batch generation and automation workflows. API integration for automated content pipelines. Workflow for consistent character across different scenes using IP-Adapter. Any other useful workflows you recommend — I am open to suggestions.