I2V workflow to run on RunPod H100 – generating 6–10 second near-photorealistic videos from a photo with action guided by a text prompt.

Job ID: 40085557

Budget: $750 – $1,500 USD

- I2V workflow to run on RunPod H100 – generating 6–10 second near-photorealistic videos from a photo with action guided by a text prompt.
- Reasonable render time for 10-second video outputs at 480p resolution.

Requirements:
- Input: 1 or 2 images (e.g., photos and/or a short reference video) + a text prompt to guide the action/motion.
- Output: Short video (6–10 seconds) delivered as a direct, clean MP4 file.
- Dialogue:
- Converts dialogue written in tagged form (e.g., Bill: "Hello!") into spoken audio, with lines assigned to the tagged person and matching mouth movements (functional and convincing lip-sync is acceptable). Other methods are fine if they produce the same result.
- Enables tone (via tags is fine) such as whisper, shout, softly, or harshly.
- Accent selection not required.
- Only one active speaker per continuous shot; cuts are allowed.

- Near-Photorealistic Quality: Motion should appear near-photorealistic at normal viewing distance; still-frame perfection is not required.
- Multi-Person Capability: Supports scenes with at least three people staying recognizable (not identical pixel-perfect faces) and correctly placed, with good face and character consistency across all frames.
- Uncensored Workflow: No content filtering or moderation layers imposed by the workflow or models.
- All download links provided.

Deliverables:
- Workflow JSON file optimized for a RunPod H100.
- A working, tested RunPod H100 setup that runs the workflow as described (RunPod will be supplied).
- Short operating guide covering settings and use.

Budget & Timing:
- $1,000 USD.
- Soon – ready to begin.