ComfyUI Workflow to creat text based Art. -- 2
Budget: $10 – $100 USD
I'm looking for an experienced ComfyUI developer to create a complete /ComfyUI workflow to create images similar to the ones attached I'm using a RTX 3060 GPU with 12 GB of ram, so it must be capable of operating within those parameters. sdxl_lightning_4step.safetensors would be my preferred model but that''s not entirely necessary if you have another that would work with my GPU.
Prompt accuracy and spelling is of paramount importance.Towards that end my advisor proposes this as a suggested workflow.
Stabilize typography by generating a clean text layout first (SVG/PNG) and then using that as a strong conditioning signal via ControlNet / IP-Adapter in ComfyUI, rather than asking the diffusion model to “invent” letters from scratch. Run a two-stage pipeline: 1) low-res layout for composition and legibility, 2) targeted high-res upscaling/refinement only on text regions using a refiner model to keep edges sharp while staying within 12 GB VRAM.
Tokenization and prompt templates would be normalized (fixed casing, spacing, and special-character handling) so the model always sees text in a consistent format.
The enclosed examples were made from an online AI platform Freepik, using several different models. Our objective is to increase the speed and accuracy, as well as to be able to batch process these so that from a list of 75 phrases we could generate approx. 300 separate images from that batch within just a few hours.
We will give a list of phrases to ComfyUI, and it will generate a batch of images based upon that list, with probably four or five variations for each phrase. All nodes, loras and instructions must be provided either by a valid download link or included with the finalized product.
All bids must be final! Any low "placeholder" bids will be rejected if you try to raise the bid later. (no exceptions). If you want to bid on this and have the experience, please begin your bid with the phrase "peanut butter". That shows me that you really have read my bid. If you do not start your bid that way I cannot accept it.
Prompt accuracy and spelling is of paramount importance.Towards that end my advisor proposes this as a suggested workflow.
Stabilize typography by generating a clean text layout first (SVG/PNG) and then using that as a strong conditioning signal via ControlNet / IP-Adapter in ComfyUI, rather than asking the diffusion model to “invent” letters from scratch. Run a two-stage pipeline: 1) low-res layout for composition and legibility, 2) targeted high-res upscaling/refinement only on text regions using a refiner model to keep edges sharp while staying within 12 GB VRAM.
Tokenization and prompt templates would be normalized (fixed casing, spacing, and special-character handling) so the model always sees text in a consistent format.
The enclosed examples were made from an online AI platform Freepik, using several different models. Our objective is to increase the speed and accuracy, as well as to be able to batch process these so that from a list of 75 phrases we could generate approx. 300 separate images from that batch within just a few hours.
We will give a list of phrases to ComfyUI, and it will generate a batch of images based upon that list, with probably four or five variations for each phrase. All nodes, loras and instructions must be provided either by a valid download link or included with the finalized product.
All bids must be final! Any low "placeholder" bids will be rejected if you try to raise the bid later. (no exceptions). If you want to bid on this and have the experience, please begin your bid with the phrase "peanut butter". That shows me that you really have read my bid. If you do not start your bid that way I cannot accept it.