Web & Mobile Generative AI Tools
Budget: $10,000 – $20,000 USD
I want to launch a cohesive suite of AI-powered creation tools that runs smoothly in any modern browser and ships as a companion iOS/Android app. The platform will revolve around four core experiences:
• Chatbot: A conversational assistant driven by a large language model (OpenAI GPT-4, Claude, or a comparable open-source model) with memory, retrieval-augmented generation, and an admin panel so I can refine prompts and monitor usage.
• Image & Video Generator: Text-to-image and text-to-video pipelines using Stable Diffusion / Stable Video Diffusion (or a stronger alternative) with prompt weighting, negative prompts, and upscaling.
• 3D Image Generator: Generation of glTF/OBJ assets ready for WebGL display, including basic texture baking.
• AI-Assisted Video Editor: Browser-based timeline that auto-detects scenes, removes silence, adds captions, and outputs H.264/HEVC.
Shared requirements
– One codebase that cleanly targets web and mobile (React + React Native, Flutter, or a similar framework).
– GPU-ready backend (Docker/Kubernetes on AWS, GCP, or Azure) that can scale each model independently.
– Secure user accounts, single sign-on, Stripe metered billing, and a token/credit system common to every tool.
– Clean componentized UI/UX; Figma mock-ups are available on request.
– Source code in a private Git repo with weekly check-ins.
Acceptance criteria
1. All four tools load and run inference end-to-end on both web and mobile.
2. Output quality meets or beats the reference models’ defaults.
3. Latency stays under five seconds for typical prompts on a T4 or comparable GPU.
4. CI/CD pipeline deploys to staging and production with zero-downtime upgrades.
5. Documentation covers environment setup, model weights, and API endpoints.
If you have proven experience deploying LLM chatbots, diffusion models, or GPU-centric cloud architecture, let’s talk; I’m ready to begin immediately and iterate quickly through milestones.
• Chatbot: A conversational assistant driven by a large language model (OpenAI GPT-4, Claude, or a comparable open-source model) with memory, retrieval-augmented generation, and an admin panel so I can refine prompts and monitor usage.
• Image & Video Generator: Text-to-image and text-to-video pipelines using Stable Diffusion / Stable Video Diffusion (or a stronger alternative) with prompt weighting, negative prompts, and upscaling.
• 3D Image Generator: Generation of glTF/OBJ assets ready for WebGL display, including basic texture baking.
• AI-Assisted Video Editor: Browser-based timeline that auto-detects scenes, removes silence, adds captions, and outputs H.264/HEVC.
Shared requirements
– One codebase that cleanly targets web and mobile (React + React Native, Flutter, or a similar framework).
– GPU-ready backend (Docker/Kubernetes on AWS, GCP, or Azure) that can scale each model independently.
– Secure user accounts, single sign-on, Stripe metered billing, and a token/credit system common to every tool.
– Clean componentized UI/UX; Figma mock-ups are available on request.
– Source code in a private Git repo with weekly check-ins.
Acceptance criteria
1. All four tools load and run inference end-to-end on both web and mobile.
2. Output quality meets or beats the reference models’ defaults.
3. Latency stays under five seconds for typical prompts on a T4 or comparable GPU.
4. CI/CD pipeline deploys to staging and production with zero-downtime upgrades.
5. Documentation covers environment setup, model weights, and API endpoints.
If you have proven experience deploying LLM chatbots, diffusion models, or GPU-centric cloud architecture, let’s talk; I’m ready to begin immediately and iterate quickly through milestones.