Python Developer for AI-Based App APIs
Budget: $15 – $25 USD
I'm working on an AI-powered app using Dify, Ollama with multiple LLMs — everything is set up and running smoothly. Now I need help from an experienced Python developer to build local APIs to complete the pipeline.
This is just a general idea — as the developer, you’re expected to handle dependencies and make technical decisions to get things working efficiently.
What I Need:
1. File to Image API
Input: PDF, DOCX, PPTX, XLSX
Process: Convert each page of the document into separate images
Requirement: Save the page number as part of the filename or metadata for tracking
Output: A set of images
2. Image to Text API using LLM Vision
Input: Images (from the previous step)
Process: Use an LLM vision model (Ollama-based) to extract text from images
Page number should be extracted from the image filename and retained
Output: Consolidated extracted text (ideally with page numbers as references)
Requirements:
Strong Python skills (FastAPI or Flask preferred)
Experience in AI projects (especially with vision models, LLMs)
Ability to manage dependencies (e.g., image processing libraries, document converters)
Familiarity with local setups (no cloud dependencies)
Experience with Dify and Ollama is a big plus
This is just a general idea — as the developer, you’re expected to handle dependencies and make technical decisions to get things working efficiently.
What I Need:
1. File to Image API
Input: PDF, DOCX, PPTX, XLSX
Process: Convert each page of the document into separate images
Requirement: Save the page number as part of the filename or metadata for tracking
Output: A set of images
2. Image to Text API using LLM Vision
Input: Images (from the previous step)
Process: Use an LLM vision model (Ollama-based) to extract text from images
Page number should be extracted from the image filename and retained
Output: Consolidated extracted text (ideally with page numbers as references)
Requirements:
Strong Python skills (FastAPI or Flask preferred)
Experience in AI projects (especially with vision models, LLMs)
Ability to manage dependencies (e.g., image processing libraries, document converters)
Familiarity with local setups (no cloud dependencies)
Experience with Dify and Ollama is a big plus