Python App for GPT-powered Computer Automation
Budget: $10,000 – $20,000 USD
I'm developing an AI agent in Python to automate tasks given by the user by taking screenshots, analyzing visually using GPT 4o Vision, applying bounding boxes and user prompts, and requesting next steps from GPT-4o. My current version handles basic use cases well, but does not function well every time, and needs more features and enhancements.
Key Requirements:
- The app must support both Windows and macOS.
- It should be equipped with a user-friendly Graphical User Interface (GUI).
Ideal Skills and Experience:
- Proficiency in Python programming.
- Experience in developing cross-platform & GUI based applications with Python.
- Firm grasp of GPT-4o and similar AI models.
- Capable of experimenting, finetuning, deploying open source models.
- Capable with Unix CLI & Ubuntu servers.
Examples of what needs to be done:
- A system to handle just-in-time OCR, allowing GPT to point to text it sees in the image, and create bounding boxes.
- More enhanced bounding box capabilities.
- More advanced "chain of thought" and "reflection" style reasoning.
- Experiment with other models (e.g. Fuyu 8B) to figure out if they can yield better results.
- Deciding when a task has been completed appropriately.
This may be a larger project, and will require a highly capable and creative person or team with experience in AI, Python, AI Agents, Automation, and more. It may require some experimentation to figure out the best/most optimal ways to complete the project as well.
Key Requirements:
- The app must support both Windows and macOS.
- It should be equipped with a user-friendly Graphical User Interface (GUI).
Ideal Skills and Experience:
- Proficiency in Python programming.
- Experience in developing cross-platform & GUI based applications with Python.
- Firm grasp of GPT-4o and similar AI models.
- Capable of experimenting, finetuning, deploying open source models.
- Capable with Unix CLI & Ubuntu servers.
Examples of what needs to be done:
- A system to handle just-in-time OCR, allowing GPT to point to text it sees in the image, and create bounding boxes.
- More enhanced bounding box capabilities.
- More advanced "chain of thought" and "reflection" style reasoning.
- Experiment with other models (e.g. Fuyu 8B) to figure out if they can yield better results.
- Deciding when a task has been completed appropriately.
This may be a larger project, and will require a highly capable and creative person or team with experience in AI, Python, AI Agents, Automation, and more. It may require some experimentation to figure out the best/most optimal ways to complete the project as well.