Multi-Task LLM Development
Budget: €8 – €30 EUR
I’m creating a language-model-driven system that can handle text generation, text summarisation and sentiment analysis in one cohesive workflow. The idea is to either fine-tune an existing open-source LLM or build the training pipeline from scratch—whichever approach gives us the fastest, most reliable path to production.
What I already have:
• A clear data schema and a sizeable corpus of domain-specific text.
• Initial benchmarks showing where current off-the-shelf models fall short for my use case.
What I need from you:
• Guidance on model selection or custom architecture if we decide to train from scratch.
• End-to-end fine-tuning or training, including data cleaning, tokenisation and evaluation with standard NLP metrics.
• An API (Python/FastAPI is fine) exposing three endpoints: /generate, /summarise and /sentiment.
• Deployment scripts (Docker or similar) so the service can be replicated easily in staging and production.
Acceptance criteria
• All three tasks must meet or exceed agreed benchmark scores on a held-out test set.
• Latency under 500 ms for an average prompt on an A100 GPU or equivalent.
• Clean, well-commented code pushed to a private Git repository, along with a short READ-ME explaining setup and usage.
If you’ve previously worked with models like GPT-J, Llama 2, Falcon, or similar, and you’re comfortable plumbing them into a production API, your expertise will be invaluable here.
What I already have:
• A clear data schema and a sizeable corpus of domain-specific text.
• Initial benchmarks showing where current off-the-shelf models fall short for my use case.
What I need from you:
• Guidance on model selection or custom architecture if we decide to train from scratch.
• End-to-end fine-tuning or training, including data cleaning, tokenisation and evaluation with standard NLP metrics.
• An API (Python/FastAPI is fine) exposing three endpoints: /generate, /summarise and /sentiment.
• Deployment scripts (Docker or similar) so the service can be replicated easily in staging and production.
Acceptance criteria
• All three tasks must meet or exceed agreed benchmark scores on a held-out test set.
• Latency under 500 ms for an average prompt on an A100 GPU or equivalent.
• Clean, well-commented code pushed to a private Git repository, along with a short READ-ME explaining setup and usage.
If you’ve previously worked with models like GPT-J, Llama 2, Falcon, or similar, and you’re comfortable plumbing them into a production API, your expertise will be invaluable here.