Llama3 Model Optimization Consultancy on GCP
Budget: $250 – $750 USD
We are seeking an expert in Ollama, NextJS, and Llama3 8B to resolve a technical issue on an already configured server. We have a machine on Google Cloud Platform (GCP) with the following specifications:
CPU: 8 cores
RAM: 32 GB
GPU: NVIDIA T4 with 15 GB of VRAM
Project Details:
Current Implementation:
The Ollama server is configured with NextJS for Ollama and the Llama3 8B model.
We have a subdomain configured for internet access.
Current Issue:
The model is experiencing an issue where responses are being truncated, preventing proper operation.
Requirement:
We are looking for a consultant with proven experience in integrating and optimizing Ollama with NextJS and Llama3 language models. The objective is to resolve the response truncation issue to ensure the server operates optimally.
CPU: 8 cores
RAM: 32 GB
GPU: NVIDIA T4 with 15 GB of VRAM
Project Details:
Current Implementation:
The Ollama server is configured with NextJS for Ollama and the Llama3 8B model.
We have a subdomain configured for internet access.
Current Issue:
The model is experiencing an issue where responses are being truncated, preventing proper operation.
Requirement:
We are looking for a consultant with proven experience in integrating and optimizing Ollama with NextJS and Llama3 language models. The objective is to resolve the response truncation issue to ensure the server operates optimally.