Llama3 Model Optimization Consultancy on GCP

Job ID: 38462436

Budget: $250 – $750 USD

We are seeking an expert in Ollama, NextJS, and Llama3 8B to resolve a technical issue on an already configured server. We have a machine on Google Cloud Platform (GCP) with the following specifications:

CPU: 8 cores
RAM: 32 GB
GPU: NVIDIA T4 with 15 GB of VRAM

Project Details:

Current Implementation:

The Ollama server is configured with NextJS for Ollama and the Llama3 8B model.
We have a subdomain configured for internet access.

Current Issue:

The model is experiencing an issue where responses are being truncated, preventing proper operation.

Requirement:
We are looking for a consultant with proven experience in integrating and optimizing Ollama with NextJS and Llama3 language models. The objective is to resolve the response truncation issue to ensure the server operates optimally.
Related categories: Linux Next.js LLaMA