LLM Observability Platform Deployment
Budget: $5,000 – $10,000 USD
I want to stand up a full observability stack for our Large Language Model workloads and run it entirely on Docker, Kubernetes, and Databricks. The system’s core purpose is real-time monitoring and alerting, so every design decision should keep fast, actionable insights front and centre.
Scope of work
• Design a high-level architecture that shows how metrics, logs, and traces flow from the LLM services into a durable data-storage layer, then surface as alerts.
• Implement that design: build the Docker images, write the Kubernetes manifests/Helm charts, and configure Databricks so raw and aggregated telemetry lands in the chosen storage backend.
• Set up alert rules and sample dashboards (optional for this phase; the immediate must-have is the data-storage component).
Key requirements
• Must run on vanilla Kubernetes (v1.27+) with Docker-compatible images.
• Data storage is mandatory; I am open to your recommendation on SQL, NoSQL, or object storage as long as it scales and plugs cleanly into Databricks for downstream analysis.
• Monitoring and alerting need to work out of the box once the cluster is up.
• All code and configuration should be delivered via a Git repo with clear README instructions.
Acceptance criteria
1. “kubectl apply -f ./manifests” spins up the stack without manual tweaks.
2. Synthetic LLM traffic appears in the datastore and triggers at least one sample alert.
3. Documentation covers deployment steps, schema, and how to add new metrics.
If you have prior experience wiring observability pipelines on Databricks or tuning Prometheus, Grafana, Loki, or similar tools inside Kubernetes, that will accelerate the hand-off, but I’m open to alternative stacks as long as the above criteria are met.
Scope of work
• Design a high-level architecture that shows how metrics, logs, and traces flow from the LLM services into a durable data-storage layer, then surface as alerts.
• Implement that design: build the Docker images, write the Kubernetes manifests/Helm charts, and configure Databricks so raw and aggregated telemetry lands in the chosen storage backend.
• Set up alert rules and sample dashboards (optional for this phase; the immediate must-have is the data-storage component).
Key requirements
• Must run on vanilla Kubernetes (v1.27+) with Docker-compatible images.
• Data storage is mandatory; I am open to your recommendation on SQL, NoSQL, or object storage as long as it scales and plugs cleanly into Databricks for downstream analysis.
• Monitoring and alerting need to work out of the box once the cluster is up.
• All code and configuration should be delivered via a Git repo with clear README instructions.
Acceptance criteria
1. “kubectl apply -f ./manifests” spins up the stack without manual tweaks.
2. Synthetic LLM traffic appears in the datastore and triggers at least one sample alert.
3. Documentation covers deployment steps, schema, and how to add new metrics.
If you have prior experience wiring observability pipelines on Databricks or tuning Prometheus, Grafana, Loki, or similar tools inside Kubernetes, that will accelerate the hand-off, but I’m open to alternative stacks as long as the above criteria are met.