Secure On-Prem ELK SIEM

Job ID: 39950084

Budget: $750 – $1,500 USD

I’m standing up a new on-premises SIEM and want it powered by the ELK stack. The core goal is security information and event management (SIEM) with log ingestion from Proxmox and other internal sources, ending in actionable Kibana dashboards.

ELK Implementation – Design, deployment, HA/DR setup, documentation & training
Scope of Work:
- Design and deploy optimal ELK architecture on Proxmox infrastructure with High Availability (HA) and Disaster Recovery (DR) setup.
- Define storage policies (ILM hot/warm/cold), backup/restore procedures, and automatic rollover configuration.
- Conduct performance, load, and fault-tolerance testing to validate stability and scalability.
- Deliver configuration documentation, SOP/runbooks, and conduct training for the operation team.
- Integrate initial log sources rebuilt to align with SOC operational requirements (Firewall, IDS/IPS, OS, DB/Apps, EDR, Network, Cloud, SaaS).
Deliverables:
- Fully configured multi-node ELK cluster (tolerates one-node failure without service disruption).
- Documented configuration, test reports, and operational training handover.
- Acceptance testing and final handover report.

Here’s what I need from you:
• Architecture & Design – size and lay out Elasticsearch, Logstash, and Kibana for high availability, including node roles, shard strategy, and network segmentation.
• Deployment – install and configure the stack on our Proxmox hosts, automate wherever practical (Ansible, Terraform, or similar) and ensure all services start cleanly after reboot.
• HA / DR – implement cluster redundancy, snapshot scheduling, and a documented fail-over procedure so we can recover quickly if a node or site goes down.
• Hardening & Tuning – role-based access control, TLS, index lifecycle management, and pipeline optimizations tailored for SIEM workloads.
• Dashboards & Alerts – create a starter set of Kibana visualizations and security alerts that help us catch critical events immediately.
• Documentation & Training – a concise runbook covering architecture, daily operations, and recovery steps, followed by a live knowledge-transfer session for our admins.

Acceptance criteria
1. Cluster survives single-node failure with no data loss.
2. Daily snapshots replicate to secondary storage and restore cleanly in a test.
3. Dashboards populate with real Proxmox log data and update in near-real time.
4. All steps are reproducible from the provided automation scripts and documentation.

If you have proven experience building resilient ELK environments for security monitoring, I’d love to work with you.