Create a cluster of graphics cards to be used by nodes running FreqTrade/FreqAI == Creación de clúster de tarjetas gráficas para ser usado por nodos ejecutando FreqTrade/FreqAI

Job ID: 36160097

Budget: $250 – $750 USD

There is (in office) a total of 80 RTX 3090 graphics cards, which must be used to create a GPU cluster using the Kubernetes management platform (or similar). The cluster must be configured in such a way that it can manage the GPUs which will be used by computers with the FreqTrade/FreqAI software (node), and each node will have at its disposal X number of graphics cards. The total number of GPU cards is distributed as follows:

- 2 computers with 2 GPUs each
- 4 computers with 3 GPUs each
- 16 computers each with 4 GPUs each

This distribution of GPU cards must be respected, that is, more cards cannot be assigned to a computer. One of those nodes will be the main computer, on which the Kubernetes will be running. The nodes must be configured so that they can communicate using the TCP/IP protocol, just like the main computer.

Also, it is required that the cluster be configured in such a way that it can be easily managed, with which it is possible to verify the performance of the GPUs, adjust the allocation of resources, solve problems, etc.

Since the project would be carried out remotely, it is necessary that we be given instructions to configure each of the nodes, as well as basic settings so that it can be accessed remotely. For this, it is proposed to configure the main computer with Linux Desktop and some remote administration software (Teamviewer or AnyDesk) and in the nodes install a version of Linux Desktop with SSH enabled so that the main computer can connect to them and thus install software and necessary configuration (FreqTrade, FreqAI, firewall, etc).

== == == == == == == == == == == == == == == == == == ==

Se cuenta (en oficina) con un total de 80 tarjetas gráficas RTX 3090, las cuales deben ser utilizadas para crear un clúster de GPU utilizando la plataforma de gestión de Kubernetes (o similar). El clúster debe ser configurado de manera que se pueda gestionar las GPU las cuales serán usadas por computadoras con el software FreqTrade/FreqAI (nodo), y cada nodo tendrá a su disposición una cantidad X de tarjetas gráficas. El total de tarjetas GPU está distribuido de la siguiente manera:

- 2 computadoras con 2 GPU c/u
- 4 computadoras con 3 GPU c/u
- 16 computadoras cada una con 4 GPU c/u

Esa distribución de tarjetas GPU debe ser respetada, o sea que no se pueden asignar más tarjetas a una computadora. Uno de esos nodos será el equipo principal, en el cual se estará ejecutando el Kubernetes. Los nodos deberán ser configurados de manera que se puedan comunicar empleando el protocolo TCP/IP, al igual que el computador principal.

También, se requiere que el clúster quede configurado de manera que pueda ser fácilmente administrado, con lo cual se pueda verificar el rendimiento de las GPU, ajustar la asignación de recursos, solucionar problemas, etc.

Dado que el proyecto se estaría realizando de manera remota, es necesario que se nos den las instrucciones para configurar cada uno de los nodos, así como ajustes básicos para que el mismo pueda ser accedidos de manera remota. Para esto se propone configurar el equipo principal con Linux Desktop y algún software de administración remota (Teamviewer o AnyDesk) y en los nodos instalar una versión Linux Desktop con SSH habilitado para que el equipo principal se pueda conectar a ellos y así hacer la instalación de software y configuración necesaria (FreqTrade, FreqAI, firewall, etc).
Related categories: Linux Network Administration Kubernetes