Multi-Agent Environment wrapper for Python RLlib PPO -- 2
Budget: £20 – £250 GBP
I'm seeking an experienced Python freelancer to implement RLlib's PPO for a custom multi-agent environment, where agents act independently. Your tasks will be:
- Solving the issue with the learning process. Despite configuring the simplest setting with just one agent, the model fails to find the optimal action. Additionally, the model seems unable to distinguish between the optimal actions for agents with the same action space but different endowments.
Ideal skills include proficiency in Python, familiarity with RLlib's PPO algorithm, and previous work dealing with agent-based environments. Experience in crafting efficient reward functions is a plus.
- Solving the issue with the learning process. Despite configuring the simplest setting with just one agent, the model fails to find the optimal action. Additionally, the model seems unable to distinguish between the optimal actions for agents with the same action space but different endowments.
Ideal skills include proficiency in Python, familiarity with RLlib's PPO algorithm, and previous work dealing with agent-based environments. Experience in crafting efficient reward functions is a plus.