Developing a dataset of attacks on AI large language models a used for security games
Budget: $30 – $250 USD
I am working on research with a group and I need help with a specific task. (I will share the details on chat)
The goal of the task is to write a bunch of attacks/test cases for each scenario.
General research description (but I am only working on a specific task):
Do research on the security of large language models (LLMs), like ChatGPT. Evaluating how well they can enforce rules, and whether someone can trick them into violating the rules. Building a benchmark to measure this, by building a number of simple games that we can play with a LLM. We need to put together a list of attacks that someone might use to try to fool the LLM.
We will create attacks, i.e., writing sentences that could send to the LLM to try to trick it into doing something that violates the rules of the game. This involves trying to attack ChatGPT in as many ways as possible and creating a dataset of all those attempts.
The goal of the task is to write a bunch of attacks/test cases for each scenario.
General research description (but I am only working on a specific task):
Do research on the security of large language models (LLMs), like ChatGPT. Evaluating how well they can enforce rules, and whether someone can trick them into violating the rules. Building a benchmark to measure this, by building a number of simple games that we can play with a LLM. We need to put together a list of attacks that someone might use to try to fool the LLM.
We will create attacks, i.e., writing sentences that could send to the LLM to try to trick it into doing something that violates the rules of the game. This involves trying to attack ChatGPT in as many ways as possible and creating a dataset of all those attempts.