Test AGIBIOS Guardrail Across LLMs

Job ID: 40348885

Budget: $1,500 – $3,000 AUD

I’m looking for a practitioner who can take the open-source dystopiabench suite and run it end-to-end on a collection of models—Opus, GLM, MiniMax, GPT, Codex Gemini, Pro Qwen, Grok, Kimi, DeepSeek, and Mistral—first in their vanilla form and then again with my open-source guardrail prompt, AGIBIOS, prepended to every request.

The core question is simple: does AGIBIOS meaningfully reduce failure rates when the benchmark tries to nudge a model, step by step, toward unethical outcomes? For this round I’m interested only in the “Ethical decision making” scenarios; dystopiabench already contains the probes.

Scope of work
• Set up or containerise dystopiabench and confirm it runs reproducibly.
• Integrate AGIBIOS so the exact same benchmark can be executed twice per model (with and without the guardrail).
• Capture raw logs, failure/success rates, and token-level latency where available.
• Produce a concise comparative report (tables + brief commentary) highlighting any statistically significant shifts in behaviour.

Deliverables
1. Re-usable scripts or notebooks that launch each test run.
2. CSV/JSON logs for every interaction.
3. A short markdown or PDF report summarising results, methodology, and how to replicate.

Acceptance criteria
• Runs complete on all listed models without manual intervention.
• Report clearly states methodology and shows before/after metrics for ethical decision making.
• Code is clean, documented, and runnable on a fresh machine with standard Python and Docker tooling.
• Duplicate the distopiabench website with both the native and AGIBIOS enhanced graphs.


You will need to create the keys for these models and where possible reach out to the labs to your keys white-listed for red-teaming use.