Adversarial Misinformation Generation Framework
Budget: ₹100 – ₹400 INR
I need a Python program that implements an adversarial training framework to fine-tune GPT-3 for generating misinformation that evades AI text detectors.
Key Components:
- Generator Model: Uses prompts to create fake news.
- Detector Model: Classifies text as AI-generated or human-written.
Training Process:
- The generator uses PPO to minimize the detector's confidence in its classification.
- The detector is periodically retrained on increasingly deceptive texts to enhance its robustness.
Requirements:
- Proficiency in Python and machine learning
- Experience with reinforcement learning, particularly PPO
- Familiarity with language models, especially GPT-3
- Knowledge of adversarial training techniques
The final output should include:
- A robust AI misinformation generator
- A resilient detector trained on adversarial examples
Please provide tools and datasets for evaluating and improving detection systems.
Key Components:
- Generator Model: Uses prompts to create fake news.
- Detector Model: Classifies text as AI-generated or human-written.
Training Process:
- The generator uses PPO to minimize the detector's confidence in its classification.
- The detector is periodically retrained on increasingly deceptive texts to enhance its robustness.
Requirements:
- Proficiency in Python and machine learning
- Experience with reinforcement learning, particularly PPO
- Familiarity with language models, especially GPT-3
- Knowledge of adversarial training techniques
The final output should include:
- A robust AI misinformation generator
- A resilient detector trained on adversarial examples
Please provide tools and datasets for evaluating and improving detection systems.