SWE-Bench AI Model Integration
Budget: $30 – $250 USD
I’m seeking an expert to help integrate and optimize AI models such as Lingxi, ExpeRepair, and Claude 4.5 Sonnet to improve performance on SWE-Bench. The objective is to create a model that exceeds current public benchmarks, with a minimum target of 60% accuracy on SWE-Bench Lite and 74% accuracy on SWE-Bench Verified. This is a focused one-week project aimed at producing high-quality, reproducible evaluation results.
The task involves integrating the existing model architectures, fine-tuning or combining them where necessary, and preparing a fully functional repository that can run both locally and on the cloud. You’ll be responsible for producing a predictions.json file, detailed benchmark reports, and clean, readable documentation outlining setup, evaluation steps, and usage instructions. The end goal is a working solution that can be validated against the official SWE-Bench leaderboard.
You’ll have access to all essential resources, including a Claude 4.5 Sonnet API key and, if preferred, Modal Cloud infrastructure for evaluation.
You have full creative and technical freedom to use whichever architecture or hybrid approach you believe will perform best. We only care about the final benchmark results and clean, reproducible documentation
Please let me know if you can do this in 1 week.
The task involves integrating the existing model architectures, fine-tuning or combining them where necessary, and preparing a fully functional repository that can run both locally and on the cloud. You’ll be responsible for producing a predictions.json file, detailed benchmark reports, and clean, readable documentation outlining setup, evaluation steps, and usage instructions. The end goal is a working solution that can be validated against the official SWE-Bench leaderboard.
You’ll have access to all essential resources, including a Claude 4.5 Sonnet API key and, if preferred, Modal Cloud infrastructure for evaluation.
You have full creative and technical freedom to use whichever architecture or hybrid approach you believe will perform best. We only care about the final benchmark results and clean, reproducible documentation
Please let me know if you can do this in 1 week.