Boost Boolean Pipeline Processing Speed Cray-3
Budget: $250 – $750 USD
I need to squeeze noticeably more raw speed out of the Boolean scalar pipeline that runs on our Cray-3 code base. At the moment the design works, but it is not hitting the latency figures we believe are achievable. The single focus is processing-speed optimisation; memory footprint and new feature work can stay as-is.
You will start by profiling the existing flow, call out the exact stages where cycles are being lost, then propose the most cost-effective micro-architectural or low-level code changes to trim those stalls. Because we have no hard performance targets yet, part of the job is to recommend a realistic benchmark suite and target numbers that make sense for this machine.
Once the targets are agreed, implement the improvements, re-run the benchmarks, and provide a concise report showing before-and-after cycle counts so we can see the real-world gains. If you need to adjust instruction scheduling, pipeline depth, branch handling, or hand-tune assembly, that is all fair game—as long as the final artefact stays fully compatible with the current toolchain.
Deliverables
• Benchmark plan with justified target speeds
• Optimised pipeline source/configuration
• Benchmark results and comparison report
• Short hand-off note explaining any maintenance implications
If you have previous experience pushing legacy supercomputer architectures—or any pipeline-heavy design—to their limit, I’d love to tap into that expertise and get this Cray-3 path humming.
You will start by profiling the existing flow, call out the exact stages where cycles are being lost, then propose the most cost-effective micro-architectural or low-level code changes to trim those stalls. Because we have no hard performance targets yet, part of the job is to recommend a realistic benchmark suite and target numbers that make sense for this machine.
Once the targets are agreed, implement the improvements, re-run the benchmarks, and provide a concise report showing before-and-after cycle counts so we can see the real-world gains. If you need to adjust instruction scheduling, pipeline depth, branch handling, or hand-tune assembly, that is all fair game—as long as the final artefact stays fully compatible with the current toolchain.
Deliverables
• Benchmark plan with justified target speeds
• Optimised pipeline source/configuration
• Benchmark results and comparison report
• Short hand-off note explaining any maintenance implications
If you have previous experience pushing legacy supercomputer architectures—or any pipeline-heavy design—to their limit, I’d love to tap into that expertise and get this Cray-3 path humming.