Building an LLM model by developing a DL framework and DL compiler from scratch
Budget: $50 – $499 USD
I’m leading a small engineering team that is building a heterogeneous deep-learning pipeline and we want an expert mentor to steer us—especially on custom kernel development. We have not yet committed to a single framework; PyTorch, TensorFlow and JAX are all possibilities, so the guidance you provide will influence that decision.
What you’ll do for us
• Meet either continuously or in weekly 1–2 hour sessions (whichever fits your schedule)
• Review the code we produce, point out performance bottlenecks and suggest remedies
• Teach us how to write, integrate and maintain custom CUDA/ROCm kernels that play nicely with TVM, Glow, XLA, TensorRT or MLIR pipelines
• Share best-practice deployment tips for CPU-GPU systems and mixed-precision execution
What we need to see from you
Reply with a concise summary of your experience solving similar problems—successful kernel integrations, compiler optimisations or large-scale model deployments you’ve led. The stronger and more recent your hands-on experience, the better the fit.
If mentoring and shaping an eager team excites you, let’s talk.
What you’ll do for us
• Meet either continuously or in weekly 1–2 hour sessions (whichever fits your schedule)
• Review the code we produce, point out performance bottlenecks and suggest remedies
• Teach us how to write, integrate and maintain custom CUDA/ROCm kernels that play nicely with TVM, Glow, XLA, TensorRT or MLIR pipelines
• Share best-practice deployment tips for CPU-GPU systems and mixed-precision execution
What we need to see from you
Reply with a concise summary of your experience solving similar problems—successful kernel integrations, compiler optimisations or large-scale model deployments you’ve led. The stronger and more recent your hands-on experience, the better the fit.
If mentoring and shaping an eager team excites you, let’s talk.