Jetson Edge LLM Deployment Specialist
Budget: ₹250,000 – ₹500,000 INR
I need to turn a Large Language Model into a practical, real-time assistant that runs directly on an NVIDIA Jetson board. The model will reason over image data at the edge, so every millisecond saved and every megabyte spared matters. Memory-optimised execution is therefore the single most important constraint, though I still expect you to keep latency low and power draw sensible.
Here is what I want to achieve. The LLM must accept a pre-processed visual input, apply prompt-engineering tricks that preserve context, and reply fast enough to be useful on-device—no cloud fallback. Smart caching, selective quantisation, and, when it genuinely pays off, lightweight fine-tuning are all on the table. I will look to you to suggest the right mix of techniques and to implement them.
You should be comfortable with TensorRT, NVIDIA Triton, or any other inference engine that squeezes the most out of Jetson GPUs; hands-on experience with model compression libraries such as bits-and-bytes, FasterTransformer, or similar will help convince me you can meet the memory target. If a custom data pipeline is needed to translate raw camera frames into embeddings the LLM can consume, please include that in your plan.
Deliverables
• A working Jetson image or container that boots straight into the optimised LLM service, ready to accept image-derived prompts and respond in real time.
• Source code, build scripts, and concise documentation describing the preprocessing, caching, and memory-saving strategies used, plus steps to reproduce results on a fresh device.
• A short benchmark report demonstrating memory footprint and end-to-end latency under typical input sizes.
Acceptance will be based on the ability to hit the promised memory budget while maintaining interactive response times on the specified Jetson hardware.
Here is what I want to achieve. The LLM must accept a pre-processed visual input, apply prompt-engineering tricks that preserve context, and reply fast enough to be useful on-device—no cloud fallback. Smart caching, selective quantisation, and, when it genuinely pays off, lightweight fine-tuning are all on the table. I will look to you to suggest the right mix of techniques and to implement them.
You should be comfortable with TensorRT, NVIDIA Triton, or any other inference engine that squeezes the most out of Jetson GPUs; hands-on experience with model compression libraries such as bits-and-bytes, FasterTransformer, or similar will help convince me you can meet the memory target. If a custom data pipeline is needed to translate raw camera frames into embeddings the LLM can consume, please include that in your plan.
Deliverables
• A working Jetson image or container that boots straight into the optimised LLM service, ready to accept image-derived prompts and respond in real time.
• Source code, build scripts, and concise documentation describing the preprocessing, caching, and memory-saving strategies used, plus steps to reproduce results on a fresh device.
• A short benchmark report demonstrating memory footprint and end-to-end latency under typical input sizes.
Acceptance will be based on the ability to hit the promised memory budget while maintaining interactive response times on the specified Jetson hardware.