Fix RAG Chatbot Tokens
Budget: $15 – $25 USD
My retrieval-augmented chatbot reads uploaded PDFs, indexes them, and answers questions through a custom Python script built with Hugging Face Transformers. Recently the system started throwing “maximum tokens” errors and dropping responses.
I’m looking for someone who can:
• Inspect the current code and reproduce the token-overflow issue.
• Adjust the chunking, prompt construction, or model parameters so each request stays within limits without sacrificing answer quality.
• Test with several sample PDFs to confirm stable, complete responses.
• Hand back the updated script plus a concise summary of what changed so I can maintain it myself.
The scope is intentionally small—just a focused fix, no large rewrites—so I expect the work to wrap up in a single short session. If you’re comfortable debugging Python and Transformers-based RAG pipelines, I’d love your help today.
I’m looking for someone who can:
• Inspect the current code and reproduce the token-overflow issue.
• Adjust the chunking, prompt construction, or model parameters so each request stays within limits without sacrificing answer quality.
• Test with several sample PDFs to confirm stable, complete responses.
• Hand back the updated script plus a concise summary of what changed so I can maintain it myself.
The scope is intentionally small—just a focused fix, no large rewrites—so I expect the work to wrap up in a single short session. If you’re comfortable debugging Python and Transformers-based RAG pipelines, I’d love your help today.