Flash Attention of LLM math algorithm concept expert
Budget: $30 – $250 USD
1. Hi, I'm seeking to comprehend this Flash Attention paper.
https://arxiv.org/pdf/2205.14135
2. The paper contains many computer science concepts, such as IO complexity and the process of storing memory into HBM versus SRAM. I suspect that I may even need to grasp CUDA programming to thoroughly understand the content. I'm keen to discuss this paper and its concepts with someone who is well-versed in these areas, and to review the paper together.
3. Ultimately, I hope to gain sufficient knowledge to implement the content into my elementary, large language model, which I may create from scratch, and witness its faster training rate, among other things.
4. If you've read up to this point, please include "I know how to do it" at the beginning of your proposal.
5. If our collaboration proves successful, we can delve into other topics, such as Flash Attention two and Mamba, as I mentioned in the video link below. However, let's begin with the Flash Attention one paper.
6. below is the full video of my question:
https://www.dropbox.com/scl/fi/ndcpuejmsn7ewlljpvd66/Screen-Recording-2024-06-04-at-2.58.37-PM.mov?rlkey=1wxbtg801x9z2kh6scjev4pg1&dl=0
https://arxiv.org/pdf/2205.14135
2. The paper contains many computer science concepts, such as IO complexity and the process of storing memory into HBM versus SRAM. I suspect that I may even need to grasp CUDA programming to thoroughly understand the content. I'm keen to discuss this paper and its concepts with someone who is well-versed in these areas, and to review the paper together.
3. Ultimately, I hope to gain sufficient knowledge to implement the content into my elementary, large language model, which I may create from scratch, and witness its faster training rate, among other things.
4. If you've read up to this point, please include "I know how to do it" at the beginning of your proposal.
5. If our collaboration proves successful, we can delve into other topics, such as Flash Attention two and Mamba, as I mentioned in the video link below. However, let's begin with the Flash Attention one paper.
6. below is the full video of my question:
https://www.dropbox.com/scl/fi/ndcpuejmsn7ewlljpvd66/Screen-Recording-2024-06-04-at-2.58.37-PM.mov?rlkey=1wxbtg801x9z2kh6scjev4pg1&dl=0
Related categories:
CUDA
C++ Programming
Computer Science
Large Language Model
Large Language Models (LLMs)