Code Extraction & Simplification for text embedding BERT Model -- 3

Job ID: 38434913

Budget: €30 – €250 EUR

<|endoftext|>

I'm looking for a skilled data scientist with experience in Python, PyTorch, and huggingface to help me extract the training code for a BERT model from a GitHub repository and simplify it in 1-2 Google Colab notebooks.

Key Requirements:
- Extract and simplify the training code of a BERT model from the given GitHub repository.
- The simplified code should cover key functionalities such as data preprocessing, model training, and model evaluation.
<|endoftext|>
Ideal Skills and Experience:
- Proficiency in Python, particularly in working with PyTorch/HuggingFace
- Strong background in NLP and working with BERT models.
- Experience with huggingface and understanding of transformer-based models.
- Ability to simplify complex code, making it more readable and understandable.

Your main task will be to simplify the existing code, ensuring its functionality is preserved while making it more accessible for users less familiar with the intricacies of the original code. If you have experience with similar projects, please provide examples of your work.

Model: https://huggingface.co/nomic-ai/nomic-embed-text-v1.5
Repo: https://github.com/nomic-ai/contrastors
Related categories: BERT Hugging Face