Build a Handwritten Text Recognition Model for a Historical Vineyard Diary Collection

Job ID: 40576149

Budget: $1,500 – $3,000 AUD

I am looking for an experienced Transkribus / Handwritten Text Recognition (HTR) specialist to help develop a high-quality transcription model for a large collection of handwritten family diaries.

The collection consists of several thousand pages of vineyard and farming diaries from South Australia, spanning approximately 1890–1947. The diaries record daily vineyard work, livestock management, crop planting, weather, visitors, grape varieties, local events and family history.

The ultimate goal is to create a reliable transcription model that can be used across the entire collection.

Rather than transcribing every page manually, I am seeking someone who can:

- Review a representative sample of diaries.
- Determine an appropriate training and validation approach.
- Train a Transkribus (or equivalent HTR) model.
- Demonstrate that the model can accurately transcribe previously unseen diary pages.

I am interested in working with someone who understands both:

- historical handwriting and transcription, and
- training and evaluation of handwritten text recognition models.

I can provide a representative subset of diaries covering different periods and handwriting styles. I expect to retain a separate set of diaries and pages that will not be used during training so that the model can be tested independently.

I am open to your recommended approach, but I would expect something along the lines of:

- A trained HTR model.
- Documentation of the training process.
- Validation results against pages not used during training.
- Recommendations for further improving the model.
- A demonstration of transcription quality on previously unseen diary pages.

This will likely be a two-phase engagement: (1) pilot assessment and proof of concept, followed by (2) full implementation if the pilot is successful.

Please tell me:

1. Your experience with Transkribus or other HTR platforms.
2. Any previous historical manuscript projects you have worked on.
3. How you would approach training and validating a model for a large collection like this.
4. How you would measure success.
5. Examples of similar projects, if available.