Flickr8K captions using tensorflow
Budget: $30 – $250 USD
• Use a pre-trained CNN as an image encoder
(e.g., VGG, Inception,.. remember to use their corresponding pre-processing)
• Pre-process the training captions
• Train a RNN decoder on a word-to-word level
(use your choice of words embedding)
• Show a few (good and bad) examples of the generated captions
(use images from the validation set for inference)
Links for the dataset:
https://github.com/jbrownlee/Datasets/releases/download/Flickr8k/Flickr8k_Dataset.zip
https://github.com/jbrownlee/Datasets/releases/download/Flickr8k/Flickr8k_text.zip
(e.g., VGG, Inception,.. remember to use their corresponding pre-processing)
• Pre-process the training captions
• Train a RNN decoder on a word-to-word level
(use your choice of words embedding)
• Show a few (good and bad) examples of the generated captions
(use images from the validation set for inference)
Links for the dataset:
https://github.com/jbrownlee/Datasets/releases/download/Flickr8k/Flickr8k_Dataset.zip
https://github.com/jbrownlee/Datasets/releases/download/Flickr8k/Flickr8k_text.zip