Online python app 70 hours of work Max ( any more please dont contact me)
Budget: $2 – $8 USD
I need to build an online python app this is the full detail of the exercixe.
The company's desire is to preserve fruit biodiversity by allowing specific treatments for each fruit species by developing intelligent picking robots.
Your start-up initially wants to make itself known by making available to the general public a mobile application that would allow users to take a picture of a fruit and obtain information on this fruit.
For the start-up, this application would make the general public aware of fruit biodiversity and set up a first version of the fruit image classification engine.
In addition, the development of the mobile application will make it possible to build a first version of the necessary Big Data architecture.
Data
Your colleague Paul tells you about the existence of a data set made up of fruit images and associated labels, which could serve as a starting point for building part of the data processing chain.
Your mission
You are therefore responsible for developing in a Big Data environment an initial data processing chain that will include preprocessing and a dimension reduction step.
There is no need to train a model at this time.
The important thing is to set up the first processing bricks that will be used when it is necessary to scale up in terms of data volume!
Constraints
During his initial brief, Paul warned you of the following points:
You will have to take into account in your developments the fact that the volume of data will increase very quickly after the delivery of this project. You will therefore develop scripts in Pyspark and use, for example, the AWS cloud to take advantage of a Big Data architecture (EC2, S3, IAM), based on an EC2 Linux server.
Implementing a Big Data architecture under (for example) AWS may require a more powerful server configuration than the one offered for free (EC2 = t2.micro, 1 GB RAM, 8 GB server disk).
This cost, which should remain less than 10 euros for reasonable use, remains your responsibility. Using a local server for design, limiting the use of the EC2 server for implementation and testing, can significantly reduce this cost.
Expected deliverables
A notebook on the cloud containing the executable Pyspark scripts (the preprocessing and a dimension reduction step).
The initial dataset images as well as the dimension reduction output (a matrix written to a CSV or other file) available in cloud storage.
A presentation support for the defense, presenting:
the different architectural bricks chosen on the cloud;
their role in Big Data architecture;
the stages of the processing chain.
USEFUL LINKS:
https://spark.apache.org/docs/latest/
https://runawayhorse001.github.io/LearningApacheSpark/pyspark.pdf
https://github.com/apache/spark/tree/master/examples/src/main/python/mllib
link of others peoples work and the data
https://drive.google.com/drive/folders/1mruf4yE_dmPej-GEqkG2DahDP04SNgbz?usp=sharing
The company's desire is to preserve fruit biodiversity by allowing specific treatments for each fruit species by developing intelligent picking robots.
Your start-up initially wants to make itself known by making available to the general public a mobile application that would allow users to take a picture of a fruit and obtain information on this fruit.
For the start-up, this application would make the general public aware of fruit biodiversity and set up a first version of the fruit image classification engine.
In addition, the development of the mobile application will make it possible to build a first version of the necessary Big Data architecture.
Data
Your colleague Paul tells you about the existence of a data set made up of fruit images and associated labels, which could serve as a starting point for building part of the data processing chain.
Your mission
You are therefore responsible for developing in a Big Data environment an initial data processing chain that will include preprocessing and a dimension reduction step.
There is no need to train a model at this time.
The important thing is to set up the first processing bricks that will be used when it is necessary to scale up in terms of data volume!
Constraints
During his initial brief, Paul warned you of the following points:
You will have to take into account in your developments the fact that the volume of data will increase very quickly after the delivery of this project. You will therefore develop scripts in Pyspark and use, for example, the AWS cloud to take advantage of a Big Data architecture (EC2, S3, IAM), based on an EC2 Linux server.
Implementing a Big Data architecture under (for example) AWS may require a more powerful server configuration than the one offered for free (EC2 = t2.micro, 1 GB RAM, 8 GB server disk).
This cost, which should remain less than 10 euros for reasonable use, remains your responsibility. Using a local server for design, limiting the use of the EC2 server for implementation and testing, can significantly reduce this cost.
Expected deliverables
A notebook on the cloud containing the executable Pyspark scripts (the preprocessing and a dimension reduction step).
The initial dataset images as well as the dimension reduction output (a matrix written to a CSV or other file) available in cloud storage.
A presentation support for the defense, presenting:
the different architectural bricks chosen on the cloud;
their role in Big Data architecture;
the stages of the processing chain.
USEFUL LINKS:
https://spark.apache.org/docs/latest/
https://runawayhorse001.github.io/LearningApacheSpark/pyspark.pdf
https://github.com/apache/spark/tree/master/examples/src/main/python/mllib
link of others peoples work and the data
https://drive.google.com/drive/folders/1mruf4yE_dmPej-GEqkG2DahDP04SNgbz?usp=sharing