machine learning python

Job ID: 30690050

Budget: $30 – $250 USD

Webscraping assignment of tapas restaurants on Yelp. Project explained below. Need it urgently!

Yelp, Inc., founded in 2004 by two former PayPal employees, develops, hosts, and markets the yelp.com website and the Yelp mobile app, which publish crowd-sourced reviews about businesses. It also operates Yelp Reservations, a table reservation service. It is headquartered in San Francisco, California.
Yelp’s website, yelp.com, is a crowd-sourced local business review and social networking site. The site has pages devoted to individual locations, such as restaurants or schools, where Yelp users can submit a review of their products or services using a one to five star rating scale. Businesses can also update contact information, hours, and other basic listing information or add special deals. In addition to writing reviews, users can react to reviews, plan events, or discuss their personal lives.
Broad description of the project
Suppose that you have a business plan for a chain of tapas restaurants in the US, serving Spanish staples such as paella, pulpo a la brasa, sangría, etc. The first step would be made in New York, where tapas food is already popular. Without leaving your office, you plan an exploratory study of the opinion of the customers of the existing tapas restaurants using the reviews published in yelp.com.
Using the search strings desc="tapas", loc="New York, NY" and attrs="RestaurantsPriceRange2.3", you end up in the webpage
https://www.yelp.com/search?find_desc=tapas&find_loc=New%20York%2C%20NY&attrs =RestaurantsPriceRange2.3
This gives you 24 pages with 10 restaurants per page. Your first finding is that the term ‘tapas’ has a broad scope in New York City, beyond the typical Spanish tapas. Since the restaurants come with tags such as ‘Spanish’, ‘Mediterranean’, ‘Seafood’, ‘Greek’, etc, you would select the restaurants with the tag ‘Spanish’, to focus the study.

Specific tasks
• Extract from these pages a list of restaurants with more than 500 reviews and tag ‘Spanish’. Include information on the neighbourhood (top right corner, below phone and address) and rating (star system).
• Summarize what you have in this list.
• For the restaurants of the list, scrape the corresponding Yelp pages, capturing the reviews and the ratings coming with the reviews.
• Your analysis may address the following questions:
– How is the distribution of the ratings?
– What is the ranking of the restaurants based on the average rating?
– What is the average length of the reviews?
– Which are the most frequent terms in the reviews?
– Which are the terms associated to high ratings?
– Which ones to low ratings?