Scalable Cluster Analysis Algorithms for Big Data Insights

Job ID: 37276303

Budget: ₹12,500 – ₹37,500 INR

QoS-based performance measurement for service management in cloud and bigdata analysis. Every day, a substantial volume of data is generated from a variety of sources, including IoT networks, smartphones, and activities on social networks. It is imperative for numerous businesses, services, and applications to make sense of this unprecedented data influx. Presently, there is a shortage of domain expertise to automate the analysis of such big data, and conventional supervised machine learning techniques face challenges due to the lack of labeled training data in this context. The objective is to create scalable and efficient algorithms capable of managing and extracting actionable insights from big data.

Cluster analysis is a useful unsupervised approach to discover the underlying groups and useful patterns in the data. Cluster analysis, when applied to any dataset, entails addressing three primary challenges: (P1) cluster assessment, which seeks to answer whether the data exhibits clusters and if so, how many; (P2) clustering, which involves partitioning the data into distinct clusters; and (P3) cluster validity, which assesses the utility of the discovered clusters and explores the possibility of uncovering better clusters that may have been overlooked. Traditional cluster analysis algorithms are not well-suited for big data due to its characteristics of high volume, diversity, and rapid velocity.

In this work, A suite of scalable algorithms will be developed to address each of the three key challenges in cluster analysis for big data. These challenges encompass cluster assessment, clustering, and cluster validity, all of which are particularly relevant for big data scenarios characterized by high dimensionality, anomalies, and streaming data.
Related categories: Machine Learning (ML) Docker Kubernetes DevOps Big Data