Document Similarity Using Machine Leaning
The client had a corpus of documents and a set of similar documents(training set). The requirement was to identify the documents in the corpus which are similar to the training set and assign the similarity score to each document. The corpus had tens of thousands of documents so the algorithm had to run in a cluster. The solution is implemented in Python. We used Jupyter Notebook, Tensorflow, and sci-kit-learn to achieve required functionality. We used the Cosine similarity algorithm to match the similarity between the documents. We used AWS Sagemaker as a notebook instance and AWS spot instances to run the classification. The project was successfully completed in less than three weeks. The use of AWS spot instances reduced project cost up to 70%.
About Me
Certified team of AWS Solutions Architect, Google Web Specialist and Salesforce developers. We hold 8+ active certifications in these technologies. We have built our knowledge and expertise to help transform enterprises in adapting and leveraging the latest in technology. We work in Cloud, AI/ML, Digital, Mobility and [login to view URL] We provide following services: -> Mobile Development -> Cloud Consulting and Services - AWS, Google and Azure -> Salesforce Development -> Digital Transformation -> Machine Learning -> Data Extraction/ETL -> Data Visualization -> Ecommerce Development -> Email-Marketing Automation -> SEO - Search Engine Optimization -> SMM Social Media Marketing
$ 25 USD/hr
NEW FREELANCER!