Computer Science Dissertation

Department of Computer Science

Covid-19 Data Analysis from Social Media

FirstName(s) LastName

Supervisor: Supervisor’s Name

A report submitted in partial fulfilment of the requirements for the degree of Bachelor of Science in Computer Science

[Date]

Declaration

I, [Firstname(s)] [Lastname], of the Department of Computer Science, confirm that this is my own work and figures, tables, equations, code snippets, artworks, and illustrations in this report are original and have not been taken from any other person’s work, except where the works of others have been explicitly acknowledged, quoted, and referenced. I understand that if failing to do so will be considered a case of plagiarism. Plagiarism is a form of academic misconduct and will be penalized accordingly. I give consent to a copy of my report being shared with future students as an exemplar. I give consent for my work to be made available more widely and public with interest in teaching, learning and research.

[Firstname(s)] [Lastname]

[DATE]

Abstract

The term big data means large datasets that are analyzed by use of computers to reveal patterns. Big data is mostly available on social media with a lot of data being exchanged between users on an everyday basis, hence the purpose of this project, which is to carry out memotion analysis on covid-19 data acquired from social media. To accomplish the purpose of this project machine learning methods will be applied to the data collected these methods include use of the Natural Language Processing toolkit, Python TextBlob and Python ensemble learning. By use of these methods, the project will accomplish sentiment analysis on this the data and the results will be classified data into three main categories namely positive, negative and neutral. Addition analysis will be carried out to establish the most commons words that appear across these different categories and across the entire dataset. From this analysis we can then come to the conclusion of what are the overall sentiments around this covid-19 pandemic, we can establish if people are panicking or hopeful and how they are handling this virus. The results acquired from this project can be used by governments, hospitals and institutions to understand what the general public feel about this covid-19 pandemic. In this way these institutions can implement different strategies to keep this covid-19 pandemic under control.

Acknowledgements

 An acknowledgements section is optional. You may like to acknowledge the support and help of your supervisor(s), friends, or any other person(s), department(s), institute(s), etc. If you have been provided specific facility from department/school acknowledged so.

Contents

Chapter 1. 9

1        Introduction. 9

1.1    Background. 10

1.2    Problem Statement. 12

1.3    Aims and Objectives. 13

1.4    Solution approach. 14

1.4.1      Data collection. 14

1.4.2      Data Cleaning and preprocessing. 15

1.4.3      Natural Language Processing toolkit. 17

1.4.4      Employing machine learning algorithms. 17

1.5    Summary of contributions and achievements. 19

Chapter 2. 20

2        Literature Review.. 20

Chapter 3. 22

3        Methodology. 22

3.1    Task Description. 22

3.2    Descriptions. 23

3.2.1      Technology. 23

3.3    Implementation. 26

3.3.1      Data collection. 27

3.3.2      Authentication. 28

3.3.3      Tweet Extraction. 29

3.3.4      Data collection main method. 30

3.4    Accessing the Data. 34

3.5    Data preprocessing. 36

3.5.1      Emoji Removal 37

3.5.2      URLs Removal 37

3.5.3      Hashtag Removal 37

3.5.4      Lowercasing tweets. 38

3.5.5      Lemmatization. 38

3.5.6      Removing stop  words. 38

3.5.7      Summary. 38

3.6    Data visualization. 39

3.6.1      Most Frequent words. 39

3.6.2      Word clouds. 40

3.7    Experiments and Designing. 41

3.7.1      Method one: Valence Aware Dictionary for Sentiment Reasoning (VADER) 41

3.7.2      Method Two: Text Blob. 42

3.7.3      Method Three: Machine Learning Algorithms. 43

3.7.3.1          Naïve Bayes Classifier. 43

Bayes Theorem.. 44

Multinomial Classifier. 45

Bernoulli Classifier. 45

3.7.3.2          Linear Models. 46

Logistic classifier. 46

3.7.3.3          Support Vector Machines. 46

Linear SVC. 47

NuSVCClassifer. 47

3.7.4      Dashboard Implementation. 48

3.7.5      Summary. 50

Chapter 4. 51

4        Results. 51

4.1    Valence Aware Dictionary for Sentiment Analysis (VADER) 51

4.2    Python TextBlob. 53

4.3    Ensemble learning. 56

4.4 Dashboard. 57

Chapter 5. 58

5.1    Discussion and Analysis. 58

5.1.1      Python TextBlob. 58

TextBlob Neutral Tweets. 58

TextBlob Positive Tweets. 59

TextBlob Negative Tweets. 60

5.1.2      Valence Aware Dictionary (VADER) 61

VADER Positive sentiments. 62

VADER Negative Tweets. 62

VADER Neutral Tweets. 64

5.1.3      Ensemble Learning. 65

Positive tweets. 65

Negative tweets. 66

5.2    Significance of Findings. 66

5.3    Limitations. 67

5.4    Summary. 67

Chapter 6. 68

6.1    Conclusion. 68

6.2    Future Works. 69

Chapter 7. 70

7.1    Reflection. 70

References. 71

List of Abbreviations

NLTK                Natural Language Tool Kit.

API                   Application Programming Interface

VADER             Valence Aware Dictionary for sEntiment Reasoning

ML                   Machine Learning

CSV                  Comma Separated Values

URLs                Uniform Resource Locator

SVC                  Support Vector Classification

Chapter 1

1    Introduction

Among the fastest growing forms of interaction, social media tops them all with large of exchange of data between users of different platforms taking place each and every day. There has been an up surge in the usage of social media platforms over the past few years with companies like tweeter recording approximately 500 million tweets being shared on their application per day which equates to roughly 6000 tweets per second(David Sayce,2020). This number has seen an increase from 2018 where only 320 million tweets where being shared per day. This sheer amount of data being exchanged has however, been brought by the coronavirus pandemic as people are now resulting to social media platforms for social interaction. This is majorly due to the limited physical interaction as people heed to the restrictions put in place by their respective governments’ to flatten the coronavirus pandemic curve.

As the project topic suggests, the aim of this project is to leverage the vast amount of data found on different social media platforms to carry out memotion analysis in order to get the sentiments surrounding this corona virus pandemic. This is to be done with various data analysis techniques such as data mining and machine learning techniques. Programming language used in this project is python which is popular in the data science community. Apart from it being popular in the data science community, python is a well maintained programming language that has outstanding modules which will aid in the successful completion of this project.

With this project both theoretical and practical advanced machine learning techniques coupled with natural language processing techniques will be developed for analysis of textual data. Data collected from social media platforms will be processed to provide meaningful datasets that will be used for the purpose of training and testing models hence developing machine learning algorithms that can be used for memotion analysis and provide good classifications of new data fed into the models. Advanced python programming language concepts will be acquired, consequently building core skills when it comes to this programming language.

1.1    Background

On the onset of the coronavirus pandemic, many activities were forced to migrate to online platforms so as to limit physical interaction, as recommend by the World Health Organization, in order to curb the spreading of the novel corona virus. With different countries being placed under lockdown, businesses closed down and activities that required people to be physically present we limited in fact face to face meet ups were to be as considered as the last option. This turn of events, led to increased use of phones and specifically increased used of social media, with top topics being about the fast spreading corona virus. A lot of data was therefore being exchanged on these platforms about this particular topic and hence the need to collect this data, analyze it, organize it into meaningful data and come with some conclusions regarding this trending topic.

In the analysis of social media data, particularly sentiment analysis, there are different of types of analysis but the most important types of analysis are:

  1. Fine-Grained Analysis

This type of analysis aids in deriving fine-tuned polarity precision, that is, the polarity categories are precise such as very positive, positive, neutral, negative, or very negative. It is mostly used by companies in review ratings but can also be used in other areas to achieve the same results.

  • Emotion detection Analysis

This type of analysis helps in emotion detection such as anger, happiness, sadness, anxiety, panic etc.

  • Aspect-Based Analysis

Similar to fine-grained analysis but it in this one it dives deeper to point out exactly what is being termed as negative. An example would be a customer review stating, “The remote power button is hard to press”, aspect based analysis will point out that the customer has commented something negative about the power button on the remote.

  • Intent Analysis

This type of analysis helps to identify the intentions of a person. A good example is determining the intention of a customer whether they are genuinely interested in a product or they are just viewing the product.

On this project we will focus majorly on the fine grained type of analysis to achieve the end goal of the project which is data analysis on covid-19 data. This means that we will have the ability to classify data from social media into categories of either positive, negative or neutral to get the general sentiment of the public around this pandemic.

There has been a lot of growth in the computer science field with branches such as data science having made a huge stride. Data science is an interdisciplinary field that encompasses scientific methods, algorithms and systems to analyze both structured and instructed data with the aim to draw conclusions and insights which can be used in multiple fields. Data science has many fields but the areas of interest in this project are the data mining and statistical analysis and machine learning also known as cognitive computer developments.

Data mining is a process of extracting and discovering patterns in large data sets involving methods at the intersection of machine learning, statistics and database systems(“Data mining”, 2021). After collection of data from social media with platforms such a Twitter making it an easy and straightforward process to collect the data by use of their application interface endpoints which have been professionally built to provide bulk, publicly available information, and for this project specifically, publicly available tweets on the corona virus pandemic. 

Machine learning which is also viewed as artificial intelligence, is the study of computer algorithms that learn patterns from large datasets and from these patterns that have been discovered, the model can make sound and reliable predictions of completely new data(Mitchell, 1997). After the data collection and data preprocessing, machine learning models built in this project will be fed these data to train on and will then be used to classify different tweets into different categories of aspect based sentiment analysis.

1.2     Problem Statement
The novel coronavirus which has become a popular topic on the various social media platforms with people on different forums sharing bulk information on this particular topic. The problem arises when this data needs to be structured to produce meaningful results in the context of rumor spreading, memotion analysis, common sense reasoning and conversation recommender systems.

For the analysis of this type of data, data derived from social media particularly textual data, emotion analysis will be considered. Memotion analysis is a type of analysis where by textual data derived from various platforms such as twitter, Facebook and Instagram, is analyzed for sentiment that is either positive, negative or neutral. Another type of this analysis is type of emotion whereby this data can be classified into sad, happy, sarcastic, funny, offensive and motivation and the corresponding intensities.

With this type of analysis it is possible to solve this data analytic problem to provide an overall sentiment on the trending topic of the novel corona virus. The results produced from this analysis will be useful to general public, researchers, medical experts and above all governments to gauge the attitude of the country towards this pandemic. Governments can also know what the general public think about restrictions put in place to curb the spread of this virus, therefore understanding what the society needs and how to handle different regions based on the sentiments that are derived from social media from that particular region.

1.3     Aims and Objectives

The project aims to:

  • To investigate the covid-19 social media data and come up with the general sentiments around the issue.

This aim will be established by specific precise objectives which are:

  • Collection of publicly available data from various social media platforms such as Twitter, Facebook and Instagram.
  • Cleaning and preparation of the data to be analyzed by use of the natural language processing toolkit found in the python programming language.
  • This data will then be use to provide general sentiments around the corona virus topic globally by use various machine learning models.

1.4     Solution approach

This provides a brief summary on how the above said aim and objective will be met in the doing of this projects. It entails the steps that will be taken to ensure flawless completion of the project.

1.4.1   Data collection

The plan for this particular step is to collect data on social media platforms. Social media applications such as Facebook and Instagram have very strict policies that guide on how data from their sites should be used. The application programming interface provided by this site are very limited, that is they provide very little data and very little capabilities hence data collection from this sites prove to be very hard. Twitter on the other hand has a very intuitive API which helps developers to integrate twitter in their applications and allows researcher to do their research with ease.

This API provides access to a variety of different resources including the following:

  • Tweets
  • Users
  • Direct Messages
  • Lists
  • Trends
  • Media
  • Places

From the list above we will be concerning ourselves with accessing the tweets from any of the below hashtags or words:

  • #coronavirus,#coronavirusoutbreak,#Covid,#coronavirusPandemic, #covid19, #covid_19, “#epitwitter”, “#ihavecorona”, “#pandemic”, “#covid__19″,”covid”, “covid19”, “corona”,”coronavirusoutbreak”,”coronadeaths”, “coronavaccine”,”covid-19″, “COVID”,”#COVID19″,”#Corona”,”#CoronaSecondWave”,”#coronaindia”,”#covid_19″,”#CovidIndia”,”#CovidHelp”,”#Covid_19″,”#COVID”,”#COVIDEmergency2021″,”#Covid19usa”,”#CoronaUpdate”,”#CoronaSecondWave”,”#Lockdown”,”#Restrictions

We shall make a query to the twitter API requesting for any tweet that contains any of the above hashtags. Although one thing you should note is that twitter has different types of accounts that allow for different capabilities, the  accounts are as follows:

  • Standard accounts v1.1

This type of account only allows a predefined amount of requests to send to the API per month and these requests are 500, 000 requests.

  • Premium V1.1

Has the capabilities of a standard account but has no predefined amount of requests and more than a seven day look back.

  • Enterprise

This type of account has enterprise level products that allows access to twitters full

Archives, coupled with this is a 30 Day Search, Power Track, Engagement API and so much more.

For this project, the standard account was considered as the best option, the only disadvantage was that the standard API only provided data that up to seven days before, which was enough to create a comprehensive dataset that would be used in the later stages of the project.

1.4.2   Data Cleaning and preprocessing

Since tweets acquired from the data collection stage come with a lot of unwanted characters, emoji’s, numbers, common words and punctuations, the data need to be cleaned to get these unwanted characters out of the way in order to remain with words that make sense which can be used to build a well-organized and labelled datasets.

The data cleaning process entails of the following steps:

  • Transforming all tweets into lowercase characters, this although overlooked by many, ensures consistency with expected output. It also helps in normalizing all the text into a common letter case.
  • Removing all forms of punctuations such as commas, full stops, exclamation marks etc.
  • Removing stop words. These are words that are used to construct a sentence such as a, the, is, and, this, that, etc.
  •  Tokenization: this simply deconstructing a sentence and building a list of words from the deconstructed sentence an example is: “He is a boy” after tokenization the sentence is placed in a list of words like this [“he”, “is”, “a”, “boy”]. These words are known as tokens.
  • Normalization: This part of data processing is extremely overlooked and yet when done it can improve the results by a very reasonable margin. Normalization is the process of transforming text into is canonical or rather standard form (Kavita, 2019). An example: the word luv or looove can be transformed to love. This particular step is very useful when working with social media content since most users of these social media applications use short forms to express themselves. Since there is no standard way to carry out this particular task, some approaches include dictionary mappings, statistical machine translation (SMT) and spelling-correction based approaches
  • Stemming and lemmatization

Stemming is define as usually refers to a crude heuristic process that chops off the ends of words in the hope of achieving this goal correctly most of the time, and often includes the removal of derivational affixes(Herinch, 2009). This put in simple terms is the chopping o of the ending of a word an example of such is removing “ing” from the word raining to get rain. Stemming results in having a standardizing dataset and also deals with sparsity issues.

Lemmatization on the other hand involves to doing things properly with the use of a vocabulary and morphological analysis of words, normally aiming to remove inflectional endings only and to return the root form of a word (Aditya, 2020). This aids in the removal of common words that have similar meaning an example would be,

  • car, car’s, car, cars’ to produce the word car
  • Text enrichment

This involves the augmenting of your original text data with new information (Kavita, 2019). It aids in providing more semantics to the original text data and as a results helps to improve the predictive power of your algorithm. This is mostly seen in word embedding layers found in deep learning to improve semantics. It also makes machine learning models to form relationships between words that have similar meaning.

Steps listed here are just but a few examples of things that one could do to the data collected to clean it up and make it more efficient when it come to the building of a reliable dataset. It should be not that these process will be used when feeding the end result algorithm new unseen data from twitter.

1.4.3   Natural Language Processing toolkit

The Natural Language Toolkit (NLTK) is a platform used for building Python programs that work with human language data for applying in statistical natural language processing (NLP).

In this step, the use of this toolkit provided as a module in python will come into play showing its strength in processing to textual data and classifying the tweets.it will help in getting an overall overview of the data that is:

  • Getting the most popular words
  • Categorizing tweets as positive or negative or Neutral
  • Plotting sentiment graphs
  • Developing word clouds for purpose of visualizing the data in picture form.

NLTK has inbuilt algorithms to make text classification easier and more direct.  This project will employ this algorithms to obtain insights from the tweets collected.

1.4.4   Employing machine learning algorithms

The NLTK library provides good insights but for the purpose of this project we will dig in even further and build machine learning models by use of different machine learning algorithms. In Machine learning there are three types of machine algorithms namely:

  • Supervised machine learning algorithms

This is the use of well labelled datasets to train algorithms that will be used to classify or predict the outcomes correctly from the patterns it has learnt from the datasets. In hindsight all that happens is that inputs and outputs are given to the algorithm to identify patterns that map out a relationship between the input and the output this learnt pattern will be later be used to predict new data to the correct outcomes.

  • Unsupervised machine learning algorithms

In this type of machine learning algorithms the datasets only consist of the inputs and no output. The algorithm has to determine relationships from the given data and come up with a meaning representation of the data. These algorithms mostly include clustering, anomaly detection and deep learning models.

  • Reinforcement machine learning algorithms

This is an area of machine learning whose focus is on how intelligent agents should take actions in an environment in order to get rewards (Carrasco, 2020). In simple terms, it is a branch of machine learning where by agent placed in a environment learns through trial and error and based on the action this agent takes, if correct there is a reward and if wrong there is a penalty, just like teaching a dog a new trick.

For this project, supervised machine learning will be used, that is training models on labelled data. Multiple algorithms are used to in order to establish which algorithm fairs well than the others. A step further will be taken to see whether all this algorithms can be used to improve the overall accuracy. Examples of algorithms used in this classification are:

  • Naïve Bayes algorithm

Under this algorithm, there are various types of Naïve Bayes algorithms but the ones to be used in this project are:

  • Multinomial Naïve Bayes
  • Bernoulli Naïve Bayes
  • Linear Regression algorithms

Under this category, the algorithms used are:

  • Logistic Regression
  • Stochastic Gradient Descent  Classifier
  • Support Vector Machine

The algorithms used under this algorithm are:

  • Support Vector Classifier
  •  Linear  Support Vector Classifier
  •  Nu Support Vector Classifier

Through these classifiers we will determine tweets sentiments and hence achieving our overall aim.

1.5     Summary of contributions and achievements

Following these steps, the project aim which is to carry out sentiment analysis on Covid-19 data obtained from social media will be met, by the use of the state of the art machine learning algorithms. This will means we can actively gauge the attitude of people regarding this issue of the corona virus pandemic.

Chapter 2

2      Literature Review

The project is focused on memotion analysis specifically, sentiment analysis. Sentiment analysis also known as opinion mining is a field that has grown rapidly over the past few years with more demand and the benefits of this type of analysis being clearly seen. Big companies have also incorporated this type of analysis every day to gauge different business aspects affecting the business. An example would be, when a company launches a product they will do a sentiment analysis to see how the product is doing in the open market. This will give the company insight as to whether the market appreciates the product based on whether the sentiments are positive, neutral or negative.

Sentiment analysis is not only used in analyzing data in the business industry, it is also used in the finance industry specifically stock market and foreign exchange market to gauge the market sentiments an example would be, When a person of interest, such a Chief Executive Officer (CEO) of a company posts something negative about a his company’s stock, the sentiment on this would be positive and hence more investors will look to invest in the stock. In this way carrying out sentiment analysis on social media data has become essential to most individuals and companies.

In this project we will focus on the health sector which has become overwhelmed by the corona virus pandemic. Governments enforcing restrictions such as the full lockdown has forced many to move to social media to air their views on this corona virus pandemic which has created a large amount of data, in simple terms big data. Big data is a collection of large datasets pertaining to a particular topic and in this case, pertaining to Covid-19 data built up from the mass participation of the individuals via use of social media platforms. In this case, Twitter is the most used platform around the world. This particular social platform provides good data since participation is from all over the world and the use of a common language, that is, English.

This analysis is aimed to provide sentiments around the corona virus pandemic which in turn will help the health sector to understand how people feel about this virus. This can be said to be an overview or an overall picture of the situation at hand. Through these sentiments this sector can establish efficient plans to reduce the public panic and if positive will help the sector maintain hope among the general public, generally this analysis will help governments, individuals and institutions navigate through this time the world is facing this pandemic.

Since this virus came into the world abruptly there has not been an established system to this particular topic, making it an amazing to topic to explore and gain much more insight from, both to health experts and developers who gain knowledge on how to analyses textual data from social media platforms.

What leverages this project is the use of multiple types of algorithm to arrive to a particular classification, this highly improves accuracy. The use of one algorithm can result in somewhat bias results since the algorithm may tend to learn negative comments more than positive comments and hence it will tend to look for negativity in every tweet. But by the use multiple algorithms, whereby these algorithms provide their output and this output is compared against each other to get best of them and establish the certainty of this selected classification.

This feature is aimed to provide very accurate results to this covid-19 analysis topic and improve on the various previous analysis used to by other developers.

Chapter 3

3      Methodology

3.1     Task Description

The main task of the project is to analyze covid-19 data acquired from social media. This is the overall task, but to achieve this main task smaller organized task must be completed. This tasks as listed in the chapter 2 under solution approach are:

  • Social media Platforms

In this particular task, a social media platform will be selected. This platform is becomes the platform from where data will be obtained. The considered platforms are: Facebook, Instagram and Twitter.

  • Data collection

After careful selection of a social media platform, collecting data from this platform is what follows next, this simply making API calls to specific API endpoints to get the desired data.

  • Data Preprocessing

Data collected will not be in a form which is acceptable by the algorithms to be built, hence this particular task will clean up the data, getting rid of any unwanted data, and transforming it into the desired structure that will be fed to the machine learning algorithms.

  • Data analysis

This task is simply for taking all the structured data from the data preprocessing task and using it for its intended purpose that is use it for sentiment analysis and prediction.

3.2     Descriptions

3.2.1   Technology

Programming language used in this project is the python programming language. This is an interpreted high level programming language used for general purpose programming. The author of this language is known as Guido Van Rossum, his main aim was to put emphasis on code readability, simplicity and straightforward such that it can be easy to understand even for individuals who are new to the computer programming field.

Python is object oriented which means a software program can be structured into simple and reusable pieces of code called classes which are then used in the definition of that class instance called objects. This helps an individual to make containerized code snippets that perform certain tasks extremely well. Being a high level and an Object oriented language makes it a simple yet powerful language to work with.

 Python is used in various fields such as:

  • Web development
  • Data science
  • Software Prototypes
  • Artificial Intelligence

In this project we will use the Python 3.7.3 which at this time is the most current version.

Another good reason to use python is, it has a lot of useful built in modules that will be use to this project. External modules can easily be fetched by simple commands typed in the command line interface. Overall it is a very well maintained language with a large community of developers.

The Natural Language Processing Toolkit (NLTK) is used for the natural language processing which is a subfield of computer science that is very much concerned with how computer interprets human language. The natural language toolkit is composed of many libraries meant to make the process of transforming the human language into computer understandable language simpler. It can also be used to run statistics on human data.

 Examples of Natural Language Processing Toolkit is capable of:

  • NLTK tokenization

This aimed at forming smaller words from a larger sentences, this small pieces are known as tokens. This task is a very common task but fundamental task in natural language processing.

  • Sentiment Intensity Analyzer

This a built into the NLTK library used for classification of data. It built into the Valence Aware Dictionary popularly known as VADER dictionary which is a lexicon and a rule-based dictionary that uses is commonly used with social media data

  • Stemming and Lemmatization

The toolkit contains tools to do the stemming and lemmatization, which makes it easier for the developer to clean up and preprocess the data.

These are just a few tools found within this vast toolkit.

Natural Language Processing Toolkit (NLTK) version is 3.4.5. is used in this project.


Scikit-learn is another library used in this project. Scikit-learn is a popular machine learning library in python used for machine learning purposes. It contains a compilation of different machine learning algorithms that are applied according to the need of the project or task being solved. It features a lot of machine learning algorithm categorized into:

  • Classification algorithms.
  • Clustering algorithms.
  • Regression algorithms.

Scikitlearn is used with other statistical libraries such as numpy and Scipy to provide accurate statistical data. The version of Scikitlearn used in this project is 0.21.2.

Numpy is another library use in this particular project. It is a popular library used for mathematical and statistical application. This particular library makes it easy for individuals to work with large multi-dimensional arrays and matrices coupled with extremely high level mathematical function to work with these arrays. The version used in this project is 1.19.5.

Pandas this is an extremely popular library in python that aims to make the analysis of fast, flexible and extremely easy to use. This library makes the manipulation of data very simple and straightforward. To be specific, the library offers data structures and operations for the manipulation of numerical tables and time series which is free software and publicly available. The version used in this project is 0.25.1.

Seaborn is a python library used in data visualization. It is based on matplotlib which is a comprehensive library for executing data visualization used for plotting both easy and complicated data. Seaborn although adds a modern interface feel to its plots, this can be said to be the major difference between the two libraries. The version used in this project is 0.9.0.

These are the major technologies used in this project. The rest although minor have a significant impact to this project.

3.3     Implementation

The execution of this project will follow the steps stated in Chapter one under solution approach. The steps can be simply represented as:

3.3.1   Data collection

The social media platform chosen for this project is twitter which provides the Twitter API for the purpose of helping developers and researchers extract data from twitter. In the language python, it provides a couple of libraries used to access the API with ease. This project will utilize the Tweepy python library to access the API and get tweets which entails specifics hashtags. The use hashtags will narrow down to tweets entailing the corona virus. To access the API, an individual must have a twitter account. This account will be useful in signing up as a developer, which helps in accessing the API since a normal account does not have the ability to access the API.

In the Twitter Dev Site all you need to do is create an application using your preferred account with the options of Standard, Premium or Enterprise. After the creation of the application, the required API Keys and authorization credentials will be provided.

From there, the creation of the code that will fetch the data is placed in a file named as DataCollection.py

import pandas as pd

import numpy as np

import time

import tweepy

import csv

import os

from os import path

from datetime import date, timedelta

import json

import config

Import the libraries needed in this particular project. Some of this libraries are inbuilt into Python. The Tweepy library is an external library, it is installed by use of Python Package manager, PIP. It is installed via the command line by use of the simple statement:

Pip install Tweepy

Imports are simply modules, packages, or libraries that an individual will need to do specific task in the application.

ACCESS_TOKEN = config.ACCESS_TOKEN ACCESS_SECRET = config.ACCESS_SECRET CONSUMER_KEY = config.CONSUMER_KEY CONSUMER_SECRET = config.CONSUMER_SECRET

The config file imported contains the developer credentials, when working with credentials they should not be directly included in the code for security reasons. This particular lines of code are just but creating variables where the credentials will be stored. These credential will then be used to authentic to the twitter API.

3.3.2   Authentication

def Auth():     auth = tweepy.auth.OAuthHandler(CONSUMER_KEY, CONSUMER_SECRET)     auth.set_access_token(ACCESS_TOKEN, ACCESS_SECRET)     api = tweepy.API(auth)     print(“[INFO] Authenticated”)       return api

In this project object oriented paradigms are heavily used and from the code snippet above, used for authentication purposes. The function Auth uses the tweepy library to access function API to authentic the user by use of the given access keys and tokens. The function returns an established connection to the API and from here requests can be made to the API. It should be noted how the Tweepy package makes the authentication simple and straightforward hence the ease to use.

3.3.3   Tweet Extraction

def extractTweet(tweets):     tweetList = []     for tweet in tweets:         tweet = [pd.to_datetime(tweet.created_at).strftime(“%Y-%m-%d”), tweet.id, tweet.id_str,tweet.full_text]         tweetList.append(tweet)     df = pd.DataFrame(tweetList)     df     df.columns = [‘created_at’, ‘id’, ‘id_str’, ‘full_text’]     return df

Function extractTweets is used to extract import details of a tweet. After a request is made to the twitter API it returns a response with JSON serialized data that contains a lot of fields. Some of these fields are not necessary to this project and hence will not be useful. This function gets the response and extract for of the most important field, which are:

  • Created at

This is when the tweet was created

  • Id

A unique identifier of the tweet

  • Id_str

This is simply the string version of the id

  • Full_text

This is the tweet itself that is has the details of what the user tweeted.

This information is then stored in a Pandas Data Frame with columns names  ‘created_at’,  ‘id’, ‘id_str’,  ‘full_text’. A pandas data frame is like an excel spreadsheet which is two dimensional, that is, it has rows and columns. These rows and columns can contain different data types. The data frame in this case is known as df.

3.3.4   Data collection main method

Other programming languages like java, C++ and C have the main method which is the method that is first execute when the program is run. It is like an entry point to your application. Python on the hand doses not have this method by default which means code can be ran outside then main method. This however does not mean that one cannot make user of the main method. In summary the Data collection main method is the entry point of this program. This means that all the code will be executed from the main method.

def main():       main_dir = os.getcwd()     print(main_dir)     hashtags = [“#coronavirus”, “#coronavirusoutbreak”,”#Covid”,”#coronavirusPandemic”, “#covid19”, “#covid_19”, “#epitwitter”, “#ihavecorona”, “#pandemic”, “#covid__19″,”covid”, “covid19”, “corona”,”coronavirusoutbreak”,”coronadeaths”, “coronavaccine”,”covid-19″, “COVID”,”#COVID19″,”#Corona”,”#CoronaSecondWave”,”#coronaindia”,”#covid_19″,”#CovidIndia”,”#CovidHelp”,”#Covid_19″,”#COVID”,”#COVIDEmergency2021″,”#Covid19usa”,”#CoronaUpdate”,”#CoronaSecondWave”,”#Lockdown”,”#Restrictions”]     filename = “twitterdata.csv”     maximumId = 99999999999999999999     print(len(hashtags))     tweetList = []     count = 0       api = Auth()

This is a snippet of the main method and in which there are different variable declaration and initialization such as:

  • Main_dir: this a variable where we are storing the current working directory in form of a string. The function os.getcwd() helps in accessing this directory
  • Hashtags: this is a list that includes all the hashtags that will be passed to the search parameter in the API request.
  • Filename: it is the name of the Comma Separated Values file commonly known as csv that we will save our information to.
  • tweetList: this just but an empty list declaration.
  • Count: is an integer value initialized to 0 that will be used as a counter in our while loop
while True:         try:             for hashtag in hashtags:                 getTweets = api.search(q=hashtag ,count=2000,max_id=str(maximumId – 1), tweet_mode=”extended”,lang=’en’)                 if not getTweets:                     print(“[INFO] No new tweets have been found”)                     break                 print(“[INFO] New Tweets Found”)                 #DailyDf = extractTweet(getTweets)                 for tweet in getTweets:                     tweetX = [pd.to_datetime(tweet.created_at).strftime(“%Y-%m-%d”), tweet.id, tweet.id_str,tweet.full_text]                     tweetList.append(tweetX)                     print(f”Length: {len(tweetList)} ID: {tweet.id}”)                   except tweepy.TweepError as e:             print(“[WARNING] Error Found: ” + str(e))             break           time.sleep(10)         count = count + 1         if count == len(hashtags):             break     Df = pd.DataFrame(tweetList)     Df     Df.columns = [‘created_at’, ‘id’, ‘id_str’, ‘full_text’]       with open(os.path.join(main_dir,’DataMined’,filename), ‘w’) as fp:         Df.to_csv(os.path.join(main_dir,’DataMined’,filename), index = None)         print(“[INFO] File Created and New data has been saved”) if __name__ == ‘__main__’:     main()

This is the rest of the code after variables initialization. It basically consists of a while loop which run continuously until all the hashtags in our list have be fed into the API request.

The API provides methods for searching tweets that are using text, language and other filters. The method that we will be using is:

Api.search()

Which accepts these parameters:

  • q: this is simply the filter string an individual wants to search for, in this case the list of hashtags.
  • Count: since the API uses pagination, count value will let it know how many tweets we want from these pages.
  • Max_id: this is the maximum id that we are going to search for.
  • Tweet-mode: This parameter is set to extended, which tells the API to return tweets which might be more than 140 characters long  and include @replies and media attachments
  • Lang: This filter tells the API to return only tweets that are in English.

After each query has been made the response returned is appended to the tweetList. When all the hashtags have been queried the while loop breaks and the list is then stored in an empty pandas data frame for better organization and structuring. This pandas data frame is then stored locally to the computer in a csv file called the twitterdata.csv.

This summarizes the data collection step. This script can be run weekly for better results since twitter allows only tweets for the past seven days to be accessed when using the standard account. This is to prevent collection of duplicate data and getting new data. Another factor to consider is storing all the tweets id and comparing to the tweets already stored, if there exists such a tweet it will not be saved to the csv file. This will allow for the script to run on a daily basis and not get any duplicate data. The only challenge this brings is that for a standard account the number of requests given will run out and as result have to wait for another month before making any more requests, hence the running of the script once a week.

In the folder structure of this project there is the main file named as the main.py. This file is where majority of tasks are completed and it has main four parts:

  • Data preprocessing
  • Data visualization
  • Sentiment analysis using Natural Language Processing Library
  • Sentiment analysis using Machine Learning algorithms
import pandas as pd
import numpy as np import io import random import string import emoji from TextBlob import TextBlob import re import json import os from collections import Counter import datetime as dt import nltk from nltk.corpus import stopwords from nltk.stem.wordnet import WordNetLemmatizer from nltk.stem.porter import PorterStemmer from nltk.sentiment.vader import SentimentIntensityAnalyzer from matplotlib import pyplot as plt from matplotlib import ticker import seaborn as sns from wordcloud import WordCloud from classifier import FinalSentiment

The import of the main.py file are quite a bunch since majority of tasks are being completed in this file. These imports are as follows:

3.4     Accessing the Data

#This function gets data from the CSV and Returns the tweets def GetData():        #Get the Tweets from the directory using the OS module        parent_dir = os.getcwd()        data_dir = “data” + “/” + “covid19_tweets.csv”             path = os.path.join(parent_dir, data_dir)          #Read the Data using the Pandas        dataFrame = pd.read_csv(path)        #The .info method shows a summary of the items        #in this case we are interested in the columns        #print(dataFrame.info())          #Select the text columns as thats where the tweet is found        mainDataFrame = dataFrame[‘text’]        #print(mainDataFrame.tail())          return mainDataFrame

The function GetData() is used to access the csv file saved after running the Datacollection.py which is responsible for getting data and saving it to a csv file.

Using the OS module which is built into python we load in the file from the system. First we get the current working directory using the os.getcwd () function and from the path we get from this function we append to it the directory where the CSV file is located. Hence the use of the function os.path.join (parent_dir, data_dir) this will return the full path to where the data is located.

Using this path we are going to load the into a pandas data frame using the pd.read_csv () which is used to read csv type of files. From the data collected there are four columns and for the data preprocessing step only the textual data will be processed, that is the tweets. The pandas library makes this simple, since it has the ability to select one column, in our code above we do this by using mainDataFrame = dataFrame[‘text’], this selects the text column and places it in a new Data Frame assigned the name mainDataFrame. This mainDataFrame is what is returned from this function, comprising of all the tweets in the CSV file.

3.5     Data preprocessing

def CleanTweets(Tweets):        #Tweets = Tweets.apply(lambda z: emoji_free_text(z))        #remove any URLS found in the Tweets        Tweets = Tweets.apply(lambda z: re.sub(r”https\S+”, “”, str(z)))        #Remove Hashtags        Tweets = Tweets.apply(lambda z: re.sub(r”#”,””,str(z)))        #Convert the tweets into lower case        Tweets = Tweets.apply(lambda z: z.lower())          #text lemmitazition        Tweets = TextLemmitization(Tweets)        #Remove any word with less characters than 2        Tweets = Tweets.apply(lambda z:’ ‘.join([word for word in z.split() if len(word) > 2]))          #Remove Punctuations        Tweets = Tweets.apply(lambda z: z.translate(str.maketrans(”, ”, string.punctuation)))          #Removing stopwords        stopWords = set(stopwords.words(‘english’))          stopWords.update([“#covid19”, “#covid19”, “#coronavirus”, “#covid_19”, “covid19″,”coronavirus”,”corona”,”covid”,”amp”])               Tweets = Tweets.apply(lambda z:’ ‘.join([name for name in z.split() if name not in stopWords]))          #Remove numbers        TweetsFinal = Tweets.apply(lambda z:’ ‘.join([number for number in z.split() if not number.isdigit()]))        Print (TweetsFinal)        return TweetsFinal

The code snippet above is a step by step process of data cleaning. To reiterate, the purpose of data cleaning is to remove unwanted characters from the data collected, especially tweets which contain a lot of these characters such as punctuations, short forms, stop words, URLS and other unwanted characters that would be of no use to the models that will be trained. If left these characters would be noise. In the function we are passing a data frame as a dependency.

3.5.1   Emoji Removal

In the first step of cleaning this data, emoji’s are being removed. This step is done by the line of code: Tweets = Tweets.apply (lambda z: emoji_free_text (z)) the apply function found in pandas is useful since it performs a custom operation for either the row or column, simply put, it can be compared to a for loop that loops across the entire data frame performing changes column wise or row wise. In this line of code we are applying a custom function to remove all the emoji’s. The code snippet below shows this custom function.

def emoji_free_text (text):     allchars = [str for str in text.decode(‘utf-8’)]     emoji_list = [c for c in allchars if c in emoji.UNICODE_EMOJI]     clean_text = ‘ ‘.join ([str for str in text.decode(‘utf-8’).split() if not any(i in str for i in emoji_list)])     return clean_text

3.5.2   URLs Removal

Tweets come with very many URLS attached to them and these URLS are not of any use when it comes to sentiment analysis. To remove these URLS the use or regular expression is involved, the re library in python makes this a simple task as shown below:

Tweets = Tweets.apply (lambda z: re.sub (r”https\S+”, “”, str(z)))

In this line of code the apply function performs a regex expression which simply removes any URL found in all the tweets.

3.5.3   Hashtag Removal

Something that is very popular in tweets are hashtags. Using a regex expression the hashtag will be removed from all the tweets found in the data frame.

Tweets = Tweets.apply (lambda z: re.sub (r”#”,””,str(z)))

3.5.4    Lowercasing tweets

This is a step which is commonly ignored but it is one that will help a lot in improving the results of the models since after lowering text words such Dad and dad become the same and hence allowing for stemming and lemmatization. Lowering of the tweets will also improve semantics since the machine will not give different meanings to same word due to different character casing. This step is implemented through:

Tweets = Tweets.apply (lambda z: z.lower())

The Python string method lower () returns a string in which all case-based characters have been lowercased. It ensure uniformity among all tweets.

3.5.5   Lemmatization

Lemmatization as define in chapter 1 under solution approach it is the process of grouping together the inflected forms of a word so they can be analysed as a single item, identified by the word’s lemma, or dictionary form. This is achieved by:

Tweets = TextLemmitization(Tweets)

This one line of code will carry out the process of lemmatization. This function is found in Natural language processing.

3.5.6   Removing stop         words

Stop words are words that most commonly occur in a language. The Natural Language Processing library has a dictionary of stop words found in the English language. To use this stop words, they have to be imported to the program and then used as follows:

       stopWords = set(stopwords.words(‘english’))

       stopWords.update([“#covid19”, “#covid19”, “#coronavirus”, “#covid_19”, “covid19″,”coronavirus”,”corona”,”covid”,”amp”])

       Tweets = Tweets.apply(lambda z:’ ‘.join([name for name in z.split() if name not in stopWords]))

In this code snippet we are adding our list of stop words that include the name covid and corona since the words are bound to be used a lot in this social media topic, they become stop words.

3.5.7   Summary

These are the major steps in cleaning the tweets to get the clean tweets.

3.6     Data visualization

3.6.1   Most Frequent words

This is the analysis of the most common words in our dataset. This will help gain a perspective of which words are most used in the topic. Both negative and positive words will be seen in this analysis.

To accomplish this step, our pandas data frame must be transformed to a list. This done using the function:

def TweetList(Tweets):        #putting the tweets into lists        tweetList = [tweet for line in Tweets for tweet in line.split()]        return tweetList

The python function using the split () splits the tweets into its respective words which are then appended to a list of words, it can be said as simple tokenization. The function TweetList (Tweets) accepts a pandas data frame and returns a list.

This list is then passed to the function Frequency () which returns a visualization of the most frequently used words. The structure of this function is as shown below:

#Function used visualize just how much the word has been used def Frequency(tweetList):        sns.set(style=”darkgrid”)        #using the Count module to see how often a particular word has been used        frequency = Counter(tweetList).most_common(20)        #Create an Empty DataFrame to store these Frequently used words        frequencyDataFrame = pd.DataFrame(frequency)         frequencyDataFrame        frequencyDataFrame.columns = [“word”,”Frequency”]          fig, ax = plt.subplots(figsize = (15, 15))        ax = sns.barplot(y=”word”, x=’Frequency’, ax = ax, data=frequencyDataFrame)        plt.savefig(‘FrequencyOfWord.png’)

Using the Counter function, the most common 20 words are searched for and the stored in the variable frequency. Frequency is a list consisting of tuples which is then append to a pandas data frame name frequencyDataFrame with two columns, the word and the frequency column.

Using the Seaborn visualization library these data is then plotted to a bar chart with is saved to the system as picture with the name FrequencyOfWord.png

3.6.2   Word clouds

This is a data visualization technique used to represent textual data, where the size of the word depends on the frequency of the word. This is a step further to show which words are frequently used in a graphical format. To implement this, the simple and easy to use word cloud library is used.

def CreateWordCloud(tweetList):        cloud = WordCloud(     background_color=’black’,     max_words=40,     max_font_size=40,     scale=5,     random_state=1,     collocations=False,     normalize_plurals=False        ).generate(‘ ‘.join(tweetList))        plt.figure(figsize = (12, 10), facecolor = None)        plt.imshow(cloud)        plt.axis(“off”)        plt.tight_layout(pad = 0)        plt.savefig(‘wordcloud.png’)

The function CreateWordCloud () accepts a list, which is fed to the function WordCloud () to generate the plot of words in varying in font size according to the frequency of that word.

3.7     Experiments and Designing

There are different algorithms used for this problem of sentiment analysis. In this experimentation phase, there will be use of three major techniques, which are:

  • Using the NLTK Valence Aware Dictionary for Sentiment Reasoning(VADER)
  • Using Text Blob
  • Using Machine Learning algorithms.

3.7.1   Method one: Valence Aware Dictionary for Sentiment Reasoning (VADER)

It is a model used for sentiment analysis which factors in both polarity, that is, if a tweet is positive, negative or neutral and also factors in intensity, that is, the strength of an emotion. VADER for sentiment analysis uses a dictionary that maps lexical features to the emotional intensities which are known as scores (Aditya, 2020). VADER has the benefit of understanding different language context an example would be the sentence, “did not hate you” this would be classified as positive since it does not only look at the word hate to classify the sentence. This brings us to the conclusion that VADER is a very efficient classifier. VADER returns a dictionary of scores categorized as:

  • Positive
  • Neutral
  • Negative
def SentimentAnalyzerMethodOne(tweets):        sia = SentimentIntensityAnalyzer()        scores = tweets.apply(lambda z: sia.polarity_scores(z))        scoreDataFrame = pd.DataFrame(list(scores))        scoreDataFrame        #Create a new column for classify the score based on the Polarity Value        scoreDataFrame[“Classified”] = scoreDataFrame[‘compound’].apply(lambda z: ‘Neutral’ if z == 0 else (‘Positive’ if z > 0 else ‘Negative’))        scoreDataFrame[“text”] = tweets        print(scoreDataFrame.head())        return scoreDataFrame
  • Compound

The code above will implement VADER. In order to use VADER, the class SentimentIntesityAnalyzer() is provided. The class is initialized as the sia object which now can access the methods found in that class. In this project, the method polarity_scores () which return the scores of each tweet is used.

The results from running this algorithm are then stored in

3.7.2   Method Two: Text Blob

TextBlob is a python Library aimed to make the language processing as simple as it can be. TextBlob has a variety of uses such as   text classification, parts of speech tagging, noun phrase extraction, translation, classification, sentiment analysis and more (Shubham, 2018). TextBlob in this project is considered as the second method in solving the problem of sentiment analysis.

The code snippet below show how this method is implemented:

def SentimentAnalyzerMethodTwo(tweets):        scores = tweets.apply(lambda z: TextBlob(z).sentiment)        #polarity = tweets.apply(lambda z: TextBlob(z).sentiment.polarity)        scoreDataFrame = pd.DataFrame(list(scores))        scoreDataFrame        scoreDataFrame.columns = [“subjectivity”, “polarity”]        #Create a new column for classify the score based on the Polarity Value        scoreDataFrame[“Classified”] = scoreDataFrame[‘polarity’].apply(lambda z: ‘Neutral’ if z == 0 else (‘Positive’ if z > 0 else ‘Negative’))        scoreDataFrame[“text”] = tweets        print(scoreDataFrame.head())          return scoreDataFrame

In the function SentimentAnalyzerMethodTwo() accepts a dataframe which consists of the tweets. The tweets are then passed through the function TextBlob() to acquire the sentiment classification. The function then returns a dataframe containing the polarity of the tweet, the subjectivity and the tweet itself.

3.7.3   Method Three: Machine Learning Algorithms

In this method, machine learning algorithms are used. The approach taken to solve this sentiment analysis is known as ensemble learning. This means model that combines multiple algorithms for the purpose of classification. Ensemble learning entails a couple of algorithms used for classification and the one with the best performance for that particular problem is selected and used for the classification. The Benefits of ensemble learning are:

  • Increase in robustness
  • Improvement in performance
  • More accurate predictions
  • Elimination of biasness and variance

This shows the benefits of using ensemble learning as opposed to the use of one algorithm. Which will most likely suffer from biasness and sometimes fail in giving an accurate prediction. This does not mean ensemble learning does not suffer from errors, it does, but it reduces possibility of these errors occurring hence the preference.

The ensemble model used in this project contains seven machine learning models namely:

  • Naïve Bayes Classifier
  • Multinomial Classifier
  • Bernoulli Classifier
  • Logistic classifier
  • Linear Classifier
  • Linear SVC Classifier
  • Nu SVC Classifier

3.7.3.1 Naïve Bayes Classifier 

This is a supervised machine learning algorithm that is probabilistic in nature. It is said to be probabilistic in nature since it classifies an object based on the probability of that object. It solely depends on the principle of Bayes’ Theorem.

Bayes Theorem

Also known as Bayes’ law, Bayes’ Rule or Bayes-price theorem depends on conditional probability(Paul, 2015), that is, it determines the probability of a hypothesis based on knowledge at hand or rather prior knowledge (James, 2003).

The formula for Bayes’ Theorem is given as follows:

Where:

P (A|B) is Posterior probability: probability of A from the observed event B

P (B|A) is Likelihood probability: Probability of the evidence given based on the probability of A being true

P (A) is prior probability: probability of an event occurring before observing the evidence.

P (B) is Marginal Probability: probability of the evidence.

From this formula above we can understand the working of the classifier and why  it is called naïve, majorly because it assumes the occurrence of an event is independent of the occurrence other events.

The Scikit library has made the integrating of this algorithm very easy and straight forward the code snippet below shows the implementation of the Naïve Bayes algorithm in this project:

def BayesClassifier(trainingData):        classifiers = nltk.NaiveBayesClassifier.train(trainingData)        saveBayesClassifier = open(“Naivebayes.pickle”, “wb”)        pickle.dump(classifiers,saveBayesClassifier)        saveBayesClassifier.close()        return classifiers

This function simply trains the Naive Bayes algorithm and Saves is as a pickle file. It then returns a trained model, which is ready to predict sentiments of a tweet.

Multinomial Classifier

It is an algorithm that uses the naïve Bayes theorem for classification. It works exactly as the naïve Bayes the only difference is the use of multiple variables in the classification. This can have some advantages since more variables are considered hence better probabilities making it a slightly better algorithm as compared to the plain Naïve Bayes algorithm. It is mostly used in the field of text analysis and prediction.

This algorithm is implement in this project as shown in the code snippet below:

def MultiNomialClassifier(trainingData):        classifiers = SklearnClassifier(MultinomialNB())        classifiers.train(trainingData)          multiNomialClassifier = open(“MultiNomialClassifier.pickle”, “wb”)        pickle.dump(classifiers,multiNomialClassifier)        multiNomialClassifier.close()          return classifiers

The function MultiNomialClassifier() accepts the training data which is well labelled. This data is used to train the Multinomial Naïve Bayes algorithm which is saved to a pickle file and then returned by the function to be used in prediction.

Bernoulli Classifier

This classifier also uses the Naive Bayes’ theorem. It is similar to the multinomial classifier, in that it uses the frequency of a word for predictors, the difference is the predictor variables are independent Boolean variables. This code snippet below shows the implementation of this algorithm:

def BernoulliClassifer(trainingData):        classifiers = SklearnClassifier(BernoulliNB())        classifiers.train(trainingData)        return classifiers

3.7.3.2 Linear Models

Logistic classifier

This is a regression model used to model the probability of a certain class of probability. It mostly used in binary classification. Hence would be a wonder pick for this type of project where we are establishing whether a tweet is positive or negative.

The code snippet below shows how this model is implemented in this project:

def LogisticRegressionClassifer(trainingData):        classifiers = SklearnClassifier(LogisticRegression())        classifiers.train(trainingData)          logisticRegressionClassifer = open(“LogisticRegressionClassifer.pickle”, “wb”)        pickle.dump(classifiers,logisticRegressionClassifer)        logisticRegressionClassifer.close()          return classifiers

The function LogisticRegressionClassifer() accepts training data, which is used to train the logistic Regression Model, after training the model is saved as a pickle file and the returned to make predictions.

3.7.3.3 Support Vector Machines

support-vector machines (SVMs, also support-vector networks) (vapnik, 1995) are very powerful and flexible machine learning algorithms used for both classification and regression. An SVM model is a representation of various classes in a hyperplane viewed in a multidimensional space. The hyperplane is made in a manner to reduce the any errors. The aim of a SVM is to divide the provided dataset into classes in order to realize the maximum marginal hyperplane.( David, 2020)

The main concepts in SVMs are:

  • Hyperplane: This is the decision space which is divided according to the different classes the objects belong to.
  • Support Vectors: These are data point that are closest to the hyperplane.
  • Margin: this is the gap between two lines on the closest data points.

In this project, two models will be used for the purpose of classification. These models are:

  • Linear SVC
  • NuSVC
Linear SVC

The objective of this model is to fit the data provided to it, and it returns a “best fit” hyperplane that categorizes the data. After getting this hyperplane, new data can be fed to the classifier to see how it will predict. This makes it very suitable for our use case. The code snippet below shows how to implement this model:

def LinearSVCClassifer(trainingData):        classifiers = SklearnClassifier(LinearSVC())        classifiers.train(trainingData)          linearSVCClassifer = open(“LinearSVCClassifer.pickle”, “wb”)        pickle.dump(classifiers,linearSVCClassifer)        linearSVCClassifer.close()          return classifiers

The scikitlearn Library has made the implementation very simple. From the function above the classifier is fed the training data and once finished it saves the model and returns it for predictions.

NuSVCClassifer

This particular model works the same as the support vector machine but it accepts different parameters and have different mathematical formulas. The parameter that is different from SVC is the nu parameter which is the upper bound on the fraction of training errors and a lower bound of the fraction of support vectors. This classifier is also capable of performing multi-class classification and is very accurate. The code snippet below will show how the nuSVClassifier is implemented in this project.

def NuSVCClassifer(trainingData):        classifiers = SklearnClassifier(NuSVC())        classifiers.train(trainingData)          nuSVCClassifer = open(“NuSVCClassifer.pickle”, “wb”)        pickle.dump(classifiers,nuSVCClassifer)        nuSVCClassifer.close()          return classifiers

The function NuSVCClassifier() accepts training data as parameter. This data is passed through the NuSVC classifier where it is trained and then saved as a pickle file. The function then returns the trained model for classification of new data.

3.7.4   Dashboard Implementation

At this stage, a dashboard is created whereby all the results are graphed and displayed. The use of python libraries such as Dash and plotly will be seen in action in this section. Dash particularly is a library built on top of plotly for visualization. It uses html, JavaScript and react js to implement the functionality.

The code snippets below shows how the graphs are plotted into the web browser to create the dashboard.

VADER = dcc.Graph(id=’VaderSentiments’,style={‘width’: ‘100%’, ‘height’: ’46vh’},figure={                            ‘data’: [                                  {‘x’: vaderCount[‘sentiment’], ‘y’: vaderCount[‘count’], ‘type’: ‘bar’, ‘name’: ‘VADER Sentiment Analyzer’},                            ],                            ‘layout’:{                            ‘plot_bgcolor’: ‘rgba(19,23,34)’,                            ‘paper_bgcolor’: ‘rgba(19,23,34)’,                            ‘title’:’VADER Sentiment Analysis’,                            }                            })

The snippet above is used to plot the results obtained from the VADER sentiment analysis. This is how to plot a graph in plotly given the values. All other graphs on the dashboard will be plotted using the same procedure. Dash makes the process simple and easy to implement. The word clouds however will use a similar but slightly different procedure to be displayed since they are inform of images.

The code snippets below show how the word clouds will be displayed:

             html.Div(className=”showcase-2″,children=[html.Div(                     children=[                     html.Div(className=”image__container”,children=[html.P(children=’Most Positive Words’),                     html.Img(className=’image’,src=’data:image/png;base64, {}’.format(encoded_positive_image.decode()))]),                                         html.Div(className=”image__container”,children=[html.P(children=’Most Negative Words’),                     html.Img(className=’image’,src=’data:image/png;base64, {}’.format(encoded_positive_image.decode()))]),                     html.Div(className=”image__container”,children= [html.P(children=’Most Frequent Words’),                     html.Img(className=’image’,src=’data:image/png;base64, {}’.format(encoded_positive_image.decode()))]                     ])

Images have to be encoded by the base64 into their binary representation and then passed to the html object Img where they are decoded and passed to the dash board. This ensures that the image maintains it original quality.

The dashboard therefore shows the overall summary of the project by displaying all the analysis on one page.

3.7.5   Summary

In this chapter, the various methods of implementation of the project have been discussed together with some code snippets to show exactly how each step was accomplished in code. Discussion about which methods have been followed to realize the aim of this project and clear and precise explanations have been provided. In the next chapter, results will be discussed to see exactly what this methodology chapter produced.

Chapter 4

4      Results

In this project, three methods were used to realize the aim of this project. These methods are:

  • Valence Aware Dictionary for Sentiment Reasoning
  • Python Library TextBlob
  • Ensemble Machine learning Algorithms

Results from this three methods will be discussed in this chapter and the most efficient will be established.

4.1     Valence Aware Dictionary for Sentiment Analysis (VADER)

VADER is a rule based model for sentiment analysis for social media text. From the name we can establish that VADER uses a dictionary of lexicons for classification. Lexicon is the vocabulary of a language. The words of the language are rated as positive or negative by allocating a specific number on a scale of -1 to 1, 0 becomes neutral.

VADER produced four types of results. This are Positive (pos), negative (neg), neutral (neu) and compound. Positive metric show how positive the statement is, the negative metric shows how the negative the statement is, the neutral metrics show the statement is neither positive nor negative and finally the compound metric which is scale at -1 to 1 show the overall measure of the statement and in our case the tweet. The Table below shows a random sample of tweets and their classification using this method:

NegativeNeutralPositiveCompoundClassifiedTweet
0.00.7580.2420.4939Positivesmelled scent hand sanitizers today someone past would think intoxicated that…
0.00.3850.6150.7906Positivepattyhajdu navdeepsbains one safe everyone safe commit ensure…
0.00.4260.5740.6597Positivepraying good health recovery chouhanshivraj covidposit
0.120.6840.1970.2263Positivehey yankees yankeespr mlb wouldnt made sense players pay respects

Out of the collected tweets which totaled to around 179,107, the overall sentiments from these were:

  • Tweets with Negative sentiments : 46,467
  • Tweets with Positive Sentiments: 69,921
  • Tweets with Neutral Sentiments: 62,720

This method therefore determine the overall sentiment regarding the covid-19 pandemic as positive. This means the people on the social media platform were hopeful during the corona virus season. This data when visualized will appear like the figure below:

4.2     Python TextBlob

TextBlob is a library meant for analysis of textual data. It is built on top of the NLTK library, it simply acts as a wrapper around the vast NLTK library. The library has many features but the major one used in this project is the classification feature. This is machine learning feature used for text classification.

TextBlob returns two values, these are polarity and objectivity. Polarity refers to whether the statement is positive or negative. It is measured by the scale -1 to 1 where -1 is negative, 0 is neutral and 1 is positive. The subjectivity refers to personal opinions, judgements and emotions. The results are arranged in that order with the text being in a column of its own.

SubjectivityPolarityClassifiedTweet
-0.250.25Negativesmelled scent hand sanitizers today someone past would think intoxicated that…
0.00.0Neutralhey yankees yankeespr mlb wouldnt made sense players pay respects
0.70.6000000001Positivepraying good health recovery chouhanshivraj covidposit
0.50.5Positivepattyhajdu navdeepsbains one safe everyone safe commit ensure…

The table above shows a sample of the tweets after being cleaned. If the polarity is equal to 0 the tweet is categorized as neutral, if the polarity is greater than 0 the tweet is categorized as positive and if the polarity is less than 0 the tweet is said to be negative. From a dataset consisting of 179,107 tweets, the overall sentiment analysis were as follows:

  • Tweets classified as positive: 69,289
  • Tweets classified as neutral: 81, 282
  • Tweets classified as negative: 28,537

This data when visualized will look like the figure below:

In this method the overall sentiment analysis are neutral, twitter users are neither positive nor negative regarding the covid19 pandemic.

4.3     Ensemble learning

The ensemble learning model which is trained on a dataset with positive and negative words classifies the tweets into two categories. These categories are either Positive or Negative. After training this model is fed the tweets collected to classify them into two categories. The results are as follows:

  • Tweets classified as positive: 73,556
  • Tweets classified as Negative: 105,552

The model is trained on single words for the purpose of generalizing this model. The figure below show a visual representation of the classification.

4.4 Dashboard

In order to get a clear analysis of the project, a dashboard is created for data visualization where all this data will be plotted and viewed so as to make comparisons with ease and to have organized work. From the dashboard, the results of the project will be viewed in an organized manner.

The results obtained from the project are displayed on the dashboard to make it organized and easier to analyze the results efficiently. From the figure above the data is organized in such a way that results from the different models used are displayed. There is also a comparison graph to see how the different models compared against each other.

Chapter 5

5.1   Discussion and Analysis

The main aim of this chapter is to discuss the results obtained from the previous chapter on results. To achieve the aim of this project, three methods have been put to use, these methods are: the use of the Valence Aware Dictionary (VADER), the use of python library TextBlob and Ensemble machine learning technique. These three methods have yielded significant results which will be review in this chapter.

5.1.1  Python TextBlob

This method is the simplest and fastest method of analysis used in this project. Simplest because the library TextBlob is just a wrapper around the NLTK library therefore abstracting a lot of functionalities. This abstraction makes very easy and straightforward to use the library. The results entail two fields the subjectivity and the polarity. From the polarity we calculate the sentiment. The results presented by this algorithm are as follows:

  • Tweets classified as positive: 69,289
  • Tweets classified as neutral: 81, 282
  • Tweets classified as negative: 28,537

From the above results, the immediate striking feature is the number of tweets classified as neutral. This means that most of the tweets analyzed by the model did not have significant words to classify them as either positive or negative, hence the high number of neutral tweets.

TextBlob Neutral Tweets

To further analyze this analysis a random sample of five tweets classified as neutral is taken to establish just how accurate this method is:

PolaritySubjectivitySentimentTweet
0.00.0Neutralrajasthan government today started plasma bank sawai man singh hospital jaipur treatment pa…
0.00.0Neutralsecond wave flandersback homework
0.00.0Neutralsouth africa update south africa july nicdsa moetitshidi whoafro…
0.00.0Neutraltamilnadu 25th july highest spike total cases chennai
0.00.0Neutralfinally oliver wack author security corruption latamoutlook partner controlrisks update…
0.00.0Neutral,sign “authorize rapid testing” i’ll deliver copy officials no…

The table above shows a sample size of five tweets taken, from the final results delivered by text blob. It should be noted that these tweets have been preprocessed hence most will not make any sense due to stripping of some words in order to remain with most essential ones. The random tweets above from a human perspective have been correctly classified since not one of them have a negative or positive word. This shows that the analysis done by this library is correct. For tweets classified as neutral.

TextBlob Positive Tweets

Taking a further look into the tweets classified as positive. A random set of five tweets will be chosen to establish whether they were classified correctly.

PolaritySubjectivitySentimentTweet
0.1250.66666Positivefirst comprehensive review wash analysis key ways wash help reduce transmission…
0.136363636363636350.45454545454545453Positivereleased two new podcast episodes week technology platforms used conduct telehealth visits c…
0.50.5Positivesonusood helped many doctors russia reach india arranging flights always…
0.1750.275Positivefull update posted lcd daily update august louisiana neworleans
0.136363636363636350.5Positivelarea della grande manchester sotto osservazione live news melbourne curfew worldwid…
0.199999999999999980.35555555555555557Positiverichlandcounty fair largely closed public exhibits livestock continue richlandsource…

From this random sample of tweets, the tweets classified as positive do not have any positivity words in them, objectivity though is quite high on most of them. From this random sample of tweets classified as positive, the model has not made the best of classification. Though under fine grained sentiment analysis we could establish that most of these tweets are just but slightly positive and positive. This is because the polarity values are close to 0.0 and are above 0.5 meaning that these tweets are not very positive.

TextBlob Negative Tweets

PolaritySubjectivitySentimentTweet
-1.01.0Negativeworst type spread covidー19 covidiots
-0.551.0Negativenigeria hapless unfortunate citizens taxed ever increasing tariffs things even without lettin…
-0.20.0Negativegovernorgreg’s actions killed still killing what’s word use people willfully intention…
-0.150.8666666666666667Negativethedailyedge seanhannity foxnews spread lies hate sexual violence murder foxnew
-0.51.0Negativesad reality hundreds cars line drive thru food drive florida
-0.250.45Negativeexpert explains wrong talk second wave

From this random sample of tweets, all of them are classified correctly as negative. This is because words such as worst, unfortunate, hate, sad and wrong are seen in this tweets. In normal human conversation this are words that would be normally used when conversing about negative things. It is therefore safe to consider that this model classified the negative tweets correctly based on the words that the tweets entailed.

The TextBlob method is therefore a significant method in the classification of social media data.

5.1.2  Valence Aware Dictionary (VADER)

Considered as a superior method to the python TextBlob, this second method of sentiment analysis comes from the NLTK library. It uses the Sentiment Intensity Analyzer to carry out sentiment analysis on social media data. The results obtained from this specific model are as follows:

  • Tweets with Negative sentiments : 46,467
  • Tweets with Positive Sentiments: 69,921
  • Tweets with Neutral Sentiments: 62,720

This means that for this method the overall sentiments about the covid-19 virus are positive. This method concludes that twitter uses are positive about the covid-19. To further understand these results random samples of tweets and their classifications are taken and studied to establish whether they have been classified correctly.

VADER Positive sentiments

CompoundClassificationTweet
0.7269  Positivesafe place visit guests said hotel meticulous applying hand sanitation als…
0.5994Positivecongratulations team royal adelaide hospital furthering development vaccine
0.4767Positiveinteresting skynews running report majority population except would exempt lo…
0.6597Positiverecovered interested learning plasma donation help others click…
0.4939Positivelockdown prompted greater community spirit involvement neighbourhood life findings th…

From these sample tweets, it can been seen that all the classified data is correct. Although some tweets do not have positive words, VADER understands the language semantics and comes up with the appropriate sentiment for that particular tweet. This establishes that VADER is indeed powerful when it comes to sentiment analysis problems.

VADER Negative Tweets

CompoundClassificationTweet
-0.4951Negativedeaths continue rise almost bad ever politicians businesses want…
-0.875Negativemunicipality reports one new fatality due compared yesterday new fatality located fa…
-0.8079Negative,“covid19 everyone’s fight covered major disasters nothing like fabehamonir vi…
-0.25Negativeleader scientific software provider discusses challenging obstacles pharma…
-0.3067Negativewithout doubt studentsfamilies clearly healthwellness factors outweighing risk ex…

From the sample above the tweets have been correctly classified. It should to be noted that that the compound value indicates how negative or positive the statement is. In this case the more negative the compound value the more negative the statement is and the less negative the compound is the less negative the statement. The table above shows a good example of how the compound value relates to the tweet. Since VADER also uses semantics in its classification, number of negative tweets have increased compared to text blob analysis.

VADER Neutral Tweets

In VADER analysis the number of neutral tweets have reduced significantly. The table below show five random neutral tweets and their classification.

CompoundClassificationTweet
0.0Neutralchange work general recruiting specifically via proactivetalent recruiting…
0.0Neutralfinancial situation changed need rethink budget check article…
0.0Neutralamong familiar faces orangecofls news conferences american sign language asl interpreters…
0.0Neutralbars close curfew starts cyrilramaphosa knows takes hour 3days get ho…
0.0Neutralprovincial department education says expecting grade learners return school in…

This sample of tweets shows just how much VADER is accurate, in that, in all these tweets there is none with a negative or positive word. It also does not contain any semantic meaning. It is then safe to conclude that these tweets have been classified correctly.

5.1.3  Ensemble Learning

This model classifies the tweets into two categories. The tweet is either positive or negative. The model is a combination of other model which are used to make the classification. The model returns the classification and the confidence of that classification. To get the best classification the model uses the confidence, if the confidence is below 0.6, that classification is discard meaning it is neutral tweet.

Positive tweets

ClassificationTweet
Poslets protect real numbers climbing fast continent lets
Possaturdayvibes current situation calls awareness less stress facilitate families our…
Posthink link stuff loony conspiracytheories think screenshot is…
Poscompanies must protect workforce’s physical mental health crisis use smart technology…
posoverview book daniel part steve gregg bible biblestudy prophecy god…

From this random sample of tweets, it is clear that this model is very far behind from the prior models, since it has some wrong classifications here. This becomes the poorest model compared to the VADER and TextBlob model.

Negative tweets

ClassificationTweet
Negjuly update tamilnadu discharge people tested actice cases chennai
Neglockdown witnessing economic angst dislocation drive discontent turmoil…
Neg  patient ends life hospital telangana tuesday xpresshyderabad
Negcase worried sending kids back school september backtoschool…
Neggoing allin remote work technical cultural changes arstechnica msservicesglobal homeoffice…

Same to the positive tweets, some of the negative tweets are wrongly classified form this small random sample of tweets. The model fails to generalize and hence classifies the wrong. For this project this becomes the poorest performance.

5.2   Significance of Findings

The results from these methods have given new insights to sentiment analysis. The findings from the different models have shown how each models differ from each other, which models are best to use, which models are reliable and which models are fast. The findings also show how different models works and how efficient the working of each model is.

The models findings show that most of social media users are neutral about the covid-19 pandemic. With some being positive about it. Like in any society there must be both groups the positive and the negative group. However, the difference between the negative and positive group is very negligible hence it can be said that social media users are in a neutral group regarding the covid-19 pandemic.

5.3   Limitations

Key limitations affecting the findings are lack of adequate data. This is mainly attributed to the use of a standard account which only allows a seven day look back. This limits the amount of data collected and a lot of duplication of tweets. When duplicate tweets are removed, the data remaining forms a small dataset. To improve the findings more data needs to be collected. Training of the ensemble model on better labelled data, this will perhaps improve the efficiency of the model. To improve the findings fine grained aspect analysis can be done, this will avoid the general sentiments and focus on finer details such as where the tweet is very positive, positive, neutral, negative and very negative.

5.4   Summary

In this chapter, results of the different models have been discussed and how they differ from each other. The most powerful models have been established while the weakest models have also been established. Sentiment analysis has been carried out on the data and the results have been discussed at length. Samples have been taken from the different models classifications and it has been established which models are reliable when it comes to sentiment analysis. Finally limitation and challenges that have faced the project have been outlined and discussed. This then marks the end of this chapter 5.

Chapter 6

6.1   Conclusion

Using data mining and machine learning techniques to analyze covid-19 data from social media is the problem this project was tackling. The aim of the project being to carry out memotion analysis on social media data to come up with the sentiments of these data .This has been achieved by the use of three methods namely VADER, Python TextBlob and Ensemble learning. Using these three techniques all tweets that were collected from twitter have been classified into their respective sentiments categories.

Data collection from twitter in form of tweets from the twitter API was successfully achieved by making use of the python library tweepy. After collection of this tweets and storing them in a CSV file, the tweets are preprocessed to produce clean tweets with no noise leaving the most significant words to be used for the purpose of sentiment analysis. This tweets are then classified into their respective sentiment categories by use of the different methods of analysis. This marks all the objectives as being accomplished using all the key methods.

Through this project the problem of analyzing social media data for sentiments has been solved by making it an easy task for an individual to easily know and understand what is the sentiment surrounding this particular topic of covid-19. General sentiments are provided to the governments, hospitals and institutions for them to know how to handle different situations in this pandemic that occur at different times. High Negativity in sentiments shows fear and panic while high positivity shows hope therefore from these metrics, individuals in charge will know how to give hope and maintain peace during these times when the world is suffering from this deadly virus.

The project has followed simple and easy to understand implementation process to achieve the main aim and the objectives. All the outcomes are as expected and saved in case there is need for any reference.

6.2   Future Works

Through this project many ideas have availed themselves. What is thought as a simple and straightforward project becomes a gateway to more ideas on how it can be grown to mega projects supporting various sectors in both the business industry and tech industry. This project with no doubt can be improved in many ways by implement unique ideas to it.

A unique idea is such as using advanced concepts in machine learning such as the use of artificial neural networks and deep neural networks.  These are a series of algorithms the mimic the working of a human brain by learning underlying relationships between data and coming up with a pattern with which they will follow to predict any new data that has not been seen before. They consists of nodes which are interconnected to form layers. When the layers exceed three layers, this now known as a deep neural network. Neural networks form complex mathematical calculations in order to achieve good accuracies these calculations are done during training. If well trained with a good dataset neural networks can predict or classify thing accurately. In our problem it can classify the tweets extremely accurately.

Another idea would be to introduce classification according to emotions, that is, classifying a tweet in terms of sadness, happiness, anxiety and other emotions. This together with sentiment analysis will give extremely informative results to the people in charge since they not only know the attitude of individuals but they also know what type of emotions are the individuals going through hence better handling of the situations.

The last idea would be to do a live twitter stream sentiment analyzer. This would work as follows, connecting to the twitter API and live streaming tweets of this particular hashtag, that is, #Covid19, and passing the tweets through the models that have been built and then displaying that data to a dash board comparing how the sentiments are changing from as tweets stream in. to further this idea most frequent words would be used and most active members listed in some sort of a dashboard.

These are just but some few ideas but there are many more ideas that can be built upon this project given time.

Chapter 7

7.1   Reflection

The journey to the completion of this project has been very educative in that I have learnt very many new things. Learning to work with the language Python, using Object Oriented Paradigms, using libraries such as NLTK and TextBlob. Understanding the working of each and how they have been abstracted for the purpose of making them easy to use even for every type of programmer across the board. This project has open my mind to how different machine learning models work and how to implement these algorithms. The underlying concept of each algorithm, the assumptions and the formulas used to calculate these the features coupled with this this project has opened my mind to a new world of what machine learning algorithms can accomplish and what this growing field in computer science has done.

The experience I have is both creative and educative. Although at times tiresome when trying to accomplish simple and easy task, the overall experience had been good and educative. As a matter of fact the experience gained from this project is far the very best from any other project. The excitement in discovering and learning new things about research, problem solving skills, that is identifying the problems and finding out ways to solve the problems. Different approach yield different results although some yield the same results. Some approaches are faster while others are slower in achieving a same task.

The only challenge I have faced in this project that I could not overcome is the data collection stage where only limited amount of data could be provided from the twitter API. But with a better account type this problem is solvable. An account such as the researchers account is very reliable only catch is that you have to be a researcher with the backing of the institution where you are doing the research for.

This project cannot be done any different way, it can only be improved upon with the implementation of new ideas to seek better results and faster build speeds because in anything computer related speed is of utmost importance.

My overall thoughts on the experience I have had from his project are splendid.

 References

Aditya, B. (2020, May 14). Stemming vs Lemmatization. Towards Data Science. https://towardsdatascience.com/stemming-vs-lemmatization-2daddabcb221

Aditya, B.(2020, May 27). Sentiment analysis using VADER. Towards Data Science. https://towardsdatascience.com/sentimental-analysis-using-vader-a3415fef7664

Data mining. (2021, April 17). In Wikipedia. https://en.wikipedia.org/wiki/Data_mining.

Frame, Paul (2015). Liberty’s Apostle. University of Wales Press.

Herinch, S. (2009, April 7). Stemming and Lemmatization    . Stemming and Lemmatization. https://nlp.stanford.edu/IR-book/html/htmledition/stemming-and-lemmatization-1.html

Hu, J.; Niu, H.; Carrasco, J.; Lennox, B.; Arvin, F. (2020, December 12). “Voronoi-Based Multi-Robot Autonomous Exploration in Unknown Environments via Deep Reinforcement Learning”. Ieeexplore. https://ieeexplore.ieee.org/abstract/document/9244647

Jinman Kim, Younhyun Jung, David Dagan Feng & Michael J. Fulham. (2020). Biomedical Information Technology (Second Edition). Academic Press

Joyce J. (2003, September 30). Bayes’ Theorem. The Stanford Encyclopedia of philosophy. https://plato.stanford.edu/archives/spr2019/index.html

Kavita, G. (2019, February 24). All you need to know about text preprocessing for NLP and Machine Learning. Towards Data Science. https://towardsdatascience.com/all-you-need-to-know-about-text-preprocessing-for-nlp-and-machine-learning-bc1c5765ff67

Kavita, G. (2019, March 7). Text processing for machine learning and NLP. Kavita-Ganesan. https://kavita-ganesan.com/text-preprocessing-tutorial/

Mitchell, Tom. (1997). Machine Learning. New York: McGraw Hill.

shubham, J. (2018, February 11). Natural language processing for beginners using text blob. Analytics Vidhya. https://www.analyticsvidhya.com/blog/2018/02/natural-language-processing-for-beginners-using-textblob/

Syce, D. (2020, March 2016). Number of tweets per day 2019. David Sayce . https://www.dsayce.com/social-media/tweets-day/

Vapnik, Vladimir N. (1995) Support Vector Machines [PDF] http://image.diku.dk/imagecanon/material/cortes_vapnik95.pdf

Still stressed from student homework?
Get quality assistance from academic writers!

Facing Problems With Your Teacher Certification Exam Study Guides, Help Is Here!