Translate

Showing posts with label Naive Bayes. Show all posts
Showing posts with label Naive Bayes. Show all posts

Sunday, October 24, 2021

Federalist Papers : Case for Naïve Bayes Text Classification

Alexander Hamilton, James Madison, and John Jay

The Federalist Papers is a collection of 85 articles and essays written by Alexander Hamilton, James Madison, and John Jay. 1787 after the UK was thrown out from US, many were in the view that 13 counties should rule independently. 

John Jay, James Madison, Alexander Hamilton wrote letters independently to pursue that the US should have a strong central government with the individual state government. Between 1787 - 1788 these papers were published under the pseudonym PUBLIS. While the authorship of 73 of The Federalist essays is fairly certain, the identities of those who wrote the twelve remaining essays are disputed by some scholars. In 1963 this dispute was fixed by Mosteller and Wallace using Bayesian Methods.

Let us see do a simple analysis of these papers by performing a analyse of the titles of these papers using the Orange Data Mining Tool. You can retrieve the sample files and the Orange workflow from dineshasanka/FederalistPapersOrangeDataMining (github.com)

Following is the Orange Data Mining workflow and let us go through important controls. 

After importing the CSV, text was preprocessed and word cloud was generated to identify the word distribution. 


Bags of Words are used to identify the keywords. Then six classifiers are used which are Neural Network, Naive Bayes, Decision Trees, Random Forest, SVM and AdaBoost. Following is the evaluation results and it shows that the Random Forest technique has the edge over the other techniques 

We can build the decision tree as shown below. 

Wednesday, January 20, 2021

Identifying Mask & Non-Mask Faces Using Orange

We have started a discussion of image processing techniques using Orange in a few blog posts previously. Let us look at another case that can be utilized in the Data Mining Tool Orange. In this time, we will look at more current problem, that is identifying the Mask & Non-Faces Using. 

As you know every prediction problem needs two solutions. First, it needs to build the model using the prediction techniques and then it needs to choose the higher accurate model and build the production application. 

Since this is a classification problem, we need a data set that is already classified. Following is the already classified images. 


PN: Today being the January 20th and the images of US president and Vice presidents images are also in the Non-Mask category. It is not a deliberate just a coincident. 

Let us build models from different classification techniques and find out what is the best technique.


In above, we have used five classification techniques, such as Naive Bayes, Random Forest, SVM, Neural Network, and Logistic Regression. Image Embedding is the special control available in Orange in order to perform the image analysis. We have included a Test and Score Control in order to verify the accuracy and other model parameters such as Precision, Recall and F1 measure etc. 


The Above results show that both Neural Network and Logistic Regression has 100% accuracy over the other techniques. 
Now let us move to the next step, which is the prediction part. 

Let us select some challenging images rather than selecting naive images for the prediction. 

The first image, yes the mouth is closed but with hands. The second image is very straight forward. Next two images are with a mask but with a transparent mask.



Let us look at the predictions. 



You will see that the image with the hand is correctly classified as the "Non-Mask" with the smiling image. Both the images with transparent images are also correctly classified as "Mask". This means that the prediction accuracy is 100%.

Friday, January 17, 2014

January 2014 Meet-up

Session #1

TITLE: Predictive Modeling with the Microsoft Naïve Bayes algorithm

Join this session where Dinesh showcases the capabilities of predictive analysis using the Microsoft Naïve Bayes algorithm. Naïve Bayes is a classification algorithm that ships with SQL Server Analysis Services and is used to mine for and predict outcomes based on selected parameters.

CATEGORY: Business Intelligence (Data Mining)

SPEAKER: Dinesh Asanka (MVP), Database Specialist (Pearson Lanka)
Linked.In | Blog | Facebook | @dineshasanka

40 minutes approx.

 

Session #2

TITLE: Writing Resilient T-SQL Code - Part II

Continuing from where he left off from November's session, Gogula will guide you through writing better T-SQL code that is more resilient to unexpected issues and common code failures. This session is based on the book Defensive Database Programming by Alex Kuznetsov.

    CATEGORY: Development

    SPEAKER: Gogula G. Aryalingam (MVP), Technical Architect (Navantis)
    Linked.In | @gogula | Blog

    40 minutes approx.

     

    Time & Location

    JANUARY 22, 2013 - 6:00 PM Onwards at MICROSOFT SRI LANKA

    11th Floor, DHPL Building, No. 42, Nawam Mawatha, Colombo 2, SRI LANKA

    An excellent opportunity to network and learn. Refreshments provided.

    Entrance FREE