Presidential Primaries Speech Analysis
With the 2016 Presidential Debates in full swing (as of when I started this little project at least), I figured it would be interesting to take a look some of the words spoken by candidates Hillary Clinton and Donald Trump given how polarizing they have been as characters. This is my first foray into doing a personal project with a dataset that I found on Kaggle. Right off the bat, I knew I wanted to pull some statistics about each candidate's speech, some sentiment analysis and produce some visuals to help capture the information in a nutshell.
The Data
The dataset contains columns for Line, Speaker, Text, Date, Party, Location and URL. For the sake of this project, I was really only concerned with Line and Speaker although I'm sure some interesting patterns might present itself based on Location.
For this project, I really wanted to start exploring Pandas as a tool so I utilized it heavily along with the matplotlib and nltk libraries. The first thing I did since I was really only concerned with Trump and Clinton was clear out all audience and non-presidential candidates related text and split the initial DataFrame into one for Trump and one for Clinton. I also went ahead and cleared punctuation here. First things first: the low hanging fruit. I wanted to look into all the words that Trump and Clinton spoke as well as their frequencies.
Basic Analysis
From the initial scrape of the debate, Hillary spoke 55,706 words while Trump spoke 39,958 with Hillary speaking 4,263 unique word and Trump speaking only 2,754 unique words. Huh. It seems Hillary has a vocabulary that is roughly 50% larger than Trump's (from the debates at least). I realize that a good chunk of these words are probably those akin to "I", "we", "they", "are", "am", etc. I looked for some ways to identify these words and quickly learned that these common words are referred to as stopwords and that the nltk library contained a collection of stopwords that I could use to run my dataset against. Now, after removing all stopwords in my dataset that matched the stopwords from nltk.corpus, I found that Hillary spoke 25,999 non-stopwords and Trump spoke 19,029 and Hillary spoke 4,141 unique words while Trump spoke 2,636. As expected, stopwords made up for a large chunk (roughly 53%) of each candidate's speech simply by virtue of them being constructs needed in proper grammar.
Let's Talk About Feelings
There's been a lot of talk about Trump's Doomsday rhetoric. Opponents allege that he paints a grim picture of America that is purposefully striking fear and doubt into voters. However, is this really true? For sentiment analysis, I used lists mined from Twitter sentiment analysis to identify positive and negative sentiments. One interesting thing that I noticed was that the word 'trump' was actually present in the positive word list so for obvious reasons, I omitted that word as a positive word to match. What I found was the following:
Hillary's positive word rate: 8.27% Hillary's negative word rate: 4.22% Trump's positive word rate: 7.87% Trump's negative word rate: 5.29%
Now this is an interesting point to talk about. It appears that the proportion of Trump's corpus that is associated with negative sentiments isn't nearly as large as some people may have felt. One theory is that the combined effect of having a lower positive rate and higher negative rate may make it seem like Trump has a more negative rhetoric. Another interesting point is that I did not measure the intensity of polarity of a word. As one could imagine a word like "abysmal" has a much larger negative connotation than "bad." There are datasets out in the public that maintain a list of words and their polarity so that may be a future direction to take with this analysis.
Clinton,min10,max500 Trump,min10,max500
The figures above plot the positively charged words (green) and negatively charged words (red) along with their frequency. I only plotted words that appeared at least 10 times (otherwise the bars were a total mess). Unsurprisingly, words like "well" and "great" occur quite frequently.
Clinton,min10,max75 Trump,min10,max75
The above graphs removed any words occurring more than 75 times in an effort to remove very commonly used positive words. What is interesting here is the distribution of positive words and what those words actually are. Trump's positive words are decently skewed towards his top few words and all his positive words are generally vague words such as "great", "good", "better", "love" and "thank." On the other hand, Hillary's distribution is a bit more spread out (once we factor out the word "well" at least. What really caught my attention is that her words lend themselves to more concrete and actionable concepts such as "affordable" and "comprehensive" (presumably referring to healthcare) as well as "reform" and "protect."
On the other hand, looking at the negatively-oriented words is equally interesting. In this case, Hillary's distribution is skewed heavily towards a few words like "problems", "issues", "hard" but still contains words that imply concepts and ideas that can be acted upon such as "terrorism", "racism", "criminal" and "debt." For Trump, it's apparent that he spoke a much greater proportion of negative words (at least for the thresholds I set) just based on where the red and green bars break. What is worth noting is that the negative words he spoke feel more polarizing; for example: "bomb", "killed", "terrible", "worst", "losing" and "disaster." Also worth noting is that very few to none of his words have any association with concepts that could be actionable unlike Hillary's.
Final Impressions
For my first real dive into some text analytics, I thought that this dataset was quite revealing because it corroborated how many people feel about each candidate's rhetoric. Trump truly did have a greater proportion of negative words to the size of his overall corpus although it might not have been as much as people suspected. However, the strength and polarity of his words may have made it seem like his speech was much more negative overall. Source code can be found here: Presidential Primaries Analysis
For anyone still reading, I'd love some feedback on any flawed methodology or future analyses that might prove insightful. One thing I'm thinking of weighting the positive/negative words via tf-idf (for those unfamiliar: term frequency–inverse document frequency which weights words proportional to their frequency in a document and inversely proportional to their frequency in the entire corpus). Until next time, thanks for reading!







Comments