An analysis of Commencement speeches in the U.S.
Every year, colleges or universities invite noted speakers such as tech leaders, politicians, famous writers, influential people in academia or entertainment etc. to address the graduating class. Through natural language processing, this project analyzes these heartening and poignant speeches, identifying both their common traits and what makes them stand out.
Learn from successful leaders in different fields about giving an inspirational or motivational speech. Three aspects are discussed:
- WHAT? ➜ Speech topics and trend of topic over the years
- WHO? ➜ Distribution of topics compared among audiences from different locations
- HOW? ➜ Evoluation of sentiment throughout the speech compared among spkearers with different professions
From topic modeling based on all speeches, 8 topics are identified:
- Family & Friends
- Women's voice
- New generation & Nation
- Tech & Business
- Hardship
- Sportsmanship
- Arts & Science
- Dream
The percent stacked area chart below shows the trend of topic over the years. Each color represents a topic.
Observations: Family & Friends is always a popular topic in the past 2 decades. However, Women's voice only starts appearing in the late 2000s. In 2004, New generation & Nation has the largest portion. Tech & Business increases in 2005 but is then replaced by Hardship in 2008. The latter two may be related to the presidential election in 2004 or the global financial crisis in 2007-2008. However, the sample size (400+ speech transcripts) is rather small for drawing any solid conclusions.
A. Different audicne (based on region)
Observations: At first glance, West and East regions have similar topic distribution with West having slightly larger portions for Hardship and Nation. There are more talks about Sports (sportsmanship) and Arts & Science in the East. Interestingly, while West and East have comparable portion for Women's voice, it is nonexistent in Central region.
B. Different speaker (based on speaker profession)
Observations: For artists and athletes, the speech topics are less diversified. Artists talk a lot about arts and hardship and athletes talk mostly about sports (not surprising). While speakers in academia and politics both have a large portion with topic of New generation and nation, the context is a little different. The former talks about new generation in terms of knowledge, and the latter is about country and history. Speakers in the entertainment industry tend to share stories and advice (included in the topic of Family & Friends). Among all professions, more (percentage wise) speakers in the publishing industry (including writers, journalists, etc) address Women's voice. And close to half of the speakers in tech or business (47%) have Dream as their speech topic.
Above comparison among different speaker profession shows that speakers in entertainment and tech industry tend to remain positive throughout the speech. However, the evolution of sentiment in speeches from lawyers, doctors, or academic researchers shows a different pattern. They may start with a positive opening and ending, but become less positive or even negative in the middle of the speech.
A network of speakers is constructed based on the overlap of words used in their speeches. It shows the relation between the speeches/speakers from another perspective other than the conventional bag of words topic modeling.
- Transcripts:
- Speaker info and institute info:
Speech
- Year
- Transcript length
- Transcript bag of words
Speaker
- Profession
- Year born
- Age when giving the speech
Institute
- Latitude, longitude (decimal degrees)
- Region (East, Central, West of the U.S.)
numpy,pandasscikit-learngensim,nltk,textblobmatplotlib,seabornnetworkx,textnets
Model detailss can be found in link to Jupyter Notebooks:
-
- Non-Negstive Matrix Factorization:
sklearn.decomposition.NMF - with TF-IDF (Term Frequency - Inverse Document Frequency):
sklearn.feature_extraction.text.TfidfVectorizer
- Non-Negstive Matrix Factorization:
-
- Latent Dirichlet Allocation:
gensim.models.LdaModel - with CountVectorizer:
sklearn.feature_extraction.text.CountVectorizer
- Latent Dirichlet Allocation:
-
- Speech clustering
- Pipline:
sklearn.decomposition.TruncatedSVD,sklearn.cluster.KMeans
-
- Valence Aware Dictionary for Sentiment Reasoning:
nltk.sentiment.vader - Polarity scores: negative, neutral, positive, compound
- Valence Aware Dictionary for Sentiment Reasoning:
-
textblob.TextBlob- Polarity (negative ⟷ positive)
- Subjectivity (facts ⟷ opinions)
-
nrclex.NRCLex- NRC emotion dictionary: positive, joy, anticipation, trust, surprise, negative, fear, anger, sadness, disgust
- LDA topic modeling: pyLDAvis
- Network analysis: networkx, textnets




