Bearbeta twittertext

Analysera data från sociala medier i R

Vivek Vijayaraghavan

Data Science Coach

Lektionsöversikt

  • Varför bearbeta tweettext?
  • Steg i bearbetning av tweettext
    • Ta bort redundant information
    • Konvertera text till ett korpus
    • Ta bort stoppord
Analysera data från sociala medier i R

Varför bearbeta tweettext?

  • Tweettext är ostrukturerad, brusig och obearbetad
  • Innehåller emojis, URL:er och siffror
  • Ren text krävs för analys och tillförlitliga resultat
Analysera data från sociala medier i R

Steg i textbearbetning

Steg 1: Ta bort redundant information

Analysera data från sociala medier i R

Steg i textbearbetning

Steg 2: Konvertera text till ett korpus

Analysera data från sociala medier i R

Steg i textbearbetning

Steg 3: Konvertera till gemener

Analysera data från sociala medier i R

Steg i textbearbetning

Steg 4: Ta bort stoppord

Analysera data från sociala medier i R

Extrahera tweettext

# Extract 1000 tweets on "Obesity" in English and exclude retweets
tweets_df <- search_tweets("Obesity", n = 1000, include_rts = F, lang = 'en')
# Extract the tweet texts and save it in a data frame
twt_txt <- tweets_df$text
Analysera data från sociala medier i R

Extrahera tweettext

head(twt_txt, 3)
[1] "@WeeaUwU for real, obesity should not be praised like it is in today's society" 

[2] "Great work by @DosingMatters in @AJHPOfficial on \"Vancomycin Vd estimation in 
adults with class III obesity\". As we continue to study/learn more about dosing in 
large body weight pts, we see that it's not a simple, one size, one level estimate 
that works https://t.co/KkYPqS6JzG" 

[3] "The Scottish Government have an ambition to halve childhood obesity by 2030. 
This means reducing obesity prevalence in 2-15yo children in Scotland to 7%. 
\n\n\U0001f449 In 2018, this figure was 16%\n\nFind out more in our latest blog: 
https://t.co/FWp56QWjQc https://t.co/XBK8Je7F1A"
Analysera data från sociala medier i R

Ta bort URL:er

# Remove URLs from the tweet text
library(qdapRegex)
twt_txt_url <- rm_twitter_url(twt_txt)
Analysera data från sociala medier i R

Ta bort URL:er

twt_txt_url[1:3]
[1] "@WeeaUwU for real, obesity should not be praised like it is in today's society"

[2] "Great work by @DosingMatters in @AJHPOfficial on \"Vancomycin Vd estimation in  adults 
with class III obesity\". As we continue to study/learn more about dosing in large body 
weight pts, we see that it's not a simple, one size, one level estimate that works"

[3] "The Scottish Government have an ambition to halve childhood obesity by 2030. 
This means reducing obesity prevalence in 2-15yo children in Scotland to 7%. 
\U0001f449In 2018, this figure was 16% Find out more in our latest blog:"  
Analysera data från sociala medier i R

Specialtecken, interpunktion & siffror

# Remove special characters, punctuation & numbers
twt_txt_chrs  <- gsub("[^A-Za-z]", " ", twt_txt_url)
Analysera data från sociala medier i R

Specialtecken, interpunktion & siffror

twt_txt_chrs[1:3]
[1] " WeeaUwU for real  obesity should not be praised like it is in today s society"

[2] "Great work by  DosingMatters in  AJHPOfficial on  Vancomycin Vd estimation in 
adults with class III obesity   As we continue to study learn more about dosing in 
large body weight pts  we see that it s not a simple  one size  one level estimate 
that works"

[3] "The Scottish Government have an ambition to halve childhood obesity by       This 
means reducing obesity prevalence in     yo children in Scotland to       In       this 
figure was     Find out more in our latest blog "   
Analysera data från sociala medier i R

Konvertera till textkarpus

# Convert to text corpus
library(tm)
twt_corpus <- twt_txt_chrs %>% 
                VectorSource() %>% 
                Corpus()
twt_corpus[[3]]$content
[1] "The Scottish Government have an ambition to halve childhood obesity by       
This means reducing obesity prevalence in     yo children in Scotland to       In  
     this figure was     Find out more in our latest blog "
Analysera data från sociala medier i R

Konvertera till gemener

  • Ett ord ska inte räknas som två olika ord enbart på grund av olika skiftläge
# Convert text corpus to lowercase
twt_corpus_lwr <- tm_map(twt_corpus, tolower) 
twt_corpus_lwr[[3]]$content
[1] "the scottish government have an ambition to halve childhood obesity by     this 
means reducing obesity prevalence in     yo children in scotland to       in    this 
figure was     find out more in our latest blog "
Analysera data från sociala medier i R

Vad är stoppord?

  • Stoppord är vanliga ord som a, an och but
# Common stop words in English
stopwords("english")

Standardstoppord på engelska

Analysera data från sociala medier i R

Ta bort stoppord

  • Stoppord behöver tas bort för att fokusera på de viktiga orden
# Remove stop words from corpus
twt_corpus_stpwd <- tm_map(twt_corpus_lwr, removeWords, stopwords("english"))
twt_corpus_stpwd[[3]]$content

[1] " scottish government   ambition  halve childhood obesity         means 
reducing obesity prevalence      yo children  scotland                figure      
find     latest blog "
Analysera data från sociala medier i R

Ta bort extra blanksteg

  • Ta bort extra blanksteg för att skapa ett rent korpus
# Remove additional spaces
twt_corpus_final <- tm_map(twt_corpus_stpwd, stripWhitespace)
twt_corpus_final[[3]]$content
[1] " scottish government ambition halve childhood obesity means reducing obesity 
prevalence yo children scotland figure find latest blog "
Analysera data från sociala medier i R

Nu kör vi en övning!

Analysera data från sociala medier i R

Preparing Video For Download...