텍스트 데이터 정리

R로 데이터 정리하기

Maggie Matsui

Content Developer @ DataCamp

텍스트 데이터란?

데이터 유형 예시 값
이름 "Veronica Hopkins", "Josiah", ...
전화번호 "6171679912", "(868) 949-4489", ...
이메일 "[email protected]", "[email protected]", ...
비밀번호 "JZY46TVG8SM", "iamjosiah21", ...
댓글/리뷰 "great service!", "This product broke after 2 days", ...
R로 데이터 정리하기

비정형 데이터의 문제

  • 형식 불일치
    • "6171679912" vs. "(868) 949-4489"
    • "9239 5849 3712 0039" vs. "4490459957881031"
  • 정보 불일치
    • +1 617-167-9912 vs. 617-167-9912
    • "Veronica Hopkins" vs. "Josiah"
  • 유효하지 않은 데이터
    • 전화번호 "0492"는 너무 짧음
    • 우편번호 "19888"는 존재하지 않음
R로 데이터 정리하기

고객 데이터

customers
# A tibble: 99 x 3
   name            company                     credit_card        
   <chr>           <chr>                       <chr>              
 1 Galena          In Magna Associates         5171 5854 8986 1916
 2 MacKenzie       Iaculis Ltd                 5128-5078-8008-5824
 3 Megan Acosta    Semper LLC                  5502 4529 0732 1744
 4 Phoebe Delacruz Sit Amet Nulla Limited      5419-7308-7424-0944
 5 Jessica         Pellentesque Sed Ltd        5419 2949 5508 9530
# ... with 95 more rows
R로 데이터 정리하기

하이픈 있는 카드 번호 감지

str_detect(customers$credit_card, "-")
FALSE TRUE FALSE TRUE FALSE TRUE TRUE FALSE FALSE TRUE TRUE TRUE FALSE ...
customers %>%
  filter(str_detect(credit_card, "-"))
   name            company                  credit_card        
 1 MacKenzie       Iaculis Ltd              5128-5078-8008-5824
 2 Phoebe Delacruz Sit Amet Nulla Limited   5419-7308-7424-0944
 3 Abel            Lorem PC                 5211-6023-0805-0217 
 ...
R로 데이터 정리하기

하이픈 치환

customers %>%

mutate(credit_card_spaces = str_replace_all(credit_card, "-", " "))
   name            company                     credit_card_spaces        
 1 Galena          In Magna Associates         5171 5854 8986 1916
 2 MacKenzie       Iaculis Ltd                 5128 5078 8008 5824
 3 Megan Acosta    Semper LLC                  5502 4529 0732 1744
 4 Phoebe Delacruz Sit Amet Nulla Limited      5419 7308 7424 0944
 5 Jessica         Pellentesque Sed Ltd        5419 2949 5508 9530
 ...
R로 데이터 정리하기

하이픈과 공백 제거

credit_card_clean <- customers$credit_card %>%

str_remove_all("-") %>% str_remove_all(" ")
customers %>% mutate(credit_card = credit_card_clean)
   name            company                credit_card     
 1 Galena          In Magna Associates    5171585489861916
 2 MacKenzie       Iaculis Ltd            5128507880085824
 3 Megan Acosta    Semper LLC             5502452907321744
 ...
R로 데이터 정리하기

유효하지 않은 카드 찾기

str_length(customers$credit_card)
16 16 16 16 16 16 16 16 16 16 16 16 12 16 16 16 16 16 16 16 16 16 16 16 16 ...
customers %>%
  filter(str_length(credit_card) != 16)
  name            company                   credit_card   
1 Jerry Russell   Sed Eu Company            516294099537
2 Ivor Christian  Ut Tincidunt Incorporated 544571330015
3 Francesca Drake Etiam Consulting          517394144089
R로 데이터 정리하기

유효하지 않은 카드 제거

customers %>%
  filter(str_length(credit_card) == 16)
   name            company                     credit_card     
 1 Galena          In Magna Associates         5171585489861916
 2 MacKenzie       Iaculis Ltd                 5128507880085824
 3 Megan Acosta    Semper LLC                  5502452907321744
 4 Phoebe Delacruz Sit Amet Nulla Limited      5419730874240944
 5 Jessica         Pellentesque Sed Ltd        5419294955089530
...
R로 데이터 정리하기

복잡한 텍스트 문제 더 다루기

  • 정규 표현식은 문자열 내에서 강력한 검색을 가능하게 하는 문자 시퀀스입니다.
  • 정규 표현식에서는 일부 문자가 특별하게 취급됩니다:
    • (, ), [, ], $, ., +, *
  • stringr 함수는 정규 표현식을 사용합니다
  • 이러한 문자를 그대로 검색하려면 fixed()를 사용합니다:
    • str_detect(column, fixed("$"))

 

 

자세히 보기: R의 stringr로 문자열 조작 & R의 중급 정규 표현식

R로 데이터 정리하기

연습해 봅시다!

R로 데이터 정리하기

Preparing Video For Download...