pandasで効率よくデータを取り込む
Amany Mahfouz
Instructor
pandas にはスプレッドシートの読み込み関数 read_excel() がありますimport pandas as pd # Excel ファイルを読み込む survey_data = pd.read_excel("fcc_survey.xlsx")# 先頭5行を表示 print(survey_data.head())
Age AttendedBootcamp ... SchoolMajor StudentDebtOwe
0 28.0 0.0 ... NaN 20000
1 22.0 0.0 ... NaN NaN
2 19.0 0.0 ... NaN NaN
3 26.0 0.0 ... Cinematography And Film 7000
4 20.0 0.0 ... NaN NaN
[5 rows x 98 columns]
read_excel() には read_csv() と共通の引数が多くありますnrows: 読み込む行数の上限skiprows: 先頭から/特定の行をスキップusecols: 列名、位置番号、または文字(例: "A:P")で列を指定
# メタデータ見出しをスキップして、列 W-AB と AR を読み込み survey_data = pd.read_excel("fcc_survey_with_headers.xlsx", skiprows=2, usecols="W:AB, AR")# データを表示 print(survey_data.head())
CommuteTime CountryCitizen ... EmploymentFieldOther EmploymentStatus Income
0 35.0 United States of America ... NaN Employed for wages 32000.0
1 90.0 United States of America ... NaN Employed for wages 15000.0
2 45.0 United States of America ... NaN Employed for wages 48000.0
3 45.0 United States of America ... NaN Employed for wages 43000.0
4 10.0 United States of America ... NaN Employed for wages 6000.0
[5 rows x 7 columns]
pandasで効率よくデータを取り込む