2017-07-31 46 views
0

我試圖從特定條件的數據框中檢索特定數據,但它顯示空的數據框。我是數據科學的新手,嘗試學習數據科學。這是我的代碼。無法檢索框架中的數據

file = open('/home/jeet/files1/files/ch03/adult.data', 'r') 
def chr_int(a): 
    if a.isdigit(): return int(a) 
    else: return 0 

data = [] 
for line in file: 
    data1 = line.split(',') 
    if len(data1) == 15: 
     data.append([chr_int(data1[0]), data1[1], 
        chr_int(data1[2]), data1[3], 
        chr_int(data1[4]), data1[5], 
        data1[6], data1[7], data1[8], 
        data1[9], chr_int(data1[10]), 
        chr_int(data1[11]), 
        chr_int(data1[12]), 
        data1[13], data1[14]]) 

import pandas as pd 
df = pd.DataFrame(data) 
df.columns = ['age', 'type-employer', 'fnlwgt', 'education','education_num', 'marital','occupation', 'relationship','race','sex','capital_gain','capital_loss','hr_per_week','country','income'] 

ml = df[(df.sex == 'Male')] # here i retrive data who is male 
ml1 = df[(df.sex == 'Male') & (df.income == '>50K\n')] 
print(ml1.head()) # here i printing that data 
fm =df[(df.sex == 'Female')] 
fm1 = df [(df.sex == 'Female') & (df.income =='>50K\n')] 

輸出:

Empty DataFrame 
Columns: [age, type-employer, fnlwgt, education, education_num, marital, occupation, relationship, race, sex, capital_gain, capital_loss, hr_per_week, country, income] 
Index: [] 

有什麼錯的代碼。爲什麼數據框是空的。

+1

你是否起訴,'收入'欄中的值是字符串,幷包含'\ n'? –

+0

是的,他們是字符串。 –

+0

然後試試這個:print(df.income.unique())。打印的值是否有'\ n'? –

回答

0

如果你仔細檢查值,您可能會看到問題:

print(df.income.unique()) 
>>> [' <=50K\n' ' >50K\n'] 

有每個值前面的空格。所以值應該被處理,以擺脫這些空間,或代碼應該像這樣修改:

ml1 = df[(df.sex == 'Male') & (df.income == ' >50K\n')] 
fm1 = df [(df.sex == 'Female') & (df.income ==' <=50K\n')]