我使用BeautifulSoup創建刮板,並請求刮擦網站的頁面以獲取匹配時間表(以及結果,如果可用)。這是我到目前爲止有:從網站格式化刮取的數據(BeautifulSoup)
def getMatches(self):
url = 'http://icc-cricket.yahoo.net/match_zone/series/fixtures.php?seriesCode=ENG_WI_2012' # change seriesCode in URL for different series.
page = requests.get(url)
page_content = page.content
soup = BeautifulSoup(page_content)
result = soup.find('div', attrs={'class':'bElementBox'})
tags = result.findChildren('tr')
for elem in tags:
x = elem.getText()
print x
而這些結果我得到:
Date & Time (GMT)fixture
Thu, May 17, 2012 10:00 AMEngland vs West Indies
3rd TESTA full scorecard will be available shortly.Venue: Edgbaston, BirminghamResult: England won by 5 wickets
Fri, May 25, 2012 11:00 AMEngland vs West Indies
2nd TESTClick here for the full scorecardVenue: Trent Bridge, NottinghamResult: England won by 9 wickets
Thu, Jun 7, 2012 10:00 AMEngland vs West Indies
1st TESTClick here for the full scorecardVenue: Lord'sResult: Match Drawn
Sat, Jun 16, 2012 9:45 AMEngland vs West Indies
1st ODIClick here for the full scorecardVenue: The Rose Bowl, SouthamptonResult: England won by 114 runs (D/L Method)
Tue, Jun 19, 2012 9:45 AMEngland vs West Indies
2nd ODIVenue: KIA Oval
Fri, Jun 22, 2012 9:45 AMEngland vs West Indies
3rd ODIVenue: Headingley Carnegie
Sun, Jun 24, 2012 12:00 AMEngland vs West Indies
1st T20Venue: Trent Bridge, Nottingham
現在,我想在一些結構化的格式對數據進行分類。一個包含
關於一場比賽的信息列表將是理想的。但我堅持如何實現這一目標。結果中的輸出字符串具有像 
這樣的字符,並且時間奇怪地排列,如AMEngland
。還有一個問題是,如果我用空格字符作爲分隔符來分割字符串,像西印度羣島這樣的國家將會被分割,並且將不會有任何統一的方式來解析它。
那麼有沒有一種方法可以統一解析這些數據,所以我可以在表單中找到。有點像:
[ {'date': match_date, 'home_team': team1, 'away_team': team2, 'venue': venue},{ same for match 2}, { match 3 }...]
我會感謝任何幫助。 :)
非常感謝。我想整天看HTML會讓我有點忘記我只能用一個簡單的正則表達式。 :) –