2017-03-16 88 views
0

我在獲取HTML表格的值時遇到問題,因爲它沒有id s。我需要獲取第二列的所有值並將它們保存到一個數組中。我使用HtmlAgilityPack和我的問題自帶選擇節點時:如何使用HtmlAgilityPack解析無HTML標識的表格

Dim doc As HtmlDocument 
Dim web As New HtmlWeb() 
Dim str As String 

doc = Web.Load("http://www.dietas.net/tablas-y-calculadoras/tabla-de-composicion-nutricional-de-los-alimentos/carnes-y-derivados/aves/pechuga-de-pollo.html#") 

Dim nodes_filas As HtmlNode() = doc.DocumentNode.SelectNodes("//table[@id='']//tr").ToArray 
Dim nodes_columnas As HtmlNode() = doc.DocumentNode.SelectNodes("//td").ToArray 

For Each row As HtmlNode In nodes_filas 
    For Each column As HtmlNode In nodes_columnas 
     str = column.InnerHtml & vbCrLf 
    Next 
Next 

這是表:

<table cellspacing="1" cellpadding="3" width="100%" border="0"> 
    <tr> 
    <td colspan="2" style="font-size:13px;color:#55711C;padding-bottom:5px;">Aporte por raci&oacute;n</td> 
    </tr> 
    <tr style="background-color:#EBEBEB"> 
    <td width="125">Energ&iacute;a [Kcal]</td> 
    <td class="td_right">145,00</td> 
    </tr> 
    <tr> 
    <td>Prote&iacute;na [g]</td> 
    <td class="td_right">22,20</td> 
    </tr> 
    <tr style="background-color:#EBEBEB"> 
    <td>Hidratos carbono [g]</td> 
    <td class="td_right">0,00</td> 
    </tr> 
    <tr> 
    <td>Fibra [g]</td> 
    <td class="td_right">0,00</td> 
    </tr> 
    <tr style="background-color:#EBEBEB"> 
    <td>Grasa total [g]</td> 
    <td class="td_right">6,20</td> 
    </tr> 
    <tr> 
    <td>AGS [g]</td> 
    <td class="td_right">1,91</td> 
    </tr> 
    <tr style="background-color:#EBEBEB"> 
    <td>AGM [g]</td> 
    <td class="td_right">1,92</td> 
    </tr> 
    <tr> 
    <td>AGP [g]</td> 
    <td class="td_right">1,52</td> 
    </tr> 
    <tr style="background-color:#EBEBEB"> 
    <td>AGP /AGS</td> 
    <td class="td_right">0,79</td> 
    </tr> 
    <tr> 
    <td>(AGP + AGM)/AGS</td> 
    <td class="td_right"> 1,80</td> 
    </tr> 
    <tr style="background-color:#EBEBEB"> 
    <td>Colesterol [mg]</td> 
    <td class="td_right">62,00</td> 
    </tr> 
    <tr> 
    <td>Alcohol [g]</td> 
    <td class="td_right">0,00</td> 
    </tr> 
    <tr style="background-color:#EBEBEB"> 
    <td>Agua [g]</td> 
    <td class="td_right">71,60</td> 
    </tr> 
</table> 

回答

0

對不起,我沒有安裝VB,但C#版本應該足以給你一個主意。您有td_right類,您可以使用lambda或xpath來查詢它。 我更喜歡lambda/linq版本,因爲我熟悉linq,並且我不需要記住XPATH語法。

LAMBDA:

public static bool HasClass(this HtmlNode node, params string[] classValueArray) 
    { 
     var classValue = node.GetAttributeValue("class", ""); 
     var classValues = classValue.Split(' '); 
     return classValueArray.All(c => classValues.Contains(c)); 
    } 

var url = "http://www.dietas.net/tablas-y-calculadoras/tabla-de-composicion-nutricional-de-los-alimentos/carnes-y-derivados/aves/pechuga-de-pollo.html#"; 
     var htmlWeb = new HtmlWeb(); 
     var htmlDoc = htmlWeb.Load(url); 
     var nodes = htmlDoc.DocumentNode.Descendants("td").Where(_ => _.HasClass("td_right")).Select(_ => _.InnerText); 

XPATH:

var nodes2 = htmlDoc.DocumentNode.SelectNodes("//td[@class='td_right']"); 
+0

完美解決!!非常感謝! –