--- title: "Getting started with zuhtml" output: rmarkdown::html_vignette vignette: > %\VignetteIndexEntry{Getting started with zuhtml} %\VignetteEngine{knitr::rmarkdown} %\VignetteEncoding{UTF-8} --- ```{r, include = FALSE} knitr::opts_chunk$set(collapse = TRUE, comment = "#>") ``` zuhtml turns real-world HTML into ordinary R values: character vectors, lists and data frames. It parses the way a browser does, so malformed markup is repaired rather than rejected. You give it a string, raw bytes, a file, a URL or a connection. ```{r} library(zuhtml) ``` ## A page to work with A small catalogue page, as it might have been saved from a site. Some markup is sloppy on purpose: unclosed `
  • `s, unquoted attributes, a stray end tag, and a product card without a price. ```{r} page <- ' Tea shop

    Free shipping over 30 EUR

    Sencha

    3.50 details

    Genmaicha

    details
    SizeGrams
    Small0100
    Large0250
    ' doc <- html_parse(page, base_url = "https://example.org/shop/") doc ``` `html_read()` does the same for a file. `base_url` is where the page came from; relative links are resolved against it. ## Selecting elements `html_elements()` finds every element that matches a CSS selector. `html_element()` finds the first match below *each* input node, and keeps a missing node where there is none. That is what keeps extracted columns aligned when some records lack a field: ```{r} cards <- html_elements(doc, ".product") cards products <- data.frame( name = html_text_clean(html_element(cards, ".name")), price = html_text_clean(html_element(cards, ".price")), url = html_url(html_element(cards, "a")) ) products ``` Genmaicha has no price, so it gets `NA` rather than shifting the column. ## Values `html_text_clean()` gives text as a reader wants it; `html_text()` gives it exactly as parsed. `html_attr()` reads attributes, and `html_serialize()` writes nodes back as HTML. ```{r} html_text_clean(html_element(doc, "title")) html_attr(html_elements(doc, "nav a"), "href") html_serialize(html_element(doc, "h2")) ``` ## Structures Links, lists and tables have their own extractors: ```{r} html_links(doc, absolute = TRUE) lapply(html_elements(doc, "nav ul"), html_list) html_tables(doc) ``` Table columns are character: `"0100"` keeps its leading zero. Convert types yourself when you know them, for example with `type.convert()`. ## What the parser repaired Real pages nearly always have markup errors, which the parser repairs. `html_problems()` lists them: ```{r} html_problems(doc) ``` ## Where next * `vignette("selectors")`: the supported CSS subset. * `vignette("tables-and-lists")`: how tables and lists are read. * `vignette("limits-and-encoding")`: resource limits, encodings, and what zuhtml does not do.