HTML
Extract elements from strings containing html. Requires WrangleWorks Account.
Extract text and links from HTML elements. Requires WrangleWorks Account.
Parameters
| Name | Description | Accepted Values | Default | Required |
|---|---|---|---|---|
| I/O | ||||
input | Name or list of input columns. | string, integer, array | — | Yes |
output | Name or list of output columns. | string, array, null | null | No |
| Options | ||||
data_type | The type of data to extract. | string; one of:
| — | Yes |
| Formatting | ||||
output_format | Format of the extract output. | string, null; one of:
| null | No |
char | Character to use when output_format is concatenate. | string | ", " | No |
| Conditions | ||||
if | Condition that determines whether the wrangle runs as a whole. Recipe variables may be referenced with ${variable}. | string | — | No |
where | Filter rows before applying the wrangle using SQL-like criteria, such as column1 = 123 OR column2 = 'abc'. | string | — | No |
where_params | Values used with where for parameterized criteria. Uses SQLite placeholder syntax such as ? or :name. | array, object | — | No |
Examples
wrangles:
- extract.html:
input: HTML
output: Text
data_type: text
| HTML |
|---|
| ` |
| Text |
|---|
wrangles:
- extract.html:
input: HTML
output: Links
data_type: links
| HTML |
|---|
| ` |
| Links |
|---|
Access
| Requirement | Value |
|---|---|
| AI-powered | No |
| Requires WrangleWorks account | Yes |
| Requires subscription | No |
| Requires external API key | No |
Technical details
| Field | Value |
|---|---|
| Catalog ID | 28 |
| Catalog key | extract.html |
| Recipe key | extract.html |
| Catalog status | active |
| Lifecycle status | active |
| Recipe Writer eligible | Yes |
| Namespace | extract |
| Documentation group | extract |
| Aliases | None |
| Runtime symbol | wrangles.recipe_wrangles.extract.html |
| Legacy UUID | 728fc87a-a20d-4efa-833a-612e0b5eadc3 |
Sources