Retrieve Link Content |search.retrieve_link_content
Retrieves targeted content from web pages using LLM URL extraction. Can optionally output a second column containing a clean, human-readable text summary of the retrieved data.
Name or list of input columns containing URLs or Scored Search Result dictionaries.
string, array
—
Yes
output
Name of the output column for the raw dictionaries. To output BOTH the raw dictionaries and the formatted text, provide a list of exactly two column names (e.g., [page_data, page_text]).
string, array, null
null
No
Options
prompt
Optional custom system prompt to guide the extraction behavior and output format.
string, null
null
No
Formatting
output_format
The desired format for the extracted content.
string; one of:
markdown
json
"json"
No
Execution
threads
Number of concurrent threads for parallel processing (default 10).
integer
10
No
Details
client
The retrieval provider to use.
string; one of:
google_url_context
"google_url_context"
No
api_key
API key for the provider. Can also be set as an environment variable (e.g., GOOGLE_API_KEY).
string, null
null
No
model_id
The specific model ID to use (default models/gemini-3-flash-preview).
string
"models/gemini-3-flash-preview"
No
Conditions
if
Condition that determines whether the wrangle runs as a whole. Recipe variables may be referenced with ${variable}.
string
—
No
where
Filter rows before applying the wrangle using SQL-like criteria, such as column1 = 123 OR column2 = 'abc'.
string
—
No
where_params
Values used with where for parameterized criteria. Uses SQLite placeholder syntax such as ? or :name.
This template extracts JSON content from a URL. Returned fields depend on the page, prompt, and retrieval model.
wrangles: -search.retrieve_link_content: input: - Product URL output: - Page Data api_key: Your Google API key client: google_url_context output_format: json prompt: Extract the product title and manufacturer.