Web tools
The web.* tools: read a public page as text, list a sitemap, and GET a public URL, with a hash of what came back.
Generated from the built-in tools registry. Do not edit by hand.
web.* tools make a plain HTTP request to a public address and return what came back. Nothing is bound and nothing is set up: a task declares the tools it calls, like any other tool. They are the same for every enterprise and know no site.
They are not a browser. No JavaScript runs, and no cookies or credentials are sent. A product a member signs in to and clicks through is a Computer-Use binding, and a system with an MCP server is a vendor tool the enterprise binds.
A request reaches public addresses only. Private, local and internal addresses are refused, including a public name that resolves to one, and every redirect is checked again before it is followed.
Each result carries a sha256 the tool computed. web.read hashes the page’s visible text, so the hash changes when the words do and not when a deploy renames a script. Content too long for one result is refused or marked incomplete, never cut: a partial result is not the page.
| Tool | What it does | Effect |
|---|---|---|
web.read |
Read one public web page as visible text, with a hash that changes only when the text does. | read |
web.sitemap |
Read an XML sitemap and return its URLs as a list. | read |
web.request |
Send one GET to a public URL and return the text or JSON body, with its status and hash. | read |
web.read
Read one public web page as visible text, with a hash that changes only when the text does.
What the model reads:
Read one public web page as text. Pass url (http or https). Returns the page's visible text, with scripts, styles and markup removed and whitespace collapsed to single spaces, its title, and sha256: the hash of exactly that text. The same page gives the same sha256 until its words change, so use it to tell whether a page changed, and as a memory.train source revision. Only HTML is read here: use web.sitemap for a sitemap and web.request for JSON or plain text. A page whose text is too long for one result is refused as web_page_too_large, with its chars and sha256 and no text: that is not the page, so do not summarise it or train from it. A status outside 2xx is refused as web_status_not_ok. Follows up to 5 redirects. Private, local and internal addresses are refused. No JavaScript runs, and no cookies or credentials are sent.
| Field | Type | Required | Description |
|---|---|---|---|
url |
string | yes | The page to read: an http or https URL on a public host. |
No other fields are accepted.
- Effect: read.
- Retry: Safe: repeating the call changes nothing. A repeat sends the request again; a GET changes nothing.
- Called by: a run.
- Returns:
{ kind: "page", url, finalUrl, status, contentType, title, text, chars, sha256 }.finalUrlis where the redirects ended.titleis null when the page has none, and is not part oftext.charsis the length oftext, andsha256is 64 hex characters over exactlytext.
| Error | When | What happens |
|---|---|---|
web_url_invalid |
url is missing, not a URL, longer than 2,048 characters, not http or https, carries a username or password, or names a port. |
Tool error |
web_address_refused |
The host is, or resolves to, a private, loopback, link-local, shared, unique-local, reserved or metadata address, or is a local or internal name. | Tool error |
web_redirect_refused |
There are more than 5 redirects, a redirect loops or names no location, or one leads to a URL or address this tool refuses. The refused address is never requested. | Tool error |
web_fetch_failed |
The name did not resolve, the connection failed, or the response used an encoding the tool does not read. | Tool error |
web_timeout |
No complete response arrived within 20 seconds. | Tool error |
web_response_too_large |
The response passed 2 MB once decompressed. Nothing past that was read. | Tool error |
web_status_not_ok |
The response status is outside 2xx, so the body is an error page and not the content. | Tool error |
web_not_html |
The response is not an HTML page. The error names its content type. | Tool error |
web_page_too_large |
The page’s text does not fit in one result. The error carries chars and sha256, and no text. |
Tool error |
Example:
{
"url": "https://example.com/pricing"
}
web.sitemap
Read an XML sitemap and return its URLs as a list.
What the model reads:
Read an XML sitemap and return its entries. Pass url. A urlset returns kind sitemap with total and urls: each loc, and its lastmod when the sitemap gives one. A sitemap index returns kind sitemap_index with sitemaps: call web.sitemap again on each loc, because this tool does not follow them. complete is true when every entry is in the list. When the list is too long for one result, complete is false, total says how many entries there are, and no entries are returned: that is not an empty sitemap. Anything that is not a sitemap is refused as web_not_sitemap, and a status outside 2xx as web_status_not_ok. Private, local and internal addresses are refused.
| Field | Type | Required | Description |
|---|---|---|---|
url |
string | yes | The sitemap’s http or https URL. |
No other fields are accepted.
- Effect: read.
- Retry: Safe: repeating the call changes nothing. A repeat sends the request again; a GET changes nothing.
- Called by: a run.
- Returns:
{ kind: "sitemap", url, finalUrl, complete, total, urls: [{ loc, lastmod? }] }, or{ kind: "sitemap_index", url, finalUrl, complete, total, sitemaps: [{ loc, lastmod? }] }. Entries are in document order, exactly as the sitemap gives them.lastmodis left out when the sitemap gives none. Whencompleteis false there is no list.
| Error | When | What happens |
|---|---|---|
web_url_invalid |
url is missing, not a URL, longer than 2,048 characters, not http or https, carries a username or password, or names a port. |
Tool error |
web_address_refused |
The host is, or resolves to, a private, loopback, link-local, shared, unique-local, reserved or metadata address, or is a local or internal name. | Tool error |
web_redirect_refused |
There are more than 5 redirects, a redirect loops or names no location, or one leads to a URL or address this tool refuses. The refused address is never requested. | Tool error |
web_fetch_failed |
The name did not resolve, the connection failed, or the response used an encoding the tool does not read. | Tool error |
web_timeout |
No complete response arrived within 20 seconds. | Tool error |
web_response_too_large |
The response passed 2 MB once decompressed. Nothing past that was read. | Tool error |
web_status_not_ok |
The response status is outside 2xx, so the body is an error page and not the content. | Tool error |
web_not_sitemap |
The document is neither a urlset nor a sitemap index. | Tool error |
Example:
{
"url": "https://example.com/sitemap.xml"
}
web.request
Send one GET to a public URL and return the text or JSON body, with its status and hash.
What the model reads:
Send one HTTP request to a public URL and return the response body as text. method is GET, the default and the only method today. Pass url, and headers only to set accept or accept-language. Never pass credentials: authorization, cookie and every other header are refused. Returns status, contentType, body, bytes and sha256 of the body bytes. A status outside 2xx is returned as a result, not an error: check status before you use body. Only text and JSON responses are returned: anything else is refused as web_not_text, and a body too long for one result as web_body_too_large, with its bytes and sha256 and no body. Follows up to 5 redirects. Private, local and internal addresses are refused. To read a page as text use web.read; for a sitemap use web.sitemap.
| Field | Type | Required | Description |
|---|---|---|---|
url |
string | yes | An http or https URL on a public host. |
method |
string, one of GET |
no | Defaults to GET, the only method today. |
headers |
object | no | Only these two. Credentials are never passed here. |
headers.accept |
string | no | The content type to ask for, such as application/json. |
headers.accept-language |
string | no | The language to ask for, such as en-GB. |
No other fields are accepted.
- Effect: read.
- Retry: Safe: repeating the call changes nothing. A repeat sends the request again; a GET changes nothing.
- Called by: a run.
- Returns:
{ url, finalUrl, status, contentType, body, bytes, sha256 }.bodyis the response as text.bytesis its size andsha256is 64 hex characters over the body bytes, both after decompression.
| Error | When | What happens |
|---|---|---|
web_url_invalid |
url is missing, not a URL, longer than 2,048 characters, not http or https, carries a username or password, or names a port. |
Tool error |
web_address_refused |
The host is, or resolves to, a private, loopback, link-local, shared, unique-local, reserved or metadata address, or is a local or internal name. | Tool error |
web_redirect_refused |
There are more than 5 redirects, a redirect loops or names no location, or one leads to a URL or address this tool refuses. The refused address is never requested. | Tool error |
web_fetch_failed |
The name did not resolve, the connection failed, or the response used an encoding the tool does not read. | Tool error |
web_timeout |
No complete response arrived within 20 seconds. | Tool error |
web_response_too_large |
The response passed 2 MB once decompressed. Nothing past that was read. | Tool error |
web_method_refused |
method is anything other than GET. |
Tool error |
web_header_refused |
headers sets anything other than accept or accept-language, or a value that is not a single line of at most 1,000 characters. |
Tool error |
web_not_text |
The response is not text or JSON, or names no content type. The error names the type. | Tool error |
web_body_too_large |
The body does not fit in one result. The error carries bytes and sha256, and no body. |
Tool error |
Example:
{
"url": "https://example.com/api/status.json"
}
Example, asking for JSON:
{
"url": "https://example.com/feed",
"method": "GET",
"headers": {
"accept": "application/json"
}
}