Web

Operates like a human

The public web, read directly: a page as text, a sitemap as a list, a URL's text or JSON. Nothing to connect.

What a member can do

A member can read the public web with three built-in tools. A task names the ones it may use:

Tool What it does
web.read Reads one page and returns the words a reader would see, its title, and a hash of exactly those words.
web.sitemap Reads a sitemap and returns its addresses as a list, each with the date the site says it last changed.
web.request Sends a GET to an address and returns the text or JSON that comes back, with its status.

What each takes and returns is in the Web tools reference.

These are plain requests, not a browser. No JavaScript runs, and no cookies or credentials are sent. For a product a member signs in to and clicks through, bind a browser. For a system with an MCP server, connect it from the Tools directory.

A hash that changes when the words do

web.read returns a sha256 of the page's visible text. Scripts, styles and tracking snippets are not part of it, so a site that is redeployed with nothing reworded returns the same hash, and a page with one sentence changed returns a new one. A task can keep the hash to tell whether a page needs reading again, and can pass it to memory.train as the revision of what it learned.

The hash follows the text, not the meaning. A page that prints the time or rotates a banner changes on every read.

Public addresses only

A request reaches public addresses, and nothing else. A private, local or internal address is refused (web_address_refused), including a public name that points at one. Every redirect is checked again before it is followed, and a redirect to such an address is refused without being requested (web_redirect_refused).

Whole, or not at all

A page or a response too long for one result is refused, with its size and its hash and none of its content (web_page_too_large, web_body_too_large). A sitemap too long for one result returns how many addresses it has, complete: false, and no list. A member never receives part of a page as though it were the page.

A page that answers with an error is refused by web.read and web.sitemap (web_status_not_ok), so an error page is not read as content. web.request returns the status for the task to judge.

Planned

Reading in parts, writing, and credentials

A long page or sitemap will be readable in parts, each carrying the hash of the whole. web.request will send POST, PUT, PATCH and DELETE, waiting for your approval unless the task is kept to reading. An enterprise will be able to attach a credential to an address on Tools, which the OS adds to requests for that address and the member never sees.

Who connects it

Nobody. There is nothing to bind and no credential to store: the tools are built in, the same for every enterprise. A task has them when it names them.

What waits for you

Nothing a web.* tool does waits for you unless the task declares a gate on it. Gate web.request and every request waits, with the address it was about to ask for, until you approve it on Gates. See Gates and tool scopes.

Setup

  1. On the task, name each tool it needs: web.read, web.sitemap, web.request.
  2. In the task's skill, say which addresses it should read, and what to do with a refusal: skip a page that is too long, and never treat a refusal as the page.
  3. Check it. Run the task. The run page shows every request, where its redirects ended, and what came back.

Troubleshooting

What you see Why What to do
web_address_refused The address is private, local or internal, or its name points at one Use the public address. A system inside your network is not reachable this way
web_redirect_refused A redirect loops, or leads somewhere the tools do not go Read the address the site redirects to, if it is public
web_url_invalid The address is not http or https, names a port, or carries a username Pass a plain https:// address
web_status_not_ok The page answered with an error status Check the address. web.request shows the status and body
web_not_html web.read was given something that is not a page Use web.sitemap for a sitemap, web.request for JSON or text
web_page_too_large, web_body_too_large The content does not fit in one result Have the task skip it. Nothing was returned, so there is nothing to summarise
A sitemap returns complete: false The list does not fit in one result Do not treat missing addresses as removed pages
web_header_refused The task set a header other than accept or accept-language Remove it. Credentials are never passed by a member
A page reads as nearly empty The site builds its content with JavaScript Bind a browser for that site
The hash changes on every read The page prints something that changes: a time, a counter, a rotating banner Compare the text, or read a steadier address