Web
Operates like a human
The public web, read directly: a page as text, a sitemap as a list, a URL's text or JSON. Nothing to connect.
What a member can do
A member can read the public web with three built-in tools. A task names the ones it may use:
| Tool | What it does |
|---|---|
web.read |
Reads one page and returns the words a reader would see, its title, and a hash of exactly those words. |
web.sitemap |
Reads a sitemap and returns its addresses as a list, each with the date the site says it last changed. |
web.request |
Sends a GET to an address and returns the text or JSON that comes back, with its status. |
What each takes and returns is in the Web tools reference.
These are plain requests, not a browser. No JavaScript runs, and no cookies or credentials are sent. For a product a member signs in to and clicks through, bind a browser. For a system with an MCP server, connect it from the Tools directory.
A hash that changes when the words do
web.read returns a sha256 of the page's visible text. Scripts, styles and tracking snippets are not part of it, so
a site that is redeployed with nothing reworded returns the same hash, and a page with one sentence changed returns a
new one. A task can keep the hash to tell whether a page needs reading again, and can pass it to memory.train as the
revision of what it learned.
The hash follows the text, not the meaning. A page that prints the time or rotates a banner changes on every read.
Public addresses only
A request reaches public addresses, and nothing else. A private, local or internal address is refused
(web_address_refused), including a public name that points at one. Every redirect is checked again before it is
followed, and a redirect to such an address is refused without being requested (web_redirect_refused).
Whole, or not at all
A page or a response too long for one result is refused, with its size and its hash and none of its content
(web_page_too_large, web_body_too_large). A sitemap too long for one result returns how many addresses it has,
complete: false, and no list. A member never receives part of a page as though it were the page.
A page that answers with an error is refused by web.read and web.sitemap (web_status_not_ok), so an error page
is not read as content. web.request returns the status for the task to judge.
Planned
Reading in parts, writing, and credentials
A long page or sitemap will be readable in parts, each carrying the hash of the whole. web.request will send
POST, PUT, PATCH and DELETE, waiting for your approval unless the task is kept to reading. An enterprise will
be able to attach a credential to an address on Tools, which the OS adds to requests for that address and the
member never sees.
Who connects it
Nobody. There is nothing to bind and no credential to store: the tools are built in, the same for every enterprise. A task has them when it names them.
What waits for you
Nothing a web.* tool does waits for you unless the task declares a gate on it. Gate web.request and every request
waits, with the address it was about to ask for, until you approve it on Gates. See
Gates and tool scopes.
Setup
- On the task, name each tool it needs:
web.read,web.sitemap,web.request. - In the task's skill, say which addresses it should read, and what to do with a refusal: skip a page that is too long, and never treat a refusal as the page.
- Check it. Run the task. The run page shows every request, where its redirects ended, and what came back.
Troubleshooting
| What you see | Why | What to do |
|---|---|---|
web_address_refused |
The address is private, local or internal, or its name points at one | Use the public address. A system inside your network is not reachable this way |
web_redirect_refused |
A redirect loops, or leads somewhere the tools do not go | Read the address the site redirects to, if it is public |
web_url_invalid |
The address is not http or https, names a port, or carries a username |
Pass a plain https:// address |
web_status_not_ok |
The page answered with an error status | Check the address. web.request shows the status and body |
web_not_html |
web.read was given something that is not a page |
Use web.sitemap for a sitemap, web.request for JSON or text |
web_page_too_large, web_body_too_large |
The content does not fit in one result | Have the task skip it. Nothing was returned, so there is nothing to summarise |
A sitemap returns complete: false |
The list does not fit in one result | Do not treat missing addresses as removed pages |
web_header_refused |
The task set a header other than accept or accept-language |
Remove it. Credentials are never passed by a member |
| A page reads as nearly empty | The site builds its content with JavaScript | Bind a browser for that site |
| The hash changes on every read | The page prints something that changes: a time, a counter, a rotating banner | Compare the text, or read a steadier address |