Self-improvement
After a run that went wrong, the OS proposes one change to the task. You decide, and an approved change is a new version you can see.
What it does
Self-improvement fixes how a task runs. After a run that gave it something to learn from, the OS reviews the run against the task's definition and may propose one change to the task: its skill, its tools, its gates or its spend. Nothing changes until you approve it.
It does not ask whether a team should exist, or chase a business number. See Retrospectives.
Which runs get a pass
A task is reviewed only if its version hands every ending to the OS's review task, with a successor on terminal:
{ "slug": "task-self-improvement", "on": "terminal" }
The tasks in Catalog carry it, so a copy of one does too. A task you create does not until you add it. A pass is a run of that review task, on your enterprise's model connection, and its cost counts like any run's.
Even then, not every run gets a pass. A run gets one when:
- it failed, or it succeeded with at least one failed tool call on its record. A clean success is left alone;
- the failure was the task's own. A run stopped because it went quiet (
run_stalled) is not, and a cancelled or blocked run gets no pass; - the task is not itself part of the loop: a task that decides recommendations, or a run carrying one out, is never reviewed;
- the task has had no pass in the last six hours, and has no recommendation waiting;
- your enterprise owns the task. A clone is not reviewed, and neither are the OS's own tasks.
When a run is passed over because it was a clean success, went quiet, belongs to the loop, or came within six hours
of the last pass, its events say so (improvement.skipped, with the reason).
A pass keeps at most one recommendation: the one it is surest will help, and never a cosmetic one. A pass that finds nothing worth changing proposes nothing, and that is a good result.
Deciding
Operations → Self-Improvement lists the recommendations waiting for you, each with its task, what it proposes, how sure the review is, and when it was raised. Self-Improvement in the sidebar counts them. They are not blockers: nothing waits on them, so Blockers does not list them. Tasks shows each task's count, and a task's page lists its own.
| Decision | What happens |
|---|---|
| Decline | Needs a reason. The recommendation is kept, declined, and the review sees it next time, so the same change is not proposed again without new evidence. |
| Approve | The OS carries it out: it writes the task's next version with the change. |
Only the enterprise that owns a task can approve a change to it.
A task given the OS's recommendation tools can decide recommendations too, and can also mark one actioned: someone else takes the work, tracked by the issue or pull request it links. Every decision records who made it, a person or a run, and Decided in the last 60 days lists them with their reasons.
Statuses
| Status | Means |
|---|---|
pending |
Waiting for a decision. While one waits, its task gets no new pass. |
declined |
Turned down, with the reason. |
actioned |
Handed to work tracked elsewhere, with its link. |
implementing |
Approved; the change is being written. |
complete |
The new version is published. |
error |
The change could not be made: the implementing run failed, was blocked or cancelled, or the change broke a rule below. You can approve or decline it again from its page. |
What the change may never do
An approved change is carried out by one of the OS's own tasks, on your enterprise's model connection. It may change only the task's skill, tools, gates and spend, and it is refused if it would:
| Refused | Code |
|---|---|
| Give the task a tool it did not have | tools_widened |
| Remove a gate | gates_removed |
| Raise the task's spend | spend_increased |
| Raise its model tier | tier_increased |
| Write a repository command into the skill | repo_command_in_skill |
| Write your own data (an id, an email address) into the skill | tenant_data_in_skill |
| Leave a skill that cannot be read | unparseable_workflow |
The goal, the success metric, the guardrails, the inputs, the workspace, the successors and the schedule are never touched. A change that needs any of those, or needs more access, is yours to make by editing the task.
The new version, and going back
An approved change is published as the task's next version, like any edit. Runs already under way keep the version they started on; the next run uses the new one. Any other recommendation still waiting on the task is declined as superseded, since it was written against the old version.
To undo a change, open the previous version in the task's Version history and save its content as a new version. There is no one-click rollback.
Over the API
With an API token scoped recommendations:read (and recommendations:write to decide):
| Call | What it does |
|---|---|
GET /v1/recommendations |
The recommendations waiting for a decision, or with view=decided, those decided in the last 60 days. |
GET /v1/recommendations/{id} |
One recommendation, in full. |
POST /v1/recommendations/{id}/decision |
Approve or decline one. A decline needs a reason. Send an idempotencyKey and a retried call returns the first decision rather than deciding twice. |