Imagine a business that compares confirmed orders for the last two complete weeks every Monday. Management wants to know whether orders are rising or falling and whether anything needs attention. An analyst prepares the figures and a short commentary. AI can help, but how much it can do depends on the information it has and the work it is allowed to take on.
We will follow this hypothetical example through six levels, adding saved instructions, data access, business context, updates to working rules, and independent execution. At each level, we will look at what changes for the analyst. Then I will show how I combine these capabilities in my own work on HeyRup.
This is my practical map, not a validated scale of ability. The capabilities can overlap, and you do not need to work through every level. Chat may be enough for a one-off piece of writing. A recurring report may justify a more deliberate approach to data sources and changes in business rules.
1. The chat window: you supply the material for each task
The analyst pastes a table into a chat and asks AI to compare the two weeks and draft a commentary. They get an initial response, ask follow-up questions, or clarify the request. Chat is useful for a task with these clear boundaries: it can help examine the material and turn it into a readable explanation.
The table still needs to contain the information required to answer the question. If it only contains two totals, AI can describe the difference. Those figures alone cannot tell it whether the analyst correctly excluded canceled orders. Nor does a decline explain whether demand weakened, a sales channel failed, or something else happened.
The analyst therefore supplies both the figures and an explanation of what they mean. They also check that the commentary distinguishes an observed change from a possible explanation. For the next report, they need to make sure AI has the necessary information again.
2. A workspace: you save the recurring instructions
After a few Mondays, it becomes clear that much of the request stays the same. The analyst saves instructions in a workspace or project: compare complete weeks, show the change as a count and a percentage, and do not explain its cause without supporting evidence. They add an example of a good commentary and the rule that canceled orders must be excluded.
Next time, they supply a new table and ask for the next report. They no longer have to explain the format or basic rules. They still provide current data and check that the saved instructions match what the business needs.
This makes the work easier to repeat. But the analyst still has to transfer the table. If that is an unnecessary manual step, connecting AI to the data source may be a useful next move.
3. Tools and integrations: AI can retrieve the material
The analyst gives AI access to an approved source containing individual orders. Instead of downloading and pasting a table, they can ask it to prepare last week's commentary. AI uses a connected tool to retrieve the data and follows the saved instructions. In this article, I use the term agent for a model that can choose and carry out steps using tools.
Access should match the task. Preparing a report may only require permission to read an approved export. The agent does not need permission to change orders or send messages to management. These boundaries should be set when the tools are connected.
A working connection does not explain what the data means. An agent might be able to read an order-status column without knowing which statuses to include. To apply the business definition, it needs somewhere to find that definition and a clear link between the rule and the data.
4. A context layer: you connect data to its meaning
The saved instruction to “count confirmed orders” was a useful start. It is too brief for working directly with business data, though. We need to specify the source, the order statuses to include, which date assigns an order to a week, and the cutoff for deciding whether it has been canceled.
In our example, the rule might be to assign orders to a week by their confirmation date and exclude any canceled by the Monday reporting cutoff. The current metric definition also needs to identify the specific data fields and statuses used in the calculation. The agent gets a way to find that definition, the corresponding data, and the person responsible for resolving ambiguity. That is the difference from a rule copied into a single prompt: several tasks can use the same definition, and there is an agreed source to follow when instructions conflict.
I call this organized material a context layer. It can span documentation, databases, and other tools. What matters is that the agent can retrieve the relevant context for the task. The report can then identify its data source, reporting period, and definition, giving the analyst a basis for checking how the result was produced.
Documentation alone does not guarantee a correct calculation. Someone still needs to verify that the query or procedure actually implements the definition. I explore how to organize this material in my article on a second brain for AI. Here, the next question is how to keep that context from going out of date.
5. A nervous system: a correction improves the next run
Suppose the analyst checks the underlying data and finds that the calculation included a canceled order. The agent traces the problem to a missing filter in the data query. If it only rewrites the commentary, it could make the same mistake next Monday.
The fix therefore needs to reach the query itself. Within its permissions, the agent fixes the query or prepares a change for approval. It adds a check to catch the same kind of mistake and records what changed and why. The business definition stays the same; the calculation is being brought into line with it.
A change to the business definition is a different situation. The approved decision and its effective date need to be recorded, and the reporting procedure updated. The next report must use the new version and flag any loss of comparability with earlier figures.
“Nervous system” is my metaphor for this connection between work and feedback. An important finding reaches the source that the next task will use. Accumulating notes is not enough: the system needs to distinguish a proposal from an approved decision and establish who may make changes. Agents can handle the recording and follow-up work under agreed rules.
6. Governed autonomy: you delegate a bounded process
Until now, the analyst has started the work each time. An agent could instead retrieve the data, calculate the figures, run checks, and draft the report every Monday on its own. A schedule alone only automates the start. The degree of autonomy also depends on which subsequent actions the agent can choose and complete without a person.
In our example, it gets specific checks: the source must cover both complete weeks, canceled orders must be excluded, and totals must reconcile with the underlying data. If a check fails, the process must stop short of a finished report. The agent passes the problem and the supporting material to the analyst.
If the checks pass, it produces a draft with links to its sources. In this setup, the analyst still reviews the commentary and decides whether to send it to management. In particular, they check the evidence behind any explanation of the change. Correct totals do not establish why orders fell.
The boundaries can differ for each task. What matters is how the result can be verified and what an error would cost. Preparing an internal draft has different consequences from changing prices or messaging a customer. Permissions and human review need to reflect that difference.
What this looks like in my work
With HeyRup, I face a similar question on a larger scale: how can I give a task in one place while drawing on information and tools spread across the project? HeyRup has an app, a public website, and a product HQ in three repositories. The app and website support fifteen languages. The work also depends on meeting decisions, ongoing conversations, and operational data.
One place to give instructions, several places to work
My main tools are ChatGPT and Codex. I give tasks through HQ_RadekDuha, my repository of working instructions and a map of projects and their sources. From one window, I direct work that agents carry out in different places. I do not edit the HeyRup repositories directly; agents make those changes.
The HQ map helps them identify where a task belongs and which rules to read. When they need a database or another tool, they use the command line (CLI), MCP servers, or native plugins. These provide different ways to retrieve data and perform permitted actions. When a task requires interacting with an application's interface, agents can also operate a browser.
Agents retrieve and record the context
Each task also builds on what came before. Notion holds my meeting transcripts and an Activity / Signal Log for important decisions and findings. Agents update that log automatically according to my instructions. During later work, they search these sources and organize relevant information. I do not go into Notion to write those entries myself.
Notion provides a record of what was discussed and decided. If a decision changes a lasting working rule, the agent also updates the relevant project HQ, which guides future tasks. A meeting record can explain why a decision was made, while the current procedure has a designated home.
ChatGPT's long-term memory retains additional context. When that is not enough, agents can follow their instructions to revisit selected conversations and records of earlier runs, known as sessions or traces. They look for relevant details, recurring problems, and lessons for future work. Depending on the task, they also use Slack, email, and analytics tools. Agents do this retrieval, so I do not have to reconstruct the project's history for every new task.
Some work continues without another prompt
I use the same instructions and sources for agent automations. They can start and carry out defined recurring work without a new chat request each time. Each automation has a scope, permissions, and a way to check its results.
I decide what I want to achieve and approve actions that affect the product, customers, or published content. An agent can independently finish an authorized, reversible step. Publication, changes to a live system, and sensitive operations require the appropriate approval. If it does not have that approval for the action, it stops. Once I have made the decision, the agent records it and updates the working rules where needed.
This is how I use AI to develop and operate HeyRup; it is not a customer-facing HeyRup feature. For me, the main benefit is the continuity: a task leads to retrieving context, doing the work, and recording what the next task will need. One window is the entry point. The connected sources, rules, and checks behind it make that possible.
How to find your own next step
Start with one recurring task where you know what a good result looks like. It might be a report, meeting preparation, or sorting emails. Describe what you do, what AI does, and where the work gets stuck. That gives you a place to look for an improvement.
The prompt below helps you examine that process. Its findings will depend on the information you supply; the model cannot independently establish how you work. A description and an example without sensitive data are enough to start. Before sharing documents, check that the tool is allowed to process them.
Help me map how I currently use AI in this specific workflow: [describe the work].
Use only information in this conversation or in materials I explicitly attached. Do not pretend you can access my memory, files, or systems you cannot see. Separate what is supported, what you infer, and what you do not know. If important information is missing, ask me no more than 8 specific questions.
Use the six levels from the article as a working map, not a certified standard:
1. Chat window: one-off task; I provide the context manually.
2. Repeatable workspace: saved instructions, templates, or a separate environment for similar tasks.
3. Tools and integrations: AI works with approved files, apps, or data.
4. Context layer: AI can find current definitions, sources of truth, decisions, and rules.
5. Live nervous system: changes, decisions, and corrections flow back into the right context for future work.
6. Governed autonomy: recurring steps run on their own within clear limits, with checks and human intervention where needed.
Assess these areas separately:
- everyday tasks and how often they repeat
- workflows and tools
- context quality and availability
- how information is updated and decisions are handed off
- output checks and the cost of an error
- autonomy and human oversight
For each area, give concrete evidence, a level or range, and the information you still do not have. Do not give me one overall score or assume a higher level is always better. Finish by recommending one practical change to this workflow. Explain how I can recognise an improvement on the next task, verify the result, and decide what should remain under human control. Choose one change and try it on the next task
If the discussion reveals that you keep rewriting the same request, save the instructions and note what you still need to add next time. If moving tables takes most of the effort, try limited access to the source. If conflicting definitions produce inconsistent results, agree on the rule first and decide where it should be retrieved from.
For the Monday report, progress can be quite concrete: fewer manual steps, figures that match the source, and commentary that does not present a guess as an explanation. Once that process works, consider giving the agent more independence. Also check what it does when data is missing or a check fails.
A finished-looking output can make it tempting to skip verification. In its AI Fluency Index, Anthropic observed less explicit fact-checking and identification of missing context in conversations that produced artifacts. The study could not see checks outside the chat and does not establish what caused the difference. I take it as a reminder to make verification part of the task.
For me, moving towards a “nervous system” means that each task can use what earlier work established. Sometimes saved instructions are enough; elsewhere, connected tools and automations are worth adding. Start where you currently repeat unnecessary work or lose important context.
Further reading: Every: The Eight Levels of AI Adoption, Anthropic: AI Fluency Index, and Simon Willison: Context engineering. These offer different perspectives on working with AI. The six levels in this article are my own synthesis.