Skip to content
Volume 1IA da Terminale
  1. The book
  2. Chapter 4

Chapter 4

From generative AI to agentic AI

It did not get smarter: it was given hands. What tools are, what the harness around the model is, and how much rope to give.

The first time you watch it work

It goes more or less like this.

You ask it for something and it does not answer. That is: it does not write you an explanation of how it is done. A line you recognise appears on screen, because it is the one from chapter 2: it is looking at what is in a folder. Then another appears: it is reading a file. Then it writes something, runs it, gets an error, reads the error, changes a line and tries again. It works. It tells you what it did.

Meanwhile you have been sitting there with your hands still on the keyboard.

It is a moment that makes an impression, and the typical reactions are equally wrong. On one side the enchantment: it does everything, I do not need to understand anything any more. On the other the knot in the stomach: what if it deletes something?

The enchantment and the knot in the stomach have the same cause, namely not knowing what is really going on. And what is really going on is simpler than it looks, because no new intelligence is involved. The machine in front of you is the same one as in the previous chapter, with the same limits: it remembers nothing from yesterday, it sees only what is on the table, and it fills the gaps by itself.

One thing only has changed, and it is the one this chapter takes apart:

It did not get smarter. It was given hands.

What you take home from this chapter

  1. how we got here, from a program that completes sentences to one that acts;
  2. what tools are, that is, the levers you put in its hands, and why the work changes in nature when reality is involved;
  3. what the harness around the model is, and why it counts more than the model;
  4. how much rope to give, when, and how to work so that being wrong is cheap.

From this chapter on I take the previous three for granted: paths, basic commands, and the fact that you know how the thing you are talking to is made. If something does not add up, it is two minutes back.


How we got here

Four steps, and to understand the fourth you need the first three. I will keep them short, because we are not studying the history of computing: we are working out why the thing in front of you behaves the way it does.

First: guessing the next word. You know the starting point from the previous chapter. A program trained to predict the most plausible continuation of a text. As useful as your phone’s autocomplete, and at first it did not look like much.

Then: answering instead of continuing. The same engine, additionally trained to do what you ask instead of carrying on your sentence. If you write “explain to me how to change a tyre”, it does not complete the sentence: it answers. It is the moment the thing stops being a toy for insiders and becomes the chat everyone knows.

Then: taking a moment before speaking. Models were taught to reason in silence before answering, trying routes and discarding them. On hard problems the difference is large, and you have already seen it in the previous chapter’s notebook.

This has a name worth recognising, because it is everywhere: chain of thought, almost always abbreviated to CoT. When a program shows you a panel labelled “reasoning” or “thinking” with the model’s attempts inside, that is the chain of thought made visible. It is worth opening every so often: it is where you see whether it understood the question or understood something else.

And finally: acting. Here is the leap. Until that moment, whatever the model produced ended up in one place only: your screen. Then someone did something conceptually banal and practically enormous, namely put levers in front of it. Not “tell me the command I should give”, but “give it, and see what happens”.

The model, inside, is the same. What changed is its environment.

It is worth stopping on this for a second, because it is the heart of the whole book. What happened is not that a machine became capable of working in your place. What happened is that a machine capable of talking was put in a room where the real things are: your files, your terminal, the network. The difference between before and after is not inside it, it is around it.

One clarification, so the sentence below is not read the wrong way. Saying that agentic AI is not a smarter model does not mean the models are standing still: they improve, and considerably. Every version reasons a little better, gets things wrong a little less, holds the thread of longer jobs. They are two different axes, and it is worth keeping them apart precisely because they multiply together: how good it is, and what it is allowed to touch. A brilliant model with no tools remains something that talks and no more. A tool in the hands of a more capable model gets further, and goes wrong in less predictable ways. The leap this chapter is about is the second axis, but the first keeps moving underneath, and it is why in a year’s time these same pages will apply to bigger things.


The levers

Let us call things by their name, which you will then hear everywhere. The levers are called tools.

A tool is an action the model can request: reading a file, writing one, running a command in the terminal, searching for something online, sending a request to a service. Every tool has a name, instructions for use and limits, exactly like the commands you learned in chapter 2.

And here is a detail worth being clear about, because it explains half the behaviour you will see.

The model touches nothing. It goes on doing the only thing it knows how to do, namely produce text. Except that part of that text is now a request: I want to use the “read file” tool on the path Documents/Invoices. It is the program around it that carries out the request, takes the result and puts it back on the table. Then it continues.

Hence a consequence that will save you a lot of irritation: the lever always works, it is the aim that can be off. When the AI makes an immaculate mess in a folder that was not the right one, nothing broke. It correctly asked for a drawer to be opened, and the drawer was the wrong drawer. It is the type 1 error from the previous chapter, the slip, and it is cured by looking at what you gave it before looking at what it handed back.

There is a second category of error, and it does not come from a bad aim. Ever since it can go and read things outside — a web page, a document someone sent you, the text of an email — that material lands on the table alongside your instructions. And it reads everything on the table. If somebody in there has written a sentence built to look like an order, it finds it next to yours and may follow it.

This is not science fiction: it is called indirect prompt injection and it is an open problem. We will come back to it in the last chapter, when we have it handle material from outside your computer. For now keep one thing in mind: the material you give it to read is not passive.

The loop

With levers in hand, the work takes on a shape with a recognisable rhythm:

it looks, decides, acts, reads how it went, corrects. And round again, until it has finished or until it gives up.

This ring is called a loop, and when you see it written agentic loop this is exactly what is being talked about: not a different technology, but the fact that the circle closes by itself.

The loop, drawn as a ring: the model at the centre, the four phases around it (looks, decides, acts, reads the result) and below, at the bottom, the real environment (a folder, a file, a terminal window) from which the return arrow starts. It has to read as a cycle that closes, not as a sequence that ends.
Figure 4.1

That fourth step, “reads how it went”, is the real difference from a chat, and it is far more important than it looks. In a normal conversation the answer ends up on your screen and you pass judgement on it, with your eyes. Here the answer ends up in the world, and the world answers back: the file is there or it is not, the command worked or gave an error, the page opens or stays blank. The feedback is not an opinion, it is a fact, and it comes back on its own.

Out of this comes the most useful criterion you take away from this chapter, and it applies every time you wonder whether something is worth delegating:

The good jobs to put in an AI’s hands are the ones where reality says whether it went well. Renaming four thousand photos is a perfect job: you open the folder and you see. Tidying an archive, checking that a site responds, extracting the amounts from a hundred PDFs and adding them up: the same. You can check in ten seconds. “Write me a text that moves the reader” is the opposite: no test, no visible error, only a judgement. It is not that it cannot be done, but the checking stays entirely on your shoulders, and has to be budgeted for beforehand.


The harness

Now the question almost nobody asks, and which in fact decides everything.

If the model is the same, why do two programs using the same AI behave so differently? Why does one ask permission at every step and the other charge ahead, one remember your preferences and the other always start from scratch, one know how to open your files and the other not?

Because around the model there is a program, and that program is not a technical detail: it is half the behaviour you see. It is called the harness — the horse’s tack — or scaffolding.

The harness decides five things:

  • which levers exist. If it has not been given the tool for sending email, it can want to as much as it likes: it does not send email;
  • what ends up on its table. Which files, which parts of the conversation, which permanent instructions. It is the context from the previous chapter, and somebody decides what goes on it;
  • when to stop and ask. Every action, only the risky ones, or never;
  • what survives the end of the session. A file of notes, some preferences, nothing;
  • where it can put its hands. One folder, the whole disk, the network too.

From this comes the most useful practical consequence of the chapter, and it goes against everything you will read elsewhere:

When you choose an agentic tool, you are choosing a harness more than a model. The frontier models, as you have seen, are all good. The harnesses are not: they are made by different people, with different ideas about how much to trust, and the difference between a productive afternoon and a wasted one lies there far more often than in the model league table.

The part you build yourself

And now the thing that makes this chapter different from the others.

A portion of that harness is not made by the supplier. You make it, and you have been making it for three chapters without my telling you.

Folders with a name that says what they contain are harness. Files called report-2026-march.pdf instead of Report March (final) FINAL.pdf are harness. The dedicated working folder, kept apart from the rest of your stuff, is harness. They are all things that reduce the number of questions somebody has to ask themselves to work in there, and that somebody is no longer only you.

There is a way of taking this to its conclusion, and it is simple to the point of seeming silly: a sheet of instructions inside the folder. A text file, written in your own language, containing what you would tell a new collaborator on their first day. What is in here. How things are to be named. Where results go. What is never to be touched.

Programmers have been doing exactly this for a couple of years, with English names and conventions of their own. The version for ordinary people is that sheet, and in the next chapter we actually write one.


How much rope

There remains the question you asked yourself watching it work the first time: how far do I let it go?

The answer is not a yes or a no, and it is the chapter’s second change of perspective. Autonomy is a dial, not a switch. The positions, from tightest to widest, are roughly these:

  • just propose: tell me what you would do, I will do it;
  • one step at a time, and ask me before each one;
  • do the whole job, but stop before the things that cannot be undone;
  • do it and report at the end.

And there is something autonomy raises along with the risk: the bill. A loop that closes by itself can run for a long time, and every turn is text passing through. If you pay as you go, the high setting is also the expensive setting.

There is no position that is right in the abstract. There is the one that is right for that job, and it is chosen with four questions that quickly become automatic:

  1. If it goes wrong, can it be undone? Renaming files inside a copy: yes, easily. Sending an email to a client: no, and there is no fixing it.
  2. Would I notice? Not “is it capable of getting it wrong”, but “if it did, would I notice?”. If the answer is no, the dial stays low whatever the tool promises.
  3. What does the worst error cost? Ten minutes lost and a client lost are two different categories, and deserve two different settings.
  4. Does the material it works on come from me or from outside? A folder you prepared yourself and a web page anyone could have written do not deserve the same degree of trust.

And one rule worth more than the four questions put together: autonomy goes up per type of job, not in general, and it goes up over time. After you have had it tidy the receipts four times while watching, on the fifth you let it go on its own with that. Which does not mean you will let it go on its own with contracts.

The perimeter

The other half is geographical, and it is the easiest to forget: where it works.

A dedicated working folder is the simplest and most effective fence there is. Not because AI is dangerous, but for the reason already given in chapter 1: it is extremely fast, obedient, and has not the faintest idea which of your folders is the one you care about. Inside a fence, an error of aim becomes a thirty-second inconvenience.

The name for this fence is one you will meet often: sandbox, the sandpit where children play without anyone getting hurt. If a program tells you it is working in a sandbox, it is telling you where the perimeter ends.

The chapter 2 commands that deserve attention apply here without re-explaining: rm, which deletes and does not go through the bin, and sudo, which suspends the protections. When you see them appear in a line about to be run, that is the moment to read before saying yes. Not out of distrust: because a line like that is the only category of error, in this whole book, that cannot be undone.

And the three habits from the end of the previous chapter stay exactly as they were. A copy, one step at a time, the plan before the execution. I will not repeat them: if they do not come to mind, they are at the end of the previous chapter.


The pocket notebook

  • agent: an AI that can act as well as answer
  • tool: a single action it is allowed, such as reading a file or running a command
  • harness (or scaffolding): the program around the model, which decides levers, context, permissions, memory and when to stop and ask
  • the loop (agentic loop): looks, decides, acts, reads how it went, corrects
  • chain of thought: the intermediate steps by which the model reaches an answer or decides the next action
  • perimeter (sandbox): where it is allowed to put its hands
  • the autonomy dial: how much rope you give it on a given job
  • the sheet of instructions: what you write inside the folder, which becomes part of the harness

Where we go now

So far we have assembled everything without building anything. You know how the place is made, you know who you are talking to, you know what giving it hands means and you know how much rope to leave it.

In the next chapter we work. We take one of the chores you wrote down at the end of chapter 1, the one with the easiest check, and we solve it together from start to finish: the folder, the sheet of instructions, the request made properly, the first attempt that does not go the way it should, and the thing that in the end works and from tomorrow saves you a piece of your week.

It will not be the next Facebook. It will be something small, yours, that nobody would ever have put on sale because it is useful to you and four other people in the world. Those are the things that change your days.