Imagine asking an AI assistant to write a function that calculates the mean of a list of numbers, with one requirement: an empty list must return None. It gives you this code:

def mean(xs):
    return sum(xs) / len(xs)

For [2, 4, 6], the result is 4. But an empty list, [], causes a division by zero. You point out that the assistant missed the requirement for empty input. It adds a check and fixes the bug.

That change solves the immediate problem. Then you close the conversation and ask it to write another function that works with lists. Will it check for empty input on its own?

The answer depends on what the fix leaves behind. We will use this bug to compare three assistants: one that revises only its current answer, one that saves a lesson, and one that modifies the program it uses to generate patches. In all three setups, the base model's weights stay fixed. These are independent teaching examples, not experiments reported in a paper or stages that a system must progress through.

What remains after the conversation closes?

The first assistant revises the function only within the conversation. It gives you the corrected code. It neither updates its model weights nor writes a record that it can read on the next task.

If the next call uses the same model with an empty context, the assistant cannot read this fix. It may still write correct code using its existing capabilities, but this conversation has added no reusable information.

This gives us a testable distinction between "fixed the code" and "learned something from the fix." The first asks whether this piece of code meets the requirements. The second also asks whether the next task can make use of this experience. Comparing the first and final drafts answers only the first question.

Leave it a lesson it can use

After fixing the bug, the second assistant also saves a record:

Before calculating statistics for a list, check whether the input is empty. What to return for empty input depends on the current task's requirements.

Now ask it to write a function that calculates variance. If the system retrieves this record at the start of a relevant task, the assistant has an extra cue to check the input. The model weights can remain completely unchanged; external memory carries the experience into the next call.

The record deliberately avoids saying "always return None for an empty list." That was the requirement for the mean function. Another interface might require an exception, in which case copying the old answer would also be wrong. The reusable part is the practice of checking; the response still depends on the new task.

Saving the record is only the first step. Will the system retrieve it? Once the assistant reads it, will it follow the advice? Will it apply the advice even when it does not fit? All of these affect later performance.

We could design a simple comparison: take two copies of the same base assistant, let one read this memory and keep it from the other, then give both a set of new tasks under the same budget. Alongside tasks that call for checking empty input, include tasks where this lesson is of no use, and observe whether it introduces extra errors.

This is a proposed experiment to test the effect of memory. We have not run it. It does make clear what to compare: how does retaining this lesson affect the next set of tasks? Clearing both copies' memories before evaluation would prevent us from measuring the effect of this lesson.

Change the program that generates patches

The third assistant can modify more of its system. It has a program responsible for repairing code, which we will call the "improver." The improver decides what information to give the model, which change to try first, and how to use a limited number of attempts.

Imagine that the original improver, A, is simple: when it sees an error, it asks the model to propose a patch. This time, A also treats its own source code as something to modify and produces a candidate improver, B. Their rules might look like this:

Improver How it starts a repair
A Give the current error to the model and ask it to propose a patch
B Read the input requirements, construct boundary cases, then use the check results to ask the model for a patch

For the mean function, B would try to include empty input in its checks, without waiting for the user to report the error. This is a candidate rule; we have not established that it works better.

There must also be an actual handoff. If B passes the acceptance checks, B generates the changes in the next round, and B's own source code remains open to modification. If B is merely saved to a file while A continues doing the work, the process that produces later changes has not changed in the way this setup describes.

This identifies a key feature of recursive self-improvement: the program responsible for producing changes is itself among the things the system can modify. Generalized Agent Iteration (GAI) uses this distinction to separate fixed improvement processes from recursive self-improvement. Our A and B are teaching examples to help explain this distinction, not experiments reported by GAI.

Permission to modify, an actual handoff, and the effects of a modification are three separate things. B's extra boundary checks might find problems that A misses. They might also take too much time, leaving fewer opportunities to propose patches.

To test whether B is better at improving programs, we could freeze versions A and B and prepare a set of new programming tasks that were not used to develop B. For each task, A and B would start from the same initial code and produce patches under the same total budget. We would then evaluate each final program using tests set aside in advance and kept hidden from both improvers. This comparison is also a proposal made in this article.

What has the system gained when the next task begins?

Returning to the same empty-list bug, we can now name what each assistant leaves behind:

Assistant What new information or program the next call can use What still needs to be tested
Revises only the current answer No new state under the setup described above Whether the revised code meets the current task's requirements
Saves a lesson A retrievable record about checking input Whether reading the record improves performance on new tasks
Modifies the improver The improvement program B, actually put into use Whether B produces more useful changes with comparable resources

These three cases help us explain where an "improvement" occurs. They do not automatically rank the assistants by capability: a relevant memory can be very useful, and a program that can modify itself can also get worse with each change.

The next time an AI fixes its own mistake, look at the moment it starts a new task. What new state does it read? Which version of the improvement program does it run? Recording those two things makes "what did it learn?" a question we can keep testing.


Further reading: the GAI paper, particularly its distinction based on whether the improver can itself be modified. This article uses that distinction only to describe system structure, not to claim a performance gain. The examples and comparison designs are teaching illustrations.

The research basis for this article comes from the survey manuscript, whose literature cutoff is September 17, 2026. This article reports no new experiments.