Objectives: A Lightweight Way to Keep Agents on Track for Long-Running Tasks
A prompt and a few files that keep a coding agent on track for long-running tasks, without building orchestration.
One of the biggest questions in AI right now is long-running agents, i.e., the ability to give an agent a lofty task and have it run without input for hours, days, or weeks until it accomplishes what you asked it to do.
The hard part is keeping it on track. It has to remember what it already proved across compactions and restarts, and it has to know when it's allowed to stop.
I ran into this in May 2026 on a public coding challenge:
I entered for the money and to see how far I could push the models at the time (GPT-5.5).
Discovering Goal Mode
My first attempt was a multi-agent research harness.
Each round went like this:
Goal mode within Codex had just come out, and I wanted to test how good a simple goal could be against my awesome custom harness.
I was extremely humbled, and I decided to see how far I could push goal mode.
Constraints of Vanilla Goal Mode
Goal mode got me most of the way, but pushing further raised a bigger question:
Four things kept coming up.
Explicit Stopping Criteria
A vague target lets the agent stop at the first easy win.
What worked was spelling out every metric on every dataset.
random10: mean >= 99.40%, p10 >= 99.25%, worst >= 99.05%
random20: mean >= 99.35%, p10 >= 99.15%, worst >= 99.00%
random40: mean >= 99.25%, median >= 99.20%, worst >= 98.95%
random80: mean >= 99.15%, median >= 99.10%, worst >= 98.90%How We Want It to Work
The biggest unlock with agents is their superhuman ability to write complex low-level code and then run a huge number of experiments.
It worked, and I wanted that process written down:
Every Experiment It Has Run
One evening of goal mode left 38 experiment scripts at the top of my repo.
Where It Is Right Now
Long runs compact and restart. Each time, the agent needs to know:
Without that, it repeats work and drifts.
The catch is getting all of this to the agent through goal mode.
Enter Objectives
An objective is a way to get all of that context to the agent, simply and efficiently, using just natural language and file system structure.
The goal prompt is the key piece.
Every objective's goal has the same shape:
<goal>
- What this objective must achieve.
</goal>
<context_refresh>
- Reread objectives/<slug>/goal.md.
- Reread objectives/<slug>/current_state.md.
- Reread the relevant objectives/<slug>/context/*.md files.
</context_refresh>
<working_strategy>
- The approach and the order of work.
</working_strategy>
<success_metrics>
- Observable signs of progress.
</success_metrics>
<non_goals>
- What this objective must not expand into.
</non_goals>
<completion_criteria>
- What must be true before the objective is done.
</completion_criteria>The context_refresh block does the heavy lifting. Each constraint gets a home:
goal.md.context/.current_state.md.The loop has no controller of its own. Every time the agent starts or compacts, it rereads the objective, and the stopping criteria decide whether the next step is another experiment or a stop.
Pros and Cons
Objectives aren't a replacement for a full orchestration system. They're the fast way to get one long task running.
If you need a general task to run for a long time with a good chance of working correctly, objectives are a good fit. If you need it cheap and efficient at scale, build the orchestration.
Try It
The objectives docs cover setup, the file format, and how to resume a run.
The objectives repository has the skill and its source.
